Datacenter GPUNVIDIA

NVIDIA DGX Spark (Project DIGITS, 128GB)

Curated Aggregate·2026-03-158 workloads · 9 records
MyAI RatingNot scoredInsufficient comparison evidence

VRAM

128 GB

TDP

240 W

MSRP

$3.0k

Perf/W

0.13 tok/s/W

Cost/1K tok

$0.99/M

Tested

2026-03-15

Quick answer

How many tokens per second does NVIDIA DGX Spark (Project DIGITS, 128GB) produce on Llama 3 70B Q4?

NVIDIA DGX Spark (Project DIGITS, 128GB) has a source-attributed result of 32.0 tok/s on Llama 3 70B Q4 (batch 1, 8192-token context, Q4_K_M; runtime not documented). This is a reference report, not an independently verified lab result. A 70B Q4_K_M artifact needs roughly 40GB for weights alone; smaller memory configurations require explicit offload or model splitting and do not establish full-GPU residency.

Source: MyAIHardware benchmark database (bench-dgx-spark-l3-70b-q4)As of 2026-03-15

Overview

AI Developer Desktop

NVIDIA DGX Spark (Project DIGITS) — compact AI developer desktop with 128GB unified LPDDR5X and GB10 Grace-Blackwell Superchip. 240W TDP. Turnkey local AI workstation.

AI Usefulness

128GB unified memory is shared by the system and GPU. Standard 405B Q4 weights exceed 200GB, so this is not a single-system full-residency configuration. Select an artifact that fits after OS, runtime and KV-cache allocations.

Verdict

NVIDIA DGX Spark (Project DIGITS, 128GB) with 128GB VRAM at 240W TDP, scored across 8 workloads with 9 benchmark records.

Reference workload

Qwen 2.5 14B

82 tok/s

Quantization

Q4_K_M

16K context · batch 1

LLM Inference Performance

04080120160Llama 370B Q4Llama 370B Q8Llama 3 8BFP16Qwen 2.514BDeepSeek-R17BMistral 7B

Benchmarks (8 workloads)

WorkloadScoreQuantContextσStatusTested
Llama 3 70B Q4

llm

32.0tok/sQ4_K_M8K, Curated Aggregate2026-03-15
Llama 3 70B Q8

llm

22.0tok/sQ8_08K, Curated Aggregate2026-03-20
Llama 3 8B FP16

llm

92.0tok/sFP168K, Curated Aggregate2026-03-25
Qwen 2.5 14B

llm

82.0tok/sQ4_K_M16K, Curated Aggregate2026-05-25
DeepSeek-R1 7B

llm

145.0tok/sQ4_K_M8K, Curated Aggregate2026-04-05
Mistral 7B

llm

145.0tok/sQ4_K_M4K, Curated Aggregate2025-04-08
Gemma 2 9B

llm

118.0tok/sQ4_K_M8K, Curated Aggregate2025-04-15
Asure 12B

llm

100.0tok/sQ4_K_M8K, Curated Aggregate2025-05-01

MyAI Score: not scored

There is insufficient comparable evidence to score NVIDIA DGX Spark (Project DIGITS, 128GB). A score needs results for multiple devices with matching workload, runtime version, quantization, context and batch size.

Workload Fit

Planning estimates for 128GB: weights plus at least 2GB or 10% overhead. Actual KV cache depends on model, context and cache format; confirm the artifact before buying.

Q4
Q8
FP16
7-8B
Excellent
Excellent
Excellent
13-14B
Excellent
Excellent
Excellent
32B
Excellent
Excellent
Excellent
70B
Excellent
Excellent
Won't fit
Reported batch-one examples
Qwen 2.5 14B82 tok/s
DeepSeek-R1 7B145 tok/s
Llama 3 8B FP1692 tok/s

Source

Re-bench at 16K ctx — Grace bandwidth wins.

View sourceHow we benchmark →

Public Trust Layer

Trust score

6/10

MyAI rating

Not scored

Runs

Not documented

Freshness

Aging

Source-linked row with explicit verification status.

Record date: 2026-05-25; 107 days old.

Open primary source

Enterprise buying note

Where to buy

Reseller compare

Retailer we'd check first

Amazon search plus reseller quotes

For datacenter and accelerator parts, Amazon is useful for spotting live listings, accessories, or used pulls, but serious procurement usually happens through integrators, brokers, or cloud partners.

Enterprise procurement, not retail

This silicon is typically acquired through an authorized OEM partner, system integrator, or hyperscaler reseller. For on-demand access, compare hourly rates at RunPod, Vast.ai, Lambda Labs, or your existing cloud provider before committing to capital expenditure.

  • +Use $3.0k only as a rough anchor. Enterprise street pricing moves with supply, warranty, and included accessories.
  • +Confirm cooling, power delivery, and return terms before you purchase. These parts often ship without consumer-friendly safeguards.
  • +If this is for production, compare against authorized reseller quotes before you commit.

There is insufficient comparable evidence for a purchase recommendation. Check model fit, software support and a current seller quote.

Affiliate note: we do not have a device-level ASIN yet, so this opens tagged Amazon search results for the exact product name.