Skip to content
AI Model Radar
IndexCollected 3 h ago

Hardware profile

Own the middle.

DGX Spark is not a faster graphics card. It is a different shape of machine: one pool of memory big enough for models a consumer GPU cannot hold at all.

EstimatedEstimated: computed from our curated model and hardware catalog — not a live reading.

Why unified memory matters

A consumer GPU keeps model weights in its own VRAM. When a model does not fit, it spills into system RAM over PCIe and slows to a crawl. A unified-memory system gives CPU and GPU one large pool — the model simply fits. Capacity, not raw speed, is what it buys you.

Capacity versus speed
MetricDGX SparkRTX 4090
Memory128 GB24 GB
Bandwidth273 GB/s1008 GB/s

Roughly five times the memory, roughly a quarter of the bandwidth. Spark runs models a discrete GPU cannot hold at all — but on a model both can hold, the GPU generates tokens faster.

NVIDIA's published claims

Inference, one unit
Models up to 200B parameters
Fine-tuning, one unit
Models up to 70B parameters
Two units linked
Models up to 405B parameters

Our own arithmetic

Not an NVIDIA statement: 200 billion parameters only fit in 128 GB at roughly 4-bit precision (200B × 0.5 bytes ≈ 100 GB of weights) before any context memory. NVIDIA does not publish the precision behind that number, so read “200B-class” as heavily quantized.

Honest caveats

Why we publish no tokens-per-second figure

Tokens per second depend on the model, its quantization, the context length and the runtime — and every runtime update moves the number. A single figure would be wrong for most readers and stale within weeks, so this site publishes none. It measures what it can measure on a schedule: memory, prices, sizes.

What can be said comes from physics, not benchmarks. Generating a token means reading most of the model's active weights out of memory, so memory bandwidth is the hard ceiling on speed. DGX Spark's 273 GB/s against roughly 1,000 GB/s on a top consumer graphics card means that on a model which fits both, the card's ceiling is about four times higher — real runtimes reach only a fraction of either ceiling, but the ratio between them holds.

The consequence: Spark buys capacity — models that would not fit a consumer card at all — not speed. If you want fast output on a model that already fits a GPU, the GPU wins; if you want a model that needs 100 GB of memory to exist on your desk at all, nothing consumer-priced competes.

Bandwidth figures are vendor specifications (NVIDIA: 273 GB/s for DGX Spark; 1,008 GB/s for the GeForce RTX 4090). The ratio is arithmetic on those two numbers, not a measurement of any model.

Two Sparks, or four?

NVIDIA's own product page states that its ConnectX networking lets up to four DGX Spark systems be connected to work with models of up to 700 billion parameters. That is a capacity claim, and it is real: four pooled machines hold four times the weights.

It is not a speed claim. The link between boxes is a network cable, orders of magnitude slower than the memory inside each box, so a model split across units generates at the pace of that link, not of the chips. And the parameter figure carries the same fine print as the single-box claim: it assumes heavy quantization, and the precision behind the number is not stated.

Honest reading: linking Sparks is for people who need a very large model to fit at all and can accept slow output. If the goal is speed on a model that fits one box, the second box buys nothing — and for an occasional big job, an hour of a rented data-center GPU is cheaper than a second machine.

Source: NVIDIA DGX Spark product page ("up to four NVIDIA DGX Spark systems … AI models of up to 700 billion parameters").

What fits in 128 GB

Context length: 8K tokens · f16 (default)

Model fit on the selected hardware
ModelFitEst. memoryEstimated memory = weights + context + runtime
Ling 3.0 FlashInclusionAI · 127B MoE · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Excellent fit~73.4 GB70.1 + 0.4 + 3.0
GLM-4.5-AirZhipu AI · 110B MoE · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Excellent fit~72.3 GB68.0 + 1.4 + 2.9
gpt-oss 120BOpenAI · 117B MoE · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~61.4 GB58.5 + 0.3 + 2.6
Qwen3-Coder NextAlibaba Qwen · 79.7B MoE · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~48.1 GB45.1 + 0.8 + 2.2
Llama 3.3 70BMeta · 70.6B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~44.1 GB39.6 + 2.5 + 2.0
Qwen AgentWorld 35B-A3BAlibaba Qwen · 34.7B MoE · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Excellent fit~22.2 GB20.6 + 0.2 + 1.5
Ornith 1.5 35B-A3BOrnith AI · 36B MoE · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~21.8 GB20.2 + 0.2 + 1.5
KAT-Coder V2.5Kwaipilot · 34.7B MoE · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Excellent fit~21.5 GB19.9 + 0.2 + 1.4
Qwen3 32BAlibaba Qwen · 32.8B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~21.8 GB18.4 + 2.0 + 1.4
Qwen3-Coder 30B-A3BAlibaba Qwen · 30.5B MoE · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~19.4 GB17.3 + 0.8 + 1.4
Gemma 4 31BGoogle · 31.3B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~20.5 GB17.1 + 2.0 + 1.4
GLM-4.7-FlashZhipu AI · 31.2B MoE · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Excellent fit~18.9 GB17.1 + 0.4 + 1.4
Granite 4.2 30BIBM · 29.3B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~19.8 GB16.5 + 2.0 + 1.3
Gemma 4 26B-A4BGoogle · 25.8B MoE · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~17.6 GB15.8 + 0.5 + 1.3
Gemma 3 27BGoogle · 27.4B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~17.7 GB15.4 + 1.0 + 1.3
Qwen3.8 27BAlibaba Qwen · 27.8B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~17.1 GB15.3 + 0.5 + 1.3
Mistral Small 3.2Mistral AI · 24B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~15.8 GB13.3 + 1.3 + 1.2
Devstral Small 2 24BMistral AI · 24B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~15.8 GB13.3 + 1.3 + 1.2
gpt-oss 20BOpenAI · 20.9B MoE · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~12.2 GB10.8 + 0.2 + 1.2
Qwen3 14BAlibaba Qwen · 14.8B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~10.8 GB8.4 + 1.3 + 1.1
Gemma 4 12BGoogle · 12B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~8.5 GB6.6 + 0.8 + 1.0
Nemotron Nano 9B v2NVIDIA · 8.9B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Excellent fit~8.9 GB6.1 + 1.8 + 1.0
Ornith 1.5 9BOrnith AI · 9.7B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~6.7 GB5.4 + 0.3 + 1.0
Granite 4.2 8BIBM · 8.8B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~7.3 GB5.0 + 1.3 + 1.0
LFM2.5 8B-A1BLiquid AI · 8.5B MoE · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~5.9 GB4.8 + 0.1 + 1.0
Qwen3 8BAlibaba Qwen · 8.2B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~6.8 GB4.7 + 1.1 + 1.0
Gemma 4 E4BGoogle · 8B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~5.7 GB4.6 + 0.1 + 1.0
Ling 3.0 TinyInclusionAI · 7.9B MoE · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Excellent fit~5.7 GB4.5 + 0.2 + 1.0
Qwen3 4BAlibaba Qwen · 4B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~4.3 GB2.3 + 1.1 + 0.9
Phi-4 MiniMicrosoft · 3.8B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~4.2 GB2.3 + 1.0 + 0.9
Granite 4.2 3BIBM · 3.7B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~3.6 GB2.1 + 0.6 + 0.9
LFM2.5 2.6BLiquid AI · 2.7B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~2.6 GB1.6 + 0.1 + 0.9
LFM2.5 1.2BLiquid AI · 1.2B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~1.7 GB0.7 + 0.1 + 0.9
Qwen3 0.6BAlibaba Qwen · 0.6B · Q8_0GGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Excellent fit~2.3 GB0.6 + 0.9 + 0.9
Qwen3.8 Flash-NextAlibaba Qwen · 180B MoE · Q4_KGGUF ↗ (External link)Original ↗ (External link)Tight fit~107.8 GB103.7 + 0.2 + 4.0
Qwen3.8 2.4T-A95BAlibaba Qwen · 2.4T-A95B · IQ4_XSGGUF ↗ (External link)Original ↗ (External link)Not recommended~1259.0 GB1220.8 + 0.7 + 37.5
DeepSeek-V4-ProDeepSeek · 1650B MoE · Q4_KGGUF ↗ (External link)Original ↗ (External link)Not recommended~816.4 GB791.3 + 0.5 + 24.6
GLM-5.3Zhipu AI · 753B · Q4_KGGUF ↗ (External link)Original ↗ (External link)Not recommended~449.8 GB435.2 + 0.7 + 13.9
GLM-5.2Zhipu AI · 753B MoE · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Not recommended~448.3 GB433.8 + 0.7 + 13.9
Ornith 1.5 397BOrnith AI · 403B MoE · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Not recommended~235.4 GB227.5 + 0.2 + 7.7
GLM-5.3-FlashZhipu AI · 321B · Q4_KGGUF ↗ (External link)Original ↗ (External link)Not recommended~192.5 GB186.0 + 0.1 + 6.4
Hy3Tencent · 299B MoE · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Not recommended~174.6 GB166.3 + 2.5 + 5.8
DeepSeek-V4-FlashDeepSeek · 304B MoE · Q4_KGGUF ↗ (External link)Original ↗ (External link)Not recommended~150.0 GB144.4 + 0.4 + 5.2
Llama 3.1 8BMeta · 8B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Unknown
Llama 3.2 3BMeta · 3B · Q4_K_MGGUF ↗ (External link)Original ↗ (External link)Ollama ↗ (External link)Unknown

All memory figures are estimates: measured quantized file size + computed context memory + runtime overhead, with a 12% safety margin on your hardware. Real usage varies with runtime version and settings.

Spark or rent?

Buying makes sense when large local models are your daily routine and your data must stay in the building. If the heavy work happens a few times a month, renting is cheaper — the break-even maths is on the rental page.

Break-even

Who it is for

  • + You run 70B-class or larger models locally, most days.
  • + Your data cannot leave your desk.
  • + Capacity matters more to you than tokens per second.
  • + You are comfortable on an ARM64 Linux machine.