28 frontier models across 13 labs, ranked by four different objectives. Auto-refreshed weekly from BenchLM. Click any header to expand. Filters apply live across every table.
0%
$30
90s
0K
🧠
Sorted by GPQA Diamond
Pure reasoning capability. Graduate-level physics, chemistry, biology.
▶
💰
Best money for GPQA
Reasoning capability per dollar. Value = GPQA² / √(blended $/M).
Reading long diffs, catching bugs, following review conventions. SWRBench-style profile.
▶
🎼
Best aggregator / orchestrator
Multi-agent conductor: long context + instruction adherence + reasoning breadth.
▶
Ranked for the NVIDIA DGX Spark — the GB10 Grace Blackwell Superchip desktop. All numbers are measured single-node, single-stream (batch=1) tokens/sec from published NVIDIA, community, and llama-benchy runs. Score blends fluency (tok/s decode), quality (GPQA/coding benchmark of the base model), and comfort of fit in 128 GB unified memory. Models that don't clear ~15 tok/s at single-stream are excluded — they run, but not fluently.
ChipGB10 Grace Blackwell
Unified memory128 GB LPDDR5X
Bandwidth273 GB/s
FP4 peak~1 PFLOPS
Networking200 GbE / NVLink-C2C
List price$3,999–$4,699
0
120GB
0%
⚡
Fluency score · best all-round on GB10
Blend of tok/s (40%), model quality (35%), fit comfort (15%), context headroom (10%). Top 3 highlighted.
▶
🚀
Sorted by raw tokens/sec
Single-stream decode speed on a single DGX Spark. Higher = more interactive chat feel.
▶
🎯
Sorted by quality (GPQA + coding)
Ignoring speed — the strongest brains you can host locally on a single GB10.
▶
Local — DGX Spark GB10
One-time $4,699 desktop
Best local reasoningQwen3.5 122B · 85% GPQA
Best local coderQwen3 Coder Next · 82% coding
Best local speedGPT-OSS 20B · 83 tok/s
Data leaves your boxNever
Peak context1M (Qwen3.5 35B YaRN)
Frontier — hosted APIs
Pay per token, no ceiling
Best reasoningGemini 3.5 Pro · 95.5% GPQA
Best coderClaude Fable 5 · 94% coding
Cheapest at frontierDeepSeek V4 Pro · $0.54/M
Data leaves your boxEvery request
Peak context2M (Gemini 3.5 Pro)
How long until the DGX Spark pays for itself?
2.0M
22
$0.15
Monthly API cost$0
Monthly Spark power$0
Monthly savings$0
Payback period—
🎯
Pick 3 local + 3 frontier · see the trade-offs
Two blocks, three dropdowns each. Rows show the six metrics that matter. The radar overlays all six picks so profile-level trade-offs (reasoning vs. speed vs. cost) become visible at a glance.
▶
Local · runs on your DGX Spark 3 slots
Model
Overall
Reasoning
Coding
Cost
Speed
Context
Frontier · hosted API 3 slots
Model
Overall
GPQA
Coding
Cost / M
Latency
Context
Overall profile radar
Six axes scored 0-100, higher is better on every axis. Cost-efficiency inverts price (local models sit near the outer ring; frontier flagships sit near the center on this axis).
⚖️
Head-to-head pairs
Each local pick matched with its closest frontier peer. Deltas show what you give up (red) or gain (green) by going local.
▶
🧭
Which should I use, by workload?
Concrete recommendations for common use cases. When local wins, when the API wins, and when to run both.
▶
NVIDIA
DGX Spark GB10
$3,999–$4,699 · 240W
ChipGrace Blackwell GB10
Memory128 GB LPDDR5X
Bandwidth273 GB/s
Peak FP4~1 PFLOPS
SoftwareCUDA / TensorRT / vLLM
OSUbuntu (DGX OS)
Best at: Prompt processing, fine-tuning, CUDA-first stacks. FP4 tensor cores handle MXFP4 kernels 5x faster than the others.
Apple
Mac Studio M4 Max (64 GB)
$2,499 base · $3,799 with 1 TB · ~150W
ChipApple M4 Max
Memory64 GB unified (max)
Bandwidth546 GB/s
GPU cores40-core
SoftwareMLX / Ollama / LM Studio
OSmacOS 15
Best at: Fastest token generation on models under ~50 GB. 546 GB/s memory bandwidth, silent operation. Apple discontinued the 128 GB M4 Max on June 25, 2026 due to the DRAM shortage; the M3 Ultra 96 GB is technically listed but ships in 13-14 weeks with resellers marked discontinued/sold out ahead of the M5 Ultra refresh.
AMD · Framework Desktop
Strix Halo (Ryzen AI Max+ 395)
$2,000–$2,600 · 120W
ChipRyzen AI Max+ 395
Memory128 GB LPDDR5X-8000
Bandwidth256 GB/s
GPURadeon 8060S (40 CU)
SoftwareROCm / Vulkan / llama.cpp
OSLinux / Windows
Best at: Value. Half the price of Spark or Mac at ~90% of Spark's token generation. Linux-first, ROCm still maturing.
📊
Same model, three machines — measured tok/s
Single-stream decode speed for each shared model. Numbers from NVIDIA, MLX, llama.cpp community benchmarks. Winner per row gets a gold badge.
▶
DGX Spark GB10
Mac M4 Max (64 GB)
Strix Halo
🧮
Which one should I buy?
Workload-based recommendations. If none of the three fits, the fourth column tells you what does.