LLM Comparison Dashboard · July 2026

Four leaderboards, one dataset

28 frontier models across 13 labs, ranked by four different objectives. Auto-refreshed weekly from BenchLM. Click any header to expand. Filters apply live across every table.

0%
$30
90s
0K
🧠

Sorted by GPQA Diamond

Pure reasoning capability. Graduate-level physics, chemistry, biology.
💰

Best money for GPQA

Reasoning capability per dollar. Value = GPQA² / √(blended $/M).
⌨️

Best coder

SWE-bench Verified + Terminal-Bench + LiveCodeBench blend. End-to-end shipping code.
🔍

Best code reviewer

Reading long diffs, catching bugs, following review conventions. SWRBench-style profile.
🎼

Best aggregator / orchestrator

Multi-agent conductor: long context + instruction adherence + reasoning breadth.
Ranked for the NVIDIA DGX Spark — the GB10 Grace Blackwell Superchip desktop. All numbers are measured single-node, single-stream (batch=1) tokens/sec from published NVIDIA, community, and llama-benchy runs. Score blends fluency (tok/s decode), quality (GPQA/coding benchmark of the base model), and comfort of fit in 128 GB unified memory. Models that don't clear ~15 tok/s at single-stream are excluded — they run, but not fluently.
ChipGB10 Grace Blackwell
Unified memory128 GB LPDDR5X
Bandwidth273 GB/s
FP4 peak~1 PFLOPS
Networking200 GbE / NVLink-C2C
List price$3,999–$4,699
0
120GB
0%

Fluency score · best all-round on GB10

Blend of tok/s (40%), model quality (35%), fit comfort (15%), context headroom (10%). Top 3 highlighted.
🚀

Sorted by raw tokens/sec

Single-stream decode speed on a single DGX Spark. Higher = more interactive chat feel.
🎯

Sorted by quality (GPQA + coding)

Ignoring speed — the strongest brains you can host locally on a single GB10.
Local — DGX Spark GB10

One-time $4,699 desktop

Best local reasoningQwen3.5 122B · 85% GPQA
Best local coderQwen3 Coder Next · 82% coding
Best local speedGPT-OSS 20B · 83 tok/s
Data leaves your boxNever
Peak context1M (Qwen3.5 35B YaRN)
Frontier — hosted APIs

Pay per token, no ceiling

Best reasoningGemini 3.5 Pro · 95.5% GPQA
Best coderClaude Fable 5 · 94% coding
Cheapest at frontierDeepSeek V4 Pro · $0.54/M
Data leaves your boxEvery request
Peak context2M (Gemini 3.5 Pro)

How long until the DGX Spark pays for itself?

2.0M
22
$0.15
Monthly API cost$0
Monthly Spark power$0
Monthly savings$0
Payback period
🎯

Pick 3 local + 3 frontier · see the trade-offs

Two blocks, three dropdowns each. Rows show the six metrics that matter. The radar overlays all six picks so profile-level trade-offs (reasoning vs. speed vs. cost) become visible at a glance.

Local · runs on your DGX Spark 3 slots

Model Overall Reasoning Coding Cost Speed Context

Frontier · hosted API 3 slots

Model Overall GPQA Coding Cost / M Latency Context

Overall profile radar

Six axes scored 0-100, higher is better on every axis. Cost-efficiency inverts price (local models sit near the outer ring; frontier flagships sit near the center on this axis).
⚖️

Head-to-head pairs

Each local pick matched with its closest frontier peer. Deltas show what you give up (red) or gain (green) by going local.
🧭

Which should I use, by workload?

Concrete recommendations for common use cases. When local wins, when the API wins, and when to run both.
NVIDIA

DGX Spark GB10

$3,999–$4,699 · 240W
ChipGrace Blackwell GB10
Memory128 GB LPDDR5X
Bandwidth273 GB/s
Peak FP4~1 PFLOPS
SoftwareCUDA / TensorRT / vLLM
OSUbuntu (DGX OS)
Best at: Prompt processing, fine-tuning, CUDA-first stacks. FP4 tensor cores handle MXFP4 kernels 5x faster than the others.
Apple

Mac Studio M4 Max (64 GB)

$2,499 base · $3,799 with 1 TB · ~150W
ChipApple M4 Max
Memory64 GB unified (max)
Bandwidth546 GB/s
GPU cores40-core
SoftwareMLX / Ollama / LM Studio
OSmacOS 15
Best at: Fastest token generation on models under ~50 GB. 546 GB/s memory bandwidth, silent operation. Apple discontinued the 128 GB M4 Max on June 25, 2026 due to the DRAM shortage; the M3 Ultra 96 GB is technically listed but ships in 13-14 weeks with resellers marked discontinued/sold out ahead of the M5 Ultra refresh.
AMD · Framework Desktop

Strix Halo (Ryzen AI Max+ 395)

$2,000–$2,600 · 120W
ChipRyzen AI Max+ 395
Memory128 GB LPDDR5X-8000
Bandwidth256 GB/s
GPURadeon 8060S (40 CU)
SoftwareROCm / Vulkan / llama.cpp
OSLinux / Windows
Best at: Value. Half the price of Spark or Mac at ~90% of Spark's token generation. Linux-first, ROCm still maturing.
📊

Same model, three machines — measured tok/s

Single-stream decode speed for each shared model. Numbers from NVIDIA, MLX, llama.cpp community benchmarks. Winner per row gets a gold badge.
DGX Spark GB10
Mac M4 Max (64 GB)
Strix Halo
🧮

Which one should I buy?

Workload-based recommendations. If none of the three fits, the fourth column tells you what does.