LLM Comparison Dashboard

Five leaderboards, one dataset

42 frontier models across 19 labs, plus 3 Tencent Hy-MT2 translation specialists in their own section. Ranked by five different objectives, auto-refreshed weekly from BenchLM. Click any header to expand. Filters apply live across every table.

0%
$30
90s
0K
🧩

Sorted by Artificial Analysis Intelligence Index

AA's headline composite (MMLU-Pro, GPQA Diamond, HLE, LiveCodeBench, SciCode, MATH-500, IFBench, AIME). Omniscience = accuracy minus hallucination rate; positive means the model knows what it doesn't know. Dashes are for models AA has not yet indexed.
▶
🧠

Sorted by GPQA Diamond

Pure reasoning capability. Graduate-level physics, chemistry, biology.
▶
💰

Best money for GPQA

Reasoning capability per dollar. Value = GPQA² / √(blended $/M).
▶
⌨️

Best coder

SWE-bench Verified + Terminal-Bench + LiveCodeBench blend. End-to-end shipping code.
▶
🔍

Best code reviewer

Reading long diffs, catching bugs, following review conventions. SWRBench-style profile.
▶
🎼

Best aggregator / orchestrator

Multi-agent conductor: long context + instruction adherence + reasoning breadth.
▶
🌐

Translation specialists

Purpose-built translation models. Ranked by FLORES-200 XX⇄XX (XCOMET-XXL) alongside price. These aren't reasoning models — the general leaderboards above are the wrong lens for them.
▶
Ranked for the NVIDIA DGX Spark — the GB10 Grace Blackwell Superchip desktop. All numbers are measured single-node, single-stream (batch=1) tokens/sec from published NVIDIA, community, and llama-benchy runs. Score blends fluency (tok/s decode), quality (GPQA/coding benchmark of the base model), and comfort of fit in 128 GB unified memory. Models that don't clear ~15 tok/s at single-stream are excluded — they run, but not fluently.
ChipGB10 Grace Blackwell
Unified memory128 GB LPDDR5X
Bandwidth273 GB/s
FP4 peak~1 PFLOPS
Networking200 GbE / NVLink-C2C
List price$3,999–$4,699
0
120GB
0%
⚡

Fluency score · best all-round on GB10

Blend of tok/s (40%), model quality (35%), fit comfort (15%), context headroom (10%). Top 3 highlighted.
▶
🚀

Sorted by raw tokens/sec

Single-stream decode speed on a single DGX Spark. Higher = more interactive chat feel.
▶
🎯

Sorted by quality (GPQA + coding)

Ignoring speed — the strongest brains you can host locally on a single GB10.
▶
🎨

Image generation on the Spark

Open-weight image models you can run on a single GB10. Click any row for install commands. Arena Elo from Pixazo's Text-to-Image leaderboard.
▶
🎬

Video generation on the Spark

Open-weight text-to-video models on GB10. Time measured for a 5-second 720p clip. Quality Elo from The Open Weights T2V leaderboard.
▶
Local — DGX Spark GB10

One-time $4,699 desktop

Best local reasoningOrnith-1.5 35B A3B · 93 quality
Best local coderQwen3.8 27B · 88% coding
Best local speedNemotron 3.5 Lightning · 108 tok/s
Data leaves your boxNever
Peak context1M (Nemotron 3.5 / Qwen3.5 YaRN)
Frontier — hosted APIs

Pay per token, no ceiling

Best reasoningFugu Ultra · 95.5% GPQA
Best coderClaude Opus 5 · 96% coding
Cheapest at frontierHunyuan Hy3 · $0.25/M blended
Data leaves your boxEvery request
Peak context2M (Gemini 3.5 Pro)

How long until the DGX Spark pays for itself?

2.0M
22
$0.15
Monthly API cost$0
Monthly Spark power$0
Monthly savings$0
Payback period—
🎯

Pick 3 local + 3 frontier · see the trade-offs

Two blocks, three dropdowns each. Rows show the six metrics that matter. The radar overlays all six picks so profile-level trade-offs (reasoning vs. speed vs. cost) become visible at a glance.
▶

Local · runs on your DGX Spark 3 slots

Model Overall Reasoning Coding Cost Speed Context

Frontier · hosted API 3 slots

Model Overall GPQA Coding Cost / M Latency Context

Overall profile radar

Six axes scored 0-100, higher is better on every axis. Cost-efficiency inverts price (local models sit near the outer ring; frontier flagships sit near the center on this axis).
⚖️

Head-to-head pairs

Each local pick matched with its closest frontier peer. Deltas show what you give up (red) or gain (green) by going local.
▶
🧭

Which should I use, by workload?

Concrete recommendations for common use cases. When local wins, when the API wins, and when to run both.
▶
NVIDIA

DGX Spark GB10

$3,999–$4,699 · 240W
ChipGrace Blackwell GB10
Memory128 GB LPDDR5X
Bandwidth273 GB/s
Peak FP4~1 PFLOPS
SoftwareCUDA / TensorRT / vLLM
OSUbuntu (DGX OS)
Best at: Prompt processing, fine-tuning, CUDA-first stacks. FP4 tensor cores handle MXFP4 kernels 5x faster than the others.
Apple

Mac Studio M4 Max (64 GB)

$2,499 base · $3,799 with 1 TB · ~150W
ChipApple M4 Max
Memory64 GB unified (max)
Bandwidth546 GB/s
GPU cores40-core
SoftwareMLX / Ollama / LM Studio
OSmacOS 15
Best at: Fastest token generation on models under ~50 GB. 546 GB/s memory bandwidth, silent operation. Apple discontinued the 128 GB M4 Max on June 25, 2026 due to the DRAM shortage; the M3 Ultra 96 GB is technically listed but ships in 13-14 weeks with resellers marked discontinued/sold out ahead of the M5 Ultra refresh.
AMD · Framework Desktop

Strix Halo (Ryzen AI Max+ 395)

$2,000–$2,600 · 120W
ChipRyzen AI Max+ 395
Memory128 GB LPDDR5X-8000
Bandwidth256 GB/s
GPURadeon 8060S (40 CU)
SoftwareROCm / Vulkan / llama.cpp
OSLinux / Windows
Best at: Value. Half the price of Spark or Mac at ~90% of Spark's token generation. Linux-first, ROCm still maturing.
📊

Same model, three machines — measured tok/s

Single-stream decode speed for each shared model. Numbers from NVIDIA, MLX, llama.cpp community benchmarks. Winner per row gets a gold badge.
▶
DGX Spark GB10
Mac M4 Max (64 GB)
Strix Halo
🧮

Which one should I buy?

Workload-based recommendations. If none of the three fits, the fourth column tells you what does.
▶
🌀

Frontier image APIs · hosted, ranked by Arena Elo

Nine hosted flagship image models across nine labs. Elo from the 2026 Text-to-Image Arena. Pricing normalized to per-image at three quality tiers.
▶