Overview

21/44
TBench Scored
21/44
SWE-Pro Scored
9
Routing Tasks
22
Open-weight Models

Recently Added

modelGPT-5.6 TerraTB 87.4 · half Sol cost
modelGPT-5.6 LunaTB 84.7 · best $/TB
modelGrok 4.5TB 83.3 · $6/M
modelMuse Spark 1.1MCP-Atlas leader
providerCognition DevinSWE-1.7

Benchmark Leaders

Terminal-Bench
shell agent loops
GPT-5.6 Sol ★ SOTA (91.9 ultra)88.8
Claude Mythos 5 (v) restricted88
GPT-5.6 Terra ★ half Sol cost87.4
GPT-5.6 Luna ★ best $/TB84.7
Claude Fable 5 ★ restored84.3
SWE-bench Pro
single-pass code review
Claude Fable 5 (v) classifier reroutes80
Claude Mythos 5 (v) restricted77.8
SWE-1.7 (m)77.8
DeepSeek V4 Pro (a)76.2
Fugu Ultra (v)73.7
MCP-Atlas
multi-tool use
Muse Spark 1.1 ★★ new leader88.1
Gemini 3.5 Flash 83.6
Claude Fable 5 83.3
Claude Opus 4.8 77.8
GLM-5.2 76.8
Speed
tokens per second
Step 3.7F387 tok/s
Luna216 tok/s
Nemotron 3 Ultra197 tok/s
Qwen3.7 Max197 tok/s
Nemotron 3U197 tok/s
SWE-Marathon & Long-Horizon Evidence
sustained multi-hour agent quality
Claude Fable 5 DeepSWE #1 ★70%
Claude Opus 4.8 Marathon ★26.0
Qwen3.7 Max 35h proven ★1,158 calls
GLM-5.2 Marathon13.0
DeepSeek V4 Pro DeepSWE ✗8%
Missing: GPT-5.6 Sol, GPT-5.6 Terra, GPT-5.6 Luna, Grok 4.5, Muse Spark 1.1, DeepSeek V4.5, Gemini 3.5 Flash, Step 3.7, Hy3, Sonnet 5, MiniMax M3, SWE-1.7 — no Marathon/DeepSWE scores yet. Sol DeepSWE v1.1 72.7%, Terra 69.6%, Luna 67.2% suggest strong long-horizon potential.
All benchmarks →

Budget Picker

Best model per spend tier
neutral ranking by budget band
Budget / TaskModelWhy
Free codingHy3Apache 2.0, OpenRouter free tier, 54 TB
Cheap agent (<$0.30/M out)DS V4 Flash$0.14/$0.28, TB 56.9, SWE-V 79.0
Value agent (<$1/M out)GLM-5.2$1.40/$4.40, TB 81, SWE-Pro 62.1
Open-weight agentMiniMax M3$0.30/$1.20, TB 66 (v), 1M ctx
Mid agent ($1-4/M out)Grok 4.5$2/$6, TB 83.3, SWE-Pro 64.7
Quality plan ($4-8/M out)GPT-5.6 Luna$1/$6, TB 84.7, SWE-Pro 62.7
Best agent (any budget)GPT-5.6 SolTB 88.8 SOTA, DeepSWE 72.7, AA Coding #1
Full decision tools →

Task Leaders

TaskLeaderScore / Note
Long-horizon multi-tool agentOpus 4.8Marathon 26.0
Terminal-agent loopsGPT-5.6 Sol★ 88.8 SOTA
Single-pass code reviewOpus 4.8SWE-Pro 69.2
Monorepo analysisGemini 3.1 Pro2M ctx
Frontend / web devGemini 3.1 ProWebDev #1
High-volume batch / cheapDS V4 Flash$0.28/M
Security reviewOpus 4.8
Planning / spec modeSonnet 5GDPval 1618
Long-horizon autonomousGPT-5.6 SolDeepSWE 72.7 ★
Full routing matrix →

Pricing Tiers

Budget (<$0.30/M out)
Agnes 2.0$0.20Cheapest reasoning
Hy3$0.21Apache 2.0
DS V4 Flash$0.28Batch floor
Value ($0.30-1.15/M out)
Step 3.5 Flash$0.30Cheapest coding/$
Nemotron 3U$2.20300+ tok/s
DS V4 Pro$0.87Single-pass
DS V4.5$1.10Open SWE-Pro 62.1
Step 3.7$1.15Free tier
Mid ($1.20-4.40/M out)
MiniMax M3$1.201M ctx, open TB leader
GLM-5.2$1.40-4.40Current main
Grok Build$2Fast coder
Ring 2.6 1T$2.501T MoE
Muse Spark 1.1$4.25MCP Atlas 88.1 leader
ERNIE 5.1$2.65Chinese only
Qwen3.7 Max$1.25-3.7535h proven
MiMo 2.5 Pro$3.20Untested
Kimi K2.7$4MCP-Mark 81.1
DS V4 Pro Max$3.20BenchLM #13
Premium ($5+/M out)
Grok 4.5$6TB 83.3, cost-efficient
Claude Haiku 4.5$5Fast, unverified
Gemini 3.5 Flash$9MCP #2 (after Spark)
Sonnet 5$10Beats Opus TB
Gemini 3.1 Pro$122M ctx
GPT-5.6 Luna$6TB 84.7, best $/TB
GPT-5.6 Terra$15TB 87.4, half Sol cost
Opus 4.8$25Marathon 26.0
GPT-5.6 Sol$30TB 88.8 SOTA
Fable 5$50Restored, classifier gap
All models →

Value Frontier

Pareto-optimal: price vs Terminal-Bench
no other model is both cheaper and higher
Hy3$0.21/MTB 54.4
MiniMax M3$1.2/MTB 66
GLM-5.2$4.4/MTB 81
Luna$6/MTB 84.7
Terra$15/MTB 87.4
GPT-5.6 Sol ★$30/MTB 88.8
Cost-Performance Map
Terminal-Bench score vs $/M output (log scale)
$0.1$0.2$0.5$1$2$5$10$25$50020406080$/M outputTB scoreHy3 — TB 54.4, $0.21/MHy3DS Flash — TB 56.9, $0.28/MDS FlashStep 3.5F — TB 51, $0.3/MStep 3.5FDS V4 Pro — TB 67.9, $0.87/MDS V4 ProNemotron — TB 56.4, $2.2/MNemotronStep 3.7F — TB 59.5, $1.15/MStep 3.7FMiniMax M3 — TB 66, $1.2/MMiniMax M3Mistral L3 — TB 12, $1.5/MMistral L3Muse Spark — TB 80, $4.25/MMuse SparkGLM-5.2 — TB 81, $4.4/MGLM-5.2Grok 4.5 — TB 83.3, $6/MGrok 4.5Luna — TB 84.7, $6/MLunaGemini 3.5F — TB 76.2, $9/MGemini 3.5FSonnet 5 ★ — TB 80.4, $10/MSonnet 5 ★Gemini 3.1P — TB 54.2, $12/MGemini 3.1POpus 4.8 — TB 79, $25/MOpus 4.8Terra — TB 87.4, $15/MTerraGPT-5.6 Sol ★ — TB 88.8, $30/MGPT-5.6 Sol ★Fable 5 ★ — TB 84.3, $50/MFable 5 ★
Full cost-performance map →

Coverage Gaps

Terminal-Bench
21 scored · 5 missing
Show missing models
ERNIE 5.1DS V4 Pro MaxDS V4.5Agnes 2.0 FlashFugu Ultra
SWE-bench Pro
21 scored · 8 missing
Show missing models
Kimi K2.7 Code (has SWE-V 62.0 not Pro)Hy3 (has SWE-V 74.0 not Pro)Llama 4 Maverick (has SWE-V 70.4 not Pro)Mistral Large 3 (has SWE-V 76.2 not Pro)Nemotron 3 Ultra (has SWE-V 71.9 not Pro)Step 3.5 Flash (has SWE-V 74.4 not Pro)MiMo v2.5 Pro, Ring 2.6 1T, Qwen3.7-Plus, Command A+, North Mini Code 1.0, Cohere North Mini Code — no SWE-Pro or SWE-V dataAgnes 2.0 Flash, ERNIE 5.1, Gemma 4 12B, DiffusionGemma 26B-A4B, Seed 2.1 Pro, Seed 2.1 Turbo — no SWE data available
MCP-Atlas
10 scored · 27 missing
Show missing models
Claude Sonnet 5Claude Mythos 5Grok 4.5Grok Build 0.1GPT-5.6 SolGPT-5.6 TerraGPT-5.6 LunaNemotron 3 UltraNemotron 3 SuperStep 3.7 FlashStep 3.5 FlashLlama 4 MaverickQwen3-Coder-NextHy3Ring 2.6 1TMiMo v2.5 ProERNIE 5.1Agnes 2.0 FlashDS V4 Pro MaxDS V4.5SWE-1.7Composer 2.5Gemma 4 12BDiffusionGemma 26B-A4BSeed 2.1 ProSeed 2.1 TurboScale MCP-Atlas leaderboard last updated April 2026 — does not include models released after that date

Data Freshness

Updated today
llm-stats.comopenrouter.aiopenai.comx.aiartificialanalysis.aimeta.comswfte.com