The Known Good Updated 16 Sep 2026 Subscribe

Leaderboard

Every tracked model that holds a Known Good Index score, ranked. Sort any column; a figure we have not ingested shows as an em dash rather than a guess.

65 ranked of 1128 tracked models All models → Index methodology → Download as CSV →
Known Good Index leaderboard

Known Good Index combines seven public evaluations — GPQA Diamond, Mock AIME 2024–25, MATH Level 5, SWE-bench Verified, LiveBench, Humanity's Last Exam, Terminal-Bench — each min-max normalised across tracked models, then averaged with equal weight and scaled 0–100. Scores ingested from Epoch AI (CC-BY) and LiveBench; price and speed from OpenRouter, metadata from models.dev. We do not run these evaluations.

# Model Creator Known Good Index Coding Blended $/1M Output tok/s Context
1 gpt-5.5-pre-release OpenAI 97.5 based on 3 of 7 evaluations
2 Google: Gemini 3.1 Pro Preview 💡 Google 97.1 based on 4 of 7 evaluations $4.50 102 1.05M
3 Google: Gemini 3.5 Flash 💡 Google 94.8 based on 3 of 7 evaluations $3.38 166 1.05M
4 DeepSeek: DeepSeek V4 Pro 0423 💡 DeepSeek 93.3 based on 3 of 7 evaluations $2.00 46 1.05M
5 OpenAI: GPT-5.5 💡 OpenAI 92.8 based on 3 of 7 evaluations $11.25 74 1.05M
6 Qwen: Qwen3.7 Max 💡 Alibaba 92.7 based on 3 of 7 evaluations $2.21 58 1.00M
7 Anthropic: Claude Opus 4.7 💡 Anthropic 92.5 based on 5 of 7 evaluations 97.2 $10.00 55 1.00M
8 MoonshotAI: Kimi K2.6 💡 Moonshot AI 92.5 based on 3 of 7 evaluations $1.71 40 262k
9 OpenAI: GPT-5.4 💡 OpenAI 91.1 based on 5 of 7 evaluations 91.9 $5.62 78 1.05M
10 Z.ai: GLM 5.2 💡 Z.ai 90.9 based on 3 of 7 evaluations $2.15 68 1.05M
11 Z.ai: GLM 5.1 💡 Z.ai 89.6 based on 3 of 7 evaluations $1.48 41 205k
12 Qwen: Qwen3.6 Max Preview 💡 Alibaba 89.5 based on 3 of 7 evaluations $2.31 50 262k
13 Anthropic: Claude Opus 4.6 💡 Anthropic 88.9 based on 5 of 7 evaluations 92.5 $10.00 34 1.00M
14 OpenAI: o3 Mini 💡 OpenAI 85.7 based on 4 of 7 evaluations $1.93 127 200k
15 gemini-3-pro-preview Google 83.8 based on 5 of 7 evaluations 75.9
16 Google: Gemini 3 Flash Preview 💡 Google 82.8 based on 4 of 7 evaluations 71.6 $1.12 75 1.05M
17 OpenAI: GPT-5.2 💡 OpenAI 81.1 based on 5 of 7 evaluations 78.6 $4.81 58 400k
18 Anthropic: Claude Sonnet 4.6 💡 Anthropic 80.5 based on 4 of 7 evaluations 72.9 $6.00 42 1.00M
19 Qwen: Qwen3.6 Plus 💡 Alibaba 78.7 based on 3 of 7 evaluations $0.73 33 1.00M
20 OpenAI: GPT-5 💡 OpenAI 78.3 based on 6 of 7 evaluations 69.0 $3.44 71 400k
21 Z.ai: GLM 5 💡 Z.ai 77.3 based on 4 of 7 evaluations 69.3 $0.93 43 205k
22 DeepSeek: R1 💡 DeepSeek 74.8 based on 4 of 7 evaluations $1.15 20 64k
23 OpenAI: GPT-5.1 💡 OpenAI 74.3 based on 6 of 7 evaluations 69.0 $3.44 70 400k
24 claude-3-7-sonnet-20250219 Anthropic 73.9 based on 5 of 7 evaluations 71.0
25 gemini-2.0-pro-exp-02-05 Google 73.7 based on 3 of 7 evaluations
26 OpenAI: o3 💡 OpenAI 73.6 based on 5 of 7 evaluations $3.50 78 200k
27 Anthropic: Claude Opus 4.5 💡 Anthropic 72.3 based on 5 of 7 evaluations 80.2 $10.00 47 200k
28 OpenAI: o1 💡 OpenAI 69.7 based on 5 of 7 evaluations $26.25 53 200k
29 QwQ-32B Alibaba 68.9 based on 3 of 7 evaluations
30 Z.ai: GLM 4.7 💡 Z.ai 68.6 based on 3 of 7 evaluations $0.74 30 205k
31 claude-opus-4-20250514 Anthropic 68.3 based on 4 of 7 evaluations
32 grok-4-0709 xAI 67.7 based on 3 of 7 evaluations
33 OpenAI: GPT-5 Nano 💡 OpenAI 67.6 based on 4 of 7 evaluations $0.14 100 400k
34 OpenAI: GPT-5 Mini 💡 OpenAI 67.2 based on 6 of 7 evaluations 51.4 $0.69 79 400k
35 claude-haiku-4-5-20251001 Anthropic 67.2 based on 4 of 7 evaluations
36 claude-sonnet-4-5-20250929 Anthropic 66.2 based on 6 of 7 evaluations 62.6
37 Qwen: Qwen3.6 35B A3B 💡 Alibaba 65.7 based on 3 of 7 evaluations $0.30 74 262k
38 Google: Gemini 2.5 Pro 💡 Google 64.7 based on 4 of 7 evaluations 43.3 $3.44 99 1.05M
39 o1-mini OpenAI 64.5 based on 4 of 7 evaluations
40 DeepSeek: DeepSeek V3 0324 DeepSeek 63.9 based on 4 of 7 evaluations $0.44 21 164k
41 DeepSeek: R1 Distill Llama 70B 💡 DeepSeek 62.5 based on 4 of 7 evaluations $0.80 20 8k
42 gemini-2.0-flash-thinking-exp-01-21 Google DeepMind,Google 62.5 based on 3 of 7 evaluations
43 DeepSeek: DeepSeek V3 DeepSeek 62.3 based on 4 of 7 evaluations $0.45 17 164k
44 OpenAI: gpt-oss-120b 💡 OpenAI 61.5 based on 3 of 7 evaluations $0.07 132 131k
45 gemini-2.0-flash-001 Google DeepMind,Google 60.8 based on 4 of 7 evaluations
46 gpt-4.5-preview OpenAI 54.0 based on 5 of 7 evaluations
47 DeepSeek-R1-Distill-Qwen-32B DeepSeek 52.6 based on 3 of 7 evaluations
48 Anthropic: Claude Opus 4.1 💡 Anthropic 49.8 based on 5 of 7 evaluations 61.6 $30.00 7 200k
49 Google: Gemma 3 27B Google 47.0 based on 4 of 7 evaluations $0.17 21 131k
50 OpenAI: GPT-4.1 OpenAI 45.7 based on 5 of 7 evaluations $3.50 58 1.05M
51 grok-2-1212 xAI 45.0 based on 4 of 7 evaluations
52 claude-3-5-sonnet-20241022 Anthropic 44.9 based on 4 of 7 evaluations
53 Qwen: Qwen3.5-9B 💡 Alibaba 43.5 based on 3 of 7 evaluations $0.11 35 262k
54 OpenAI: gpt-oss-20b 💡 OpenAI 41.6 based on 3 of 7 evaluations $0.06 112 131k
55 Microsoft: Phi 4 Microsoft 41.3 based on 4 of 7 evaluations $0.09 57 16k
56 mistral-large-2411 Mistral 37.4 based on 4 of 7 evaluations
57 Meta: Llama 3.3 70B Instruct Meta 34.0 based on 4 of 7 evaluations $0.15 39 131k
58 mistral-small-2503 Mistral 33.0 based on 4 of 7 evaluations
59 claude-3-opus-20240229 Anthropic 32.4 based on 4 of 7 evaluations
60 mistral-small-2501 Mistral 31.2 based on 4 of 7 evaluations
61 OpenAI: GPT-4o-mini OpenAI 30.9 based on 4 of 7 evaluations $0.26 59 128k
62 claude-3-5-haiku-20241022 Anthropic 29.6 based on 4 of 7 evaluations
63 OpenAI: GPT-4o OpenAI 26.7 based on 6 of 7 evaluations 27.2 $4.38 42 128k
64 Google: Gemma 2 27B Google 21.4 based on 4 of 7 evaluations $0.65 38 8k
65 gemma-2-9b-it Google 12.8 based on 4 of 7 evaluations
💡 A lightbulb marks a model its provider publishes as a reasoning model
The Known Good
How to read this table. The Known Good Index is our own 0–100 composite of seven published evaluations; it is not any publisher's own scale. A model is scored only when it has at least one science, one mathematics and one coding result, so a model measured in a single area is absent rather than flattered. Blended price is USD per 1M tokens at a 3:1 input:output blend — we publish blended price rather than cost-per-task, because cost-per-task needs token counts from an evaluation suite we do not run. Full provenance is on the methodology and attribution pages.