Leaderboard
Every tracked model that holds a Known Good Index score, ranked. Sort any column; a figure we have not ingested shows as an em dash rather than a guess.
Known Good Index leaderboard
Known Good Index combines six public evaluations — GPQA Diamond, MATH / AIME, SWE-bench Verified, LiveBench, Humanity's Last Exam, Terminal-Bench — each min-max normalised across tracked models, then averaged with equal weight and scaled 0–100. Scores ingested from Epoch AI (CC-BY) and LiveBench; price and speed from OpenRouter, metadata from models.dev. We do not run these evaluations.
| #▲ | Model | Creator | Known Good Index | Coding | Blended $/1M | Output tok/s | Context |
|---|---|---|---|---|---|---|---|
| 1 |
|
OpenAI | 97.9 based on 3 of 6 evaluations | — | — | — | — |
| 2 |
|
95.9 based on 4 of 6 evaluations | — | $4.50 | 104 | 1.05M | |
| 3 |
|
95.2 based on 3 of 6 evaluations | — | $3.38 | 119 | 1.05M | |
| 4 |
|
Anthropic | 93.8 based on 5 of 6 evaluations | 100.0 | $10.00 | 52 | 1.00M |
| 5 |
|
Alibaba | 93.2 based on 3 of 6 evaluations | — | $2.21 | 40 | 1.00M |
| 6 |
|
DeepSeek | 93.2 based on 3 of 6 evaluations | — | $0.54 | 45 | 1.05M |
| 7 |
|
Moonshot AI | 92.8 based on 3 of 6 evaluations | — | $1.16 | 66 | 262k |
| 8 |
|
Z.ai | 91.3 based on 3 of 6 evaluations | — | $1.07 | 55 | 1.05M |
| 9 |
|
Alibaba | 90.4 based on 3 of 6 evaluations | — | $2.34 | 46 | 262k |
| 10 |
|
OpenAI | 89.6 based on 5 of 6 evaluations | 88.9 | $5.62 | 82 | 1.05M |
| 11 |
|
Anthropic | 87.9 based on 5 of 6 evaluations | 89.5 | $10.00 | 37 | 1.00M |
| 12 |
|
Z.ai | 87.8 based on 3 of 6 evaluations | — | $1.48 | 53 | 205k |
| 13 |
|
83.4 based on 4 of 6 evaluations | 77.4 | $1.12 | 89 | 1.05M | |
| 14 |
|
83.2 based on 5 of 6 evaluations | 73.6 | — | — | — | |
| 15 |
|
OpenAI | 81.5 based on 3 of 6 evaluations | — | $1.93 | 132 | 200k |
| 16 |
|
OpenAI | 80.4 based on 5 of 6 evaluations | 76.2 | $4.81 | 54 | 400k |
| 17 |
|
Anthropic | 79.7 based on 4 of 6 evaluations | 70.9 | $6.00 | 41 | 1.00M |
| 18 |
|
Alibaba | 77.6 based on 3 of 6 evaluations | — | $0.73 | 47 | 1.00M |
| 19 |
|
Z.ai | 76.6 based on 4 of 6 evaluations | 67.4 | $1.35 | 37 | 205k |
| 20 |
|
OpenAI | 73.9 based on 6 of 6 evaluations | 67.9 | $3.44 | 62 | 400k |
| 21 |
|
OpenAI | 73.4 based on 5 of 6 evaluations | 67.2 | $3.44 | 57 | 400k |
| 22 |
|
Anthropic | 71.5 based on 5 of 6 evaluations | 77.9 | — | — | — |
| 23 |
|
Anthropic | 69.5 based on 4 of 6 evaluations | 71.0 | — | — | — |
| 24 |
|
DeepSeek | 68.1 based on 3 of 6 evaluations | — | $1.15 | 49 | 164k |
| 25 |
|
DeepSeek | 68.1 based on 3 of 6 evaluations | — | — | — | — |
| 26 |
|
Z.ai | 68.0 based on 3 of 6 evaluations | — | $0.74 | 66 | 205k |
| 27 |
|
xAI | 67.4 based on 3 of 6 evaluations | — | — | — | — |
| 28 |
|
OpenAI | 67.0 based on 4 of 6 evaluations | — | $3.50 | 88 | 200k |
| 29 |
|
64.2 based on 4 of 6 evaluations | 42.1 | $3.44 | 96 | 1.05M | |
| 30 |
|
OpenAI | 63.1 based on 4 of 6 evaluations | — | $26.25 | 24 | 200k |
| 31 |
|
Anthropic | 62.2 based on 3 of 6 evaluations | — | — | — | — |
| 32 | GD gemini-2.0-flash-thinking-exp-01-21 | Google DeepMind,Google | 62.0 based on 3 of 6 evaluations | — | — | — | — |
| 33 |
|
OpenAI | 61.1 based on 3 of 6 evaluations | — | $0.07 | 172 | 131k |
| 34 |
|
OpenAI | 60.2 based on 5 of 6 evaluations | 50.2 | $0.69 | 100 | 400k |
| 35 |
|
DeepSeek | 59.6 based on 3 of 6 evaluations | — | $0.48 | 26 | 164k |
| 36 |
|
Anthropic | 59.0 based on 5 of 6 evaluations | 61.1 | — | — | — |
| 37 |
|
OpenAI | 57.1 based on 3 of 6 evaluations | — | $0.14 | 121 | 400k |
| 38 |
|
Anthropic | 56.1 based on 3 of 6 evaluations | — | — | — | — |
| 39 |
|
OpenAI | 55.5 based on 3 of 6 evaluations | — | — | — | — |
| 40 | GD gemini-2.0-flash-001 | Google DeepMind,Google | 53.0 based on 3 of 6 evaluations | — | — | — | — |
| 41 |
|
DeepSeek | 52.5 based on 3 of 6 evaluations | — | $0.80 | 24 | 8k |
| 42 |
|
Anthropic | 49.2 based on 5 of 6 evaluations | 60.3 | — | — | — |
| 43 |
|
OpenAI | 47.5 based on 4 of 6 evaluations | — | — | — | — |
| 44 |
|
DeepSeek | 44.2 based on 3 of 6 evaluations | — | $0.35 | 21 | 164k |
| 45 |
|
Anthropic | 40.5 based on 3 of 6 evaluations | — | — | — | — |
| 46 |
|
xAI | 38.3 based on 3 of 6 evaluations | — | — | — | — |
| 47 |
|
36.6 based on 3 of 6 evaluations | — | $0.17 | 24 | 262k | |
| 48 |
|
36.6 based on 3 of 6 evaluations | — | — | — | — | |
| 49 |
|
OpenAI | 36.0 based on 4 of 6 evaluations | — | $3.50 | 33 | 1.05M |
| 50 |
|
Microsoft | 32.9 based on 3 of 6 evaluations | — | $0.09 | 56 | 16k |
| 51 |
|
Mistral | 32.8 based on 3 of 6 evaluations | — | — | — | — |
| 52 |
|
Meta | 31.2 based on 3 of 6 evaluations | — | $0.20 | 54 | 131k |
| 53 |
|
Anthropic | 30.4 based on 3 of 6 evaluations | — | — | — | — |
| 54 |
|
Mistral | 28.1 based on 3 of 6 evaluations | — | — | — | — |
| 55 |
|
Mistral | 26.2 based on 3 of 6 evaluations | — | — | — | — |
| 56 |
|
Anthropic | 23.4 based on 3 of 6 evaluations | — | — | — | — |
| 57 |
|
OpenAI | 22.9 based on 3 of 6 evaluations | — | $0.26 | 37 | 128k |
| 58 |
|
OpenAI | 21.1 based on 5 of 6 evaluations | 27.2 | $4.38 | 41 | 128k |
| 59 |
|
18.9 based on 3 of 6 evaluations | — | $0.65 | 14 | 8k | |
| 60 |
|
18.9 based on 3 of 6 evaluations | — | — | — | — | |
| 61 |
|
9.6 based on 3 of 6 evaluations | — | — | — | — |
💡 Reasoning models are indicated by a lightbulb
The Known Good
How to read this table. The Known Good Index is our own 0–100 composite of six published evaluations; it is not any publisher's own scale. A model is scored only when it has at least one science, one mathematics and one coding result, so a model measured in a single area is absent rather than flattered. Blended price is USD per 1M tokens at a 3:1 input:output blend — we publish blended price rather than cost-per-task, because cost-per-task needs token counts from an evaluation suite we do not run. Full provenance is on the methodology and attribution pages.