Leaderboard
Every tracked model that holds a Known Good Index score, ranked. Sort any column; a figure we have not ingested shows as an em dash rather than a guess.
Known Good Index leaderboard
Known Good Index combines seven public evaluations — GPQA Diamond, Mock AIME 2024–25, MATH Level 5, SWE-bench Verified, LiveBench, Humanity's Last Exam, Terminal-Bench — each min-max normalised across tracked models, then averaged with equal weight and scaled 0–100. Scores ingested from Epoch AI (CC-BY) and LiveBench; price and speed from OpenRouter, metadata from models.dev. We do not run these evaluations.
| #▲ | Model | Creator | Known Good Index | Coding | Blended $/1M | Output tok/s | Context |
|---|---|---|---|---|---|---|---|
| 1 |
|
OpenAI | 97.5 based on 3 of 7 evaluations | — | — | — | — |
| 2 |
|
97.1 based on 4 of 7 evaluations | — | $4.50 | 101 | 1.05M | |
| 3 |
|
94.8 based on 3 of 7 evaluations | — | $3.38 | 162 | 1.05M | |
| 4 |
|
DeepSeek | 93.3 based on 3 of 7 evaluations | — | $1.30 | 42 | 1.05M |
| 5 |
|
OpenAI | 92.8 based on 3 of 7 evaluations | — | $11.25 | 102 | 1.05M |
| 6 |
|
Alibaba | 92.7 based on 3 of 7 evaluations | — | $2.21 | 64 | 1.00M |
| 7 |
|
Anthropic | 92.5 based on 5 of 7 evaluations | 97.2 | $10.00 | 50 | 1.00M |
| 8 |
|
Moonshot AI | 92.5 based on 3 of 7 evaluations | — | $1.71 | 52 | 262k |
| 9 |
|
OpenAI | 91.1 based on 5 of 7 evaluations | 91.9 | $5.62 | 80 | 1.05M |
| 10 |
|
Z.ai | 90.9 based on 3 of 7 evaluations | — | $1.48 | 62 | 1.05M |
| 11 |
|
Z.ai | 89.6 based on 3 of 7 evaluations | — | $1.48 | 39 | 205k |
| 12 |
|
Alibaba | 89.5 based on 3 of 7 evaluations | — | $2.31 | 46 | 262k |
| 13 |
|
Anthropic | 88.9 based on 5 of 7 evaluations | 92.5 | $10.00 | 35 | 1.00M |
| 14 |
|
OpenAI | 85.7 based on 4 of 7 evaluations | — | $1.93 | 126 | 200k |
| 15 |
|
83.8 based on 5 of 7 evaluations | 75.9 | — | — | — | |
| 16 |
|
82.8 based on 4 of 7 evaluations | 71.6 | $1.12 | 86 | 1.05M | |
| 17 |
|
OpenAI | 81.1 based on 5 of 7 evaluations | 78.6 | $4.81 | 50 | 400k |
| 18 |
|
Anthropic | 80.5 based on 4 of 7 evaluations | 72.9 | $6.00 | 40 | 1.00M |
| 19 |
|
Alibaba | 78.7 based on 3 of 7 evaluations | — | $0.73 | 35 | 1.00M |
| 20 |
|
OpenAI | 78.3 based on 6 of 7 evaluations | 69.0 | $3.44 | 66 | 400k |
| 21 |
|
Z.ai | 77.3 based on 4 of 7 evaluations | 69.3 | $0.93 | 45 | 205k |
| 22 |
|
DeepSeek | 74.8 based on 4 of 7 evaluations | — | $1.15 | 20 | 64k |
| 23 |
|
OpenAI | 74.3 based on 6 of 7 evaluations | 69.0 | $3.44 | 42 | 400k |
| 24 |
|
Anthropic | 73.9 based on 5 of 7 evaluations | 71.0 | — | — | — |
| 25 |
|
73.7 based on 3 of 7 evaluations | — | — | — | — | |
| 26 |
|
OpenAI | 73.6 based on 5 of 7 evaluations | — | $3.50 | 80 | 200k |
| 27 |
|
Anthropic | 72.3 based on 5 of 7 evaluations | 80.2 | $10.00 | 48 | 200k |
| 28 |
|
OpenAI | 69.7 based on 5 of 7 evaluations | — | $26.25 | 53 | 200k |
| 29 |
|
Alibaba | 68.9 based on 3 of 7 evaluations | — | — | — | — |
| 30 |
|
Z.ai | 68.6 based on 3 of 7 evaluations | — | $0.74 | 33 | 205k |
| 31 |
|
Anthropic | 68.3 based on 4 of 7 evaluations | — | — | — | — |
| 32 |
|
xAI | 67.7 based on 3 of 7 evaluations | — | — | — | — |
| 33 |
|
OpenAI | 67.6 based on 4 of 7 evaluations | — | $0.14 | 93 | 400k |
| 34 |
|
OpenAI | 67.2 based on 6 of 7 evaluations | 51.4 | $0.69 | 78 | 400k |
| 35 |
|
Anthropic | 67.2 based on 4 of 7 evaluations | — | — | — | — |
| 36 |
|
Anthropic | 66.2 based on 6 of 7 evaluations | 62.6 | — | — | — |
| 37 |
|
Alibaba | 65.7 based on 3 of 7 evaluations | — | $0.30 | 74 | 262k |
| 38 |
|
64.7 based on 4 of 7 evaluations | 43.3 | $3.44 | 102 | 1.05M | |
| 39 |
|
OpenAI | 64.5 based on 4 of 7 evaluations | — | — | — | — |
| 40 |
|
DeepSeek | 63.9 based on 4 of 7 evaluations | — | $0.44 | 24 | 164k |
| 41 |
|
DeepSeek | 62.5 based on 4 of 7 evaluations | — | $0.80 | 21 | 8k |
| 42 | GD gemini-2.0-flash-thinking-exp-01-21 | Google DeepMind,Google | 62.5 based on 3 of 7 evaluations | — | — | — | — |
| 43 |
|
DeepSeek | 62.3 based on 4 of 7 evaluations | — | $0.46 | 22 | 164k |
| 44 |
|
OpenAI | 61.5 based on 3 of 7 evaluations | — | $0.07 | 167 | 131k |
| 45 | GD gemini-2.0-flash-001 | Google DeepMind,Google | 60.8 based on 4 of 7 evaluations | — | — | — | — |
| 46 |
|
OpenAI | 54.0 based on 5 of 7 evaluations | — | — | — | — |
| 47 |
|
DeepSeek | 52.6 based on 3 of 7 evaluations | — | — | — | — |
| 48 |
|
Anthropic | 49.8 based on 5 of 7 evaluations | 61.6 | $30.00 | 8 | 200k |
| 49 |
|
47.0 based on 4 of 7 evaluations | — | $0.17 | 19 | 131k | |
| 50 |
|
OpenAI | 45.7 based on 5 of 7 evaluations | — | $3.50 | 57 | 1.05M |
| 51 |
|
xAI | 45.0 based on 4 of 7 evaluations | — | — | — | — |
| 52 |
|
Anthropic | 44.9 based on 4 of 7 evaluations | — | — | — | — |
| 53 |
|
Alibaba | 43.5 based on 3 of 7 evaluations | — | $0.11 | 32 | 262k |
| 54 |
|
OpenAI | 41.6 based on 3 of 7 evaluations | — | $0.06 | 118 | 131k |
| 55 |
|
Microsoft | 41.3 based on 4 of 7 evaluations | — | $0.09 | 54 | 16k |
| 56 |
|
Mistral | 37.4 based on 4 of 7 evaluations | — | — | — | — |
| 57 |
|
Meta | 34.0 based on 4 of 7 evaluations | — | $0.15 | 44 | 131k |
| 58 |
|
Mistral | 33.0 based on 4 of 7 evaluations | — | — | — | — |
| 59 |
|
Anthropic | 32.4 based on 4 of 7 evaluations | — | — | — | — |
| 60 |
|
Mistral | 31.2 based on 4 of 7 evaluations | — | — | — | — |
| 61 |
|
OpenAI | 30.9 based on 4 of 7 evaluations | — | $0.26 | 45 | 128k |
| 62 |
|
Anthropic | 29.6 based on 4 of 7 evaluations | — | — | — | — |
| 63 |
|
OpenAI | 26.7 based on 6 of 7 evaluations | 27.2 | $4.38 | 43 | 128k |
| 64 |
|
21.4 based on 4 of 7 evaluations | — | $0.65 | 30 | 8k | |
| 65 |
|
12.8 based on 4 of 7 evaluations | — | — | — | — |
💡 A lightbulb marks a model its provider publishes as a reasoning model
The Known Good
How to read this table. The Known Good Index is our own 0–100 composite of seven published evaluations; it is not any publisher's own scale. A model is scored only when it has at least one science, one mathematics and one coding result, so a model measured in a single area is absent rather than flattered. Blended price is USD per 1M tokens at a 3:1 input:output blend — we publish blended price rather than cost-per-task, because cost-per-task needs token counts from an evaluation suite we do not run. Full provenance is on the methodology and attribution pages.