The Known Good Updated 4 Sep 2026 Subscribe

Leaderboard

Every tracked model that holds a Known Good Index score, ranked. Sort any column; a figure we have not ingested shows as an em dash rather than a guess.

65 ranked of 1094 tracked models All models → Index methodology → Download as CSV →
Known Good Index leaderboard

Known Good Index combines seven public evaluations — GPQA Diamond, Mock AIME 2024–25, MATH Level 5, SWE-bench Verified, LiveBench, Humanity's Last Exam, Terminal-Bench — each min-max normalised across tracked models, then averaged with equal weight and scaled 0–100. Scores ingested from Epoch AI (CC-BY) and LiveBench; price and speed from OpenRouter, metadata from models.dev. We do not run these evaluations.

# Model Creator Known Good Index Coding Blended $/1M Output tok/s Context
1 gpt-5.5-pre-release OpenAI 97.5 based on 3 of 7 evaluations
2 Google: Gemini 3.1 Pro Preview 💡 Google 97.1 based on 4 of 7 evaluations $4.50 101 1.05M
3 Google: Gemini 3.5 Flash 💡 Google 94.8 based on 3 of 7 evaluations $3.38 162 1.05M
4 DeepSeek: DeepSeek V4 Pro 0423 💡 DeepSeek 93.3 based on 3 of 7 evaluations $1.30 42 1.05M
5 OpenAI: GPT-5.5 💡 OpenAI 92.8 based on 3 of 7 evaluations $11.25 102 1.05M
6 Qwen: Qwen3.7 Max 💡 Alibaba 92.7 based on 3 of 7 evaluations $2.21 64 1.00M
7 Anthropic: Claude Opus 4.7 💡 Anthropic 92.5 based on 5 of 7 evaluations 97.2 $10.00 50 1.00M
8 MoonshotAI: Kimi K2.6 💡 Moonshot AI 92.5 based on 3 of 7 evaluations $1.71 52 262k
9 OpenAI: GPT-5.4 💡 OpenAI 91.1 based on 5 of 7 evaluations 91.9 $5.62 80 1.05M
10 Z.ai: GLM 5.2 💡 Z.ai 90.9 based on 3 of 7 evaluations $1.48 62 1.05M
11 Z.ai: GLM 5.1 💡 Z.ai 89.6 based on 3 of 7 evaluations $1.48 39 205k
12 Qwen: Qwen3.6 Max Preview 💡 Alibaba 89.5 based on 3 of 7 evaluations $2.31 46 262k
13 Anthropic: Claude Opus 4.6 💡 Anthropic 88.9 based on 5 of 7 evaluations 92.5 $10.00 35 1.00M
14 OpenAI: o3 Mini 💡 OpenAI 85.7 based on 4 of 7 evaluations $1.93 126 200k
15 gemini-3-pro-preview Google 83.8 based on 5 of 7 evaluations 75.9
16 Google: Gemini 3 Flash Preview 💡 Google 82.8 based on 4 of 7 evaluations 71.6 $1.12 86 1.05M
17 OpenAI: GPT-5.2 💡 OpenAI 81.1 based on 5 of 7 evaluations 78.6 $4.81 50 400k
18 Anthropic: Claude Sonnet 4.6 💡 Anthropic 80.5 based on 4 of 7 evaluations 72.9 $6.00 40 1.00M
19 Qwen: Qwen3.6 Plus 💡 Alibaba 78.7 based on 3 of 7 evaluations $0.73 35 1.00M
20 OpenAI: GPT-5 💡 OpenAI 78.3 based on 6 of 7 evaluations 69.0 $3.44 66 400k
21 Z.ai: GLM 5 💡 Z.ai 77.3 based on 4 of 7 evaluations 69.3 $0.93 45 205k
22 DeepSeek: R1 💡 DeepSeek 74.8 based on 4 of 7 evaluations $1.15 20 64k
23 OpenAI: GPT-5.1 💡 OpenAI 74.3 based on 6 of 7 evaluations 69.0 $3.44 42 400k
24 claude-3-7-sonnet-20250219 Anthropic 73.9 based on 5 of 7 evaluations 71.0
25 gemini-2.0-pro-exp-02-05 Google 73.7 based on 3 of 7 evaluations
26 OpenAI: o3 💡 OpenAI 73.6 based on 5 of 7 evaluations $3.50 80 200k
27 Anthropic: Claude Opus 4.5 💡 Anthropic 72.3 based on 5 of 7 evaluations 80.2 $10.00 48 200k
28 OpenAI: o1 💡 OpenAI 69.7 based on 5 of 7 evaluations $26.25 53 200k
29 QwQ-32B Alibaba 68.9 based on 3 of 7 evaluations
30 Z.ai: GLM 4.7 💡 Z.ai 68.6 based on 3 of 7 evaluations $0.74 33 205k
31 claude-opus-4-20250514 Anthropic 68.3 based on 4 of 7 evaluations
32 grok-4-0709 xAI 67.7 based on 3 of 7 evaluations
33 OpenAI: GPT-5 Nano 💡 OpenAI 67.6 based on 4 of 7 evaluations $0.14 93 400k
34 OpenAI: GPT-5 Mini 💡 OpenAI 67.2 based on 6 of 7 evaluations 51.4 $0.69 78 400k
35 claude-haiku-4-5-20251001 Anthropic 67.2 based on 4 of 7 evaluations
36 claude-sonnet-4-5-20250929 Anthropic 66.2 based on 6 of 7 evaluations 62.6
37 Qwen: Qwen3.6 35B A3B 💡 Alibaba 65.7 based on 3 of 7 evaluations $0.30 74 262k
38 Google: Gemini 2.5 Pro 💡 Google 64.7 based on 4 of 7 evaluations 43.3 $3.44 102 1.05M
39 o1-mini OpenAI 64.5 based on 4 of 7 evaluations
40 DeepSeek: DeepSeek V3 0324 DeepSeek 63.9 based on 4 of 7 evaluations $0.44 24 164k
41 DeepSeek: R1 Distill Llama 70B 💡 DeepSeek 62.5 based on 4 of 7 evaluations $0.80 21 8k
42 gemini-2.0-flash-thinking-exp-01-21 Google DeepMind,Google 62.5 based on 3 of 7 evaluations
43 DeepSeek: DeepSeek V3 DeepSeek 62.3 based on 4 of 7 evaluations $0.46 22 164k
44 OpenAI: gpt-oss-120b 💡 OpenAI 61.5 based on 3 of 7 evaluations $0.07 167 131k
45 gemini-2.0-flash-001 Google DeepMind,Google 60.8 based on 4 of 7 evaluations
46 gpt-4.5-preview OpenAI 54.0 based on 5 of 7 evaluations
47 DeepSeek-R1-Distill-Qwen-32B DeepSeek 52.6 based on 3 of 7 evaluations
48 Anthropic: Claude Opus 4.1 💡 Anthropic 49.8 based on 5 of 7 evaluations 61.6 $30.00 8 200k
49 Google: Gemma 3 27B 💡 Google 47.0 based on 4 of 7 evaluations $0.17 19 131k
50 OpenAI: GPT-4.1 OpenAI 45.7 based on 5 of 7 evaluations $3.50 57 1.05M
51 grok-2-1212 xAI 45.0 based on 4 of 7 evaluations
52 claude-3-5-sonnet-20241022 Anthropic 44.9 based on 4 of 7 evaluations
53 Qwen: Qwen3.5-9B 💡 Alibaba 43.5 based on 3 of 7 evaluations $0.11 32 262k
54 OpenAI: gpt-oss-20b 💡 OpenAI 41.6 based on 3 of 7 evaluations $0.06 118 131k
55 Microsoft: Phi 4 Microsoft 41.3 based on 4 of 7 evaluations $0.09 54 16k
56 mistral-large-2411 Mistral 37.4 based on 4 of 7 evaluations
57 Meta: Llama 3.3 70B Instruct Meta 34.0 based on 4 of 7 evaluations $0.15 44 131k
58 mistral-small-2503 Mistral 33.0 based on 4 of 7 evaluations
59 claude-3-opus-20240229 Anthropic 32.4 based on 4 of 7 evaluations
60 mistral-small-2501 Mistral 31.2 based on 4 of 7 evaluations
61 OpenAI: GPT-4o-mini OpenAI 30.9 based on 4 of 7 evaluations $0.26 45 128k
62 claude-3-5-haiku-20241022 Anthropic 29.6 based on 4 of 7 evaluations
63 OpenAI: GPT-4o OpenAI 26.7 based on 6 of 7 evaluations 27.2 $4.38 43 128k
64 Google: Gemma 2 27B Google 21.4 based on 4 of 7 evaluations $0.65 30 8k
65 gemma-2-9b-it Google 12.8 based on 4 of 7 evaluations
💡 A lightbulb marks a model its provider publishes as a reasoning model
The Known Good
How to read this table. The Known Good Index is our own 0–100 composite of seven published evaluations; it is not any publisher's own scale. A model is scored only when it has at least one science, one mathematics and one coding result, so a model measured in a single area is absent rather than flattered. Blended price is USD per 1M tokens at a 3:1 input:output blend — we publish blended price rather than cost-per-task, because cost-per-task needs token counts from an evaluation suite we do not run. Full provenance is on the methodology and attribution pages.