The Known Good Updated 25 Jul 2026

Leaderboard

Every tracked model that holds a Known Good Index score, ranked. Sort any column; a figure we have not ingested shows as an em dash rather than a guess.

61 ranked of 908 tracked models All models → Index methodology → Download as CSV →
Known Good Index leaderboard

Known Good Index combines six public evaluations — GPQA Diamond, MATH / AIME, SWE-bench Verified, LiveBench, Humanity's Last Exam, Terminal-Bench — each min-max normalised across tracked models, then averaged with equal weight and scaled 0–100. Scores ingested from Epoch AI (CC-BY) and LiveBench; price and speed from OpenRouter, metadata from models.dev. We do not run these evaluations.

# Model Creator Known Good Index Coding Blended $/1M Output tok/s Context
1 gpt-5.5-pre-release OpenAI 97.9 based on 3 of 6 evaluations
2 Google: Gemini 3.1 Pro Preview 💡 Google 95.9 based on 4 of 6 evaluations $4.50 104 1.05M
3 Google: Gemini 3.5 Flash 💡 Google 95.2 based on 3 of 6 evaluations $3.38 119 1.05M
4 Anthropic: Claude Opus 4.7 💡 Anthropic 93.8 based on 5 of 6 evaluations 100.0 $10.00 52 1.00M
5 Qwen: Qwen3.7 Max 💡 Alibaba 93.2 based on 3 of 6 evaluations $2.21 40 1.00M
6 DeepSeek: DeepSeek V4 Pro 💡 DeepSeek 93.2 based on 3 of 6 evaluations $0.54 45 1.05M
7 MoonshotAI: Kimi K2.6 💡 Moonshot AI 92.8 based on 3 of 6 evaluations $1.16 66 262k
8 Z.ai: GLM 5.2 💡 Z.ai 91.3 based on 3 of 6 evaluations $1.07 55 1.05M
9 Qwen: Qwen3.6 Max Preview 💡 Alibaba 90.4 based on 3 of 6 evaluations $2.34 46 262k
10 OpenAI: GPT-5.4 💡 OpenAI 89.6 based on 5 of 6 evaluations 88.9 $5.62 82 1.05M
11 Anthropic: Claude Opus 4.6 💡 Anthropic 87.9 based on 5 of 6 evaluations 89.5 $10.00 37 1.00M
12 Z.ai: GLM 5.1 💡 Z.ai 87.8 based on 3 of 6 evaluations $1.48 53 205k
13 Google: Gemini 3 Flash Preview 💡 Google 83.4 based on 4 of 6 evaluations 77.4 $1.12 89 1.05M
14 gemini-3-pro-preview Google 83.2 based on 5 of 6 evaluations 73.6
15 OpenAI: o3 Mini 💡 OpenAI 81.5 based on 3 of 6 evaluations $1.93 132 200k
16 OpenAI: GPT-5.2 💡 OpenAI 80.4 based on 5 of 6 evaluations 76.2 $4.81 54 400k
17 Anthropic: Claude Sonnet 4.6 💡 Anthropic 79.7 based on 4 of 6 evaluations 70.9 $6.00 41 1.00M
18 Qwen: Qwen3.6 Plus 💡 Alibaba 77.6 based on 3 of 6 evaluations $0.73 47 1.00M
19 Z.ai: GLM 5 💡 Z.ai 76.6 based on 4 of 6 evaluations 67.4 $1.35 37 205k
20 OpenAI: GPT-5.1 💡 OpenAI 73.9 based on 6 of 6 evaluations 67.9 $3.44 62 400k
21 OpenAI: GPT-5 💡 OpenAI 73.4 based on 5 of 6 evaluations 67.2 $3.44 57 400k
22 claude-opus-4-5-20251101 Anthropic 71.5 based on 5 of 6 evaluations 77.9
23 claude-3-7-sonnet-20250219 Anthropic 69.5 based on 4 of 6 evaluations 71.0
24 DeepSeek: R1 💡 DeepSeek 68.1 based on 3 of 6 evaluations $1.15 49 164k
25 deepseek-r1 DeepSeek 68.1 based on 3 of 6 evaluations
26 Z.ai: GLM 4.7 💡 Z.ai 68.0 based on 3 of 6 evaluations $0.74 66 205k
27 grok-4-0709 xAI 67.4 based on 3 of 6 evaluations
28 OpenAI: o3 💡 OpenAI 67.0 based on 4 of 6 evaluations $3.50 88 200k
29 Google: Gemini 2.5 Pro 💡 Google 64.2 based on 4 of 6 evaluations 42.1 $3.44 96 1.05M
30 OpenAI: o1 💡 OpenAI 63.1 based on 4 of 6 evaluations $26.25 24 200k
31 claude-opus-4-20250514 Anthropic 62.2 based on 3 of 6 evaluations
32 gemini-2.0-flash-thinking-exp-01-21 Google DeepMind,Google 62.0 based on 3 of 6 evaluations
33 OpenAI: gpt-oss-120b 💡 OpenAI 61.1 based on 3 of 6 evaluations $0.07 172 131k
34 OpenAI: GPT-5 Mini 💡 OpenAI 60.2 based on 5 of 6 evaluations 50.2 $0.69 100 400k
35 DeepSeek: DeepSeek V3 0324 DeepSeek 59.6 based on 3 of 6 evaluations $0.48 26 164k
36 claude-sonnet-4-5-20250929 Anthropic 59.0 based on 5 of 6 evaluations 61.1
37 OpenAI: GPT-5 Nano 💡 OpenAI 57.1 based on 3 of 6 evaluations $0.14 121 400k
38 claude-haiku-4-5-20251001 Anthropic 56.1 based on 3 of 6 evaluations
39 o1-mini OpenAI 55.5 based on 3 of 6 evaluations
40 gemini-2.0-flash-001 Google DeepMind,Google 53.0 based on 3 of 6 evaluations
41 DeepSeek: R1 Distill Llama 70B 💡 DeepSeek 52.5 based on 3 of 6 evaluations $0.80 24 8k
42 claude-opus-4-1-20250805 Anthropic 49.2 based on 5 of 6 evaluations 60.3
43 gpt-4.5-preview OpenAI 47.5 based on 4 of 6 evaluations
44 DeepSeek: DeepSeek V3 DeepSeek 44.2 based on 3 of 6 evaluations $0.35 21 164k
45 claude-3-5-sonnet-20241022 Anthropic 40.5 based on 3 of 6 evaluations
46 grok-2-1212 xAI 38.3 based on 3 of 6 evaluations
47 Google: Gemma 3 27B Google 36.6 based on 3 of 6 evaluations $0.17 24 262k
48 gemma-3-27b-it Google 36.6 based on 3 of 6 evaluations
49 OpenAI: GPT-4.1 OpenAI 36.0 based on 4 of 6 evaluations $3.50 33 1.05M
50 Microsoft: Phi 4 Microsoft 32.9 based on 3 of 6 evaluations $0.09 56 16k
51 mistral-large-2411 Mistral 32.8 based on 3 of 6 evaluations
52 Meta: Llama 3.3 70B Instruct Meta 31.2 based on 3 of 6 evaluations $0.20 54 131k
53 claude-3-opus-20240229 Anthropic 30.4 based on 3 of 6 evaluations
54 mistral-small-2503 Mistral 28.1 based on 3 of 6 evaluations
55 mistral-small-2501 Mistral 26.2 based on 3 of 6 evaluations
56 claude-3-5-haiku-20241022 Anthropic 23.4 based on 3 of 6 evaluations
57 OpenAI: GPT-4o-mini OpenAI 22.9 based on 3 of 6 evaluations $0.26 37 128k
58 OpenAI: GPT-4o OpenAI 21.1 based on 5 of 6 evaluations 27.2 $4.38 41 128k
59 Google: Gemma 2 27B Google 18.9 based on 3 of 6 evaluations $0.65 14 8k
60 gemma-2-27b-it Google 18.9 based on 3 of 6 evaluations
61 gemma-2-9b-it Google 9.6 based on 3 of 6 evaluations
💡 Reasoning models are indicated by a lightbulb
The Known Good
How to read this table. The Known Good Index is our own 0–100 composite of six published evaluations; it is not any publisher's own scale. A model is scored only when it has at least one science, one mathematics and one coding result, so a model measured in a single area is absent rather than flattered. Blended price is USD per 1M tokens at a 3:1 input:output blend — we publish blended price rather than cost-per-task, because cost-per-task needs token counts from an evaluation suite we do not run. Full provenance is on the methodology and attribution pages.