The Known Good Updated 25 Jul 2026
Evaluations· Epoch AI· CC-BY 4.0· Known Good Index component

Humanity's Last Exam

Expert-written questions spanning many specialist fields. Aggregated from Epoch AI's AI Benchmarking Hub (hle_external.csv, column 'Accuracy'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations.

Publisher's page Licence & attribution How we use this score
45 models scored Higher is better Published by Epoch AI CC-BY 4.0
Humanity's Last Exam

Humanity's Last Exam as measured and published by Epoch AI (CC-BY 4.0). Ingested, normalised only where it feeds an index, and never re-run by us. Higher is better.

16 of 45 models +Add model from specific provider
💡 Reasoning models are indicated by a lightbulb
The Known Good
Every model we hold a Humanity's Last Exam score for
45 of 45 models
# Model Creator Humanity's Last Exam Known Good Index Measured
1 Google: Gemini 3.1 Pro Preview 💡 Google 46.4 96 25 Jul 2026
2 OpenAI: GPT-5.4 Pro 💡 OpenAI 44.3 25 Jul 2026
3 muse-spark Meta 40.6 25 Jul 2026
4 gemini-3-pro-preview Google 37.5 83 25 Jul 2026
5 OpenAI: GPT-5.4 💡 OpenAI 36.2 90 25 Jul 2026
6 Anthropic: Claude Opus 4.7 💡 Anthropic 36.2 94 25 Jul 2026
7 Anthropic: Claude Opus 4.6 💡 Anthropic 34.4 88 25 Jul 2026
8 OpenAI: GPT-5 Pro 💡 OpenAI 31.6 25 Jul 2026
9 OpenAI: GPT-5.2 💡 OpenAI 27.8 80 25 Jul 2026
10 OpenAI: GPT-5 💡 OpenAI 25.3 73 25 Jul 2026
11 claude-opus-4-5-20251101-thinking Anthropic 25.2 25 Jul 2026
12 MoonshotAI: Kimi K2.5 💡 Moonshot AI 24.4 25 Jul 2026
13 OpenAI: GPT-5.1 💡 OpenAI 23.7 74 25 Jul 2026
14 Google: Gemini 2.5 Pro Preview 06-05 💡 Google 21.6 25 Jul 2026
15 OpenAI: o3 💡 OpenAI 20.3 67 25 Jul 2026
16 OpenAI: GPT-5 Mini 💡 OpenAI 19.4 60 25 Jul 2026
17 Gemini 2.5 Pro Experimental (March 2025) Google 18.2 25 Jul 2026
18 OpenAI: o4 Mini 💡 OpenAI 18.1 25 Jul 2026
19 Google: Gemini 2.5 Pro Preview 05-06 💡 Google 17.8 25 Jul 2026
20 claude-opus-4-5-20251101 Anthropic 14.2 72 25 Jul 2026
21 claude-sonnet-4-5-20250929-thinking Anthropic 13.7 25 Jul 2026
22 Google: Gemini 2.5 Flash 💡 Google 12.1 25 Jul 2026
23 claude-opus-4-1-20250805-thinking Anthropic 11.5 25 Jul 2026
24 Gemini 2.5 Flash Preview (May 2025) Google 11.0 25 Jul 2026
25 Anthropic: Claude Opus 4 💡 Anthropic 10.7 25 Jul 2026
26 Google: Gemini 3.1 Flash Lite Preview 💡 Google 8.6 25 Jul 2026
27 Z.ai: GLM 4.5 💡 Z.ai 8.3 25 Jul 2026
28 Z.ai: GLM 4.5 Air 💡 Z.ai 8.1 25 Jul 2026
29 OpenAI: o1-pro 💡 OpenAI 8.1 25 Jul 2026
30 Claude 3.7 Sonnet (Thinking) Anthropic 8.0 25 Jul 2026
31 OpenAI: o1 💡 OpenAI 8.0 63 25 Jul 2026
32 claude-opus-4-1-20250805 Anthropic 7.9 49 25 Jul 2026
33 Anthropic: Claude Sonnet 4 💡 Anthropic 7.8 25 Jul 2026
34 claude-sonnet-4-5-20250929 Anthropic 7.5 59 25 Jul 2026
35 gpt-5.1-instant OpenAI 6.8 25 Jul 2026
36 Gemini 2.0 Flash Thinking (January 2025) Google DeepMind,Google 6.6 25 Jul 2026
37 Meta: Llama 4 Maverick Meta 5.7 25 Jul 2026
38 gpt-4.5-preview OpenAI 5.4 47 25 Jul 2026
39 OpenAI: GPT-4.1 OpenAI 5.4 36 25 Jul 2026
40 gemini-1.5-pro-002 Google 4.6 25 Jul 2026
41 Mistral: Mistral Medium 3 Mistral 4.5 25 Jul 2026
42 Nova Pro Amazon 4.4 25 Jul 2026
43 Claude 3.5 Sonnet (October 2024) Anthropic 4.1 25 Jul 2026
44 Nova Lite Amazon 3.6 25 Jul 2026
45 OpenAI: GPT-4o OpenAI 2.7 21 25 Jul 2026
The Known Good
Provenance. These figures are published by Epoch AI and redistributed here under CC-BY 4.0. The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is min-max normalised against every other model holding the same evaluation, and nothing else is done to it. This evaluation is one of the six components of the Known Good Index. Full licence detail is on attribution, and every score here is in the CSV download.