The Known Good Updated 4 Sep 2026 Subscribe
Evaluations· Epoch AI· CC-BY 4.0· Known Good Index component

Humanity's Last Exam

Expert-written questions spanning many specialist fields. Aggregated from Epoch AI's AI Benchmarking Hub (hle_external.csv, column 'Accuracy'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations.

Publisher's page Licence & attribution How we use this score
45 models scored Higher is better Published by Epoch AI CC-BY 4.0
Every model we hold a Humanity's Last Exam score for
All 45 of 45 models
# Model Creator Humanity's Last Exam Known Good Index Measured
1 Google: Gemini 3.1 Pro Preview 💡 Google 46.4 97 4 Sep 2026
2 OpenAI: GPT-5.4 Pro 💡 OpenAI 44.3 4 Sep 2026
3 muse-spark Meta 40.6 4 Sep 2026
4 gemini-3-pro-preview Google 37.5 84 4 Sep 2026
5 OpenAI: GPT-5.4 💡 OpenAI 36.2 91 4 Sep 2026
6 Anthropic: Claude Opus 4.7 💡 Anthropic 36.2 92 4 Sep 2026
7 Anthropic: Claude Opus 4.6 💡 Anthropic 34.4 89 4 Sep 2026
8 OpenAI: GPT-5 Pro 💡 OpenAI 31.6 4 Sep 2026
9 OpenAI: GPT-5.2 💡 OpenAI 27.8 81 4 Sep 2026
10 OpenAI: GPT-5 💡 OpenAI 25.3 78 4 Sep 2026
11 claude-opus-4-5-20251101-thinking Anthropic 25.2 4 Sep 2026
12 MoonshotAI: Kimi K2.5 💡 Moonshot AI 24.4 4 Sep 2026
13 OpenAI: GPT-5.1 💡 OpenAI 23.7 74 4 Sep 2026
14 Google: Gemini 2.5 Pro Preview 06-05 💡 Google 21.6 4 Sep 2026
15 OpenAI: o3 💡 OpenAI 20.3 74 4 Sep 2026
16 OpenAI: GPT-5 Mini 💡 OpenAI 19.4 67 4 Sep 2026
17 Gemini 2.5 Pro Experimental (March 2025) Google 18.2 4 Sep 2026
18 OpenAI: o4 Mini 💡 OpenAI 18.1 4 Sep 2026
19 Google: Gemini 2.5 Pro Preview 05-06 💡 Google 17.8 4 Sep 2026
20 Anthropic: Claude Opus 4.5 💡 Anthropic 14.2 72 4 Sep 2026
21 claude-sonnet-4-5-20250929-thinking Anthropic 13.7 4 Sep 2026
22 Google: Gemini 2.5 Flash 💡 Google 12.1 4 Sep 2026
23 claude-opus-4-1-20250805-thinking Anthropic 11.5 4 Sep 2026
24 Gemini 2.5 Flash Preview (May 2025) Google 11.0 4 Sep 2026
25 Anthropic: Claude Opus 4 💡 Anthropic 10.7 4 Sep 2026
26 Google: Gemini 3.1 Flash Lite Preview 💡 Google 8.6 4 Sep 2026
27 Z.ai: GLM 4.5 💡 Z.ai 8.3 4 Sep 2026
28 Z.ai: GLM 4.5 Air 💡 Z.ai 8.1 4 Sep 2026
29 OpenAI: o1-pro 💡 OpenAI 8.1 4 Sep 2026
30 Claude 3.7 Sonnet (Thinking) Anthropic 8.0 4 Sep 2026
31 OpenAI: o1 💡 OpenAI 8.0 70 4 Sep 2026
32 Anthropic: Claude Opus 4.1 💡 Anthropic 7.9 50 4 Sep 2026
33 Anthropic: Claude Sonnet 4 💡 Anthropic 7.8 4 Sep 2026
34 claude-sonnet-4-5-20250929 Anthropic 7.5 66 4 Sep 2026
35 gpt-5.1-instant OpenAI 6.8 4 Sep 2026
36 Gemini 2.0 Flash Thinking (January 2025) Google DeepMind,Google 6.6 4 Sep 2026
37 Meta: Llama 4 Maverick Meta 5.7 4 Sep 2026
38 gpt-4.5-preview OpenAI 5.4 54 4 Sep 2026
39 OpenAI: GPT-4.1 OpenAI 5.4 46 4 Sep 2026
40 gemini-1.5-pro-002 Google 4.6 4 Sep 2026
41 Mistral: Mistral Medium 3 Mistral 4.5 4 Sep 2026
42 Nova Pro Amazon 4.4 4 Sep 2026
43 Claude 3.5 Sonnet (October 2024) Anthropic 4.1 4 Sep 2026
44 Nova Lite Amazon 3.6 4 Sep 2026
45 OpenAI: GPT-4o OpenAI 2.7 27 4 Sep 2026
The Known Good
Provenance. These figures are published by Epoch AI and redistributed here under CC-BY 4.0. The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is min-max normalised against every other model holding the same evaluation, and nothing else is done to it. This evaluation is one of the seven components of the Known Good Index. Full licence detail is on attribution, and every score here is in the CSV download.