Humanity's Last Exam
Expert-written questions spanning many specialist fields. Aggregated from Epoch AI's AI Benchmarking Hub (hle_external.csv, column 'Accuracy'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations.
45 models scored
Higher is better
Published by Epoch AI
CC-BY 4.0
Humanity's Last Exam
Humanity's Last Exam as measured and published by Epoch AI (CC-BY 4.0). Ingested, normalised only where it feeds an index, and never re-run by us. Higher is better.
16 of 45 models ▾
☰⚙⇩
+Add model from specific provider
💡 Reasoning models are indicated by a lightbulb
The Known Good
Every model we hold a Humanity's Last Exam score for
45 of 45 models
| # | Model | Creator | Humanity's Last Exam | Known Good Index | Measured |
|---|---|---|---|---|---|
| 1 |
|
46.4 | 96 | 25 Jul 2026 | |
| 2 |
|
OpenAI | 44.3 | — | 25 Jul 2026 |
| 3 |
|
Meta | 40.6 | — | 25 Jul 2026 |
| 4 |
|
37.5 | 83 | 25 Jul 2026 | |
| 5 |
|
OpenAI | 36.2 | 90 | 25 Jul 2026 |
| 6 |
|
Anthropic | 36.2 | 94 | 25 Jul 2026 |
| 7 |
|
Anthropic | 34.4 | 88 | 25 Jul 2026 |
| 8 |
|
OpenAI | 31.6 | — | 25 Jul 2026 |
| 9 |
|
OpenAI | 27.8 | 80 | 25 Jul 2026 |
| 10 |
|
OpenAI | 25.3 | 73 | 25 Jul 2026 |
| 11 |
|
Anthropic | 25.2 | — | 25 Jul 2026 |
| 12 |
|
Moonshot AI | 24.4 | — | 25 Jul 2026 |
| 13 |
|
OpenAI | 23.7 | 74 | 25 Jul 2026 |
| 14 |
|
21.6 | — | 25 Jul 2026 | |
| 15 |
|
OpenAI | 20.3 | 67 | 25 Jul 2026 |
| 16 |
|
OpenAI | 19.4 | 60 | 25 Jul 2026 |
| 17 |
|
18.2 | — | 25 Jul 2026 | |
| 18 |
|
OpenAI | 18.1 | — | 25 Jul 2026 |
| 19 |
|
17.8 | — | 25 Jul 2026 | |
| 20 |
|
Anthropic | 14.2 | 72 | 25 Jul 2026 |
| 21 |
|
Anthropic | 13.7 | — | 25 Jul 2026 |
| 22 |
|
12.1 | — | 25 Jul 2026 | |
| 23 |
|
Anthropic | 11.5 | — | 25 Jul 2026 |
| 24 |
|
11.0 | — | 25 Jul 2026 | |
| 25 |
|
Anthropic | 10.7 | — | 25 Jul 2026 |
| 26 |
|
8.6 | — | 25 Jul 2026 | |
| 27 |
|
Z.ai | 8.3 | — | 25 Jul 2026 |
| 28 |
|
Z.ai | 8.1 | — | 25 Jul 2026 |
| 29 |
|
OpenAI | 8.1 | — | 25 Jul 2026 |
| 30 |
|
Anthropic | 8.0 | — | 25 Jul 2026 |
| 31 |
|
OpenAI | 8.0 | 63 | 25 Jul 2026 |
| 32 |
|
Anthropic | 7.9 | 49 | 25 Jul 2026 |
| 33 |
|
Anthropic | 7.8 | — | 25 Jul 2026 |
| 34 |
|
Anthropic | 7.5 | 59 | 25 Jul 2026 |
| 35 |
|
OpenAI | 6.8 | — | 25 Jul 2026 |
| 36 | GD Gemini 2.0 Flash Thinking (January 2025) | Google DeepMind,Google | 6.6 | — | 25 Jul 2026 |
| 37 |
|
Meta | 5.7 | — | 25 Jul 2026 |
| 38 |
|
OpenAI | 5.4 | 47 | 25 Jul 2026 |
| 39 |
|
OpenAI | 5.4 | 36 | 25 Jul 2026 |
| 40 |
|
4.6 | — | 25 Jul 2026 | |
| 41 |
|
Mistral | 4.5 | — | 25 Jul 2026 |
| 42 | Az Nova Pro | Amazon | 4.4 | — | 25 Jul 2026 |
| 43 |
|
Anthropic | 4.1 | — | 25 Jul 2026 |
| 44 | Az Nova Lite | Amazon | 3.6 | — | 25 Jul 2026 |
| 45 |
|
OpenAI | 2.7 | 21 | 25 Jul 2026 |
The Known Good
Provenance. These figures are published by Epoch AI and redistributed here under CC-BY 4.0.
The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is
min-max normalised against every other model holding the same evaluation, and nothing else is done to it.
This evaluation is one of the six components of the
Known Good Index. Full licence detail is on attribution, and every score here is
in the CSV download.