LiveBench Reasoning
LiveBench reasoning subset. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'Reasoning average'). The Known Good aggregates published results; it runs no evaluations.
52 models scored
Higher is better
Published by LiveBench
Apache-2.0
LiveBench Reasoning
LiveBench Reasoning as measured and published by LiveBench (Apache-2.0). Ingested, normalised only where it feeds an index, and never re-run by us. Higher is better.
16 of 52 models ▾
☰⚙⇩
+Add model from specific provider
💡 Reasoning models are indicated by a lightbulb
The Known Good
Every model we hold a LiveBench Reasoning score for
52 of 52 models
| # | Model | Creator | LiveBench Reasoning | Known Good Index | Measured |
|---|---|---|---|---|---|
| 1 |
|
OpenAI | 95.8 | 74 | 25 Jul 2026 |
| 2 |
|
OpenAI | 91.6 | 63 | 25 Jul 2026 |
| 3 |
|
89.8 | — | 25 Jul 2026 | |
| 4 |
|
OpenAI | 89.6 | 82 | 25 Jul 2026 |
| 5 |
|
Anthropic | 87.8 | 70 | 25 Jul 2026 |
| 6 |
|
Alibaba | 83.5 | — | 25 Jul 2026 |
| 7 |
|
DeepSeek | 83.2 | 68 | 25 Jul 2026 |
| 8 |
|
DeepSeek | 83.2 | 68 | 25 Jul 2026 |
| 9 | GD gemini-2.0-flash-thinking-exp-01-21 | Google DeepMind,Google | 78.2 | 62 | 25 Jul 2026 |
| 10 |
|
OpenAI | 72.3 | 56 | 25 Jul 2026 |
| 11 |
|
OpenAI | 71.1 | 47 | 25 Jul 2026 |
| 12 |
|
DeepSeek | 67.6 | 52 | 25 Jul 2026 |
| 13 |
|
DeepSeek | 65.8 | 60 | 25 Jul 2026 |
| 14 |
|
60.1 | — | 25 Jul 2026 | |
| 15 | GD gemini-2.0-flash-exp | Google DeepMind,Google | 59.1 | — | 25 Jul 2026 |
| 16 |
|
Alibaba | 57.7 | — | 25 Jul 2026 |
| 17 | GD gemini-exp-1206 | Google DeepMind,Google | 57.0 | — | 25 Jul 2026 |
| 18 |
|
DeepSeek | 56.8 | 44 | 25 Jul 2026 |
| 19 |
|
Anthropic | 56.7 | 40 | 25 Jul 2026 |
| 20 |
|
OpenAI | 55.8 | 21 | 25 Jul 2026 |
| 21 | GD gemini-2.0-flash-001 | Google DeepMind,Google | 55.2 | 53 | 25 Jul 2026 |
| 22 |
|
xAI | 54.8 | 38 | 25 Jul 2026 |
| 23 |
|
DeepSeek | 52.2 | — | 25 Jul 2026 |
| 24 |
|
Alibaba | 51.4 | — | 25 Jul 2026 |
| 25 |
|
Meta | 50.8 | 31 | 25 Jul 2026 |
| 26 |
|
50.1 | — | 25 Jul 2026 | |
| 27 |
|
Microsoft | 47.8 | 33 | 25 Jul 2026 |
| 28 | D7 Dracarys2-72B-Instruct | Unknown | 47.4 | — | 25 Jul 2026 |
| 29 |
|
Perplexity | 46.2 | — | 25 Jul 2026 |
| 30 |
|
44.9 | — | 25 Jul 2026 | |
| 31 |
|
Mistral | 44.8 | 28 | 25 Jul 2026 |
| 32 | DL Dracarys2-Llama-3.1-70B-Instruct | Unknown | 44.7 | — | 25 Jul 2026 |
| 33 |
|
43.8 | 37 | 25 Jul 2026 | |
| 34 |
|
43.8 | 37 | 25 Jul 2026 | |
| 35 |
|
Mistral | 43.5 | 33 | 25 Jul 2026 |
| 36 | L1 learnlm-1.5-pro-experimental | Unknown | 43.4 | — | 25 Jul 2026 |
| 37 |
|
Alibaba | 42.1 | — | 25 Jul 2026 |
| 38 |
|
Anthropic | 40.6 | 30 | 25 Jul 2026 |
| 39 | Az amazon.nova-lite-v1:0 | Amazon | 36.7 | — | 25 Jul 2026 |
| 40 |
|
Mistral | 36.4 | 26 | 25 Jul 2026 |
| 41 |
|
OpenAI | 32.8 | 23 | 25 Jul 2026 |
| 42 | Az amazon.nova-pro-v1:0 | Amazon | 32.6 | — | 25 Jul 2026 |
| 43 |
|
28.1 | 19 | 25 Jul 2026 | |
| 44 |
|
Anthropic | 28.1 | 23 | 25 Jul 2026 |
| 45 |
|
28.1 | 19 | 25 Jul 2026 | |
| 46 |
|
Microsoft | 26.8 | — | 25 Jul 2026 |
| 47 | Az amazon.nova-micro-v1:0 | Amazon | 25.1 | — | 25 Jul 2026 |
| 48 | CC c4ai-command-r-plus-08-2024 | Cohere,Cohere for AI | 24.8 | — | 25 Jul 2026 |
| 49 |
|
Cohere | 21.9 | — | 25 Jul 2026 |
| 50 | AI OLMo-2-1124-13B-Instruct | Allen Institute for AI,University of Washington,New York University (NYU) | 16.3 | — | 25 Jul 2026 |
| 51 |
|
Microsoft | 15.9 | — | 25 Jul 2026 |
| 52 |
|
15.2 | 10 | 25 Jul 2026 |
The Known Good
Provenance. These figures are published by LiveBench and redistributed here under Apache-2.0.
The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is
min-max normalised against every other model holding the same evaluation, and nothing else is done to it.
Full licence detail is on attribution, and every score here is
in the CSV download.