The Known Good Updated 25 Jul 2026
Evaluations· LiveBench· Apache-2.0· Tracked, not in the index

LiveBench Mathematics

LiveBench mathematics subset. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'Mathematics average'). The Known Good aggregates published results; it runs no evaluations.

Publisher's page Licence & attribution How we use this score
52 models scored Higher is better Published by LiveBench Apache-2.0
LiveBench Mathematics

LiveBench Mathematics as measured and published by LiveBench (Apache-2.0). Ingested, normalised only where it feeds an index, and never re-run by us. Higher is better.

16 of 52 models +Add model from specific provider
💡 Reasoning models are indicated by a lightbulb
The Known Good
Every model we hold a LiveBench Mathematics score for
52 of 52 models
# Model Creator LiveBench Mathematics Known Good Index Measured
1 OpenAI: GPT-5.1 💡 OpenAI 94.5 74 25 Jul 2026
2 gemini-2.5-pro-exp-03-25 Google 90.2 25 Jul 2026
3 DeepSeek: R1 💡 DeepSeek 80.7 68 25 Jul 2026
4 deepseek-r1 DeepSeek 80.7 68 25 Jul 2026
5 OpenAI: o1 💡 OpenAI 80.3 63 25 Jul 2026
6 claude-3-7-sonnet-20250219 Anthropic 79.0 70 25 Jul 2026
7 QwQ-32B Alibaba 77.8 25 Jul 2026
8 OpenAI: o3 Mini 💡 OpenAI 77.3 82 25 Jul 2026
9 gemini-2.0-flash-thinking-exp-01-21 Google DeepMind,Google 75.8 62 25 Jul 2026
10 DeepSeek: DeepSeek V3 0324 DeepSeek 73.5 60 25 Jul 2026
11 gemini-exp-1206 Google DeepMind,Google 72.4 25 Jul 2026
12 gemini-2.0-pro-exp-02-05 Google 71.0 25 Jul 2026
13 gpt-4.5-preview OpenAI 69.3 47 25 Jul 2026
14 gemini-2.0-flash-001 Google DeepMind,Google 65.6 53 25 Jul 2026
15 o1-mini OpenAI 62.0 56 25 Jul 2026
16 DeepSeek: DeepSeek V3 DeepSeek 60.5 44 25 Jul 2026
17 gemini-2.0-flash-exp Google DeepMind,Google 60.4 25 Jul 2026
18 DeepSeek-R1-Distill-Qwen-32B DeepSeek 59.4 25 Jul 2026
19 qwen2.5-max Alibaba 58.4 25 Jul 2026
20 QwQ-32B-Preview Alibaba 58.3 25 Jul 2026
21 DeepSeek: R1 Distill Llama 70B 💡 DeepSeek 58.1 52 25 Jul 2026
22 gemini-2.0-flash-lite Google 58.1 25 Jul 2026
23 learnlm-1.5-pro-experimental Unknown 57.8 25 Jul 2026
24 gemini-2.0-flash-lite-preview-02-05 Google 55.5 25 Jul 2026
25 Google: Gemma 3 27B Google 55.4 37 25 Jul 2026
26 gemma-3-27b-it Google 55.4 37 25 Jul 2026
27 grok-2-1212 xAI 54.9 38 25 Jul 2026
28 Dracarys2-72B-Instruct Unknown 54.7 25 Jul 2026
29 claude-3-5-sonnet-20241022 Anthropic 52.3 40 25 Jul 2026
30 OpenAI: GPT-4o OpenAI 49.5 21 25 Jul 2026
31 Qwen2.5 Coder 32B Instruct Alibaba 46.6 25 Jul 2026
32 claude-3-opus-20240229 Anthropic 43.6 30 25 Jul 2026
33 mistral-large-2411 Mistral 42.5 33 25 Jul 2026
34 Meta: Llama 3.3 70B Instruct Meta 42.2 31 25 Jul 2026
35 Microsoft: Phi 4 Microsoft 42.0 33 25 Jul 2026
36 Perplexity: Sonar Perplexity 41.6 25 Jul 2026
37 Dracarys2-Llama-3.1-70B-Instruct Unknown 40.3 25 Jul 2026
38 mistral-small-2501 Mistral 39.9 26 25 Jul 2026
39 mistral-small-2503 Mistral 39.4 28 25 Jul 2026
40 amazon.nova-pro-v1:0 Amazon 38.0 25 Jul 2026
41 amazon.nova-lite-v1:0 Amazon 36.7 25 Jul 2026
42 OpenAI: GPT-4o-mini OpenAI 36.3 23 25 Jul 2026
43 claude-3-5-haiku-20241022 Anthropic 35.5 23 25 Jul 2026
44 amazon.nova-micro-v1:0 Amazon 34.5 25 Jul 2026
45 Google: Gemma 2 27B Google 26.5 19 25 Jul 2026
46 gemma-2-27b-it Google 26.5 19 25 Jul 2026
47 c4ai-command-r-plus-08-2024 Cohere,Cohere for AI 21.3 25 Jul 2026
48 gemma-2-9b-it Google 19.8 10 25 Jul 2026
49 c4ai-command-r-08-2024 Cohere 19.4 25 Jul 2026
50 Phi-3-small-8k-instruct Microsoft 17.6 25 Jul 2026
51 Phi-3-mini-4k-instruct Microsoft 15.7 25 Jul 2026
52 OLMo-2-1124-13B-Instruct Allen Institute for AI,University of Washington,New York University (NYU) 13.6 25 Jul 2026
The Known Good
Provenance. These figures are published by LiveBench and redistributed here under Apache-2.0. The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is min-max normalised against every other model holding the same evaluation, and nothing else is done to it. Full licence detail is on attribution, and every score here is in the CSV download.