The Known Good Updated 25 Jul 2026
Evaluations· LiveBench· Apache-2.0· Known Good Index component

LiveBench

Contamination-free objective benchmark, refreshed monthly. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'Global average'). The Known Good aggregates published results; it runs no evaluations.

Publisher's page Licence & attribution How we use this score
52 models scored Higher is better Published by LiveBench Apache-2.0
LiveBench

LiveBench as measured and published by LiveBench (Apache-2.0). Ingested, normalised only where it feeds an index, and never re-run by us. Higher is better.

16 of 52 models +Add model from specific provider
💡 Reasoning models are indicated by a lightbulb
The Known Good
Every model we hold a LiveBench score for
52 of 52 models
# Model Creator LiveBench Known Good Index Measured
1 gemini-2.5-pro-exp-03-25 Google 82.3 25 Jul 2026
2 OpenAI: GPT-5.1 💡 OpenAI 78.8 74 25 Jul 2026
3 claude-3-7-sonnet-20250219 Anthropic 76.1 70 25 Jul 2026
4 OpenAI: o3 Mini 💡 OpenAI 75.9 82 25 Jul 2026
5 OpenAI: o1 💡 OpenAI 75.7 63 25 Jul 2026
6 QwQ-32B Alibaba 72.0 25 Jul 2026
7 DeepSeek: R1 💡 DeepSeek 71.6 68 25 Jul 2026
8 deepseek-r1 DeepSeek 71.6 68 25 Jul 2026
9 gpt-4.5-preview OpenAI 69.0 47 25 Jul 2026
10 gemini-2.0-flash-thinking-exp-01-21 Google DeepMind,Google 66.9 62 25 Jul 2026
11 DeepSeek: DeepSeek V3 0324 DeepSeek 66.9 60 25 Jul 2026
12 gemini-2.0-pro-exp-02-05 Google 65.1 25 Jul 2026
13 gemini-exp-1206 Google DeepMind,Google 64.1 25 Jul 2026
14 qwen2.5-max Alibaba 62.3 25 Jul 2026
15 gemini-2.0-flash-001 Google DeepMind,Google 61.5 53 25 Jul 2026
16 DeepSeek: DeepSeek V3 DeepSeek 60.5 44 25 Jul 2026
17 gemini-2.0-flash-exp Google DeepMind,Google 59.3 25 Jul 2026
18 claude-3-5-sonnet-20241022 Anthropic 59.0 40 25 Jul 2026
19 o1-mini OpenAI 57.8 56 25 Jul 2026
20 OpenAI: GPT-4o OpenAI 55.3 21 25 Jul 2026
21 DeepSeek: R1 Distill Llama 70B 💡 DeepSeek 54.5 52 25 Jul 2026
22 grok-2-1212 xAI 54.3 38 25 Jul 2026
23 gemini-2.0-flash-lite Google 54.3 25 Jul 2026
24 gemini-2.0-flash-lite-preview-02-05 Google 53.2 25 Jul 2026
25 Dracarys2-72B-Instruct Unknown 52.6 25 Jul 2026
26 learnlm-1.5-pro-experimental Unknown 52.2 25 Jul 2026
27 Meta: Llama 3.3 70B Instruct Meta 50.2 31 25 Jul 2026
28 Google: Gemma 3 27B Google 50.0 37 25 Jul 2026
29 gemma-3-27b-it Google 50.0 37 25 Jul 2026
30 claude-3-opus-20240229 Anthropic 49.2 30 25 Jul 2026
31 mistral-large-2411 Mistral 48.4 33 25 Jul 2026
32 Perplexity: Sonar Perplexity 46.9 25 Jul 2026
33 Qwen2.5 Coder 32B Instruct Alibaba 46.2 25 Jul 2026
34 Dracarys2-Llama-3.1-70B-Instruct Unknown 46.2 25 Jul 2026
35 DeepSeek-R1-Distill-Qwen-32B DeepSeek 45.5 25 Jul 2026
36 mistral-small-2503 Mistral 44.0 28 25 Jul 2026
37 amazon.nova-pro-v1:0 Amazon 43.5 25 Jul 2026
38 claude-3-5-haiku-20241022 Anthropic 43.5 23 25 Jul 2026
39 mistral-small-2501 Mistral 42.5 26 25 Jul 2026
40 Microsoft: Phi 4 Microsoft 41.6 33 25 Jul 2026
41 OpenAI: GPT-4o-mini OpenAI 41.3 23 25 Jul 2026
42 QwQ-32B-Preview Alibaba 40.2 25 Jul 2026
43 Google: Gemma 2 27B Google 38.2 19 25 Jul 2026
44 gemma-2-27b-it Google 38.2 19 25 Jul 2026
45 amazon.nova-lite-v1:0 Amazon 36.4 25 Jul 2026
46 c4ai-command-r-plus-08-2024 Cohere,Cohere for AI 31.8 25 Jul 2026
47 amazon.nova-micro-v1:0 Amazon 29.6 25 Jul 2026
48 gemma-2-9b-it Google 28.7 10 25 Jul 2026
49 c4ai-command-r-08-2024 Cohere 27.5 25 Jul 2026
50 Phi-3-small-8k-instruct Microsoft 24.0 25 Jul 2026
51 Phi-3-mini-4k-instruct Microsoft 22.4 25 Jul 2026
52 OLMo-2-1124-13B-Instruct Allen Institute for AI,University of Washington,New York University (NYU) 22.1 25 Jul 2026
The Known Good
Provenance. These figures are published by LiveBench and redistributed here under Apache-2.0. The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is min-max normalised against every other model holding the same evaluation, and nothing else is done to it. This evaluation is one of the six components of the Known Good Index. Full licence detail is on attribution, and every score here is in the CSV download.