The Known Good Updated 25 Jul 2026
Evaluations· LiveBench· Apache-2.0· Tracked, not in the index

LiveBench Data Analysis

LiveBench data-analysis subset. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'Data analysis average'). The Known Good aggregates published results; it runs no evaluations.

Publisher's page Licence & attribution How we use this score
52 models scored Higher is better Published by LiveBench Apache-2.0
LiveBench Data Analysis

LiveBench Data Analysis as measured and published by LiveBench (Apache-2.0). Ingested, normalised only where it feeds an index, and never re-run by us. Higher is better.

16 of 52 models +Add model from specific provider
💡 Reasoning models are indicated by a lightbulb
The Known Good
Every model we hold a LiveBench Data Analysis score for
52 of 52 models
# Model Creator LiveBench Data Analysis Known Good Index Measured
1 gemini-2.5-pro-exp-03-25 Google 79.9 25 Jul 2026
2 claude-3-7-sonnet-20250219 Anthropic 74.0 70 25 Jul 2026
3 OpenAI: GPT-5.1 💡 OpenAI 72.1 74 25 Jul 2026
4 OpenAI: o3 Mini 💡 OpenAI 70.6 82 25 Jul 2026
5 DeepSeek: R1 💡 DeepSeek 69.8 68 25 Jul 2026
6 deepseek-r1 DeepSeek 69.8 68 25 Jul 2026
7 gemini-2.0-flash-thinking-exp-01-21 Google DeepMind,Google 69.4 62 25 Jul 2026
8 gemini-2.0-pro-exp-02-05 Google 68.0 25 Jul 2026
9 qwen2.5-max Alibaba 67.9 25 Jul 2026
10 gemini-2.0-flash-001 Google DeepMind,Google 67.5 53 25 Jul 2026
11 OpenAI: o1 💡 OpenAI 65.5 63 25 Jul 2026
12 gemini-2.0-flash-lite Google 65.5 25 Jul 2026
13 QwQ-32B Alibaba 65.0 25 Jul 2026
14 gpt-4.5-preview OpenAI 64.3 47 25 Jul 2026
15 gemini-exp-1206 Google DeepMind,Google 63.2 25 Jul 2026
16 gemini-2.0-flash-exp Google DeepMind,Google 61.7 25 Jul 2026
17 DeepSeek: DeepSeek V3 DeepSeek 60.9 44 25 Jul 2026
18 OpenAI: GPT-4o OpenAI 60.9 21 25 Jul 2026
19 DeepSeek: DeepSeek V3 0324 DeepSeek 60.4 60 25 Jul 2026
20 o1-mini OpenAI 57.9 56 25 Jul 2026
21 claude-3-opus-20240229 Anthropic 57.9 30 25 Jul 2026
22 gemini-2.0-flash-lite-preview-02-05 Google 57.5 25 Jul 2026
23 DeepSeek: R1 Distill Llama 70B 💡 DeepSeek 55.9 52 25 Jul 2026
24 Dracarys2-72B-Instruct Unknown 55.5 25 Jul 2026
25 claude-3-5-sonnet-20241022 Anthropic 55.0 40 25 Jul 2026
26 learnlm-1.5-pro-experimental Unknown 55.0 25 Jul 2026
27 grok-2-1212 xAI 54.5 38 25 Jul 2026
28 Dracarys2-Llama-3.1-70B-Instruct Unknown 54.0 25 Jul 2026
29 mistral-small-2501 Mistral 53.7 26 25 Jul 2026
30 Google: Gemma 3 27B Google 51.5 37 25 Jul 2026
31 gemma-3-27b-it Google 51.5 37 25 Jul 2026
32 mistral-small-2503 Mistral 50.5 28 25 Jul 2026
33 mistral-large-2411 Mistral 50.1 33 25 Jul 2026
34 OpenAI: GPT-4o-mini OpenAI 50.0 23 25 Jul 2026
35 Qwen2.5 Coder 32B Instruct Alibaba 49.9 25 Jul 2026
36 Meta: Llama 3.3 70B Instruct Meta 49.5 31 25 Jul 2026
37 claude-3-5-haiku-20241022 Anthropic 48.5 23 25 Jul 2026
38 amazon.nova-pro-v1:0 Amazon 48.3 25 Jul 2026
39 Google: Gemma 2 27B Google 47.9 19 25 Jul 2026
40 gemma-2-27b-it Google 47.9 19 25 Jul 2026
41 DeepSeek-R1-Distill-Qwen-32B DeepSeek 45.4 25 Jul 2026
42 Microsoft: Phi 4 Microsoft 45.2 33 25 Jul 2026
43 c4ai-command-r-plus-08-2024 Cohere,Cohere for AI 38.1 25 Jul 2026
44 Perplexity: Sonar Perplexity 37.9 25 Jul 2026
45 amazon.nova-lite-v1:0 Amazon 37.2 25 Jul 2026
46 gemma-2-9b-it Google 36.4 10 25 Jul 2026
47 Phi-3-mini-4k-instruct Microsoft 34.7 25 Jul 2026
48 amazon.nova-micro-v1:0 Amazon 34.0 25 Jul 2026
49 c4ai-command-r-08-2024 Cohere 33.3 25 Jul 2026
50 QwQ-32B-Preview Alibaba 31.6 25 Jul 2026
51 Phi-3-small-8k-instruct Microsoft 30.3 25 Jul 2026
52 OLMo-2-1124-13B-Instruct Allen Institute for AI,University of Washington,New York University (NYU) 20.6 25 Jul 2026
The Known Good
Provenance. These figures are published by LiveBench and redistributed here under Apache-2.0. The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is min-max normalised against every other model holding the same evaluation, and nothing else is done to it. Full licence detail is on attribution, and every score here is in the CSV download.