The Known Good Updated 4 Sep 2026 Subscribe
Evaluations· LiveBench· Apache-2.0· Tracked, not in the index

LiveBench Data Analysis

LiveBench data-analysis subset. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'Data analysis average'). The Known Good aggregates published results; it runs no evaluations.

Publisher's page Licence & attribution How we use this score
49 models scored Higher is better Published by LiveBench Apache-2.0
Every model we hold a LiveBench Data Analysis score for
All 49 of 49 models
# Model Creator LiveBench Data Analysis Known Good Index Measured
1 gemini-2.5-pro-exp-03-25 Google 79.9 4 Sep 2026
2 claude-3-7-sonnet-20250219 Anthropic 74.0 74 4 Sep 2026
3 OpenAI: GPT-5.1 💡 OpenAI 72.1 74 4 Sep 2026
4 OpenAI: o3 Mini 💡 OpenAI 70.6 86 4 Sep 2026
5 DeepSeek: R1 💡 DeepSeek 69.8 75 4 Sep 2026
6 gemini-2.0-flash-thinking-exp-01-21 Google DeepMind,Google 69.4 62 4 Sep 2026
7 gemini-2.0-pro-exp-02-05 Google 68.0 74 4 Sep 2026
8 qwen2.5-max Alibaba 67.9 4 Sep 2026
9 gemini-2.0-flash-001 Google DeepMind,Google 67.5 61 4 Sep 2026
10 OpenAI: o1 💡 OpenAI 65.5 70 4 Sep 2026
11 gemini-2.0-flash-lite Google 65.5 4 Sep 2026
12 QwQ-32B Alibaba 65.0 69 4 Sep 2026
13 gpt-4.5-preview OpenAI 64.3 54 4 Sep 2026
14 gemini-exp-1206 Google DeepMind,Google 63.2 4 Sep 2026
15 gemini-2.0-flash-exp Google DeepMind,Google 61.7 4 Sep 2026
16 DeepSeek: DeepSeek V3 DeepSeek 60.9 62 4 Sep 2026
17 OpenAI: GPT-4o OpenAI 60.9 27 4 Sep 2026
18 DeepSeek: DeepSeek V3 0324 DeepSeek 60.4 64 4 Sep 2026
19 o1-mini OpenAI 57.9 65 4 Sep 2026
20 claude-3-opus-20240229 Anthropic 57.9 32 4 Sep 2026
21 gemini-2.0-flash-lite-preview-02-05 Google 57.5 4 Sep 2026
22 DeepSeek: R1 Distill Llama 70B 💡 DeepSeek 55.9 63 4 Sep 2026
23 Dracarys2-72B-Instruct Unknown 55.5 4 Sep 2026
24 claude-3-5-sonnet-20241022 Anthropic 55.0 45 4 Sep 2026
25 learnlm-1.5-pro-experimental Unknown 55.0 4 Sep 2026
26 grok-2-1212 xAI 54.5 45 4 Sep 2026
27 Dracarys2-Llama-3.1-70B-Instruct Unknown 54.0 4 Sep 2026
28 mistral-small-2501 Mistral 53.7 31 4 Sep 2026
29 Google: Gemma 3 27B 💡 Google 51.5 47 4 Sep 2026
30 mistral-small-2503 Mistral 50.5 33 4 Sep 2026
31 mistral-large-2411 Mistral 50.1 37 4 Sep 2026
32 OpenAI: GPT-4o-mini OpenAI 50.0 31 4 Sep 2026
33 Qwen2.5 Coder 32B Instruct Alibaba 49.9 4 Sep 2026
34 Meta: Llama 3.3 70B Instruct Meta 49.5 34 4 Sep 2026
35 claude-3-5-haiku-20241022 Anthropic 48.5 30 4 Sep 2026
36 amazon.nova-pro-v1:0 Amazon 48.3 4 Sep 2026
37 Google: Gemma 2 27B Google 47.9 21 4 Sep 2026
38 DeepSeek-R1-Distill-Qwen-32B DeepSeek 45.4 53 4 Sep 2026
39 Microsoft: Phi 4 Microsoft 45.2 41 4 Sep 2026
40 c4ai-command-r-plus-08-2024 Cohere,Cohere for AI 38.1 4 Sep 2026
41 Perplexity: Sonar Perplexity 37.9 4 Sep 2026
42 amazon.nova-lite-v1:0 Amazon 37.2 4 Sep 2026
43 gemma-2-9b-it Google 36.4 13 4 Sep 2026
44 Phi-3-mini-4k-instruct Microsoft 34.7 4 Sep 2026
45 amazon.nova-micro-v1:0 Amazon 34.0 4 Sep 2026
46 c4ai-command-r-08-2024 Cohere 33.3 4 Sep 2026
47 QwQ-32B-Preview Alibaba 31.6 4 Sep 2026
48 Phi-3-small-8k-instruct Microsoft 30.3 4 Sep 2026
49 OLMo-2-1124-13B-Instruct Allen Institute for AI,University of Washington,New York University (NYU) 20.6 4 Sep 2026
The Known Good
Provenance. These figures are published by LiveBench and redistributed here under Apache-2.0. The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is min-max normalised against every other model holding the same evaluation, and nothing else is done to it. Full licence detail is on attribution, and every score here is in the CSV download.