The Known Good Updated 25 Jul 2026
Evaluations· LiveBench· Apache-2.0· Tracked, not in the index

LiveBench Instruction Following

LiveBench instruction-following subset. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'IF Average'). The Known Good aggregates published results; it runs no evaluations.

Publisher's page Licence & attribution How we use this score
52 models scored Higher is better Published by LiveBench Apache-2.0
LiveBench Instruction Following

LiveBench Instruction Following as measured and published by LiveBench (Apache-2.0). Ingested, normalised only where it feeds an index, and never re-run by us. Higher is better.

16 of 52 models +Add model from specific provider
💡 Reasoning models are indicated by a lightbulb
The Known Good
Every model we hold a LiveBench Instruction Following score for
52 of 52 models
# Model Creator LiveBench Instruction Following Known Good Index Measured
1 OpenAI: GPT-5.1 💡 OpenAI 93.3 74 25 Jul 2026
2 gemini-2.0-flash-001 Google DeepMind,Google 85.8 53 25 Jul 2026
3 OpenAI: o3 Mini 💡 OpenAI 84.4 82 25 Jul 2026
4 gemini-2.0-pro-exp-02-05 Google 83.4 25 Jul 2026
5 Meta: Llama 3.3 70B Instruct Meta 82.7 31 25 Jul 2026
6 gemini-2.0-flash-thinking-exp-01-21 Google DeepMind,Google 82.5 62 25 Jul 2026
7 gemini-2.0-flash-exp Google DeepMind,Google 81.9 25 Jul 2026
8 QwQ-32B Alibaba 81.8 25 Jul 2026
9 OpenAI: o1 💡 OpenAI 81.5 63 25 Jul 2026
10 DeepSeek: DeepSeek V3 0324 DeepSeek 81.5 60 25 Jul 2026
11 claude-3-7-sonnet-20250219 Anthropic 81.2 70 25 Jul 2026
12 gemini-2.5-pro-exp-03-25 Google 80.6 25 Jul 2026
13 DeepSeek: R1 💡 DeepSeek 80.5 68 25 Jul 2026
14 deepseek-r1 DeepSeek 80.5 68 25 Jul 2026
15 gemini-2.0-flash-lite-preview-02-05 Google 78.3 25 Jul 2026
16 gemini-exp-1206 Google DeepMind,Google 77.3 25 Jul 2026
17 gemini-2.0-flash-lite Google 76.6 25 Jul 2026
18 Perplexity: Sonar Perplexity 76.2 25 Jul 2026
19 qwen2.5-max Alibaba 75.3 25 Jul 2026
20 DeepSeek: DeepSeek V3 DeepSeek 75.2 44 25 Jul 2026
21 Google: Gemma 3 27B Google 74.9 37 25 Jul 2026
22 gemma-3-27b-it Google 74.9 37 25 Jul 2026
23 gpt-4.5-preview OpenAI 72.3 47 25 Jul 2026
24 DeepSeek: R1 Distill Llama 70B 💡 DeepSeek 69.9 52 25 Jul 2026
25 grok-2-1212 xAI 69.6 38 25 Jul 2026
26 claude-3-5-sonnet-20241022 Anthropic 69.3 40 25 Jul 2026
27 OpenAI: GPT-4o OpenAI 68.6 21 25 Jul 2026
28 learnlm-1.5-pro-experimental Unknown 68.2 25 Jul 2026
29 mistral-large-2411 Mistral 67.9 33 25 Jul 2026
30 amazon.nova-pro-v1:0 Amazon 67.1 25 Jul 2026
31 o1-mini OpenAI 65.4 56 25 Jul 2026
32 Dracarys2-72B-Instruct Unknown 65.2 25 Jul 2026
33 claude-3-opus-20240229 Anthropic 63.9 30 25 Jul 2026
34 mistral-small-2503 Mistral 63.7 28 25 Jul 2026
35 Dracarys2-Llama-3.1-70B-Instruct Unknown 63.2 25 Jul 2026
36 claude-3-5-haiku-20241022 Anthropic 61.9 23 25 Jul 2026
37 OLMo-2-1124-13B-Instruct Allen Institute for AI,University of Washington,New York University (NYU) 60.6 25 Jul 2026
38 mistral-small-2501 Mistral 59.5 26 25 Jul 2026
39 Qwen2.5 Coder 32B Instruct Alibaba 58.7 25 Jul 2026
40 Microsoft: Phi 4 Microsoft 58.4 33 25 Jul 2026
41 Google: Gemma 2 27B Google 58.1 19 25 Jul 2026
42 gemma-2-27b-it Google 58.1 19 25 Jul 2026
43 c4ai-command-r-plus-08-2024 Cohere,Cohere for AI 57.6 25 Jul 2026
44 OpenAI: GPT-4o-mini OpenAI 56.8 23 25 Jul 2026
45 DeepSeek-R1-Distill-Qwen-32B DeepSeek 55.7 25 Jul 2026
46 c4ai-command-r-08-2024 Cohere 55.6 25 Jul 2026
47 amazon.nova-lite-v1:0 Amazon 54.1 25 Jul 2026
48 gemma-2-9b-it Google 52.6 10 25 Jul 2026
49 amazon.nova-micro-v1:0 Amazon 48.0 25 Jul 2026
50 Phi-3-small-8k-instruct Microsoft 47.2 25 Jul 2026
51 Phi-3-mini-4k-instruct Microsoft 39.1 25 Jul 2026
52 QwQ-32B-Preview Alibaba 35.6 25 Jul 2026
The Known Good
Provenance. These figures are published by LiveBench and redistributed here under Apache-2.0. The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is min-max normalised against every other model holding the same evaluation, and nothing else is done to it. Full licence detail is on attribution, and every score here is in the CSV download.