LiveBench Instruction Following
LiveBench instruction-following subset. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'IF Average'). The Known Good aggregates published results; it runs no evaluations.
52 models scored
Higher is better
Published by LiveBench
Apache-2.0
LiveBench Instruction Following
LiveBench Instruction Following as measured and published by LiveBench (Apache-2.0). Ingested, normalised only where it feeds an index, and never re-run by us. Higher is better.
16 of 52 models ▾
☰⚙⇩
+Add model from specific provider
💡 Reasoning models are indicated by a lightbulb
The Known Good
Every model we hold a LiveBench Instruction Following score for
52 of 52 models
| # | Model | Creator | LiveBench Instruction Following | Known Good Index | Measured |
|---|---|---|---|---|---|
| 1 |
|
OpenAI | 93.3 | 74 | 25 Jul 2026 |
| 2 | GD gemini-2.0-flash-001 | Google DeepMind,Google | 85.8 | 53 | 25 Jul 2026 |
| 3 |
|
OpenAI | 84.4 | 82 | 25 Jul 2026 |
| 4 |
|
83.4 | — | 25 Jul 2026 | |
| 5 |
|
Meta | 82.7 | 31 | 25 Jul 2026 |
| 6 | GD gemini-2.0-flash-thinking-exp-01-21 | Google DeepMind,Google | 82.5 | 62 | 25 Jul 2026 |
| 7 | GD gemini-2.0-flash-exp | Google DeepMind,Google | 81.9 | — | 25 Jul 2026 |
| 8 |
|
Alibaba | 81.8 | — | 25 Jul 2026 |
| 9 |
|
OpenAI | 81.5 | 63 | 25 Jul 2026 |
| 10 |
|
DeepSeek | 81.5 | 60 | 25 Jul 2026 |
| 11 |
|
Anthropic | 81.2 | 70 | 25 Jul 2026 |
| 12 |
|
80.6 | — | 25 Jul 2026 | |
| 13 |
|
DeepSeek | 80.5 | 68 | 25 Jul 2026 |
| 14 |
|
DeepSeek | 80.5 | 68 | 25 Jul 2026 |
| 15 |
|
78.3 | — | 25 Jul 2026 | |
| 16 | GD gemini-exp-1206 | Google DeepMind,Google | 77.3 | — | 25 Jul 2026 |
| 17 |
|
76.6 | — | 25 Jul 2026 | |
| 18 |
|
Perplexity | 76.2 | — | 25 Jul 2026 |
| 19 |
|
Alibaba | 75.3 | — | 25 Jul 2026 |
| 20 |
|
DeepSeek | 75.2 | 44 | 25 Jul 2026 |
| 21 |
|
74.9 | 37 | 25 Jul 2026 | |
| 22 |
|
74.9 | 37 | 25 Jul 2026 | |
| 23 |
|
OpenAI | 72.3 | 47 | 25 Jul 2026 |
| 24 |
|
DeepSeek | 69.9 | 52 | 25 Jul 2026 |
| 25 |
|
xAI | 69.6 | 38 | 25 Jul 2026 |
| 26 |
|
Anthropic | 69.3 | 40 | 25 Jul 2026 |
| 27 |
|
OpenAI | 68.6 | 21 | 25 Jul 2026 |
| 28 | L1 learnlm-1.5-pro-experimental | Unknown | 68.2 | — | 25 Jul 2026 |
| 29 |
|
Mistral | 67.9 | 33 | 25 Jul 2026 |
| 30 | Az amazon.nova-pro-v1:0 | Amazon | 67.1 | — | 25 Jul 2026 |
| 31 |
|
OpenAI | 65.4 | 56 | 25 Jul 2026 |
| 32 | D7 Dracarys2-72B-Instruct | Unknown | 65.2 | — | 25 Jul 2026 |
| 33 |
|
Anthropic | 63.9 | 30 | 25 Jul 2026 |
| 34 |
|
Mistral | 63.7 | 28 | 25 Jul 2026 |
| 35 | DL Dracarys2-Llama-3.1-70B-Instruct | Unknown | 63.2 | — | 25 Jul 2026 |
| 36 |
|
Anthropic | 61.9 | 23 | 25 Jul 2026 |
| 37 | AI OLMo-2-1124-13B-Instruct | Allen Institute for AI,University of Washington,New York University (NYU) | 60.6 | — | 25 Jul 2026 |
| 38 |
|
Mistral | 59.5 | 26 | 25 Jul 2026 |
| 39 |
|
Alibaba | 58.7 | — | 25 Jul 2026 |
| 40 |
|
Microsoft | 58.4 | 33 | 25 Jul 2026 |
| 41 |
|
58.1 | 19 | 25 Jul 2026 | |
| 42 |
|
58.1 | 19 | 25 Jul 2026 | |
| 43 | CC c4ai-command-r-plus-08-2024 | Cohere,Cohere for AI | 57.6 | — | 25 Jul 2026 |
| 44 |
|
OpenAI | 56.8 | 23 | 25 Jul 2026 |
| 45 |
|
DeepSeek | 55.7 | — | 25 Jul 2026 |
| 46 |
|
Cohere | 55.6 | — | 25 Jul 2026 |
| 47 | Az amazon.nova-lite-v1:0 | Amazon | 54.1 | — | 25 Jul 2026 |
| 48 |
|
52.6 | 10 | 25 Jul 2026 | |
| 49 | Az amazon.nova-micro-v1:0 | Amazon | 48.0 | — | 25 Jul 2026 |
| 50 |
|
Microsoft | 47.2 | — | 25 Jul 2026 |
| 51 |
|
Microsoft | 39.1 | — | 25 Jul 2026 |
| 52 |
|
Alibaba | 35.6 | — | 25 Jul 2026 |
The Known Good
Provenance. These figures are published by LiveBench and redistributed here under Apache-2.0.
The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is
min-max normalised against every other model holding the same evaluation, and nothing else is done to it.
Full licence detail is on attribution, and every score here is
in the CSV download.