The Known Good Updated 4 Sep 2026 Subscribe
Evaluations· Epoch AI· CC-BY 4.0· Known Good Index component

Terminal-Bench

Agentic terminal tasks. Epoch publishes one row per agent harness; the best score per model is kept, because the harness is not a property of the model. Aggregated from Epoch AI's AI Benchmarking Hub (terminalbench_external.csv, column 'Accuracy mean'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations.

Publisher's page Licence & attribution How we use this score
52 models scored Higher is better Published by Epoch AI CC-BY 4.0
Terminal-Bench

Terminal-Bench as measured and published by Epoch AI (CC-BY 4.0). Ingested, normalised only where it feeds an index, and never re-run by us. Higher is better.

💡 A lightbulb marks a model its provider publishes as a reasoning model
The Known Good
Every model we hold a Terminal-Bench score for
All 52 of 52 models
# Model Creator Terminal-Bench Known Good Index Measured
1 OpenAI: GPT-5.5 💡 OpenAI 84.7 93 4 Sep 2026
2 OpenAI: GPT-5.4 💡 OpenAI 81.8 91 4 Sep 2026
3 Anthropic: Claude Opus 4.7 💡 Anthropic 80.2 92 4 Sep 2026
4 Google: Gemini 3.1 Pro Preview 💡 Google 80.2 97 4 Sep 2026
5 Anthropic: Claude Opus 4.6 💡 Anthropic 79.8 89 4 Sep 2026
6 OpenAI: GPT-5.3-Codex 💡 OpenAI 78.4 4 Sep 2026
7 Gemini 3 Pro Google 69.4 4 Sep 2026
8 OpenAI: GPT-5.2-Codex 💡 OpenAI 66.5 4 Sep 2026
9 OpenAI: GPT-5.2 💡 OpenAI 64.9 81 4 Sep 2026
10 gemini-3-flash Google 64.3 4 Sep 2026
11 Anthropic: Claude Opus 4.5 💡 Anthropic 63.1 72 4 Sep 2026
12 gemini-3-pro-preview Google 61.8 84 25 Aug 2026
13 OpenAI: GPT-5.1-Codex-Mini 💡 OpenAI 61.6 4 Sep 2026
14 OpenAI: GPT-5.1-Codex-Max 💡 OpenAI 60.4 4 Sep 2026
15 claude-opus-4-5-20251101_128K Anthropic 59.1 4 Sep 2026
16 OpenAI: GPT-5.1-Codex 💡 OpenAI 57.8 4 Sep 2026
17 SpaceXAI: Grok 4.20 💡 xAI 57.3 4 Sep 2026
18 Anthropic: Claude Sonnet 4.6 💡 Anthropic 53.4 80 4 Sep 2026
19 Z.ai: GLM 5 💡 Z.ai 52.4 77 4 Sep 2026
20 Google: Gemini 3 Flash Preview 💡 Google 51.0 83 4 Sep 2026
21 OpenAI: GPT-5 💡 OpenAI 49.6 78 4 Sep 2026
22 OpenAI: GPT-5.1 💡 OpenAI 47.6 74 4 Sep 2026
23 Anthropic: Claude Sonnet 4.5 💡 Anthropic 46.5 4 Sep 2026
24 MiniMax: MiniMax M2.7 💡 MiniMax 45.1 4 Sep 2026
25 OpenAI: GPT-5 Codex 💡 OpenAI 44.3 4 Sep 2026
26 MoonshotAI: Kimi K2.5 💡 Moonshot AI 43.2 4 Sep 2026
27 claude-sonnet-4-5-20250929 Anthropic 42.8 66 4 Sep 2026
28 MiniMax: MiniMax M2.5 💡 MiniMax 42.7 4 Sep 2026
29 Claude 4.5 Sonnet Anthropic 42.7 4 Sep 2026
30 DeepSeek: DeepSeek V3.2 💡 DeepSeek 39.6 4 Sep 2026
31 Anthropic: Claude Opus 4.1 💡 Anthropic 38.0 50 4 Sep 2026
32 MiniMax: MiniMax M2.1 💡 MiniMax 36.6 4 Sep 2026
33 MoonshotAI: Kimi K2 Thinking 💡 Moonshot AI 35.7 4 Sep 2026
34 Anthropic: Claude Haiku 4.5 💡 Anthropic 35.5 4 Sep 2026
35 OpenAI: GPT-5 Mini 💡 OpenAI 34.8 67 4 Sep 2026
36 Z.ai: GLM 4.7 💡 Z.ai 33.4 69 4 Sep 2026
37 Google: Gemini 2.5 Pro 💡 Google 32.6 65 4 Sep 2026
38 MiniMax: MiniMax M2 💡 MiniMax 30.0 4 Sep 2026
39 claude-haiku-4-5-20251001 Anthropic 29.8 67 4 Sep 2026
40 Kimi-K2-Instruct Moonshot AI 27.8 4 Sep 2026
41 grok-4-0709 xAI 27.2 68 4 Sep 2026
42 Qwen 3 Coder 480B Alibaba 27.2 4 Sep 2026
43 grok-code-fast-1 xAI 25.8 4 Sep 2026
44 Qwen3-Coder-480B-A35B-Instruct Alibaba 25.4 4 Sep 2026
45 Z.ai: GLM 4.6 💡 Z.ai 24.5 4 Sep 2026
46 Qwen: Qwen3.6 35B A3B 💡 Alibaba 23.0 66 4 Sep 2026
47 OpenAI: GPT-5 Nano 💡 OpenAI 21.8 68 4 Sep 2026
48 OpenAI: gpt-oss-120b 💡 OpenAI 18.7 62 4 Sep 2026
49 Google: Gemini 2.5 Flash 💡 Google 17.1 4 Sep 2026
50 gemini-2.5-flash-preview-09-2025 Google 17.1 4 Sep 2026
51 Qwen: Qwen3.5-9B 💡 Alibaba 9.2 44 4 Sep 2026
52 OpenAI: gpt-oss-20b 💡 OpenAI 3.4 42 4 Sep 2026
The Known Good
Provenance. These figures are published by Epoch AI and redistributed here under CC-BY 4.0. The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is min-max normalised against every other model holding the same evaluation, and nothing else is done to it. This evaluation is one of the seven components of the Known Good Index. Full licence detail is on attribution, and every score here is in the CSV download.