The Known Good Updated 4 Sep 2026 Subscribe
Evaluations· Epoch AI· CC-BY 4.0· Known Good Index component

SWE-bench Verified

Real GitHub issues resolved against a verified test harness. Aggregated from Epoch AI's AI Benchmarking Hub (swe_bench_verified.csv, column 'mean_score'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations.

Publisher's page Licence & attribution How we use this score
32 models scored Higher is better Published by Epoch AI CC-BY 4.0
SWE-bench Verified

SWE-bench Verified as measured and published by Epoch AI (CC-BY 4.0). Ingested, normalised only where it feeds an index, and never re-run by us. Higher is better.

💡 A lightbulb marks a model its provider publishes as a reasoning model
The Known Good
Every model we hold a SWE-bench Verified score for
All 32 of 32 models
# Model Creator SWE-bench Verified Known Good Index Measured
1 Anthropic: Claude Opus 4.7 💡 Anthropic 83.5 92 4 Sep 2026
2 gpt-5.5-pre-release OpenAI 80.6 97 4 Sep 2026
3 Google: Gemini 3.5 Flash 💡 Google 79.3 95 4 Sep 2026
4 Anthropic: Claude Opus 4.6 💡 Anthropic 78.7 89 4 Sep 2026
5 Z.ai: GLM 5.2 💡 Z.ai 78.7 91 4 Sep 2026
6 DeepSeek: DeepSeek V4 Pro 0423 💡 DeepSeek 77.6 93 4 Sep 2026
7 Qwen: Qwen3.7 Max 💡 Alibaba 77.3 93 4 Sep 2026
8 OpenAI: GPT-5.4 💡 OpenAI 76.9 91 4 Sep 2026
9 Qwen: Qwen3.6 Max Preview 💡 Alibaba 76.7 89 4 Sep 2026
10 MoonshotAI: Kimi K2.6 💡 Moonshot AI 76.7 92 4 Sep 2026
11 Anthropic: Claude Opus 4.5 💡 Anthropic 76.7 72 4 Sep 2026
12 Google: Gemini 3.1 Pro Preview Custom Tools 💡 Google 75.6 4 Sep 2026
13 Google: Gemini 3 Flash Preview 💡 Google 75.4 83 4 Sep 2026
14 Anthropic: Claude Sonnet 4.6 💡 Anthropic 75.2 80 4 Sep 2026
15 OpenAI: GPT-5.3-Codex 💡 OpenAI 74.8 4 Sep 2026
16 Z.ai: GLM 5.1 💡 Z.ai 74.2 90 4 Sep 2026
17 MoonshotAI: Kimi K2.5 💡 Moonshot AI 73.8 4 Sep 2026
18 OpenAI: GPT-5.2 💡 OpenAI 73.8 81 4 Sep 2026
19 OpenAI: GPT-5 💡 OpenAI 73.6 78 4 Sep 2026
20 Anthropic: Claude Opus 4.1 💡 Anthropic 73.3 50 4 Sep 2026
21 gemini-3-pro-preview Google 72.9 84 4 Sep 2026
22 Z.ai: GLM 5 💡 Z.ai 72.1 77 4 Sep 2026
23 claude-sonnet-4-5-20250929 Anthropic 71.3 66 4 Sep 2026
24 claude-opus-4-20250514 Anthropic 70.7 68 4 Sep 2026
25 OpenAI: GPT-5.1 💡 OpenAI 68.0 74 4 Sep 2026
26 OpenAI: GPT-5 Mini 💡 OpenAI 64.7 67 4 Sep 2026
27 OpenAI: o3 💡 OpenAI 62.3 74 4 Sep 2026
28 claude-3-7-sonnet-20250219 Anthropic 61.0 74 4 Sep 2026
29 Qwen: Qwen3.6 Plus 💡 Alibaba 57.9 79 4 Sep 2026
30 Google: Gemini 2.5 Pro 💡 Google 57.6 65 4 Sep 2026
31 OpenAI: GPT-4.1 OpenAI 48.5 46 4 Sep 2026
32 OpenAI: GPT-4o OpenAI 31.0 27 4 Sep 2026
The Known Good
Provenance. These figures are published by Epoch AI and redistributed here under CC-BY 4.0. The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is min-max normalised against every other model holding the same evaluation, and nothing else is done to it. This evaluation is one of the seven components of the Known Good Index. Full licence detail is on attribution, and every score here is in the CSV download.