The Known Good Updated 25 Jul 2026
Evaluations· Epoch AI· CC-BY 4.0· Known Good Index component

SWE-bench Verified

Real GitHub issues resolved against a verified test harness. Aggregated from Epoch AI's AI Benchmarking Hub (swe_bench_verified.csv, column 'mean_score'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations.

Publisher's page Licence & attribution How we use this score
32 models scored Higher is better Published by Epoch AI CC-BY 4.0
SWE-bench Verified

SWE-bench Verified as measured and published by Epoch AI (CC-BY 4.0). Ingested, normalised only where it feeds an index, and never re-run by us. Higher is better.

16 of 32 models +Add model from specific provider
💡 Reasoning models are indicated by a lightbulb
The Known Good
Every model we hold a SWE-bench Verified score for
32 of 32 models
# Model Creator SWE-bench Verified Known Good Index Measured
1 Anthropic: Claude Opus 4.7 💡 Anthropic 83.5 94 25 Jul 2026
2 gpt-5.5-pre-release OpenAI 80.6 98 25 Jul 2026
3 Google: Gemini 3.5 Flash 💡 Google 79.3 95 25 Jul 2026
4 Anthropic: Claude Opus 4.6 💡 Anthropic 78.7 88 25 Jul 2026
5 Z.ai: GLM 5.2 💡 Z.ai 78.7 91 25 Jul 2026
6 DeepSeek: DeepSeek V4 Pro 💡 DeepSeek 77.6 93 25 Jul 2026
7 Qwen: Qwen3.7 Max 💡 Alibaba 77.3 93 25 Jul 2026
8 OpenAI: GPT-5.4 💡 OpenAI 76.9 90 25 Jul 2026
9 Qwen: Qwen3.6 Max Preview 💡 Alibaba 76.7 90 25 Jul 2026
10 MoonshotAI: Kimi K2.6 💡 Moonshot AI 76.7 93 25 Jul 2026
11 claude-opus-4-5-20251101 Anthropic 76.7 72 25 Jul 2026
12 Google: Gemini 3.1 Pro Preview Custom Tools 💡 Google 75.6 25 Jul 2026
13 Google: Gemini 3 Flash Preview 💡 Google 75.4 83 25 Jul 2026
14 Anthropic: Claude Sonnet 4.6 💡 Anthropic 75.2 80 25 Jul 2026
15 OpenAI: GPT-5.3-Codex 💡 OpenAI 74.8 25 Jul 2026
16 Z.ai: GLM 5.1 💡 Z.ai 74.2 88 25 Jul 2026
17 MoonshotAI: Kimi K2.5 💡 Moonshot AI 73.8 25 Jul 2026
18 OpenAI: GPT-5.2 💡 OpenAI 73.8 80 25 Jul 2026
19 OpenAI: GPT-5 💡 OpenAI 73.6 73 25 Jul 2026
20 claude-opus-4-1-20250805 Anthropic 73.3 49 25 Jul 2026
21 gemini-3-pro-preview Google 72.9 83 25 Jul 2026
22 Z.ai: GLM 5 💡 Z.ai 72.1 77 25 Jul 2026
23 claude-sonnet-4-5-20250929 Anthropic 71.3 59 25 Jul 2026
24 claude-opus-4-20250514 Anthropic 70.7 62 25 Jul 2026
25 OpenAI: GPT-5.1 💡 OpenAI 68.0 74 25 Jul 2026
26 OpenAI: GPT-5 Mini 💡 OpenAI 64.7 60 25 Jul 2026
27 OpenAI: o3 💡 OpenAI 62.3 67 25 Jul 2026
28 claude-3-7-sonnet-20250219 Anthropic 61.0 70 25 Jul 2026
29 Qwen: Qwen3.6 Plus 💡 Alibaba 57.9 78 25 Jul 2026
30 Google: Gemini 2.5 Pro 💡 Google 57.6 64 25 Jul 2026
31 OpenAI: GPT-4.1 OpenAI 48.5 36 25 Jul 2026
32 OpenAI: GPT-4o OpenAI 31.0 21 25 Jul 2026
The Known Good
Provenance. These figures are published by Epoch AI and redistributed here under CC-BY 4.0. The publisher's own page is here. We did not run this evaluation and we do not adjust the published numbers — where a score feeds an index it is min-max normalised against every other model holding the same evaluation, and nothing else is done to it. This evaluation is one of the six components of the Known Good Index. Full licence detail is on attribution, and every score here is in the CSV download.