Trends
One line per model creator, stepped at each release that raised that creator's best Known Good Index score. It holds flat until a better model ships, so a step is a real improvement rather than a rolling average. Click a legend entry to hide a creator.
Frontier quality over time
Known Good Index by creator, stepped at each release that raised their best score
Known Good Index combines seven published evaluations: GPQA Diamond, Mock AIME 2024–25, MATH Level 5, SWE-bench Verified, LiveBench, Humanity's Last Exam and Terminal-Bench, each normalised 0–100. Scores are ingested from their publishers — Epoch AI (CC-BY) and LiveBench — and we do not run these evaluations. Each step is a new release that raised that creator's best score.
| Creator | Steps plotted | First plotted release | Latest step | Best index |
|---|---|---|---|---|
| OpenAI | 7 | OpenAI: GPT-4o · May 2024 | gpt-5.5-pre-release · April 2026 | 97 |
| 5 | gemma-2-9b-it · June 2024 | Google: Gemini 3.1 Pro Preview · February 2026 | 97 | |
| DeepSeek | 3 | DeepSeek-R1-Distill-Qwen-32B · January 2025 | DeepSeek: DeepSeek V4 Pro 0423 · April 2026 | 93 |
| Alibaba | 4 | QwQ-32B · March 2025 | Qwen: Qwen3.7 Max · May 2026 | 93 |
| Anthropic | 5 | claude-3-opus-20240229 · February 2024 | Anthropic: Claude Opus 4.7 · April 2026 | 92 |
| Z.ai | 4 | Z.ai: GLM 4.7 · December 2025 | Z.ai: GLM 5.2 · June 2026 | 91 |
| xAI | 2 | grok-2-1212 · December 2024 | grok-4-0709 · July 2025 | 68 |