The Known Good Updated 25 Jul 2026

Evaluations

Every number on this site traces back to an evaluation someone else published and someone else ran. This page lists all 13 of them, who publishes each one, and the licence the results are redistributed under.

13 of 13 evaluations
In the Known Good Index — 6 evaluations
GPQA Diamond
Epoch AI · CC-BY 4.0
Graduate-level science questions written to be Google-proof. Aggregated from Epoch AI's AI Benchmarking Hub (gpqa_diamond.csv, column 'mean_score'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations.
169 models scored Higher is better Index component
Humanity's Last Exam
Epoch AI · CC-BY 4.0
Expert-written questions spanning many specialist fields. Aggregated from Epoch AI's AI Benchmarking Hub (hle_external.csv, column 'Accuracy'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations.
45 models scored Higher is better Index component
LiveBench
LiveBench · Apache-2.0
Contamination-free objective benchmark, refreshed monthly. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'Global average'). The Known Good aggregates published results; it runs no evaluations.
52 models scored Higher is better Index component
MATH / AIME
Epoch AI · CC-BY 4.0
Competition mathematics, from the OTIS Mock AIME 2024-2025 set. Aggregated from Epoch AI's AI Benchmarking Hub (otis_mock_aime_2024_2025.csv, column 'mean_score'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations.
142 models scored Higher is better Index component
SWE-bench Verified
Epoch AI · CC-BY 4.0
Real GitHub issues resolved against a verified test harness. Aggregated from Epoch AI's AI Benchmarking Hub (swe_bench_verified.csv, column 'mean_score'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations.
32 models scored Higher is better Index component
Terminal-Bench
Epoch AI · CC-BY 4.0
Agentic terminal tasks. Epoch publishes one row per agent harness; the best score per model is kept, because the harness is not a property of the model. Aggregated from Epoch AI's AI Benchmarking Hub (terminalbench_external.csv, column 'Accuracy mean'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations.
55 models scored Higher is better Index component
Also tracked — 7 evaluations
LiveBench Coding
LiveBench · Apache-2.0
LiveBench coding subset. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'Coding average'). The Known Good aggregates published results; it runs no evaluations.
52 models scored Higher is better
LiveBench Data Analysis
LiveBench · Apache-2.0
LiveBench data-analysis subset. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'Data analysis average'). The Known Good aggregates published results; it runs no evaluations.
52 models scored Higher is better
LiveBench Instruction Following
LiveBench · Apache-2.0
LiveBench instruction-following subset. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'IF Average'). The Known Good aggregates published results; it runs no evaluations.
52 models scored Higher is better
LiveBench Language
LiveBench · Apache-2.0
LiveBench language subset. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'Language average'). The Known Good aggregates published results; it runs no evaluations.
52 models scored Higher is better
LiveBench Mathematics
LiveBench · Apache-2.0
LiveBench mathematics subset. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'Mathematics average'). The Known Good aggregates published results; it runs no evaluations.
52 models scored Higher is better
LiveBench Reasoning
LiveBench · Apache-2.0
LiveBench reasoning subset. Published by LiveBench (Apache-2.0) and redistributed by Epoch AI under CC-BY 4.0 in live_bench_external.csv (column 'Reasoning average'). The Known Good aggregates published results; it runs no evaluations.
52 models scored Higher is better
MATH Level 5
Epoch AI · CC-BY 4.0
Hardest tier of the MATH dataset. Aggregated from Epoch AI's AI Benchmarking Hub (math_level_5.csv, column 'mean_score'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations.
101 models scored Higher is better
We do not run any of these evaluations. Scores are ingested from the publishers listed above — Epoch AI (CC-BY, credit required) and LiveBench among them — and normalised only when they feed an index. The six evaluations marked Index component are the ones combined into the Known Good Index; a model needs at least one science, one mathematics and one coding result among them before we publish an index score for it. See methodology for the normalisation and attribution for every licence in full.