Quality
Quality of leading AI models, aggregated from published independent evaluations
Known Good Index v1.2 incorporates 6 evaluations: GPQA Diamond, MATH / AIME, SWE-bench Verified, LiveBench, Humanity's Last Exam, Terminal-Bench. Normalised 0–100. Scores ingested from Epoch AI (CC-BY) and LiveBench.
Coding Index
Software engineering and agentic terminal ability
Coding Index, normalised 0–100 from the published evaluations that feed it. Ingested from Epoch AI (CC-BY) and LiveBench. We do not run these evaluations.
Image & Video Leaderboards
Elo from blind human preference votes, with 95% confidence intervals
Elo rating from blind pairwise preference votes. Ingested from Arena, updated daily. Higher is better.
Speech Leaderboards
Text-to-speech quality from blind listening preference votes
Elo rating from blind A/B listening tests. Ingested from TTS Arena, updated daily. Higher is better.
Capability Indices
Capability-specific indices built from the relevant evaluations
Mathematics Index, normalised 0–100 from the published evaluations that feed it. Ingested from Epoch AI (CC-BY) and LiveBench. We do not run these evaluations.
Quality Breakdown
Individual evaluations behind the index, as measured by their publishers
Graduate-level science questions written to be Google-proof. Aggregated from Epoch AI's AI Benchmarking Hub (gpqa_diamond.csv, column 'mean_score'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations. Higher is better.
Arena Elo
Human preference rating from blind pairwise votes
Elo rating from blind pairwise preference votes. Ingested from Arena, updated daily. Higher is better.
Openness Index
Weight availability and licence permissiveness
How many tracked models sit in each openness band. Scores weight availability, licence permissiveness, and commercial-use restrictions, derived from models.dev and Epoch AI metadata, not from evaluations. Bars are bands, not models, so bar colour carries no creator here.
Context Window
Maximum input tokens accepted
Maximum input context, from provider metadata via models.dev and OpenRouter. Higher is better.
Output Tokens
Maximum tokens a model will emit in a single response
Provider-declared maximum completion length, from provider metadata via models.dev and OpenRouter. Higher allows longer single-pass generation.
Price and Cost
Published API pricing, blended and by token type
USD per 1M tokens at a 3:1 input:output blend. Live from OpenRouter, cross-checked against models.dev. Lower is better. Excludes 18 models published at no cost (free and promotional tiers).
Speed & Latency
Throughput and time-to-first-token, measured across hosting providers
Median output tokens per second across all tracked providers. Derived from OpenRouter endpoint telemetry. Higher is better.
API Provider Performance
How inference hosts compare when serving the same model
Median tokens per second per host, serving the same model (gpt-oss-120b). Derived from OpenRouter endpoint telemetry. Higher is better.
Trends
How frontier model quality has moved over time, by creator
Known Good Index v1.2 incorporates 6 evaluations: GPQA Diamond, MATH / AIME, SWE-bench Verified, LiveBench, Humanity's Last Exam, Terminal-Bench. Each step is a new release that raised that creator's best score. Scores ingested from Epoch AI (CC-BY) and LiveBench.