The Known Good Updated 9 Sep 2026 Subscribe

Reference data for AI models

Reference data and field notes for IT, security, and AI. Compare every major model on quality, price, and speed — aggregated from public benchmarks, every figure traceable to its source and dated where a date exists.

Browse the leaderboardHow we source data
Highlights
Quality

Known Good Index · Higher is better · Ingested from Epoch AI (CC-BY) and LiveBench

Price

Blended USD per 1M tokens · Lower is better · Published rates via OpenRouter

Find your best-fit model
Tailored recommendations weighted for your priorities across quality, speed, and cost
Compare any two models head to head
Side-by-side quality, pricing, context, and provider availability
Download the raw data
Every figure as CSV or JSON, free, with source attribution — no account needed
Changeloglatest 12
Price change · 9 Sep
DeepSeek: DeepSeek V4 Pro 0423 price down 9%
Price change · 9 Sep
DeepSeek: DeepSeek V4 Pro 0813 price down 37%
Price change · 9 Sep
Qwen: Qwen3 30B A3B Instruct 2507 price down 41%
Price change · 9 Sep
Qwen: Qwen3 30B A3B Instruct 2507 price up 69%
Price change · 9 Sep
DeepSeek: DeepSeek V4 Pro 0813 price up 81%
Price change · 9 Sep
Qwen: Qwen3 235B A22B Instruct 2507 price up 88%
New model tracked · 9 Sep
New model seen in arena leaderboards: gpt-image-2.5-sunburst
New model tracked · 9 Sep
New model seen in arena leaderboards: gpt-image-2.5-flare
Arena update · 9 Sep
Arena Elo refreshed: 1062 ratings across 7 arenas
Benchmark update · 9 Sep
Benchmark refresh: 1051 scores across 13 evaluations
Price change · 9 Sep
MoonshotAI Kimi Latest price down 6%
Price change · 9 Sep
DeepSeek: DeepSeek V4 Flash 0423 price up 9%
Jump to section

Quality

Quality of leading AI models, aggregated from published independent evaluations

Known Good IndexCoding Index
Known Good Index

Known Good Index v1.3 incorporates 7 evaluations: GPQA Diamond, Mock AIME 2024–25, MATH Level 5, SWE-bench Verified, LiveBench, Humanity's Last Exam, Terminal-Bench. Normalised 0–100. Scores ingested from Epoch AI (CC-BY) and LiveBench.

💡 A lightbulb marks a model its provider publishes as a reasoning model
The Known Good
Known Good Index v1.3 combines seven public benchmarks, each min-max normalised across tracked models then weighted equally. We do not run these evaluations ourselves — scores are ingested from Epoch AI (CC-BY) and LiveBench. See the Index methodology for weights, normalisation, and per-eval provenance.

Coding Index

Software engineering and agentic terminal ability

Coding IndexSWE-bench VerifiedTerminal-BenchLiveBench Coding
Coding Index

Coding Index, built from SWE-bench Verified, LiveBench Coding and Terminal-Bench, each min-max normalised across tracked models and averaged, scaled 0–100. Ingested from Epoch AI (CC-BY 4.0) and LiveBench (Apache-2.0). We do not run these evaluations.

💡 A lightbulb marks a model its provider publishes as a reasoning model
The Known Good

Image & Video Leaderboards

Elo from blind human preference votes, with 95% confidence intervals

Text to ImageImage EditingText to VideoImage to Video

Speech Leaderboards

Text-to-speech quality from blind listening preference votes

Text to Speech Arena

Elo rating from blind A/B listening tests. Ingested from TTS Arena, updated daily. Higher is better.

The Known Good

Capability Indices

Capability-specific indices built from the relevant evaluations

MathematicsReasoning
Mathematics Index

Mathematics Index, built from Mock AIME 2024–25 and MATH Level 5, each min-max normalised across tracked models and averaged, scaled 0–100. Ingested from Epoch AI (CC-BY 4.0). We do not run these evaluations.

💡 A lightbulb marks a model its provider publishes as a reasoning model
The Known Good

Quality Breakdown

Individual evaluations behind the index, as measured by their publishers

GPQA DiamondMock AIME 2024–25MATH Level 5SWE-bench VerifiedLiveBenchHumanity's Last ExamTerminal-Bench
GPQA Diamond

Graduate-level science questions written to be Google-proof. Aggregated from Epoch AI's AI Benchmarking Hub (gpqa_diamond.csv, column 'mean_score'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations. Higher is better.

💡 A lightbulb marks a model its provider publishes as a reasoning model
The Known Good

Arena Elo

Human preference rating from blind pairwise votes

OverallStyle controlled
Text Arena

Elo rating from blind pairwise preference votes. Ingested from Arena, updated daily. Higher is better.

The Known Good

Openness Index

Weight availability and licence permissiveness

Openness Index — distribution

How many tracked models sit in each openness band. Scores weight availability, licence permissiveness, and commercial-use restrictions, derived from models.dev and Epoch AI metadata, not from evaluations. Bars are bands, not models, so bar colour carries no creator here.

All 5 bands · 1,109 models classified
Each bar is an openness band; the value is how many tracked models sit in itBands: 100 open weights, permissive · 90 open weights, permissive, no retrievable repository · 60 open weights, restricted use · 40 open weights, non-commercial · 5 proprietary
The Known Good

Context Window

Maximum input tokens accepted

Context windowMax output tokens
Context Window

Maximum input context, from provider metadata via models.dev and OpenRouter. Higher is better.

The Known Good

Output Tokens

Maximum tokens a model will emit in a single response

Max output tokensContext window
Max Output Tokens

Provider-declared maximum completion length, from provider metadata via models.dev and OpenRouter. Higher allows longer single-pass generation.

The Known Good

Price and Cost

Published API pricing, blended and by token type

Blended priceInput priceOutput priceCache hit priceQuality per dollar
Blended Price

USD per 1M tokens at a 3:1 input:output blend. Live from OpenRouter, cross-checked against models.dev. Lower is better. Ranked against the 450 models priced above zero, which is the population the chart draws; excludes 32 published at no cost (free and promotional tiers).

The Known Good
We publish blended price rather than cost-per-task. Cost-per-task requires token counts from a proprietary evaluation suite we do not run. Blended price is computed from published rates and is fully reproducible — see pricing methodology.

Speed & Latency

Throughput and time-to-first-token, measured across hosting providers

Output speedLatency (TTFT)

API Provider Performance

How inference hosts compare when serving the same model

Output speedLatencyPriceContext served
Output Speed by Provider: gpt-oss-120b

Tokens per second per host, the most recent figure OpenRouter publishes for each host, for hosts serving the same model (gpt-oss-120b). Higher is better.

The Known Good
Field notes

Stop Saying Vibe Coding Is Easy

26 Aug · Field note

New York's 911 Failure Didn't Look Like a Failure. That's the Problem.

20 Aug · Field note

What the Teams Vishing Research Looks Like From a 911 Center

31 Jul · Field note

Hugging Face Wasn't the Target. It Was in the Way.

21 Jul · Field note

Delaware County's Cyberattack and the Missing 911 Boundary

20 Jul · Field note

When the Router Exfiltrates Its Own Configuration

18 Jul · Field note

AI Worked Both Sides of the Ledger This Week. The Lesson Isn't "Patch Faster."

17 Jul · Field note

Pennsylvania Says Its Statewide 911 Disruption Wasn't a Cyberattack. The Alternative May Be Harder to Defend Against.

15 Jul · Field note

PamStealer Skips the Process Chains Defenders Watch. Not All of Them.

2 Jul · Field note

Your Agent Framework Is a Pile of API Keys on a Public IP

1 Jul · Field note

There's No Patch for FortiBleed. Public-Sector Networks Are Where That Hurts Most.

27 Jun · Field note

Four Failures, One Excuse: "Patch Faster" Was Never the Answer

23 Jun · Field note

Vibe Coding Isn't the Problem. Not Understanding the Stack Is.

20 Jun · Field note

I Handed Claude Code the Keys. Turns Out I'm Not the Only One Using Them.

16 Jun · Field note

The Circuit Nobody Could Find: How Exact-Match Searching Nearly Cost Me the Audit

12 Jun · Field note

The HTTP/2 Bomb Sat in Plain Sight for a Decade. An AI Just Had to Read the Code.

4 Jun · Field note

The Copilot Meter Didn't Raise the Price. It Showed You the Bill.

3 Jun · Field note

Designing AI for a Teen Discord Server Without Turning It Into a Surveillance Machine

29 May · Field note

WhatsApp Says No One Can Read Your Messages. A Federal Agent Spent 10 Months Disagreeing.

22 May · Field note

The Patch Queue Is the New Vulnerability

18 May · Field note

The Week the Toolchain Became the Kill Chain

17 May · Field note

ShinyHunters Didn't Breach 9,000 Schools. They Breached One Vendor. Your Institution Inherited the Rest.

8 May · Field note

The Floor Has Been Hit: Navigating the 2026 Systems Engineering Realignment

3 May · Field note
Every number is sourced
OpenRouter — pricing & provider speedmodels.dev — metadataEpoch AI — benchmarks (CC-BY)LiveBench — contamination-free evalsArena — text, image & video EloTTS Arena — speech Elo Methodology & attribution →
The Known Good News
A daily brief on AI — new models, price moves, security incidents, and the occasional argument. No accounts, no personal tracking, one-click unsubscribe.