The Known Good Updated 25 Jul 2026

Reference data for AI models

Reference data and field notes for IT, security, and AI. Compare every major model on quality, price, and speed — aggregated from public benchmarks, refreshed continuously, traceable to origin.

Browse the leaderboardHow we source data
Highlights
Quality

Known Good Index · Higher is better · Ingested from Epoch AI (CC-BY) and LiveBench

Price

Blended USD per 1M tokens · Lower is better · Published rates via OpenRouter

Find your best-fit model
Tailored recommendations weighted for your priorities across quality, speed, and cost
Compare any two models head to head
Side-by-side quality, pricing, context, and provider availability
Download the raw data
Every figure as CSV or JSON, free, with source attribution — no account needed
Changeloglive
Price change · 25 Jul
Qwen: Qwen3 30B A3B Instruct 2507 price down 44%
Arena update · 25 Jul
Arena Elo refreshed: 1002 ratings across 7 arenas
Arena update · 25 Jul
Arena Elo refreshed: 1002 ratings across 7 arenas
Benchmark update · 25 Jul
Benchmark refresh: 875 scores across 13 evaluations
Price change · 25 Jul
MoonshotAI: Kimi K2.6 price down 15%
Arena update · 25 Jul
Arena Elo refreshed: 1002 ratings across 7 arenas
Benchmark update · 25 Jul
Benchmark refresh: 875 scores across 13 evaluations
New model tracked · 25 Jul
New model seen in arena leaderboards: deepseek-r1-0528
New model tracked · 25 Jul
New model seen in arena leaderboards: deepseek-r1
Arena update · 25 Jul
Arena Elo refreshed: 1002 ratings across 7 arenas
Benchmark update · 25 Jul
Benchmark refresh: 875 scores across 13 evaluations
New model tracked · 25 Jul
New model seen in arena leaderboards: chatgpt-image-latest-high-fidelity (20251216)
Jump to section

Quality

Quality of leading AI models, aggregated from published independent evaluations

Known Good IndexCoding IndexAgentic Index
Known Good Index

Known Good Index v1.2 incorporates 6 evaluations: GPQA Diamond, MATH / AIME, SWE-bench Verified, LiveBench, Humanity's Last Exam, Terminal-Bench. Normalised 0–100. Scores ingested from Epoch AI (CC-BY) and LiveBench.

22 of 61 ranked +Add model from specific provider
💡 Reasoning models are indicated by a lightbulb
The Known Good
Known Good Index v1.2 combines six public benchmarks, each min-max normalised across tracked models then weighted equally. We do not run these evaluations ourselves — scores are ingested from Epoch AI (CC-BY) and LiveBench. See the Index methodology for weights, normalisation, and per-eval provenance.

Coding Index

Software engineering and agentic terminal ability

Coding IndexSWE-bench VerifiedTerminal-BenchLiveBench CodingAgentic Index
Coding Index

Coding Index, normalised 0–100 from the published evaluations that feed it. Ingested from Epoch AI (CC-BY) and LiveBench. We do not run these evaluations.

19 of 19 ranked +Add model from specific provider
💡 Reasoning models are indicated by a lightbulb
The Known Good

Image & Video Leaderboards

Elo from blind human preference votes, with 95% confidence intervals

Text to ImageImage EditingText to VideoImage to Video
Text to Image Arena

Elo rating from blind pairwise preference votes. Ingested from Arena, updated daily. Higher is better.

12 of 79 models +Add model from specific provider
The Known Good

Speech Leaderboards

Text-to-speech quality from blind listening preference votes

Text to Speech Elo
Text to Speech Arena

Elo rating from blind A/B listening tests. Ingested from TTS Arena, updated daily. Higher is better.

12 of 42 models +Add model from specific provider
The Known Good

Capability Indices

Capability-specific indices built from the relevant evaluations

MathematicsScienceReasoningAgentic
Mathematics Index

Mathematics Index, normalised 0–100 from the published evaluations that feed it. Ingested from Epoch AI (CC-BY) and LiveBench. We do not run these evaluations.

22 of 166 ranked +Add model from specific provider
💡 Reasoning models are indicated by a lightbulb
The Known Good

Quality Breakdown

Individual evaluations behind the index, as measured by their publishers

GPQA DiamondMATH / AIMESWE-bench VerifiedLiveBenchHumanity's Last ExamTerminal-Bench
GPQA Diamond

Graduate-level science questions written to be Google-proof. Aggregated from Epoch AI's AI Benchmarking Hub (gpqa_diamond.csv, column 'mean_score'), used under CC-BY 4.0. The Known Good aggregates published results; it runs no evaluations. Higher is better.

16 of 908 models +Add model from specific provider
💡 Reasoning models are indicated by a lightbulb
The Known Good

Arena Elo

Human preference rating from blind pairwise votes

OverallStyle controlled
Text Arena

Elo rating from blind pairwise preference votes. Ingested from Arena, updated daily. Higher is better.

12 of 381 models +Add model from specific provider
The Known Good

Openness Index

Weight availability and licence permissiveness

Openness Index
Openness Index — distribution

How many tracked models sit in each openness band. Scores weight availability, licence permissiveness, and commercial-use restrictions, derived from models.dev and Epoch AI metadata, not from evaluations. Bars are bands, not models, so bar colour carries no creator here.

908 of 908 models +Add model from specific provider
Each bar is an openness band; the value is how many tracked models sit in itBands: 100 open weights, permissive · 90 open weights, permissive, no retrievable repository · 60 open weights, restricted use · 40 open weights, non-commercial · 5 proprietary
The Known Good

Context Window

Maximum input tokens accepted

Context windowMax output tokens
Context Window

Maximum input context, from provider metadata via models.dev and OpenRouter. Higher is better.

12 of 908 models +Add model from specific provider
The Known Good

Output Tokens

Maximum tokens a model will emit in a single response

Max output tokensContext window
Max Output Tokens

Provider-declared maximum completion length, from provider metadata via models.dev and OpenRouter. Higher allows longer single-pass generation.

12 of 908 models +Add model from specific provider
The Known Good

Price and Cost

Published API pricing, blended and by token type

Blended priceInput priceOutput priceCache hit priceQuality per dollar
Blended Price

USD per 1M tokens at a 3:1 input:output blend. Live from OpenRouter, cross-checked against models.dev. Lower is better. Excludes 18 models published at no cost (free and promotional tiers).

13 of 908 models +Add model from specific provider
The Known Good
We publish blended price rather than cost-per-task. Cost-per-task requires token counts from a proprietary evaluation suite we do not run. Blended price is computed from published rates and is fully reproducible — see pricing methodology.

Speed & Latency

Throughput and time-to-first-token, measured across hosting providers

Output speedLatency (TTFT)
Output Speed

Median output tokens per second across all tracked providers. Derived from OpenRouter endpoint telemetry. Higher is better.

12 of 908 models +Add model from specific provider
The Known Good

API Provider Performance

How inference hosts compare when serving the same model

Output speedLatencyPriceContext served
Output Speed by Provider: gpt-oss-120b

Median tokens per second per host, serving the same model (gpt-oss-120b). Derived from OpenRouter endpoint telemetry. Higher is better.

10 of 17 providers +Add model from specific provider
The Known Good
Field notes

Hugging Face Wasn't the Target. It Was in the Way.

21 Jul · Field note

Delaware County's Cyberattack and the Missing 911 Boundary

20 Jul · Field note

When the Router Exfiltrates Its Own Configuration

18 Jul · Field note

AI Worked Both Sides of the Ledger This Week. The Lesson Isn't "Patch Faster."

17 Jul · Field note

Pennsylvania Says Its Statewide 911 Disruption Wasn't a Cyberattack. The Alternative May Be Harder to Defend Against.

15 Jul · Field note

PamStealer Skips the Process Chains Defenders Watch. Not All of Them.

2 Jul · Field note

Your Agent Framework Is a Pile of API Keys on a Public IP

1 Jul · Field note

There's No Patch for FortiBleed. Public-Sector Networks Are Where That Hurts Most.

27 Jun · Field note

Four Failures, One Excuse: "Patch Faster" Was Never the Answer

23 Jun · Field note

Vibe Coding Isn't the Problem. Not Understanding the Stack Is.

20 Jun · Field note

I Handed Claude Code the Keys. Turns Out I'm Not the Only One Using Them.

16 Jun · Field note

The Circuit Nobody Could Find: How Exact-Match Searching Nearly Cost Me the Audit

12 Jun · Field note

The HTTP/2 Bomb Sat in Plain Sight for a Decade. An AI Just Had to Read the Code.

4 Jun · Field note

The Copilot Meter Didn't Raise the Price. It Showed You the Bill.

3 Jun · Field note

Designing AI for a Teen Discord Server Without Turning It Into a Surveillance Machine

29 May · Field note

WhatsApp Says No One Can Read Your Messages. A Federal Agent Spent 10 Months Disagreeing.

22 May · Field note

The Patch Queue Is the New Vulnerability

18 May · Field note

The Week the Toolchain Became the Kill Chain

17 May · Field note

ShinyHunters Didn't Breach 9,000 Schools. They Breached One Vendor. Your Institution Inherited the Rest.

8 May · Field note

The Floor Has Been Hit: Navigating the 2026 Systems Engineering Realignment

3 May · Field note
Every number is sourced
OpenRouter — pricing & provider speedmodels.dev — metadataEpoch AI — benchmarks (CC-BY)LiveBench — contamination-free evalsArena — text, image & video EloTTS Arena — speech Elo Methodology & attribution →
Get the field notes
New models, price moves, and the occasional argument. One email per meaningful update — no accounts, no tracking.