The Known Good Updated 25 Jul 2026

About The Known Good

Reference data and field notes for IT, security, and AI. We aggregate what other people publish about AI models — prices, benchmark results, throughput, human preference ratings — put it on one comparable footing, and show our working. We do not run our own evaluations, and every figure names its source.

What this is

A reference site, not a review site. The question it answers is "what does the published evidence say about this model, today, and where did that come from" — not "which model is best".

Everything is aggregated. Benchmark scores come from their publishers, prices and throughput from the marketplaces that serve the models, Elo from arenas that collect human votes. Our contribution is putting them on one scale, matching model identities across sources that disagree about names, and never publishing a number we cannot point back to a source.

How it works

Scheduled jobs ingest each source into a staging table, validate every record, sanity-check the row count against the last good run, and only then swap it into the live tables inside one transaction. A source that returns nothing is rejected and the live data is left untouched.

A build step then renders every page you see to flat HTML from a PostgreSQL database. Nothing is computed when you load a page — the interactive bits run in your browser against the same public JSON you can download. There are no accounts, no logins and nothing to sign in to.

Coverage
Models tracked
908
from 83 model creators
With an index score
61
models with a science, mathematics
and coding result
Evaluations ingested
13
published by other people, run by
other people
Hosting providers
72
endpoints tracked for price,
speed and latency
Data freshness — last run of every source
Source Status Rows written Last run Last success Note
blog ok 20 25 Jul 2026 23:05 UTC25 Jul 2026 23:05 UTC feed items=20 valid=20 rejected=0 duplicate_guid=0 | covers=20/20 (NULL cover renders the gradient) | sanity gate: staged 20 rows vs last good 20 | upserted 20 posts (new=0 refr...
elo ok 1,002 25 Jul 2026 14:33 UTC25 Jul 2026 14:33 UTC lmarena/text: collapsed 735 duplicate overall row(s) into 378 models; tts-arena: ranks were 0-based (0..41) and were shifted to 1..42 to match every other arena_kind; staged 100...
index ok 1,443 25 Jul 2026 04:02 UTC25 Jul 2026 04:02 UTC known_good: 61 models qualifying (>=1 each of science, mathematics, coding; of 6 evaluations) | coding: 19 models (>=2 of 3) | math: 166 models (>=1 of 2) | science: 169 models ...
pricing ok 345 25 Jul 2026 23:00 UTC25 Jul 2026 23:01 UTC openrouter: 345 records, 345 valid, 0 rejected | models.dev: 172 providers, 5759 entries, 0 rejected | models.dev matched 345 of 345 models | staged: 345 models, price 340 canon...
quality ok 875 25 Jul 2026 04:02 UTC25 Jul 2026 04:02 UTC attribution: Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved from https://epoch.ai/benchmarks. Used under CC-BY 4.0. attribution: LiveBench scores redis...
The Known Good
Reading this table. ok means the run validated and was swapped into the live tables. rejected means the run came back empty or implausibly small and was thrown away — the live data is the previous good run, unchanged, and the pages are stale rather than wrong. That is the trade we want: bad upstream data must never destroy good data. If a source has been rejected for a while, the pages fed by it will say so rather than quietly filling the gap.

What we will not do

  • Claim we tested anything. We do not run evaluations. Every chart names the publisher whose measurement it shows.
  • Publish cost per task. It needs token counts from an evaluation suite we do not run, so we publish blended price and explain the blend instead. See methodology.
  • Take money to move a ranking. The index is equal-weighted, which leaves nothing to negotiate over.
  • Guess. A missing figure renders as an em dash, never as a zero and never as an estimate.
  • Run accounts. There is a newsletter. There is nothing to log into.

Contact and corrections

Corrections come first. If a figure looks wrong, start at the evaluation page behind it — every score links to the evaluation, and every evaluation links to the publisher's own page, so you can check our number against theirs in two clicks. If ours is the one that is wrong, tell us and we will fix it and say what changed in the changelog.

The project also writes up what it finds at Field Notes, which is the best place to reach whoever is arguing about a number this week.

Newsletter and privacy

The newsletter is the only thing we ask for an address for: new models, price moves, and the occasional argument. It is double opt-in — you confirm by email before we ever send anything — delivered through a third-party email provider, and used for nothing else. No accounts, no advertising, no third-party analytics tracking you across the site.

Provenance, in one line

Prices and provider telemetry from OpenRouter. Metadata from models.dev. Benchmark results from their publishers, including Epoch AI under CC-BY with the credit that licence requires. Elo from Arena and TTS Arena. Everything listed, licence by licence, on attribution, and downloadable in full from the data page.