About The Known Good
Reference data and field notes for IT, security, and AI. We aggregate what other people publish about AI models — prices, benchmark results, throughput, human preference ratings — put it on one comparable footing, and show our working. We do not run our own evaluations, and every figure names its source.
A reference site, not a review site. The question it answers is "what does the published evidence say about this model, today, and where did that come from" — not "which model is best".
Everything is aggregated. Benchmark scores come from their publishers, prices and throughput from the marketplaces that serve the models, Elo from arenas that collect human votes. Our contribution is putting them on one scale, matching model identities across sources that disagree about names, and never publishing a number we cannot point back to a source.
Scheduled jobs ingest each source into a staging table, validate every record, sanity-check the row count against the last good run, and only then swap it into the live tables inside one transaction. A source that returns nothing is rejected and the live data is left untouched.
A build step then renders every page you see to flat HTML from a PostgreSQL database. Nothing is computed when you load a page — the interactive bits run in your browser against the same public JSON you can download. There are no accounts, no logins and nothing to sign in to.
and coding result
other people
speed and latency
| Source | Status | Rows written | Last run | Last success | Note |
|---|---|---|---|---|---|
blog |
ok | 20 | 25 Jul 2026 23:05 UTC | 25 Jul 2026 23:05 UTC | feed items=20 valid=20 rejected=0 duplicate_guid=0 | covers=20/20 (NULL cover renders the gradient) | sanity gate: staged 20 rows vs last good 20 | upserted 20 posts (new=0 refr... |
elo |
ok | 1,002 | 25 Jul 2026 14:33 UTC | 25 Jul 2026 14:33 UTC | lmarena/text: collapsed 735 duplicate overall row(s) into 378 models; tts-arena: ranks were 0-based (0..41) and were shifted to 1..42 to match every other arena_kind; staged 100... |
index |
ok | 1,443 | 25 Jul 2026 04:02 UTC | 25 Jul 2026 04:02 UTC | known_good: 61 models qualifying (>=1 each of science, mathematics, coding; of 6 evaluations) | coding: 19 models (>=2 of 3) | math: 166 models (>=1 of 2) | science: 169 models ... |
pricing |
ok | 345 | 25 Jul 2026 23:00 UTC | 25 Jul 2026 23:01 UTC | openrouter: 345 records, 345 valid, 0 rejected | models.dev: 172 providers, 5759 entries, 0 rejected | models.dev matched 345 of 345 models | staged: 345 models, price 340 canon... |
quality |
ok | 875 | 25 Jul 2026 04:02 UTC | 25 Jul 2026 04:02 UTC | attribution: Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved from https://epoch.ai/benchmarks. Used under CC-BY 4.0. attribution: LiveBench scores redis... |
What we will not do
- Claim we tested anything. We do not run evaluations. Every chart names the publisher whose measurement it shows.
- Publish cost per task. It needs token counts from an evaluation suite we do not run, so we publish blended price and explain the blend instead. See methodology.
- Take money to move a ranking. The index is equal-weighted, which leaves nothing to negotiate over.
- Guess. A missing figure renders as an em dash, never as a zero and never as an estimate.
- Run accounts. There is a newsletter. There is nothing to log into.
Contact and corrections
Corrections come first. If a figure looks wrong, start at the evaluation page behind it — every score links to the evaluation, and every evaluation links to the publisher's own page, so you can check our number against theirs in two clicks. If ours is the one that is wrong, tell us and we will fix it and say what changed in the changelog.
The project also writes up what it finds at Field Notes, which is the best place to reach whoever is arguing about a number this week.
Newsletter and privacy
The newsletter is the only thing we ask for an address for: new models, price moves, and the occasional argument. It is double opt-in — you confirm by email before we ever send anything — delivered through a third-party email provider, and used for nothing else. No accounts, no advertising, no third-party analytics tracking you across the site.
Provenance, in one line
Prices and provider telemetry from OpenRouter. Metadata from models.dev. Benchmark results from their publishers, including Epoch AI under CC-BY with the credit that licence requires. Elo from Arena and TTS Arena. Everything listed, licence by licence, on attribution, and downloadable in full from the data page.