Quality
Known Good Index and the capability indices, side by side
Each index is min-max normalised across tracked models and scaled 0–100 from published evaluations. We do not run these evaluations.
| Metric | claude-haiku-4-5-20251001 | grok-4-0709 |
|---|---|---|
| Known Good Index | 56 | 67 |
| Mathematics Index | 82 | 84 |
| Science Index | 71 | 91 |
| Agentic Index | 30 | 27 |
| Openness Index | 5 | 5 |
Quality Breakdown
Every evaluation both models report a published score for
Bar = claude-haiku-4-5-20251001 · marker = grok-4-0709 · scores ingested from Epoch AI (CC-BY 4.0)
Right-hand figures read claude-haiku-4-5-20251001 / grok-4-0709. claude-haiku-4-5-20251001 leads on 1 of 3, grok-4-0709 on 2.
Arena Elo
Human preference rating from blind pairwise votes
Elo from blind pairwise votes, ingested from Arena. Higher is better.
| Metric | claude-haiku-4-5-20251001 | grok-4-0709 |
|---|---|---|
| Text Arena | 1,393 | 1,410 |
| Text Arena (style controlled) | 1,412 | 1,410 |
Price & Cost
Published API pricing by token type
USD per 1M tokens. Blended is a 3:1 input:output mix. Live from OpenRouter, cross-checked against models.dev. Lower is better.
| Metric | claude-haiku-4-5-20251001 | grok-4-0709 |
|---|---|---|
| Input $/1M | — | — |
| Output $/1M | — | — |
| Blended $/1M | — | — |
| Cache read $/1M | — | — |
Context Window
Maximum input accepted and maximum output emitted
Provider-declared limits, from metadata via models.dev and OpenRouter. Higher is better.
| Metric | claude-haiku-4-5-20251001 | grok-4-0709 |
|---|---|---|
| Context window | — | — |
| Max output tokens | — | — |
Speed & Latency
Median throughput and time to first token across hosting providers
Medians across all tracked providers, derived from OpenRouter endpoint telemetry.
| Metric | claude-haiku-4-5-20251001 | grok-4-0709 |
|---|---|---|
| Output tokens/s | — | — |
| Time to first token | — | — |
Providers
Which hosts serve each model, and at what price and speed
No hosting providers are currently tracked for either model.