Quality
Known Good Index and the capability indices, side by side
Each index is min-max normalised across tracked models and scaled 0–100 from published evaluations. We do not run these evaluations.
| Metric | DeepSeek: R1 Distill Llama 70B | OpenAI: GPT-4.1 |
|---|---|---|
| Known Good Index | 52 | 36 |
| Mathematics Index | 71 | 61 |
| Science Index | 52 | 66 |
| Reasoning Index | 59 | 36 |
| Agentic Index | — | 33 |
| Openness Index | 60 | 5 |
Quality Breakdown
Every evaluation both models report a published score for
Bar = DeepSeek: R1 Distill Llama 70B · marker = OpenAI: GPT-4.1 · scores ingested from Epoch AI (CC-BY 4.0)
Right-hand figures read DeepSeek: R1 Distill Llama 70B / OpenAI: GPT-4.1. DeepSeek: R1 Distill Llama 70B leads on 2 of 3, OpenAI: GPT-4.1 on 1.
Arena Elo
Human preference rating from blind pairwise votes
Elo from blind pairwise votes, ingested from Arena. Higher is better.
Neither model currently appears in an arena we ingest.
Price & Cost
Published API pricing by token type
USD per 1M tokens. Blended is a 3:1 input:output mix. Live from OpenRouter, cross-checked against models.dev. Lower is better.
| Metric | DeepSeek: R1 Distill Llama 70B | OpenAI: GPT-4.1 |
|---|---|---|
| Input $/1M | $0.80 | $2.00 |
| Output $/1M | $0.80 | $8.00 |
| Blended $/1M | $0.80 | $3.50 |
| Cache read $/1M | — | $0.500 |
Context Window
Maximum input accepted and maximum output emitted
Provider-declared limits, from metadata via models.dev and OpenRouter. Higher is better.
| Metric | DeepSeek: R1 Distill Llama 70B | OpenAI: GPT-4.1 |
|---|---|---|
| Context window | 8k | 1.05M |
| Max output tokens | 8k | 33k |
Speed & Latency
Median throughput and time to first token across hosting providers
Medians across all tracked providers, derived from OpenRouter endpoint telemetry.
| Metric | DeepSeek: R1 Distill Llama 70B | OpenAI: GPT-4.1 |
|---|---|---|
| Output tokens/s | 24 | 33 |
| Time to first token | 0.8s | 0.8s |
Providers
Which hosts serve each model, and at what price and speed
| Provider | DeepSeek: R1 Distill Llama 70B blended | DeepSeek: R1 Distill Llama 70B tok/s | OpenAI: GPT-4.1 blended | OpenAI: GPT-4.1 tok/s |
|---|---|---|---|---|
| Novita | $0.80 | 19 | — | — |
| OpenAI · first party | — | — | $3.50 | 59 |
| Azure · first party | — | — | $3.50 | 42 |
Per-host pricing and telemetry from OpenRouter. A dash means that host does not serve that model.