Quality
Known Good Index and the capability indices, side by side
Each index is min-max normalised across tracked models and scaled 0–100 from published evaluations. We do not run these evaluations.
| Metric | DeepSeek: DeepSeek V4 Pro | OpenAI: GPT-5 |
|---|---|---|
| Known Good Index | 93 | 73 |
| Coding Index | — | 67 |
| Mathematics Index | 97 | 96 |
| Science Index | 94 | 90 |
| Reasoning Index | — | 71 |
| Agentic Index | 89 | 67 |
| Openness Index | 60 | 5 |
Quality Breakdown
Every evaluation both models report a published score for
Bar = DeepSeek: DeepSeek V4 Pro · marker = OpenAI: GPT-5 · scores ingested from Epoch AI (CC-BY 4.0)
Right-hand figures read DeepSeek: DeepSeek V4 Pro / OpenAI: GPT-5. DeepSeek: DeepSeek V4 Pro leads on 3 of 3, OpenAI: GPT-5 on 0.
Arena Elo
Human preference rating from blind pairwise votes
Elo from blind pairwise votes, ingested from Arena. Higher is better.
| Metric | DeepSeek: DeepSeek V4 Pro | OpenAI: GPT-5 |
|---|---|---|
| Text Arena | 1,449 | — |
| Text Arena (style controlled) | 1,457 | — |
Price & Cost
Published API pricing by token type
USD per 1M tokens. Blended is a 3:1 input:output mix. Live from OpenRouter, cross-checked against models.dev. Lower is better.
| Metric | DeepSeek: DeepSeek V4 Pro | OpenAI: GPT-5 |
|---|---|---|
| Input $/1M | $0.43 | $1.25 |
| Output $/1M | $0.87 | $10.00 |
| Blended $/1M | $0.54 | $3.44 |
| Cache read $/1M | $0.004 | $0.125 |
Context Window
Maximum input accepted and maximum output emitted
Provider-declared limits, from metadata via models.dev and OpenRouter. Higher is better.
| Metric | DeepSeek: DeepSeek V4 Pro | OpenAI: GPT-5 |
|---|---|---|
| Context window | 1.05M | 400k |
| Max output tokens | 384k | 128k |
Speed & Latency
Median throughput and time to first token across hosting providers
Medians across all tracked providers, derived from OpenRouter endpoint telemetry.
| Metric | DeepSeek: DeepSeek V4 Pro | OpenAI: GPT-5 |
|---|---|---|
| Output tokens/s | 45 | 57 |
| Time to first token | 2.0s | 2.6s |
Providers
Which hosts serve each model, and at what price and speed
| Provider | DeepSeek: DeepSeek V4 Pro blended | DeepSeek: DeepSeek V4 Pro tok/s | OpenAI: GPT-5 blended | OpenAI: GPT-5 tok/s |
|---|---|---|---|---|
| BaseTen | $2.17 | 110 | — | — |
| Fireworks | $2.17 | 66 | — | — |
| Novita | $1.46 | 66 | — | — |
| DeepInfra | $1.62 | 65 | — | — |
| GMICloud | $0.85 | 62 | — | — |
| Venice | $2.06 | 61 | — | — |
| Baidu · first party | $0.78 | 58 | — | — |
| Together | $2.17 | 53 | — | — |
| Alibaba · first party | $1.77 | 52 | — | — |
| AtlasCloud | $2.10 | 46 | — | — |
| Wafer | $1.50 | 45 | — | — |
| DeepSeek · first party | $0.54 | 44 | — | — |
| SiliconFlow | $1.91 | 41 | — | — |
| StreamLake | $0.84 | 41 | — | — |
| CoreWeave | $2.17 | 28 | — | — |
| Ionstream | $1.41 | 20 | — | — |
| DigitalOcean | $1.74 | 6 | — | — |
| Parasail | $2.17 | — | — | — |
| OpenAI · first party | — | — | $3.44 | 84 |
| Azure · first party | — | — | $3.44 | 40 |
Per-host pricing and telemetry from OpenRouter. A dash means that host does not serve that model.