Quality
Known Good Index and the capability indices, side by side
Each index is min-max normalised across tracked models and scaled 0–100 from published evaluations. We do not run these evaluations.
| Metric | DeepSeek: DeepSeek V3 0324 | gemini-2.0-flash-thinking-exp-01-21 |
|---|---|---|
| Known Good Index | 60 | 62 |
| Mathematics Index | 57 | 58 |
| Science Index | 67 | 54 |
| Reasoning Index | 65 | 66 |
| Openness Index | 100 | 5 |
Quality Breakdown
Every evaluation both models report a published score for
Bar = DeepSeek: DeepSeek V3 0324 · marker = gemini-2.0-flash-thinking-exp-01-21 · scores ingested from Epoch AI (CC-BY 4.0), LiveBench (Apache-2.0)
Right-hand figures read DeepSeek: DeepSeek V3 0324 / gemini-2.0-flash-thinking-exp-01-21. DeepSeek: DeepSeek V3 0324 leads on 3 of 9, gemini-2.0-flash-thinking-exp-01-21 on 6.
Arena Elo
Human preference rating from blind pairwise votes
Elo from blind pairwise votes, ingested from Arena. Higher is better.
| Metric | DeepSeek: DeepSeek V3 0324 | gemini-2.0-flash-thinking-exp-01-21 |
|---|---|---|
| Text Arena | 1,375 | — |
| Text Arena (style controlled) | 1,396 | — |
Price & Cost
Published API pricing by token type
USD per 1M tokens. Blended is a 3:1 input:output mix. Live from OpenRouter, cross-checked against models.dev. Lower is better.
| Metric | DeepSeek: DeepSeek V3 0324 | gemini-2.0-flash-thinking-exp-01-21 |
|---|---|---|
| Input $/1M | $0.27 | — |
| Output $/1M | $1.12 | — |
| Blended $/1M | $0.48 | — |
| Cache read $/1M | $0.135 | — |
Context Window
Maximum input accepted and maximum output emitted
Provider-declared limits, from metadata via models.dev and OpenRouter. Higher is better.
| Metric | DeepSeek: DeepSeek V3 0324 | gemini-2.0-flash-thinking-exp-01-21 |
|---|---|---|
| Context window | 164k | — |
| Max output tokens | 66k | — |
Speed & Latency
Median throughput and time to first token across hosting providers
Medians across all tracked providers, derived from OpenRouter endpoint telemetry.
| Metric | DeepSeek: DeepSeek V3 0324 | gemini-2.0-flash-thinking-exp-01-21 |
|---|---|---|
| Output tokens/s | 26 | — |
| Time to first token | 1.1s | — |
Providers
Which hosts serve each model, and at what price and speed
| Provider | DeepSeek: DeepSeek V3 0324 blended | DeepSeek: DeepSeek V3 0324 tok/s | gemini-2.0-flash-thinking-exp-01-21 blended | gemini-2.0-flash-thinking-exp-01-21 tok/s |
|---|---|---|---|---|
| Crusoe | $0.75 | 37 | — | — |
| Novita | $0.48 | 30 | — | — |
| DeepInfra | $0.41 | 25 | — | — |
| SiliconFlow | $0.44 | 21 | — | — |
Per-host pricing and telemetry from OpenRouter. A dash means that host does not serve that model.