Quality
Known Good Index and the capability indices, side by side
Each index is min-max normalised across tracked models and scaled 0–100 from published evaluations. We do not run these evaluations.
| Metric | Microsoft: Phi 4 | Z.ai: GLM 5 |
|---|---|---|
| Known Good Index | 33 | 77 |
| Coding Index | — | 67 |
| Mathematics Index | 39 | 80 |
| Science Index | 53 | 92 |
| Reasoning Index | 47 | — |
| Agentic Index | — | 67 |
| Openness Index | 60 | 60 |
Quality Breakdown
Every evaluation both models report a published score for
Bar = Microsoft: Phi 4 · marker = Z.ai: GLM 5 · scores ingested from Epoch AI (CC-BY 4.0)
Right-hand figures read Microsoft: Phi 4 / Z.ai: GLM 5. Microsoft: Phi 4 leads on 0 of 2, Z.ai: GLM 5 on 2.
Arena Elo
Human preference rating from blind pairwise votes
Elo from blind pairwise votes, ingested from Arena. Higher is better.
| Metric | Microsoft: Phi 4 | Z.ai: GLM 5 |
|---|---|---|
| Text Arena | 1,217 | 1,445 |
| Text Arena (style controlled) | 1,256 | 1,457 |
Price & Cost
Published API pricing by token type
USD per 1M tokens. Blended is a 3:1 input:output mix. Live from OpenRouter, cross-checked against models.dev. Lower is better.
| Metric | Microsoft: Phi 4 | Z.ai: GLM 5 |
|---|---|---|
| Input $/1M | $0.07 | $0.95 |
| Output $/1M | $0.14 | $2.55 |
| Blended $/1M | $0.09 | $1.35 |
| Cache read $/1M | — | $0.200 |
Context Window
Maximum input accepted and maximum output emitted
Provider-declared limits, from metadata via models.dev and OpenRouter. Higher is better.
| Metric | Microsoft: Phi 4 | Z.ai: GLM 5 |
|---|---|---|
| Context window | 16k | 205k |
| Max output tokens | 16k | 131k |
Speed & Latency
Median throughput and time to first token across hosting providers
Medians across all tracked providers, derived from OpenRouter endpoint telemetry.
| Metric | Microsoft: Phi 4 | Z.ai: GLM 5 |
|---|---|---|
| Output tokens/s | 55 | 38 |
| Time to first token | 0.2s | 2.5s |
Providers
Which hosts serve each model, and at what price and speed
| Provider | Microsoft: Phi 4 blended | Microsoft: Phi 4 tok/s | Z.ai: GLM 5 blended | Z.ai: GLM 5 tok/s |
|---|---|---|---|---|
| DeepInfra | $0.09 | 38 | $0.97 | 29 |
| Amazon Bedrock · first party | — | — | $1.55 | 116 |
| GMICloud | — | — | $0.93 | 60 |
| Phala | — | — | $1.77 | 57 |
| Chutes | — | — | $1.35 | 49 |
| Venice | — | — | $1.55 | 40 |
| Baidu · first party | — | — | $1.08 | 39 |
| Parasail | — | — | $1.55 | 39 |
| AtlasCloud | — | — | $1.50 | 36 |
| Novita | — | — | $1.55 | 36 |
| StreamLake | — | — | $0.93 | 36 |
| Z.AI · first party | — | — | $1.55 | 33 |
| SiliconFlow | — | — | $1.35 | 30 |
| DigitalOcean | — | — | $1.16 | 8 |
Per-host pricing and telemetry from OpenRouter. A dash means that host does not serve that model.