stealth/ox-alpha on 2026-08-28. The benchmark context below is historical; no reproducible Ox Alpha score was ever published. See what the removal means.Ox Alpha Benchmarks: Where the Z.ai GLM Family Sits
Z.ai has confirmed that Ox Alpha is a new GLM-family iteration. The exact checkpoint is still not published, so the honest market comparison is a benchmark snapshot of the named GLM-5.3 baseline and its closest peers—not a fabricated Ox Alpha score.
The benchmark status that matters
Ox Alpha has not published a reproducible benchmark score. The responsible comparison is against the GLM-5.3 family and live market baselines, while waiting for the released weights and an independent rerun.
Artificial Analysis market snapshot
The table below uses the Artificial Analysis composite fields exposed in OpenRouter's live model API. Intelligence, Coding, and Agentic are index scores, not percentages of tasks solved. A dash means no score was published for that model in the live snapshot.
| Model | Intelligence | Coding | Agentic | Input / 1M | Output / 1M |
|---|---|---|---|---|---|
| Claude Opus 5Anthropic | 63.1 | 78.0 | 59.2 | $5.00 | $25.00 |
| Claude Fable 5Anthropic | 62.1 | 76.5 | 56.6 | $10.00 | $50.00 |
| GPT-5.6 SolOpenAI | 60.9 | 77.4 | 57.8 | $5.00 | $30.00 |
| Grok 4.6 FlagshipxAI | 60.9 | 76.8 | 58.7 | $2.00 | $6.00 |
| Kimi K3 (Moonshot)Moonshot AI | 59.7 | 76.2 | 54.3 | $3.00 | $15.00 |
| GLM-5.3 (Z.ai)Z.ai (Zhipu) | 59.5 | 74.8 | 59.1 | $1.40 | $4.40 |
| GLM-5 (Z.ai)Z.ai (Zhipu) | 59.5 | 74.8 | 59.1 | $1.00 | $3.20 |
| Qwen 3.8 Max (2.4T MoE)Alibaba Qwen | 58.1 | 71.8 | 58.4 | $2.00 | $6.00 |
| GLM-5.3-Flash (Z.ai)Z.ai (Zhipu) | 57.5 | 71.5 | 58.2 | $0.075 | $0.25 |
| Claude Opus 4.8 (OpenRouter)Anthropic | 57.3 | 74.3 | 49.4 | $5.00 | $25.00 |
| GPT-5.6 TerraOpenAI | 56.6 | 76.7 | 50.2 | $2.00 | $12.00 |
| GPT-5.5 StandardOpenAI | 56.3 | 74.9 | 47.4 | $5.00 | $30.00 |
| Gemini 3.7 FlashGoogle | 56.0 | 76.1 | 45.1 | $1.50 | $6.00 |
| Grok 4.5xAI | 55.8 | 72.4 | 48.9 | $2.00 | $6.00 |
| Claude Sonnet 5Anthropic | 55.3 | 71.5 | 49.7 | $2.00 | $10.00 |
| DeepSeek V4 ProDeepSeek | 53.2 | 68.8 | 49.6 | $1.32 | $3.96 |
| GPT-5.4 WorkhorseOpenAI | 53.1 | 71.1 | 44.2 | $2.50 | $15.00 |
| GLM-5.2 (Z.ai)Z.ai (Zhipu) | 52.6 | 68.8 | 45.7 | $1.40 | $4.40 |
| GPT-5.6 LunaOpenAI | 52.3 | 71.4 | 46.9 | $0.20 | $1.20 |
| Gemini 3.5 Flash (OpenRouter)Google | 52.0 | 70.1 | 39.7 | $1.50 | $9.00 |
| Gemini 3.6 FlashGoogle | 51.6 | 69.2 | 40.5 | $1.50 | $7.50 |
| Claude Sonnet 4.6 (OpenRouter)Anthropic | 48.4 | 63.0 | 42.1 | $3.00 | $15.00 |
| Gemini 3.1 ProGoogle | 47.7 | 68.8 | 23.0 | $2.00 | $12.00 |
| MiniMax M3 (OpenRouter)MiniMax | 45.4 | 58.6 | 36.1 | $0.30 | $1.20 |
| Kimi K2.6Moonshot AI | 45.1 | 61.8 | 31.2 | $1.50 | $6.00 |
| Kimi K2.7 Code (OpenRouter)Moonshot AI | 43.0 | 60.8 | 30.3 | $0.67 | $3.40 |
| DeepSeek V4 FlashDeepSeek | 42.1 | 56.2 | 33.7 | $0.44 | $1.32 |
| GLM-5.1 (Z.ai)Z.ai (Zhipu) | 41.0 | 55.8 | 30.6 | $1.40 | $4.40 |
| GPT-5.4 miniOpenAI | 40.9 | 56.1 | 31.5 | $0.75 | $4.50 |
| Grok Build 0.1xAI | 40.7 | 51.5 | 28.9 | $1.00 | $2.00 |
| GPT-5.4 nanoOpenAI | 39.7 | 56.1 | 29.7 | $0.20 | $1.25 |
| MiniMax M2.7 (OpenRouter)MiniMax | 38.9 | 52.6 | 25.9 | $0.30 | $1.20 |
| Grok 4.3xAI | 37.9 | 42.2 | 24.2 | $1.25 | $2.50 |
| Gemini 3.5 Flash-LiteGoogle | 37.4 | 49.3 | 27.2 | $0.30 | $2.50 |
| GLM 4.7 (Zhipu)Z.ai (Zhipu) | 34.5 | 45.3 | 26.2 | $0.60 | $2.20 |
| Qwen3.5 397B A17B (OpenRouter)Alibaba Qwen | 34.3 | 48.2 | 19.8 | $0.39 | $2.34 |
| Mistral Medium 3.5Mistral AI | 30.4 | 46.9 | 19.2 | $1.50 | $7.50 |
| Claude Haiku 4.5Anthropic | 29.9 | 43.9 | 16.5 | $1.00 | $5.00 |
| Gemma 4 31B InstructGoogle | 29.7 | 43.4 | 14.4 | $0.25 | $0.75 |
| GLM-4.6 (Z.ai)Z.ai (Zhipu) | 29.3 | 45.8 | 18.6 | $0.60 | $2.20 |
| Gemini 2.5 ProGoogle | 25.9 | 33.3 | 7.2 | $1.25 | $10.00 |
| Gemini 3.1 Flash Lite (OpenRouter)Google | 25.6 | 34.7 | 6.5 | $0.25 | $1.50 |
| Cerebras — GPT OSS 120BCerebras CS-3 | 24.1 | 30.4 | 13.4 | $0.35 | $0.75 |
| Qwen3 Coder Next (OpenRouter)Alibaba Qwen | 21.3 | 36.2 | 8.9 | $0.12 | $0.80 |
| Devstral 2 (2512) (OpenRouter)Mistral AI | 19.2 | 31.3 | 10.6 | $0.44 | $2.20 |
| o3-miniOpenAI | 15.7 | 16.3 | 1.7 | $1.10 | $4.40 |
| DeepSeek V3 (Chat)DeepSeek | 15.2 | 21.2 | 1.6 | $0.14 | $0.28 |
| o3 (Reasoning Frontier)OpenAI | 14.5 | 16.2 | 2.9 | $10.00 | $40.00 |
| Llama 4 Maverick (400B MoE)Meta Llama | 14.5 | 16.3 | 1.2 | $0.45 | $1.25 |
| Together AI — Llama 4 MaverickTogether AI | 14.5 | 16.3 | 1.2 | $0.35 | $0.90 |
| Llama 4 Scout (109B MoE)Meta Llama | 10.3 | 8.2 | 1.1 | $0.15 | $0.45 |
| GPT-4o (Omni)OpenAI | — | 11.4 | 1.0 | $2.50 | $10.00 |
| GPT-4o miniOpenAI | — | 11.4 | 1.0 | $0.15 | $0.60 |
| o1 (Reasoning)OpenAI | — | 39.7 | — | $15.00 | $60.00 |
| GPT-4 TurboOpenAI | — | 21.5 | — | $10.00 | $30.00 |
| Llama 3.3 70B InstructMeta Llama | — | 11.9 | — | $0.18 | $0.59 |
| Llama 3.1 8B InstructMeta Llama | — | 5.4 | — | $0.05 | $0.08 |
| OpenRouter Free Tier (Llama 3.3)OpenRouter | — | 11.9 | — | $0.00 | $0.00 |
| DeepSeek V3.2 (OpenRouter)DeepSeek | — | 44.2 | — | $0.26 | $0.38 |
| DeepInfra — Llama 3.3 70BDeepInfra | — | 11.9 | — | $0.10 | $0.30 |
| Qwen3.8-Flash (QwenCloud preview)Alibaba Qwen · Vendor-reported task scores, not third-party composite indices. | — | — | — | $0.16 | $0.47 |
| Ox Alpha (Z.ai GLM preview)Z.ai (Zhipu) · REMOVED from OpenRouter on 2026-08-28 (stealth/ox-alpha no longer listed). No third-party composite score was ever published. Do not substitute GLM-5.3 or GLM-5.3-Flash results as an Ox Alpha benchmark. | — | — | — | $0.00 | $0.00 |
Snapshot verified 26 Aug 2026. Source: OpenRouter live model API.
How to read the comparison
GLM-5.3 is already competitive
Its 59.5 Intelligence, 74.8 Coding, and 59.1 Agentic scores place the named Z.ai baseline near the frontier in this snapshot, especially for long-horizon agent work.
The Flash price is the disruption
Z.ai lists GLM-5.3-Flash at a temporary $0.075/$0.25/$0.015 per 1M input/output/cache tokens. The model has no published composite score yet, so price is verified while quality remains to be measured.
Ox Alpha is not a benchmark result
The free preview has no reproducible public score. Until the weights are released and an independent harness reruns the tests, GLM-5.3-Flash is a comparison proxy—not proof that every Ox Alpha response will match it.
Practical model choice
- Choose Ox Alpha for disposable experiments while the preview remains free, but do not send secrets or assume the route will persist.
- Choose GLM-5.3 for a named Z.ai baseline when you need a published API model and a current third-party benchmark snapshot.
- Choose GLM-5.3-Flash for price-sensitive coding tests if the temporary promotion is available in your route; measure quality, retries, and output length yourself.
- Keep a frontier fallback such as Claude Opus 5 or GPT-5.6 Sol for tasks where failure cost is higher than token cost.
Identity source: TechCrunch's report of Z.ai's confirmation. Pricing source: official Z.ai pricing. Benchmark source: OpenRouter live API.