stealth/ox-alpha from its catalog on 2026-08-28. Availability is now unconfirmed; the free preview as documented here has ended on that route. See Ox Alpha Is Gone from OpenRouter — What It Means.Ox Alpha: Z.ai's GLM-Family Preview, Explained
Z.ai has now confirmed that the anonymous Ox Alpha release is a new GLM-family iteration. The exact checkpoint is still not formally named, but the public evidence points most strongly toward a GLM-5.3-Flash-era model. Here is what is verified, priced, and still unproven.
Input
$0 / 1M
Output
$0 / 1M
Context
1,048,576
Max output
131,072
What actually launched?
Command Code's current model page lists Ox Alpha under the provider namespace stealth with the model ID stealth/ox-alpha. The listing says it is designed for coding, sustained agentic work, long-horizon software engineering, complex reasoning, and production workloads. Z.ai confirmed to Bloomberg, as reported by TechCrunch, that Ox Alpha is a new iteration of its GLM series and that weights would be released.
The leading exact-model hypothesis is GLM-5.3-Flash, because the current named Z.ai model shares the 1M context, multimodal inputs, mandatory reasoning, and 131K output shape. That label is not yet a published Ox Alpha model card, and Ox Alpha has no independent benchmark score.
The important capabilities
Long context, not a promise of long-context quality
The original preview listing published a 1,048,576-token context limit with up to 131,072 completion tokens. That makes large repositories, long agent traces, and document-heavy workflows plausible test cases. It does not, by itself, prove reliable retrieval across a million-token prompt.
Reasoning was mandatory in the preview contract
The initial preview contract marked reasoning as mandatory and exposed low, high, and max effort levels. Expect a reasoning-heavy route to spend part of its completion budget before the visible answer; compare workloads by total tokens and latency, not visible output alone.
The preview was multimodal and tool-capable
The original listing exposed text, image, and video input, plus tools, structured response controls, and reasoning parameters. That points toward coding agents and visual software workflows, although no independent tool-call or vision evaluation is published yet.
Treat the free tier as a test window
Command Code currently prices Ox Alpha at $0 for input, output, and cache reads while the stealth preview lasts. That is the hosted preview price—not a published Z.ai API price for the Ox Alpha name. The exact checkpoint, rate limits, availability, and data handling can still change. Do not send secrets, customer data, or regulated information until you have verified the active route's policy.
Where to use Ox Alpha free right now
OpenCode
Open ↗Free coding agent
Install OpenCode, then check `/models` for the free `x-preview-f-free` entry. The live alias is operationally related to the preview, but OpenCode has not published a direct Ox Alpha label, so verify the current mapping before starting a long run.
Command Code
Open ↗Free model on paid harness
Command Code's official Ox Alpha page lists `stealth/ox-alpha` at $0 input, output, and cache-read cost while the preview lasts. The Command Code harness currently requires Go or a higher plan.
Free availability is temporary and capacity-limited. Check each linked model page or live model picker before sending work.
How to evaluate Ox Alpha this week
- Start with disposable workloads. Use synthetic repositories, public documents, and non-sensitive agent tasks.
- Measure total economics. Log input, visible output, reasoning tokens, latency, tool-call errors, and retry rates.
- Test the failure modes. Long-context recall, structured output, tool selection, code repair, and multi-step task completion matter more than a single chat demo.
- Keep a fallback. Route production traffic to a named model if Ox Alpha disappears, changes behavior, or hits a free-tier limit.
Bottom line
Ox Alpha matters because it combined a 1M-token, reasoning-heavy coding profile with a zero-dollar preview and then turned out to be a Z.ai GLM-family release. The strongest current description is: a free hosted preview, probably closest to GLM-5.3-Flash, with no standalone price card or independent benchmark score yet. Use the named GLM-5.3 baseline for reproducible comparisons and treat the preview as disposable until the released weights and route terms settle.
Sources: OpenCode live model list,Command Code Ox Alpha page, andofficial Z.ai pricing,TechCrunch's reveal report, andZ.ai's GLM-5.3-Flash announcement. Catalog facts last checked 26 Aug 2026.