Ox Alpha, OpenRouter’s No.1 Anonymous Model, Is GLM-5.3-Flash Under MIT
The anonymous ox-alpha model that appeared free on August 20 and topped OpenRouter usage within six days is GLM-5.3-Flash. Zhipu released the 320B-A18B weights under MIT the same day, and the free window is now closed.
- OpenRouter’s top-usage anonymous model
ox-alphaturned out to be GLM-5.3-Flash. - The 320B weights shipped under MIT, but the FP8 checkpoint alone is 306GiB.
- The free tier ended; input runs $0.075 per million tokens through September 9.
On August 20, 2026, a model carrying nothing but the name ox-alpha appeared free on OpenRouter, OpenCode, and Cline, with no vendor attached. Six days later, on August 26, China’s Zhipu AI (Z.ai) confirmed it was GLM-5.3-Flash and published the weights on Hugging Face under an MIT license the same day.
Shipping a model anonymously first to collect usage data is called a stealth launch. What made this one stand out was the volume. By OpenRouter’s count, ox-alpha processed roughly 42 trillion tokens in six days and at one point ran more than double the usage of DeepSeek-family models. OpenRouter had not seen numbers like this since DeepSeek Flash held the top slot for 56 days.
What happened during the six days
Free access is the obvious explanation, and it is part of the story. But free models do not automatically reach No.1. ox-alpha launched with a 1.05M-token context window, up to 131K output tokens, and text, image, and video all accepted as input. Nowhere else offered that combination at no cost.
ox-alpha listed free on OpenRouter, OpenCode, Cline, and Nous Research portal+.
Vendor undisclosed
Takes the No.1 usage slot on OpenRouter. Cache hit rate around 77.9%, knowledge cutoff confirmed as November 2025
Eleven tokenizer fingerprint signals match the GLM family. Numeric probing consumes 29 tokens (same as GLM; DeepSeek uses 98)
Z.ai confirms it as GLM-5.3-Flash. MIT weights published on Hugging Face the same day, free slot closed
The identification came down to the tokenizer. Every model splits text into tokens using a different vocabulary, so counting how many tokens a fixed string consumes narrows down the family. When the community fed in numeric strings, ox-alpha spent 29 tokens, the same figure GLM models produce. DeepSeek spends 98 on the same input. All eleven signals pointed to GLM.
In its announcement, Z.ai said the model was served entirely on domestic Chinese AI chips during the preview period. Zhipu made a similar point about Huawei Ascend training when GLM-5.1 took first place on SWE-Bench Pro.
320B parameters, 18B switched on
GLM-5.3-Flash has 320 billion parameters but activates only 18 billion of them per generated token. It is a mixture-of-experts design: many expert subnetworks sit in the model and the router turns on only the ones matching the input. The compute cost lands near that of an 18B model, which is where the low price comes from, while quality draws on knowledge distributed across all 320B.
It is also the first natively multimodal model in the GLM-5 line, taking images and video directly. The intended use is feeding in screen captures so the model can judge rendered output and interaction results. Attention mixes sparse and linear variants; Z.ai says the combination cuts the KV cache, the memory that accumulates as a conversation grows longer, to 1/4.44 of previous GLM models. For agents that hold long conversations, that ratio is where serving cost is decided.
The official comparison chart Z.ai published puts six models side by side across six benchmarks. GLM-5.3-Flash leads exactly one of them.

The single first place is GDPVal-AA v2. It scored 1773 against 1675 for the runner-up, DeepSeek-V4-Vision-Exp. Artificial Analysis ran that particular evaluation.
The coding numbers look different. On Terminal Bench 2.1 it scored 84.3, behind GPT-5.6 Terra at 87.4, Gemini 3.7 Flash at 85.8, and Claude Opus 4.8 at 85.0. Fourth of six. DeepSWE v1.1 comes in at 63.4, under GPT-5.6 Terra’s 69.6 and Gemini 3.7 Flash’s 65.3. Agents’ Last Exam scores 26.3, lower than every comparison model except its own predecessor GLM-5.2 at 20.4.
This is a model that pulled level with frontier systems, not one that beat them. The model card says "approaching Claude Opus 4.8," not surpassing it. Against Claude Opus 4.8 the results split by task: GLM-5.3-Flash leads on DeepSWE (63.4 vs 58.0) and AutomationBench (48.8 vs 41.0), and trails on Terminal Bench (84.3 vs 85.0) and HLE tool use (55.3 vs 57.9).
Attach prices instead of scores and the arithmetic changes. Claude Opus 4.8 costs $5 per million input tokens and $25 per million output. GLM-5.3-Flash list price is $0.15 input and $0.50 output, 1/50th on output. The "10x cheaper" line in the model card is measured against the previous GLM-5.2, not against Claude.
Against that predecessor the gap is large. AutomationBench went from 26.2 to 48.8 and DeepSWE from 46.2 to 63.4. Both benchmarks measure whether a model can finish a task across repeated tool calls.
Independent evaluation tells a different story
Artificial Analysis measured the model separately, and its strengths and weaknesses separate cleanly. It scored 57 on Intelligence Index v4.1.1, third among 108 models in its class. The median for open-weight models of comparable size is 27, so this is more than double.
Intelligence Index (class median 27)
Output tokens per second (median 67)
Output tokens to finish the index (median 110M)
The two trailing numbers are the problem. At 48.7 tokens per second it runs slower than the 67-token class median. Completing the same evaluation consumed 150 million output tokens against a median of 110 million. This is a verbose model. On an API billed by output tokens, a low unit price does not shrink the invoice by as much as the rate card suggests. Time to first token was 1.52 seconds, faster than the 2.14-second median.
That profile favors background work where accuracy matters more than latency, and works against interactive use where someone is waiting at a screen.
What it costs to use today
The free slot is closed. ox-alpha was billed as a one-week free preview, and revealing the identity moved it to a regular paid listing. Any configuration left running since August 20 is now being charged.
| Route | Who can use it | Price (per 1M tokens) | Conditions |
|---|---|---|---|
OpenRouter z-ai/glm-5.3-flash | Anyone with an account | $0.075 input, $0.25 output, $0.015 cache read | 50% off through September 9, 24:00 UTC+8, then double. Six providers (Z.ai, NovitaAI, GMICloud, Cloudflare, DeepInfra, io.net) |
Z.ai API glm-5.3-flash | Developers with an API key | $0.15 input, $0.50 output, $0.03 cached input | OpenAI-compatible format. Function calling, streaming, and structured JSON output supported |
| GLM Coding Plan | Subscribers | From $18/month | 3x the GLM-5.3 quota, an extra 50% off during off-peak hours |
| Self-hosting | Weight download | Free (MIT) | FP8 checkpoint around 306GiB, BF16 roughly double. vLLM supports Hopper and newer only |
Both the international Z.ai API and OpenRouter list no regional restrictions, and the OpenRouter provider list includes global operators such as Cloudflare and DeepInfra. Note that mainland China’s open.bigmodel.cn and the international z.ai are separate platforms, so create the account on the latter.
Released under MIT and runnable on your own server are two different claims. Commercial internal deployment and modification are legally unencumbered, and that part is real. But the FP8 checkpoint alone is around 306GiB before the KV cache is added on top. You need something on the order of an 8x H200 node, and the vLLM implementation supports only Hopper-generation NVIDIA GPUs and newer. llama.cpp has no glm5_next support yet, so there is no official GGUF either. If the plan was to load this onto a laptop, that does not work today.
Alibaba released Qwen3.8-Flash-Next as open weights the same day, previewing the Qwen4 architecture. That one ships under qwen-community-1.0 rather than MIT, and its hosted API is not open yet. Two Chinese labs published two open-weight releases within a day of each other. On license freedom and on whether you can use it today, GLM is ahead.
If you wired ox-alpha into a coding agent last week, the number to check before swapping the model string to z-ai/glm-5.3-flash is that week’s output token count. If the model really emits 36% more tokens than the class median while doing the same work, as Artificial Analysis measured, the $0.075 input rate gets offset by exactly that much on the real invoice. Pricing through September 9 is discounted, which distorts the comparison, so run the math at the doubled list price and put that figure next to your current model’s monthly bill.