Notes from five days with the free OpenRouter preview
I spent 243 million tokens with a model that had no name.
I used Ox Alpha for coding, technical editing, and local-inference work. Then Z.ai revealed what it was and released the weights.
stealth/ox-alpha→zai-org/GLM-5.3-Flash
On August 20, a new endpoint appeared on OpenRouter. It had a one-million-token context window, accepted images and video, supported tool calls, and cost nothing. The listing gave no developer name. It was simply stealth/ox-alpha.
I added it to Claude Code and used it for the personal research work already on my desk. It reviewed a book about inference systems, implemented schedulers and allocators against adversarial harnesses, helped bring up a local inference stack, and rebuilt a data-heavy website. My rough impression after those sessions was that it worked better than Sonnet, but was still a little behind Opus or Sol.
That comparison is subjective. The useful part is that OpenRouter retained enough per-generation data to check what I had actually asked the model to do and which tasks reached a verifiable result.
The OpenRouter management export contains 2,118 Ox Alpha generations across five active days: 242.1 million prompt tokens and 1.56 million completion tokens. That large input number is mostly long Claude Code histories being sent again on each turn, not 242 million tokens of new material; 83.2% of the prompt volume hit cache. The route produced 23.1 output tokens per second on average and the bill stayed at zero.
How Ox Alpha appeared
OpenCode said it had capacity for 100 trillion tokens a day. OpenRouter described the model as suitable for coding and sustained agent work. Neither post named a lab, so the first few days were split between people trying the model and people trying to identify it.
“Ox Alpha (stealth model) is free for the next week — 1M Context — Multi-modal — Zero Data Retention… We have capacity for 100T tokens per day.”
This was the post that made the preview hard to ignore: a large context window, no charge, and a capacity claim big enough to invite stress testing.
OpenRouter introduced Ox Alpha as a stealth model for coding and sustained agent work, with text, image, and video input.
The capabilities narrowed the list of plausible model families before any tokenizer probing began.
Asking the model to identify itself was not useful evidence; language models will confidently repeat names suggested by their prompts or training data. The better clues came from the API: exact context and output limits, reasoning controls, GLM-family error signatures, and tokenizer counts that matched GLM after accounting for the wrapper.
1,048,5761,048,576exact131,072131,072exactlow / high / maxlow / high / maxexactOx + wrapperGLM-5.3matchtext · image · videoGLM-5V lineagealigned13011301alignedTechCrunch covered the attribution debate on August 23. Two days later, a Hugging Face repository appeared with 62 weight shards and a new Glm5Next architecture. On August 26, Z.ai confirmed that it had tested GLM-5.3-Flash anonymously as Ox Alpha on OpenCode and OpenRouter.
- Anonymous route opens
Free, multimodal, 1M context.
- Fingerprinting begins
Tokenizer, errors, controls.
- Mystery becomes news
Community theory goes mainstream.
- Weights appear
62 shards, MIT, native vision.
- Z.ai confirms it
Ox Alpha was GLM-5.3-Flash.
Once the identity was settled, I wanted to know whether my positive impression was supported by the work in the logs.
What the logs actually show
The aggregate token count says nothing about whether a task finished. I downloaded the retained history for each of the 13 named sessions and checked the tool output for passing harnesses, completed builds, HTTP responses, errors, and recovery loops. The raw prompts, paths, and code remain private; the table below summarizes the workloads.
Schedulers, speculative rollback, prefix sharing, a temporal allocator, and distributed garbage collection.
At least five bounded implementations ended with their supplied harness at 0 violations; several also passed adversarial checks.
Backfill, categorization, trend analysis, and a React UI over 65,767 stories.
Generator completed, production build passed, and the local preview returned HTTP 200.
Content critique, cross-chapter consistency, new chapters, appendices, links, and exercises.
The trace shows useful review work and clean links, then repeated blank turns, timeouts, and “why are you stuck?” prompts.
vLLM Metal installation, a 22 GiB model download, server debugging, and a Claude Code smoke benchmark.
Installation and download succeeded; the agent benchmark still failed at a 16,385-token request against a 16,384-token window when the sampled trace ended.
These were personal research and hobby projects, not production systems. A passing harness tells me that the implementation satisfied a deliberately constrained specification. It does not tell me that the code would survive real traffic, hostile inputs, maintenance by a team, or a security review. I find the harness results useful because they are concrete, but I do not treat them as a code-quality benchmark.
The strongest results came from bounded repository tasks with an explicit oracle. The allocator and distributed-GC implementations passed their main tests on the first run. In the harder scheduler sessions, the model wrote additional stress tests, discovered that some failures came from mistakes in those tests rather than in the implementation, and kept the supplied harness green. That ability to distinguish “my test is wrong” from “my code is wrong” was one of the better signals in the entire trial.
$ python3 test_harness.py
INFERENCE SCHEDULER BENCHMARK — HARD MODE
Requests: 4 Completed: 4
Violations: 0
✓ ALL INVARIANTS PASSED
$ python3 /tmp/stress_scheduler.py
ALL STRESS CHECKS PASSED
The website task was another clean result, but for a different reason. The model had existing data, a visible UI, a build command, and a local preview. It reorganized the trends view around categories, regenerated the dataset, produced a successful production build, and returned an HTTP 200 from the preview. Again, this is evidence of task completion on a hobby project. It is not evidence that I would merge the result into a production product without review.
The longer sessions were less reliable. The book review used 1,400 requests and nearly 195 million prompt tokens. It produced useful editorial analysis and caught cross-chapter inconsistencies, but the history also contains repeated requests to continue after blank responses, stopped background workers, and several timeouts. The size of that session is also why the headline token count needs context: Claude Code kept replaying an increasingly large book and conversation history. Most of the traffic was cached input, not new output from the model.
The local-inference session was the clearest example of capable work without a completed outcome. Ox Alpha installed vLLM Metal, downloaded roughly 22 GiB of Qwen weights, diagnosed a KV-cache allocation failure, and got far enough to send a Claude Code task through the local endpoint. The final request was 16,385 tokens against a 16,384-token window—an off-by-one mismatch between the client’s estimate and the rendered prompt. The model proposed a clamping proxy, but the sampled trace ended before an end-to-end benchmark passed.
This is why I still place it between Sonnet and Opus/Sol for this kind of personal research. I cannot make a claim about production code quality from hobby projects. What I can say is that Ox Alpha completed several bounded tasks with machine-checkable results, while needing more intervention when a session involved background processes, repeated tool failures, or several hours of accumulated state.
One user called it “quite effective” at tracking down complex bugs, then documented repeated partial fixes and hung streams during a longer orchestration task.
That report is close to what I saw: strong implementation work, less dependable recovery after the session became complicated.
Another user described a good multimodal debugging result, followed by 400 errors when trying to continue the conversation.
The overloaded free route made it difficult to separate model failures from serving failures.
The benchmark numbers moved
The first viral number was 80% on DeepSWE. It came from ten tasks and was posted with a warning about variance. The larger run landed around 63%. A community evaluation over all 113 tasks later reported 58.4%. Z.ai’s release reports 63.4% with its own mini-swe-agent settings.
The first ten-task DeepSWE sample scored 80%. Davis warned in the same post that the estimate could have substantial variance.
“Ended at ~63% NOT the 80% my first subset test got, which makes way more sense.”
The figures use different samples, settings, and harnesses. An agentic score describes the checkpoint together with its scaffold, tools, token budget, timeout, prompt, and task distribution. Quoting the score without that context makes the comparison much less useful.
The correction from 80% to roughly 63% is more informative than the original headline. It shows how quickly a small task sample can overstate a coding model, especially when a few successes move the percentage by ten points. The 58.4% community run and Z.ai’s 63.4% result are close enough to place the model in a serious range, but far enough apart to make the harness configuration part of the result.
Vendor-reported scores from Z.ai’s launch post. Use the tabs to change benchmarks; evaluation settings differ.
Z.ai’s table places GLM-5.3-Flash close to the leading models across terminal coding, software engineering, agent tasks, and tool-assisted reasoning. It does not lead every column. For an engineer choosing a model, the remaining question is how those scores transfer to a particular repository and tool setup. My logs are one small answer: good results when success was explicit, weaker reliability when the work depended on a long chain of operational state.
What Z.ai released
On August 26 the anonymous service became an inspectable artifact. The Hugging Face repository is MIT-licensed and roughly 328 GB across 62 FP8 shards. The config describes 320 billion total parameters, 18 billion activated for each token, 45 decoder layers, 288 routed experts, eight selected per token, and a 24-layer native vision tower. Z.ai calls it the first natively multimodal model in the GLM-5 family and says pretraining used 30 trillion multimodal tokens.
The central trick is a repeated 3:1 rhythm: three linear-attention layers, then one sparse-attention layer for global retrieval. IndexPool compresses four indexer key vectors into one before the sparse search. Manifold-Constrained Hyper-Connections, or mHC, manage multiple residual streams while constraining how they mix.
The motivation is straightforward. Full attention becomes expensive as the context grows because every new token may attend over an enormous history. Linear attention keeps most layers cheap, while the periodic sparse-attention layer provides a route back to distant context. IndexPool tries to make even that retrieval step cheaper by reducing the number of keys the indexer has to examine. Whether the approximation preserves the right details on a particular million-token workload is something independent evaluations still need to test.
config.json
The layer schedule is in config.json.
The list continues this hybrid schedule across 45 decoder layers.
Inspect the full config ↗"layer_types": [
"linear_attention",
"linear_attention",
"linear_attention",
"deepseek_sparse_attention",
"linear_attention",
"linear_attention",
"linear_attention",
"deepseek_sparse_attention"
]
18B active does not mean an 18B download
Mixture-of-experts sparsity reduces the weights touched during a token’s forward pass. It does not make the other experts disappear from storage. The full expert bank still has to live in memory, on disk with offload, or across a distributed serving cluster. At the published FP8 size, four 80 GB accelerators cannot even hold the raw 328 GB checkpoint; a practical deployment also needs room for runtime state and KV cache.
The “Flash” name refers to how much compute is used for a token, not to a small checkpoint. Running it on consumer hardware will require aggressive quantization and offload. Production deployments will have to deal with expert placement, interconnect, KV-cache layout, and long-context prefill.
The open repository is valuable even for engineers who never plan to host the full checkpoint. The config settles questions that were impossible to answer from the Ox Alpha endpoint, and the weights make quantization, expert offload, kernel profiling, and independent reproduction possible. The cost of doing those experiments is still substantial; “open weights” removes an access restriction, not the hardware requirement.
Why the stealth launch worked
Z.ai says the anonymous preview became the week’s most popular model and that the traffic was served on Chinese AI accelerators. The setup gave the company several useful things at once.
Users formed an opinion before they knew the vendor. Free access generated a large set of real workloads. The route exercised the production serving stack under heavy demand, and the attribution puzzle supplied publicity that a normal model announcement would not have received.
There is a less flattering reading too. Engineers were sending real code to an unnamed provider, and the serving instability of the free preview was mixed into early judgments about model quality. The experiment worked because people accepted that uncertainty in exchange for free access. A production evaluation should be much stricter about data handling, provider identity, retention policy, and uptime.
The free endpoint made it easy to try the model. The MIT-licensed weights make it possible to inspect, quantize, profile, and evaluate the same model without depending on the preview route. For engineers, that is what gives the release a life beyond the launch week.
Should you use GLM-5.3-Flash?
If you are choosing a model for personal coding or research, it belongs on the shortlist. The hosted price is low, the context window is unusually large, tool calling worked well in my bounded tasks, and the model can now be tested through a stable public slug instead of the temporary stealth route.
If you are evaluating it for a production agent, my results are a reason to run your own test suite—not a reason to skip one. Measure completion at the task level, not the generation level. Include repository setup, tool errors, retries, background processes, and recovery after a bad edit. The distinction matters because the model looked strongest when the verifier was immediate and weakest when success depended on maintaining state over a long session.
Try it now
Use the OpenRouter endpoint, keep tests in the loop, and compare the result with the model you already know.
Evaluate the whole loop
Track task completion, recovery attempts, latency, cache-adjusted cost, and human interventions—not just accepted patches.
Plan for infrastructure
The 18B-active label does not remove the 328 GB checkpoint, runtime memory, or long-context KV-cache requirement.
The evaluation I would run next
I would take a small set of real tasks from the same repository—one new feature, one bug with a known root cause, one refactor, one test failure, and one long-running task with an external service. I would run them through GLM-5.3-Flash, Sonnet, and Opus/Sol with the same tool permissions and token budget. The primary metric would be independently verified task completion. Secondary metrics would include human interventions, failed tool calls, wall-clock time, cached and uncached cost, and whether a correction remains fixed ten turns later.
That would say more about the model for my work than another broad leaderboard score. It would also make the comparison behind my “better than Sonnet, behind Opus/Sol” impression reproducible instead of anecdotal.
A minimal request
OpenRouter now lists the released model as z-ai/glm-5.3-flash. The endpoint reports the same 1,048,576-token context, image and video input, mandatory reasoning, and low/high/max effort controls.
const response = await fetch(
"https://openrouter.ai/api/v1/chat/completions",
{
method: "POST",
headers: {
"Authorization": `Bearer ${OPENROUTER_API_KEY}`,
"Content-Type": "application/json"
},
body: JSON.stringify({
model: "z-ai/glm-5.3-flash",
messages: [{
role: "user",
content: "Inspect this repository and propose a test plan."
}],
temperature: 1,
top_p: 0.95,
reasoning: { effort: "max" },
stream: true
})
}
);
The model is no longer the mystery
The anonymous preview established that GLM-5.3-Flash could attract users without a familiar lab name attached. The release gives engineers a better basis for judgment: a stable API route, published architecture, open weights, deployment recipes, and enough community testing to know which questions are still unresolved.
My own evidence is narrower. Across personal research projects, Ox Alpha was very good at bounded tasks with clear verification and less dependable over long operational sessions. That makes it interesting, useful, and worth evaluating. It does not make the hobby projects proof of production readiness.
If you try it, keep the task verifier outside the model, retain your traces, and report the failures as carefully as the successes. The most valuable follow-up to the stealth launch would be a set of reproducible workload reports from engineers who know exactly what “done” means in their own systems.
Sources and further reading
The primary artifacts, reporting, community tests, and deployment recipes used for this article. Vendor claims and anecdotes remain labeled.