Anonymous AI Model Surpasses GPT and Claude in Coding Tests, Technical Clues Point to Z.AI's Unreleased Flagship

Deep News
Aug 21

An unidentified AI model has quietly surfaced on the model distribution platform OpenRouter. Independent researcher testing reveals that the model, called "stealth/ox-alpha," not only beats mainstream models like GPT and Claude on certain coding benchmarks, but its technical characteristics also strongly point to Z.AI's next-generation multimodal flagship, which has yet to be officially released.

According to a test report published by technical researcher Ben Davis on August 21, ox-alpha was launched on OpenRouter on August 20 and currently offers free access for one week. The model supports text, image, and video inputs, possesses reasoning capabilities, and features a context window of 1.048 million tokens.

Davis stated on X that he is "99% certain" ox-alpha belongs to the GLM-5.x series from Z.AI, citing multiple pieces of evidence including the video encoder, tokenizer, output style, and audio refusal behavior. His tests also showed that ox-alpha outperformed GPT and Claude in several comparative evaluations.

However, it must be emphasized that the above conclusions are based on personal testing and technical inference, and have not yet received official confirmation from Z.AI.

Video Encoder Emerges as Key Evidence

Davis's testing traced ox-alpha's technical origins across multiple dimensions, including video encoder, tokenizer, audio interface, and output style. Among these, the high degree of matching in the video encoder is considered the strongest evidence.

The tests showed that across four controlled video sets, ox-alpha's video token consumption was completely identical to GLM-5V-Turbo, with both displaying the same three characteristics: frame-rate-independent frame sampling, a duration scaling ratio of approximately 147 tokens per second, and a per-frame resolution scaling mechanism.

In contrast, candidate models such as MiMo v2.5, Qwen 3.8 Max, and GLM-4.6V all exhibited clearly different video encoding characteristics.

Tokenizer testing also pointed toward GLM. The report stated that after testing 25 different prompts, ox-alpha's token count matched GLM-5.3 exactly, with only a fixed +75 token hidden wrapping difference, suggesting the two may share the same vocabulary.

Additionally, ox-alpha's refusal of audio input is consistent with GLM-5V, while MiMo v2.5, listed as a primary competitive candidate, supports audio input, further weakening the latter's possibility.

In terms of output style, ox-alpha uses approximately 1.3 emojis per thousand characters, which is relatively close to the GLM/Qwen series. By comparison, Claude, GPT-5.6, and Grok showed near-zero emoji usage under the same test conditions.

Has Coding Ability Already Surpassed GLM-5.3?

Regarding capability testing, the report cited DeepSWE benchmark data showing that ox-alpha passed 8 out of 10 deterministic tasks, achieving a Pass@1 rate of 80%.

For comparison, Claude Fable 5 had a pass rate of 65%, GLM-5.3 and Grok 4.6 both stood at 62%, and GPT-5.6-sol came in at 52%. However, since the number of test runs for different models was not entirely consistent, and ox-alpha currently has a relatively small sample size, this result still requires more independent testing for verification.

Notably, in the "meriyah-explicit-resource-declarations" task, ox-alpha passed on the first attempt, while GLM-5.3, GPT-5.6-sol, and Grok 4.6 had previously all failed with 0/4 results. At the same time, ox-alpha maintained a perfect record of all 51,469 regression tests passing.

The report also documented an agentic task involving 69 tool calls. Throughout the entire process, the model made only a single error, experienced no retry loops, and maintained low inference overhead. Based on this, Davis believes that ox-alpha's capability performance is already clearly stronger than GLM-5.3, making it more likely to be a next-generation model checkpoint rather than a simple variant version.

Why the Evidence Points to Z.AI

Beyond technical fingerprints, Davis also made cross-judgments based on model release timing and Z.AI's previous testing methods.

Z.AI released the text-only version of GLM-5.3 on August 14, and a unified vision flagship model has been a focus of community attention. More importantly, Z.AI has a precedent for testing models through stealth channels, as Pony Alpha was ultimately confirmed to be related to GLM-5.

In terms of model scale, the report stated that ox-alpha's decoding speed differs from GLM-5V-Turbo by approximately 6%. The latter has 744B total parameters and 40B active parameters, leading Davis to speculate that ox-alpha may adopt a similar-scale MoE architecture.

The report further argued that if the model indeed has approximately 40B active parameters, then the operator's claimed daily serving capacity of 100 trillion tokens would be more feasible from both technical and cost perspectives.

Meanwhile, the report also systematically ruled out other potential sources. Xiaomi's MiMo showed clear differences from ox-alpha in video encoder and audio interface; DeepSeek has not previously released video capabilities, with different tokenizer and model release methods; Google, Qwen, xAI, OpenAI, and Anthropic were all considered mismatched with ox-alpha in terms of tokenizer, output style, or video encoder.

Free Testing May Run Through August 27

It is worth noting that ox-alpha's current free access window may last until August 27.

Davis pointed out that some previous stealth models were also formally claimed by relevant Chinese AI labs after their free testing periods ended.

Currently, Z.AI has not made an official response regarding ox-alpha's identity. If the model is ultimately confirmed to belong to Z.AI's next-generation multimodal flagship, then this "stealth test" on OpenRouter could serve as a public preview before its official release.

Until official confirmation arrives, the true origin of ox-alpha remains an open question, but from video encoder to tokenizer to model behavior, multiple independent tests currently all point in the same direction.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10