GLM-5.3 outperforms ox-alpha, the mystery model going viral, across three benchmark measurements
Since August 20, a mystery model called ox-alpha has been given away for free while the community guesses its identity. The leading theory points to the GLM family, owned by Z.ai, as its maker. We measured both directly: on Javanese speech levels and Indonesian slang, GLM-5.3 outperformed ox-alpha across every arm, and their answers matched only 71% of the time. Maybe related, but clearly not identical twins.
Key numbers
Where these come fromMeasured from export-lajur-uji-2026-08-24-jv-A-ringkas.csv and 4 more
A mystery model going viral
On August 20, 2026, a model called ox-alpha appeared for free on OpenRouter and OpenCode: a million-token context, no company name attached, and within three days it was already used by hundreds of thousands of people.
The community raced to guess its identity through tokenizer probes and code-style fingerprints, and the leading theory points to the GLM family owned by Z.ai. We were curious about a question no probe had answered: can this mystery model handle Javanese, and if it really is related to GLM, does the official GLM-5.3 behave the same way?
One thing needs to be clear from the start. Ox-alpha is not on the ibahasa.com/benchmark leaderboard, and it never will be. Stealth models swap identities within days, so we measured it in a trial lane: a separate track that uses the same questions and scoring rubric as the official board, but keeps the results quarantined, never mixed with the official numbers. GLM-5.3 is different: it is a stable, official model, and we will run it on the official board in early September. The numbers in this article can then be compared directly against its official score.
How did we measure it?
We used two sets from our benchmark. First, the Javanese speech-level set (jv-A): 38 multiple-choice questions on picking the right register, plus 8 open-ended questions asking the model to rewrite a ngoko sentence into krama alus or the reverse. Open-ended answers were scored one by one by a Javanese speaker with a rubric weighted per word. An answer that left the main pronoun or verb in ngoko was marked failed and scored zero. Second, the Indonesian slang set (id-B): 60 multiple-choice questions on understanding colloquial language. As a size comparison, we included hy-mt2, a small 1.8-billion-parameter model.
Neither set's questions nor its answer key has ever been published, so none of the three models could have seen them during training. The whole measurement cost about 3 cents (USD), and ox-alpha was free during its promotional period.
Benchmark results
| Measurement | GLM-5.3 | ox-alpha | hy-mt2 (1.8B) |
|---|---|---|---|
| jv-A multiple choice (38 items) | 32 | 27 | 13 |
| jv-A krama speech-level production (8 items, summed score) | 7.67 | 6.72 | 0 |
| id-B Indonesian slang (60 items) | 57 | 49 | 34 |
In percentages, both multiple-choice sets tell the same story from two angles. On Javanese speech levels, GLM-5.3 answered 84.2% correctly and ox-alpha 71.1%. On Indonesian slang, GLM-5.3 reached 95.0% and ox-alpha 81.7%. The small model hy-mt2 scored 34.2% on the Javanese questions and 56.7% on the slang ones.
The chart above shows two gaps with different meanings. The gap between the GLM-5.3 and ox-alpha bars is thin on both sets: the two are in the same class, with GLM-5.3 always slightly ahead. The gap to the hy-mt2 bar is wide, and widest on the Javanese set: the small model still manages roughly half correct on Indonesian slang, but drops close to random guessing once the questions turn to speech levels. Regional-language ability, it turns out, is the slowest to reach small models.
On krama production, the story changes again: GLM-5.3 and ox-alpha both came close to a perfect score. Out of 8 sentences, our reviewer accepted nearly all of them, something no model on the official board has managed outside its very top tier, while hy-mt2 failed all eight sentences.
So, is ox-alpha GLM in disguise?
If ox-alpha is GLM-5.3 wearing a mask, the two should answer almost identically. They do not.
11 questions in the red segment are the core finding: on 11 of 38 questions, the two models picked different letters, which is why the match rate stops at 71%. Ox-alpha got 11 wrong, GLM-5.3 got 6 wrong, and their mistakes overlapped on only 3 questions. Two models that are truly identical, run twice, typically match above 90%, and their donut ring would come out nearly one solid color.
Our conclusion: behaviorally, ox-alpha is not the official GLM-5.3. Being related is still possible, for instance a multimodal variant or a differently tuned model built on the same base, and the two models' equally high krama-production scores support that. But identical twins are ruled out by the data itself.
Evidence that fell apart in our own hands
3 questions where both models picked the same wrong letter first read to us as a family fingerprint. Then we checked those questions against 14 official-board models, and 2 of those 3 fingerprints collapsed: on one question, 12 of the 14 board models picked the exact same wrong letter, and on the other, 5 more models picked it too. Both turned out to be distractor traps that snag almost every model, not a sign of kinship. One fingerprint remains: a single question where the wrong answer picked by ox-alpha and GLM-5.3 was only ever picked by one other model on the entire board.
We also found a flaw in our own instrument during this benchmark session. On one production item in Javanese speech levels, ox-alpha wrote a nearly perfect krama sentence but left the word "aku" (I) in ngoko, and our rubric turned out to have no criterion for pronouns on that item: the mistake slipped through with no penalty. We will revise the rubric through our item-versioning process, and we note this here so readers know ox-alpha's score of 6.72 is a little more generous than it should be.
Conclusion
The 5-answer gap on jv-A multiple choice can still arise by chance (a two-proportion test gives z around 1.4). So the honest conclusion is not that GLM-5.3 is more skilled at speech levels, but that GLM-5.3 is consistently ahead across all three measurements, and only on id-B is the gap statistically decisive.
Ox-alpha's identity itself remains a theory. Our data refutes the two being identical twins, but it does not prove who built ox-alpha. Every number above also comes from just one run per model. GLM-5.3's official score on the board in early September will test how stable all of this really is.
The derived data for the public slice is available in the trial-model-probes dataset linked below this article. It contains per-model summaries and a per-question match flag. If ox-alpha's identity is ever announced, these numbers will already be waiting to be compared.
Correction, August 26, 2026
The first version of this article said that GLM-5.3's Indonesian slang score was "far above the highest score on our official board (42 out of 60)". That was wrong. We took the number 42 from an export column that counts only the 42-item public slice, not the full 60 items. gpt-5.6-luna actually scored 60 out of 60 on this set, which puts GLM-5.3's 57 level with deepseek-v4-flash and below gemma-4-31b. Ox-alpha's 49 sits near the bottom of the board, above only hunyuan-a13b and three locally run models.
We also corrected "four measurements" to three, matching the table itself: Javanese speech-level multiple choice, krama production, and Indonesian slang. The article's main finding, the 71 percent answer match that rules out ox-alpha being a GLM-5.3 twin, is unaffected by this correction.
Data & provenance
- export-lajur-uji-2026-08-24-jv-A-ringkas.csv
- export-lajur-uji-2026-08-24-jv-A-kecocokan.csv
- export-lajur-uji-2026-08-24-id-B-ringkas.csv
- export-lajur-uji-2026-08-24-id-B-kecocokan.csv
- export-lajur-uji-PROVENANCE.md
Limitations
One run per model per arm, from a trial lane kept separate from the official board. The jv-A multiple-choice set has 38 items and id-B has 60, so a gap of a few answers can still arise by chance, and we state that limit in the body. Open-ended answers were scored by a single Javanese speaker with a weighted rubric. Ox-alpha's identity remains a community theory, not a fact we have confirmed.
Revision history
- Published
- Last updated
Terms in this article
- token
- The chunk of text a model counts in, roughly a syllable up to a short word. AI services are priced per million tokens, with input and output billed separately, so a rambling answer genuinely costs more than a concise one.
Written by the author for this article, not taken from a dictionary entry.