The two smallest models we have ever tested, LFM2.5 at 2.6 billion parameters and hy-mt2 at 1.8 billion, faced off on 60 Indonesian slang questions. The scores ended in a dead heat, but the two failed in opposite ways, and one of them tried calling Google mid-exam.
34/60hy-mt2 (1.8B)32/60LFM2.5 (2.6B)358xToken ratio
We tested ox-alpha and GLM-5.3 on Javanese speech levels and Indonesian slang in a trial lane separate from the official leaderboard. GLM-5.3 led on all four measurements, and a 71% answer match shows the two are not the same model.
57/60GLM-5.349/60ox-alpha34/60hy-mt2 (1.8B)
We tested 10 AI models on 135 Javanese words and predicted the results would split into two clusters with an empty gap between them. The last 3 models were picked precisely because they were the likeliest to break that prediction, and 2 of them landed squarely inside the gap we had called empty.
10Models tested135Javanese words1 of 4Predictions right
Three earlier articles ended with the same admission: every time we read one model's answers and fixed the key, that model gained the most. This time the protocol was designed to make that impossible, and across 135 Javanese words and 6 AI models, the protocol held. What we did not expect: once the instrument stopped moving, the AI models themselves turned out to be unsteady. At temperature 0, which should produce identical answers, 36 percent of answers changed between two consecutive calls.
135Javanese words6AI models36%Answers changed
Rankings on our Javanese benchmark have so far been set by a human-written answer key. To check whether that ranking is real, we added a second instrument that never sees the key: IndoBERT, a small 40 MB model that runs on our own machine. All four models came back in the exact same order. But 63 percent of the similarity it measured turned out to be sentence framing, not meaning, and that has to be measured first before a single number can be read.
43Javanese words tested40 MBLocal IndoBERT, second instrument63%Similarity that was framing
Four AI models were asked the meaning of 43 Javanese words absent from NusaX, NusaWrites and Javanese Wiktionary. Mistral Small 3 got 5 percent right, DeepSeek V4 Flash 53 percent, GPT-5.6 Luna 81 percent, Gemini 3.5 Flash 91 percent. The instrument itself was patched twice mid-measurement, and each patch raised scores without a single call being repeated.
43Javanese words91%Max Score4 ModelsMistral, DeepSeek, GPT, Gemini
Example sentences from our Javanese dictionary became the test items, and their Indonesian glosses became the answer key. Two AI models were measured. In a blind adjudication of 20 pairs, our glosses won 4, the model won 5, and the remaining 11 were tied. An answer key that is no better than what it grades measures agreement on word choice, not ability.
11 of 20Blind comparisons tied0 of 20Exact match with our gloss2AI Models measured
Our slang test offered free answers to anyone reading options. Yet, AI models grabbed at most 1.00 out of 4.1 items.
38 questionsSlang set tested4.1 questionsFree answers available1.00 questionActually taken by model
Mistral Small 3 and Llama 3.1 8B Instruct are sold at the exact same price and are equally fast. On 69 items testing normalization of abbreviated writing to standard Indonesian, one scored 56.52%, the other 26.09%.
69 itemsItems56.52%Mistral Small 326.09%Llama 3.1 8B
The hypothesis was reasonable: let a contextual language model pick the best correction candidate based on sentence meaning. Tested A/B with a single switch and an answer key of 1,500 pairs, suggestion accuracy instead fell from 88.3 percent to 77.1 percent, while the model consumed 98.5 percent of analysis time.
1,500Answer-keyed pairs88.3%Correct suggestions, reranking off77.1%Correct suggestions, reranking on
Two Anthropic models were tested on 36 Javanese speech level questions. Claude Sonnet 4.6 missed 3, Claude Haiku 4.5 missed 10.
36 questionsMultiple-choice questions91.67%Claude Sonnet 4.6 score72.22%Claude Haiku 4.5 score
Seven models tested on one and the same set. An 876-fold gap in output tokens for a single letter answer, and the numbers themselves move between runs.
875.8Tokens, most wasteful model1.0Tokens, leanest model876xGap on the same set