Three earlier articles ended with the same admission: every time we read one model's answers and fixed the key, that model gained the most. This time the protocol was designed to make that impossible, and across 135 Javanese words and 6 AI models, the protocol held. What we did not expect: once the instrument stopped moving, the AI models themselves turned out to be unsteady. At temperature 0, which should produce identical answers, 36 percent of answers changed between two consecutive calls.
135Javanese words6AI models36%Answers changed
Rankings on our Javanese benchmark have so far been set by a human-written answer key. To check whether that ranking is real, we added a second instrument that never sees the key: IndoBERT, a small 40 MB model that runs on our own machine. All four models came back in the exact same order. But 63 percent of the similarity it measured turned out to be sentence framing, not meaning, and that has to be measured first before a single number can be read.
43Javanese words tested40 MBLocal IndoBERT, second instrument63%Similarity that was framing
Four AI models were asked the meaning of 43 Javanese words absent from NusaX, NusaWrites and Javanese Wiktionary. Mistral Small 3 got 5 percent right, DeepSeek V4 Flash 53 percent, GPT-5.6 Luna 81 percent, Gemini 3.5 Flash 91 percent. The instrument itself was patched twice mid-measurement, and each patch raised scores without a single call being repeated.
43Javanese words91%Max Score4 ModelsMistral, DeepSeek, GPT, Gemini
Example sentences from our Javanese dictionary became the test items, and their Indonesian glosses became the answer key. Two AI models were measured. In a blind adjudication of 20 pairs, our glosses won 4, the model won 5, and the remaining 11 were tied. An answer key that is no better than what it grades measures agreement on word choice, not ability.
11 of 20Blind comparisons tied0 of 20Exact match with our gloss2AI Models measured
Our slang test offered free answers to anyone reading options. Yet, AI models grabbed at most 1.00 out of 4.1 items.
38 questionsSlang set tested4.1 questionsFree answers available1.00 questionActually taken by model
Mistral Small 3 and Llama 3.1 8B Instruct are sold at the exact same price and are equally fast. On 69 items testing normalization of abbreviated writing to standard Indonesian, one scored 56.52%, the other 26.09%.
69 itemsItems56.52%Mistral Small 326.09%Llama 3.1 8B
The hypothesis was reasonable: let a contextual language model pick the best correction candidate based on sentence meaning. Tested A/B with a single switch and an answer key of 1,500 pairs, suggestion accuracy instead fell from 88.3 percent to 77.1 percent, while the model consumed 98.5 percent of analysis time.
1,500Answer-keyed pairs88.3%Correct suggestions, reranking off77.1%Correct suggestions, reranking on
Two Anthropic models were tested on 36 Javanese speech level questions. Claude Sonnet 4.6 missed 3, Claude Haiku 4.5 missed 10.
36 questionsMultiple-choice questions91.67%Claude Sonnet 4.6 score72.22%Claude Haiku 4.5 score
Seven models tested on one and the same set. An 876-fold gap in output tokens for a single letter answer, and the numbers themselves move between runs.
875.8Tokens, most wasteful model1.0Tokens, leanest model876xGap on the same set
Before writing grammar rules from scratch, we mined LanguageTool's entire rule base to see what Indonesian could borrow. A third of it is unusable for us, and there is no Indonesian module at all.
35Languages with a module0Indonesian modules32.8%Rules impossible to reuse
A 5.8% pass rate. That number is the real price of verification, and the reason "just let AI handle it" is not an answer.
744Corrections mined43Passed verification5.8%Pass rate
We pointed our own spelling checker at writing people actually produced, not at test sentences. It flagged 3,740 things, and most were not errors. This is the record of eight rounds of fixing it.
3,740Flags, round 12,554Flags, round 732%Drop