Three earlier articles ended with the same admission: every time we read one model's answers and fixed the key, that model gained the most. This time the protocol was designed to make that impossible, and across 135 Javanese words and 6 AI models, the protocol held. What we did not expect: once the instrument stopped moving, the AI models themselves turned out to be unsteady. At temperature 0, which should produce identical answers, 36 percent of answers changed between two consecutive calls.
135Javanese words6AI models36%Answers changed
Rankings on our Javanese benchmark have so far been set by a human-written answer key. To check whether that ranking is real, we added a second instrument that never sees the key: IndoBERT, a small 40 MB model that runs on our own machine. All four models came back in the exact same order. But 63 percent of the similarity it measured turned out to be sentence framing, not meaning, and that has to be measured first before a single number can be read.
43Javanese words tested40 MBLocal IndoBERT, second instrument63%Similarity that was framing
Four AI models were asked the meaning of 43 Javanese words absent from NusaX, NusaWrites and Javanese Wiktionary. Mistral Small 3 got 5 percent right, DeepSeek V4 Flash 53 percent, GPT-5.6 Luna 81 percent, Gemini 3.5 Flash 91 percent. The instrument itself was patched twice mid-measurement, and each patch raised scores without a single call being repeated.
43Javanese words91%Max Score4 ModelsMistral, DeepSeek, GPT, Gemini
Example sentences from our Javanese dictionary became the test items, and their Indonesian glosses became the answer key. Two AI models were measured. In a blind adjudication of 20 pairs, our glosses won 4, the model won 5, and the remaining 11 were tied. An answer key that is no better than what it grades measures agreement on word choice, not ability.
11 of 20Blind comparisons tied0 of 20Exact match with our gloss2AI Models measured
Mistral Small 3 and Llama 3.1 8B Instruct are sold at the exact same price and are equally fast. On 69 items testing normalization of abbreviated writing to standard Indonesian, one scored 56.52%, the other 26.09%.
69 itemsItems56.52%Mistral Small 326.09%Llama 3.1 8B
Two Anthropic models were tested on 36 Javanese speech level questions. Claude Sonnet 4.6 missed 3, Claude Haiku 4.5 missed 10.
36 questionsMultiple-choice questions91.67%Claude Sonnet 4.6 score72.22%Claude Haiku 4.5 score
A map of our five internal tools and how each one verifies Indonesian, including how we test AI, with a link to the raw data.
3×Runs per figure5Internal tools