Measured Collection

Research

We test how well AI actually understands Indonesian and its regional languages, map the words that trip up native speakers, and publish the numbers along with the raw data and their limits, including when the results are not what we hoped for.

6 articlesEvery article links to its raw data
FilteredDictionaryClear filter

6 articles

Empirical

36% of an AI's answers changed even though the question was identical

Three earlier articles ended with the same admission: every time we read one model's answers and fixed the key, that model gained the most. This time the protocol was designed to make that impossible, and across 135 Javanese words and 6 AI models, the protocol held. What we did not expect: once the instrument stopped moving, the AI models themselves turned out to be unsteady. At temperature 0, which should produce identical answers, 36 percent of answers changed between two consecutive calls.

135Javanese words6AI models36%Answers changed
Empirical

We pit IndoBERT against our own human answer key on the Javanese benchmark

Rankings on our Javanese benchmark have so far been set by a human-written answer key. To check whether that ranking is real, we added a second instrument that never sees the key: IndoBERT, a small 40 MB model that runs on our own machine. All four models came back in the exact same order. But 63 percent of the similarity it measured turned out to be sentence framing, not meaning, and that has to be measured first before a single number can be read.

43Javanese words tested40 MBLocal IndoBERT, second instrument63%Similarity that was framing
Empirical

From 5% to 91%: results of 4 AI models defining 43 Javanese words

Four AI models were asked the meaning of 43 Javanese words absent from NusaX, NusaWrites and Javanese Wiktionary. Mistral Small 3 got 5 percent right, DeepSeek V4 Flash 53 percent, GPT-5.6 Luna 81 percent, Gemini 3.5 Flash 91 percent. The instrument itself was patched twice mid-measurement, and each patch raised scores without a single call being repeated.

43Javanese words91%Max Score4 ModelsMistral, DeepSeek, GPT, Gemini
Empirical

Our answer key is no better than the AI's answers in a Javanese benchmark trial

Example sentences from our Javanese dictionary became the test items, and their Indonesian glosses became the answer key. Two AI models were measured. In a blind adjudication of 20 pairs, our glosses won 4, the model won 5, and the remaining 11 were tied. An answer key that is no better than what it grades measures agreement on word choice, not ability.

11 of 20Blind comparisons tied0 of 20Exact match with our gloss2AI Models measured