Seven AI models tested on a single letter answer with an 876-fold token gap
Seven models tested on one and the same set. An 876-fold gap in output tokens for a single letter answer, and the numbers themselves move between runs.
We test how well AI actually understands Indonesian and its regional languages, map the words that trip up native speakers, and publish the numbers along with the raw data and their limits, including when the results are not what we hoped for.
Seven models tested on one and the same set. An 876-fold gap in output tokens for a single letter answer, and the numbers themselves move between runs.
Before writing grammar rules from scratch, we mined LanguageTool's entire rule base to see what Indonesian could borrow. A third of it is unusable for us, and there is no Indonesian module at all.
A 5.8% pass rate. That number is the real price of verification, and the reason "just let AI handle it" is not an answer.
We pointed our own spelling checker at writing people actually produced, not at test sentences. It flagged 3,740 things, and most were not errors. This is the record of eight rounds of fixing it.
Six of our 40 Javanese test items were dropped. Not because models answered them wrongly, but because the textbook rule contradicted what speakers actually accept.
We matched 72,454 lemmas against two million Indonesian sentences. Half never appear, and our own counter-list turned out to be flawed.
A map of our five internal tools and how each one verifies Indonesian, including how we test AI, with a link to the raw data.