The book says wrong, the speaker says natural
When we built Javanese speech-level test items, we thought the source for the answer key was obvious: the grammar book. Speech levels have written rules, and written rules can become an answer key.
Six of 40 items were voided because of that belief.
Those items asked which form was correct, with the key taken from the book's rules. Everything was tidy on paper. Once checked against native speakers, the key didn't hold. The form the book called "wrong" turned out to be widely accepted, in some cases even sounding more natural than the form the book prescribed, to the point where the speakers we asked needed a few seconds to even recall that the book's version existed.
We didn't fix the key on those six items. We dropped them. Fixing the key wasn't enough, because the problem wasn't the key, it was the question itself.
Two kinds of "correct" that must not be mixed
From there we separated two things we'd previously treated as one.
The first, prescriptive correctness, anchors to an official document. That's valid for spelling rules: a document sets them, there's one correct form, and anyone can point to it to settle an argument.
The second, speaker acceptability, anchors to native speakers. This is what actually governs Javanese speech levels, since acceptability there is far looser than any table in a book.
An item that mixes the two can't be interpreted at all. If an AI answers "wrong" on an item like that, there's no way to tell whether it genuinely failed to understand the language, or answered exactly like a native speaker while the answer key followed the book. Those two possibilities need very different responses, yet the score looks identical.
The question itself had to change
The fix is small on the surface, large in consequence. We stopped asking "which one is correct" and started asking "which one sounds more natural." The first assumes a single correct form exists, the second measures a relative judgment, which is what a speaker is actually doing when they hear two sentences.
The side effect was interesting: this comparative format turned out to be far more stable to rate. A speaker who hesitates to call a sentence "correct" or "wrong" almost never hesitates to say one sentence sounds more natural than another.
Why this matters beyond Javanese
The same problem waits in every regional language worth measuring. Sundanese, Minangkabau, Buginese, all of them have written traditions that don't always match how their speakers actually talk today.
If test items are built from the grammar book alone, what gets measured is how well an AI has memorized that book, not how well it understands the language. Those are two different things, and the second matters far more to people who actually use the language every day.
This isn't a claim that Javanese grammar books are wrong. Books describe norms, speakers follow habits, and the two can differ without either being mistaken. A language whose norms match its habits exactly is probably a language that's stopped being spoken.
The full method for how we built these items, and all the raw data, are in the technical version below.
Behind this finding
The technical version has the raw numbers, the test setup, and everything that cannot be concluded from them.