When the grammar book is wrong and the native speaker is right
Also available, written for a general reader6 of 40 Javanese Test Items Voided Because We Trusted the Book Too MuchWe assumed the grammar book was the absolute standard until we found 6 of our 40 test items were irrelevant to native speakers. This finding forced us to distinguish between codified rules and living language usage.
Key numbers
Where these come fromMeasured from meta.json run pilot Jawa: parameter reproduksi and 1 more
When we wrote Javanese test items, the source of the answers looked obvious: the grammar book. Speech levels have written rules, and written rules can serve as an answer key.
Six items died because of that assumption.
The answer key was wrong, not the model
Those items asked which form was correct, with keys taken from prescriptive rules. On paper everything was tidy.
Checked against speaker acceptability, the keys did not hold. The form the book called wrong turned out to be widely accepted. In several cases it sounded more natural than the prescribed form, and the speakers we asked needed a few seconds to recall that the textbook version even existed.
Those six items were not fixed. They were discarded. Correcting the key was not enough, because the problem was not in the key but in the question.
Two kinds of correctness that must not be mixed
Out of that we made explicit a distinction we had been treating as one thing.
Prescriptive correctness anchors to an official document. This is valid for spelling rules: a document defines them, the correct form is singular, and anyone can point at it to settle an argument.
Speaker acceptability anchors to native speakers. This is what governs Javanese speech levels, where acceptability is far looser than any table.
A set that mixes the two cannot be interpreted at all. When a model answers "wrongly", you have no way to tell whether it failed to understand the language or answered exactly as a native speaker would while your key followed the book. Those two possibilities demand opposite responses, and the score looks identical.
The question itself had to change
The fix is small on the surface and large in consequence. We stopped asking "which one is correct" and started asking "which one is most appropriate", or asking for a location: which element does not fit.
The first presumes a single correct form exists. The second measures relative judgement, which is what a speaker actually performs when hearing two sentences.
The side effect was interesting. The comparative format turned out to be far more stable to rate, because a speaker who hesitates to call a sentence right or wrong almost never hesitates to call one sentence more natural than another.
It is also why we now record norm_basis as an explicit column on every set rather than as an unwritten understanding. Unwritten understandings survive until the next person adds an item, then vanish.
What this does not show
This is not a claim that Javanese grammar books are wrong. Books describe norms, speakers follow habits, and the two can differ without either being mistaken. A language whose norms match its habits exactly is probably a language that has stopped being used.
The six-out-of-40 figure does not generalise either. It is a count from one authoring round, with one author. Another round with another author could land somewhere else.
And because our set uses a single dialect variant, the line we drew between accepted and not is Yogya-Solo's line. East Javanese or Banyumasan speakers may draw it elsewhere. That is a limitation we are recording, not one we have resolved.
Data & provenance
Limitations
One dialect variant only (Yogya-Solo); speakers from East Javanese or Banyumasan areas may judge differently, and that has not been tested. Verification was done by one person with no comparison rater. Six items dropped out of 40 is a count from one authoring round, not a generalisable error rate.
Revision history
- Published
- Last updated