The hypothesis was reasonable: let a contextual language model pick the best correction candidate based on sentence meaning. Tested A/B with a single switch and an answer key of 1,500 pairs, suggestion accuracy instead fell from 88.3 percent to 77.1 percent, while the model consumed 98.5 percent of analysis time.
1,500Answer-keyed pairs88.3%Correct suggestions, reranking off77.1%Correct suggestions, reranking on
Before writing grammar rules from scratch, we mined LanguageTool's entire rule base to see what Indonesian could borrow. A third of it is unusable for us, and there is no Indonesian module at all.
35Languages with a module0Indonesian modules32.8%Rules impossible to reuse
A 5.8% pass rate. That number is the real price of verification, and the reason "just let AI handle it" is not an answer.
744Corrections mined43Passed verification5.8%Pass rate
We pointed our own spelling checker at writing people actually produced, not at test sentences. It flagged 3,740 things, and most were not errors. This is the record of eight rounds of fixing it.
3,740Flags, round 12,554Flags, round 732%Drop