Measured Collection

Research

We test how well AI actually understands Indonesian and its regional languages, map the words that trip up native speakers, and publish the numbers along with the raw data and their limits, including when the results are not what we hoped for.

6 articlesEvery article links to its raw data
FilteredCorrection engineClear filter

6 articles

Empirical

Language models consumed 98.5% of analysis time when added to a spell checker

The hypothesis was reasonable: let a contextual language model pick the best correction candidate based on sentence meaning. Tested A/B with a single switch and an answer key of 1,500 pairs, suggestion accuracy instead fell from 88.3 percent to 77.1 percent, while the model consumed 98.5 percent of analysis time.

1,500Answer-keyed pairs88.3%Correct suggestions, reranking off77.1%Correct suggestions, reranking on