Insight

How far can AI and our own language tools be trusted?

We test our own correction engine and various outside AI models against Indonesian and regional languages, then publish the numbers as they are. On this page, we sum up each finding in a single picture. The full method and raw data are one click away on every card.

2 findingsEvery finding links to its technical version
FilteredCorpusClear filter

2 findings

CorpusEmpirical
Made it into the test set 43Discarded 701
Of the 744 corrections we mined

We Mined 744 Wikipedia Corrections, Only 43 Survived

Wikipedia's edit history looks like a ready-made source of spelling corrections. Of the 744 we mined, only 43 were genuinely fit to become test items. The rest weren't real corrections.

August 15, 2026 · 2 menit bacaRead the finding
DictionaryStructured Data
38,559
Never appear 38,559Appear in corpus 33,895
Of our 72,454 dictionary lemmas

Only 899 of Our 3,289 New-Word Candidates Survived the Filter

We looked for words Indonesians use often that are absent from our dictionary. We found 3,289 candidate new words, but once checked closely, only 899 genuinely qualified. The rest were brand names, place names, or words barely used at all.

August 13, 2026 · 3 menit bacaRead the finding