Insight

How far can AI and our own language tools be trusted?

We test our own correction engine and various outside AI models against Indonesian and regional languages, then publish the numbers as they are. On this page, we sum up each finding in a single picture. The full method and raw data are one click away on every card.

2 findingsEvery finding links to its technical version
FilteredCorrection engineClear filter

2 findings

Correction engineEmpirical
3,741 Round 1 2,729 Round 6
Number of engine flags, round to round

Passed the Exam, Failed on Real Writing

We pointed our own spell checker at real writing people produced, not test sentences. It flagged 3,741 things, and most weren't errors. This is the record of six rounds of fixing it, including a number we miscounted ourselves.

August 17, 2026 · 3 menit bacaRead the finding
CorpusEmpirical
Made it into the test set 43Discarded 701
Of the 744 corrections we mined

We Mined 744 Wikipedia Corrections, Only 43 Survived

Wikipedia's edit history looks like a ready-made source of spelling corrections. Of the 744 we mined, only 43 were genuinely fit to become test items. The rest weren't real corrections.

August 15, 2026 · 2 menit bacaRead the finding