Insight

How far can AI and our own language tools be trusted?

We test our own correction engine and various outside AI models against Indonesian and regional languages, then publish the numbers as they are. On this page, we sum up each finding in a single picture. The full method and raw data are one click away on every card.

2 findingsEvery finding links to its technical version
FilteredOur own failuresClear filter

2 findings

BenchmarkEmpirical
Wording 1 Wording 3
Gemma 4 31B
32
60
Qwen3-30B-A3B
40
46
Ling 3.0 Flash
39
41
Mistral Small 3
39
38
Correct answers out of the exact same 69 sentences

Change the Prompt, and the AI Ranking Flips

We put 4 AI models through the same 69 sentences using three ways of asking. Gemma 4 31B, built by Google, came last under the first way and first under the third, with the questions and the answer key never changing.

September 15, 2026 · 4 menit bacaRead the finding →
Correction engineEmpirical
3,741 Round 1 2,729 Round 6
Number of engine flags, round to round

Passed the Exam, Failed on Real Writing

We pointed our own spell checker at real writing people produced, not test sentences. It flagged 3,741 things, and most weren't errors. This is the record of six rounds of fixing it, including a number we miscounted ourselves.

August 17, 2026 · 3 menit bacaRead the finding →