Insight

How far can AI and our own language tools be trusted?

We test our own correction engine and various outside AI models against Indonesian and regional languages, then publish the numbers as they are. On this page, we sum up each finding in a single picture. The full method and raw data are one click away on every card.

9 findingsEvery finding links to its technical version

9 findings

BenchmarkEmpirical
Cost rank Score rank
Gemma 4 31B
1
3
Qwen3-30B-A3B
2
7
Ling 3.0 Flash
3
6
GPT-5.6 Luna
4
2
DeepSeek V4 Flash
5
4
Mistral Small 3
6
9
Hunyuan A13B
7
10
Claude Haiku 4.5
8
5
Nemotron 3.5 Lightning
9
8
Gemini 3.5 Flash
10
1
Cost rank against score rank, 1 is best

The Smartest AI on Javanese Words Is 187 Times More Wasteful

We asked 10 AI models the meaning of 135 Javanese words. Gemini 3.5 Flash answered the most correctly, 128 out of 135, and is at the same time the worst value of the lot: 187 times more expensive per correct answer than a model that trails it by just 11 words.

September 22, 2026 · 4 menit bacaRead the finding →
BenchmarkEmpirical
Wording 1 Wording 3
Gemma 4 31B
32
60
Qwen3-30B-A3B
40
46
Ling 3.0 Flash
39
41
Mistral Small 3
39
38
Correct answers out of the exact same 69 sentences

Change the Prompt, and the AI Ranking Flips

We put 4 AI models through the same 69 sentences using three ways of asking. Gemma 4 31B, built by Google, came last under the first way and first under the third, with the questions and the answer key never changing.

September 15, 2026 · 4 menit bacaRead the finding →
BenchmarkEmpirical
11
Different 11Same 27
Answers from ox-alpha and GLM-5.3 on the same 38 questions

Two days before Z.ai spoke up, our data already said this was not GLM-5.3

On August 24 we wrote that ox-alpha was not GLM-5.3 but a relative of it. On August 26, Z.ai announced the model was GLM-5.3-Flash. What took us there was not a hunch, but 38 Javanese questions.

August 30, 2026 · 3 menit bacaRead the finding →
Correction engineStructured Data
32.8
Unusable 32.8Adaptable 64.3
Of the 2,909 LanguageTool rules we mined

Our Closest Language Relative on LanguageTool Only Has 44 Rules

Before writing grammar rules from scratch, we first checked what could be borrowed from LanguageTool, the largest open-source grammar checker there is. It has modules for 35 languages. The world's fourth most spoken language isn't one of them.

August 21, 2026 · 2 menit bacaRead the finding →
BenchmarkEmpirical
Claude Haiku 4.5
72.22
Claude Sonnet 4.6
91.67
Correct answers out of 36 Javanese politeness-level questions

Tested on Javanese Speech Levels, Anthropic's Cheapest AI Got It Wrong 10 Times

Claude Haiku, Anthropic's cheapest model, answered 10 of 36 Javanese politeness-level questions wrong. Claude Sonnet, its expensive counterpart, missed only 3.

August 19, 2026 · 3 menit bacaRead the finding →
Correction engineEmpirical
3,741 Round 1 2,729 Round 6
Number of engine flags, round to round

Passed the Exam, Failed on Real Writing

We pointed our own spell checker at real writing people produced, not test sentences. It flagged 3,741 things, and most weren't errors. This is the record of six rounds of fixing it, including a number we miscounted ourselves.

August 17, 2026 · 3 menit bacaRead the finding →
CorpusEmpirical
Made it into the test set 43Discarded 701
Of the 744 corrections we mined

We Mined 744 Wikipedia Corrections, Only 43 Survived

Wikipedia's edit history looks like a ready-made source of spelling corrections. Of the 744 we mined, only 43 were genuinely fit to become test items. The rest weren't real corrections.

August 15, 2026 · 2 menit bacaRead the finding →
JavaneseEmpirical
Voided, book vs speaker 6Kept in the set 34
Out of the 40 Javanese items we built

6 of 40 Javanese Test Items Voided Because We Trusted the Book Too Much

We built Javanese test items from grammar-book rules. Six of 40 items were dropped, not because AI answered wrong, but because the form the book called wrong turned out to be widely accepted by native speakers.

August 14, 2026 · 3 menit bacaRead the finding →
DictionaryStructured Data
38,559
Never appear 38,559Appear in corpus 33,895
Of our 72,454 dictionary lemmas

Only 899 of Our 3,289 New-Word Candidates Survived the Filter

We looked for words Indonesians use often that are absent from our dictionary. We found 3,289 candidate new words, but once checked closely, only 899 genuinely qualified. The rest were brand names, place names, or words barely used at all.

August 13, 2026 · 3 menit bacaRead the finding →