Insight

How far can AI and our own language tools be trusted?

We test our own correction engine and various outside AI models against Indonesian and regional languages, then publish the numbers as they are. On this page, we sum up each finding in a single picture. The full method and raw data are one click away on every card.

1 findingsEvery finding links to its technical version
FilteredAnthropicClear filter

1 findings

BenchmarkEmpirical
Claude Haiku 4.5
72.22
Claude Sonnet 4.6
91.67
Correct answers out of 36 Javanese politeness-level questions

Tested on Javanese Speech Levels, Anthropic's Cheapest AI Got It Wrong 10 Times

Claude Haiku, Anthropic's cheapest model, answered 10 of 36 Javanese politeness-level questions wrong. Claude Sonnet, its expensive counterpart, missed only 3.

August 19, 2026 · 3 menit bacaRead the finding