Tested on Javanese Speech Levels, Anthropic's Cheapest AI Got It Wrong 10 Times
Claude Haiku, Anthropic's cheapest model, answered 10 of 36 Javanese politeness-level questions wrong. Claude Sonnet, its expensive counterpart, missed only 3.
We test our own correction engine and various outside AI models against Indonesian and regional languages, then publish the numbers as they are. On this page, we sum up each finding in a single picture. The full method and raw data are one click away on every card.
5 findings
Claude Haiku, Anthropic's cheapest model, answered 10 of 36 Javanese politeness-level questions wrong. Claude Sonnet, its expensive counterpart, missed only 3.
We pointed our own spell checker at real writing people produced, not test sentences. It flagged 3,741 things, and most weren't errors. This is the record of six rounds of fixing it, including a number we miscounted ourselves.
Wikipedia's edit history looks like a ready-made source of spelling corrections. Of the 744 we mined, only 43 were genuinely fit to become test items. The rest weren't real corrections.
We built Javanese test items from grammar-book rules. Six of 40 items were dropped, not because AI answered wrong, but because the form the book called wrong turned out to be widely accepted by native speakers.
We looked for words Indonesians use often that are absent from our dictionary. We found 3,289 candidate new words, but once checked closely, only 899 genuinely qualified. The rest were brand names, place names, or words barely used at all.