BenchmarkEmpirical
We asked 10 AI models the meaning of 135 Javanese words. Gemini 3.5 Flash answered the most correctly, 128 out of 135, and is at the same time the worst value of the lot: 187 times more expensive per correct answer than a model that trails it by just 11 words.
September 22, 2026 · 4 menit bacaRead the finding →
BenchmarkEmpirical
We put 4 AI models through the same 69 sentences using three ways of asking. Gemma 4 31B, built by Google, came last under the first way and first under the third, with the questions and the answer key never changing.
September 15, 2026 · 4 menit bacaRead the finding →
BenchmarkEmpirical
On August 24 we wrote that ox-alpha was not GLM-5.3 but a relative of it. On August 26, Z.ai announced the model was GLM-5.3-Flash. What took us there was not a hunch, but 38 Javanese questions.
August 30, 2026 · 3 menit bacaRead the finding →
BenchmarkEmpirical
Claude Haiku, Anthropic's cheapest model, answered 10 of 36 Javanese politeness-level questions wrong. Claude Sonnet, its expensive counterpart, missed only 3.
August 19, 2026 · 3 menit bacaRead the finding →