Why This Matters
People pick AI models the way they pick a mobile data plan: look at the unit price, take the cheapest. For language work that habit can mislead, and we wanted to know by how much.
We asked 10 AI models from 10 companies the most basic question there is: what does this Javanese word mean in Indonesian? 135 words, exactly the same question for every model, and an answer key already checked by a speaker.
A question like this is good to test because the answer cannot be half right. Unlike translating a sentence, where one sentence can be rendered many equally good ways, the meaning of a word either lands or it does not.
Then we counted two things. How many answers were correct, and how much each correct answer cost. That second figure is rarely computed, and it tells a completely different story from the first.
What We Found
The 10 models are spread wide apart. The highest got 128 words right, the lowest 22, and the rest fill almost every point in between: Ling 3.0 Flash at 70 words, Qwen3-30B-A3B at 52. There are no two neatly separated groups, only one long series running from nearly perfect to nearly clueless.
Gemini 3.5 Flash answered the most questions correctly, 128 of 135 words. It is also the worst value in the whole panel.
Gemma spends $0.0000154 per correct answer. Gemini spends $0.0028749. Both answer the same questions, and the gap between their results is 11 words out of 135.
Put the two rankings side by side and the relationship almost disappears:
3
7
6
2
4
9
10
5
8
1
Gemma 4 31B sits at rank 1 for cost and rank 3 for score. Gemini is the exact mirror image: rank 1 for score and rank 10 for cost. Being cheap per call is no guarantee either. Nemotron 3.5 Lightning is cheap every time it is called, yet it ends up the second most expensive model per correct answer.
Here is why. Nemotron spends tokens reasoning before it answers, and it reasons in English even when the question is in Indonesian. On a 600-token budget every answer came back empty, because the budget ran out before the answer appeared. We raised it to 4,000 tokens so the model could finish. Its real cost was 30 times the estimate we had calculated from its official price table.
That is exactly what makes unit-price comparison misleading: what gets billed is not the question but every word the model produces, including the words that never reach a reader.
To use these numbers in practice, the question is not which model is smartest but how many correct answers the job needs and what the budget is. For work whose output a person checks afterwards anyway, a model that scores slightly lower and costs far less is usually the more sensible choice.
The Limits
These numbers must not be read as a ranking of intelligence. We called the top 3 models 3 times with the same 135 questions, and their own scores moved by as much as 4 words without a single question changing. Any difference smaller than that is not a difference in ability, it is ordinary wobble.
The 135 words we used are not rare ones either. All of them were already on our publicly readable dictionary pages before this testing, so a high score does not necessarily mean a model knows Javanese. It may simply have read our dictionary. To separate the two, we are preparing a control set of 192 words we have never published anywhere.
And model prices change at any time. Every cost figure on this page was measured in August 2026, and a price table can shift whenever a provider decides. What we hope lasts longer is not the number but the way of counting it: compare the cost of a correct result, not the cost of asking once.
How we built the answer keys, the full results for all 10 models, and all of the raw data are in the technical version of this article.
Behind this finding
The technical version has the raw numbers, the test setup, and everything that cannot be concluded from them.