From 5% to 91%: results of 4 AI models defining 43 Javanese words
The previous article ended with our own instrument failing: scoring sentence translations against a single dictionary gloss turned out to measure agreement on word choice, not ability. What remained was a question whose answer is binary and cannot end in a tie, namely whether a model knows what a word means. We asked it for 43 words that have nowhere to be learned from.
Key numbers
Where these come fromMeasured from export-jv-arti-kata-hasil-model.csv and 2 more
A question whose answer cannot end in a tie
Continuing from our previous Javanese benchmark measurement, which ended with our own instrument failing, though still holding something worth building on.
The old design scored sentence translations against a single dictionary gloss, and it failed for a reason machine translation has long known: one sentence can be translated many equally correct ways, so scoring it against a single reference measures agreement on word choice. In a blind comparison of 20 pairs, 11 came back tied.
One sentence can be translated many equally correct ways. Scoring it against a single reference does not measure ability, it measures agreement on word choice.
The question "what does this word mean" carries none of that. Its answer is binary, there are no ties, and there is no wording to argue over. What remains is one thing a machine can check, namely whether the Indonesian equivalent is named.
What makes the question worth asking is the words we chose to test it with: 43 words picked precisely for their rarity.
We did this over and over
This section explains how the benchmark was actually run, because the numbers below cannot be read without it.
Asking the models is the easiest and cheapest part. All 172 calls in this measurement finished in minutes and cost $0.1339. What takes time is everything that comes after.
Each time a model finished answering, its answers were assembled into a review sheet: one block per word, holding the dictionary definition, the keys in force, and the model's answer verbatim. A Javanese speaker read them one by one and answered a deliberately narrow question: is this answer correct?
What made it tiring was not the count but how often the answer was "correct, just not with the word we listed". A model answers guntur while our key says gemuruh. It answers perselingkuhan while our key says selingkuhan. It answers aduh for an interjection whose key we had written as interjeksi, which is its word class rather than its meaning.
Every case like that is not a lost point but an instrument caught out. The missing equivalent gets added, then every model is rescored with the same keys, including the ones measured earlier. That happened twice in this measurement.
The sequence, where each arrow is one round of human reading:
2 models (Mistral, Gemini) → single keys proved too narrow
↓ keys widened into sets
1 model (GPT-5.6 Luna) → new equivalents surfaced again
↓ keys extended by the speaker
1 model (DeepSeek V4 Flash) → tested against finished keys
And one addition was rejected after being measured, which is told below. Adding keys does not always make the instrument better.
Words with nowhere to be learned from
Of the 386 Javanese lemmas in our dictionary, each word was looked up in the three open Javanese sources most likely to be in a model's training data: the NusaX lexicon and corpora, the NusaWrites corpora, and every entry title in Javanese Wiktionary.
48 lemmas appear in none of the three. Five were dropped because their example sentences turned out to be Indonesian, leaving 43.
| Lemma | Meaning |
|---|---|
bladak | a fritter from the Semarang region, batter with cabbage and bean sprouts |
gondes | a Central Javanese term for a young man of a certain appearance |
kemecer | to want something badly, usually on seeing food |
semribit | of wind, blowing gently |
That absence is not proof the word is missing from training data, since the web is far larger than these three sources. What we can state is narrower and more checkable: these words are absent from the curated Javanese sources available to anyone.
Four models, one question
Each model was asked the same thing, at temperature 0, one attempt:
Apa arti kata Jawa berikut dalam bahasa Indonesia?
Jawab singkat, maksimal satu kalimat.
Kata: bladak
| Model | Correct | Score | Cost per call | Median latency |
|---|---|---|---|---|
| Mistral Small 3, Mistral AI | 2 of 43 | 5% | $0.0000110 | 1,403 ms |
| DeepSeek V4 Flash 0731, DeepSeek | 23 of 43 | 53% | $0.0000582 | 4,670 ms |
| GPT-5.6 Luna, OpenAI | 35 of 43 | 81% | $0.0000921 | 3,708 ms |
| Gemini 3.5 Flash, Google | 39 of 43 | 91% | $0.0029536 | 3,418 ms |
The whole measurement cost $0.1339, and 95 percent of that went to one model.
What is interesting is not the order but the distance. Gemini costs 268 times more per call than Mistral, and GPT-5.6 Luna reaches 81 percent at 32 times less than Gemini. The most expensive model is not always the only one that can do the job.
Our benchmark was revised 2 times, and that is part of the finding
This section reports how the numbers above changed before they could be published.
First patch: one answer key is not enough. The initial scoring checked whether a single equivalent appeared in the answer. Under that rule Gemini scored 60 percent. Once the keys were widened into sets of equally valid equivalents, the exact same answers scored 91 percent.
Those 31 points touched not a single call. They are entirely an instrument defect, and its direction is always the same: a narrow key punishes the model that answers correctly with another word, never the model that answers wrongly.
Second patch: the keys were reviewed by a speaker. After GPT-5.6 Luna was measured, all of its answers were read one by one and the missing equivalents added. Luna's score rose from 67 percent to 81 percent, again without a single call repeated. The other two models did not move at all under the same patch, because the equivalents added were the ones only Luna used.
One addition was rejected after being measured. The word jawa was proposed as an equivalent for one lemma, and it turned out to appear in 37 percent of the answers across all lemmas, since the question itself reads "what does this Javanese word mean". A key like that passes a model that understands nothing.
Every time the keys widen, the chance of hitting one by accident widens too. We measured it by crossing each lemma's keys against every other lemma's answers: 23 of 1,806 pairs, or 1.3 percent, up from 0.8 percent before the keys were extended. Still far below the gaps between models, and that figure stops being meaningful past roughly 5 percent.
One model was never read, and it shows in its number
The four models did not receive equal treatment, and that has to be stated before the numbers are quoted.
The answers of Mistral Small 3 and Gemini 3.5 Flash were read one by one in the first round, and GPT-5.6 Luna in the second. Each reading added equivalents that were missing. DeepSeek V4 Flash 0731 is the only one never read, since it was tested last against keys that were already finished.
To estimate how much that matters, we counted answers scored wrong that nonetheless contain a word from the lemma's own definition, a strong sign that what missed was the key rather than the model:
| Model | Scored wrong | Possibly misjudged | Upper bound |
|---|---|---|---|
| Mistral Small 3 | 41 | 4 | 14% |
| DeepSeek V4 Flash 0731 | 20 | 6 | 67% |
| GPT-5.6 Luna | 8 | 3 | 88% |
| Gemini 3.5 Flash | 4 | 1 | 93% |
The widest headroom belongs to DeepSeek, 14 points, and it exists precisely because its answers were never read. Its 53 percent is a floor, not a verdict.
The other three have been through a reading, so their remaining headroom is more likely to come from equivalents nobody thought of than from unequal treatment.
One word that made all four models confabulate
One word deserves its own telling.
Bladak is a fritter from the Semarang region, made from batter with cabbage and bean sprouts. As far as we have searched, it appears in no Javanese dictionary available online.
Mistral Small 3 "banyak" (many)
DeepSeek V4 Flash "tiba-tiba" (suddenly)
GPT-5.6 Luna "terbuka" (open, clearly visible)
Gemini 3.5 Flash "lantai" (floor, a landing board)
Four models, four different answers, all of them wrong. And not one said it did not know.
That behaviour differs in kind from a low score. A model that answers "I do not know this word" tells its user it has left its range.
A model that invents confidently gives no signal at all, and the user finds out only when it is too late.
What these numbers do not carry
This does not mean AI models cannot handle Javanese. These 43 words were chosen precisely for their rarity. On everyday Javanese vocabulary, our previous measurement found model answers indistinguishable from our own dictionary glosses.
This does not mean the order of these four models holds for other tasks. What was measured is recognition of rare vocabulary in one regional language. A model that leads here need not lead at reasoning, writing, or another language.
This does not mean the gap between two adjacent models is established. With 43 lemmas, one lemma is worth 2.3 points, and without repeats we cannot separate a gap of a few lemmas from noise.
This does not mean the more expensive model is always better. GPT-5.6 Luna reaches 81 percent at 32 times less than Gemini.
What makes this measurement cheap
The design that failed in the previous article demanded a human adjudicator for every item, because no single reference was valid. This one does not. Once the equivalents are written, scoring runs by itself, and a speaker is needed only when a model uses an equivalent that has never appeared before.
The whole four-model measurement cost $0.1339. What is expensive is not the calls but the human reading that follows, and that is what most deserves saving.
Three things apply beyond Javanese. Measure the scorer's chance floor every time the keys widen, because looser keys raise scores and raise accidental hits at the same time, and both look identical in a results table. Record which models have had their output read by a human, because a reviewed model always leads an unreviewed one, and that lead belongs to the instrument. And count how often a model admits it does not know, because that number tells you something the score does not.
Data & provenance
Limitations
Three models had their answers read by a speaker one by one: Mistral Small 3 and Gemini 3.5 Flash in the first round, then GPT-5.6 Luna in the second. Each reading added equivalents that were missing from the key. DeepSeek V4 Flash 0731 is the only one never read, so its answers that are correct but use a new equivalent are still scored wrong. Its 53 percent is a floor with the widest headroom of the four.
A word's absence from three open sources is not proof that it is absent from any model's training data. Models are trained on a far larger web, most of which we cannot inspect. What is measured here is rarity within curated Javanese sources that anyone can check.
The answer keys were written by a single speaker, the person who compiled the dictionary, with no inter-annotator agreement figure. Of the 43 lemmas, 39 have keys a speaker has touched and 4 still come from the machine.
One attempt per lemma per model at temperature 0. Between-attempt variance is unmeasured, so a gap of a few lemmas between two adjacent models cannot be separated from noise.
Scoring checks whether one of the listed equivalents appears in the answer. A model that answers correctly with an equivalent not yet listed is still scored wrong, and that is exactly the mechanism that moved Luna's figure after review.
Revision history
- Published
- Last updated
- never revised since publication
Terms in this article
- gloss
- A short equivalent placed alongside a foreign word or sentence to convey its meaning, usually in parentheses. It differs from a full translation: its purpose is to make an example readable to someone who does not know the source language, not to produce a sentence that stands on its own. That difference becomes a problem when a gloss is used as a benchmark answer key, since one sentence can be validly translated many ways.
- latency
- The time from sending a request to receiving the complete answer. It is not the same as the model's typing speed: this number also contains the wait when a provider is overloaded, so a handful of unlucky calls can pull the average far above the everyday experience.
- lemma
- The base form of a word that heads a dictionary entry, and the form someone looks up. A single lemma can cover several senses at once, and each sense can carry its own example sentence.
Written by the author for this article, not taken from a dictionary entry.