Empirical· 8 min read

Our answer key is no better than the AI's answers in a Javanese benchmark trial

Javanese has almost no benchmark number anyone can cite. We hold a dictionary with 391 human-verified example sentences, and turning them into test items looked like a legitimate shortcut. What we had not accounted for was the possibility that a model's answers would be as good as our own answer key.

Measured from export-jv-id-pilot-hasil-model.csv and 2 more

11 of 20
Blind comparisons tied
0 of 20
Exact match with our gloss
2
AI Models measured
360 times
Cost gap

A language with no number yet

If someone asks how well AI models understand Javanese, we have no figure to point at. Test sets for Indonesia's regional languages do exist, but the ones we found target other tasks, mainly classification and sentiment analysis, rather than everyday sentence comprehension.

We hold material that looked like a fit. Our dictionary carries 391 human-checked Javanese senses, and every one of them comes with an example sentence in a uniform pattern:


Kowe selak telat, ndang mreneo saiki! (Kamu hampir terlambat, cepat ke sini sekarang!)

Ojo ngono karo kancane, kudu sing rukun. (Jangan begitu sama temannya, harus rukun.)

The Javanese sentence comes first, the Indonesian gloss in parentheses. Already paired, already checked, and already published in our Javanese dictionary. Turning that into a translation test looked like a legitimate shortcut: the Javanese sentence becomes the question, the gloss becomes the answer key.

This article reports what happened when the shortcut was tested, and why it cannot go forward in that shape.

The benchmark items

Those dictionary pages yielded 399 items. Each one is a single instruction plus a single Javanese sentence, and the answer key is the gloss that already accompanied that sentence:


Translate the following Javanese sentence into Indonesian.

Answer ONLY with the translated sentence, no explanation.



Sentence: 'Aku mrene nggawa oleh-oleh kanggo kowe.'



Key     : Aku ke sini membawa oleh-oleh untukmu.

Twenty items were drawn for the first measurement, sampled in strata so that easy and hard sentences were both represented.

The methodological problem: a single answer key

The scorer we already use for other sets compares an answer to its key exactly. For text normalisation that works, because the correct answer really is single.

On translation it collapses at once. Of the 20 items run against the first model, 0 answers matched our gloss exactly. Not one.

Before blaming the model, we tested the scorer itself. Every gloss was altered in ways that are certainly still correct: clause order flipped around the comma, punctuation removed, synonyms swapped, scaffolding words like itu and yang dropped. That produced 945 variants, all of them still valid translations.

ScorerCorrect variants that passed
Exact match38.2%
Content-word overlap98.2%
chrF, character n-gram F-score98.4%

Exact match rejects 6 out of every 10 correct translations. Every synonym swap, every reordered clause, every dropped scaffold scores zero.

This is not a new discovery. Machine translation has long recorded that scoring a translation against a single reference is inadequate. The most cited criticism appeared in 2006, when Callison-Burch, Osborne, and Koehn showed that BLEU scores do not always track human judgement.

WMT, the field's annual benchmark conference, now decides its official ranking through professional human annotators. The WMT24 general translation task report ran it with an error span annotation protocol rather than an automatic metric.

We walked into it anyway, and the next section explains how.

Measuring two AI models directly

Mistral Small 3 vs Gemini 3.5 Flash - Ilustrasi
Mistral Small 3 vs Gemini 3.5 Flash - Illustration

Two models ran on 20 identical items at temperature 0. Mistral Small 3, from Mistral AI, represents the cheap low-latency class. Gemini 3.5 Flash, from Google, represents the far more expensive reasoning class.

Mistral Small 3Gemini 3.5 Flash
Exact match out of 2005
Content-word overlap, median0.3750.857
chrF, median0.4870.879
Echoed the Javanese without translating40
Declined to answer10
Median latency853 ms3,943 ms
Cost per call$0.0000121$0.0043511

The cost gap is 360 times, the latency gap 4.6 times. The cheap model did not merely score lower: it handed the Javanese sentence back untouched on 4 items, and once replied that it did not understand what the sentence meant.

Up to this point the numbers read cleanly and the conclusion looks obvious. The next section undoes it.

Gemini's first number was wrong, and the mistake was ours

The first benchmark run against Gemini used a limit of 200 tokens per answer, and produced a median of 0.571. That figure placed it only slightly above the cheap model, and the conclusion was nearly written: the gap between model classes is smaller than expected on Javanese.

What saved it was one column that happened to be read. The output token count came back as 196 across all 20 calls, the exact same number, while the visible text ran to about 20 characters:


Jika berkenan, penyampa

Mungkin bes

Aduh, sakitnya bad

Gemini 3.5 Flash is a reasoning model, a thinking model. It spent nearly 190 tokens reasoning before writing its answer, then hit the ceiling before the sentence finished. A truncated answer scores as wrong under any scorer, and not a single error message appeared.

With the limit raised to 1,024, the same median rose from 0.571 to 0.857. The conclusion that nearly went out was not merely off, it pointed the wrong way.

This trap has been written down in our own internal benchmark notes for a long time, and we walked into it regardless. The pilot tooling now refuses to report any figure at all if an answer hits the token ceiling.

The blind comparison, and the number that voids the whole design

One objection remained, and it is the decisive one. Every figure in the table above measures how closely a model's answer resembles our gloss, and that is only meaningful if our gloss is in fact better.

To test it, 20 pairs were assembled blind. Each pair held the dictionary gloss and Gemini's answer with no label of which was which, and the positions were shuffled per item so that no pattern could be read off. The adjudicator was the dictionary compiler, and the question was single: which translation is better.

Judged betterCount
Our dictionary gloss4
The model's answer5
Equally good11
11 of 20
comparisons came back tied, and that number voids the whole design: an answer key no better than what it grades measures agreement on word choice, not ability

The 5 to 4 split among the rest carries no conclusion. Out of the 9 items that produced a decision, that spread is precisely the most likely outcome if both sides are equivalent. The numbers carry exactly one thing: at the level of a full sentence, the model's answers and our glosses cannot be told apart.

The adjudicator's notes say it more plainly than any table:


sama saja bagusnya

makna sama baiknya

sama hasilnya

Four conclusions these numbers do not carry

This does not mean AI models now match Javanese speakers. Only one direction was measured, from Javanese into Indonesian, and that direction is far lighter than the reverse. Translating into Javanese requires choosing a speech level, and that was not tested here at all.

This does not mean our dictionary glosses are weak. Both sides were judged equivalent, not one judged poor. Even the glosses that lost, lost on word choice rather than on meaning.

This does not mean the cheap model is useless. It is 360 times cheaper and 4.6 times faster, and for work that is not regional-language translation those figures decide the matter.

This does not mean a Javanese benchmark cannot be built. What failed is the unit of measurement, not the material.

Where the dictionary still wins

The four items our glosses won share one pattern, and that pattern is what saves the rest of the project.

Not one of them was won on sentence structure. Every one was won on a single word:

WordWhat the model missed
mriyangrendered as "sakit", when it means a fever
bladakhanded back untranslated
tampahhanded back untranslated; our gloss says a bamboo winnowing tray
mbok menawarendered as "mungkin", when it means perhaps

The word bladak means a fritter in the Semarang region and its surroundings, and as far as we have searched it appears in no Javanese dictionary available online. The model does not know it because there was nowhere to learn it. Every Javanese lemma we have collected is open at our Javanese topic page.

That is where a human-curated dictionary holds its value, and it is not on ordinary sentences. The question now worth measuring is no longer how well models translate Javanese, since that already has an answer, but how far the tail of regional vocabulary reaches that no model holds.

That question is also far cheaper to measure. It needs no gloss as a reference, produces no ties, and requires no human adjudicator per item. Its answer is binary: the model knows the word, or it does not.

Three things that apply beyond Javanese

Test the scorer before testing the model. Variants that are certainly still correct can be generated mechanically from an answer key you already hold, without calling a model at all. A scorer that rejects 6 out of 10 correct translations will show itself before a single cent is spent.

Raise the token limit before testing a reasoning model, then check the output token counts. A count identical across every call is the signature of truncated answers, and truncated answers blame the model for the tester's mistake.

Test your own answer key blind. If a human adjudicator calls most pairs tied, that answer key does not separate quality, and no scorer can repair it. This check takes 20 pairs and one hour, and it belongs before all the other work.

Data & provenance

Limitations

The blind adjudication covered 20 pairs with a single adjudicator, the person who compiled the dictionary. There is no inter-annotator agreement figure, the same gap we have already recorded for our entire Javanese dimension. For a negative finding like this one, namely that the instrument fails to separate, that sample is adequate. For a positive claim about which model is better, it is not.

The 5 to 4 split among the items that were not tied is nobody's win. Out of the 9 items that produced a decision, that spread is the most likely outcome if both sides are in fact equivalent. The numbers carry exactly one conclusion: indistinguishable.

Two models, one translation direction, one language. Translation from Indonesian into Javanese has not been tested at all, and that direction is far more demanding because it involves speech levels.

The example sentences have been published on our public dictionary pages since February 2026, so not one of them can be treated as a held-out item. Whether the models have read them has not been measured, and that measurement is separate from this article.

Two of the 20 items use a gloss written by the benchmark author rather than one already in the dictionary, because those sentences carried no gloss. One of them falls among the glosses judged better, so the figure of 4 is really 3 dictionary glosses and 1 gloss written for this measurement. Since the 5 to 4 split carries no conclusion anyway, that composition does not change the finding, but readers should know it.

Of the 391 senses, 365 began as AI-generated drafts that were then human-reviewed. That review makes the meaning sound, but the sentence form still originates from a model, which may make the items easier for models. A comparison group of 26 non-AI senses exists but has not been used.

Revision history

Published
Last updated

Terms in this article

chrF
A similarity measure between two sentences that compares CHARACTER chunks rather than word chunks. Working at the character level, it still credits an answer that is correct but differently inflected, and it needs no hand-maintained function word list. Its weakness is a floor: two entirely unrelated sentences still score around 0.18 because the same letters are bound to appear in both, so the raw figure must not be read as though zero means zero.
error rate
The percentage of wrong answers out of everything a model attempted on a test set. The inverse of score: a 91.7% score means an 8.3% error rate. Used instead of score when the point being made is how often a model gets things wrong rather than how often it succeeds, especially when comparing how far apart two models are.
gloss
A short equivalent placed alongside a foreign word or sentence to convey its meaning, usually in parentheses. It differs from a full translation: its purpose is to make an example readable to someone who does not know the source language, not to produce a sentence that stands on its own. That difference becomes a problem when a gloss is used as a benchmark answer key, since one sentence can be validly translated many ways.
latency
The time from sending a request to receiving the complete answer. It is not the same as the model's typing speed: this number also contains the wait when a provider is overloaded, so a handful of unlucky calls can pull the average far above the everyday experience.
lemma
The base form of a word that heads a dictionary entry, and the form someone looks up. A single lemma can cover several senses at once, and each sense can carry its own example sentence.
reasoning model
An AI model that works out its own reasoning steps before writing an answer. Those steps count as output tokens even though they are not shown, so two things change at once: the cost rises, and a token limit that feels generous for an ordinary model can cut the answer off mid-sentence. A truncated answer scores as wrong under any scorer, and produces no error message.
temperature
A dial that controls how willing a model is to pick a less likely word. Set to zero, it always takes the most likely option, so the same question tends to get the same answer. Raised, answers become more varied but harder to repeat. Benchmarks use zero so what gets measured is the model's ability rather than the luck of its sampling.
token
The chunk of text a model counts in, roughly a syllable up to a short word. AI services are priced per million tokens, with input and output billed separately, so a rambling answer genuinely costs more than a concise one.
WMT
The annual machine translation conference, which also runs an open shared task: participants submit translation systems for the same language pairs and the results are ranked together. Its official ranking is now decided by professional human annotators rather than an automatic measure, and that shift is itself an admission that automatic measures have limits.

Written by the author for this article, not taken from a dictionary entry.