Our prediction failed: there is no dividing line between AI that knows Javanese words and AI that does not
Instead of proving a chasm between AI that knows Javanese and AI that is blind to it, our final step knocked down that boundary line.
Key numbers
Where these come fromMeasured from export-jv-B-hasil-model.csv and 3 more
What we predicted, written before a single call
In our previous article we tested 6 AI models on 135 Javanese words. After that article went out we added Nemotron 3.5 Lightning from NVIDIA as a seventh model, with no plan beyond widening the panel. The scores of those 7 models had a suspicious shape:
The figures in that chart have been recomputed against the final keys so they can be checked against the data we published. When the prediction was written, Nemotron sat one lemma lower, and the shape of the gap was the same.
5 models clustered at the top, 2 at the bottom, and not a single model landed between 20 and 63 percent. The tempting reading: a model either has Javanese vocabulary or it does not, with nothing in between.
The problem is that with 7 models a shape like that can come from chance. So before adding any further model we wrote 4 predictions, locked the document with a date on it, and only then called the models:
- The gap between 20 and 63 percent stays empty; the 3 new models land in one cluster or the other, not between them.
- Qwen3-30B-A3B lands in the top cluster even though its architecture is shaped exactly like Nemotron, which got only 27 lemmas right.
- Ling 3.0 Flash lands in the top cluster, and because its price is the lowest it takes the best cost per correct answer away from Gemma 4 31B.
- Hunyuan A13B lands below 20 percent even though it is the largest of the three new models.
We also wrote down what we did not claim, and that part matters most. We had no basis for aiming at the middle. Our earlier testing had already found that a model's published specification does not predict its behaviour, and a 4-lemma pilot gave us no signal either: the results were 0, 1 and 3 out of 4, with nothing in the middle. All we could do was spread the sample.
3 models we chose to break our own prediction
All three come from companies that had never appeared in our panel, and each one tests something different.
Qwen3-30B-A3B from Alibaba was chosen because its architecture is shaped exactly like Nemotron 3.5 Lightning, which had just joined the panel: both are mixture-of-experts models with 30 billion total parameters and 3 billion active. Nemotron answered 27 lemmas correctly. If Qwen lands far above it, the shape of the architecture is not what decides.
Ling 3.0 Flash from Ant Group barely appears in any benchmark, and it costs $0.0000117 per call, among the cheapest we have ever called.
Hunyuan A13B from Tencent is a mixture-of-experts model with 80 billion total parameters and 13 billion active, far larger than the other two.
We rejected 4 other candidates before running anything. ByteDance Seed 1.6 Flash produced 426 reasoning tokens and then returned an empty answer. Qwen3-14B spent 10,124 milliseconds on a one-word question. StepFun and Xiaomi ran at 4,190 and 3,683 milliseconds without offering a new angle to test.
The result: the gap filled
3 models, 405 calls, $0.0051 in total.
| Model | Correct out of 135 | Percent correct | Semantic percentile |
|---|---|---|---|
| Gemini 3.5 Flash | 128 | 94.8% | 94.7% |
| GPT-5.6 Luna | 118 | 87.4% | 90.7% |
| Gemma 4 31B | 117 | 86.7% | 92.0% |
| DeepSeek V4 Flash | 114 | 84.4% | 92.9% |
| Claude Haiku 4.5 | 86 | 63.7% | 84.9% |
| Ling 3.0 Flash | 70 | 51.9% | 81.8% |
| Qwen3-30B-A3B | 52 | 38.5% | 72.6% |
| Nemotron 3.5 Lightning | 27 | 20.0% | 64.8% |
| Mistral Small 3 | 24 | 17.8% | 49.9% |
| Hunyuan A13B | 22 | 16.3% | 63.1% |
4 models in that table did not exist in our panel when the previous article was published. Nemotron 3.5 Lightning joined first and helped form the very gap we set out to test; the other 3, shown in bold, were chosen after the prediction was locked and precisely in order to break it.
Ling landed at 51.9 percent and Qwen at 38.5 percent. Both fell inside the gap we had called empty, on the first attempt.
The 10-model series now rises smoothly from 16.3 to 94.8 percent with no break anywhere. What looked like two clusters turned out to be the shape of our own model list, not the shape of the actual distribution.
Structure visible in a small panel can come from how the sample was picked, and the only way to find out is to try to break it on purpose.
The first prediction failed, and with it the reading that a model either has Javanese vocabulary or does not, with nothing in between. The fourth prediction was right: Hunyuan A13B landed at 16.3 percent, the lowest of the 10 models, despite being the largest of the three new ones. A large company with a large model is no guarantee of Javanese vocabulary. Predictions two and three are covered in the next two sections, and both missed.
What still stands: architecture shape decides nothing
The second prediction missed on its number, but the question it was built to answer came back with a cleaner answer than we expected.
We predicted that Qwen3-30B-A3B would land in the top cluster. It landed at 52 lemmas, far below that. Yet Qwen and Nemotron 3.5 Lightning carry exactly the same architecture shape, 30 billion total parameters with 3 billion active, and they are nearly 2 times apart: 52 lemmas against 27 lemmas.
So the conclusion we were after still holds, only the direction of our guess was wrong. The number of active parameters explains nothing about which model knows what gangsal means and which does not. Hunyuan A13B confirms it from the other side: the largest of the three new models landed lowest.
How far scores move when the questions are repeated
Before comparing models that sit close together, we needed to know how far a score moves when nothing changes at all. We called the top 3 models 2 more times, with the same 135 questions, at the setting that should make a model always pick the most likely answer.
| Model | Three runs | Spread | Answer text differs | Verdict flips |
|---|---|---|---|---|
| GPT-5.6 Luna | 118 · 121 · 119 | 3 | 74% | 8% |
| Gemma 4 31B | 117 · 115 · 116 | 2 | 65% | 4% |
| DeepSeek V4 Flash | 114 · 110 · 110 | 4 | 84% | 7% |
The widest spread is 4 lemmas, exactly the figure we reported earlier for the weakest model in the panel. That number answers a question we left open at the time: a spread that size is not a property of weak models alone, since all three models in this table sit in the top cluster.
The consequence is immediate. Luna to Gemma is 1 lemma, Gemma to DeepSeek is 3 lemmas, and both are smaller than the wobble of a single model on its own. The order inside the top cluster cannot be read as an order of ability. What can be read is the distance between clusters, for instance Gemini at 128 lemmas against Ling at 70 lemmas.
The highest-scoring model costs the most per correct answer
A figure people rarely compute: not cost per call, but cost for each answer that turns out to be correct.
| Model | Correct | Cost per correct answer | Cost rank | Score rank |
|---|---|---|---|---|
| Gemma 4 31B | 117 | $0.0000154 | 1 | 3 |
| Qwen3-30B-A3B | 52 | $0.0000175 | 2 | 7 |
| Ling 3.0 Flash | 70 | $0.0000226 | 3 | 6 |
| GPT-5.6 Luna | 118 | $0.0000478 | 4 | 2 |
| DeepSeek V4 Flash | 114 | $0.0000537 | 5 | 4 |
| Mistral Small 3 | 24 | $0.0000612 | 6 | 9 |
| Hunyuan A13B | 22 | $0.0001207 | 7 | 10 |
| Claude Haiku 4.5 | 86 | $0.0002924 | 8 | 5 |
| Nemotron 3.5 Lightning | 27 | $0.0009019 | 9 | 8 |
| Gemini 3.5 Flash | 128 | $0.0028749 | 10 | 1 |
Gemini 3.5 Flash answers the most questions correctly and is at the same time the worst value in the panel: 187 times more expensive per correct answer than Gemma 4 31B, for 11 more lemmas. The full range across this panel is 187 times, and its order has almost nothing to do with the order of the scores.
We deliberately did not draw the cost as bars. The distance from cheapest to most expensive is 187 times, and on an ordinary axis 9 of the 10 bars would shrink until they cannot be read, so the chart would hide the very thing worth seeing. What can be drawn is the ranking, and there the two turn out to be almost unrelated:
3
7
6
2
4
9
10
5
8
1
Our third prediction was about this table, and it missed twice. Ling 3.0 Flash is indeed the cheapest model per call, but it answered only 70 lemmas correctly, and Gemma 4 31B kept the best cost per correct answer. Price per call decides the ranking only when the scores are comparable, and here they are not.
Nemotron 3.5 Lightning shows why an estimate taken from a price table cannot be trusted. Its published specification states that reasoning is off by default. It reasons, in English, for a question asked in Indonesian, and on a 600-token budget every answer came back empty. We raised its budget to 4,000 tokens so it could finish reasoning. Its real cost was 30 times the price-table estimate.
Our own answer keys were tested, and did not move
Adding 4 models meant adding 540 answers no speaker had ever read. Some of them are correct in a way not yet listed in the keys, and that is a real risk: if our keys were still loose, the numbers across the whole panel would shift.
A speaker reviewed all four, and the review added 6 keys across 4 lemmas, from 540 to 546. Not a single key was removed. We recomputed the scores of all 10 models against the new keys.
The 6 models whose numbers we had already published did not move by a single lemma. The added keys only rescued answers from models that had just joined the panel. That figure matters to us, because our earliest testing did the opposite: the instrument moved from 83.8 to 90.6 percent in 6 days with no model changing at all, and on the set of 43 rare words 1 model gained 14 lemmas on its own after review. 6 rounds of checking appear to be enough to make the keys stop moving.
Answer keys cannot separate the bottom 3 models; meaning similarity can
The bottom 3 models are practically tied when judged by the keys: Nemotron 27, Mistral 24, Hunyuan 22. Their differences fall below the between-run spread, so the three cannot be ordered at all.
A meaning-similarity column separates them clearly. That column measures how close a model's answer sits to the dictionary definition, reported as a percentile against a chance score we compute by crossing every answer against the definitions of all other lemmas.
Mistral Small 3 sits at 49.9 percent, which is exactly the zero point, the same as answering at random. Hunyuan sits at 63.1 percent and Nemotron at 64.8 percent, both clearly above it. Wrong answers have degrees too: some miss narrowly and some never touch the question at all, and key-based scoring alone throws that difference away.
Placed side by side, the two measures track each other in the top cluster and then part ways at the bottom.
94.7
90.7
92.0
92.9
84.9
81.8
72.6
64.8
49.9
63.1
The semantic column separates correct from incorrect verdicts with an AUC of 0.896 across all 135 lemmas.
What must not be concluded
This does not mean a high-scoring model knows Javanese. All 135 lemmas in this set were already published on our publicly indexed dictionary pages before the testing, so a high score cannot be told apart from having read our dictionary. To answer that we are preparing a control set of 192 lemmas we have never used or published.
This does not mean the order in the first table is a ranking. Differences smaller than 4 lemmas cannot be read, and that covers most of the distances inside the top cluster.
This does not mean a cheap model is always the economical one. Nemotron 3.5 Lightning is cheap per token and still ends up the second most expensive model per correct answer, because it burns reasoning tokens that produce no answer.
This does not mean our prediction failed because it was badly designed. The prediction was written so that it could fail, and the candidates were picked in order to make it fail. Had the gap stayed empty, we would be reporting that the distribution is split. What makes this result worth reading is not the direction of the finding, but that the direction was fixed before the data existed.
Every model answer, the answer keys, and the raw answers from all three repeat runs can be downloaded from the links below.
Data & provenance
Limitations
The 135 lemmas in this set are not rare words. All of them were already published on our publicly indexed dictionary pages before this testing, so a high score cannot be told apart from having read our dictionary. A control set built from lemmas we have never published is being prepared to answer that.
Between-run variance was measured on 3 models only, all of them from the top cluster. The other 7 models were called once per lemma, so their scores carry an uncertainty we have not measured.
The 4 models that newly joined the panel went through 1 round of speaker review, while the 6 earlier models went through 6 rounds. The answer keys have indeed stopped moving, but the two groups did not go through the same amount of checking.
No model family is represented at two different sizes in this panel, so the question of whether size or training data is what decides cannot be answered from the numbers here.
This article is an early record from the long benchmark-nusantara journey, which as this piece goes out is still heading toward an academic paper. Numbers on the public site's benchmark page and model page will keep changing as new models join and the methodology gets refined, so if they later differ from what is written here, that is not an error, it is a sign the project is still moving.
Revision history
- Published
- Last updated
Terms in this article
- lemma
- The base form of a word that heads a dictionary entry, and the form someone looks up. A single lemma can cover several senses at once, and each sense can carry its own example sentence.
- token
- The chunk of text a model counts in, roughly a syllable up to a short word. AI services are priced per million tokens, with input and output billed separately, so a rambling answer genuinely costs more than a concise one.
Written by the author for this article, not taken from a dictionary entry.