Benchmark Javanese

Everyday Javanese word meanings

Explaining everyday Javanese word meanings verified by native speakers, part of it held back.Read more
Part of Benchmark Nusantara

Our own collection of benchmarks for Indonesian and its regional languages. Each set stands on its own, shares the same method and guards, and its data is published in one repository.

199 itemsLast tested
GitHub

14

Models tested

199/ 64

Questions / held back

34

Runs recorded

This set has 1 other set of results, from earlier versions of the questions or the answer key. We do not draw them on the same chart, because a score computed against different questions is not the same measurement.

Quick summary

Highest score

94.8%

128 of 135 answered correctly

Gemini 3.5 FlashGoogle
The gap to second place is 12 items, and this page makes no claim about the order between models whose run-to-run ranges overlap.

Cheapest

$0.000012

per correct answer

Gemma 4 31BGoogle
Computed only among models scoring at least half of the top score on this board, because a model that gets a lot wrong always looks cheapest. This model scored 85.9%.

Lowest latency

816 ms

median wait

Qwen3 30B A3B Instruct 2507Alibaba
It scored 37.8%, so fast here does not mean accurate.

Results

How many questions each model answered correctly, from the exact same set. The bars show the count of correct answers, not a general ranking of model ability.

word meaning

The model is asked to explain what a single word means in free text. An answer counts as correct if it contains every word of one of our approved keys.

temperature 0token cap variesKeterangan
14 models · 135 items scored

A model whose name does not end in a date is called through an id that does not pin the version. The provider may swap the model behind that id at any time without notice, so the figures apply to whichever version was active when the run happened.

How these figures are computed
phrasing arti-kata-v1Keteranganscoring kunci-alias-v1Keteranganquestion version a7ba9fa24e64Keterangananswer key version 873b6f168502Keterangan

Per-model detail

The same figures as the chart above, plus the cost per correct answer, the wait time, and the date of the last run. Any column that cannot be guessed from its name carries its own explanation.

word meaning
Results per model on the word meaning task: correct answers, cost per correct answer, wait time, and the date of the last run.
ModelCorrectKeteranganHeld backKeterangan$/correctKeteranganLatencyKeteranganLast runKeterangan
Gemini 3.5 FlashGoogle128/13594.8%58/6490.6%$0.002985$0.55522,711 msmedian
Gemma 4 31BGoogle116/13585.9%51/6479.7%$0.000012$0.00211,063 msmedian
GPT-5.6 LunaOpenAI112/13583.0%53/6482.8%$0.000051$0.00841,438 msmedian
DeepSeek V4 Flash 0731DeepSeek107/13579.3%49/6476.6%$0.000054$0.00853,123 msmedian
Claude Haiku 4.5Anthropic86/13563.7%41/6464.1%$0.000292$0.03711,635 msmedian
Ling-3.0-flashAnt Group69/13551.1%27/6442.2%$0.000036$0.00351,390 msmedian
Llama3 8B CPT Sahabat-AI v1 Instructlocal61/13545.2%27/6442.2%-528 msmedian
Gemma-SEA-LION v4.5 E2B-ITlocal60/13544.4%27/6442.2%-734 msmedian
Solar Pro 4Upstage52/13538.5%24/6437.5%$0.000013$0.00101,142 msmedian
Qwen3 30B A3B Instruct 2507Alibaba51/13537.8%23/6435.9%$0.000020$0.0015816 msmedian
Llama 3.1 8B Instructlocal26/13519.3%11/6417.2%-532 msmedian
Nemotron 3.5 LightningNVIDIA25/13518.5%9/6414.1%$0.001081$0.03683,088 msmedian
Mistral Small 3Mistral25/13518.5%6/649.4%$0.000070$0.0022843 msmedian
Hunyuan A13B InstructTencent21/13515.6%12/6418.8%$0.000118$0.00391,747 msmedian

Models marked local run on our own hardware, so there is no bill to record and the wait time measures our machine, not a service anyone else can buy. The score is still comparable: the items, the key, and the temperature are identical to every other row here.

Every figure in this table can be recomputed from the raw data. Open the data

Cost and tokensCheapest per correct answer: Gemma 4 31B ($0.000012). Most expensive: Gemini 3.5 Flash ($0.002985).

All cost figures are in US dollars.

Cost per correct answer

The model with the highest score is not always the cheapest one. This divides the cost of one full test by its number of correct answers.

Cost of one full run

The raw figure, before dividing by the number of correct answers. Every model answered the exact same questions, so this is directly comparable: it is what you pay to run this benchmark once on each model.

Tokens used

Input and output are kept apart because they are priced differently, often tenfold. A long output bar means the model talks a lot, and that is where almost all of the cost difference between models comes from. A model with no bar has no token record, which is not the same as zero.

inputoutput
Nemotron 3.5 Lightning8,445.0 · 155,249.0
Gemini 3.5 Flash4,437.0 · 60,950.0
DeepSeek V4 Flash 07316,313.0 · 36,231.0
Ling-3.0-flash10,649.0 · 27,904.0
GPT-5.6 Luna6,029.0 · 12,966.0
Claude Haiku 4.59,041.0 · 5,613.0
Hunyuan A13B Instruct7,669.0 · 4,928.0
Solar Pro 415,811.0 · 4,563.0
Mistral Small 337,698.0 · 3,509.0
Gemma 4 31B7,027.0 · 3,431.0
Llama3 8B CPT Sahabat-AI v1 Instruct8,062.0 · 3,281.0
Qwen3 30B A3B Instruct 25077,669.0 · 3,100.0
Llama 3.1 8B Instruct8,062.0 · 2,910.0
Gemma-SEA-LION v4.5 E2B-IT6,228.0 · 2,594.0
Tokens used

Earlier versions

The questions or the key changed after these numbers were measured, so they cannot be compared directly with the results above. We keep them rather than delete them.

Gemini 3.5 Flash127/135 94.1%
GPT-5.6 Luna121/135 89.6%
GPT-5.6 Luna119/135 88.1%
Gemma 4 31B116/135 85.9%
GPT-5.6 Luna116/135 85.9%
Gemma 4 31B115/135 85.2%
Gemma 4 31B113/135 83.7%
Claude Haiku 4.582/135 60.7%
Ling-3.0-flash69/135 51.1%
Solar Pro 452/135 38.5%
Mistral Small 321/135 15.6%
word meaning · arti-kata-v1 · 135 items · 4e682041

Benchmark