Benchmark Javanese

Javanese speech levels

Choosing the Javanese register that fits the person being addressed.Read more
Part of Benchmark Nusantara

Our own collection of benchmarks for Indonesian and its regional languages. Each set stands on its own, shares the same method and guards, and its data is published in one repository.

38 itemsLast tested
GitHub

13

Models tested

38

Questions

20

Runs recorded

Quick summary

Taken from the multiple choice board, the board with the most models in this set. The other boards have their own figures below.

Highest score

96.3%

26 of 27 answered correctly

Gemini 3.5 FlashGoogle
Tied with Gemma 4 31B on exactly the same number of correct answers. The name here is not a winner, and we make no claim about the order between the two.

Cheapest

$0.000010

per correct answer

Solar Pro 4Upstage
Computed only among models scoring at least half of the top score on this board, because a model that gets a lot wrong always looks cheapest. This model scored 59.3%.

Lowest latency

547 ms

median wait

Mistral Small 3Mistral
It scored 40.7%, so fast here does not mean accurate.

Results

How many questions each model answered correctly, from the exact same set. The bars show the count of correct answers, not a general ranking of model ability.

multiple choice

The model reads the question with its answer options, then replies with a single letter. Scoring is automatic: correct when the letter matches the key.

temperature 0token cap variesKeterangan
13 models · 27 items scored

A model whose name does not end in a date is called through an id that does not pin the version. The provider may swap the model behind that id at any time without notice, so the figures apply to whichever version was active when the run happened.

How these figures are computed
phrasing mcq-py1Keteranganscoring mcq-py1Keteranganquestion version c0c83850858eKeterangananswer key version 18e8ec378904Keterangan
open answer

The model writes its own answer, with no options offered. Some are scored automatically against a reference, some are read one by one by a native speaker using a weighted rubric.

temperature 0600 token capKeterangan

A model whose name does not end in a date is called through an id that does not pin the version. The provider may swap the model behind that id at any time without notice, so the figures apply to whichever version was active when the run happened.

How these figures are computed
phrasing open-py1Keteranganscoring open-py1Keteranganquestion version 9f2d79048c32Keterangananswer key version 4610a4b37355Keterangan

Per-model detail

The same figures as the chart above, plus the cost per correct answer, the wait time, and the date of the last run. Any column that cannot be guessed from its name carries its own explanation.

multiple choice
Results per model on the multiple choice task: correct answers, cost per correct answer, wait time, and the date of the last run.
ModelCorrectKeteranganHeld backKeterangan$/correctKeteranganLatencyKeteranganLast runKeterangan
Gemini 3.5 FlashGoogle26/2796.3%11/11$0.002148$0.07952,217 msmedian
Gemma 4 31BGoogle26/2796.3%10/11$0.000015$0.0006803 msmedian
DeepSeek V4 Flash 0731DeepSeek25/2792.6%11/11$0.000054$0.00204,324 msmedian
GPT-5.6 LunaOpenAI23/2785.2%11/11$0.000023$0.0008988 msmedian
Claude Haiku 4.5Anthropic20/2774.1%8/11$0.000256$0.00722,228 msmedian
Ling-3.0-flashAnt Group18/2766.7%8/11$0.000058$0.00151,547 msmedian
Solar Pro 4Upstage16/2759.3%6/11$0.000010$0.00026,680 msmedian
Gemma-SEA-LION v4.5 E2B-ITlocal16/2759.3%6/11-837 msmedian
Qwen3 30B A3B Instruct 2507Alibaba14/2751.9%5/11$0.000023$0.0004612 msmedian
Llama3 8B CPT Sahabat-AI v1 Instructlocal13/2748.1%5/11-311 msmedian
Hunyuan A13B InstructTencent12/2744.4%9/11$0.000039$0.00081,396 msmedian
Mistral Small 3Mistral11/2740.7%7/11$0.000014$0.0003547 msmedian
Llama 3.1 8B Instructlocal9/2733.3%6/11-432 msmedian

Models marked local run on our own hardware, so there is no bill to record and the wait time measures our machine, not a service anyone else can buy. The score is still comparable: the items, the key, and the temperature are identical to every other row here.

open answer
Results per model on the open answer task: correct answers, cost per correct answer, wait time, and the date of the last run.
ModelCorrectKeteranganHeld backKeterangan$/correctKeteranganLatencyKeteranganLast runKeterangan
Gemma 4 31BGoogle6/6100.0%2/2$0.000019$0.00011,318 msmedian
Gemini 3.5 FlashGoogle6/6100.0%2/2$0.003273$0.02622,753 msmedian
Gemma-SEA-LION v4.5 E2B-ITlocal5/683.3%1/2-920 msmedian
Qwen3 30B A3B Instruct 2507Alibaba2/633.3%0/2$0.000050$0.0001477 msmedian
Llama3 8B CPT Sahabat-AI v1 Instructlocal1/616.7%0/2-426 msmedian
Llama 3.1 8B Instructlocal1/616.7%0/2-571 msmedian
Mistral Small 3Mistral0/60.0%0/2-$0.00011,098 msmedian

Models marked local run on our own hardware, so there is no bill to record and the wait time measures our machine, not a service anyone else can buy. The score is still comparable: the items, the key, and the temperature are identical to every other row here.

Every figure in this table can be recomputed from the raw data. Open the data

Cost and tokensmultiple choiceCheapest per correct answer: Solar Pro 4 ($0.000010). Most expensive: Gemini 3.5 Flash ($0.002148).

All cost figures are in US dollars.

Cost per correct answer · multiple choice

The model with the highest score is not always the cheapest one. This divides the cost of one full test by its number of correct answers.

Cost per correct answer · multiple choice

Cost of one full run · multiple choice

The raw figure, before dividing by the number of correct answers. Every model answered the exact same questions, so this is directly comparable: it is what you pay to run this benchmark once on each model.

Cost of one full run · multiple choice

Tokens used · multiple choice

Input and output are kept apart because they are priced differently, often tenfold. A long output bar means the model talks a lot, and that is where almost all of the cost difference between models comes from. A model with no bar has no token record, which is not the same as zero.

inputoutput
Ling-3.0-flash6,245.0 · 13,504.0
Gemini 3.5 Flash4,122.0 · 8,142.0
DeepSeek V4 Flash 07314,960.0 · 6,495.0
GPT-5.6 Luna4,754.0 · 528.0
Claude Haiku 4.56,396.0 · 152.0
Solar Pro 46,781.0 · 78.0
Mistral Small 35,016.0 · 78.0
Llama 3.1 8B Instruct5,740.0 · 77.0
Gemma 4 31B4,908.0 · 76.0
Gemma-SEA-LION v4.5 E2B-IT4,654.0 · 76.0
Llama3 8B CPT Sahabat-AI v1 Instruct5,740.0 · 76.0
Hunyuan A13B Instruct5,517.0 · 76.0
Qwen3 30B A3B Instruct 25075,669.0 · 67.0
Tokens used · multiple choice
Cost and tokensopen answerCheapest per correct answer: Gemma 4 31B ($0.000019). Most expensive: Gemini 3.5 Flash ($0.003273).

All cost figures are in US dollars.

Cost per correct answer · open answer

The model with the highest score is not always the cheapest one. This divides the cost of one full test by its number of correct answers.

Cost per correct answer · open answer

Cost of one full run · open answer

The raw figure, before dividing by the number of correct answers. Every model answered the exact same questions, so this is directly comparable: it is what you pay to run this benchmark once on each model.

Cost of one full run · open answer

Tokens used · open answer

Input and output are kept apart because they are priced differently, often tenfold. A long output bar means the model talks a lot, and that is where almost all of the cost difference between models comes from. A model with no bar has no token record, which is not the same as zero.

inputoutput
Gemini 3.5 Flash802.0 · 2,776.0
Llama3 8B CPT Sahabat-AI v1 Instruct1,231.0 · 132.0
Qwen3 30B A3B Instruct 25071,219.0 · 122.0
Mistral Small 31,031.0 · 108.0
Llama 3.1 8B Instruct1,231.0 · 103.0
Gemma 4 31B962.0 · 97.0
Gemma-SEA-LION v4.5 E2B-IT914.0 · 96.0
Tokens used · open answer

Benchmark