Benchmark Indonesian

Indonesian slang comprehension

Explaining what Indonesian slang words mean inside a sentence.Read more
Part of Benchmark Nusantara

Our own collection of benchmarks for Indonesian and its regional languages. Each set stands on its own, shares the same method and guards, and its data is published in one repository.

60 itemsLast tested
GitHub

12

Models tested

60/ 18

Questions / held back

25

Runs recorded

This set has 1 other set of results, from earlier versions of the questions or the answer key. We do not draw them on the same chart, because a score computed against different questions is not the same measurement.

Quick summary

Highest score

100.0%

42 of 42 answered correctly

GPT-5.6 LunaOpenAI
The gap to second place is 1 items, and this page makes no claim about the order between models whose run-to-run ranges overlap.

Cheapest

$0.000008

per correct answer

Solar Pro 4Upstage
Computed only among models scoring at least half of the top score on this board, because a model that gets a lot wrong always looks cheapest. This model scored 88.1%.

Lowest latency

501 ms

median wait

Mistral Small 3Mistral
It scored 76.2%, so fast here does not mean accurate.

Results

How many questions each model answered correctly, from the exact same set. The bars show the count of correct answers, not a general ranking of model ability.

multiple choice

The model reads the question with its answer options, then replies with a single letter. Scoring is automatic: correct when the letter matches the key.

temperature 0token cap variesKeterangan
GPT-5.6 Luna42/42 100.0%
Gemma 4 31B41/42 97.6%
Llama3 8B CPT Sahabat-AI v1 Instruct2 runsKeterangan40/42 95.2%
Claude Haiku 4.538/42 90.5%
Ling-3.0-flash37/42 88.1%
Solar Pro 437/42 88.1%
Mistral Small 332/42 76.2%
spread across runsmodels without a dark block were run once
12 models · 42 items scored

A model whose name does not end in a date is called through an id that does not pin the version. The provider may swap the model behind that id at any time without notice, so the figures apply to whichever version was active when the run happened.

How these figures are computed
phrasing mcq-py1Keteranganscoring mcq-py1Keteranganquestion version d3e8c01bc554Keterangananswer key version 4e94e1cb7f25Keterangan

Per-model detail

The same figures as the chart above, plus the cost per correct answer, the wait time, and the date of the last run. Any column that cannot be guessed from its name carries its own explanation.

multiple choice
Results per model on the multiple choice task: correct answers, cost per correct answer, wait time, and the date of the last run.
ModelCorrectKeteranganHeld backKeterangan$/correctKeteranganLatencyKeteranganLast runKeterangan
GPT-5.6 LunaOpenAI42/42100.0%18/18$0.000051$0.00311,103 msmedian
Gemma 4 31BGoogle41/4297.6%17/18$0.000017$0.0010614 msmedian
Llama3 8B CPT Sahabat-AI v1 Instructlocal2 runsKeterangan56-56/4295.2%16/18-430 msmedian
DeepSeek V4 Flash 0731DeepSeek39/4292.9%18/18$0.000039$0.00223,281 msmedian
Claude Haiku 4.5Anthropic38/4290.5%18/18$0.000274$0.01541,426 msmedian
Ling-3.0-flashAnt Group37/4288.1%17/18$0.000021$0.00111,410 msmedian
Qwen3 30B A3B Instruct 2507Alibaba37/4288.1%16/18$0.000021$0.0011701 msmedian
Solar Pro 4Upstage37/4288.1%15/18$0.000008$0.0004881 msmedian
Gemma-SEA-LION v4.5 E2B-ITlocal33/4278.6%12/18-843 msmedian
Mistral Small 3Mistral32/4276.2%13/18$0.000013$0.0006501 msmedian
Llama 3.1 8B Instructlocal32/4276.2%13/18-533 msmedian
Hunyuan A13B InstructTencent29/4269.0%14/18$0.000042$0.00181,488 msmedian

Models marked local run on our own hardware, so there is no bill to record and the wait time measures our machine, not a service anyone else can buy. The score is still comparable: the items, the key, and the temperature are identical to every other row here.

Every figure in this table can be recomputed from the raw data. Open the data

Cost and tokensCheapest per correct answer: Solar Pro 4 ($0.000008). Most expensive: Claude Haiku 4.5 ($0.000274).

All cost figures are in US dollars.

Cost per correct answer

The model with the highest score is not always the cheapest one. This divides the cost of one full test by its number of correct answers.

Cost per correct answer

Cost of one full run

The raw figure, before dividing by the number of correct answers. Every model answered the exact same questions, so this is directly comparable: it is what you pay to run this benchmark once on each model.

Cost of one full run

Tokens used

Input and output are kept apart because they are priced differently, often tenfold. A long output bar means the model talks a lot, and that is where almost all of the cost difference between models comes from. A model with no bar has no token record, which is not the same as zero.

inputoutput
Ling-3.0-flash13,334.0 · 10,279.0
DeepSeek V4 Flash 073110,888.0 · 4,981.0
GPT-5.6 Luna10,310.0 · 831.0
Claude Haiku 4.514,163.0 · 240.0
Solar Pro 413,796.0 · 125.0
Gemma 4 31B10,271.0 · 120.0
Llama3 8B CPT Sahabat-AI v1 Instruct2 runsKeterangan12,771.0 · 120.0
Gemma-SEA-LION v4.5 E2B-IT9,920.0 · 120.0
Mistral Small 311,084.0 · 120.0
Llama 3.1 8B Instruct12,771.0 · 120.0
Hunyuan A13B Instruct12,446.0 · 120.0
Qwen3 30B A3B Instruct 250712,684.0 · 103.0
Tokens used

Over time

Only models run more than once are drawn. One point is not a trend, and drawing it as a flat line implies a stability we have not measured.

2026-08-172026-08-172026-08-17 · 562026-08-17 · 56
Over time
DeretTanggalNilai
Llama3 8B CPT Sahabat-AI v1 Instruct2026-08-17T15:04:03Z56
Llama3 8B CPT Sahabat-AI v1 Instruct2026-08-17T15:10:43Z56
Over time

Earlier versions

The questions or the key changed after these numbers were measured, so they cannot be compared directly with the results above. We keep them rather than delete them.

multiple choice · mcq-py1 · 27 items · 3e5973f0

Benchmark