Benchmark Indonesian

Indonesian text normalisation

Turning text-message style writing into standard Indonesian.Read more
Part of Benchmark Nusantara

Our own collection of benchmarks for Indonesian and its regional languages. Each set stands on its own, shares the same method and guards, and its data is published in one repository.

69 itemsLast tested
GitHub

11

Models tested

69/ 20

Questions / held back

33

Runs recorded

Quick summary

Taken from the open answer, built-in prompt board, the board with the most models in this set. The other boards have their own figures below.

Highest score

59.2–87.8%

The best score depends on how the question is phrased

Lowest
59.2% · Built-in prompt · Qwen3 30B A3B Instruct 2507
Highest
87.8% · Detailed prompt forbidding word swaps · Gemma 4 31B

Full comparison of all three phrasings is below.

Cheapest

$0.000014

per correct answer

Mistral Small 3Mistral
Computed only among models scoring at least half of the top score on this board, because a model that gets a lot wrong always looks cheapest. This model scored 53.1%.

Lowest latency

962 ms

median wait

Mistral Small 3Mistral
It scored 53.1%, so fast here does not mean accurate.

Results

How many questions each model answered correctly, from the exact same set. The bars show the count of correct answers, not a general ranking of model ability.

This set was run with 3 different instructions

The items and the answer key are identical across all 3 runs, 69 open answer items that do not move by a single word. Only the instruction sent to the model differs. Gemma 4 31B moves 44.9 points purely from changing the instruction. The winner changes too, from Qwen3 30B A3B Instruct 2507 to Gemma 4 31B. So this set has no single leaderboard, and the numbers must not be averaged across instructions.

Built-in promptDetailed prompt forbidding word swapsPrompt stating the key's rule
Gemma 4 31B42.9% · 87.8% · 87.8%
GPT-5.6 Luna44.9% · 77.6% · 83.7%
DeepSeek V4 Flash 073149.0% · 71.4% · 75.5%
Gemma-SEA-LION v4.5 E2B-IT46.9% · 73.5% · 71.4%
Ling-3.0-flash55.1% · 65.3% · 59.2%
Qwen3 30B A3B Instruct 250759.2% · 61.2% · 63.3%
Solar Pro 438.8% · 57.1% · 61.2%
Mistral Small 353.1% · 55.1% · 46.9%
Llama 3.1 8B Instruct32.7% · 40.8% · 24.5%
Hunyuan A13B Instruct24.5% · 38.8% · 22.4%
This set was run with 3 different instructions

What each instruction says

Built-in prompt

The prompt reads “Rewrite the sentence in standard Indonesian”, taken as it stands from inside the item. It is loose: it never says what must be left alone, so a model is free to read it as licence to swap words for synonyms.

open-py1
Detailed prompt forbidding word swaps

The prompt forbids swapping words for synonyms and forbids changing the level of formality. That sounds reasonable, but our own answer key swaps words in 40 places, so this prompt contradicts the key it is scored against.

open-py2
Prompt stating the key's rule

The prompt asks for every non-standard word to be made standard, including abbreviations written without vowels, and for already-standard words to be left alone. It was written after the key's rule was counted word by word.

open-py3
Show 3 separate boards, one per instruction
open answerBuilt-in prompt

The model writes its own answer, with no options offered. Some are scored automatically against a reference, some are read one by one by a native speaker using a weighted rubric.

temperature 0token cap variesKeterangan
11 models · 49 items scored

A model whose name does not end in a date is called through an id that does not pin the version. The provider may swap the model behind that id at any time without notice, so the figures apply to whichever version was active when the run happened.

How these figures are computed
phrasing open-py1Keteranganscoring open-py1Keteranganquestion version 5ca9dfe11dc2Keterangananswer key version eda1d1cddc75Keterangan
open answerDetailed prompt forbidding word swaps

The model writes its own answer, with no options offered. Some are scored automatically against a reference, some are read one by one by a native speaker using a weighted rubric.

temperature 0token cap variesKeterangan
11 models · 49 items scored

A model whose name does not end in a date is called through an id that does not pin the version. The provider may swap the model behind that id at any time without notice, so the figures apply to whichever version was active when the run happened.

How these figures are computed
phrasing open-py2Keteranganscoring open-py1Keteranganquestion version 5ca9dfe11dc2Keterangananswer key version eda1d1cddc75Keterangan
open answerPrompt stating the key's rule

The model writes its own answer, with no options offered. Some are scored automatically against a reference, some are read one by one by a native speaker using a weighted rubric.

temperature 0token cap variesKeterangan
11 models · 49 items scored

A model whose name does not end in a date is called through an id that does not pin the version. The provider may swap the model behind that id at any time without notice, so the figures apply to whichever version was active when the run happened.

How these figures are computed
phrasing open-py3Keteranganscoring open-py1Keteranganquestion version 5ca9dfe11dc2Keterangananswer key version eda1d1cddc75Keterangan

Per-model detail

The same figures as the chart above, plus the cost per correct answer, the wait time, and the date of the last run. Any column that cannot be guessed from its name carries its own explanation.

Show 3 separate tables, one per instruction
open answerBuilt-in prompt
Results per model on the open answer task: correct answers, cost per correct answer, wait time, and the date of the last run.
ModelCorrectKeteranganHeld backKeterangan$/correctKeteranganLatencyKeteranganLast runKeterangan
Qwen3 30B A3B Instruct 2507Alibaba29/4959.2%11/20$0.000024$0.00102,043 msmedian
Ling-3.0-flashAnt Group27/4955.1%12/20$0.000050$0.00201,766 msmedian
Mistral Small 3Mistral26/4953.1%13/20$0.000014$0.0005962 msmedian
DeepSeek V4 Flash 0731DeepSeek24/4949.0%11/20$0.000116$0.00403,644 msmedian
Gemma-SEA-LION v4.5 E2B-ITlocal23/4946.9%12/20-890 msmedian
GPT-5.6 LunaOpenAI22/4944.9%9/20$0.000069$0.00211,034 msmedian
Gemma 4 31BGoogle21/4942.9%11/20$0.000051$0.00161,330 msmedian
Solar Pro 4Upstage19/4938.8%9/20$0.000017$0.00052,684 msmedian
Llama3 8B CPT Sahabat-AI v1 Instructlocal17/4934.7%5/20-390 msmedian
Llama 3.1 8B Instructlocal16/4932.7%10/20-528 msmedian
Hunyuan A13B InstructTencent12/4924.5%5/20$0.000114$0.00191,485 msmedian

Models marked local run on our own hardware, so there is no bill to record and the wait time measures our machine, not a service anyone else can buy. The score is still comparable: the items, the key, and the temperature are identical to every other row here.

open answerDetailed prompt forbidding word swaps
Results per model on the open answer task: correct answers, cost per correct answer, wait time, and the date of the last run.
ModelCorrectKeteranganHeld backKeterangan$/correctKeteranganLatencyKeteranganLast runKeterangan
Gemma 4 31BGoogle43/4987.8%16/20$0.000026$0.00161,373 msmedian
GPT-5.6 LunaOpenAI38/4977.6%14/20$0.000081$0.00422,199 msmedian
Gemma-SEA-LION v4.5 E2B-ITlocal36/4973.5%15/20-863 msmedian
DeepSeek V4 Flash 0731DeepSeek35/4971.4%15/20$0.000095$0.00474,526 msmedian
Ling-3.0-flashAnt Group32/4965.3%12/20$0.000040$0.00181,767 msmedian
Qwen3 30B A3B Instruct 2507Alibaba30/4961.2%13/20$0.000027$0.00121,656 msmedian
Llama3 8B CPT Sahabat-AI v1 Instructlocal29/4959.2%11/20-415 msmedian
Solar Pro 4Upstage28/4957.1%13/20$0.000013$0.00055,339 msmedian
Mistral Small 3Mistral27/4955.1%11/20$0.000017$0.0006650 msmedian
Llama 3.1 8B Instructlocal20/4940.8%10/20-505 msmedian
Hunyuan A13B InstructTencent19/4938.8%10/20$0.000078$0.00231,493 msmedian

Models marked local run on our own hardware, so there is no bill to record and the wait time measures our machine, not a service anyone else can buy. The score is still comparable: the items, the key, and the temperature are identical to every other row here.

open answerPrompt stating the key's rule
Results per model on the open answer task: correct answers, cost per correct answer, wait time, and the date of the last run.
ModelCorrectKeteranganHeld backKeterangan$/correctKeteranganLatencyKeteranganLast runKeterangan
Gemma 4 31BGoogle43/4987.8%17/20$0.000027$0.0016952 msmedian
GPT-5.6 LunaOpenAI41/4983.7%14/20$0.000074$0.00412,020 msmedian
DeepSeek V4 Flash 0731DeepSeek37/4975.5%14/20$0.000096$0.00494,094 msmedian
Gemma-SEA-LION v4.5 E2B-ITlocal35/4971.4%15/20-930 msmedian
Qwen3 30B A3B Instruct 2507Alibaba31/4963.3%15/20$0.000028$0.00131,636 msmedian
Solar Pro 4Upstage30/4961.2%13/20$0.000012$0.00052,275 msmedian
Ling-3.0-flashAnt Group29/4959.2%12/20$0.000083$0.00341,936 msmedian
Mistral Small 3Mistral23/4946.9%15/20$0.000016$0.0006671 msmedian
Llama3 8B CPT Sahabat-AI v1 Instructlocal20/4940.8%9/20-418 msmedian
Llama 3.1 8B Instructlocal12/4924.5%11/20-534 msmedian
Hunyuan A13B InstructTencent11/4922.4%11/20$0.000099$0.00221,497 msmedian

Models marked local run on our own hardware, so there is no bill to record and the wait time measures our machine, not a service anyone else can buy. The score is still comparable: the items, the key, and the temperature are identical to every other row here.

Every figure in this table can be recomputed from the raw data. Open the data

Cost and tokensopen answerBuilt-in promptCheapest per correct answer: Mistral Small 3 ($0.000014). Most expensive: DeepSeek V4 Flash 0731 ($0.000116).

All cost figures are in US dollars.

Cost per correct answer · Built-in prompt

The model with the highest score is not always the cheapest one. This divides the cost of one full test by its number of correct answers.

Cost per correct answer · Built-in prompt

Cost of one full run · Built-in prompt

The raw figure, before dividing by the number of correct answers. Every model answered the exact same questions, so this is directly comparable: it is what you pay to run this benchmark once on each model.

Cost of one full run · Built-in prompt

Tokens used · Built-in prompt

Input and output are kept apart because they are priced differently, often tenfold. A long output bar means the model talks a lot, and that is where almost all of the cost difference between models comes from. A model with no bar has no token record, which is not the same as zero.

inputoutput
Ling-3.0-flash10,971.0 · 25,604.0
DeepSeek V4 Flash 07318,996.0 · 12,280.0
GPT-5.6 Luna8,826.0 · 2,099.0
Llama3 8B CPT Sahabat-AI v1 Instruct10,752.0 · 950.0
Llama 3.1 8B Instruct10,752.0 · 945.0
Qwen3 30B A3B Instruct 250710,617.0 · 928.0
Mistral Small 39,348.0 · 873.0
Hunyuan A13B Instruct10,341.0 · 869.0
Solar Pro 412,501.0 · 853.0
Gemma-SEA-LION v4.5 E2B-IT8,106.0 · 687.0
Gemma 4 31B9,098.0 · 685.0
Tokens used · Built-in prompt
Cost and tokensopen answerDetailed prompt forbidding word swapsCheapest per correct answer: Solar Pro 4 ($0.000013). Most expensive: DeepSeek V4 Flash 0731 ($0.000095).

All cost figures are in US dollars.

Cost per correct answer · Detailed prompt forbidding word swaps

The model with the highest score is not always the cheapest one. This divides the cost of one full test by its number of correct answers.

Cost per correct answer · Detailed prompt forbidding word swaps

Cost of one full run · Detailed prompt forbidding word swaps

The raw figure, before dividing by the number of correct answers. Every model answered the exact same questions, so this is directly comparable: it is what you pay to run this benchmark once on each model.

Cost of one full run · Detailed prompt forbidding word swaps

Tokens used · Detailed prompt forbidding word swaps

Input and output are kept apart because they are priced differently, often tenfold. A long output bar means the model talks a lot, and that is where almost all of the cost difference between models comes from. A model with no bar has no token record, which is not the same as zero.

inputoutput
Ling-3.0-flash13,455.0 · 22,398.0
DeepSeek V4 Flash 073110,998.0 · 14,585.0
GPT-5.6 Luna10,689.0 · 5,256.0
Llama 3.1 8B Instruct13,098.0 · 955.0
Llama3 8B CPT Sahabat-AI v1 Instruct13,098.0 · 939.0
Qwen3 30B A3B Instruct 250712,963.0 · 930.0
Mistral Small 311,349.0 · 875.0
Hunyuan A13B Instruct12,687.0 · 871.0
Solar Pro 414,502.0 · 849.0
Gemma 4 31B10,432.0 · 688.0
Gemma-SEA-LION v4.5 E2B-IT9,900.0 · 684.0
Tokens used · Detailed prompt forbidding word swaps
Cost and tokensopen answerPrompt stating the key's ruleCheapest per correct answer: Solar Pro 4 ($0.000012). Most expensive: Hunyuan A13B Instruct ($0.000099).

All cost figures are in US dollars.

Cost per correct answer · Prompt stating the key's rule

The model with the highest score is not always the cheapest one. This divides the cost of one full test by its number of correct answers.

Cost per correct answer · Prompt stating the key's rule

Cost of one full run · Prompt stating the key's rule

The raw figure, before dividing by the number of correct answers. Every model answered the exact same questions, so this is directly comparable: it is what you pay to run this benchmark once on each model.

Cost of one full run · Prompt stating the key's rule

Tokens used · Prompt stating the key's rule

Input and output are kept apart because they are priced differently, often tenfold. A long output bar means the model talks a lot, and that is where almost all of the cost difference between models comes from. A model with no bar has no token record, which is not the same as zero.

inputoutput
Ling-3.0-flash12,351.0 · 42,114.0
DeepSeek V4 Flash 073110,240.0 · 14,955.0
GPT-5.6 Luna9,999.0 · 5,122.0
Llama 3.1 8B Instruct12,477.0 · 971.0
Llama3 8B CPT Sahabat-AI v1 Instruct12,477.0 · 942.0
Qwen3 30B A3B Instruct 250712,342.0 · 916.0
Gemma 4 31B9,761.0 · 901.0
Mistral Small 310,521.0 · 879.0
Hunyuan A13B Instruct12,066.0 · 871.0
Solar Pro 413,812.0 · 860.0
Gemma-SEA-LION v4.5 E2B-IT9,210.0 · 687.0
Tokens used · Prompt stating the key's rule

Benchmark