Benchmark
All AI Models in the Ibahasa Benchmark
Every model below has answered Indonesian or regional-language questions we wrote ourselves. Open one to see how it did across all our benchmarks.Read more
The only regional language measured so far is Javanese. Others such as Sundanese, Madurese, Balinese, Batak, Buginese and Minangkabau are still being collected: the sets do not exist yet, so there are no figures. We publish numbers only for what has actually been tested.
- 4 benchmarksClaude Haiku 4.5KeteranganThis model is called through a provider id that does not pin a version. Pinned ids carry a date at the end, like model-name-0731. Without that date, the provider can swap the model behind the same id at any time without telling anyone. The figures here describe the version that served during the test.Anthropicanthropic/claude-haiku-4.5context 200Ktext · image · file → text
- 6 benchmarksDeepSeek V4 Flash 0731DeepSeekdeepseek/deepseek-v4-flash-0731context 1311Ktext → text
- 6 benchmarksGPT-5.6 LunaKeteranganThis model is called through a provider id that does not pin a version. Pinned ids carry a date at the end, like model-name-0731. Without that date, the provider can swap the model behind the same id at any time without telling anyone. The figures here describe the version that served during the test.OpenAIopenai/gpt-5.6-lunacontext 1050Kfile · image · text → text
- 3 benchmarksGemini 3.5 FlashKeteranganThis model is called through a provider id that does not pin a version. Pinned ids carry a date at the end, like model-name-0731. Without that date, the provider can swap the model behind the same id at any time without telling anyone. The figures here describe the version that served during the test.Googlegoogle/gemini-3.5-flashcontext 1049Ktext · image · video · file · audio → text
- 6 benchmarksGemma 4 31BKeteranganThis model is called through a provider id that does not pin a version. Pinned ids carry a date at the end, like model-name-0731. Without that date, the provider can swap the model behind the same id at any time without telling anyone. The figures here describe the version that served during the test.Googlegoogle/gemma-4-31b-itcontext 262Kimage · text · video → text
- localKeteranganA model we run on our own PC or internal server, not through a paid provider. Its score is still comparable with the others because the questions, the key, and the temperature are identical. Cost and latency are not comparable: there is no bill to record, and the wait time measures our machine rather than a service anyone can buy.6 benchmarksGemma-SEA-LION v4.5 E2B-ITcontext 131K4.6BQ4_K_Maudio · thinking · tools · vision
- 5 benchmarksHunyuan A13B InstructKeteranganThis model is called through a provider id that does not pin a version. Pinned ids carry a date at the end, like model-name-0731. Without that date, the provider can swap the model behind the same id at any time without telling anyone. The figures here describe the version that served during the test.Tencenttencent/hunyuan-a13b-instructcontext 131Ktext → text
- 5 benchmarksLing-3.0-flashKeteranganThis model is called through a provider id that does not pin a version. Pinned ids carry a date at the end, like model-name-0731. Without that date, the provider can swap the model behind the same id at any time without telling anyone. The figures here describe the version that served during the test.Ant Groupinclusionai/ling-3.0-flashcontext 262Ktext → text
- localKeteranganA model we run on our own PC or internal server, not through a paid provider. Its score is still comparable with the others because the questions, the key, and the temperature are identical. Cost and latency are not comparable: there is no bill to record, and the wait time measures our machine rather than a service anyone can buy.6 benchmarksLlama 3.1 8B Instructcontext 131K8.0BQ4_K_Mtools
- localKeteranganA model we run on our own PC or internal server, not through a paid provider. Its score is still comparable with the others because the questions, the key, and the temperature are identical. Cost and latency are not comparable: there is no bill to record, and the wait time measures our machine rather than a service anyone can buy.6 benchmarksLlama3 8B CPT Sahabat-AI v1 Instructcontext 8K8.0BQ4_K_M
- 6 benchmarksMistral Small 3Mistralmistralai/mistral-small-24b-instruct-2501context 33Ktext → text
- 2 benchmarksNemotron 3.5 LightningKeteranganThis model is called through a provider id that does not pin a version. Pinned ids carry a date at the end, like model-name-0731. Without that date, the provider can swap the model behind the same id at any time without telling anyone. The figures here describe the version that served during the test.NVIDIAnvidia/nemotron-3.5-lightningcontext 1000Ktext → text
- 5 benchmarksQwen3 30B A3B Instruct 2507Alibabaqwen/qwen3-30b-a3b-instruct-2507context 262Ktext → text
- 6 benchmarksSolar Pro 4KeteranganThis model is called through a provider id that does not pin a version. Pinned ids carry a date at the end, like model-name-0731. Without that date, the provider can swap the model behind the same id at any time without telling anyone. The figures here describe the version that served during the test.Upstageupstage/solar-pro4context 524Ktext → text