We Test How AI Understands Indonesian and Its Regional Languages
ibahasa tests AI's understanding of Indonesian ourselves, from the standard register to slang and regional dialects, then publishes the results together with the raw data and their limits. The measuring stick is a living dictionary we also built ourselves: a record of Indonesian as it is actually used.
We Measure How AI Understands Indonesian
We measure how well AI models actually understand Indonesian and its regional languages once they step outside the standard register: slang, dialects, cultural expressions, down to terms used only in a single region. The measuring stick is a corpus of Indonesian we built ourselves and had reviewed by people, part of the living dictionary we also maintain (more on that below).
We look for the answer ourselves and publish it one report at a time, together with the raw data and the things that cannot be concluded from it.
Alongside the written reports there is a live version of the same question. Battle AI puts two models side by side on an Indonesian-language prompt and lets you vote blind on which answered better, then reveals both identities and their benchmark scores.
Since When
This research began on 1 August 2026, and it grew out of a fortunate coincidence rather than a grand plan drawn up years in advance. Two technical projects we were building for entirely different reasons ended up converging on the same point.
Benchmark Orchestration Engine
To score answers from many AI models consistently, we first built our own orchestration engine: a system that sends hundreds of questions to different models, records their answers, and tallies the scores. It started as an internal tool, not something built to be published.
Spelling-Checker Editor
At the same time, we were developing an Indonesian spelling-checker editor, trying to measure how far a language model could be trusted to catch spelling mistakes in real text.
If we already had a tool for measuring this, why keep the results to ourselves?
That question is what pushed us to build this research page and publish the results in the open. The raw data lives on Zenodo and Hugging Face Datasets; the readable version is written up as research and insight articles on this page.
Research data and benchmark data are two separate, versioned Zenodo archives, each with its own DOI. Benchmark data has its own Zenodo record and GitHub repository. The benchmark side is also mirrored at Software Heritage, an independent archive dedicated to long-term preservation of open-source data and code.
Rules We Hold Ourselves To
- Test data is written from our own verified corpus, not lifted from benchmarks already circulating online. A model cannot have memorised what did not exist before we wrote it.
- A portion of those items is deliberately never published, so it stays usable for checking whether a model has seen the material before.
- Every published figure links to its data file in a public repository, pinned to the version it was cited at rather than to whatever the file looks like today.
- Findings that contradict what we expected get published too. Several of the reports there are simply records of us being wrong.
Who Runs This
Fullstack engineer. I build the dictionary and the tooling on this site myself, so when something is wrong, that is on me too. Every article links to its raw data so you can check it yourself.
This is not academic work and it has not been peer reviewed. The method is written out in the open precisely so that it can be argued with. If a number here looks wrong to you, that is a message worth sending.
Open Knowledge, Protected Community
Our dictionary is a corpus of Indonesian we curate ourselves: the lemmas and senses, together with community contributions, are licensed under Creative Commons Attribution-NonCommercial-ShareAlike 4.0 (CC BY-NC-SA 4.0). Some entries are sourced from KBBI instead; we do not hold the rights to that material and do not license it onward. ibahasa is an independent project, not affiliated with Badan Pengembangan dan Pembinaan Bahasa or with the official KBBI.
View CC BY-NC-SA 4.0 License