Frequently Asked Questions

Everything you need to know about ibahasa: how we work, what we cover, and how you can get involved.

About ibahasa

What is ibahasa?
ibahasa began as a living dictionary of Indonesian: a descriptive record of how the language is actually used day to day, and that is what makes it useful as a measuring stick. That corpus is now the foundation of our research: measuring how well AI models actually understand Indonesian once they step outside the standard register, then publishing the findings as research reports, together with the raw data and their limits.
Why did ibahasa start as a dictionary, and why is the focus research now?
The dictionary is done: a corpus of Indonesian we curate ourselves, reviewed by people before publication. Once that corpus matured, we found a second use for it, as a measuring stick to test AI. Our active work today is on that side: the research and benchmarks we publish in the open. If you find a dictionary entry that is wrong, tell us. That is still our most reliable way of catching mistakes.

Research & AI Testing

Where do the numbers in your research articles come from?
From tests we run ourselves, not from figures quoted off someone else. Every report links to the raw data file it used, in a public repository, pinned to the version the number was cited at rather than to whatever the file looks like today. If a figure cannot be traced back to a file, it does not get published.
What is Benchmark Nusantara?

It is the name of our own collection of benchmarks for Indonesian and its regional languages. Each set measures one ability, shares the same method and guards, and all of the data is published in a repository of the same name. That name is also what you cite. The position behind it fits in one sentence: the main output of a benchmark is not a leaderboard but a calibrated statement about how far its figures can be trusted. The list of sets is at ibahasa.com/en/benchmark.

Can I download the research data?

Yes, in two repositories, both CC BY 4.0. Benchmark data is at github.com/ibahasa/benchmark-nusantara, everything else at github.com/ibahasa/research-data. They are separate because benchmark data changes every time a model is run, while the rest must stay frozen so that an article citing it remains checkable. Each dataset ships with a MANIFEST.json listing every file, its size, and its sha256 value, so you can verify a download. Third-party sources keep their own terms and are attributed there.

Why is part of the test data withheld?
Because if we published all of it, all of it could end up in the training data of the next model, and after that we would have no way of telling whether a model actually understands the language or has simply read the answers. The withheld portion is never used in quizzes, games, or any public feature either.

Still Have Questions?

Check out our About page for more on our research, or visit our community for direct support.