Transparency

Benchmark run log

Every time we run an AI model against a benchmark set, the row is recorded here exactly as it happened. Not a rollup like the /benchmark leaderboards, but one row per call, so any summary figure on this site can be traced back to the actual run that produced it.Read more
134 test runs14 models tested6 benchmark setsMost recent run
GitHub

1 runs
  1. 03:19Ling-3.0-flashNormalisasi teks Indonesiaopen answerDetailed prompt forbidding word swaps32/49

1 runs
  1. 19:42Ling-3.0-flashTingkat tutur bahasa Jawamultiple choice18/27

1 runs
  1. 16:57Gemma-SEA-LION v4.5 E2B-ITTingkat tutur bahasa Jawaopen answer5/6

The number at the far right of each row is the score: correct answers out of items scored. It can be smaller than the set's full item count when some items are held back or still awaiting human judgment.

This page is the raw list, not the analysis. For results per set with their limits, see the benchmark leaderboard. For how scoring works, see our methodology.