How many answers were correct
Each bar is one benchmark: how many items it answered correctly out of every item in that set. We wrote the questions and answer keys ourselves and had them checked by two native speakers, and every model answers exactly the same items.
How far the cost pays off
Sorted by the cheapest correct answer first. The left column shows how many items were answered correctly, so a top row with a long left bar means both cheap and accurate.
Tokens used
Input and output are kept apart because they are priced differently, often tenfold. A long output bar means the model talks a lot, and that is where nearly all the cost difference comes from.
Wait time per call
Median wait for a single call on each benchmark. Median rather than mean, so one stalled call does not move the number.
Test coverage
This model has been tested on 5 of the 6 benchmarks we publish. The rest have not been run, so this page cannot say anything about how it does there.
Earlier question versions
The questions or the key changed after these were measured. Kept as history, not as current results.show 2 rowsclose
Log10 calls to this model, spread across 5 days.OpenClose
This model is called through a provider id that may not pin a version, so the dates below state which weights were actually tested. Rows marked with a hyphen have a date we inferred from file history rather than recorded at call time, so we do not show a clock time for them.
- 22:32Pemahaman slang Indonesiamultiple choice37/42
- 19:57Arti kata Jawa sehari-hariword meaning51/135
- 06:02Tingkat tutur bahasa Jawaopen answer2/6
- 03:02Pemahaman slang Indonesiamultiple choice23/27
- 02:56Koreksi ejaan Indonesiaopen answer21/38
- 22:44Tingkat tutur bahasa Jawamultiple choice14/27
- 18:15Normalisasi teks Indonesiaopen answerPrompt stating the key's rule31/49
- 16:09Normalisasi teks Indonesiaopen answerDetailed prompt forbidding word swaps30/49
- 15:31Normalisasi teks Indonesiaopen answerBuilt-in prompt29/49
- 22:15Arti kata Jawa sehari-hariword meaning50/135