OpenAIModel

GPT-6 Astra

How GPT-6 Astra performed across every Indonesian and regional-language benchmark we publish.Read more
Last runon Pemahaman slang Indonesiaversion not pinnedKeterangan

Benchmarks tested

3

of 6 benchmarks we publish

The rest have not been run, so this page can say nothing about how it does on 3 other benchmarks.

Runs completed

4

complete runs on record

One run means the model answered every item in a set, start to finish. A set run more than once contributes more than one.

Total cost

$0.4790

for every run above

The real amount billed by the provider, summed, not an estimate from per-token pricing. Not yet divided by correct answers.

Total tokens

23,459

input and output combined

Counted only from runs that recorded tokens. Runs on the older path did not store them, and the figure is not guessed.

How many answers were correct

Each bar is one benchmark: how many items it answered correctly out of every item in that set. We wrote the questions and answer keys ourselves and had them checked by two native speakers, and every model answers exactly the same items.

4 benchmarks
Pemahaman slang Indonesia42/42 100.0%
Tingkat tutur bahasa Jawaopen answer6/6 100.0%
Tingkat tutur bahasa Jawamultiple choice26/27 96.3%
Arti kata Jawa langka33/43 76.7%
Share of items answered correctly across 4 benchmarks

How far the cost pays off

Sorted by the cheapest correct answer first. The left column shows how many items were answered correctly, so a top row with a long left bar means both cheap and accurate.

4 with cost on record
Correct answersCost per correct answer
Tingkat tutur bahasa Jawamultiple choice96.3%$0.001961
Pemahaman slang Indonesia100.0%$0.002465
Tingkat tutur bahasa Jawaopen answer100.0%$0.005116
Arti kata Jawa langka76.7%$0.006669
Correct answers against cost per correct answer

Tokens used

Input and output are kept apart because they are priced differently, often tenfold. A long output bar means the model talks a lot, and that is where nearly all the cost difference comes from.

17,349 in, 6,110 out
inputoutput
Pemahaman slang Indonesia10,310 · 847
Tingkat tutur bahasa Jawaopen answer968 · 625
Tingkat tutur bahasa Jawamultiple choice4,754 · 500
Arti kata Jawa langka1,317 · 4,138
Tokens for one full run, per benchmark

Wait time per call

Median wait for a single call on each benchmark. Median rather than mean, so one stalled call does not move the number.

4 with wait time on record
Tingkat tutur bahasa Jawamultiple choice1.93 s
Pemahaman slang Indonesia2.76 s
Arti kata Jawa langka4.31 s
Tingkat tutur bahasa Jawaopen answer5.01 s
Median wait for one call, per benchmark

Test coverage

This model has been tested on 3 of the 6 benchmarks we publish. The rest have not been run, so this page cannot say anything about how it does there.

3 of 6 benchmarks we publish
Log4 calls to this model, spread across 1 days.

This model is called through a provider id that may not pin a version, so the dates below state which weights were actually tested. Rows marked with a hyphen have a date we inferred from file history rather than recorded at call time, so we do not show a clock time for them.

4 runs
  1. 19:51Pemahaman slang Indonesiamultiple choice42/42
  2. 10:06Tingkat tutur bahasa Jawaopen answer6/6
  3. 10:05Tingkat tutur bahasa Jawamultiple choice26/27
  4. 10:01Arti kata Jawa langkaword meaning33/43

See all models