GPT-6 Astra and Luna tie on Indonesian slang items
Two models answered every public slang item correctly. The only thing that told them apart was the half of the set we never published.
Key numbers
Where these come fromMeasured from export-id-B-ringkas-model.csv and 1 more
A perfect score that did not win
We ran GPT-6 Astra through our Indonesian slang benchmark, 60 multiple-choice items asking what a slang term means in the sentence it appears in. 42 of those items are published. The other 18 are withheld, never released in any dataset, so that a model trained after the fact cannot have read them.
Astra answered all 42 public items correctly. Nothing in the panel has done better. It then missed 1 of the 18 withheld items and finished the whole set at 59 of 60, 1 item behind GPT-5.6 Luna, the cheaper OpenAI model it was built to replace, which answered all 60.
| Model | Public | Withheld | Whole set | Cost |
|---|---|---|---|---|
| GPT-5.6 Luna | 42/42 | 18/18 | 60/60 | $0.0031 |
| GPT-6 Astra | 42/42 | 17/18 | 59/60 | $0.1455 |
| Gemma 4 31B | 41/42 | 17/18 | 58/60 | $0.0010 |
| DeepSeek V4 Flash | 39/42 | 18/18 | 57/60 | $0.0022 |
| Claude Haiku 4.5 | 38/42 | 18/18 | 56/60 | $0.0154 |
| GLM 5.3 | 40/42 | 16/18 | 56/60 | $0.0174 |
1 item apart on 60 is not a real gap. A two-proportion test puts it at z = 1.00 and p ≈ 0.32, which is another way of saying we cannot tell these two models apart on this set at all. Anyone reporting this as Luna beating Astra would be reading noise as a result.
The two models finished in the same order on our rare Javanese word test, where Luna answered 35 of 43 and Astra 33. That gap was not significant either. Two runs leaning the same way is not evidence of a real difference between these models, and we are not presenting it as one.
The public half stopped measuring
The more durable finding is what the top of that table looks like. Two models scored 42 of 42 on the public items and a third missed a single item. At that ceiling the published half of the set has no discriminating power left. Three of the strongest models are separated by a single public item between them, which is well inside the margin where one rerun could reorder them.
Everything that still separates the leaders comes from the 18 withheld items, and even there the signal is thin. Luna, DeepSeek V4 Flash, and Claude Haiku 4.5 all answered 18 of 18, yet they finished 60, 57, and 56 on the whole set. The withheld half is doing nearly all of the ranking work while being the smaller half.
A benchmark whose published half everyone passes is a benchmark that has stopped asking a question.
This is the argument for keeping a withheld slice that we have made before, arriving from the opposite direction. The usual worry is contamination, where published items leak into training data and stop measuring anything. Here the public items have not leaked, they have simply been solved. Both roads end in the same place, and only the withheld items are still asking anything.
Expensive, and not fast either
Cost separates these models far more sharply than accuracy does. Astra's run cost $0.1455 at OpenAI's list price of $10 per million input tokens and $50 per million output tokens. Luna's identical run cost $0.0031, 47 times less, for 1 more correct answer. Gemma 4 31B finished 1 item behind Astra at $0.0010, roughly 145 times cheaper.
Speed follows the same pattern we measured on Javanese vocabulary in that earlier test. Astra's median response time on this set was 2,759 ms, against 1,103 ms for Luna, 1,214 ms for GLM 5.3, and 614 ms for Gemma 4 31B. Only DeepSeek V4 Flash and GLM 5.3 Flash, both of which emit long reasoning traces, were slower. Astra is terse in what it writes and still takes longer to write it.
The limits of a single item
Astra did not fail this test. It matched the best public score in our panel and finished 1 withheld item off a perfect run, on a set where the strongest models are now stacked on top of each other. Read strictly, this run says Astra is at the top of the Indonesian slang panel and indistinguishable from the models around it.
It also says something about the instrument rather than the models. When the published half of a benchmark can no longer tell the leaders apart, the honest response is to say so in the same breath as the scores, and to keep building items the leaders have not already solved.
Data & provenance
Limitations
60 multiple-choice items, one run per model, no repeats. A one-item difference at this sample size is inside ordinary run-to-run noise, so the ranking between the top three models should be read as a tie, not an order.
The public half of this set is close to saturated. Two models answered all 42 public items correctly and a third missed a single item, which means the public figures alone can no longer rank the strongest models. Any future comparison at this level depends on the withheld half, and the withheld half is 18 items.
Item text and answer keys for the withheld half are never published, including the single item Astra missed. That protects the measurement, and it also means readers cannot check that specific item themselves.
This set covers Indonesian slang comprehension in multiple-choice form. It says nothing about generating slang, about regional varieties outside the corpus the items were drawn from, or about the reasoning and agentic tasks OpenAI markets Astra for.
Revision history
- Published
- Last updated
- never revised since publication