Which AI model handles Formal Register best?
This measures one thing only: how well a model answers in the Formal Register category, judged by readers. A model near the bottom here may well be the stronger one for anything else. Every number comes from the arena itself, not from a separate test.
Live standings· not the final ranking
| # | Model | Played | W / L / T | Pts |
|---|---|---|---|---|
| 1 | DeepSeek V3.2DeepSeek | 12of 603 pairs | 6 / 4 / 2 | 20 |
| 2 | Gemma 4 31BGoogle | 13of 603 pairs | 6 / 5 / 2 | 20 |
| 3 | GPT-4o Mini (Legacy)OpenAI | 12of 571 pairs | 4 / 3 / 5 | 17 |
| 4 | MiniMax M2.5MiniMax | 7of 247 pairs | 5 / 1 / 1 | 16 |
| 5 | Llama 3.3 70BMeta | 9of 603 pairs | 5 / 3 / 1 | 16 |
| 6 | Mistral Small 3.2 24BMistral | 11of 603 pairs | 4 / 3 / 4 | 16 |
| 7 | Gemini 3.5 Flash LiteGoogle | 15of 603 pairs | 4 / 7 / 4 | 16 |
| 8 | GPT-5.6 LunaOpenAI | 5of 217 pairs | 0 / 4 / 1 | 1 |
| 9 | Reka Flash 3Lainnya | 1of 13 pairs | 0 / 1 / 0 | 0 |
| 10 | IBM: Granite 4.1 8BLainnya | 3of 603 pairs | 0 / 3 / 0 | 0 |
| · | Claude 4.6 SonnetAnthropic | 0of 603 pairs | 0 / 0 / 0 | 0 |
| · | MiMo-V2.5Lainnya | 0of 85 pairs | 0 / 0 / 0 | 0 |
| · | Mistral Small 3Mistral | 0of 603 pairs | 0 / 0 / 0 | 0 |
| · | UI-TARS 7BLainnya | 0of 603 pairs | 0 / 0 / 0 | 0 |
- Standing
- : Position by points, exactly as in a league table. It moves with every vote. A dot instead of a number means the model is in the arena but has never been judged, so it has no standing yet.
- Played
- : How many times this model has actually been judged, out of the pairs it appears in. Read the points next to it together with this: three points from one match says far less than three points from thirty.
- Wins, losses, ties
- : Raw counted votes. Calibration pairs, whose answer we already know, are excluded because their votes never feed the ranking.
- Points
- : Three for a win, one for a draw, as in a league table. Equal points are broken by wins minus losses, then by matches played; both are already visible in the W / L / T column. Points do not account for how strong the opponent was: the final rating does, and it replaces this column once enough matches have been played.
- Which side an answer appears on is decided by the server and randomised, so a habit of picking one side never becomes a result.
What this does not say
A position here is not a verdict on a model. It is one narrow measurement: how well it answers in the Formal Register category, judged by people reading two unlabelled answers.
Models are built for different things, and are trained on very different amounts of Indonesian. One that sits low here may be the far stronger choice for English, for code, for reasoning, or for longer work, none of which this arena touches. A low position means it was less convincing on these questions, in this variety, to the people who voted. Nothing beyond that.
Where these numbers come from
Every row above was produced by people reading two unlabelled answers and picking one. There is no other source.
- counted votes
- 44
- people
- 3
- models in the arena
- 14