Model ranking

Which AI model handles Formal Register best?

This measures one thing only: how well a model answers in the Formal Register category, judged by readers. A model near the bottom here may well be the stronger one for anything else. Every number comes from the arena itself, not from a separate test.

Prescriptive rule

Live standings· not the final ranking

10%
15/150 matches
#ModelPlayedW / L / TPts
1DeepSeek V3.2DeepSeek12of 603 pairs6 / 4 / 220
2Gemma 4 31BGoogle13of 603 pairs6 / 5 / 220
3GPT-4o Mini (Legacy)OpenAI12of 571 pairs4 / 3 / 517
4MiniMax M2.5MiniMax7of 247 pairs5 / 1 / 116
5Llama 3.3 70BMeta9of 603 pairs5 / 3 / 116
6Mistral Small 3.2 24BMistral11of 603 pairs4 / 3 / 416
7Gemini 3.5 Flash LiteGoogle15of 603 pairs4 / 7 / 416
8GPT-5.6 LunaOpenAI5of 217 pairs0 / 4 / 11
9Reka Flash 3Lainnya1of 13 pairs0 / 1 / 00
10IBM: Granite 4.1 8BLainnya3of 603 pairs0 / 3 / 00
·Claude 4.6 SonnetAnthropic0of 603 pairs0 / 0 / 00
·MiMo-V2.5Lainnya0of 85 pairs0 / 0 / 00
·Mistral Small 3Mistral0of 603 pairs0 / 0 / 00
·UI-TARS 7BLainnya0of 603 pairs0 / 0 / 00
Standing
: Position by points, exactly as in a league table. It moves with every vote. A dot instead of a number means the model is in the arena but has never been judged, so it has no standing yet.
Played
: How many times this model has actually been judged, out of the pairs it appears in. Read the points next to it together with this: three points from one match says far less than three points from thirty.
Wins, losses, ties
: Raw counted votes. Calibration pairs, whose answer we already know, are excluded because their votes never feed the ranking.
Points
: Three for a win, one for a draw, as in a league table. Equal points are broken by wins minus losses, then by matches played; both are already visible in the W / L / T column. Points do not account for how strong the opponent was: the final rating does, and it replaces this column once enough matches have been played.
Which side an answer appears on is decided by the server and randomised, so a habit of picking one side never becomes a result.

What this does not say

A position here is not a verdict on a model. It is one narrow measurement: how well it answers in the Formal Register category, judged by people reading two unlabelled answers.

Models are built for different things, and are trained on very different amounts of Indonesian. One that sits low here may be the far stronger choice for English, for code, for reasoning, or for longer work, none of which this arena touches. A low position means it was less convincing on these questions, in this variety, to the people who voted. Nothing beyond that.

Where these numbers come from

Every row above was produced by people reading two unlabelled answers and picking one. There is no other source.

counted votes
44
people
3
models in the arena
14
Add your vote