Which AI model handles Slang best?
This measures one thing only: how well a model answers in the Slang category, judged by readers. A model near the bottom here may well be the stronger one for anything else. Every number comes from the arena itself, not from a separate test.
Live standings· not the final ranking
| # | Model | Played | W / L / T | Pts |
|---|---|---|---|---|
| 1 | Claude 4.6 Sonnet (OR)Anthropic | 11of 777 pairs | 6 / 1 / 4 | 22 |
| 2 | DeepSeek V4 FlashDeepSeek | 10of 507 pairs | 6 / 1 / 3 | 21 |
| 3 | DeepSeek V3.2DeepSeek | 12of 777 pairs | 6 / 5 / 1 | 19 |
| 4 | Llama 3.3 70BMeta | 9of 777 pairs | 4 / 4 / 1 | 13 |
| 5 | UI-TARS 7BLainnya | 8of 782 pairs | 2 / 4 / 2 | 8 |
| 6 | Gemma 4 31BGoogle | 5of 777 pairs | 1 / 0 / 4 | 7 |
| 7 | Voxtral Small 24B 2507Mistral | 7of 782 pairs | 2 / 4 / 1 | 7 |
| 8 | Gemini 3.5 Flash LiteGoogle | 4of 777 pairs | 1 / 1 / 2 | 5 |
| 9 | Mistral Small 3Mistral | 7of 782 pairs | 0 / 2 / 5 | 5 |
| 10 | Gemma 3 12BGoogle | 3of 782 pairs | 1 / 2 / 0 | 3 |
| 11 | Qwen3 30B InstructQwen | 7of 782 pairs | 0 / 4 / 3 | 3 |
| 12 | GPT OSS 20BOpenAI | 1of 221 pairs | 0 / 1 / 0 | 0 |
| · | Tencent: Hy3Lainnya | 0of 65 pairs | 0 / 0 / 0 | 0 |
| · | Xiaomi: MiMo-V2.5-ProLainnya | 0of 782 pairs | 0 / 0 / 0 | 0 |
- Standing
- : Position by points, exactly as in a league table. It moves with every vote. A dot instead of a number means the model is in the arena but has never been judged, so it has no standing yet.
- Played
- : How many times this model has actually been judged, out of the pairs it appears in. Read the points next to it together with this: three points from one match says far less than three points from thirty.
- Wins, losses, ties
- : Raw counted votes. Calibration pairs, whose answer we already know, are excluded because their votes never feed the ranking.
- Points
- : Three for a win, one for a draw, as in a league table. Equal points are broken by wins minus losses, then by matches played; both are already visible in the W / L / T column. Points do not account for how strong the opponent was: the final rating does, and it replaces this column once enough matches have been played.
- Which side an answer appears on is decided by the server and randomised, so a habit of picking one side never becomes a result.
What this does not say
A position here is not a verdict on a model. It is one narrow measurement: how well it answers in the Slang category, judged by people reading two unlabelled answers.
Models are built for different things, and are trained on very different amounts of Indonesian. One that sits low here may be the far stronger choice for English, for code, for reasoning, or for longer work, none of which this arena touches. A low position means it was less convincing on these questions, in this variety, to the people who voted. Nothing beyond that.
Where these numbers come from
Every row above was produced by people reading two unlabelled answers and picking one. There is no other source.
- counted votes
- 42
- people
- 4
- models in the arena
- 14