BenchmarkEmpirical
Wording 1
Wording 3
Gemma 4 31B
32
60
60
Qwen3-30B-A3B
40
46
46
Ling 3.0 Flash
39
41
41
Mistral Small 3
39
38
38
Change the Prompt, and the AI Ranking Flips
We put 4 AI models through the same 69 sentences using three ways of asking. Gemma 4 31B, built by Google, came last under the first way and first under the third, with the questions and the answer key never changing.
September 15, 2026 · 4 menit bacaRead the finding →