Empirical· 7 min read

How good are NVIDIA's model scores compared to other models

In mid-August we ran the paid version of Nemotron 3.5 Lightning on the official board, on Javanese word-meaning questions, and it landed among the lowest scores in the panel. NVIDIA has also been giving away some of its larger models for free, so we asked a fairer question for them: how do these models score on everyday Indonesian, and does a much larger size pay off on questions like that.

Measured from export-lajur-uji-2026-08-26-id-B-ringkas.csv and 1 more

34/42
Nemotron Super
26/42
Nemotron 3.5
41/42
Gemma 4 31B
$0
Test cost

NVIDIA already appeared on our board, and it did not go well

nvidia nemotron AI
Nvidia nemotron AI

On August 12 and 16, 2026, we ran the paid version of Nemotron 3.5 Lightning on the official board, on two Javanese word-meaning sets. It scored 34 of 199 lemmas on one set and 3 of 43 on the other, among the lowest of any model we have tested. Javanese word meanings are genuinely hard for almost every model, so that figure is not a verdict on NVIDIA in general.

Over the past few months NVIDIA has given away some of its larger models for free through OpenRouter, including ones sized at 120 and 550 billion parameters. So we asked a fairer question for them: how do they score on everyday Indonesian, and does the much larger size pay off on questions like that.

How did we measure it?

We used our Indonesian slang set: multiple-choice questions on the meaning of everyday slang words and expressions, from youth abbreviations to terms that have shifted meaning. We used 42 public-slice questions here, the portion safe to share because it is not used for contamination testing. The remaining 18 questions we hold back and never send to models that retain user prompts, because those reserve questions are exactly how we check for contamination. The answer key has never been published.

All three models ran in the trial lane, a separate track that uses the same questions and scoring rubric as the official board but keeps the results quarantined. Scores in this article do not count toward the ibahasa.com/benchmark leaderboard. The cost was zero rupiah, since every model tested was being given away free at the time.

Benchmark results

ModelSizeCorrect out of 42Status
GPT-5.6 Lunaundisclosed42paid, official board
Gemma 4 31B31 billion41paid, official board
DeepSeek V4 Flashundisclosed39paid, official board
Nemotron 3 Super120 billion34free
SEA-LION 4.6B4.6 billion33local, on our own machine
Laguna S 2.1undisclosed32free
Hunyuan A13B13 billion active29paid, official board
Nemotron 3.5 Lightning30 billion (3 billion active)26free
Gemma 4 31B (paid)
41
DeepSeek V4 Flash (paid)
39
Nemotron 3 Super 120B (free)
34
SEA-LION 4.6B (local)
33
Nemotron 3.5 Lightning (free)
26
Correct out of 42 Indonesian slang questions

Nemotron 3 Super answered 34 questions correctly, and that is a respectable score: it beats the paid Hunyuan and matches SEA-LION, which we run ourselves in-house. Nemotron 3.5 Lightning answered 26 correctly, the lowest of any model that has ever entered this set.

Size does not decide the score

The most interesting part of the table above is not the ranking, it is the size column. Nemotron 3 Super carries 120 billion parameters and answered 34 questions correctly. Gemma 4 31B, roughly a quarter the size of Nemotron 3 Super, answered 41. That 7-question gap is fairly decisive statistically (z around 2.5), and the direction runs against the common assumption that a bigger model must understand more.

An even more surprising comparison sits one row below. SEA-LION is 4.6 billion parameters, quantized down to 4 bits to fit on an ordinary computer, and runs on our own machine with no internet connection. It answered 33 questions correctly, practically tied with Nemotron 3 Super, a model 26 times larger on paper and running in an NVIDIA data center. For questions like these, what decides is not how big the model is, but how much everyday Indonesian it has actually read.

Total Active per answer
Nemotron 3 Super
120
12
Gemma 4 31B
31
31
Nemotron 3.5 Lightning
30
3
SEA-LION 4.6B
4.6
4.6
Model size, total versus actually active when answering (billion parameters)

The chart above explains most of the puzzle. Nemotron 3 Super's 120 billion is the size of the whole model, while only 12 billion actually does the work when answering one question, smaller than Gemma's full 31 billion firing at once. Nemotron 3.5 Lightning is even more extreme: 3 billion active out of 30 billion total. The short gray bar is what we are actually pitting against Gemma, not the long red one.

Before concluding large models are wasted, it is worth looking at what they were built for. Nemotron 3 Super is genuinely large, but only 12 billion of its parameters activate when answering, and it uses a hybrid architecture that is not a pure transformer. It is designed for agentic reasoning and automating technical support tickets. Nemotron 3.5 Lightning is optimized for writing code and enterprise use, with 3 billion active parameters out of 30 billion total. Nemotron 3 Ultra, with 55 billion active parameters, is aimed at advanced reasoning and coordinating long-running agents. All three are built to reason and complete tasks, not to guess the meaning of slang. The low score on Indonesian slang is therefore not a design failure, it is a model used outside its main expertise.

The model we could not test

NVIDIA also gives away Nemotron 3 Ultra for free, a 550-billion-parameter model that was supposed to be the focus of this article. We never managed to get a single answer out of it. Every API request was rejected with a message that the shared free-tier quota was full, even after we added an eight-second delay between questions and raised the retry delay up to forty seconds.

A model that cannot be called produces no data, and a model with no data does not go in our table. This is itself worth recording as a hidden cost of free models: not just accuracy, but whether the model can be reached when needed. That free quota is shared with the entire world, and your turn depends on how busy it happens to be at the time.

A mistake of our own in these numbers

This article uses scores on the 42 public-slice questions, and that figure differs from what we wrote in the article two days ago. The cause was our own mistake: our export file stores the public-slice score and the held-out-slice score in separate columns, and we had been reading the first column as the overall score. As a result, one comparison in the earlier article was wrong and has now been corrected, and we have fixed the column description in the public dataset file so no one else makes the same reading mistake.

We mention it here because the Gemma score of 41 out of 42 in the table above is the same figure we previously misread. It is correct now, and the way to check it is open: the summary file is linked below this article.

Limits of the conclusion

Each model ran once, on one type of question, in one language. So the conclusion stays limited to understanding Indonesian slang on the day of measurement. The gap between the two NVIDIA models themselves, 34 versus 26, is not yet decisive enough to rule out chance (z around 1.9), while the gap from both of them to Gemma and GPT-5.6 is wide enough to discuss. We have also not yet tested these NVIDIA models on Javanese speech levels, the hardest part of our instrument, and the paid version's result on Javanese word meanings suggests Javanese speech levels are a considerably harder test.

Data & provenance

Limitations

One run per model on 42 multiple-choice questions from the public slice of our slang set. The questions and answer key have never been published. The numbers here come from the trial lane, kept separate from the official board, so they do not move anyone's ranking on ibahasa.com/benchmark. Nemotron 3 Ultra, 550 billion parameters, could not be tested at all because its free quota stayed full, and a model with no data does not go in the table.

Revision history

Published
Last updated
never revised since publication

Terms in this article

agentic
The trait of an AI model that acts on its own rather than only answering once: deciding the next step, calling tools, or pursuing a multi-step goal without being re-prompted at each step. An agentic failure means the model acted wrongly while pursuing that goal, not just picked the wrong answer.

Written by the author for this article, not taken from a dictionary entry.