Loading prices...
All news
A large featureless clay boulder beside a small grooved cube glowing and lifted off the floor

A 30B model beat Nvidia's flagship by 31 points on Nvidia's own problem

Nvidia published the engineering post behind yesterday's supply chain announcement, and it contains a benchmark that argues against the way most people shop for models. On Nvidia's own allocation task, a 30-billion-parameter model beats the company's flagship by 31.2 percentage points.

Accuracy on the development benchmark for allocation decisions:

  • Nemotron 3.5 Lightning, 30 billion parameters, post-trained on the decisions: 86.7%.
  • Nemotron 3 Ultra, the much larger flagship, is 31.2 points behind: 55.5%.
  • The same 30B model before post-training is 69.2 points behind: 17.5%.

Nvidia gives the gaps rather than the absolute scores for the two losing models, so we subtracted: 86.7 minus 31.2 is 55.5, and 86.7 minus 69.2 is 17.5. The flagship, running on a task with a documented right answer, gets about half of them. The same small model that scores 86.7% after post-training scored 17.5% before it, which puts almost all of the capability in the data rather than in the model itself.

What the model is actually deciding

The task is worth describing, because it is not a chatbot problem. Every week Nvidia decides how much of each scarce component goes to which manufacturing site, running through the current quarter and the next. Contract manufacturers cannot start assembly until every part has arrived, so whatever comes early waits for whatever comes late, and Nvidia measures the wait as Time of Ownership.

The arithmetic under that is heavy. One compute tray of the eighteen in a Grace Blackwell NVL72 rack needs two Grace CPUs, four Blackwell GPUs and thirty-two HBM3e memory stacks. Multiply by eighteen and a single rack needs 36 CPUs, 72 GPUs and 576 memory stacks, which is where the 72 in the product name comes from. Nvidia says the supply chain it built for Vera Rubin is twice the size of that one.

The quantitative half runs on cuOpt, posed as a mixed-integer linear program that minimises Time of Ownership. It also returns which constraint is binding, so a planner reads that capacity in Taiwan held the number down this week and memory supply did not. Then Nvidia and Palantir back-tested the solver against what the humans actually decided, and the humans kept winning.

Planners were working from information the solver could not see: emails exchanged with partners that week, severe weather in the forecast for a key region or an ongoing geopolitical event, the transcript from the last supplier debrief call, and years of accumulated expertise.

NVIDIA, Developer blog, 10 September 2026

NVIDIA developer blog, 10 September 2026

Where the capability really sits

So they built the workflow around the planners instead of over them. The system records the decision, the reasoning behind it, the expected result and the outcome that followed, and that record is what the small LLM was post-trained on. Every accepted, edited or overridden recommendation goes back into the same store for the next training run.

This is the specific answer to the question we raised about the sovereign stack yesterday. Palantir claimed capabilities that exceed the frontier while the open model underneath had already lost its leaderboard position. The engineering post makes the better version of that argument: on a narrow task with proprietary data, the leaderboard is not the thing being measured, and a 30B model with the right examples beats a frontier model without them.

Read the 17.5% figure as the honest catch. A company that wants this result needs years of recorded decisions with reasoning and outcomes attached, which most companies do not have and cannot buy. What Nvidia is selling here is the machinery for capturing that record. The record itself has to be earned week by week.

This article is for informational purposes only and does not constitute investment advice.

Published: 21:45 · 10.09.2026
Maks

Author

Maks

Trading man

I've been interested in the cryptocurrency market for a long time, am a trader, and write articles and news about my experience and crypto in simple terms.

Comments (0)

No comments yet — be the first!

The market talks all day. We write when it says something

Short, and it tells you why it came