Loading prices...
All news
Four glowing steel slabs behind a crowded bench of small glass figures lit from within

Four B200s serve forty people, and that is the optimised number

Nvidia published the serving numbers for Nemotron 3 Ultra on four B200 accelerators, and the useful figure is not the speedup in the headline. It is how many people one box holds.

Measured at a fixed interactivity target of 20 milliseconds between tokens:

  • Baseline serving stack: 718 output tokens per second across the four GPUs.
  • NIM 2.0.12 optimised stack: 1,997 tokens per second on the same hardware.
  • At 50 tokens per second per user, that is 14 concurrent users against 40.

Forty concurrent users on four Blackwell accelerators, after every optimisation Nvidia's own engineers could apply. Before the optimisations, fourteen. This is the model that sits at the base of the Palantir sovereign stack announced this week.

The headline understates its own table, which is unusual enough to note. Divide 1,997 by 718 and the ratio is 2.78, while the title says 2.5x. Nvidia is rounding down, most likely because the Pareto curve gives different ratios at different latency targets and 50 tokens per user is only one point on it.

The workload is doing some of the work

The workload matters as much as the hardware. The benchmark uses a 64,000-token context window with 400 output tokens and 76% of the key-value cache reused between requests. Reuse that high is realistic for agentic traffic, where the same long prompt comes back step after step, and it is doing real work in the result. On traffic without that reuse the same stack would show a smaller gain.

Nvidia says so itself, in the sentence most vendors leave out.

The published curves are a starting point, not a promise that every application will see the same result.

NVIDIA, Technical blog, 10 September 2026

NVIDIA technical blog, 10 September 2026

550 billion parameters, 55 billion firing

The benchmark command in the post names the model in full: nemotron-3-ultra-550b-a55b. That is 550 billion parameters with 55 billion active per token, a mixture-of-experts design where tensor parallelism spreads the weights across all four GPUs and only a tenth of them fire on any given token.

Put this beside the other number Nvidia published on the same day. On the company's own weekly material allocation task, a 30-billion-parameter Nemotron beat this 550-billion flagship by 31 points. Nvidia gives no comparable serving figure for the small model, so we cannot say how many users it holds per box, and that comparison is the one an operations team would actually want.

The engineering here is real: autotuned kernels for the mixture-of-experts and Mamba layers, prefix caching with partial matching, scheduler and memory tuning, and speculative decoding. None of it is a trick, and all of it comes packaged so a team does not start from a blank runtime configuration.

What the post quietly establishes is the shape of the cost. Serving a frontier open model to forty people at a responsive speed takes four of the most expensive accelerators on the market and a stack tuned by the people who built the chip. Anyone planning to run one of these in-house should benchmark their own traffic before believing a curve, which is exactly what Nvidia recommends with AIPerf.

This piece is informational, not a recommendation to buy, sell, or hold any asset.

Published: 04:30 · 11.09.2026
Maks

Author

Maks

Trading man

I've been interested in the cryptocurrency market for a long time, am a trader, and write articles and news about my experience and crypto in simple terms.

Comments (0)

No comments yet — be the first!

The market talks all day. We write when it says something

Short, and it tells you why it came