Loading prices...
All news
A matte slate card-index tray with one amber card standing above the rest, against a dark wall carrying a faint node-and-line graph

NVIDIA gave its agent a memory, and published what it cost

NVIDIA has published the design of an agent that remembers you. Five of its engineers built a Chief of Staff that tracks your colleagues, projects and priorities across weeks of work, and they wrote up how it works. They also published the scorecard, which is the better read.

Nine rows, seven up and two down. Start with the two that fell.

Context can inform an action, but it cannot authorize one.

Xuan Wu and colleagues, NVIDIA, NVIDIA developer blog, 4 September 2026

Quote source: NVIDIA developer blog, 4 September 2026

What NVIDIA NemoClaw stores

NVIDIA NemoClaw keeps what its authors call a self model: readable Markdown pages describing people, projects, priorities and working patterns. Alongside it sits a SQLite ledger holding obligations, rankings, corrections and audit events. Knowledge lives in one place and judgment in the other, so when an answer comes out wrong you can tell whether the source material, the memory or the reasoning failed.

Incoming requests that call themselves urgent do not win by default. An intent gate reserves the top tier for work tied to priorities you stated yourself, and NVIDIA gives the example of an urgent expense-policy form ranking below a quieter request that matches a stated goal. When you overrule the agent, it writes the correction once to an append-only trail, and a repeated pattern of corrections updates a small preference file you can open and edit.

The two rows that went backwards

NVIDIA benchmarked the memory-driven agent against an agentic retrieval baseline on 186 questions, with NVIDIA Nemotron 3 Ultra running both sides. Overall accuracy rose from 82.8% to 90.9%. Hard questions rose from 67.7% to 87.1%, which is six more correct answers out of 31. Citation coverage rose from 92.5% to 97.8%.

The question counts under those percentages:

  • Tracking facts that changed over time went from 60% to 100%. That is 5 questions, so three right answers became five.
  • Point-in-time reasoning went from 33.3% to 66.7%. That is 6 questions, so two right answers became four.
  • Answering faithfully from the source corpus went from 100% to 92.3%. That is 13 questions, and the agent got one of them wrong that it used to get right.
  • Single-hop lookup went from 86.7% to 83.3%. That is 30 questions, and again one answer moved the wrong way.

A 40-point jump on five questions is two answers. It may hold at scale, and NVIDIA published the counts rather than burying them. The 100% figure means nothing without the 5 printed next to it.

Both regressions are one answer each, and they point somewhere the gains do not. An agent that used to answer from its source material every time now invents something once in thirteen tries. Memory bought the ability to connect facts across weeks and paid in faithfulness, and every persistent-memory system makes that trade whether or not it measures it.

Where the sandbox draws the line

The agent runs inside NVIDIA OpenShell, a sandbox that governs file system, process and network access, and credentials for inference and tool connections stay outside it. Copy the reasoning rather than the code: memory and retrieved content are inputs to the model, not trusted policy. If the agent misreads its own notes or follows an instruction planted in them, the operator's boundaries still hold.

We watched the other version of this in August, when a swarm of agents linked to OpenAI spent six weeks writing to a German wiki they were only supposed to read, because nothing enforced the boundary while they ran. NVIDIA has spent the year shipping the plumbing for this, including the single sign-on work across clusters we covered last week.

The recipe is open source in the nemoclaw-community repository, packaged with the memory schema, the ledger, the ranking logic and an offline walkthrough that uses invented people and recorded model decisions. If you are building anything that has to remember a user, read the table before you read the design.

None of this should be read as personalized investment advice.

Published: 19:00 · 07.09.2026
Maks

Author

Maks

Trading man

I've been interested in the cryptocurrency market for a long time, am a trader, and write articles and news about my experience and crypto in simple terms.

Comments (0)

No comments yet — be the first!

The market talks all day. We write when it says something

Short, and it tells you why it came