Nvidia has presented a simulation system that learns from real videos and robot motion data, then predicts what happens next when an action is performed. The idea is to manufacture experience rather than collect text. As the reserve of written data runs dry, this approach shifts the question: the next scarce resource would no longer be content, but consequence.
We described this morning the rush on printed books from before 2022, a symptom of the exhaustion of easy textual data. Here is the other answer to that problem, radically different: rather than seeking more text, produce experience.
The problem it solves
A language model learns by reading. That is effective for everything expressed in words, and it is precisely what explains its limits that we have documented: its clumsiness at basic manipulations, and more generally the fact that it has never touched anything.
For a robot, this limit is disqualifying. Grasping a fragile object, opening a door that resists, recovering from a botched move: none of that is learned by reading descriptions. You have to have tried, failed, and started again. Yet physical experience is extraordinarily costly to collect: each trial takes real time, wears out equipment, and unfolds in only one world at a time.
The idea: simulating consequences
The approach presented involves learning from two complementary sources: real video, including highly technical sequences such as surgical gestures, and robot motion data, meaning the exact position of each joint at every moment.
From there, the system learns to predict what happens next. If an arm exerts such force on such an object, what occurs? The object slips, tips over, deforms, resists. The system becomes a kind of consequence engine.
The benefit is considerable: a robot can then train in this simulation, thousands of times in parallel, without breaking hardware or waiting for real time. It no longer learns from descriptions of the world, but in an approximate world where actions have outcomes.
Until now, a training data point was a document: a text, an image, a sound. Here, the data point becomes an environment with consequences. This is not a matter of vocabulary. It means the next strategic resource may not be content, but control over reliable experiences: licensed human archives, verifiable synthetic practice, or proprietary physical environments. In this scenario, simulator makers and robot fleet operators become parts of the model value chain, on par with publishers.
The well-known limit: the gap from reality
We should raise the main objection right away, which roboticists have long known as the simulation-to-reality gap.
A robot trained in a simulator excels in that simulator. Faced with the world, it discovers what the simulation did not model: a reflection that tricks its camera, dust that changes friction, an object whose material does not behave as expected. The more realistic the simulation, the smaller the gap, but it never disappears.
And there is a subtler risk, which should remind us of something. If models learn mainly in worlds generated by other models, we find the logic of model collapse that we described for text: a system that feeds on its own approximations ends up drifting from reality without noticing. Simulation is a formidable accelerator, provided it remains anchored by real physical data.
What this signals
This direction confirms a fundamental shift in the sector, often summed up as physical AI. After models that write and code, the goal is to get systems working that act in the material world: logistics, industry, healthcare.
For a chipmaker, the interest is obvious: every deployed robot consumes compute, both in training and at inference. It is the same logic we described regarding the compute race: whoever sells the infrastructure benefits from uses multiplying everywhere.
For employment, this raises a question we addressed in our article on jobs. We wrote there that anchoring in the physical world protects certain activities, because robotics progresses far more slowly than software. That slowness stems precisely from the difficulty of accumulating experience. If simulation solves part of the problem, the timeline could accelerate. Still, going from a successful move in simulation to a reliable robot in a cluttered warehouse represents years yet.
What to take away
The interesting piece of information is not that a chipmaker releases yet another simulator. It is what this says about the state of the raw material: easy text is exhausted, and the industry is actively looking for where the next deposit will come from.
Two answers now coexist, and they are revealing. Buying up paper printed before 2022 to recover guaranteed human language. Or manufacturing artificial experience by simulating the consequences of actions. One looks backward, toward what we have already written. The other looks toward a synthetic world where learning no longer depends on us. It is not certain that the second is reassuring, but it at least has the merit of destroying no libraries.