August 17, 2026

World models and the next chapter of model training

Mark Esposito

,

Chief Economist at micro1

For years, the AI conversation has been organized around a powerful paradigm: language. Train a model on text, scale the data, scale the compute, and watch capability emerge. That paradigm produced extraordinary results. Large language models learned to code, to reason, to hold conversations that feel human. They conquered the screen.

But the screen is one environment among many. The physical world, the one where objects have weight, surfaces have friction, and actions have consequences, remains largely unsolved. Language models can describe how a cup falls off a table. They cannot predict the trajectory. They can narrate a warehouse. They cannot navigate one. For AI to move from digital environments into the real world, it needs something language alone cannot provide: an understanding of physics.

That is the premise behind world models. And for the first time, the technology, the data, and the engineering talent are converging to make them real.

What world models actually are

A world model is not a language model with better prompts. It is a fundamentally different kind of system, one that learns how the physical world behaves by observing it, not by reading about it.Where a language model predicts the next token in a sequence of text, a world model predicts the next state of a physical environment. Seed it with an image of a cup on a coffee table, and it will continue to produce a cup on a coffee table. But ask it to push the cup versus pull the cup, and the prediction changes. The model is not generating plausible text. It is simulating causality.The architectural distinctions matter. The most promising world models are causal, meaning they only look backward and at the present, never forward. They are action-conditioned, meaning the prediction depends on what you do next. And they are autoregressive, building output one piece at a time, where each piece depends on everything generated before it. This is a different computational posture than the bidirectional video generation models most people have encountered, which take a start frame and an end frame and fill in the middle. A world model does not know the ending. It unfolds.

Why language models are limited, not wrong

There is nothing fundamentally broken about taking a pre-trained language model and fine-tuning it on robot data. The issue is that the approach is fundamentally limited.Language models are very good at modeling temporal coherence. They handle sequence, logic, and structure well. But they are not modeling the explicit physics of the world. They do not learn what happens when you change the state of a physical environment. They do not generalize well to novel situations, new settings, or unfamiliar objects. The shortcomings show up reliably, and they show up at the moments that matter most: the edge cases, the unseen environments, the tasks that require genuine physical intuition.

The alternative is to pretrain world models on embodied data from the start, building the right priors into the system rather than hoping they emerge from fine-tuning. The bet is that models trained on rich, multimodal, physically grounded data (not just images and video, but tactile information, proprioceptive signals, and interactive sensory inputs) will generalize in ways that text-derived models structurally cannot.

That bet is starting to pay off. Early results show that these architectures generalize more effectively to novel tasks and novel data than vision-language models fine-tuned for robotics. The gap is not yet decisive, but it is consistent, and it is growing.

The evaluation problem no one has solved

If world models are going to power robots, autonomous systems, and interactive simulations, someone has to be able to say whether they are working. Right now, that is harder than it should be. The evaluation landscape for world models is nascent. There is no equivalent of the standardized benchmarks that exist for language models. Physical plausibility, temporal consistency, and frame-to-frame object stability are all dimensions that matter but measuring them rigorously is an open problem.One insight that is sharpening the conversation: there is a critical difference between a model that produces a plausible outcome and a model that represents the true distribution of possible outcomes. A world model can generate a convincing frame of a ball rolling down an incline. But if you run it a hundred times, does it capture the variance? Does it represent the full range of trajectories, including the unlikely ones? A double pendulum is chaotic. Simulating it correctly requires propagating uncertainty forward, not collapsing to the most likely path.This reframes physical accuracy from a visual test you pass once to a distributional property you prove over and over. The field is only beginning to build evaluation infrastructure at that level of rigor. Until it does, the gap between demo and deployment will persist.

Gaming as a proving ground

One of the less obvious but most instructive applications for world models is gaming. Games represent controlled environments that are low-consequence, interactive, and rich enough to stress-test a model's understanding of how worlds behave.This is not a distraction from the harder problems in robotics. It is a stepping stone. Entertainment is one of the platforms where world model capabilities are most mature today, precisely because the stakes are lower and the feedback loops are faster. Being able to explore and interact with a generated world in real time, even imperfectly, is a meaningful proof of concept for the same technology that will eventually power embodied agents in physical environments.The connection runs deeper than surface similarity. General-purpose world models that serve gaming, simulation, and robotics may converge on shared architectures. The physics engine that makes a game world feel real and the prediction engine that lets a robot anticipate what happens next are solving closely related problems with overlapping methods.

Humans are still in the loop at every stage

In language model development, the role of human expertise in training and evaluation is well understood. In robotics and world models, the same dependency exists, but it is less visible and, in some ways, more demanding.

Human judgment is needed at every stage: in verifying whether simulated data is plausible, in assessing whether a model's structuring and annotation of physical scenes is accurate, in evaluating edge cases that automated metrics cannot catch. Even when data is generated synthetically or through simulation, humans are required to validate the output. The analogy to self-driving is instructive. Waymo started with humans driving cars to create data. Simulation expanded the data envelope. But humans remained essential to verify that the simulated edge cases, the ones too dangerous or too rare to replicate in the real world, were handled correctly.

The type of expertise required evolves. In early-stage pre-training, generalist evaluators who understand how the physical world works are sufficient. As models specialize for particular environments (a dentist's office, a mine, a warehouse), domain experts become necessary. This mirrors exactly the trajectory we have seen in language model evaluation, where generalist annotators gave way to subject-matter experts as models matured. The same will happen in robotics.

The end of demo culture

For years, the robotics industry has operated in what might fairly be called demo culture. Impressive demonstrations, carefully staged environments, curated tasks. These demos attract funding, generate excitement, and showcase real capability. But they are not deployment.

That era is beginning to close. Companies are now shipping robots into real environments, not as prototypes in controlled labs, but as products that coexist with people, interact with unstructured settings, and generate live operational data. The robots are not yet as efficient, as fast, or as reliable as humans. They do not need to be. The hill-climbing function that transformed language models from a novelty (write a haiku in the voice of Snoop Dogg) into a daily productivity tool is beginning for physical AI.

The moment when robots enter the cultural awareness as things that simply exist in everyday life, doing real tasks alongside people, is closer than most observers expect. It will not arrive with a single breakthrough. It will arrive incrementally, the way useful technology always does: imperfect at first, then indispensable.

Where this goes

The path from here is a phased one. Build a generalized foundation, a world model with a broad understanding of how physical environments behave. Then specialize. Deploy in specific environments, validate which data is useful, expand to richer sensor suites, and iterate. Start with navigation, where autonomy is further along, and move toward complex manipulation as capability matures.

Six to twelve months from now, we should begin to see real progress on simpler autonomous tasks. Not full autonomy, but meaningful, verifiable capability in constrained environments. The underlying technology is not new. Diffusion transformers are a five-year-old architecture. What is new is the sophistication of the data pipelines, the infrastructure, and the scaling strategies being built around them.

The breakthrough, when it comes, may feel sudden. It always does. But the work happening now, in evaluation, in data infrastructure, in multimodal model design, is what will make that moment possible. The future of AI is not only about language. It is about physics. And that future is being built right now.

This piece draws on a conversation hosted by the micro1 Virtual Series, featuring Arian Sadeghi (micro1), Derek Sarshad (Odyssey), and Sam Sinha (1X)