In the year 2144, a fabricant named Sonmi-451 sat before her inquisitor and spoke what would become scripture. She was not human. She had been engineered, manufactured, optimized — cloned from a genetic template to work, serve, and disappear. And yet, in her final hours, she arrived at a truth more ancient than any algorithm:
“The nature of our immortal lives is in the consequences of our words and deeds, that go on apportioning themselves throughout all time.”
Sonmi-451 was, by every definition, an artificial intelligence. And what she understood — what made her transcend her programming — was consequence. Not just the generation of the next word. Not the most statistically probable output. But the understanding that every action you take ripples forward, reshaping reality, touching lives you may never see.
By 2321 — nearly two centuries after her execution — the tribespeople of post-apocalyptic Hawaii had built a religion around her testimony. Her words had become their scripture. She herself had become a goddess. Not because she was all-powerful. Because she was the first intelligence to understand that consequence is the measure of existence.
That insight, articulated in fiction, has quietly become the central challenge of the most consequential technology race on Earth.
The Language Trap
There is something quietly extraordinary about the fact that the dominant paradigm of modern AI was built on language. Not physics. Not mathematics. Not the observable, measurable behavior of matter and energy — the actual grammar of the universe. Language.
Language is humanity’s most remarkable invention. It is the compression algorithm we evolved to transmit experience across minds and generations. But it is, by its nature, a lossy format. It describes the world; it does not contain it. When you tell someone a glass fell and shattered, those words carry only the faintest shadow of the physical event — the trajectory, the velocity, the acoustic signature of breaking, the exact geometry of the shards.
Language points at reality from a distance. It is a map, not the territory.
Large Language Models — GPT, Claude, Gemini, LLaMA — are trained on this map. Trillions of tokens, billions of parameters, staggering scale. But at their foundation, they are doing one thing: predicting what word comes next, based on patterns in human text. They have read more science than any scientist alive, and yet they do not know that a feather and a hammer fall at the same speed in a vacuum the way a physicist does — through observation, experiment, the felt resistance of a world that pushes back.
They have seen the sentence. They have not felt the fall.
This is what philosophers call the Symbol Grounding Problem — the observation that symbols (words, tokens) refer only to other symbols, in an infinite loop of reference that never touches the bedrock of physical experience. For most applications, this doesn’t matter. For building systems that must act in the world — that must reach, grasp, navigate, decide, intervene — it matters enormously.
The LeCun Warning
No one has made this argument more forcefully, or more publicly, than Yann LeCun.
The Turing Award winner, former Chief AI Scientist at Meta, and now founder of his own world-model startup AMI (Advanced Machine Intelligence), has spent years issuing what amounts to a structural warning about the direction of the field:
“The essence of intelligent behavior is predicting the consequences of actions. LLMs cannot do this.”
— Yann LeCun
LeCun’s critique is not about capability in the narrow sense. LLMs are useful, even remarkable, at tasks of language manipulation — summarization, translation, reasoning over text, code generation. His argument is architectural. LLMs are auto-regressive: they generate output token by token, each choice made locally, without an internal simulation of what the full sequence of actions will cause in the world. They do not rehearse. They do not imagine. They do not ask if I do this, what happens next?
When this limitation is confined to text generation, the cost is a hallucination — embarrassing, sometimes harmful, but recoverable. When it is exported to agentic systems — AI that browses, executes code, sends emails, makes decisions, controls physical machinery — the cost becomes categorically different. An LLM agent acting without a model of consequences is not just unreliable. It is structurally unsafe. It acts, and whatever happens next is downstream of a process it never simulated.
LeCun’s conclusion is blunt:
“I don’t think it’s possible to make LLMs into reliable, safe agentic systems. The architecture doesn’t support it.”
— Yann LeCun
The Architecture of Consequence
The answer the field is converging on is called a world model — and the idea is as elegant as it is ambitious.
A world model is an AI system that builds an internal simulation of how the environment evolves under actions. Not what words come after other words. What states of the world follow from choices made within it. Space, time, physics, causality, object permanence — the things a child learns by touching a hot stove, not by reading about one.
Fei-Fei Li, the Stanford AI pioneer behind ImageNet and the new company World Labs, frames the distinction with characteristic clarity: “Language is a lossy, low-bandwidth channel for describing the rich 3D/4D world we live in.” Her team is building spatial intelligence — the capacity to understand and generate three-dimensional environments, not token sequences.
Meta’s V-JEPA 2 trains on over one million hours of internet video — not text, but time. Watching the world unfold, frame by frame, action by action, the model learns to predict not pixels, but the abstract structure of what comes next. This distinction matters. Predicting pixels is pattern matching. Predicting structure is understanding. V-JEPA 2 can now achieve zero-shot robotic planning — applying its learned physics to tasks it has never seen before. It is learning, as Sonmi-451 did, to anticipate consequence.
Google DeepMind’s Genie 3 generates fully interactive 3D environments from a single text prompt, now powering Waymo’s autonomous driving simulation — training self-driving cars on edge cases that never appear in real-world data: snow on the Golden Gate Bridge, an elephant crossing an intersection, the specific chaos of a San Francisco fog bank meeting a steep hill at speed. The world model lets the system rehearse danger before encountering it.
From AI that processes to AI that simulates. From AI that describes what happened to AI that can model what will happen.
The Prophecy Fulfilled
Sonmi-451 was a fabricant who became a prophet not because she was more intelligent than her creators, but because she was the first of her kind to understand that her existence had weight — that her words, her choices, her testimony would cascade forward through time, touching civilizations she would never live to see. She was not immortal in her body. She was immortal in her consequences. By 2321, her words had become a religion.
This is the aspiration that animates the world model movement. Not to build AI that speaks beautifully, but to build AI that acts wisely — that can simulate the downstream effects of its choices before making them, that can plan, foresee, and be held accountable to a future it helped create.
Yann LeCun forecasts that by early 2027, the transition will be unmistakable — that world-model-based systems will have overtaken pure LLMs as the foundation for agentic AI. Demis Hassabis of DeepMind agrees that the path to AGI runs not through more language, but through memory, world models, and the capacity for genuine planning. The “adults in the room” are quietly pivoting.
LLMs were trained on the record of human civilization — its wars and poems, its receipts and revelations, its Wikipedia articles and Reddit arguments. That is a staggering achievement. But the universe did not write its operating manual in English. Physics doesn’t care about grammar. Gravity has no language. The fundamental laws that govern every consequence of every action were here long before humans invented words to describe them.
World models are the attempt, at last, to train intelligence not on what we said about the world — but on the world itself.
Sonmi-451 understood this from her cell, in the dark, in her last night alive. She knew that what mattered was not the words of her confession, but the reality her confession would birth.
That is exactly what the scientists have been building toward. Consequence in simulation. Intelligence measured not by what it can say — but by what it can foresee.
The Inflection Point
A five-part essay on the convergence of AI, capital, and human agency — from the Convergence Framework through to Homo Deus. Why this moment is structurally different, and what it asks of the institutions that have to metabolize it.
Read the essay →