For years, the AI revolution has been dominated by language.
We taught machines to read enormous amounts of text. Then we taught them to write. Then they learned to reason, code, analyze images and generate videos. Every major breakthrough seemed to push AI closer to understanding the information humans communicate through language.
But there is a problem with language.
The real world isn't made of words.
A chair isn't a paragraph describing a chair. A building isn't a collection of sentences. A road isn't a text prompt. A robot navigating a warehouse can't simply understand the world by knowing what the word "warehouse" means.
It needs to understand space, objects, movement, physics and cause and effect.
And that could be why one of the next major battles in AI won't be about who has the best language model.
It could be about who builds the best model of the physical world.
AMD's agreement to acquire World Labs for $8.2 billion is a major signal that this idea is moving from research into the center of the AI industry. World Labs, founded by AI researcher Fei-Fei Li, develops “world models” designed to generate and reason about interactive 3D environments. The acquisition is expected to bring that technology and research team into AMD's broader AI strategy.
And I think people are underestimating what this could eventually mean.
A language model predicts what comes next. A world model could predict what happens next.
There is an important difference between knowing information and understanding an environment.
An AI can tell you that dropping a glass will probably cause it to fall.
But a system operating in a physical environment needs much more than that.
It needs to know where the glass is.
Where the floor is.
How far the glass is from the edge of the table.
How quickly an object is moving.
What happens if another object gets in the way.
What actions are possible.
And what the environment will look like after an action is taken.
That is closer to how humans interact with reality.
We don't experience the world as a giant database of sentences. We constantly build internal predictions about what is around us and what might happen next.
If AI can begin doing something similar, the consequences could be enormous.
This could be the missing piece for robotics
One reason robotics has remained much harder than generating text is that physical environments are incredibly complicated.
A language model can generate an answer in a fraction of a second.
A robot can't simply “generate” the correct movement.
It has to understand the environment and physically execute the action.
Consider something as simple as picking up a cup.
A human immediately understands that the cup is an object, that it occupies space, that it can be grasped, that it has weight and that moving it changes its position.
A robot has to convert all of that into data and actions.
A better world model could allow machines to simulate possible outcomes before actually performing them.
Instead of blindly trying something, the AI could effectively ask:
“If I move this object here, what happens next?”
That is a much more powerful form of intelligence.
And this is why 3D AI could become bigger than image generation
AI-generated images are impressive because they produce something that looks real.
But a world model is trying to go further.
The goal isn't simply to generate a picture of a room.
It is to represent a room as an environment.
An object shouldn't merely look like it is sitting on a table. The system should understand that it is sitting on the table.
If the table moves, the object should move with it.
If the object falls, the system should understand where it could land.
If a person walks through the room, the environment should change accordingly.
That distinction is enormous.
A generated image is a visual output.
A world model could become a simulation of reality.
And simulations are useful for much more than entertainment.
Think about self-driving cars
A self-driving system doesn't just need to recognize a car.
It needs to understand what that car is likely to do.
Is it slowing down?
Is it changing lanes?
Is the driver about to turn?
Is the pedestrian going to cross the road?
Is that object actually an obstacle?
The system needs to constantly predict what happens next.
A better world model could potentially make those predictions more sophisticated.
The same principle applies to drones, robots, industrial machines and eventually humanoid robots.
The AI doesn't just need to recognize the world.
It needs to anticipate it.
This could also change how AI learns
There is another reason world models are interesting.
Text-based AI learns heavily from human-created information.
Books.
Websites.
Documents.
Code.
Conversations.
But the physical world contains vastly more information than humans have written down.
There are relationships between objects, physical interactions, movements and environmental changes that don't exist as text.
A model capable of learning from simulated environments could potentially experience enormous numbers of situations without needing a human to describe each one.
A robot could practice navigating a warehouse thousands of times inside a simulation before entering the real warehouse.
A self-driving system could experience millions of simulated road scenarios.
A robotic arm could practice manipulating objects without breaking actual objects.
This could make simulation an enormous training ground for AI.
And suddenly, the company that owns the best world model may have something much more valuable than a cool graphics demo.
It could have a training environment for physical intelligence.
This is where AMD's move becomes particularly interesting
AMD isn't simply buying another chatbot company.
It's bringing World Labs into a hardware company whose business increasingly depends on supplying the computing infrastructure required for AI.
That combination is important.
The next stage of AI may require enormous amounts of computation because models won't just process text.
They may have to continuously simulate environments, objects and physical interactions.
That means the competition isn't only about who has the smartest model.
It is also about who has enough computing power to run these models efficiently.
This is one reason the AI hardware race is becoming so intense.
Bain estimates that the global AI buildout could require roughly $6 trillion of investment through 2030, creating pressure for AI to generate entirely new sources of value rather than simply improving existing productivity.
If world models and physical AI become major applications, they could become part of the answer.
The AI race may therefore be moving from screens into reality
The first AI revolution happened on screens.
Chatbots.
Search.
Coding assistants.
Image generators.
Video generators.
Agents.
But the next one could happen outside the screen.
Robots moving through warehouses.
Machines operating factories.
Vehicles navigating roads.
Drones exploring environments.
AI systems designing and testing physical products.
Virtual worlds that behave more like actual worlds.
And eventually, machines that can learn about reality by interacting with it.
That's a much bigger change than making chatbots better at writing emails.
It means AI would no longer be primarily a technology for manipulating information.
It would become a technology for understanding and acting on the physical world.
And that could make today's AI benchmarks look almost irrelevant
Right now we constantly compare models using tests.
Who scores higher?
Who reasons better?
Who writes better code?
Who has the largest context window?
Those measurements matter.
But imagine an AI that scores slightly worse on a language benchmark but can accurately predict how objects move through a warehouse.
Which one would a robotics company rather have?
The answer is obvious.
As AI moves into the physical world, intelligence will increasingly be measured by something different:
Can the system understand what is happening around it and predict what happens next?
That is a much harder problem.
And it may be one of the most important problems in AI.
The next ChatGPT moment might not look like ChatGPT
That's why I think the World Labs acquisition deserves more attention than simply another billion-dollar AI acquisition headline.
It represents a possible shift in what we're asking AI to understand.
We spent the first part of the AI boom teaching machines how humans communicate.
Now we're beginning to teach them how the world itself works.
If that succeeds, the consequences will go far beyond chatbots.
Because a machine that understands language can tell you how to build something.
A machine that understands the world might eventually build it.
And that could be the point where AI stops being primarily something we use through a screen and becomes something that increasingly operates alongside us in the real world.