From Language Models to World Models: The Next Evolution of Artificial Intelligence

Why ChatGPT was only the beginning of the AI revolution

The extraordinary success of ChatGPT just in a few years transformed the public perception of artificial intelligence more profoundly than any technological breakthrough since the emergence of the modern Internet. Within only a few months, millions of people who had never previously interacted with advanced AI systems found themselves conversing naturally with machines capable of answering questions, writing essays, generating software code, translating between languages, analysing scientific papers and producing creative content with remarkable fluency. What had long remained the preserve of research laboratories suddenly became an everyday tool for students, researchers, businesses, governments and ordinary citizens across the world. The release of ChatGPT did not simply introduce a new application; it fundamentally altered society’s understanding of what artificial intelligence had already become.

This remarkable progress was driven by the emergence of Large Language Models (LLMs), enormous neural networks trained on unprecedented quantities of digital text collected from books, scientific publications, websites, software repositories and many other sources. Through exposure to trillions of words, these systems gradually learned the statistical relationships that govern human language, allowing them to generate coherent responses, perform sophisticated reasoning tasks and reproduce many aspects of written human communication. Their achievements demonstrated that language itself contains an extraordinary amount of structured knowledge about the world, and that sufficiently large computational systems can learn to manipulate that knowledge with a level of fluency previously thought to require human intelligence.

Yet the spectacular capabilities of language models have also made their limitations increasingly visible. Although an LLM can describe how to assemble a piece of furniture, explain the laws of motion or discuss the mechanics of a robotic arm, it does not directly experience physical reality. It has never observed gravity acting upon an object, manipulated tools with robotic hands or navigated through a three-dimensional environment. Its understanding of the world is mediated entirely through language, images and other symbolic representations created by human beings. Consequently, even the most advanced conversational systems continue to operate primarily within the domain of symbols rather than within the physical world itself.

This distinction has become one of the central questions shaping the next phase of artificial intelligence research. If future AI systems are expected not merely to converse with people but also to assist in scientific discovery, operate autonomous robots, manage industrial processes, explore hazardous environments or interact safely with the physical world, then linguistic competence alone will no longer be sufficient. Such systems must develop an internal understanding of objects, space, movement, causality and physical interaction. They must learn not only to describe reality but also to anticipate how reality changes over time. In other words, they require something that language models were never originally designed to provide: an internal representation of the world itself.

This emerging research direction has given rise to one of the most significant conceptual developments in contemporary artificial intelligence: the creation of World Models. Rather than focusing exclusively on predicting the next word in a sentence, these new architectures seek to construct computational representations of the surrounding environment, enabling machines to reason about physical processes, anticipate future events and simulate the consequences of possible actions before they occur. The ambition is not simply to produce more knowledgeable conversational systems but to develop intelligent agents capable of understanding how the world functions in much the same way that humans gradually learn through observation, interaction and experience.

For many researchers, this transition represents the natural continuation of the AI revolution initiated by large language models. ChatGPT demonstrated that machines could become remarkably proficient at understanding and generating language. World Models attempt to extend that achievement into a much broader form of intelligence by allowing artificial systems to acquire an intuitive understanding of the physical environment in which language itself ultimately acquires meaning. If successful, they could transform not only robotics and autonomous vehicles but also scientific simulation, industrial automation, healthcare, engineering, education and countless other domains that depend upon understanding the relationship between thought and action.

Seen from a historical perspective, the emergence of World Models may eventually be regarded as one of the defining moments in the evolution of artificial intelligence. Just as the introduction of transformer architectures fundamentally reshaped machine learning by enabling the rise of modern language models, the development of systems capable of modelling the physical world may mark the beginning of a new stage in which artificial intelligence gradually expands beyond the interpretation of language towards the comprehension of reality itself. Whether this transition ultimately leads to more general forms of machine intelligence remains uncertain, but it already signals a profound shift in the ambitions of AI research. The next frontier is no longer simply to build machines that speak more convincingly. It is to build machines that understand the world that language describes.

The Limits of Large Language Models

The astonishing achievements of Large Language Models have sometimes created the impression that artificial intelligence is approaching a point at which every aspect of human cognition can be reproduced simply by increasing computational power, expanding training datasets and refining existing architectures. Indeed, the rapid progression from GPT-2 to GPT-4, Claude, Gemini and other frontier systems has repeatedly demonstrated that scaling language models produces remarkable improvements across an extraordinarily wide range of cognitive tasks. Mathematical reasoning, software development, scientific explanation, multilingual translation and complex document analysis have all benefited from larger models trained on increasingly diverse sources of textual information. Such successes naturally encouraged the belief that language itself might provide a sufficient foundation upon which more general forms of intelligence could eventually emerge.

However, many researchers have argued that this conclusion reflects a misunderstanding of both language and intelligence. Language undoubtedly encodes a vast amount of knowledge about the world, accumulated across centuries of human observation and cultural development. Through books, conversations, scientific publications and historical records, people describe objects, explain physical phenomena, discuss intentions, analyse causes and recount the consequences of actions. An artificial intelligence trained upon these texts therefore acquires an extraordinarily rich statistical representation of how humans talk about reality. Yet talking about the world and directly understanding the world are not necessarily equivalent cognitive achievements.

To appreciate this distinction, it is useful to consider the way in which human beings themselves acquire knowledge during early childhood. Long before children are capable of speaking fluently, they begin constructing internal models of the physical environment through continuous interaction with the world around them. They learn that unsupported objects fall under gravity, that solid walls cannot be passed through, that fragile objects break under sufficient force and that other people possess intentions that influence their behaviour. These forms of understanding emerge not primarily through verbal explanation but through direct sensory experience involving vision, touch, movement and repeated interaction with the environment. Language subsequently provides names and concepts for experiences that have already become deeply embedded within an internal representation of reality.

Large Language Models follow an almost opposite developmental trajectory. Rather than beginning with perception and interaction, they begin with language alone. Everything they know about gravity, friction, geometry, human movement or physical causality has been inferred indirectly from written descriptions created by people who themselves acquired that knowledge through embodied experience. As a consequence, LLMs frequently demonstrate an impressive ability to discuss physical processes while simultaneously lacking the intuitive understanding that humans develop through direct engagement with the material world. They may correctly explain the principles governing balance or momentum, yet fail in situations requiring genuine physical intuition because their reasoning remains fundamentally grounded in symbolic associations rather than experiential representations.

This limitation becomes particularly evident whenever artificial intelligence must move beyond conversation and interact directly with its environment. A conversational assistant can accurately explain how to pick up a glass resting on a table, describing the sequence of movements required and even discussing the physics involved. A household robot attempting the same task faces an entirely different challenge. It must estimate the exact position of the object, calculate an appropriate trajectory, determine the orientation of its mechanical fingers, apply sufficient force to secure the glass without crushing it and continuously adjust its movements in response to changing sensory information. Each of these decisions depends upon an internal understanding of space, geometry, physical interaction and temporal dynamics that cannot be derived solely from linguistic knowledge.

Similar limitations appear across numerous domains extending far beyond robotics. Autonomous vehicles must anticipate the behaviour of pedestrians, cyclists and other drivers while continuously predicting how the surrounding environment will evolve over the next few seconds. Scientific simulations require an understanding of causal relationships governing physical systems rather than merely textual descriptions of those systems. Industrial automation depends upon precise reasoning about mechanical interactions, spatial constraints and dynamic processes unfolding in real time. Even seemingly simple tasks, such as organising objects on a shelf or navigating through an unfamiliar building, involve forms of physical reasoning that remain fundamentally different from generating coherent language.

These observations do not diminish the extraordinary accomplishments of Large Language Models. On the contrary, they clarify the remarkable achievement they already represent. LLMs have demonstrated that language contains sufficient structure to support highly sophisticated reasoning across an enormous range of intellectual activities. Nevertheless, they also reveal that language alone may not constitute a complete foundation for general intelligence. Human cognition integrates language with perception, memory, action, physical intuition and continuous interaction with the surrounding world. If artificial intelligence is eventually to approach comparable flexibility, it may likewise require architectures capable of combining linguistic reasoning with richer representations of reality itself.

It is precisely this challenge that has inspired the emergence of World Models. Rather than replacing Large Language Models, these new systems seek to complement them by providing a computational framework through which artificial intelligence can gradually learn not merely to describe the world but to model its underlying structure, anticipate its evolution and reason about the consequences of action within it. In many respects, this represents one of the most important conceptual transitions currently taking place in artificial intelligence research, shifting attention from machines that understand language towards machines that understand the reality from which language ultimately emerges.

What Is a World Model? Teaching Machines to Build an Internal Reality

The concept of a World Model is, in many respects, both simple and profoundly ambitious. At its core lies the idea that an intelligent system should not merely react to information presented at a particular moment, but should gradually construct an internal representation of the external world that allows it to interpret events, anticipate future developments and plan its actions accordingly. Rather than treating each observation as an isolated piece of data, a World Model attempts to organise experience into a coherent mental structure describing how objects, people, environments and physical processes relate to one another over time. In essence, it seeks to give artificial intelligence something resembling what cognitive scientists often describe as a mental model of reality.

This notion is far from new within the broader history of artificial intelligence. Researchers have long recognised that intelligence requires more than pattern recognition alone. During the earliest decades of AI, many symbolic systems attempted to represent aspects of the external world through logical rules and explicit knowledge bases. These approaches achieved notable successes within carefully defined environments but struggled to cope with the complexity and unpredictability of the real world. The rise of deep learning shifted attention towards statistical learning from enormous datasets, enabling machines to recognise images, translate languages and eventually generate human-like text with unprecedented accuracy. Yet despite these advances, the fundamental question remained unresolved: how can a machine develop an internal understanding of reality rather than merely learning statistical correlations within data?

World Models represent one of the most promising contemporary answers to that question. Instead of learning exclusively from written language, these systems are trained using streams of sensory information that more closely resemble the experiences through which humans and animals acquire knowledge. Videos, three-dimensional environments, depth information, motion sequences, robotic sensor readings and other forms of multimodal data provide continuous records of how the physical world evolves. By analysing these observations, the AI gradually begins to infer regularities governing movement, spatial relationships, object permanence, physical interaction and causal structure. Over time, these regularities are integrated into an internal computational representation capable of predicting what is likely to happen next under changing circumstances.

One useful way to understand this distinction is to compare reading a description of a city with actually navigating through its streets. A traditional language model resembles someone who has read thousands of books describing the city in extraordinary detail. Such a person may know its history, recognise the names of its districts and explain how different neighbourhoods are connected. A World Model, by contrast, resembles someone who has physically explored the city, learning how distances feel, how traffic flows, where obstacles arise and how movement through space unfolds in practice. Both forms of knowledge are valuable, yet they differ fundamentally in the kinds of reasoning they support. The first excels at explanation, while the second enables prediction, navigation and action.

This predictive capability is perhaps the defining characteristic of World Models. Rather than simply identifying what exists within an environment, they attempt to forecast how that environment is likely to change as events unfold. If an object begins to fall from a table, the system predicts its trajectory before it reaches the floor. If a pedestrian steps onto a road, it anticipates the possible paths both the pedestrian and approaching vehicles may follow. If a robot reaches towards an object, it estimates how the object will respond to different movements before the action is executed. These predictions allow the machine to evaluate multiple possible futures internally, selecting actions based upon simulated consequences rather than relying exclusively upon trial and error in the real world.

This ability to simulate future outcomes closely resembles an essential feature of human cognition. Much of human intelligence depends upon imagining possible scenarios before acting. People routinely consider alternative courses of action, estimate likely consequences and mentally rehearse complex tasks without physically performing them. A surgeon visualises an operation before making the first incision; an architect imagines how a building will stand before construction begins; a driver anticipates the behaviour of surrounding vehicles before changing lanes. Such mental simulation dramatically reduces risk while improving planning and decision-making. World Models aspire to provide machines with a computational analogue of this capacity, allowing them to “think ahead” before interacting with their environment.

Importantly, this does not imply that World Models possess consciousness, subjective experience or human-like understanding in the philosophical sense. Their internal representations remain mathematical structures learned through optimisation rather than lived experience. Nevertheless, these representations enable forms of reasoning that extend beyond the capabilities of purely language-based systems. By modelling the dynamics of the external world rather than merely describing it, they allow artificial intelligence to bridge the gap between perception and action, transforming static knowledge into predictive intelligence capable of supporting increasingly autonomous behaviour.

For this reason, many researchers regard World Models as one of the most significant conceptual developments since the emergence of transformer-based language models. They shift the objective of artificial intelligence from predicting sequences of symbols towards modelling the underlying processes that generate those symbols in the first place. Language remains enormously important, but it becomes only one expression of a richer computational understanding grounded in space, time, causality and interaction. If Large Language Models taught machines to communicate with remarkable fluency, World Models seek to teach them why the world behaves as it does.

From Predicting Words to Predicting Reality

The extraordinary success of Large Language Models rests upon a deceptively simple learning principle. During training, the model repeatedly encounters sequences of words in which one element has been removed, and its task is to predict what is most likely to appear next. This process is repeated trillions of times across vast collections of books, articles, conversations, software code and countless other forms of digital text. Although the task appears almost trivial in isolation, its repetition on an unprecedented scale enables the model to acquire an astonishingly sophisticated understanding of grammar, semantics, reasoning and human communication. In effect, the ability to predict the next word gradually gives rise to many of the intellectual capabilities that have made systems such as ChatGPT so transformative.

Yet from the perspective of World Models, predicting language represents only a special case of a much broader cognitive challenge. The physical world also unfolds as a sequence of events, but instead of predicting the next word, an intelligent agent must predict the next state of reality. If a ball rolls across a table, where will it move next? If a person reaches towards a door, what action are they likely to perform? If two vehicles approach an intersection simultaneously, how will their trajectories interact? These questions concern not linguistic continuation but physical evolution. They require an understanding of motion, geometry, intention and causality that extends beyond the statistical regularities embedded within written language.

This distinction may appear subtle, yet it represents one of the most profound conceptual shifts currently taking place within artificial intelligence research. A language model learns primarily by identifying patterns within symbolic sequences. A World Model learns by identifying patterns within the evolution of reality itself. The objects, movements and interactions observed through cameras, sensors or simulated environments become analogous to the words and sentences that fuelled the first generation of conversational AI. Instead of learning how one sentence follows another, the machine begins learning how one physical state gives rise to the next. Intelligence thus becomes increasingly grounded in the dynamics of the environment rather than solely in the structure of human language.

Central to this transition is the concept of causality. Language often describes causal relationships, but descriptions remain one step removed from the underlying processes they represent. A sentence may explain that dropping a glass causes it to shatter upon the floor, yet the sentence itself does not embody the physics governing gravity, momentum, impact and material fracture. Humans understand these relationships because they have repeatedly observed and experienced similar events throughout their lives. World Models attempt to reconstruct this form of understanding computationally by learning directly from sequences of observations in which causes and consequences unfold naturally over time. Rather than memorising verbal explanations, they infer the regularities connecting events within the world itself.

This emphasis on causality distinguishes World Models from many earlier forms of machine learning. Traditional pattern recognition systems often excelled at identifying correlations without necessarily understanding why those correlations existed. A model might learn that dark clouds frequently precede rainfall or that certain traffic conditions often lead to congestion, yet remain unable to distinguish genuine causal mechanisms from coincidental statistical associations. By modelling how environments evolve through time, World Models aspire to capture deeper structural relationships governing physical processes. Although this objective remains technically challenging, it represents an important step towards more robust and reliable forms of machine reasoning.

Equally significant is the emergence of counterfactual reasoning, the capacity to consider hypothetical alternatives before acting. Humans constantly ask themselves questions such as “What would happen if I took another route?”, “What if I used a different tool?” or “What if I waited a few more minutes?” Such reasoning depends upon mentally simulating futures that have not yet occurred. World Models seek to endow machines with an analogous capability by allowing them to evaluate multiple possible outcomes internally before selecting a course of action. This capacity is particularly valuable for autonomous systems operating in complex environments where experimentation through physical trial and error would be inefficient, expensive or potentially dangerous.

The implications extend far beyond robotics. Scientific research increasingly depends upon simulations capable of predicting climate behaviour, molecular interactions or biological processes before costly experiments are conducted. Engineers routinely evaluate digital prototypes before manufacturing physical products. Urban planners model traffic flows and infrastructure projects years before construction begins. In each of these domains, the ability to simulate reality accurately provides enormous practical advantages. World Models therefore represent not merely a new form of artificial intelligence but a general computational framework for reasoning about dynamic systems whose behaviour unfolds across space and time.

From a historical perspective, this transition from predicting language to predicting reality may eventually be viewed as one of the defining moments in the evolution of machine intelligence. The first generation of foundation models demonstrated that statistical prediction could produce machines capable of extraordinary linguistic competence. The emerging generation of World Models extends that principle beyond language, applying predictive learning to the physical environment itself. If successful, this evolution could fundamentally redefine the objectives of artificial intelligence, shifting its centre of gravity from communication towards understanding, anticipation and purposeful interaction with the world.

How Humans Learn About the World: Inspiration from Cognitive Science

One of the most distinctive characteristics of the current generation of World Models is that they draw inspiration not only from computer science but also from cognitive science, developmental psychology and neuroscience. While the first wave of Large Language Models demonstrated that extraordinary capabilities could emerge from statistical learning over vast collections of text, many researchers argue that human intelligence develops according to a far richer process, one in which language represents only a relatively late stage in cognitive development. To build machines capable of interacting naturally with the physical world, it therefore becomes essential to understand how human beings themselves acquire their intuitive knowledge of reality.

The developmental journey of an infant provides perhaps the clearest illustration of this process. During the earliest months of life, children possess only limited linguistic ability, yet they rapidly begin constructing increasingly sophisticated internal models of their surroundings. Long before they can explain the concept of gravity, they learn that unsupported objects fall. Before understanding geometry, they recognise that some spaces are too narrow to pass through. Before mastering the vocabulary of intention, they begin interpreting the movements and expressions of other people as purposeful actions rather than random events. This accumulation of knowledge does not arise primarily from verbal instruction but from continuous interaction with a rich and dynamic environment in which perception, movement and feedback are inseparably linked.

Developmental psychologists have long described this process as the gradual construction of an internal model of reality. Rather than memorising isolated facts, children organise experience into coherent frameworks that allow them to predict what is likely to happen under different circumstances. They discover that objects continue to exist even when temporarily hidden from view, that physical actions have predictable consequences and that living beings behave differently from inanimate objects. Every successful prediction reinforces these internal representations, while every unexpected event forces them to revise their understanding. Learning therefore becomes a continuous dialogue between expectation and observation, through which increasingly accurate models of the world emerge over time.

This principle has acquired growing importance within contemporary neuroscience through theories collectively known as predictive processing or predictive coding. According to these frameworks, the human brain functions not merely as a passive receiver of sensory information but as an active prediction engine. At every moment, the brain continuously generates expectations about incoming sensory experiences and compares those expectations with reality. Whenever discrepancies arise, known as prediction errors, the internal model is updated to better reflect the external world. Perception, under this view, is not simply the registration of sensory inputs but an ongoing process of hypothesis testing in which the brain constantly refines its representation of reality.

Although predictive processing remains an active area of scientific research, its central insight has profoundly influenced the thinking of several leading artificial intelligence researchers. If biological intelligence depends fundamentally upon building predictive models of the environment, then perhaps artificial intelligence should follow a similar trajectory. Rather than relying exclusively upon language prediction, future AI systems might benefit from learning to predict the evolution of the physical world itself. Such systems would gradually acquire intuitive knowledge not because they had memorised explicit rules, but because they had repeatedly refined internal models capable of anticipating future observations with increasing accuracy.

This perspective helps explain why many researchers regard World Models as a natural continuation rather than a replacement of Large Language Models. Human cognition integrates multiple forms of intelligence simultaneously. People combine linguistic reasoning with visual perception, spatial awareness, motor coordination, social understanding and physical intuition without consciously separating these domains. A child does not first complete the study of language before beginning to understand the physical world; both capacities develop together through continuous interaction with the environment. World Models seek to restore this broader conception of intelligence by complementing linguistic knowledge with computational representations of physical experience.

An important consequence of this approach is that learning increasingly depends upon experience rather than description. Reading thousands of books about riding a bicycle does not teach a person how to maintain balance. Watching repeated demonstrations, attempting the movement and correcting mistakes through sensory feedback gradually produce an intuitive understanding that cannot easily be expressed through language alone. The same principle applies to countless everyday activities, from catching a ball to navigating unfamiliar environments or manipulating delicate objects. These abilities rely upon forms of embodied prediction that emerge through interaction rather than explicit instruction. By exposing AI systems to videos, simulations and sensory streams instead of text alone, researchers hope to cultivate comparable forms of computational intuition.

It is important, however, not to overstate the analogy between biological and artificial intelligence. Human learning remains profoundly shaped by consciousness, emotion, motivation, biological evolution and social interaction in ways that current AI systems do not replicate. World Models are not attempts to recreate the human brain neuron by neuron, nor do they imply that machines experience the world subjectively. Instead, they borrow selected computational principles that appear to contribute to efficient learning in biological systems. The objective is not to imitate humanity in every respect but to identify those mechanisms that enable robust prediction, flexible adaptation and effective interaction with complex environments.

For this reason, the growing dialogue between neuroscience and artificial intelligence represents one of the most fascinating developments in contemporary research. As AI increasingly seeks to understand the world rather than merely describe it, insights from the study of human cognition become progressively more relevant. At the same time, advances in computational modelling provide neuroscientists with new tools for testing theories concerning perception, learning and predictive reasoning. The relationship is therefore becoming increasingly reciprocal, with each discipline informing and enriching the other. World Models stand at the centre of this convergence, embodying the idea that the future of artificial intelligence may depend as much upon understanding human cognition as upon increasing computational power.

JEPA, World Labs and the New Generation of AI Research

The growing interest in World Models is not confined to theoretical discussions within academic research. During the past few years, several of the world’s most influential AI laboratories have begun developing new architectures specifically designed to move beyond the limitations of purely language-based systems. Although these initiatives differ considerably in their technical implementation, they share a common ambition: to enable machines to construct richer internal representations of the physical world and to reason about future events before they occur. Among the most influential of these efforts are Meta’s Joint Embedding Predictive Architecture (JEPA), World Labs, founded by Fei-Fei Li, and the research programme pursued by Yann LeCun through AMI Labs. Together, they illustrate the remarkable diversity of approaches currently shaping the next phase of artificial intelligence.

Perhaps the most widely discussed of these initiatives is JEPA, developed under the scientific leadership of Yann LeCun, Meta’s Chief AI Scientist and one of the pioneers of modern deep learning. LeCun has long argued that predicting individual words, although extraordinarily successful for language modelling, cannot by itself produce the kind of robust understanding required for general intelligence. Instead, he proposes that intelligent systems should learn by predicting abstract representations of future observations rather than attempting to reconstruct every detail of incoming sensory data. This distinction may appear technical, yet it reflects a fundamentally different philosophy of learning.

Traditional generative models often attempt to predict every pixel of an image or every word within a sentence. JEPA, by contrast, seeks to predict the underlying conceptual structure that connects successive observations. If a person walks behind a tree and momentarily disappears from view, the system does not need to reconstruct every visual detail hidden by the tree. Instead, it learns that the person continues to exist and is likely to emerge on the other side because its internal model captures the broader causal structure of the scene. In this way, prediction becomes less concerned with superficial appearance and more focused on understanding the persistent organisation of the environment itself. This ability to reason about hidden states and incomplete observations resembles an important characteristic of human perception, where much of what we “see” consists not of direct sensory input but of expectations generated by internal models.

A complementary perspective has emerged through World Labs, the company co-founded by Fei-Fei Li, whose pioneering contributions to computer vision helped lay many of the foundations upon which modern visual AI has been built. Fei-Fei Li has consistently argued that intelligence requires more than recognising isolated images; it requires understanding the three-dimensional structure of the world within which those images are embedded. Human beings perceive depth, spatial relationships, object permanence and environmental continuity almost effortlessly, yet traditional computer vision systems have often treated each image as an independent snapshot. World Labs seeks to overcome this limitation by developing AI systems capable of constructing coherent three-dimensional representations from visual experience, enabling machines to reason about space in ways that more closely resemble human perception.

This emphasis on three-dimensional understanding has profound practical implications. A robot navigating a cluttered room, an autonomous drone inspecting industrial infrastructure or a digital assistant helping architects design buildings must all reason about geometry, perspective and spatial relationships rather than merely recognising objects in two-dimensional photographs. By enabling machines to reconstruct the structure of their surroundings, World Models open the possibility of AI systems that interact naturally with complex environments instead of responding only to isolated visual inputs. In this respect, World Labs extends the achievements of computer vision towards a broader conception of spatial intelligence.

Meanwhile, AMI Labs, another initiative associated with Yann LeCun’s long-term research vision, explores how autonomous systems might learn to predict the consequences of their own actions before carrying them out. This capability represents one of the defining characteristics of intelligent behaviour. Humans rarely act without first considering likely outcomes, mentally rehearsing possible scenarios and selecting among alternative strategies. An autonomous machine operating within an unpredictable environment requires comparable abilities if it is to perform safely and effectively. Rather than relying upon exhaustive trial and error, it must learn to simulate future possibilities internally, evaluating different courses of action before interacting with the physical world. This approach has particular relevance for robotics, autonomous vehicles and other applications where mistakes may carry significant physical consequences.

Despite their different technical strategies, these research programmes share several important assumptions. All recognise that intelligence depends upon prediction rather than simple reaction. All emphasise the importance of learning from continuous sensory experience rather than text alone. And all seek to equip machines with internal representations capable of supporting planning, reasoning and adaptation within dynamic environments. Collectively, they signal a broader shift within artificial intelligence research away from systems specialised in processing language towards architectures designed to understand the underlying structure of reality itself.

Historically, this convergence is particularly significant because it demonstrates that the transition towards World Models is not the vision of a single research group but an emerging direction embraced by many of the field’s leading scientists. Just as the transformer architecture rapidly became the dominant foundation for language models after demonstrating its effectiveness, the coming decade may witness increasing convergence around predictive world modelling as a central component of more general forms of machine intelligence. Whether a single architecture ultimately prevails or multiple complementary approaches continue to coexist remains uncertain. What already appears clear, however, is that the centre of gravity in artificial intelligence research is beginning to shift from understanding language towards understanding the world that language seeks to describe.

Physical Intelligence and the Future of Robotics

One of the clearest demonstrations of why World Models matter can be found in the field of robotics. While conversational artificial intelligence has transformed the way humans interact with digital systems, robots operate under a fundamentally different set of constraints. Language unfolds within a symbolic environment in which mistakes are often harmless and can usually be corrected through further conversation. The physical world offers no such flexibility. A robot that misunderstands the position of an object, misjudges the force required to grasp a fragile glass or incorrectly predicts the movement of a nearby person may immediately generate costly or even dangerous consequences. For machines acting in physical environments, understanding reality is not simply advantageous; it is an essential prerequisite for safe and effective behaviour.

This distinction highlights one of the most important limitations of purely language-based intelligence. A conversational system may produce an impeccable explanation of how to prepare a meal, assemble industrial equipment or assist an elderly person in daily activities. Yet the successful execution of these tasks requires an entirely different form of cognition. A domestic robot must continuously estimate distances, recognise changing spatial relationships, interpret human gestures, predict the consequences of its own movements and adapt to unexpected events occurring in real time. Unlike written language, the physical world is uncertain, dynamic and only partially observable. Intelligence within such environments depends not merely upon knowing facts but upon maintaining an evolving internal model capable of anticipating what is likely to happen next.

Consider the apparently simple act of picking up a ceramic mug resting on a kitchen table. For a human being, this task is performed almost automatically, requiring little conscious reflection. Beneath this apparent simplicity, however, lies an extraordinarily complex sequence of cognitive operations. The brain estimates the weight of the mug from previous experience, identifies its orientation, calculates the trajectory of the arm, adjusts grip strength according to the material, compensates for gravity and continuously updates these calculations through visual and tactile feedback as the movement unfolds. None of these operations depends primarily upon linguistic reasoning. They rely instead upon deeply embedded predictive models acquired through years of interaction with the physical world.

Replicating this form of intelligence has proved to be one of the greatest challenges in robotics. For decades, engineers attempted to solve such problems through carefully programmed rules describing every possible situation a robot might encounter. While successful in highly controlled industrial environments, these systems struggled whenever confronted with the complexity of ordinary human surroundings. Homes, hospitals, warehouses, farms and public spaces contain countless variations that cannot realistically be anticipated through explicit programming alone. Consequently, robots require a more flexible form of intelligence capable of generalising from previous experience while adapting continuously to new circumstances. World Models offer one of the most promising pathways towards achieving this objective because they allow robots to develop predictive representations of their environment rather than relying exclusively upon predefined instructions.

The implications extend far beyond domestic assistance. Modern manufacturing increasingly demands collaborative robots capable of working safely alongside human workers without rigid physical separation. Autonomous agricultural machines must navigate changing terrain while recognising crops, obstacles and weather conditions. Search-and-rescue robots may operate within damaged buildings where every movement alters the surrounding environment. Surgical robots require extraordinary precision while interacting with living tissues that respond dynamically throughout medical procedures. In each of these examples, successful performance depends upon understanding the evolving state of the world rather than executing fixed sequences of programmed actions. Intelligence becomes inseparable from the ability to predict physical consequences before they occur.

Autonomous vehicles provide another compelling illustration of this principle. Driving safely requires far more than recognising traffic signs or identifying nearby objects. Human drivers constantly anticipate the intentions of pedestrians, estimate how other vehicles are likely to behave, infer potential hazards from incomplete information and prepare for situations that have not yet materialised. A child standing near the edge of a pavement, for example, may suddenly enter the road. An experienced driver begins adjusting speed before the child moves because an internal model of human behaviour suggests that such an event is possible. This form of anticipatory reasoning cannot be reduced to simple pattern recognition; it depends upon understanding causal relationships unfolding across space and time. World Models seek to provide autonomous systems with precisely this capacity for predictive situational awareness.

Perhaps even more significantly, the emergence of physical intelligence promises to redefine the relationship between robots and humans. Rather than functioning as highly specialised machines restricted to narrow industrial tasks, future robots may gradually evolve into adaptable assistants capable of learning new skills through observation, simulation and interaction. Instead of programming every movement manually, people may simply demonstrate a task, allowing the robot’s internal World Model to infer the underlying principles governing successful performance. Such systems would not merely imitate isolated actions but would understand why those actions achieve particular objectives, enabling them to adapt learned skills to entirely new situations.

From a historical perspective, this transition represents a decisive stage in the evolution of intelligent machines. The twentieth century largely focused on automating repetitive physical labour through precisely engineered mechanical systems. The first decades of the twenty-first century demonstrated that machines could also automate many forms of symbolic reasoning through language models. World Models now suggest the possibility of combining these achievements, allowing artificial intelligence to bridge the long-standing divide between cognitive reasoning and physical action. If this vision is realised, robotics may undergo a transformation comparable to that experienced by natural language processing following the emergence of transformer architectures. Machines will no longer merely execute programmed instructions or respond to verbal commands; they will increasingly understand the physical environments in which those instructions acquire practical meaning.

Simulation, Causality and Planning: Teaching Machines to Think Before They Act

Among the many capabilities that distinguish biological intelligence from conventional computer systems, one of the most remarkable is the ability to imagine possible futures before taking action. Human beings rarely respond to the world through immediate reaction alone. Instead, they continuously construct mental simulations of alternative scenarios, evaluating likely consequences before making decisions. A chess player mentally explores multiple sequences of moves before selecting the strongest strategy. An engineer visualises how a bridge will respond to changing loads long before construction begins. A doctor considers alternative diagnoses while anticipating the probable effects of different treatments. Even the most ordinary decisions, crossing a busy street, rearranging furniture or planning a journey, depend upon an internal capacity to predict future events before they unfold in reality.

This ability to simulate the future represents one of the central ambitions of World Models. Rather than simply recognising objects or interpreting language, these systems aim to create internal environments within which possible actions can be explored computationally before being executed in the physical world. Such internal simulation dramatically reduces the need for costly or dangerous trial-and-error learning because mistakes can be identified within the model itself rather than through direct interaction with reality. The machine effectively rehearses its decisions before acting, much as humans often imagine the consequences of different choices before committing themselves to a particular course of action.

The importance of this capability becomes particularly evident when considering environments where errors carry significant consequences. A warehouse robot may need to determine the safest route through a crowded storage facility without colliding with workers or equipment. An autonomous spacecraft cannot afford repeated experimental manoeuvres while navigating millions of kilometres from Earth. Medical decision-support systems must evaluate treatment strategies without exposing patients to unnecessary risk. In each of these situations, the capacity to simulate multiple possible futures internally offers enormous practical advantages. World Models seek to provide machines with precisely this predictive capability, allowing them to reason about actions before transforming those actions into reality.

Underlying this process is the concept of causality, one of the most fundamental principles in both science and human cognition. Causality concerns not merely recognising that two events frequently occur together, but understanding why one event produces another. For centuries, scientific progress has depended upon identifying causal mechanisms governing natural phenomena, from planetary motion and biological evolution to electricity and molecular chemistry. Human reasoning likewise relies heavily upon causal understanding. People naturally distinguish between coincidence and genuine cause, allowing them to generalise knowledge to unfamiliar situations rather than relying solely upon repeated statistical associations.

Artificial intelligence has historically found this distinction particularly challenging. Traditional machine learning excels at identifying correlations within enormous datasets, yet correlations alone do not necessarily reveal the underlying mechanisms responsible for observed behaviour. Two variables may appear strongly associated while sharing no direct causal relationship whatsoever. A World Model attempts to move beyond this limitation by learning how physical systems evolve over time, gradually uncovering the structural relationships connecting actions and consequences. Instead of simply recognising recurring patterns, it seeks to infer the dynamic processes generating those patterns. This richer representation supports more reliable prediction when circumstances change or when the system encounters situations absent from its training data.

Closely related to causality is the emergence of planning, another defining characteristic of intelligent behaviour. Planning requires the ability to organise sequences of actions extending beyond the immediate present while continually adapting those sequences as new information becomes available. Humans routinely perform such planning almost unconsciously, coordinating countless short-term and long-term objectives simultaneously. Preparing a meal involves anticipating the cooking times of different ingredients; organising scientific research requires projecting years into the future; managing a city demands planning infrastructure capable of serving generations yet to come. These activities depend upon maintaining internal models that connect present decisions with future outcomes across extended periods of time.

For autonomous artificial intelligence, planning represents one of the essential ingredients of genuine autonomy. A conversational system can respond intelligently to isolated questions without maintaining an extensive understanding of future objectives. An autonomous scientific laboratory, industrial robot or planetary exploration vehicle, by contrast, must continually evaluate long sequences of interconnected decisions while adjusting its strategy in response to changing environmental conditions. World Models provide the computational framework within which such planning becomes possible because they allow future states of the environment to be estimated before irreversible actions are taken.

This capability also explains why many researchers regard World Models as an important stepping stone towards increasingly autonomous AI agents. As artificial intelligence evolves beyond single conversational exchanges towards systems capable of completing complex multi-step tasks independently, prediction, causality and planning become progressively more important than language generation alone. The future AI assistant will not simply answer questions; it may coordinate research projects, manage industrial processes, supervise robotic teams or support emergency response operations. Achieving such capabilities requires intelligence grounded in an understanding of how actions transform the world over time rather than merely how sentences follow one another in written language.

Viewed historically, this transition from reactive computation to predictive simulation may prove as significant as the transition from rule-based programming to machine learning itself. Earlier generations of AI primarily learned to classify, recognise or generate information. World Models aspire to something considerably more ambitious: enabling machines to reason about the unfolding dynamics of reality, to anticipate the consequences of their own behaviour and to choose actions based upon internally simulated futures rather than immediate observation alone. In doing so, they bring artificial intelligence one step closer to one of the defining characteristics of biological cognition: the ability to think before acting.

Why World Models Could Transform Every Industry

The emergence of World Models is often discussed in connection with robotics, autonomous vehicles and advanced research laboratories, yet their long-term significance extends far beyond these highly visible applications. Throughout the history of technology, some innovations have transformed only particular sectors of the economy, while others have functioned as general-purpose technologies, reshaping virtually every field in which they were adopted. Electricity, the steam engine, digital computing and the Internet all belong to this latter category because they fundamentally altered the way societies produced knowledge, organised work and created economic value. Many researchers now believe that World Models possess the potential to become a similarly transformative technology, not because they perform one task exceptionally well, but because they introduce a new computational capability applicable across an extraordinary diversity of human activities.

The defining characteristic of a World Model is its capacity to construct predictive representations of dynamic environments. Any discipline that depends upon understanding how systems evolve over time may therefore benefit from this approach. Unlike traditional software, which generally follows predetermined rules, or language models, which primarily interpret symbolic information, World Models reason about processes, interactions and future states. They enable machines not merely to answer questions about the present but to explore alternative futures before decisions are implemented. This shift from static information processing towards predictive simulation has implications wherever uncertainty, planning and adaptation play central roles.

Scientific research offers one of the clearest examples of this transformation. Modern science increasingly depends upon computational simulation to investigate phenomena that cannot easily be observed directly. Climate scientists model atmospheric systems spanning entire continents. Biologists simulate protein folding and cellular interactions. Physicists explore the behaviour of subatomic particles through complex computational environments. Astronomers recreate the evolution of galaxies across billions of years. In each case, researchers seek to understand dynamic processes rather than isolated observations. World Models could significantly enhance these efforts by enabling AI systems to construct increasingly accurate representations of physical reality, generating hypotheses, identifying unexpected causal relationships and suggesting experiments that human researchers might never have considered. Rather than replacing scientific reasoning, such systems may become extraordinarily powerful partners in the process of discovery itself.

Industrial manufacturing represents another domain likely to experience profound change. Contemporary factories already employ sophisticated digital twins that simulate production lines before physical modifications are introduced. World Models could extend these capabilities dramatically by allowing AI systems to anticipate equipment failures, optimise logistics, redesign manufacturing processes and coordinate autonomous robots operating within constantly changing environments. Instead of reacting to problems after they occur, intelligent factories may increasingly predict disruptions before they emerge, continuously adapting production schedules and resource allocation in response to evolving conditions. Such predictive intelligence has the potential to improve efficiency while simultaneously reducing waste, energy consumption and operational risk.

Healthcare provides equally compelling opportunities. Medical diagnosis rarely depends upon a single observation but upon understanding how biological systems evolve over time. Diseases develop through complex causal processes involving genetics, physiology, environmental influences and individual behaviour. World Models may eventually assist physicians by simulating the probable progression of illnesses under different treatment strategies, integrating medical imaging, laboratory data, patient histories and physiological monitoring into coherent predictive representations. Rather than simply recognising patterns associated with existing diagnoses, these systems could help anticipate future developments, supporting earlier intervention and more personalised medical care. Importantly, such technologies would augment rather than replace clinical expertise, providing healthcare professionals with richer predictive insights while preserving human judgement at the centre of medical decision-making.

Urban planning and public infrastructure constitute another area in which predictive world modelling may prove particularly valuable. Modern cities function as extraordinarily complex adaptive systems in which transportation, energy networks, housing, environmental conditions and human behaviour interact continuously. Decisions concerning infrastructure investment often produce consequences extending decades into the future. World Models could enable planners to simulate alternative development strategies before implementation, exploring how new transport systems influence economic activity, how climate adaptation measures affect urban resilience or how demographic changes reshape demand for public services. Such capabilities align closely with the broader movement towards digital twins of cities and regions, where computational models increasingly support evidence-based governance and long-term strategic planning.

Education itself may also evolve through the integration of World Models. Traditional educational technologies have largely focused on delivering information more efficiently or adapting learning materials to individual students. Predictive world modelling opens the possibility of immersive learning environments in which students explore dynamic simulations rather than passively receiving explanations. Future educational systems may allow learners to experiment with ecological systems, historical events, engineering designs or biological processes within richly interactive virtual worlds that respond realistically to their decisions. Learning would become increasingly experiential, encouraging curiosity, experimentation and systems thinking rather than memorisation alone.

Perhaps the most remarkable aspect of these diverse applications is that they all depend upon the same underlying computational principle: understanding how reality changes through time. Whether modelling molecular interactions, managing autonomous factories, assisting surgeons or designing resilient cities, World Models provide a common framework for reasoning about dynamic systems whose future behaviour cannot be captured through static databases or textual descriptions alone. This remarkable versatility explains why many researchers consider them one of the most important technological frontiers currently emerging within artificial intelligence.

From a historical perspective, the significance of World Models may ultimately resemble that of the Internet during the 1990s or cloud computing during the early twenty-first century. Initially, their applications may appear concentrated within specialised industries requiring advanced technical expertise. Gradually, however, the underlying capability of predictive world representation could become embedded within countless digital services, often without users even recognising the sophisticated computational models operating beneath the surface. In this sense, the true impact of World Models may not lie in any single application but in the quiet transformation of how intelligent systems interact with reality across virtually every domain of human activity.

Towards Artificial General Intelligence?

No discussion of World Models would be complete without addressing one of the most frequently debated questions in contemporary artificial intelligence: do they bring us closer to Artificial General Intelligence (AGI)? Although the term AGI remains the subject of considerable disagreement among researchers, it is commonly used to describe hypothetical systems capable of performing the broad range of cognitive tasks that humans accomplish with flexibility, adaptability and minimal task-specific training. Unlike specialised AI designed for particular applications, an AGI would be expected to learn continuously, transfer knowledge between domains, solve unfamiliar problems and operate successfully within environments it had never previously encountered.

Large Language Models have already demonstrated that many intellectual activities previously regarded as uniquely human can emerge through large-scale statistical learning. Their ability to reason across disciplines, generate software, interpret scientific literature and engage in sophisticated dialogue has significantly expanded expectations concerning machine intelligence. Yet many leading researchers have cautioned that linguistic competence alone should not be mistaken for general intelligence. Human cognition depends upon a far richer integration of perception, physical interaction, memory, planning, social understanding, emotional regulation and causal reasoning than current conversational systems possess. Consequently, while language models may represent an essential component of future AGI architectures, they are unlikely to constitute the entire solution.

It is within this context that World Models have attracted such intense scientific interest. By providing machines with internal representations of the physical world, these systems address one of the most significant limitations of language-based intelligence: the absence of grounded understanding. Human beings do not merely manipulate symbols; they connect language to lived experience accumulated through continuous interaction with their environment. Words acquire meaning because they refer to objects, actions, places, relationships and events that individuals have directly perceived or experienced. World Models attempt to establish a computational equivalent of this grounding by allowing artificial intelligence to learn from observation, movement and physical interaction rather than from text alone.

Some researchers therefore regard World Models as a critical missing ingredient on the path towards more general forms of machine intelligence. If intelligence fundamentally depends upon constructing predictive representations of reality, then systems capable of integrating language, perception, action and simulation may eventually display forms of flexibility unavailable to architectures based solely upon text prediction. Such machines would not simply answer questions about the world; they would possess internal computational structures enabling them to anticipate change, adapt to unfamiliar environments and plan coherent sequences of action across extended periods of time. These capabilities closely resemble several characteristics traditionally associated with general intelligence.

Nevertheless, it would be misleading to conclude that World Models alone will inevitably produce AGI. Human cognition encompasses many dimensions that remain only partially understood by contemporary science. Social reasoning, abstract conceptual thought, creativity, emotional intelligence, ethical judgement, long-term motivation and consciousness continue to present profound theoretical and practical challenges. Building accurate models of the physical environment addresses one important aspect of intelligence, but it does not automatically solve every problem associated with creating machines capable of matching the extraordinary versatility of the human mind. As with previous milestones in AI research, World Models should therefore be viewed as an important step rather than a definitive destination.

Indeed, many of the scientists leading this field have consistently emphasised the need for humility regarding future predictions. Yann LeCun has frequently argued that current AI systems remain far from possessing the common-sense understanding displayed even by young children. Fei-Fei Li has repeatedly stressed that intelligence emerges through the interaction of perception, embodiment and experience rather than through isolated technical advances. These perspectives suggest that the road towards more general artificial intelligence is likely to involve the gradual integration of multiple complementary capabilities rather than a single revolutionary breakthrough. Language models, World Models, memory systems, planning architectures, autonomous agents and multimodal perception may ultimately converge into increasingly sophisticated cognitive ecosystems rather than remaining separate research directions.

This possibility points towards one of the most intriguing developments currently emerging within artificial intelligence. Instead of replacing Large Language Models, World Models are increasingly being viewed as complementary components within future AI architectures. A conversational model provides extraordinary linguistic competence; a World Model contributes physical intuition and predictive simulation; long-term memory preserves accumulated experience; planning systems coordinate complex objectives; specialised reasoning modules support scientific or mathematical analysis. Together, these elements may eventually produce systems whose overall capabilities exceed those of any individual component considered in isolation. Intelligence, in this view, becomes an emergent property arising from the interaction of multiple cognitive subsystems rather than from the scaling of a single architecture.

Whether such systems ultimately deserve the label of Artificial General Intelligence remains a question for future generations of researchers. Definitions will undoubtedly evolve as technological capabilities continue to expand. Yet regardless of terminology, the emergence of World Models clearly marks a significant conceptual transition within AI research. The field is progressively moving beyond machines that excel primarily at interpreting language towards systems that aspire to understand the broader reality within which language, action and experience are inseparably connected. In that sense, World Models represent not simply another technical innovation but one of the most important milestones in the long intellectual journey towards more comprehensive forms of machine intelligence.

The Beginning of the Era of World Understanding

Looking back at the history of artificial intelligence, it becomes apparent that the field has advanced through a succession of conceptual revolutions rather than through a single continuous technological progression. The earliest decades focused on symbolic reasoning, attempting to reproduce intelligence through explicit logical rules and carefully constructed knowledge bases. The rise of machine learning shifted attention towards systems capable of discovering statistical regularities directly from data. Deep learning subsequently demonstrated that sufficiently large neural networks could achieve extraordinary performance across perception, language and pattern recognition. Finally, the emergence of Large Language Models transformed public understanding of artificial intelligence by revealing that machines could communicate, reason and generate knowledge with a fluency that had previously seemed unattainable. Each of these transitions expanded the boundaries of what researchers believed artificial systems might ultimately achieve.

The emergence of World Models now signals the beginning of another transformation whose significance may prove equally profound. Rather than concentrating exclusively on language, future AI systems increasingly seek to construct internal representations of the physical world itself. This shift reflects a recognition that intelligence is not simply the ability to manipulate symbols but the capacity to understand how reality evolves through time, how actions generate consequences and how future events can be anticipated before they occur. Language remains an indispensable component of intelligent behaviour, yet it is gradually becoming one element within a much richer cognitive architecture that integrates perception, prediction, planning and interaction with the surrounding environment.

Perhaps the most important contribution of World Models is that they redefine the objective of artificial intelligence. During the first generation of conversational AI, success was measured largely by the quality of linguistic interaction. Machines were expected to answer questions, generate coherent text, summarise information and assist with intellectual tasks traditionally associated with written communication. World Models expand this ambition considerably. Their purpose is not merely to help machines describe reality but to enable them to understand the structure of reality sufficiently well to anticipate its future evolution. This apparently subtle distinction has far-reaching implications because prediction, planning and causal reasoning underpin almost every form of intelligent behaviour observed in biological systems.

The consequences of this transition extend across virtually every domain in which artificial intelligence is expected to play an increasingly important role. Robots operating alongside human beings require intuitive physical understanding rather than language alone. Autonomous vehicles must continuously predict the behaviour of dynamic environments. Scientific research increasingly depends upon sophisticated computational simulations capable of exploring hypotheses before physical experiments are undertaken. Healthcare, engineering, industrial manufacturing, environmental management and public governance all involve decisions whose quality depends upon anticipating the future rather than merely analysing the present. World Models therefore provide a common computational framework capable of supporting these diverse applications through a shared emphasis on predictive understanding.

Equally significant is the interdisciplinary character of this new research direction. The development of World Models has encouraged an unprecedented convergence between artificial intelligence, neuroscience, developmental psychology, cognitive science, robotics and computational physics. Rather than viewing intelligence solely as a problem of statistical optimisation, researchers increasingly examine how biological organisms learn through perception, movement and interaction with complex environments. This dialogue between disciplines is enriching both fields simultaneously. Artificial intelligence gains new theoretical inspiration from the study of human cognition, while computational modelling offers neuroscientists increasingly sophisticated tools for exploring the mechanisms underlying learning, prediction and perception. Such reciprocal exchange may ultimately prove as important as any individual technological breakthrough.

At the same time, it is important to recognise that the emergence of World Models does not imply that the fundamental challenges of artificial intelligence have been solved. Machines remain far from possessing the flexibility, common sense and adaptive creativity that characterise even young children. Many aspects of human cognition, including consciousness, emotional understanding, ethical judgement, social reasoning and long-term autonomous learning, continue to present profound scientific questions whose solutions remain uncertain. World Models should therefore be understood not as the culmination of AI research but as one important stage within a much longer process of exploration. They provide new capabilities while simultaneously revealing further questions concerning the nature of intelligence itself.

For historians of technology, however, their significance is already becoming increasingly apparent. Future generations may look back upon the period following the success of ChatGPT as the moment when the ambitions of artificial intelligence fundamentally changed. The objective was no longer simply to build machines capable of producing convincing language but to create systems capable of constructing internal models of reality, reasoning about physical processes and interacting intelligently with the world that surrounds them. In retrospect, this transition may prove comparable to the emergence of transformer architectures or deep learning itself: a conceptual shift that redirected an entire field towards new scientific horizons.

From the IDHUS perspective, the rise of World Models marks one of the defining milestones in the evolution of modern AI. It represents the point at which researchers increasingly recognised that genuine intelligence cannot be separated from an understanding of the environment in which that intelligence operates. Words, images and symbols remain indispensable, but they acquire their deepest meaning only because they refer to an underlying physical reality governed by space, time, causality and interaction. By seeking to model that reality directly, World Models open the possibility of artificial systems capable not merely of explaining the world but of participating within it with increasing competence and autonomy.

Whether this research direction ultimately leads to Artificial General Intelligence remains impossible to predict. History repeatedly reminds us that scientific progress rarely follows the linear trajectories imagined by contemporary observers. Unexpected discoveries, new theoretical insights and entirely unforeseen technological innovations will undoubtedly continue to reshape the landscape of artificial intelligence over the coming decades. Nevertheless, one conclusion already appears difficult to dispute. The era inaugurated by ChatGPT demonstrated that machines could learn to understand language. The era now beginning with World Models seeks to teach machines something far more ambitious: how to understand the world itself. That aspiration may well define the next great chapter in the history of artificial intelligence.