Su Hao's WAIC Keynote: The Mission of Physical AI is to Return Humanity to Itself

Deep News
Jul 17

At the main forum of the 2026 World Artificial Intelligence Conference (WAIC) and the High-Level Meeting on Global AI Governance held this afternoon, Su Hao, Haoqing Distinguished Professor at Fudan University and the inaugural Dean of the Fudan University Institute for General Physical AI, delivered a keynote speech titled "Physical AI: From Illusion to Reality." The following is the full text of his address.

Physical AI: From Illusion to Reality

Distinguished leaders and guests, good afternoon. It is an honor to share some personal perspectives at the WAIC. The theme of this forum is "Cornerstones," and what I wish to discuss today is precisely the most fundamental cornerstone of intelligence—physical AI.

Let me state my conclusion upfront: I am a firm optimist regarding physical AI. Seeing the illusion clearly is to hasten the arrival of reality—hence the title, "From Illusion to Reality."

From Large Models to Physical AI

The progress of large models is evident, but why do even the most intelligent models suffer from hallucinations? I believe the most fundamental reason is that language is merely a projection of the world. Humans first experience the physical world before compressing that experience into language; models have only ever learned this shadow, never the entity that casts it. A model can fluently describe "a cup will break if dropped on the floor," but it has never had the chance to feel the weight of a cup. Knowledge without an anchor in reality means the model has no way of knowing when it is wrong. You can argue eloquently on a screen, but you cannot persuade gravity.

The hallucinations of large models largely stem from their lack of a "body." To move beyond illusion, one must step beyond the boundaries of the "digital world," to experience firsthand and submit oneself to the judgment of the "physical world"—making predictions, taking actions, and being corrected by reality. This ancient process is called experimentation, and it is also the key to physical AI.

The Core Deficiency of Physical AI: Aggregating Scattered Knowledge

So, what is truly lacking for physical AI?

What we need is a model capable of aggregating all of humanity's knowledge about the physical world. From the source of cognition, physical knowledge has at least six foundational layers.

The lower three layers belong to the objective world—they exist whether there is an "I" or not.

The first is knowledge about objects: the world is composed of independent, persistent objects—a ball rolls under the sofa, out of sight, but it is still there. Recognizing "what exists" comes first.

The second is knowledge about states: what is the condition of these things at this moment? The ball is under the sofa; its weight, hardness, or softness are at this layer.

The third is knowledge about dynamics: how does the world change on its own? Balls roll, water flows, released objects fall. "Force" resides at this layer—you cannot see it with your eyes.

The upper three layers arise because of the "I"—each layer is the lower three bound to a "subject."

The fourth is knowledge about function: what is this thing's use to me? A handle is for gripping, a cup is for holding. The same chair is "sittable" for a human, but not for an ant. Objects become "tools" because of a subject.

The fifth is knowledge about goals: what state do I want the world to be in? The ball is under the sofa; I want it back in my hand. Where the ball is, is a state; where I want it to be, is a goal. The difference between "is" and "ought to be" for a target state encompasses the entire subject.

The sixth is knowledge about behavior: knowing what to do and being able to do it with one's hands are two different things—carrying a cup of water across a room without spilling requires finesse; tying shoelaces or using chopsticks requires dexterity. These skills are not written in books; they are learned through physical action.

From objects to behavior, the higher one goes, the less it is learned by "seeing," and the more it must be "done" firsthand in the physical world. Every infant's first two years are spent climbing this ladder step by step—what Piaget called the "sensorimotor stage": the foundation of human intelligence is built with our hands.

But models have no such childhood—they learn primarily from records left by humans. The trouble is that these six layers of knowledge are scattered across disconnected carriers, with less and less recorded the higher one goes. Internet videos are the most abundant but largely remain at the lower end of the ladder—appearance and motion are visible, but force and tactile sensation are out of reach. Textbook equations are the most precise, describing dynamics, but only for an idealized world. Real machine data contains force feedback and operational demonstrations, reaching the upper layers, but it is scarce.

Language models are fortunate: the internet has already aggregated linguistic knowledge for them. Physical knowledge lacks this fortune. As Polanyi said, we know more than we can tell. Therefore, the upper half of the ladder has not yet been systematically recorded—you cannot develop complete physical AI just by browsing webpages. Aggregation is the scientific engineering task for our generation: to fuse the breadth of video, the precision of equations, the authenticity of real machine data, and the nuance of intuition into a single model, calibrating and complementing each other to complete the ladder. The gap between illusion and reality for physical AI is precisely this ladder.

A Critical Need of Our Era and How It Will Arrive

Why must we climb this ladder?

The answer is not in the corpus or the server room, but in the real world.

Today's AI can write poetry, code, and create presentations—but it cannot help an elderly person turn over in bed. Most of the value of intelligence remains in the digital "bit" world, yet humanity's most pressing needs are in the physical "atom" world. Look around: an aging population is a demographic fact. The number of people needing care is increasing, while the number of people available to provide care is decreasing. On the other hand, dangerous and labor-intensive tasks in high-altitude, underground, or high-temperature environments also face labor shortages. The demand exists; what is lacking is manpower. Physical AI is not a replacement for humans but a collaborator—filling the manpower gap, handing over physical tasks like turning patients or moving objects to machines, returning caregivers' time to companionship and care; handing over dangerous operations to machines, allowing humans to step back behind the safety line to take on judgment and creation. The mission of physical AI is to return humanity to itself.

How will it arrive? As a technical professional, here is my personal assessment. It will not be "achieved" overnight at a product launch; it will be more like the electrification of the past—first lighting up factories and warehouses, then entering shops and hospitals, and finally reaching millions of households. The value of physical AI will be released along the way, not just at the destination.

The most perilous stretch on this path is the reliability gap between a demonstration and a product. Bridging it relies not on hype, but on foundational work and data accumulation—especially the half involving force and interaction. Bridging it relies on industry standards, supply chains, and the hardest thing to accumulate: societal trust. Trust can only be earned bit by bit through reliability; and this trust is also the most fundamental cornerstone for the implementation of physical AI.

Three Predictions

Finally, I leave three predictions for future verification.

First, the breakthrough for physical AI lies not in model architecture, but in knowledge aggregation. Aggregation is inherently a task that transcends the boundaries of any single institution—videos are on the internet, equations are in textbooks, force data is in various labs, and operational intuition resides in billions of workers. No one can gather this ladder alone; it requires collaboration across the entire industry and society: co-building data, co-establishing standards, and sharing infrastructure like simulation and evaluation. When this knowledge is truly fused into a single model, the physical world will have its own "internet moment"—the so-called GPT moment would merely be a byproduct.

Second, the industry's focus will shift from "how stunning the demo is" to "how reliable the operation is." In engineering, there is a concept called "the number of nines": from 99% to 99.9%, each additional "9" increases the difficulty exponentially; the gap between a demo and a product lies precisely in those last few "nines." Generality is the destination, but reliability is the starting point—teams willing to do the hard work on those "nines" will go the farthest.

Third, physical AI will transform AI from a "reader" of science into a "creator" of knowledge. Today's AI has read almost all human papers but has hardly ever conducted an experiment with its own hands; yet new knowledge is born precisely from experimentation. When AI possesses hands that can perceive, operate, and verify the real world, it will be able to propose hypotheses, conduct experiments, and make corrections around the clock. The discovery of new materials and drugs could accelerate by several orders of magnitude as a result.

There are no shortcuts from illusion to reality. It relies not on louder narratives, but on reverence for the physical world and the diligent effort of climbing that ladder step by step. The physical world is intelligence's oldest teacher, its most honest examiner, and its ultimate cornerstone. We choose to submit our answers to it.

Thank you.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10