The AI Race Enters the Physical World

And with it a new era in US-China competition begins.
September 22, 2026

Listen to this article

Audio file

Audio is generated automatically and may contain minor inaccuracies.

When German Chancellor Friedrich Merz visited China in February, Chinese President Xi Jinping greeted him with a demonstration of humanoid robots. Dancing, sparring, and moving in perfect synchrony, the platoon of mechanical beings offered more than a technological showcase. The display reflected what a concerted industrial and political system can build and deploy at scale.

The carefully scripted performance by itself says little about the operational maturity of China's robotics sector. More significant was what it intimated: multiple AI systems operating together in the physical world, backed by the manufacturing capacity required to scale them. That is where AI is heading.

Two Trajectories

The field is developing along different paths. The first extends the current course of large language models trained on ever-larger troves of text, code, and images, with progress driven by compute and data. The second concerns systems that can model the physical world and act within it, learning from sensory input to predict, plan, and adjust. The first path has produced a recent wave of investment. The second is less mature and addresses the different problem of acting in the world rather than reasoning about it. Which path dominates, and how far the two converge, will contribute to shaping the next phase of strategic competition.

The United States leads on the first path, but China is ahead in building the substrate for the second. Whichever side can integrate both and translate them into deployed capability at scale will gain a structural industrial and military advantage in a contest in which strategic primacy and technological supremacy are becoming one and the same.

The current AI cycle has been driven by large language models (LLMs) drawing their power from scale. This has delivered rapid progress in coding, search, writing, classification, and decision support, and has concentrated advantages in the parts of the tech stack where the United States leads, even as China catches up. These include frontier model development, advanced semiconductor design, hyperscale cloud infrastructure, and private capital. Stanford's 2026 AI Index shows US-based institutions producing 50 notable AI models in 2025 compared to China's 30, a narrower gap than the 40-to-15 ratio one year earlier. In investment, the asymmetry is even starker. US private AI funding reached $285.9 billion in 2025, roughly 23 times China's $12.4 billion, with generative AI capturing nearly half of that total and growing more than 200% year on year.

The structural limitation of the LLM paradigm is that it is based on models that operate on representations of the world rather than within it. They parse instructions, generate plans, and infer patterns across symbolic data, and benchmark results register the gains. But real-world robustness diverges from benchmark performance because reasoning about the world is not the same as operating in it.

A language model can describe how to assemble a machine, generate the instructions, and even optimize the workflow on paper. It cannot learn from the act of assembly, friction, error, or variation in the environment. That is the gap world models, the second trajectory, set out to close. They learn by acting, adjusting behavior as conditions change, taking feedback from sensors, and refining their internal representation over time. This gives them access to causal signals of success or failure, stability or breakdown, that static datasets do not provide.

Technical literature converges on a working definition of world models. They are internal representations of environmental dynamics that allow an agent to perceive, predict, and act. Fei-Fei Li's World Labs uses a different label, "spatial intelligence", for essentially the same objective: building frontier models that can perceive, generate, and interact with three-dimensional environments.

Capital is already moving in this direction. Yann LeCun, formerly Meta's chief AI scientist, raised $1.03 billion at launch for his new venture AMI, describing its mission as building world-model-based AI "that understands the real world", with applications in industry, robotics, healthcare, and automation. World Labs announced a separate $1 billion round in February to advance spatial intelligence and build world models for robotics and scientific discovery. These are large commitments to a still-emerging technical direction, and they signal that world models and “embodied AI” are now treated as strategically consequential.

To be clear, none of this makes, or will make, LLMs obsolete. The frontier is moving toward hybrid systems that combine language models with perception, planning, and control. LLMs handle semantic reasoning, task decomposition, and high-level planning, whereas world models support control and adaptation in dynamic settings. The question is how the balance between symbolic processing and physical substrate evolves, and who controls the inputs each requires.

The Stack Beneath the Models

Most of the debate around AI competition still focuses on model capabilities and compute investment. As the field pushes beyond language into embodied applications, the new emerging question is: Who controls the respective substrates?

For the inputs LLMs demand, the United States holds a commanding position. NVIDIA controlled an estimated 80% to 90% of the AI GPU segment in 2025, a dominance built on hardware and the software ecosystem that has made its chips the default for frontier model training across the industry. The US position is also being strengthened. The OpenAI–Oracle–SoftBank Stargate venture is building toward 10 gigawatts of data center capacity, and TSMC is expanding its $165 billion investment in Arizona to anchor the compute stack on American soil. For this generation of AI, chips are the crucial asset for controlling the speed of progress.

World models demand a different order of input: streams of sensory data generated by physical systems working in real environments. This kind of data accumulates only through deployment. Every movement, adjustment, error, and correction feeds back into learning. In this domain, China has built a structural advantage that is, for now, still widening. According to the International Federation of Robotics, the country installed a record 295,000 industrial robots in 2024, 54% of global deployments, and nearly nine times as many as the United States. The gap in the installed base number is starker still. Fully 2,027,200 factory robots operate in China, roughly five times the 393,700 in the United States. Chinese firms such as Unitree are commercializing humanoid platforms at prices that no Western equivalent can match.

This does not mean the terrain is uncontested. Software for training, sensing, and control remains dominated by firms based outside China, most importantly NVIDIA, whose ecosystem underpins nearly every major robotics company in the country. High-precision mechanical components are another dependency. Bosch Rexroth, Schaeffler, THK, and NSK control roughly 90% of the high-end ball screw and precision motion market on which Chinese manufacturers rely. At the level of autonomous function, the gap is wider still. Chinese humanoids are mostly deployed in narrow tasks and site-specific trials, a far cry from full autonomy in unpredictable environments.

Another layer that cuts across the two trajectories is critical minerals and energy. Rare earths are embedded in AI infrastructure, from the permanent magnets in data center cooling systems to the actuators that give robots their range of motion. China is the dominant refiner of 19 of the 20 minerals in the International Energy Agency's Global Critical Minerals Outlook 2025, with an average market share around 70%. In permanent magnets the concentration is near total. China's share of sintered magnet production has climbed from roughly 50% two decades ago to 94% today. In its recent trade confrontation with Washington, Beijing has already demonstrated its willingness to use this dominance as a strategic lever.

Regarding energy, electricity demand from AI data centers surged 50% in 2025 and is on track to triple by 2030. The United States and China are scaling rapidly, but the constraints they face differ structurally. In the former, PJM Interconnection, the largest US grid operator, serving more than 65 million people, announced in December 2025 that its capacity auction for the 2027-2028 delivery year had, for the first time since the market was created in 2007, failed to procure enough capacity to meet its projected reliability requirement, leaving a shortfall of approximately 6,623 megawatts. This reflects mounting demand pressures across the PJM system: data centers accounted for approximately 94% of its projected peak-load growth between 2024 and 2030.

China faces bottlenecks of its own: grid congestion, renewable curtailment, and continued dependence on coal. But critically, it can direct its energy system politically, accelerating investment in renewables, nuclear power, transmission, and generation capacity in ways that a fragmented American grid cannot easily match.

Since the two problems of critical minerals and energy converge, the competition in each sector overlaps. Gallium, essential to advanced semiconductors and power electronics in AI systems, is expected to see demand rise above 10% of current global supply by 2030. China refines 99% of the element, reflecting the country’s better position for now.

Industrial Implications

The industrial consequence of embodied AI matters less for the automation of existing work than for the location in the global economy of created and captured value. The first wave of AI productivity, in coding assistants, document generation, and customer service, captured value for software-intensive industries and concentrated returns in firms that already operated digitally. The next wave promises to embed intelligence into factories, warehouses, hospitals, construction sites, and energy infrastructure. In other words, advantage will shift from those who build the best models to those who can deploy them at scale in physical environments and harvest the operational data they generate.

China seems to hold an advantage here. Its firms accounted for roughly 90% of global humanoid robot shipments in 2025, led by AgiBot (more than 5,000 units), Unitree (more than 4,200), and UBTech. Over 150 humanoid companies now operate in the country. A manufacturing base of that scale is also a deployment environment for embodied AI and can generate the sensor data those same systems require to improve. Every robot installed in a Chinese factory is a node in a feedback loop that has no comparable equivalent anywhere else.

The political architecture to consolidate this advantage is now also in place. China's 15th Five-Year Plan (2026–2030), released earlier this year, elevates humanoid robotics and embodied intelligence to strategic priorities, making them crucial parts of a general economic framework. In December 2025, the Ministry of Industry and Information Technology established a Humanoid Robot and Embodied Intelligence Standardization Technical Committee, which in March issued the first national standard system covering the humanoid lifecycle. Beijing is also leading the International Electrotechnical Commission standard-setting for elder-care robots and shaping international norms for robot safety, interoperability, and data governance.

This is where China shows the inherent advantage stemming from centralized political control of the economy. The Chinese Communist Party is constructing a vertically integrated system in which model development, hardware, deployment, and standard setting reinforce one another. Western firms operating in any single layer of that stack, be it chips, software, components, or robots, cannot count on the same level of long-term, top-down planning and coordination.

Of course, infrastructure and doctrine are not the same thing. A country can accumulate robots, data, and industrial capacity and still fail to translate them into operational advantage if regulation, organizational culture, or bureaucratic incentives do not adapt in step. Beijing's ability to move fast does not solve problems of jointness, logistics, and political reliability, and industrial iteration does not automatically translate into effectiveness. Institutional adaptation, just as much as industrial capability, will decide much of how this advantage plays out in practice.

For the United States, the binding constraint is deployment. Producing more frontier models matters less if the operational substrate on which embodied systems are tested, refined, and scaled is not at hand. Figure, 1X, Physical Intelligence, and Waymo have made meaningful progress in building that substrate domestically, but their efforts remain in early stages, and the American deployment base is nowhere near the scale its frontier-model position would seem to require.

Europe’s exposure is different still. Its precision-component suppliers remain indispensable inputs to Chinese robotics, and that is precisely what makes those companies targets of Beijing's stated localization objectives. The 15th Five-Year Plan's explicit identification of "high-precision ball screws" and "high-parameter transmission devices" as priorities for domestic breakthrough is, in effect, a notice to compete for a layer of the tech stack that policymakers on both sides of the Atlantic have yet to take seriously.

Military Implications

Hard power is the other domain where embodied AI promises to tilt the balance. In defense terms, systems that can act, perceive, and adapt within physical environments at scale amount, in effect, to the dehumanization of war. Autonomous systems capable of operating in contested space, air, sea, undersea, and ground, with progressively less human cognitive load required to employ them, are already at the center of a revolution whose implications are still difficult to grasp, a process that Ukraine has visibly accelerated. Beyond the obvious immediate tactical returns, the effects will reach into the building blocks of strategic posture: military doctrine, force structure, and deterrence.

The Pentagon has recognized the shift. The Replicator initiative, launched in August 2023 under then-Deputy Secretary of Defense Kathleen Hicks, sought to field thousands of expendable autonomous systems by August 2025, explicitly framed as a response to the People’s Liberation Army's (PLA) quantitative advantage. The results have been uneven. By the deadline, the program had reportedly fielded only hundreds rather than thousands, encountered persistent software integration problems, and was eventually restructured under a new Defense Autonomous Warfare Group within Special Operations Command. The Trump administration's fiscal 2027 budget request nevertheless signals that the underlying logic is intact: $53.6 billion for autonomy, drone platforms, and contested logistics, the largest single investment in drone warfare in American history, and a more than 24,000% increase in funding for the new command over the previous year. General David Petraeus (ret.) recently warned, however, that money and platforms count for little without doctrine, training, and organizational structures to absorb them. The warning applies as much to Beijing as to Washington.

A key question, then, is whether the United States can harness the data, hardware, and component base it currently controls. Autonomous systems depend on the inputs that define the embodied AI race: sensor data from deployment, precision components from concentrated supply chains, and rare-earth-dependent actuators and motors. The dual-use character of Chinese civilian robotics amplifies the asymmetry. Every humanoid platform working in a Chinese warehouse, every quadruped tested in commercial trials, and every autonomous vehicle logging miles in a Chinese city contributes to a data substrate the PLA can draw on, directly or indirectly. The same Unitree quadrupeds that performed for Chancellor Merz have already been documented in PLA exercises, in armed configurations and as logistical platforms. The line between commercial and military embodied AI in the Chinese system is porous by design.

The impact of that is beginning to surface ever more sharply. It comprises drone swarms operating under loose human supervision, persistent surveillance grids composed of low-cost autonomous platforms, logistics moved by uncrewed vessels and aircraft, and area-denial architectures that fuse sensors, effectors, and decision-support systems into integrated kill webs. Ukraine has demonstrated, at considerable cost, that inexpensive autonomous systems can shift the balance of force against a larger and better-equipped adversary. The recent Iran war showed that cheap unmanned systems can saturate air defenses, threaten critical infrastructure, and impose disproportionate costs, making counter-drone capability a structural component of modern defense. Both theaters suggest that embodied AI is beginning to shift the cost-exchange ratio between offense and defense, and mass and precision, in ways that legal, tactical, and strategic frameworks have not yet caught up with.

Europe, for its part, retains defense industries with world-class capabilities in precision components, sensors, and integration, exactly the inputs embodied military AI requires. But it has neither the model-development scale of the United States nor the deployment substrate of China, and it remains dependent on transatlantic security architecture in domains where the underlying technology is shifting fastest. The current push toward European defense autonomy will have to confront the fact that capability in embodied AI cannot be imported because much of it is generated through deployment.

The Next Phase of AI Competition

The implications of all this reach well beyond technology policy. As AI moves into physical environments, the perimeter of national security widens. Robotics, sensors, advanced materials, energy infrastructure, logistics networks, and manufacturing capacity are now inside the core architecture of the competition.

The "AI bubble" question fits here, though not quite in the form usually posed. Current valuations reflect high expectations about AI's economic impact and the ongoing strategic contest between Washington and Beijing. The relevant issue is less about expectations that are correct in aggregate and more about the allocation of capital to the right layer of the tech stack. If investors stay focused on language models while the strategic center of gravity shifts toward embodied ones, an adjustment inside the AI sector will occur rather than the popping of a general AI bubble.

The show put on for the German chancellor was a staged act, but the system behind it is not. Coordinated machines are being produced by a manufacturing base built to scale them and deployed within a political system that treats their proliferation as a strategic objective. That chain connects, with little mediation, to industrial capacity, military capability, and the structure of economic power.

The implication for policymakers is that AI strategy can no longer be reduced to models, chips, and compute. It also depends on the capacity to deploy, integrate, and iterate systems in real environments, military applications included. Deployment substrate, industrial integration, and the conversion of technological progress into operational capability are shared problems, particularly among Western democracies, and these challenges will be addressed together or not at all. Meeting them will require the kind of long-horizon coordination that Beijing's system produces by design and open ones do not. In fact, the deeper dilemma reaches beyond AI, as democracies attempt to build a new compact between public and private interest, one that keeps them competitive without eroding the freedoms they exist to protect.

The views expressed herein are those solely of the author(s). GMF as an institution does not take positions.