Skip to content

The Generalization Gap: Why World-Action Models Are Embodied AI's Real Bottleneck

#embodied-ai #world-action-models #robot-learning #humanoid-robots #data-collection

The gap between demo and deployment ​

The second World Humanoid Robot Games wrapped on August 26, and the headline numbers are wild. Tiangong Ultra won the 100-meter final in 8.64 seconds, beating Usain Bolt's 9.58-second human record. Last year the same robot ran 21.50 seconds. In twelve months, the field cut the time by more than half.

Then the robots hit the crash pads. Several sprinted straight into the barriers at the finish line. One was carried off the track on a stretcher. Watch the footage and you'll see the same thing I did: the robots don't stop, they collide. In the 400-meter small-class race, Tiangong Omni crossed the line with its arms raised beside its face, a "face-covering sprint" that its own engineers say emerged from simulation learning, not design. The robot discovered it felt more stable that way. 每经头条 (National Business Daily) reported that this year's track events all required fully autonomous operation, no remote control. The robots had to locate themselves, plan routes, and convert those plans into joint actions on their own.

The same week, 铅笔道 (Pencil) published an investigation that put a number on the industry's dirty secret. If you define a "real commercial order" as one where the customer bought on production need, the robot was delivered and accepted, and it works without engineers standing by, then at least 80% of the humanoid robot orders you've read about are fake. The reporter called that estimate conservative.

These two stories are the same story. Robots can do spectacular things once. They cannot yet do ordinary things reliably, every day, in environments that change. The technical term is generalization, and it's the wall the entire field is running into.

In-context learning comes to manipulation ​

Zero-WAM, a new paper on arXiv, attacks the hardest version of this problem: zero-shot cross-task generalization, where a policy must execute manipulation tasks it never saw during training. The authors borrow the trick that made LLMs work. In a language model, you don't retrain for a new task. You specify the task in the prompt, and in-context learning does the rest. Generalization becomes a problem of task specification.

The question is what the task specification should be for a robot hand. The Zero-WAM answer: a human video. Language loses too much. A video of a person folding laundry carries the grip, the wrist rotation, the order of operations, the visual feedback loops. The paper builds a causal video-action model that takes a human video prompt plus the robot's current observation and predicts the next chunk of actions.

Training this thing required data that didn't exist. Paired human-robot in-context learning data is scarce, so the authors built an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos. The result is HumanGen: 74.2K human-robot pairs across 8.6K manipulation tasks. The scale matters because it's the first dataset that makes video-prompted manipulation training practical at all.

They also added an in-context future chunk prediction (IFP) objective. This is the subtle part. A naive model can cheat by memorizing shortcuts from seen tasks and ignoring the video prompt. IFP forces the policy to actually draw task information from the video. On seven unseen tasks in the RoboTwin 2.0 simulator, Zero-WAM hit 47.0% average success, an absolute gain of 29.5 points over the strongest video-action baseline. In the real world, it followed human video guidance for multi-object scenes, long-horizon tasks, and fine-grained insertion.

Read those numbers honestly. 47% means more than half of unseen tasks still fail. This is a research result, not a product. But the 29.5-point jump over the previous best approach says the direction is right.

The object-centric alternative ​

Zero-WAM operates on 2D video, and that's the weakness a second paper, 4DGS-WAM, is attacking. World action models that work on 2D frames produce nice visuals, but they lack explicit spatial structure for individual objects. They also reprocess the same static background every frame, which is pure waste.

4DGS-WAM lifts observations into a 4D Gaussian Splatting representation that separates dynamic objects from the static background. A policy model predicts the actor's future actions, and a world model predicts how the observed Gaussian splats of those dynamic objects will transform. The background doesn't need regeneration because it's already been seen. Future prediction focuses only on what changes. The paper evaluates on KITTI-MOT for short-horizon prediction and past reconstruction.

The tradeoff between the two approaches is real. Video-based WAMs like Zero-WAM scale with the kind of data you can harvest from the internet and crowdsourcing. Object-centric models like 4DGS-WAM are more compute-efficient per scene and give you object permanence, but they're harder to scale across scenes. My guess is the eventual winner is a hybrid: video-conditioned policies that internally build object-centric representations. Nobody has built that yet.

Quick Take: The field is converging on a simple bet. If you can specify a manipulation task in a human video, a policy trained with the right objective can execute it without ever seeing that task in training.

Why data is the real bottleneck ​

The reason both papers needed new datasets is the same reason Figure launched Index. Embodied AI has no internet. 杨海波, co-founder of 光轮智能 (Guanglun Intelligent), puts the gap this way: embodied intelligence needs at least 1000x more data than a large language model. The world's 8.1 billion people produce an estimated 100 billion hours of physical interaction every day. Almost none of it is recorded.

Figure's Index app, which went global on August 26 after four months of testing, is a direct attempt to buy that data, as 潮涌AI detailed. Ordinary people record first-person video of themselves doing chores, upload it, and get paid. In testing, Index covered 108 countries, accumulated 16 million videos, and paid creators $15 million. Figure has committed over $1 billion to data and compute in the next 12 months. Each 1,000 hours of data averages 373 unique tasks, 1,146 manipulated objects, and 116 distinct environments.

The model side explains why Figure needs this. Its Helix system pairs a 7B-parameter vision-language model with an 80M-parameter vision-motion network. The 80M network is small enough to run real-time control on the robot; the 7B model handles task understanding. That split only works if the motion network has seen enough real human manipulation, and the only place to get it at scale is from people doing chores in their actual kitchens.

Key numbers

  • 16M first-person videos collected by Figure's Index app across 108 countries in four months of testing
  • $15M paid to creators so far, with $1B+ committed to data and compute over the next 12 months
  • 74.2K human-robot in-context learning pairs in Zero-WAM's HumanGen dataset, spanning 8.6K tasks
  • 47.0% average success on seven unseen tasks, +29.5 points over the strongest video-action baseline
  • 55.5% of disclosed Chinese humanoid procurement value went to education and research in H1 2026

The data economy race ​

Chinese companies see the same wall, but they're climbing it differently. The contrast is stark.

PlayerStrategyScaleConstraint
Figure (Index)Crowdsourced first-person video, paid per upload16M videos, 108 countriesPrivacy, per-hour cost
AgiBot (智元)Physical data factory + open-source release850TB, 217 tasks, 3000+ objectsHeavy capex per new site
Unitree (宇树)Full-stack hardware + open teleop dataset340 hours, 1.89M trajectoriesGrowth capped by community
UBTech (优必选)Centralized collection, internal closed loopRoboMIND: 279 tasksSelf-funded, low openness
Robotera (自变量)Real-scene data + bodyless collection rigsClaims ~60% lower data costStill asset-heavy

AgiBot runs a 2,000-square-meter collection facility covering 217 tasks, and its open datasets now supply more than 80% of the real-robot training data for NVIDIA's GR00T N1. Unitree open-sourced 340 hours of whole-body teleoperation data in March. Robotera's QUANXTA Zero system claims to cut data collection cost for simple tasks by about 60%.

Every one of these is a different bet on the same problem. Figure bets that economic incentives plus global reach beat everything. AgiBot bets on controlled quality at scale. Robotera bets on cheaper collection hardware. The uncomfortable pattern is that nobody has solved the standards problem: every player collects in its own format, with its own annotation rules, and the resulting datasets can't talk to each other. The industry is building ponds, not an ocean.

Who's actually buying ​

The procurement data explains why the data race is so frantic. In 2025, China saw 48 disclosed humanoid robot orders above 10 million yuan, totaling more than 5.7 billion yuan. Impressive, until you ask what "order" means. Intent agreements, framework contracts, formal purchases, production, delivery, acceptance, revenue recognition: that's seven or eight distinct stages, and public reporting compresses them all into one word.

The H1 2026 winning-bid data is more honest. Across 218 disclosed projects worth 1.72 billion yuan, education and research institutions took 55.5% of the value. Government and public platforms took 20.6%. Industrial and technical enterprises took 21.1%. In July alone, universities and vocational schools accounted for over 90% of winning bid value. Unitree's own revenue mix confirms it: 73.6% of humanoid revenue in the first nine months of 2025 came from research and education. Real manufacturing and inspection work was 2.64%.

铅笔道 also documented three ways orders get inflated. Local governments place procurement orders as a disguised form of industrial subsidy. Component suppliers and robot makers buy from each other to prop up shipment numbers. Overseas integrators ask for sample units with vague promises of thousand-unit contracts that never materialize. One supply chain finance firm cross-checked logistics and invoicing data and found that a large share of signed contracts generated zero engineering service fees within six months. The machines were never delivered, even as demos.

The one real industrial signal is State Grid. Its 2026 embodied intelligence plan calls for 8,500 units at 6.8 billion yuan, with the three main categories breaking down as follows:

CategoryUnitsBudgetUnit price
Humanoid live-line operation robots5002.5B yuan5M yuan
Quadruped inspection robots5,0001.5B yuan300K yuan
Dual-arm inspection robots3,0001.8B yuan600K yuan

State Grid expects each device to save 500,000 to 800,000 yuan per year in labor costs, a two-to-three-year payback. But of that 6.8 billion, what has actually cleared bidding so far is a ~70 million yuan framework agreement for distribution-network live-line robots and a 128 million yuan quadruped pilot. The rest is planning. The economics also explain the skepticism: a UBTech humanoid averages 760,000 yuan per unit, and the industry consensus is that mass replacement needs prices under 150,000. Humanoid joints last hundreds of hours; industrial robot joints exceed 10,000. Battery life is 2 to 4 hours; factories run 24/7.

The cloud can't fix this ​

Days later, 亿欧网 (EqualOcean) published a sharp takedown of the "four clouds": Alibaba, Huawei, Tencent, and Baidu. All four have spent eighteen months packaging AI capabilities for embodied intelligence. Huawei's CloudRobo bundles data synthesis, model development, simulation, and deployment. Alibaba's Qwen-Robot series includes Manip, Nav, and World models trained on 38,100+ hours of open-source data. Tencent offers a full stack plus an Embodied-AI-as-a-Service subscription. Baidu pairs its Baige compute platform with a "data supermarket."

EqualOcean's argument is that all four are selling shovels while the gold is somewhere else. The data flywheel that matters spins at the deployment site: robots run, fail, generate correction data, and that data flows back into the model. That reflux belongs to whoever owns the body and the site. Cloud providers, by their own design, don't. Huawei's own materials admit the industry data gap exceeds 95%. Alibaba's Manip was trained entirely on open-source data, which means it has zero first-party friction from real robot deployments. Baidu's embodied data comes mostly from outsourced collection, not from robots failing in the field.

The deeper point applies beyond the clouds. The metrics that will decide this industry over the next 18 months are unit task cost, effective working hours, human takeover rate, and replication cycle. All four are measured at the deployment site. None of them can be improved from a cloud console.

Capital is still placing bets ​

None of this has stopped the money. The Information reported that SoftBank is in talks to acquire a majority stake in Norway's 1X Technologies at around a $6 billion valuation. It would be SoftBank's largest direct controlling investment in humanoid robots, and it caps a wild valuation arc: $820 million in January 2025, a failed attempt to raise $1 billion at a $10 billion valuation last fall, and now $6 billion. 1X's Neo humanoid, priced at $20,000, took over 10,000 pre-orders in its first week last October. As of now, zero units have been delivered. SoftBank's own history here is complicated: it killed Pepper, sold most of Boston Dynamics, and only returned to the table after buying ABB's robotics division for $5.4 billion in October 2025. OpenAI also explored acquiring 1X last year before talks stalled.

The same pattern shows in Chery's Mojia Robotics, which announced IPO preparations at the World Robot Conference. Phoenix Weekly reported that Mojia was founded in January 2025, raised its angel round at a 2.5 billion yuan post-money valuation, and has signed 1,030 police robots with 110 delivered. Its revenue for the first three quarters of 2025 was 3.03 million yuan against a 5.1 million yuan net loss. The company is targeting 10,000 global deliveries in 2027. That gap between valuation and revenue is the whole industry in miniature.

What the community is saying. I've watched the blood-drawing machine and the dancing robots that keep circulating on Reddit, and the whiplash between those clips and the procurement data is hard to overstate. The viral demos show what a robot can do once, in a favorable setup. The order books show what robots do reliably, without an engineer in the loop. Those are different capabilities, and the industry is still mostly selling the first while the market is waiting for the second. As Robotera founder 王潜 put it: the hardware is ready, but the brain hasn't caught up.

Common pitfalls ​

Five mistakes keep showing up across the papers, the procurement data, and the platform strategies.

Counting intent agreements as orders. Framework contracts and purchase intentions are not revenue. Check delivery, acceptance, and whether the robot runs without an engineer on site. 铅笔道's 80% figure exists because the industry let one word cover eight different stages.

Using language as the task interface for manipulation. Language is too lossy for grip, wrist rotation, and contact dynamics. Zero-WAM's whole premise is that human video is the right specification. If your policy ignores the visual prompt and memorizes task shortcuts, you'll get great training numbers and zero generalization.

Training video-action models on 2D frames without object structure. You reprocess static background every frame and lose object permanence