Why Robot Training Data Is Becoming Big Tech’s Next Race

The next major robotics race may not be decided by who builds the most impressive machine. Instead, the winner could be whoever collects the most useful training data.

That idea drew fresh attention on September 4, 2026, when TechCrunch reported that XDOF was in late-stage talks for a Series B round at a valuation of about $1.2 billion. The terms weren’t final, and neither XDOF nor the reported lead investor confirmed the deal. Even so, the report points to a meaningful change in where the robotics industry believes value is being created.

Robots have a data problem

Large language models benefited from vast quantities of text and code. Robots don’t have an equivalent internet-scale dataset that shows them how to grasp a cup, fold fabric, open a drawer, or recover after making a mistake.

Robotics companies have to collect physical demonstrations instead. People operate robots while cameras record the scene, sensors track movement and contact, and specialists label the results. Compared with gathering digital information, the process is slower, more expensive, and much harder to standardize.

XDOF says it is building infrastructure for this work, including production-scale datasets, robotic systems, and tools for robotics laboratories and companies developing physical AI. In practice, the company is aiming to become part of the supply chain behind robot intelligence rather than selling one more robot.

Why teleoperation matters

One of XDOF’s strongest ties to the robotics research community is GELLO, a low-cost teleoperation framework created by co-founders Philipp Wu and Yide Shentu alongside collaborators at UC Berkeley. It allows a person to control a robot arm with a mechanically similar controller, which makes complex movements easier to demonstrate.

Infographic showing how teleoperation data is collected, processed, and used to train physical AI robots.

The original GELLO research centered on reducing the barriers to collecting high-quality demonstrations. That matters because imitation-learning systems can acquire skills from human examples, but their performance depends heavily on the quality, variety, and scale of those examples.

Infrastructure before autonomy

For years, robotics companies have competed over motors, grippers, batteries, and humanoid designs. Now the competition is shifting one layer deeper. The question is who can build the data pipelines, annotation systems, evaluation tools, and operating procedures needed to turn physical activity into useful machine-learning material.

Better data won’t instantly produce household robots. Real-world environments are still unpredictable, and collecting demonstrations at scale remains difficult. But the shift suggests that physical AI is starting to resemble the broader AI market, where specialized infrastructure companies became essential as models grew larger and more expensive to train.

XDOF’s reported financing interest is one sign of a broader recognition that robots need a data industry around them. If that infrastructure matures, the next wave of robotics progress may come as much from better training loops as it does from better hardware.

Leave a Comment

Related Posts