The current era of artificial intelligence is defined by a stark divide between the digital and the physical. While large language models have effectively indexed the sum of human knowledge by scraping the open internet, physical AI remains trapped in a data drought. Robotics engineers possess the compute power and the transformer architectures, but they lack the equivalent of a Common Crawl for the physical world. There is no massive, standardized dataset of how to fold a shirt or open a box that a model can simply ingest to achieve general-purpose dexterity. This gap has turned data acquisition into the primary bottleneck for the next generation of autonomous machines.

The Financial Velocity of Robot Data

XDOF has emerged from stealth mode to position itself as the primary infrastructure layer for this missing data supply chain. The company is currently in discussions for a Series B funding round that would value the startup at approximately 1.2 billion dollars. This round is being led by venture capital firm 8VC and comes at an extraordinary pace. Only three months after exiting stealth, XDOF is already scaling its valuation following a 70 million dollar Series A in June, which saw participation from heavyweights including Thrive Capital, Andreessen Horowitz, Lux, and Spark Capital.

The aggressive valuation is backed by rapid commercial traction. XDOF reports an annualized revenue run rate approaching 50 million dollars. While the company initially had no immediate plans for further fundraising after its Series A, the sheer velocity of its growth metrics prompted venture capitalists to proactively seek a new round. Founded in 2024 by UC Berkeley researchers Philipp Wu, who serves as CEO, and Fred Shentu, the CTO, XDOF specializes in real-world teleoperation data collection. The company currently maintains partnerships with 20 customers, including several of the world's leading frontier AI research labs.

From Model Architecture to Data Pipelines

The strategic pivot XDOF represents is a shift in how the industry views the path to general-purpose robotics. For years, the focus remained on refining model architectures and reward functions. However, the industry is realizing that the quality and volume of training data are the true determinants of performance. XDOF defines itself as the Scale AI or Mercor of the robotics world. Just as Scale AI fueled the LLM boom by providing the labeling and curation infrastructure that AI labs could not build internally, XDOF provides the outsourced pipeline for the physical world.

To break the data bottleneck, XDOF employs a dual-track collection strategy. The first is the GELLO system, a low-cost teleoperation interface that allows human operators to remotely control robot arms, creating high-fidelity demonstrations of complex tasks. The second is egocentric data collection, where humans wear sensors while performing everyday activities like folding laundry or organizing boxes. By capturing the world from the perspective of the actor, XDOF creates a bridge between human intuition and robotic execution. This methodology is the foundation for the ABC dataset, a massive high-quality robot training set being developed in collaboration with the UC Berkeley AI Lab, which aims to be the largest of its kind.

This approach transforms the competitive landscape. While competitors like Mecka AI focus on similar niches, and established platforms like Scale AI and Micro1 are expanding into robotics, XDOF is building a moat based on human infrastructure. The company does not just provide software; it manages and trains a global fleet of teleoperators and sensor-equipped agents. This suggests that the ultimate winner in physical AI will not necessarily be the team with the best algorithm, but the team that controls the most efficient pipeline for converting human movement into machine-readable tokens.

For developers and robotics firms, this signals a move toward the commoditization of robot data. The cost and time required to build an in-house data collection environment are becoming prohibitive compared to purchasing refined, high-quality datasets from a specialized provider. As the ABC dataset becomes available, the industry will likely see a surge in model performance that correlates directly with the diversity of human movement data rather than the number of parameters in the model.

The race for physical AI has shifted from a battle of code to a battle of logistics.