Teahose.
SIGN IN
NEW HERE — WHAT TEAHOSE DOES
We read the entire AI & tech firehose — so you don't have to.
PODPodcastsAll-In, No Priors, Acquired…
NEWNewslettersStratechery, Newcomer…
PAPPapersPhysical AI research
PHProduct Huntdaily launches
VCInvestor ScoutSequoia, a16z, Benchmark…
CLAUDE DISTILLS →
7 reads, 30 sec each — free, 6 AM ET.
+ a live graph of the companies, people & themes underneath.
HOME/GUIDES/ROBOTICS DATA
GUIDE

Robotics Training Data: The Companies Feeding Physical AI (2026)

Every ranking page on this topic is written by a vendor selling data services. This one isn’t. Here is the full robotics training data landscape — the companies ranked by disclosed funding, the real per-hour prices, the data pyramid, and what the people building robot foundation models actually say about how much data robots need — drawn from 1,779 expert podcast, newsletter and research summaries.

Bryan Altman
Bryan Altman
Founder, Teahose · angel investor & builder
Updated 2026-08-24

Key takeaways

  • A real industry now exists to collect and sell robotics training data — and nearly its entire collection-specialist cohort formed in roughly eighteen months. The funded specialists: Lightwheel (≈$280M reported across 2026 rounds, at a reported $2B valuation), Tacta Systems ($75M), XDOF ($70M), Mecka ($68M), Micro1 ($35M Series A), Config ($35M), Human Archive ($8.2M), Hub.xyz, Build AI, plus Encord ($110M) on the infrastructure layer and Scale AI's generalist data engine. The full funding-ranked table, with sources, is below.
  • The market is far smaller than the noise suggests — but bigger than its only public estimate. Micro1's CEO, via MIT Technology Review in April 2026, put annual spend on real-world training data at more than $100 million — yet Mecka alone reports a $100M run rate and Lightwheel reported RMB 550M ($77M) of new orders in Q1 2026, so the honest read is low hundreds of millions and growing fast. For scale: one text-data vendor, Mercor, was reported at $614M gross revenue in H1 2026 alone, and Morgan Stanley's $5 trillion humanoid TAM assigns only ~6% to software, data and services combined.
  • Across the 1,779 expert podcast, newsletter and research summaries we've analyzed, the data problem is the most discussed bottleneck in robotics — teleoperation appears in 137 summaries, world models in 120, imitation learning in 60, and egocentric data in 40. Methodology: a count of how many of our expert summaries contain each term, August 2026. It measures share of expert discussion, not market share.
  • Nobody agrees how much data robots need — stated requirements span five orders of magnitude. NVIDIA's fine-tuning mix used 4 hours of teleoperation; Ant Lingbo's chief scientist puts the "robot GPT-1 moment" at 1,000,000 hours; and Ken Goldberg's yardstick for the gap is that vision-language models already trained on ~100,000 years of human-equivalent internet data while the largest reported robot teleoperation dataset is about one year. The chart below puts every expert-stated number on one log scale.
  • Prices exist, but the honest unit is the demonstration, not the hour. Public sources converge on roughly $1–15 per usable simple demonstration, while per-hour quotes diverge 4–10×: egocentric wearable capture runs ~$25–60 per collected hour, real-robot teleoperation ~$72–145 in China and ~$90–150 fully loaded in the US. The workers filming are paid $1–15 per hour offshore — and free open corpora (Build AI's 100,405 hours under Apache 2.0, AgiBot's million-trajectory datasets) keep resetting the floor.
  • Every other page ranking for this topic is written by a data vendor selling its own services. This one is written by a media and intelligence layer with no data business — which is exactly why it can tell you who the buyers are, what things cost, and where the bear case bites.

Robot training data requirements on one log scale — from 4 hours of teleoperation in NVIDIA's fine-tuning mix to a 1,000,000-hour robot GPT-1 threshold, with Ken Goldberg's 100,000-year internet-corpus yardstick shown for scale, as stated by named experts across Teahose's corpus and Science Robotics
Robot training data requirements on one log scale — from 4 hours of teleoperation in NVIDIA's fine-tuning mix to a 1,000,000-hour robot GPT-1 threshold, with Ken Goldberg's 100,000-year internet-corpus yardstick shown for scale, as stated by named experts across Teahose's corpus and Science Robotics

Tracking this market? Paste any robotics data company's website into Teahose Lookalikes and we'll show you its closest peers in our graph — and email you when new funding, product and partnership signals land on them.

Large language models had the internet. Robots have almost nothing: the experience data a robot needs — synchronized vision, motion, force and action — was never written down anywhere. That single fact created a new industry, a new gig-economy job category, and one of the sharpest strategic debates in AI. Yet search for "robotics training data" and every result is a vendor pitching its own pipeline. What follows is the neutral version: the companies (ranked by disclosed funding), the buyers, the prices, and the argument — grounded in what the people building robot foundation models actually say on the record, across the 1,779 expert summaries in our corpus. It is research, not investment advice.

What is robotics training data — and why there's no internet for it

Robot training data is the recorded experience used to train robot foundation models, and its defining property is scarcity. Skild AI co-founder Deepak Pathak, on the NVIDIA AI Podcast, states the founding problem of the entire category in three sentences:

"Robotics is a data problem. Unlike language or vision, there is not much data in robotics. There is no internet of robot data."

What robots need is not video in general but specific data: first-person viewpoints with the camera where the robot's head will be, hands and objects in close-up, and — for the highest tier — synchronized action labels (joint positions, gripper state, force) that say not just what happened but what the body did. YouTube is third-person, edited, and unlabeled. That is why this data must be manufactured, and why manufacturing it became a business.

The modalities stack up: RGB video, depth, IMU (motion), force/torque, tactile signals, audio, and language annotations describing tasks and failures. Delivery has quietly standardized too — managed programs typically ship data in RLDS or HDF5 formats, priced per hour of collected data rather than per clip or per episode.

The robot data pyramid: the one mental model that organizes everything

Every serious participant in this market carries some version of the same picture: a pyramid with internet-scale video at the base and real robot data at the tip. The framing was popularized by NVIDIA's GR00T N1 paper and is credited to UT Austin professor Yuke Zhu. What is striking in our corpus is that four independent experts describe the same pyramid — and then disagree about which layer does the work.

Xie Chen (谢晨), the ex-NVIDIA founder of simulation-data company Lightwheel (光轮智能), laid it out on Zhang Xiaojun's data-survey episode — the single richest conversation on this topic we have summarized (quotes translated from Mandarin):

"The smallest data volume will be from real deployed robots… The middle portion will be simulation-generated data. And at the bottom will be internet and human first-person perspective data. These bottom two — simulation and human first-person data — don't require a hardware body to generate, and their scalability is far superior." [00:51:19]

1X CEO Bernt Børnich told the All In robotics panel the same thing from a humanoid maker's seat — and drew the radical conclusion that the robot's body shape is itself a data strategy:

"The bottom layer in the pyramid, which is this video data, general video data, is absolutely ludicrously immense compared to anything else… Our cross embodiment is not another robot. Our cross embodiment is the human."

And Xie Chen's version carries the most useful pricing insight in the entire corpus — the counterintuitive fact that failure data is worth more than perfect demonstrations:

"People might think a perfect pizza-making video would be the most expensive. But actually, it's not. If you dropped a few pieces of vegetables and picked them back up, that's worth more. It's similar to human learning — the experience of failing then succeeding is the most valuable." [02:00:01]

That is not a throwaway line; it is load-bearing for how datasets get built. AgiBot's million-trajectory open dataset deliberately keeps failed demonstrations, labeled with an error_cause field, rather than filtering them out — the design choice that separates it from most success-only corpora.

The companies: a field guide to who collects what

There are four distinct sourcing models, and knowing which model a company runs tells you most of what you need to know about it.

Model A — gig-worker egocentric video. Pay distributed contributors to film daily life through head-mounted cameras. Mecka (New York and Toronto; four co-founders from fintech and e-commerce — CEO Josh Gao's own line, per Upstarts Media: "None of us have a robotics background, which is pretty funny") collects human-motion data through body sensors and iPhones; Fortune reports it projects a ~$100M annual run rate on signed contracts (while its CEO declined to name customers), and Upstarts Media reported its one publicly identified customer: 1X Technologies, which used Mecka household-motion data in building its World Model. Micro1 (Palo Alto) runs thousands of contract workers in 50+ countries recording household tasks — MIT Technology Review's April 2026 feature on it is the definitive portrait of this labor market, and CNN reports its contributors collectively submit more than 160,000 hours of video a month, the largest publicly reported collection volume anywhere. Hub.xyz (San Francisco; Y Combinator, Spring 2026) collects egocentric video through a global contributor network — its site advertises 540,000+ hours across 150 countries — with field operations running in Brazil. DoorDash launched a standalone app, Tasks, paying gig workers to record everyday physical tasks — the clearest signal an incumbent labor marketplace sees the same opportunity. Smaller entrants include Kled, Luel (a rights-cleared marketplace building on the academic Ego4D datasets) and Waffle Video.

Model B — instrument the worker, not the task. Capture skilled labor where it already happens. Tacta Systems (Palo Alto, $75M from America's Frontier Fund, SBVA, SoftBank, Toyota's Woven Capital and others) puts a sensorized glove — force, motion, video, temperature — on workers on real production lines; co-founder Andreas Bibl previously founded LuxVue, acquired by Apple in 2014. Build AI (founder Eddy Xu, a teenage dropout running a public benefit corporation) paid 14,000+ factory workers across Southeast Asia to wear camera glasses — and then released the resulting 100,405-hour Egocentric-100K corpus under an open Apache 2.0 license (access-gated on Hugging Face). Human Archive (YC W26, $8.2M) deploys 1,000+ multimodal headsets through India's gig economy — and is also the category's cautionary tale: TechCrunch reports India's Ministry of Electronics and IT is examining consent mechanisms under the country's data-protection law, and two major Indian gig platforms — Urban Company and Pronto — declined to partner over privacy concerns.

Model C — teleoperation and robot-native trajectories. XDOF (launched June 2026 with $70M from Thrive, Spark, a16z, Lux and WndrCo; ~20 customers including unnamed frontier labs) is the flagship: its founders built GELLO, the low-cost Berkeley teleoperation rig, and it spans all three collection philosophies — teleop on deployed robots, GELLO-rig teleop, and wearable sensors. With researchers from Berkeley, CMU, MIT and Amazon it released an Apache-2.0 dataset of 130,000+ teleoperated episodes (~3,600 hours) across ~195 bimanual tasks. Config (Seoul and San Jose, $35M total; $27M seed led by Samsung Venture Investment with Hyundai's, LG's and SK Telecom's venture arms and angel Pieter Abbeel) focuses on bimanual manipulation data. Scale AI runs the best-documented generalist operation — more on it in the buyers section. The UMI research lineage (a handheld gripper plus GoPro, "in-the-wild robot teaching without in-the-wild robots") sits underneath this whole model; NVIDIA's Jim Fan calls UMI "perhaps one of the greatest papers ever written in robotics data," and its first author, Cheng Chi, co-founded Sunday Robotics — whose Skill Capture Glove (roughly $200 to produce, against ~$20,000 for a teleoperation rig) has collected a company-stated 10 million demonstration episodes from 500+ wearers with zero robots in the loop; Sunday raised $165M at a $1.15B valuation in March 2026 largely on the strength of that data engine.

Model D — simulation and synthetic data. Lightwheel (Beijing; founded by Xie Chen above) builds SimReady assets and simulation-generated corpora, reportedly raising at least RMB 2B (≈$280M) across three 2026 rounds — including an Ant Group-led strategic round at a reported $2B valuation — with RMB 550M ($77M) of new orders reported in Q1 2026 alone. NVIDIA is the gravitational center: Isaac Sim, the Cosmos world models, and GR00T-Dreams, which NVIDIA says generated the synthetic training data for GR00T N1.5 in 36 hours versus roughly three months of manual collection. Genesis AI ($105M seed, Eclipse and Khosla) and Hillbot (built on the open-source ManiSkill engine) sell the same thesis. Crypto-incentivized networks round out the map: PrismaX ($11M, a16z's crypto accelerator) pays contributors for teleoperation and visual data; FrodoBots/BitRobot ($8M, with Protocol Labs) turned $250 sidewalk robots into a "drive to earn" game and released ~2,000 hours of real urban teleoperation data.

One pattern worth naming: most of these founders don't come from robotics. Fintech exits (Mecka), a display-technology founder (Tacta), consumer-data operators (Hub.xyz), a teenager (Build AI). (Human Archive's young founding team, out of robotics and tactile-sensing research, is the exception that runs a gig network.) The robotics-native hardware teams — XDOF (Berkeley GELLO) and Proception (ex-Tesla Optimus engineers, whose $11M First Round-led seed arrived the same month it settled Tesla's trade-secret suit) — build hardware and rigs, not gig networks. Collecting human data at scale is a logistics-and-marketplace business wearing a robotics costume, and marketplace operators are winning it.

Robotics data companies ranked by disclosed funding

This is the table the vendor-written pages won't publish. Data-layer specialists only — the labs that buy data are in the next section. Sorted by total disclosed funding, largest first.

#CompanyWhat they sellTotal disclosed fundingHow we sourced the number
1LightwheelSimulation data, SimReady assets, evaluation≈$280M+ (at least RMB 2B across three 2026 rounds)Chinese tech press (Sina Finance, Gasgoo): ≈RMB 1B March A++/A+++ round, an Ant Group-led strategic round in May at a reported ~$2B valuation, and a further ≈RMB 1B round in June — reported, not audited
2EncordData infrastructure, annotation, curation$110MCompany announcement of $60M Series C (Feb 2026, Wellington-led) plus prior disclosed rounds
3Tacta SystemsGlove-captured manipulation data + robot hands$75MCompany press release, July 2025 ($11M seed, Matter Venture Partners + $64M Series A, America's Frontier Fund/SBVA)
4XDOFTeleoperation + wearable-sensor data for labs$70MTechCrunch launch coverage, June 17, 2026 (Thrive, Spark, a16z, Lux, WndrCo)
5MeckaHuman-motion data (body sensors + iPhones)≈$68MFortune, June 1, 2026 ($25M Series A + $35M follow-on, Framework Ventures) plus the $8M seed (led by Neo) reported by Upstarts Media, Aug 2025
6Micro1Gig-worker egocentric video, 50+ countries$35M Series ASeptember 2025 round at a reported $500M valuation; MIT Technology Review profiled the operation in April 2026
7ConfigBimanual manipulation data + policy training$35MTechCrunch, May 2026 ($27M seed led by Samsung Venture Investment at a $200M+ valuation; $35M total stated in round coverage)
8Build AIFactory-worker POV video — openly licensed≈$15MRound coverage (Humanoids Daily) naming Abstract, Pear and HF0, joined by Balaji Srinivasan, Guillermo Rauch and Thomas Wolf
9PrismaXCrypto-incentivized teleop + visual data network$11MThe Robot Report, June 2025 (a16z CSX-led seed)
10ProceptionDexterous robot hands + glove-captured demos$11MTechCrunch, June 2026 (seed led by First Round Capital, with YC and BoxGroup)
11Human ArchiveMultimodal headset data via India's gig economy$8.2MTechCrunch, May 2026 (Wing Venture Capital and NVP-led seed)
12FrodoBots / BitRobotSidewalk-robot teleop network (crypto)$8MChainwire, February 2025 ($6M seed round; $8M total, with Protocol Labs)
13Hub.xyzEgocentric video, global contributor network$1.7MCrunchbase / YC profile (pre-seed)

Disclosed rounds only, compiled from company press releases and named-outlet reporting as of August 24, 2026. Rumored or in-talks rounds are excluded; so are companies with no public funding disclosure (Kled, Luel, Waffle Video, Objectways) and Scale AI, whose robotics business is a product line inside a company Meta paid $14.3B for a 49% stake in. (The same caveat applies in miniature to Micro1, whose Series A funds a broader AI-data business of which robotics collection is one line — it stays in the table because, unlike Scale, its disclosed robotics-era funding is separable and its egocentric program is its flagship growth story.) Sunday Robotics ($200M raised at a $1.15B valuation) is excluded for the opposite reason: its glove-collected corpus feeds its own robot — it sells robots, not data. Every figure above can be traced to the source named in its row.

Two reading notes. First, the entire data-vendor layer above sums to roughly $730M of disclosed funding — less than Skild AI raised in its single January 2026 Series C. Second, the biggest number in the table (Lightwheel) rests on Chinese tech-press reporting rather than audited filings; when a market is this young, funding tables that don't say how they know each number are advertising, not analysis.

A table like this ages in weeks, not years. Watch the physical AI theme on Teahose and new rounds, launches and partnerships from these companies land in your inbox as our system extracts them from expert coverage.

Who's buying — and why the demand side is deliberately invisible

The single strangest fact about this market: two separately-funded startups both say frontier labs pay them, and neither can name most of them. XDOF claims ~20 customers "including several frontier AI labs." Mecka reports a ~$100M run rate from signed contracts with exactly one customer on the public record — 1X Technologies. Beyond that, the only disclosed customer list anywhere is on Scale AI's own product page: Physical Intelligence, Generalist, Cobot and Dyna.

What the buyers do say in public is how they get data themselves:

  • Scale AI is the generalist incumbent's play, and the numbers are its own: more than 150,000 hours of physical-AI data delivered in 2025 (per July 2026 trade reporting), ten new robotics customers, a currently claimed pace of "1000+ hours of diverse demonstration data collected and uploaded daily," collection running through data factories, residential settings and commercial deployments, plus a partnership with Universal Robots to collect force-feedback data directly on production arms. Scale's September 2025 framing of the scarcity is the most-quoted stat in the field: all open-source robotics datasets combined offered "only about 5,000 hours of interaction data." (Its robotics push also reads as a response to losing LLM-data neutrality after Meta's $14.3B investment — the story we cover in Scale AI competitors.)
  • Figure solved data sourcing with real estate instead of vendors: its September 2025 Brookfield partnership opens 100,000+ residential units and 500M+ square feet of commercial space for egocentric human-video collection, and Figure says its Helix navigation model trained on 100% human video, zero robot demonstrations (Project Go-Big). More in Figure AI valuation.
  • Tesla made the same pivot the hard way: Business Insider reported in August 2025 that Optimus data collection moved away from mocap suits and VR teleoperation toward workers wearing multi-camera helmet rigs — the most vertically-integrated humanoid program concluding that human video at scale beats high-fidelity teleop on cost per useful hour.
  • Physical Intelligence (≈$1.1B raised; $600M round at a $5.6B valuation in late 2025, with Bloomberg reporting talks at $11B in March 2026 — no close confirmed as of late August 2026) collects its own data at what it calls unprecedented scale and appears on Scale's customer list — the realistic answer is labs do both. Its π0 model trained on 10,000+ hours of robot data across 7 robot configurations; its π0.5 generalization result needed only ~400 hours across ~100 home environments.
  • Google DeepMind trains Gemini Robotics on ALOHA rigs and partner platforms; Toyota Research Institute and Boston Dynamics run teleop-on-Atlas as an explicit data loop; AgiBot collects on fleets of a hundred identical robots and open-sources the result; Skild AI (company-stated total funding above $2B; valued at $14B in its January 2026 Series C) pre-trains on video and simulation precisely to minimize how much real-robot data it needs.

The pattern: vendors are one leg of a three-legged sourcing strategy (buy, build, simulate), and no lab wants competitors to know its mix — hence the NDAs. As Xie Chen puts it, the data company and the model company are becoming "symbiotic": labs need external evaluation and data; data companies need model feedback to know what to collect next.

What robot training data costs

No page ranking for this query publishes real prices — and the numbers that do circulate usually omit the label that matters most: geography. Here is everything that is public, with its provenance attached.

Start with the unit the market actually agrees on. Across every public source — vendor rate cards, cost breakdowns, unit-economics posts — a simple usable single-arm demonstration converges on roughly $1–15. Per-hour prices for the same work diverge 4–10×, because vendors assume anywhere from 6 to 60 demonstrations per hour and rarely say which. A per-hour quote that doesn't state throughput, geography and acceptance rate isn't a price; it's three hidden assumptions wearing a dollar sign.

Buy-side reference points (every figure vendor- or trade-press-reported; treat as list prices, not market quotes):

What you're buyingReported buyer price per collected hourProvenance
Egocentric wearable (no robot)$25–60vendor rate card, APAC managed programs; Chinese platforms quote ¥200–400 (~$29–58)
Single-arm teleoperation (real robot)¥500–1,000 (~$72–145) in China; ~$90–150 fully loaded in the US, with one published US benchmark at $118Chinese trade press, August 2026; Robotics Center of Silicon Valley benchmark, March 2026
Bimanual (ALOHA-style)$40–80 offshore; roughly 2–4× single-arm rates onshorevendor rate card; ~3–15 demos/hour, 70–85% episode acceptance
Full humanoid multi-sensor$80–150, or $50–150 per usable demonstrationvendor rate card; 1–3 demos/hour

Reported figures as published in 2026; delivery in RLDS/HDF5 is the quiet standard. One APAC vendor lists simple single-arm teleoperation as low as $15–30 per collected hour — a real offshore rate for the simplest tasks, but it sits below the wage alone of a US teleoperator (Tesla's data-collection-operator role paid $25–48/hour; a current US robot-teleoperator posting offers $30–55/hour), so it should never be read as a general market price.

Two haircuts separate a collected hour from a usable one. Chinese trade reporting — the only independent full cost stack published in any language — finds that 100 collected hours yield about 50 usable, and that a 12-hour collection shift produces under 6 effective hours; US vendors claim 60–90% acceptance. Cost per usable hour therefore runs roughly 1.2–2× the collected-hour figure — which is why finished, QA'd US teleoperation data realistically lands around $120–200 per usable hour.

Sell-side — what the people filming get paid (all reported, with outlets): MIT Technology Review documents Micro1 contributors earning around $15/hour — among them a Nigerian medical student, good income locally; on a different program, TechCrunch reports Human Archive paying Indian collectors a base rate of about $1/hour, against ₹250–400 (~$2.63–4.20) per hour at competitors; a remote-contributor job posting for egocentric collection across eight countries offered $6/hour.

And nobody neutral publishes a number. No academic paper discloses teleop labor cost — ALOHA and DROID publish hardware prices ($20–32K per rig) and demonstration counts, never dollars — and most vendors publish no rates at all. Every per-hour figure above comes from an interested party measuring its own market: the $118/hour US benchmark is a data-services vendor's self-published figure (it does hold up against a bottom-up build from public wages and hardware prices), and the cheap offshore quotes are marketing. The wage-to-price spread — $1–15/hour to the worker, $25–150+/hour to the buyer — is what funds rigs, robots, QA, rejected episodes and margin, and no company discloses how that spread splits.

How much data is enough? The 250,000× question

The chart at the top of this page is the honest answer: stated requirements for robot data span a factor of 250,000 — and the gap to internet scale is starker still. Every figure on it is a direct claim by a named expert in an episode we've summarized, or a published calculation.

At the small end, Rhoda AI's Jagdeep Singh — a $450M Series A at a $1.7B valuation to train robots on internet video — turns the hours arms race into a punchline:

"You'll see some people claim that we trained on 70,000 hours of robot data. Then you'll see the next guy saying we train on 270,000 hours… and then you come to Rhoda, which is later in time, and we say we train on 10 to 20 hours of data." [00:20:55]

NVIDIA's Jim Fan, presenting at Sequoia's AI Ascent, put real proportions on the pyramid — and reported a scaling law:

"We pre-train on 21k hours of in-the-wild egocentric human data with zero robot data whatsoever… Then in action fine-tuning we collect only 50 hours of high-precision mocap data and four hours of tele-op. That's four hours of tele-op, less than 0.1% of our training mix." [00:11:31]

"We discovered this neural scaling law for dexterity… a clean log-linear mathematical equation, six years after the original neural scaling law for language models." [00:12:29]

At the large end, Ant Lingbo chief scientist Shen Yujun defines the threshold that matters (translated):

"I think it should start at one million hours — when you can start saying the embodied data and the internet data are relatively balanced." [00:21:51]

"The robot data is about two orders of magnitude smaller than internet data." [00:08:55]

His own corpus is at 60,000 hours; Generalist — the largest reported anywhere — trained GEN-0 on 270,000 hours and says its corpus now exceeds half a million, with co-founder Pete Florence publicly claiming collection above 10,000 hours per week. At that pace a single company crosses Shen's 1,000,000-hour threshold around 2027 — the "GPT-1 moment" yardstick is roughly a year of collection away, not a philosophical horizon.

And above them all sits the field's most-quoted — and most-misquoted — number. UC Berkeley's Ken Goldberg, in a 2025 Science Robotics editorial (invoked as an impossibility argument in the Moneyball for Physical AI newsletter we've summarized), calculates that the internet corpus vision-language models already train on equals roughly 100,000 years of human experience — while the largest reported robot teleoperation dataset is on the order of one year. The 100,000 years is not a collection target; it is the size of the gap. And Goldberg's own conclusion is that nobody should try to teleoperate across it: he argues for closing it with rigorous engineering plus "real data flywheels" — robots gathering data while doing paid work.

Two corpus findings cut against the raw-hours framing entirely. Physical Intelligence's π0.5 generalized to unseen homes after training on roughly 100 distinct home environments (about 400 hours of mobile-manipulation data) — you don't need every home on earth before generalization kicks in, and 97.6% of π0.5's first-phase training data came from robots other than the one being evaluated. And the Moneyball analysis carries the sharpest anti-volume stat in our corpus — a language-model scaling result it applies to robotics by analogy: repeating just 0.1% of a corpus 100 times collapses an 800M-parameter model's downstream performance to that of a 400M baseline. Diversity and novelty, not hours, are the binding constraint — which is exactly why "hours collected" is quietly being retired as the industry's KPI.

Human video vs. teleoperation vs. simulation: the three-way fight

This is the debate that decides which companies above survive, and our corpus has both poles on the record.

The case for human video. Jim Fan's teleoperation kill shot is the cleanest unit-economics argument in robotics:

"For tele-op, it's upper bounded by 24 hours per robot per day, the fundamental physical limit. And actually, who am I kidding? It's more like three hours per robot per day and only when the robot god is merciful, because they throw tantrums all the time." [00:07:15]

He predicts teleoperation drops "to almost negligible amount" within two years, with egocentric video as "the main diet for robotics." The Sunday Robotics founders add the quality argument nobody else makes — teleop data is secretly worse because the operator can't feel:

"There's not a good teleoperation system that can let you feel how much force the robot is feeling. So basically when you're teleoperating, your hand is numb." [00:36:03]

The case against. Skild's Abhinav Gupta, on the NVIDIA AI Podcast, in one line:

"If we can learn from videos, all of us would be Federers… If it was sufficient, I could dunk a basketball, but I can't."

And Physical Intelligence's Sergey Levine — the field's most consistent real-data advocate — reframes the whole question in his long-form interview: the point is not choosing a data source but reaching the flywheel, where deployed robots gather their own experience. His hardware-cost curve explains why that's close: the research arm he used in 2014 cost $400,000; his Berkeley lab bought $30,000 arms; Physical Intelligence's current arms cost about $3,000 each. Cheap bodies mean fleets; fleets mean autonomous data. The open question he flags is the one that should decide your view of every company in this guide: will robots need 90% demonstrations and 10% autonomous experience — or 10-90? "It does change the correct approach pretty dramatically."

Simulation's honest scorecard. Sim has won locomotion and visual diversity — NVIDIA generated GR00T N1.5's synthetic data in 36 hours versus a claimed three months of manual collection, and its GR00T N1 paper describes ~780,000 simulated trajectories generated in about 11 hours. But contact physics resists: a Physical Intelligence researcher notes that for clothing, friction and viscosity "you might not even be able to build such a simulator," and the strongest 2026 academic results are co-training — sim and real together beating either alone. Meanwhile the most contrarian position in our corpus belongs to Google DeepMind's Jie Tan: within two to three years, he argues, generative world models replace traditional physics simulation for scene generation entirely — "500 different home scenarios" becomes 500 prompts.

The bear case: what if collecting data is a bad business?

The strongest arguments against this category come from inside it. XDOF CEO Philipp Wu — who runs the best-funded pure data vendor — told TechCrunch that "simply creating the data itself is a poor business model," which is why XDOF is racing up-stack into cleaning, annotation and tooling. The bear case has four legs:

  1. The market is small. No credible sizing exists: the one public estimate (Micro1's CEO, via MIT Technology Review: over $100M/year) sits awkwardly beside vendors' own claims — Mecka's ~$100M run rate, Lightwheel's RMB 550M of Q1 2026 orders — but even a generous low-hundreds-of-millions read against ~$730M of vendor funding is an inverted ratio; the buyers' patience is doing a lot of work. The sharpest tell comes from a vendor itself: Config, freshly valued above $200M as the self-styled "TSMC of robot data," publicly targets just $10M ARR by the end of 2027. And even Morgan Stanley's $5 trillion humanoid bull case budgets only ~6% for software, data and services combined.
  2. Free data keeps landing. Scale framed the entire open corpus at ~5,000 hours in September 2025; three months later Build AI released 100,405 hours of factory-floor egocentric video under Apache 2.0, and AgiBot open-sourced over a million real-robot trajectories. Different data types, same effect: the volume tier deflates toward zero.
  3. The premium tier may be structurally thin. If NVIDIA's <0.1%-teleop mix generalizes, the expensive embodiment-specific apex of the pyramid — the part with real pricing power — is a niche, not a market.
  4. Deployment eats collection. The end-state consensus across our corpus, from Kyle Vogt ("the majority of data collection will come from robots and less from people getting data") to Karol Hausman ("that second stage is actually going to be the stage where we get most data from"), is that fleets of working robots replace paid collection. The counterpoint, from the Moneyball analysis: today's deployable niches are deliberately low-variance, so their data is low-entropy — "the AGI revolution will not be supervised with Sweatshop Teleop," but the flywheel's data may not generalize either.

The synthesis our corpus supports: raw hours commoditize; access and curation don't. The durable positions are (a) owning access to people and places at scale — Figure×Brookfield's 500M square feet, Micro1's 50-country gig network, Tacta's production lines, DoorDash's driver base — and (b) owning the quality layer: curation, evaluation, failure-data labeling. Xie Chen's most repeated observation is that his clients' real bottleneck isn't training data at all — "they cannot scale their evaluation anymore."

Bull or bear, the answer will show up in the signals before it shows up in the narrative. Watch Scale AI — or any company in this guide — and every funding, product and hiring move reaches your inbox the day we extract it.

Where to hear it from the founders

Podcast coverage of this category is itself a signal — and it is wildly uneven. Physical Intelligence's founders are everywhere (Sergey Levine's Dwarkesh Podcast episode is the biggest-reach conversation in the space; Chelsea Finn spoke at Y Combinator's Startup School in August 2026). Skild's founders did a joint NVIDIA AI Podcast episode, and Deepak Pathak appeared on Bloomberg Tech the week of its mega-round. A Tacta cofounder gave a deeply technical interview on the Soft Robotics Podcast in mid-August 2026. Genesis AI's Théophile Gervet does French business media (BFM's Tech&Co). And at the other extreme: as of August 2026 we could not find a single long-form founder interview from Mecka or Hub.xyz anywhere — the two youngest companies exist in the record only through funding coverage. For a category selling picks and shovels to the loudest gold rush in tech, the sellers are strikingly quiet.

Bottom line: Robotics training data is a real, funded industry — its specialist cohort barely eighteen months old — solving the defining bottleneck of physical AI. But it is a market measured in low hundreds of millions a year wearing trillion-dollar branding, its cheap tiers are commoditizing, and its own leaders say raw collection won't be the durable business. Watch the data companies anyway: as one of our sources puts it, every change in the model paradigm shows up in the data paradigm first — "the data we are collecting today will likely translate into model breakthroughs three to six months from now."

Live from the Teahose intel graph

Robotics Data Companies in the Teahose Graph

Vector similarity against the Teahose company graph · same engine as the /similar lookalikes tool · updates as new companies are extracted daily

  1. 01XDOF89% match
  2. 02Human ArchiveAI / Robotics Data87% match
  3. 03MeckaRobotics85% match
  4. 04Open X-Embodiment CollaborationRobotics84% match
  5. 05Assured Robot IntelligenceRobotics82% match
  6. 06PiAI81% match
  7. 07Micro1AI81% match
  8. 08Open X-Embodiment ConsortiumAI / Robotics81% match
  9. 09Physical IntelligenceAI79% match
  10. 10Ropediarobotics79% match
  11. 11Origin LabAI79% match
  12. 12Generalist AIRobotics79% match
  13. 13DexmalRobotics79% match
  14. 14Columbia University RoboPIL LabRobotics78% match
  15. 15Physical IntelligenceRobotics78% match
Updated continuously as new signals landRun this for any company

Related

Physical AI companies · What is physical AI · Scale AI competitors · Scale AI valuation · Physical Intelligence valuation · Skild AI valuation · Figure AI valuation · Humanoid robot companies · Robotics startups · Robotics investors & VC firms · Unitree stock · AgiBot stock · Live physical AI theme

Corpus figures from Teahose's analysis of 1,779 expert podcast, newsletter and research summaries, August 2026 — counts measure share of expert discussion, not market share. Funding figures are disclosed rounds only, as of August 24, 2026, sourced per the table above. Mandarin-language quotes are translated from our transcripts. None of this is investment advice.

Sources

Frequently Asked Questions

What is robotics training data?

Robotics training data is the recorded experience used to train robot foundation models — the "brains" that let robots see, decide and act. Unlike text models, which pre-trained on an internet that already existed, robots need data that mostly does not exist yet: synchronized video, depth, motion, force and action labels showing how physical tasks are performed. It spans four broad tiers — internet video, first-person (egocentric) human video, simulation-generated data, and real robot trajectories collected by teleoperation — and an entire industry has formed since 2024 to collect and sell it.

How is robot training data collected?

Four main ways. (1) Egocentric human video: people wear head-mounted cameras or sensor rigs while doing everyday tasks — the method used by Mecka, Hub.xyz, Micro1 and Build AI. (2) Instrumented tools: sensorized gloves or handheld grippers capture manipulation directly from human hands, with no robot in the loop — Tacta Systems, Sunday Robotics’ Skill Capture Glove, and the UMI research lineage. (3) Teleoperation: human operators drive real robots through tasks, producing robot-native trajectories — the method XDOF and Scale AI industrialize. (4) Simulation: physics engines and world models generate synthetic trajectories — NVIDIA’s Isaac/Cosmos stack and companies like Lightwheel. Most serious labs combine all four.

What is egocentric data collection for robotics?

Egocentric data is first-person video recorded from the point of view of the person doing a task — the camera sits where the robot’s head will be. It matters because scraped internet video is third-person, edited, and lacks the close-up hand-object detail robots need. The research case is now strong: NVIDIA pre-trained on 20,854 hours of egocentric human video and reported a clean log-linear scaling law between hours and validation loss (the EgoScale paper, February 2026), and Figure says it trained its Helix navigation model on 100% human video with zero robot demonstrations. That evidence is why a dozen startups now pay people to film chores, factory shifts and errands through head-mounted cameras.

Which companies sell robotics training data?

The funded, publicly-verifiable specialists as of August 2026 include Mecka (human-motion data via body sensors and iPhones, ~$68M raised), XDOF (teleoperation and wearable-sensor data, $70M), Tacta Systems (sensorized gloves on real production lines, $75M), Encord (data infrastructure and annotation, $110M), Micro1 (gig workers in 50+ countries recording household tasks, $35M Series A — and, per CNN, over 160,000 hours of video submitted monthly), Config (bimanual manipulation data, $35M), Human Archive (multimodal headset data via India’s gig economy, $8.2M), Hub.xyz (egocentric video via a global contributor network, YC-backed), Build AI (factory-worker POV video, openly licensed), PrismaX and FrodoBots/BitRobot (crypto-incentivized collection networks). Scale AI runs the generalist robotics data engine with the only disclosed customer list: Physical Intelligence, Generalist, Cobot and Dyna.

How much does robot training data cost?

Almost no ranking page will tell you, and every published figure is vendor- or trade-press-reported, so provenance matters. The unit the market actually agrees on is the demonstration: across public sources, a simple usable single-arm demonstration converges on roughly $1–15. Per collected hour, quotes diverge with geography and modality: egocentric wearable programs run about $25–60; real-robot single-arm teleoperation is quoted at ¥500–1,000 per hour in China (about $72–145) and around $90–150 fully loaded in the US, with one published US benchmark at $118 per hour; bimanual ALOHA-style collection is listed at $40–80 offshore; full humanoid multi-sensor capture at $80–150, or $50–150 per usable demonstration. Acceptance rates as low as ~50% mean cost per usable hour runs 1.2–2× the collected-hour figure. Meanwhile the workers filming earn $1–15 per hour offshore and $25–55 per hour in the US. Delivery in RLDS or HDF5 is the quiet standard, and there is no public exchange for robot data — treat every number as a reported list price, not a market quote.

How much training data do robots actually need?

Stated requirements disagree by five orders of magnitude — a factor of 250,000 — and that spread is the honest answer. Rhoda AI claims 10–20 hours of robot data suffice after internet-video pre-training. NVIDIA’s GR00T fine-tuning mix contained just 4 hours of teleoperation — under 0.1% of the total — on top of 21,000 hours of egocentric human video. Ant Lingbo’s chief scientist puts the "robot GPT-1 moment" at 1,000,000 hours of embodied data (his own corpus: 60,000). Generalist says its corpus already exceeds 500,000 hours, growing above 10,000 hours per week. And for scale: UC Berkeley’s Ken Goldberg calculates that vision-language models already trained on the equivalent of roughly 100,000 years of human experience, while the largest reported robot teleoperation dataset amounts to about one year. Nobody has run the experiment that settles it.

Is simulation or synthetic data enough to train robots?

Not on its own — the 2026 consensus is co-training, not substitution. Simulation has clearly won for locomotion and for visual diversity: NVIDIA generated the synthetic data for GR00T N1.5 in 36 hours versus a claimed three months of manual collection, and its GR00T N1 paper describes 780,000 simulated trajectories generated in about 11 hours. But contact-rich manipulation resists it. A Physical Intelligence researcher notes that for deformable objects — clothing, friction, viscosity — "you might not even be able to build such a simulator," and Skild AI’s Abhinav Gupta has the sharpest version of the video-only counterargument: "If we can learn from videos, all of us would be Federers… If it was sufficient, I could dunk a basketball, but I can’t." Every credible lab buys real data and generates synthetic data.

What is the robot data pyramid?

A mental model — popularized by NVIDIA’s GR00T work and UT Austin professor Yuke Zhu — that organizes robot training data by scale and cost. At the base: internet and human video (massive, cheap, not robot-specific). In the middle: simulation and synthetic data. At the apex: real robot trajectories from teleoperation (smallest, most expensive, most embodiment-specific). Quantity shrinks and cost rises as you climb. The strategic fight in 2026 is over how thin the expensive apex can get: NVIDIA’s EgoScale recipe — as Jim Fan described it at Sequoia’s AI Ascent — needed roughly 50 hours of human mocap and 4 hours of robot teleoperation on top of ~21,000 hours of egocentric video (the EgoScale paper reports the 20,854-hour pre-training corpus and its scaling law) — which, if it generalizes, shrinks the premium tier of this market to a rounding error.

Who buys robotics training data?

Robot foundation-model labs and humanoid makers — but most purchases are deliberately invisible. The only disclosed customer list is on Scale AI’s own product page — Physical Intelligence, Generalist, Cobot and Dyna — and the one other publicly confirmed pairing is Mecka supplying household-motion data to 1X Technologies. XDOF says it has about 20 customers "including several frontier AI labs" it cannot name; Mecka reports a ~$100M annual run rate from signed contracts while disclosing only that single customer. The biggest buyers also build in-house: Figure partnered with Brookfield to film egocentric video across 100,000+ residential units, Tesla shifted Optimus data collection from mocap suits to camera-rigged workers, and Physical Intelligence, AgiBot and Toyota Research Institute run their own robot fleets as data engines.

What is teleoperation data and why is it so expensive?

Teleoperation data is produced by a human driving a real robot through a task, recording perfectly embodiment-matched state-action trajectories — the highest-quality and most expensive tier of robot data. The economics are brutal for a physical reason NVIDIA’s Jim Fan states plainly: teleop is "upper bounded by 24 hours per robot per day, the fundamental physical limit. And actually… it’s more like three hours per robot per day." Every hour requires a trained operator, a working robot and QA review. There is also a subtler quality problem — Sunday Robotics’ founders note that "when you’re teleoperating, your hand is numb": operators feel no force feedback, so the demonstrations themselves are degraded. This is why the market has moved toward instrumenting humans instead of driving robots.

Do people really get paid to film themselves doing chores for robots?

Yes — it became a real gig-economy category in 2025–2026. MIT Technology Review documented thousands of contract workers in 50+ countries recording household tasks with iPhones strapped to their foreheads for Micro1, at rates around $15/hour; separately, TechCrunch reports Human Archive paying Indian collectors a ~$1/hour base rate, with competitors at roughly $2.60–4.20 per hour; DoorDash launched a standalone app, Tasks, paying gig workers to record everyday physical tasks; and Build AI paid 14,000+ factory workers across Southeast Asia to wear camera glasses on shift. The ethical tension is real and publicly debated: regulators in India are examining consent mechanisms, and the recurring framing in coverage is that these workers are filming the exact labor the robots are being trained to replace.

Is selling robot training data a good business?

The honest two-sided answer: demand is real and growing — three separately-funded startups (XDOF, Mecka, Micro1) all report frontier-lab customers, and Scale delivered over 150,000 hours of physical-AI data in 2025 — but the structural bears have receipts. The market is small — the only public estimate, from Micro1’s own CEO via MIT Technology Review, is over $100M per year, and even reading vendors’ own claims generously it is low hundreds of millions — versus $614M in reported gross revenue at a single text-data vendor in half of 2026. Prices are deflating, free corpora keep landing (Build AI’s 100,405-hour Apache-2.0 release; AgiBot’s million-trajectory open datasets), and XDOF’s own CEO has said simply creating data "is a poor business model," which is why vendors are racing up-stack into curation, annotation and evaluation. The winners likely own access to people and places, or own the quality layer — not raw hours.

Robotics Training Data: The Companies Feeding Physical AI (2026) | Teahose