Free public instrument from GAGE
The Embodied AI Data Ledger
As of September 2026 this ledger holds 37 records across 8 jurisdictions: 21 verified at a primary source, 6 reported by a secondary source, 2 announced with no document yet, 3 searched and absent, and 5 open questions. China carries 6 data supply records; the European Union carries 0 data supply records and 3 verified access rules.
- 37
- Records
- 21
- Verified at the source
- 5
- Announced or absent
- 5
- Open questions
8 jurisdictions, 74 primary sources
6 more reported, primary not reached
2 announced with no document, 3 searched and absent
Posed, sourced, not answered
As of 13 September 2026 the GAGE Embodied AI Data Ledger records 21 verified instruments, 6 reported at a secondary source, 2 announcements without a document, 3 absences and 5 open questions on the data that trains robots, across 8 jurisdictions. Counts are a floor, not a ceiling: a training ground or a standard the ledger has not found is not on it.
Why this ledger exists
The bottleneck moved from the model to the hours
Embodied AI models learn from recorded reality: hours of robots and people doing things, with every sensor captured. Who produces those hours, what standard says an hour is any good, and what the law lets a robot maker do with hours recorded in each place are three different questions, and they are being answered at three different speeds. This ledger keeps the record of all three, as verdicts about documents rather than opinions about robots.
Every row is fetched at its source where the source could be reached, and says so where it could not. The verdict language is the discipline: a reader learns five words once and never has to guess what a row claims.
Verified
The document exists. The ledger fetched it at its publisher and quotes it.
Reported, primary not reached
A reliable secondary source carries the figure or the program, and the primary document could not be reached. Printed with this label, never as verified.
Announced, no document yet
A body has said it will act. No document exists yet, so the row records the statement and nothing more.
Absent
The ledger searched and found no instrument. The record says where it looked and when.
Open question
No settled answer exists. The ledger poses the question, links the live debate, and does not answer it.
Side by side
Who builds, who regulates, who does neither
China
13 records
- Verified
- 5
- Reported, primary not reached
- 6
- Announced
- 1
- Absent
- 0
- Open question
- 1
6 supply, 5 standards, 1 rule, 1 open.
European Union
6 records
- Verified
- 3
- Reported, primary not reached
- 0
- Announced
- 1
- Absent
- 1
- Open question
- 1
1 supply, 1 standard, 3 rules, 1 open.
United States
7 records
- Verified
- 6
- Reported, primary not reached
- 0
- Announced
- 0
- Absent
- 1
- Open question
- 0
3 supply, 1 standard, 3 rules, 0 open.
United Kingdom
1 record
- Verified
- 0
- Reported, primary not reached
- 0
- Announced
- 0
- Absent
- 1
- Open question
- 0
0 supply, 1 standard, 0 rules, 0 open.
Germany
1 record
- Verified
- 1
- Reported, primary not reached
- 0
- Announced
- 0
- Absent
- 0
- Open question
- 0
1 supply, 0 standards, 0 rules, 0 open.
Japan
1 record
- Verified
- 1
- Reported, primary not reached
- 0
- Announced
- 0
- Absent
- 0
- Open question
- 0
1 supply, 0 standards, 0 rules, 0 open.
South Korea
1 record
- Verified
- 1
- Reported, primary not reached
- 0
- Announced
- 0
- Absent
- 0
- Open question
- 0
1 supply, 0 standards, 0 rules, 0 open.
Global
7 records
- Verified
- 4
- Reported, primary not reached
- 0
- Announced
- 0
- Absent
- 0
- Open question
- 3
2 supply, 2 standards, 0 rules, 3 open.
Figures of record
Every number on this ledger, with who measured it and when
47 figures, each one printed in the unit its publisher used, beside the publisher and the date it was true. Nothing here is summed across sources, converted between units, or forecast.
- 31 organizations
Drafting organizations listed on the platform entry. China has published a national guiding document on embodied AI data quality for real data, GB/Z 218.1-2026
National standards platform, State Administration for Market Regulation, entry for GB/Z 218.1-2026, primary source, as of .
- 11,499,999 euro
EU contribution to euROBIN, the closest network. No EU or member state program funds robot training data collection as its object
CORDIS, euROBIN project entry, primary source, as of .
- 25,000 dollars
Texas civil penalty ceiling per violation, per the mirror. US state biometric laws: Washington excludes video recordings from the definition, Texas needs consent for face geometry, and the Illinois text could not be reached
Texas Business and Commerce Code 503.001, mirror at texas.public.law, secondary source, as of .
- 320,000 trajectories
Trajectories in the physical AI dataset described on Hugging Face, at least. NVIDIA's synthetic supply: 780,000 trajectories in eleven hours, and an open simulated set of about 273,000
NVIDIA, physical AI dataset post on the Hugging Face blog, primary source, as of .
- 273,000 trajectories
Simulated trajectories in the open GR00T cross embodiment set, approximately. NVIDIA's synthetic supply: 780,000 trajectories in eleven hours, and an open simulated set of about 273,000
NVIDIA, PhysicalAI Robotics GR00T X Embodiment Sim dataset card, Hugging Face, primary source, as of .
- 7 companies
Private companies at the 29 August meeting that asked for embodied AI data infrastructure. China's data regulator says it will advance embodied AI data standards, twelve days after seven companies asked
National Data Administration, digital economy private enterprise symposium release, primary source, as of .
- 1,000,000 hours
Upper bound of usable high quality real data worldwide, per the report. The hours gap: a CAICT report puts the need on the order of ten million hours and the world's usable supply at under a million
Xinhua summary of the CAICT training ground report, carried by IT Home, secondary source, as of .
- 10,000,000 hours
Effective real data the report says one foundation model needs, order of magnitude. The hours gap: a CAICT report puts the need on the order of ten million hours and the world's usable supply at under a million
Xinhua summary of the CAICT training ground report, carried by IT Home, secondary source, as of .
- 40 organizations
Drafting organizations, at least, per the trade paper. An industry standard on embodied AI dataset quality, YD/T 6771-2026, takes effect on 1 November 2026, per a trade paper the ledger could not trace to the MIIT notice
People's Posts and Telecommunications News, reprinted on the Shenzhen government portal, secondary source, as of .
- 5 gyms
Locations the company named as its year end target. Europe's nearest equivalent to a training ground is one company: NEURA Robotics has ten gyms under development
NEURA Robotics, release on NEURA Gym and RWTH Aachen, primary source, as of .
- 10 gyms
Gyms under development, per the company. Europe's nearest equivalent to a training ground is one company: NEURA Robotics has ten gyms under development
NEURA Robotics, release on NEURA Gym and RWTH Aachen, primary source, as of .
- 86 percent
Share of grounds that include industrial manufacturing scenes. China has over 70 embodied AI training grounds in use, per a CAICT report the ledger could only reach through state media
Xinhua summary carried by IT Home, secondary source, as of .
- 46 training grounds
Under construction or planned, per the Xinhua summary; CCTV said forty odd. China has over 70 embodied AI training grounds in use, per a CAICT report the ledger could only reach through state media
Xinhua summary carried by IT Home, secondary source, as of .
- 70 training grounds
Training grounds built and in use, at least. China has over 70 embodied AI training grounds in use, per a CAICT report the ledger could only reach through state media
CCTV report reproduced on the Digital China government portal, secondary source, as of .
- 15 applications
Applications to the NEDO foundation model program. Japan funds a robotics data platform and a physical AI foundation model program through NEDO, with no published data volume
NEDO, call page for the AI robot and physical AI multimodal foundation model program, primary source, as of .
- 20,000,000 data items per year
Annual data output, approximately. Beijing's Shijingshan humanoid data training center reports close to 270 robots producing about 20 million data items a year
Beijing Daily, reprinted on the Beijing municipal government portal, secondary source, as of .
- 270 robots
Robot units training at full load, approximately. Beijing's Shijingshan humanoid data training center reports close to 270 robots producing about 20 million data items a year
Beijing Daily, reprinted on the Beijing municipal government portal, secondary source, as of .
- 25,000,000 euro
Indicative budget of the EuroHPC call to network AI Factory data labs. The EU promised data labs for the fourth quarter of 2025 and no page of its own names one as open
EuroHPC Joint Undertaking, call to strengthen the European AI ecosystem, primary source, as of .
- 88 percent
New work item approval vote, per the Beijing portal. ISO is drafting a humanoid robot datasets standard, ISO/CD 26264-1, at committee draft stage
Beijing science and technology portal, on the ISO new work item, secondary source, as of .
- 30 scenes
Typical scenes, at least. Beijing's humanoid robot innovation center runs a data and training base with over 120 robots in more than 30 scenes
Yicheng Times, reprinted on the Beijing E-Town government portal, secondary source, as of .
- 120 robots
Robot units at the base, at least. Beijing's humanoid robot innovation center runs a data and training base with over 120 robots in more than 30 scenes
Yicheng Times, reprinted on the Beijing E-Town government portal, secondary source, as of .
- 38,000,000 pounds
Robotics Adoption Hubs competition, up to. The United Kingdom funds robot adoption, 38 million pounds of it, and nothing on robot training data
Innovate UK, Robotics Adoption Hubs competition, primary source, as of .
- 40,000,000 pounds
Robotics adoption funding named in the action plan update. The United Kingdom funds robot adoption, 38 million pounds of it, and nothing on robot training data
GOV.UK, AI Opportunities Action Plan: one year on, primary source, as of .
- 6,000,000 data items per year
Annual data output reported eight months earlier, at least. Beijing's Shijingshan humanoid data training center reports close to 270 robots producing about 20 million data items a year
Beijing municipal science and technology portal, secondary source, as of .
- 30,000,000 euro
France 2030 robotics research program. No EU or member state program funds robot training data collection as its object
French Ministry of the Economy, robotics research program release, primary source, as of .
- 18 months
Drafting cycle stated in the plan notice. Two national standards for humanoid robot datasets are in drafting in China, on an eighteen month cycle from June 2025
National standards platform, plan 20253226-T-604, Humanoid robot datasets part 1, primary source, as of .
- 9 sectors
Sectors on that list. Open question: are a training ground's recordings important data under Chinese law? No catalog names robot data
Beijing Municipal Government, on the 2025 free trade zone data negative list, primary source, as of .
- 612 data fields
Data fields on Beijing's 2025 free trade zone negative list. Open question: are a training ground's recordings important data under Chinese law? No catalog names robot data
Beijing Municipal Government, on the 2025 free trade zone data negative list, primary source, as of .
- 6,500 hours
Human demonstration hours the company says that equals. Open question: do synthetic hours count? NVIDIA generated the equivalent of 6,500 human hours in eleven, and no standard says what an hour is
NVIDIA newsroom, Isaac GR00T N1 release, primary source, as of .
- 780,000 trajectories
Synthetic trajectories generated in eleven hours, per the press release. Open question: do synthetic hours count? NVIDIA generated the equivalent of 6,500 human hours in eleven, and no standard says what an hour is
NVIDIA newsroom, Isaac GR00T N1 release, primary source, as of .
- 780,000 trajectories
Synthetic trajectories generated in eleven hours, per the press release. NVIDIA's synthetic supply: 780,000 trajectories in eleven hours, and an open simulated set of about 273,000
NVIDIA newsroom, Isaac GR00T N1 release, primary source, as of .
- 100 robots
Robots that recorded it. The largest open real robot dataset is Chinese, over a million trajectories, and licensed for non commercial use only
AgiBot World Beta dataset card, Hugging Face, primary source, as of .
- 2,976.4 hours
Total duration of AgiBot World Beta. The largest open real robot dataset is Chinese, over a million trajectories, and licensed for non commercial use only
AgiBot World Beta dataset card, Hugging Face, primary source, as of .
- 1,003,672 trajectories
Trajectories in AgiBot World Beta. The largest open real robot dataset is Chinese, over a million trajectories, and licensed for non commercial use only
OpenDriveLab, AgiBot World repository, primary source, as of .
- 100 robots
Heterogeneous robots deployed in the first phase, at least. Shanghai's Zhangjiang humanoid training ground opened its first phase with over 100 heterogeneous robots
Liberation Daily, reprinted on the Shanghai municipal government portal, secondary source, as of .
- 209,403 videos
Grasp videos in the AI Hub dataset, approximately. South Korea's public robot dataset: 209,403 grasp videos that only Korean nationals may request, for non commercial use
AI Hub, robot behaviour data for small object grasping, primary source, as of .
- 10,000 people
Head count of sensitive personal information at or above which an assessment is required. Hours recorded in China leave the country under the 2024 cross border provisions and the 2025 network data regulations
Cyberspace Administration of China, Provisions on Promoting and Regulating Cross Border Data Flows, primary source, as of .
- 1,000,000 people
Head count of personal information at or above which an outbound security assessment is required. Hours recorded in China leave the country under the 2024 cross border provisions and the 2025 network data regulations
Cyberspace Administration of China, Provisions on Promoting and Regulating Cross Border Data Flows, primary source, as of .
- 564 scenes
Scenes. DROID: 76,000 demonstration trajectories, 350 hours, one robot arm, under CC BY 4.0
DROID project page, primary source, as of .
- 350 hours
Hours of interaction data. DROID: 76,000 demonstration trajectories, 350 hours, one robot arm, under CC BY 4.0
DROID project page, primary source, as of .
- 76,000 trajectories
Demonstration trajectories. DROID: 76,000 demonstration trajectories, 350 hours, one robot arm, under CC BY 4.0
DROID project page, primary source, as of .
- 1,286.3 hours
Hours in Ego-Exo4D. The largest open supply of hours is human, not robot: Ego4D holds over 3,670 hours of first person video
Ego-Exo4D project site, primary source, as of .
- 527 skills
Open X-Embodiment paper, arXiv 2310.08864, primary source, as of .
- 22 embodiments
Robot embodiments. Open X-Embodiment: over a million real robot trajectories from 22 embodiments, pooled from 60 datasets, under CC BY 4.0
Open X-Embodiment paper, arXiv 2310.08864, primary source, as of .
- 1,000,000 trajectories
Real robot trajectories, at least, per the project page. Open X-Embodiment: over a million real robot trajectories from 22 embodiments, pooled from 60 datasets, under CC BY 4.0
Open X-Embodiment project page, primary source, as of .
- 3,670 hours
Hours in Ego4D, at least. The largest open supply of hours is human, not robot: Ego4D holds over 3,670 hours of first person video
Ego4D project site, primary source, as of .
- 100 hours
Hours in EPIC-KITCHENS-100. The largest open supply of hours is human, not robot: Ego4D holds over 3,670 hours of first person video
EPIC-KITCHENS project site, primary source, as of .
The ledger
Every record, newest first
37 records. Each row opens a page carrying the answer, the verdict and what it means, the key facts, the figures with their sources, what it changes for a robot maker, and the sources it was verified against.
Answers
What people ask about embodied AI data
Who is collecting the data that trains robots?
14 data supply records are on the ledger as of September 2026, counted by jurisdiction: China 6, European Union 1, United States 3, Germany 1, Japan 1, South Korea 1, Global 2. China has over 70 embodied AI training grounds in use, per a CAICT report the ledger could only reach through state media. The largest open real robot dataset is Chinese and non commercial, the largest open pool of hours is human first person video, and no public program in the European Union or the United Kingdom records robot data at all.
How many hours of robot training data exist?
The same August 2026 training ground report from the China Academy of Information and Communications Technology estimates that an embodied AI foundation model of about 55 billion parameters would need effective real data on the order of ten million hours, while usable high quality real data worldwide is on the order of hundreds of thousands to a million hours. The ledger reached the figures through a Xinhua summary, not the report. Every figure the ledger prints is on the figures of record list on this page with its publisher and its date, and none is summed across sources.
Is there a standard for embodied AI training data?
10 standard records: The United Kingdom funds robot adoption, 38 million pounds of it, and nothing on robot training data (absent); No United States federal standard or rule governs robot training data; one NIST project names datasets as a deliverable (absent); The published international standards on data quality for machine learning, the ISO/IEC 5259 series, are general and name no robot (verified); ISO is drafting a humanoid robot datasets standard, ISO/CD 26264-1, at committee draft stage (verified); The EU promised data labs for the fourth quarter of 2025 and no page of its own names one as open (announced, no document yet); An industry standard on embodied AI dataset quality, YD/T 6771-2026, takes effect on 1 November 2026, per a trade paper the ledger could not trace to the MIIT notice (reported, primary not reached); Two national standards for humanoid robot datasets are in drafting in China, on an eighteen month cycle from June 2025 (verified); China has published a national guiding document on embodied AI data quality for real data, GB/Z 218.1-2026 (verified); China's June 2026 dataset action plan names real machine interaction data for embodied AI (verified); China's data regulator says it will advance embodied AI data standards, twelve days after seven companies asked (announced, no document yet). The international work item is at committee draft stage and led from China; the published international series on data quality for machine learning is general and names no robot.
What does the EU Data Act mean for robots?
Regulation (EU) 2023/2854 has applied since 12 September 2025. It gives the user of a connected product a right to the data it generates and to share it with a third party, bars that third party from using the data to build a competing product, and requires products placed on the market after 12 September 2026 to be designed so the data is accessible. The Commission's own explainer names robots among connected products.
Can Chinese training ground data train a robot sold in Europe?
Nobody has settled it, which is why it is an open question on this ledger rather than a rule. 5 open questions are recorded: does a person in a training ground's footage have a say? Three regimes, three different answers, none written for robots; who owns the hours a robot records at a customer's site? Only the EU has a rule, and it is about access, not ownership; do synthetic hours count? NVIDIA generated the equivalent of 6,500 human hours in eleven, and no standard says what an hour is; can hours recorded in a Chinese training ground train a robot sold in the European Union?; are a training ground's recordings important data under Chinese law? No catalog names robot data. Each one links the documents that would answer it if any did.
What is an embodied AI training ground?
A site, usually publicly funded, where many robots of different bodies perform scripted and unscripted tasks in built scenes while every sensor is recorded, to produce the hours of real interaction data that embodied AI models learn from. China counts them in the dozens and publishes a national report on them; Europe's nearest equivalent is a company's gym; the United States funds project calls for data rather than sites.
What do the verdicts mean?
Verified: the document exists and the ledger fetched it at its publisher. Reported: a reliable secondary source carries it and the primary could not be reached. Announced: a body said it will act and no document exists. Absent: the ledger searched and found nothing, and the search is written into the record. Open: a question nobody has settled, posed and not answered.
Every surface
Cut the ledger the way you need it
By jurisdiction
By kind
Every record page
- EAD-2026-0037: The United Kingdom funds robot adoption, 38 million pounds of it, and nothing on robot training data
- EAD-2026-0036: South Korea's public robot dataset: 209,403 grasp videos that only Korean nationals may request, for non commercial use
- EAD-2026-0035: Japan funds a robotics data platform and a physical AI foundation model program through NEDO, with no published data volume
- EAD-2026-0034: NVIDIA's synthetic supply: 780,000 trajectories in eleven hours, and an open simulated set of about 273,000
- EAD-2026-0033: DROID: 76,000 demonstration trajectories, 350 hours, one robot arm, under CC BY 4.0
- EAD-2026-0032: No US export rule names robot training data; the AI diffusion rule that reached model weights is under rescission
- EAD-2026-0031: US state biometric laws: Washington excludes video recordings from the definition, Texas needs consent for face geometry, and the Illinois text could not be reached
- EAD-2026-0030: A Senate bill, not a law: the Humanoid ROBOT Act of 2025 would order a report on data of US persons obtained through humanoid robots
- EAD-2026-0029: The US answer to a training ground is a manufacturing institute's data call: the ARM Institute's AI Data Foundry
- EAD-2026-0028: No United States federal standard or rule governs robot training data; one NIST project names datasets as a deliverable
- EAD-2026-0027: Open question: does a person in a training ground's footage have a say? Three regimes, three different answers, none written for robots
- EAD-2026-0026: Open question: who owns the hours a robot records at a customer's site? Only the EU has a rule, and it is about access, not ownership
- EAD-2026-0025: Open question: do synthetic hours count? NVIDIA generated the equivalent of 6,500 human hours in eleven, and no standard says what an hour is
- EAD-2026-0024: The largest open supply of hours is human, not robot: Ego4D holds over 3,670 hours of first person video
- EAD-2026-0023: Open X-Embodiment: over a million real robot trajectories from 22 embodiments, pooled from 60 datasets, under CC BY 4.0
- EAD-2026-0022: The published international standards on data quality for machine learning, the ISO/IEC 5259 series, are general and name no robot
- EAD-2026-0021: ISO is drafting a humanoid robot datasets standard, ISO/CD 26264-1, at committee draft stage
- EAD-2026-0020: Europe's nearest equivalent to a training ground is one company: NEURA Robotics has ten gyms under development
- EAD-2026-0019: Open question: can hours recorded in a Chinese training ground train a robot sold in the European Union?
- EAD-2026-0018: The AI Act reaches robots as machinery through the old Machinery Directive, and learned safety behaviour needs third party assessment from January 2027
- EAD-2026-0017: Under the GDPR, factory footage of a worker is personal data, and biometric data only when processed to identify them
- EAD-2026-0016: No EU or member state program funds robot training data collection as its object
- EAD-2026-0015: The EU promised data labs for the fourth quarter of 2025 and no page of its own names one as open
- EAD-2026-0014: The EU Data Act reaches robots as connected products: user access since September 2025, access by design for products placed on the market after 12 September 2026
- EAD-2026-0013: The largest open real robot dataset is Chinese, over a million trajectories, and licensed for non commercial use only
- EAD-2026-0012: Open question: are a training ground's recordings important data under Chinese law? No catalog names robot data
- EAD-2026-0011: Hours recorded in China leave the country under the 2024 cross border provisions and the 2025 network data regulations
- EAD-2026-0010: Shanghai's Zhangjiang humanoid training ground opened its first phase with over 100 heterogeneous robots
- EAD-2026-0009: Beijing's Shijingshan humanoid data training center reports close to 270 robots producing about 20 million data items a year
- EAD-2026-0008: Beijing's humanoid robot innovation center runs a data and training base with over 120 robots in more than 30 scenes
- EAD-2026-0007: An industry standard on embodied AI dataset quality, YD/T 6771-2026, takes effect on 1 November 2026, per a trade paper the ledger could not trace to the MIIT notice
- EAD-2026-0006: Two national standards for humanoid robot datasets are in drafting in China, on an eighteen month cycle from June 2025
- EAD-2026-0005: China has published a national guiding document on embodied AI data quality for real data, GB/Z 218.1-2026
- EAD-2026-0004: The hours gap: a CAICT report puts the need on the order of ten million hours and the world's usable supply at under a million
- EAD-2026-0003: China has over 70 embodied AI training grounds in use, per a CAICT report the ledger could only reach through state media
- EAD-2026-0002: China's June 2026 dataset action plan names real machine interaction data for embodied AI
- EAD-2026-0001: China's data regulator says it will advance embodied AI data standards, twelve days after seven companies asked
Take the data
The whole dataset, free, in two formats
Licensed CC BY 4.0. Use it in an article, a paper, a slide or a product. The only condition is attribution, and the citation page gives you the line to paste.
- ledger.jsonEvery field of every record, the shape documented on the data page.
- ledger.csvOne row per record, figures flattened, for a spreadsheet or a stats package.
How the ledger is built, what the verdicts mean, and what the gate refuses: the method page. Every change, dated: the changelog. The kinds on the shelf: data supply, data standard, access rule, open question. Something missing or wrong is a bug, and we want to hear about it. The jobs that name these standards are counted on the Physical AI Hiring Index, and the safety standards stack for the robots themselves is on the Robot Compliance Navigator.
Cite this page
Free to reuse under CC BY 4.0, with attribution.
- In a sentence
- According to the GAGE Embodied AI Data Ledger (as of 13 September 2026), embodied ai data ledger.
- APA
- GAGE (Global Academy of Generative-AI Education). (2026). Embodied AI Data Ledger. Embodied AI Data Ledger. Retrieved 13 September 2026, from https://www.gage.academy/tools/embodied-ai-data-ledger
- MLA
- "Embodied AI Data Ledger." Embodied AI Data Ledger, GAGE (Global Academy of Generative-AI Education), 13 September 2026, https://www.gage.academy/tools/embodied-ai-data-ledger.
- Chicago
- GAGE (Global Academy of Generative-AI Education). "Embodied AI Data Ledger." Embodied AI Data Ledger. Last modified 13 September 2026. https://www.gage.academy/tools/embodied-ai-data-ledger.
- Permalink
- https://www.gage.academy/tools/embodied-ai-data-ledger
Last updated . Every record re verified . The ledger is checked every Monday, and the same day for any announcement by a data regulator.
31 organizations
GAGE briefings tell you which AI regulation deadlines are coming, what they actually require of you, and when a program opens.