ACE-Data-0 Turns Real Homes Into Data Engines for Embodied AI
Introduction
Embodied AI does not only need better models; it needs data that reflects how people actually act in the physical world. A system that learns household behavior must connect what a person sees, how the body moves, how the hands manipulate objects, how objects change state, and when contact, sound or touch occurs. ACE-Data-0 addresses this gap by capturing these signals together rather than treating them as separate fragments.
Key points
- A human-centric capture engine: The paper introduces the Ambient Capture Engine, or ACE, which turns real home environments into spatially calibrated and temporally synchronized recording spaces.
- Two complementary scales: ACE uses a table-scale setup for detailed hand-object manipulation and a room-scale setup for whole-body motion, locomotion and interaction across a furnished home.
- Synchronized multimodal streams: The dataset includes egocentric video, multi-view exocentric video, full-body motion, articulated hand motion, object geometry and 6-DoF trajectories, audio and tactile signals.
- Dataset scale: ACE-Data-0 contains 150 hours of recordings, 17 million video frames, 200 task categories, 50 participants, 2 environments and 75,000 interaction episodes.
- Natural variation by design: Instead of giving step-by-step instructions, the data collection uses goal-level instructions, allowing participants to complete tasks in their own ways.
Why it matters
Many embodied AI datasets observe only part of the perception-action loop. Some focus on hand manipulation but miss whole-body context; others provide external videos but lack first-person perception; still others record motion while leaving contact and object state under-specified. ACE-Data-0 is important because it aligns these pieces along the same timeline, making it possible to study how perception, kinematics and contact evolve together.
The paper also introduces a hierarchical benchmark that moves from low-level signals to scene components and then to interactions. According to the authors, evaluations of current state-of-the-art methods reveal substantial weaknesses under contact, occlusion, egomotion and long temporal horizons. In that sense, the dataset is not only a training resource but also a stress test for whether embodied models can handle realistic physical interaction.
For imitation learning, world models, vision-language-action systems and home robotics, ACE-Data-0 offers a more complete form of human demonstration data. Its current scope—50 participants in 2 environments—also makes clear that this is a starting point rather than an endpoint. The next question is whether such high-quality synchronized capture can scale to more homes, more objects and more open-ended tasks.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...