Yukang Cao1*, Haozhe Xie1*, Beichen Wen2*, Runmao Yao1, Yinghao Liu2, Yue Huang2, Zhichao Liao2, Yunxiang Wang1, Haiheng Liu1, Xingshun Tian2, Dawei Su2, Long Zhuo2, Dacheng Tao2, Xiaogang Wang2, Liang Pan2, Ziwei Liu1
1 S-Lab, Nanyang Technological University 2 ACE Robotics
* Equal Contributions Project Advisor Project Lead
1 Introduction
Humans spend the majority of their lives in built environments, continuously interacting with the objects around them, such as opening cabinets, pouring water, folding laundry, and assembling furniture. Effortless as they appear, these everyday interactions embody a form of intelligence that existing models have yet to attain: the seamless coordination of perception, whole-body movement, dexterous manipulation, and physical sensing [30, 74]. Reproducing this intelligence is a central pursuit of embodied AI, yet unlike the AI breakthroughs that preceded it, this pursuit inherits no data. Language and vision models [20, 83, 80, 8] rose on archives that humanity had spent centuries accumulating. Physical skills, exercised without deliberation, have never been written down: how a hand closes around a cup, with what force a fragile glass is held, by what coordination of vision and balance an object is carried across a room. The data for embodied intelligence must therefore be built rather than found, by instrumenting everyday life itself and recording human-object interaction (HOI) as it naturally unfolds.
Constructing such a dataset, however, demands far more than pointing a camera at daily life. It requires a holistic recording of the interaction, spanning what the human is doing, how the body and hands move, how the object state evolves in response, what audio and tactile signals the interaction produces, and how the scene appears from both the human’s first-person perspective and third-person viewports. More importantly, such a dataset for learning embodied AI demands all of these signals to be recorded synchronously and in registration with one another, a requirement that no existing dataset satisfies.
Specifically, existing datasets fall short in three aspects, as summarized in Table 1: (1) Fragmented modalities. Large-scale egocentric datasets (e.g., Ego-Exo4D [34], EPIC-Kitchens [19], Xperience-10M [82]) offer naturalistic behavior at scale, yet provide neither ground-truth body and object motion nor synchronized third-person observation. Conversely, motion-captured HOI datasets, such as BEHAVE [5], GRAB [88], ARCTIC [22], OakInk2 [110], and HOT3D [3], supply accurate poses but omit the egocentric perspective. More importantly, audio or tactile sensing is absent from nearly all of them. (2) Unnatural environments. Physically annotated datasets [7, 15, 88, 22, 26] are captured almost exclusively in laboratories, whose sparse layouts eliminate precisely the occlusions, spatial constraints, and object diversity that render real homes challenging. (3) Short horizons. The vast majority of HOI clips span seconds and depict one simple movement [7, 15, 52, 22]. Genuine household activities, in contrast, are goal-directed. They may unfold over minutes or hours, chain together sub-tasks and multiple objects, and require the human to move across the scene rather than acting in a single place. Consequently, the complete perception-action loop of everyday interaction has remained beyond the reach of existing benchmarks.
To bridge this gap, we develop the Ambient Capture Engine (ACE), a capture system that turns real home environments into recording studios while preserving their lived-in realism. Fine-grained dexterous manipulation and room-scale activity impose conflicting requirements on sensor placement and coverage, so we build ACE as two complementary configurations. The first is table-scale: densely arranged close-range cameras and high-resolution tactile sensing resolve the local details of hand-object manipulation. The second is room-scale: sensors span an entire furnished environment to record whole-body motion, locomotion, and interactions distributed across the scene. While differing in sensor placement and spatial scope, both scales record synchronized egocentric video, multi-view exocentric video, optical full-body motion, and per-object 6-DoF trajectories, which are all registered into a common spatio-temporal frame.
With ACE, we collect ACE-Data-0: 150 hours of daily living interactions spanning 200 task categories (e.g., cooking, tidying, and drinking), performed by 50 participants across 2 environments, amounting to 17M video frames and 75,000 interaction episodes. Rather than executing scripts, participants pursue goal-level instructions in their own manner: planning, hesitating, and improvising as they would at home. Crucially, this freedom sacrifices no measurement fidelity: every moment of activity is observed simultaneously from more than 8 viewpoints and grounded in metric body, hand, and object states. Upon these raw signals, we provide rich annotations for every sequence, including camera calibrations and synchronized timelines; full-body and hand poses and their projections onto every frame of every camera; per-object mesh models, 6-DoF poses, bounding boxes, and motion trails; and textual descriptions of the ongoing events. A large fraction of these annotations is derived automatically from the tracked physical states, instead of being estimated via existing pipelines.
Building upon ACE-Data-0, we further establish a perceptually hierarchical benchmark, comprising three tracks that advance from signals, to components, to interactions. The first track targets low-level signals: cross-modal prediction of tactile signals from visual observations. The second track focuses on scene components: methods reconstruct the pose of the human body and hands, evaluated against our tracked ground-truth. The third track evaluates the interactions: models estimate the hand motion during human-object interactions from egocentric and exocentric videos. Additionally, our collected paired viewpoints allow for a direct comparison between these two. Evaluations of more than 30 representative methods expose failure modes characteristic of long-horizon home-scene data, stemming from the diversity of interactions and tasks, their complexity and interleaving, heavy occlusion, extreme viewpoints, and continual movement through the scene. Notably, these three tracks mirror the perceptual capabilities an embodied agent must chain together: sensing contact, estimating scene state, and mastering hand-object interactions. The value of ACE-Data-0 to such agents, moreover, extends beyond evaluation to training itself: because our egocentric viewpoint, multi-view exocentric videos, motion trajectories, and contact-level supervision are synchronized rather than assembled from disparate sources, every training signal refers to the same physical moment, a property essential for robot learning, from imitation and policy learning to world modeling [104, 102, 6, 87].
In general, our contributions can be summarized as follows:
- •
We present ACE, a “human-centric ambient capture as embodied data engine paradigm” realized as two complementary configurations, a table-scale setup for fine-grained local manipulation and a room-scale setup for whole-body motion in larger scenes. Within these two configurations, ACE records both temporally and spatially aligned egocentric video, multi-view exocentric video, human body and hand motion, object motion, audio, and tactile signals.
- •
We contribute ACE-Data-0, a large-scale, long-horizon home-scene HOI dataset comprising 17M frames and 75,000 interaction episodes, together with rich and high-quality annotations.
- •
We establish a three-level benchmark that advances from signals to components and ultimately to interactions, and provide evaluations of more than 30 state-of-the-art methods that present the open challenges for embodied perception and robot learning.
| Dataset | Year | Hours | #Frames | #Subj | #Obj | #Tasks | Ego | #Exo | Body | Hand | Obj. 6D | Tactile | Audio | Sync | Setup | LH |
| Egocentric video | ||||||||||||||||
| Ego4D [33] | 2022 | 3670 | – | 931 | – | Open | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | – | In-the-wild | ✓ |
| EPIC-KITCHENS-100 [19] | 2021 | 100 | 20M | 37 | – | Open | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | – | Kitchens | ✓ |
| HoloAssist [95] | 2023 | 166 | – | 222 | 16 | 20 | ✓ | ✗ | ✗ | (✓) | ✗ | ✗ | ✓ | – | Desktop | ✓ |
| Ego-Exo4D [34] | 2024 | 1286 | – | 740 | – | 8 Dom. | ✓ | 4–5 | (✓) | ✗ | ✗ | ✗ | ✓ | ✓ | In-the-wild | ✓ |
| EgoLife [100] | 2025 | 266 | – | 6 | – | Open | ✓ | ✓ | ✗ | ✗ | ✗ | ✗ | ✓ | ✗ | Shared house | ✓ |
| HD-EPIC [73] | 2025 | 41 | 4.46M | 9 | – | 69 | ✓ | ✗ | ✗ | ✗ | (✓) | ✗ | ✓ | – | Home kitchens | ✓ |
| Hand–object interaction | ||||||||||||||||
| ContactPose [7] | 2020 | – | 2.9M | 50 | 25 | 2 | ✗ | 3 | ✗ | ✓ | ✓ | ✓ | ✗ | ✗ | Table-top | ✗ |
| HO-3D [36] | 2020 | – | 78K | 10 | 10 | – | ✗ | 1–5 | ✗ | ✓ | ✓ | ✗ | ✗ | ✗ | Table-top | ✗ |
| GRAB [88] | 2020 | – | 1.6M | 10 | 51 | 4 | ✗ | ✗ | ✓ | ✓ | ✓ | (✓) | ✗ | – | Mocap lab | ✗ |
| DexYCB [15] | 2021 | – | 582K | 10 | 20 | 1 | ✗ | 8 | ✗ | ✓ | ✓ | ✗ | ✗ | ✗ | Table-top | ✗ |
| H2O [52] | 2021 | – | 571K | 4 | 8 | 36 | ✓ | 4 | ✗ | ✓ | ✓ | ✗ | ✗ | ✓ | Table-top | ✗ |
| OakInk [101] | 2022 | – | 230K | 12 | 100 | 5 | ✗ | 4 | ✗ | ✓ | ✓ | ✗ | ✗ | ✓ | Table-top | ✗ |
| HOI4D [61] | 2022 | 22 | 2.4M | 9 | 800 | 54 | ✓ | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ | – | Indoor rooms | ✗ |
| ARCTIC [22] | 2023 | – | 2.1M | 10 | 11 | 2 | ✓ | 8 | ✓ | ✓ | ✓ | ✗ | ✗ | ✓ | Mocap lab | ✗ |
| TACO [60] | 2024 | – | 5.2M | 14 | 196 | 151 | ✓ | 12 | ✗ | ✓ | ✓ | ✗ | ✗ | ✓ | Table-top | ✗ |
| HOT3D [3] | 2024 | 13.9 | 3.7M | 19 | 33 | Open | ✓ | ✗ | ✗ | ✓ | ✓ | ✗ | ✗ | – | Lab rooms | ✗ |
| OakInk2 [110] | 2024 | – | 4.0M | 9 | 75 | 150 | ✓ | 3 | (✓) | ✓ | ✓ | ✗ | ✗ | ✓ | Table-top | (✓) |
| GigaHands [26] | 2025 | 34 | 183M | 56 | 417 | Open | ✗ | 51 | ✗ | (✓) | (✓) | ✗ | ✗ | ✗ | Table-top | ✗ |
| Full-body HOI, human-scene interaction, and daily motion | ||||||||||||||||
| BEHAVE [5] | 2022 | – | 15K | 8 | 20 | – | ✗ | 4 | ✓ | ✗ | ✓ | ✗ | ✗ | ✓ | Lab rooms | ✗ |
| InterCap [41] | 2022 | – | 67K | 10 | 10 | – | ✗ | 6 | ✓ | (✓) | ✓ | ✗ | ✗ | ✓ | Lab room | ✗ |
| CHAIRS [44] | 2022 | 17.3 | – | 46 | 81 | 32 | ✗ | 4 | ✓ | ✓ | ✓ | ✗ | ✗ | ✓ | Lab room | – |
| EgoBody [113] | 2022 | – | 220K | 36 | – | 5 Cat. | ✓ | 3–5 | (✓) | (✓) | ✗ | ✗ | ✗ | ✓ | Indoor rooms | ✗ |
| Aria Digital Twin [69] | 2023 | 6.6 | – | – | 398 | Open | ✓ | ✗ | (✓) | ✗ | ✓ | ✗ | ✓ | – | Apartment, office | ✗ |
| OMOMO [54] | 2023 | 10 | – | 17 | 15 | – | ✗ | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ | – | Lab room | ✗ |
| HIMO [65] | 2024 | – | 4.1M | 34 | 53 | – | ✓ | ✗ | ✓ | ✓ | ✓ | ✗ | ✗ | – | Mocap lab | ✗ |
| TRUMANS [45] | 2024 | 15 | 1.6M | 7 | 20 | – | ✗ | ✗ | ✓ | ✗ | (✓) | ✗ | ✗ | – | Scene mockups | ✗ |
| ParaHome [50] | 2024 | 8.1 | – | 38 | 22 | – | ✗ | 70 | ✓ | ✓ | ✓ | ✗ | ✗ | ✓ | Home room | ✓ |
| Nymeria [66] | 2024 | 300 | – | 264 | – | 20 Scen. | ✓ | ✓ | ✓ | ✗ | ✗ | ✗ | ✓ | ✓ | In-the-wild | ✓ |
| HuMoTo [64] | 2025 | 2.2 | 236K | 1 | 63 | – | ✗ | ✗ | ✓ | ✓ | ✓ | ✗ | ✗ | – | Scene | ✗ |
| Robot data and human demonstrations for robots | ||||||||||||||||
| BridgeData V2 [93] | 2023 | – | 60K Traj | – | 100+ | 13 Skills | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | – | Table-top | ✗ |
| RH20T [23] | 2023 | – | 110K Traj | – | – | 147 | ✓ | ✓ | ✗ | ✗ | ✗ | (✓) | ✓ | ✓ | Table-top | ✗ |
| DexCap [94] | 2024 | – | – | – | – | – | ✗ | ✓ | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | Table-top | ✗ |
| DROID [49] | 2024 | 350 | 76K Traj | – | – | 86 | ✗ | 3 | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | In-the-wild | ✗ |
| AgiBot World [1] | 2025 | 2976.4 | 1M Traj | – | 3000+ | 217 | ✗ | ✓ | ✗ | ✗ | ✗ | ✓ | ✗ | ✗ | Staged scenes | ✓ |
| EgoDex [38] | 2025 | 829 | 90M | – | – | 194 | ✓ | ✗ | ✗ | (✓) | ✗ | ✗ | ✗ | – | Table-top | ✗ |
| Galaxea Open-World [46] | 2025 | 500 | 100K Traj | – | 1600 | 150 | ✗ | ✓ | ✗ | ✗ | ✗ | ✗ | ✗ | ✓ | Open-world | (✓) |
| EgoScale [114] | 2026 | 20,854 | – | – | 43,237 | 6015 | ✓ | ✗ | – | – | – | ✗ | ✗ | – | Table-top | ✗ |
| Xperience-10M [82] | 2026 | 10,000 | 2.88B | – | – | – | ✓ | ✗ | (✓) | (✓) | ✗ | ✗ | ✓ | – | In-the-wild | ✗ |
| Ours | ||||||||||||||||
| ACE-Data-0 | 2026 | 150 | 17M | 50 | 50 | 200 | ✓ | 8 | ✓ | ✓ | ✓ | ✓ | ✓ | ✓ | Home (Room + Table-top) | ✓ |
2 Related Work
2.1 Multi-modal Datasets & Benchmarks
Everyday interaction is inherently multi-modal. A single physical event is reflected simultaneously in visual appearance, body and hand motion, changes in object state, acoustic cues, and physical contact. These signals are complementary rather than interchangeable: vision captures what is externally observable, kinematics describes how the human and objects evolve, audio reveals impacts and state transitions, and touch provides direct evidence of contact and force. Accordingly, datasets for embodied perception have progressively moved beyond isolated RGB observations toward richer combinations of geometry, motion, contact, and semantic annotations.
Early physically grounded datasets primarily focused on the local geometry of grasping and manipulation. ContactPose [7] associates 3D hand pose with object-surface contact over approximately 2.9M multi-view RGB-D images from 50 participants and 25 objects, while DexYCB [15] provides 582K frames from eight synchronized RGB-D cameras with MANO hand [81] annotations and YCB object poses. H2O [52] extends this setting to bimanual interaction through synchronized first- and third-person RGB-D observations, 3D hand keypoints, and object poses. H2O-3D [37] further contributes more than 76K annotated images for challenging two-hand-object pose estimation under severe self-occlusion. At a larger semantic scale, HOI4D [61] contains 2.4M egocentric RGB-D frames involving roughly 800 object instances and supports category-level object pose tracking, action segmentation, and dynamic 4D interaction understanding. A complementary line of work expands the captured state from the hands to the full human-object system. GRAB [88] records detailed whole-body, hand, face, and object motion for interactions with 51 objects. BEHAVE [5] and InterCap [41] jointly recover humans and rigid objects from multi-view RGB-D observations, enabling the study of body-scale coordination, contact, and joint reconstruction.
More recent datasets increasingly emphasize dexterity, compositionality, and task structure. ARCTIC [22] captures approximately 2.1M images of bimanual interaction with articulated objects from one egocentric and eight allocentric views, grounded by optical motion capture. TACO [60] organizes 5.2M frames around compositional tool-action-object relationships, whereas OakInk2 [110] represents complex bimanual activity through a hierarchy of affordances, motion primitives, and task-level structures. At a larger scale, GigaHands [26] collects roughly 183M frames from 56 participants and 417 objects using a 51-camera system, substantially increasing the diversity of bimanual activities and language annotations. HOT3D [3] provides approximately 3.7M egocentric multi-view images from 19 participants and 33 scanned objects, together with metric trajectories for hands, objects, cameras, and gaze. Collectively, these datasets have substantially advanced hand pose estimation [72, 112, 105], human-object reconstruction [41, 17], contact reasoning [7, 22, 68], and interaction understanding [60, 26, 3]. Their sensing configurations, however, are typically optimized for a particular spatial scale or source of supervision. Hand-centric datasets resolve detailed local geometry but usually confine activity to a compact workspace, whereas whole-body datasets preserve global coordination but may offer less detailed hand articulation or omit the actor’s first-person view. Existing multi-modal resources also predominantly combine RGB, depth, pose, geometry, gaze, and language [12]. Audio is rarely aligned with metric human and object motion at the interaction level, and tactile sensing is scarcer still, despite directly revealing contact onset, pressure, and slip. Moreover, annotations may combine direct measurement with offline fitting or model-based reconstruction, limiting their uniformity under sustained occlusion, rapid motion, or long-duration activity. High-fidelity capture is therefore often achieved in controlled spaces that simplify the clutter, furniture occlusion, and spatial constraints of domestic interaction.
ACE-Data-0 complements these efforts by treating synchronized multisensory grounding as the central unit of data. The table-scale configuration resolves fine-grained hand-object manipulation, whereas the room-scale configuration captures full-body motion and interactions distributed across a furnished domestic environment. Both configurations provide egocentric video, multi-view exocentric video, full-body and articulated hand motion, object 6-DoF trajectories, audio, and tactile signals through a shared synchronization and calibration pipeline. The modalities therefore describe the same physical event on a common timeline and in a common spatial frame, rather than serving as independently produced annotations. This unified grounding supports a coherent benchmark progression from low-level signal inference, through scene component recovery, to interaction understanding.
2.2 Egocentric Datasets & Benchmarks
Egocentric video observes activity from the viewpoint of the acting subject rather than that of an external camera. This perspective naturally emphasizes action-relevant objects, hand proximity, gaze allocation, and the immediate visual consequences of self-motion, making it particularly relevant to embodied agents. At the same time, it introduces distinctive challenges: the camera moves continuously, motion blur is frequent, the hands often occlude manipulated objects, and much of the actor’s body remains outside the field of view.
Large-scale datasets first established the breadth and naturalism of egocentric perception. Ego4D [33] collects 3,670 hours of unscripted first-person video from 931 camera wearers across 74 locations, with benchmarks spanning episodic memory, forecasting, social interaction, manipulation, and audio-visual understanding. HoloAssist [95] focuses on interactive task assistance, recording 166 hours of instructor-performer activity with synchronized RGB-D, gaze, head and hand pose, IMU, audio, and dialogue; these signals support mistake detection, intervention prediction, and hand-motion forecasting. Ego-Exo4D [34] further connects first- and third-person perspectives through 1,286 hours of synchronized skilled activity from 740 participants across 123 scenarios, together with audio, gaze, language, camera poses, and 3D scene information. Its paired-view design enables cross-view activity understanding, proficiency estimation, and pose recovery on the same underlying events.
Manipulation-oriented datasets trade some behavioral breadth for stronger physical supervision. H2O [52] and HOI4D [61] associate first-person RGB-D observations with hand and object states, while HOT3D [3] provides calibrated egocentric multi-view recordings with metric trajectories for hands, objects, cameras, and gaze. EgoDex [38] scales human dexterous demonstrations to 829 hours across 194 task types using wearable hand tracking, creating a large source of action-relevant human manipulation data. PH2D [79] further connects human first-person demonstrations with humanoid policy learning through a robot-compatible action representation. A complementary source of first-person experience comes from robot demonstration data. Fourier ActionNet [25] records more than 30K teleoperated trajectories, totaling approximately 140 hours of dexterous bimanual manipulation across multiple humanoid platforms; it pairs observations from the robot’s first-person cameras with language annotations and executable robot actions. Human egocentric datasets and robot demonstrations therefore provide complementary forms of supervision: the former preserve natural human strategies and embodiment-independent behavior, whereas the latter provide direct action labels in the target robotic embodiment.
Recent human-to-robot learning methods increasingly combine these complementary data sources. DexMV [78], EgoMimic [47], and EgoBridge [77] explicitly align or adapt human demonstrations to robotic embodiments. EgoVLA [102], UniVLA [9], and In-N-On [10] instead learn transferable action representations from mixtures of human videos and robot demonstrations. EgoScale [114] and Emergence of Human to Robot Transfer [48] further suggest that transfer benefits from increasing the scale and diversity of human experience, robot tasks, and embodiments. These developments reinforce the value of first-person data, but also expose a recurring trade-off between behavioral diversity and physical completeness. Large naturalistic collections capture broad activity distributions and extended temporal context, yet do not continuously measure full-body state, object trajectories, and physical contact throughout every sequence. More instrumented datasets provide stronger hand or object supervision, but often remain local in spatial scope and omit some combination of synchronized room-scale external views, audio, or tactile sensing.
ACE-Data-0 addresses this gap by embedding egocentric observation within a shared physical representation of the surrounding event. The wearable cameras preserve the close-range perspective available to an embodied agent, while synchronized exocentric views recover body and scene context that is frequently invisible from the first-person view. Both perspectives are registered with measured headset, human, and object motion, enabling direct comparison and fusion under the same physical ground truth. Audio and tactile streams further connect visual observation to the acoustic and contact consequences of the same interaction.
2.3 Long-horizon Datasets & Benchmarks
Long-horizon interaction [110, 50, 2, 90] is defined not merely by recording duration, but by persistent dependencies among actions, objects, and scene states. In household tasks, earlier decisions alter later possibilities, objects move in and out of relevance, and sub-tasks may be interrupted, reordered, or resumed elsewhere. Locomotion and manipulation are likewise interdependent: the actor must move through the environment while retaining object locations, intermediate states, and the remaining goal. These properties require models to maintain task and scene memory rather than treating an activity as a sequence of independent atomic actions.
Several human-centered datasets have begun to preserve this richer temporal structure. OakInk2 [110] organizes bimanual manipulation into hierarchical task components rather than isolated grasps, supporting the study of affordances, motion primitives, and complex task completion. OMOMO [54] models the temporal coupling between object motion and full-body human motion over extended interactions, demonstrating how object trajectories constrain human behavior beyond a single contact event. ParaHome [50] records 207 household activity sequences, totaling approximately 486 minutes, from 38 participants using 70 synchronized RGB cameras and wearable motion capture; it emphasizes continuous, concurrent, and multi-object interaction in domestic settings. HuMoTo [64] contains 735 motion-capture sequences involving 63 objects and 72 articulated parts, with task designs that emphasize purposeful progression and coherent multi-object activity. A parallel line of work synthesizes extended human motion under semantic, kinematic, collective, or musical conditioning [35, 56, 13], which likewise depends on motion data that remains coherent over long durations. Together, these datasets mark an important transition from atomic HOI clips toward scene evolution and task-level temporal structure.
Robot-learning datasets address the same challenge from the control side. Mobile ALOHA [27] combines navigation and bimanual manipulation in 290 whole-body teleoperation demonstrations across seven tasks. AgiBot World [1] provides over one million trajectories across 217 tasks, while Galaxea Open-World [46] contains approximately 100K trajectories spanning 150 tasks. RoboCOIN [97] includes roughly 180K demonstrations across 421 tasks and 15 robotic embodiments. RoboMIND 2.0 [39] further scales to approximately 310K trajectories and 739 tasks across six embodiments, with RGB-D observations, mobile manipulation, digital-twin assets, and a subset of tactile episodes. Real-to-sim reconstruction offers a complementary route to scaling such data: HSImul3R [14] reconstructs physical scenes with physics in the loop, producing simulation-ready environments in which robot trajectories can be generated at low cost. Recent methods make the computational demands of long-horizon behavior explicit. SayCan [2] connects high-level task reasoning with executable skills, MEM [90] introduces multi-scale embodied memory, and WholeBodyVLA [43] integrates vision-language-action learning with whole-body loco-manipulation. DynamicVLA [98] instead targets manipulation of moving objects, where temporal anticipation and closed-loop adaptation become necessary. EMMA [116] and HoMMI [99] further use egocentric human demonstrations to reduce the cost of collecting mobile-robot trajectories.
Human-centered and robot-centered datasets nevertheless provide different forms of supervision. Human recordings preserve natural task decomposition, flexible sub-task ordering, hesitation, and recovery, while remaining largely independent of any particular robot embodiment. Robot datasets provide executable controls and precisely aligned observations, but their trajectories are tied to specific kinematics, sensors, and teleoperation interfaces. Few existing resources combine the natural behavioral structure of human demonstrations with continuous measurement of corresponding body, object, scene, and contact states. This separation also limits diagnosis: long-horizon robot benchmarks often emphasize final task success, whereas human activity datasets commonly focus on recognition, segmentation, or motion reconstruction. Final success alone cannot reveal whether failure originated from missed contact, inaccurate object-state estimation, lost task context, incorrect sub-task ordering, or poor motion execution.
ACE-Data-0 is designed to expose these intermediate structures in goal-directed household activity. Participants receive goal-level instructions rather than fixed atomic scripts, allowing object choice, task ordering, movement paths, hesitation, and recovery to emerge naturally. Because visual observations, human and object motion, audio, and tactile signals remain synchronized throughout each task, the dataset bridges geometric HOI benchmarks, long-form egocentric video, and robot trajectory corpora, supporting analysis of not only what occurs, but also the signals, states, and interaction dynamics through which a long-horizon task is carried out.
3 Ambient Capture Engine
In this section, we present the Ambient Capture Engine (ACE) designed for collecting synchronized multi-modal data. Specifically, we first provide an overview of the system and its two scales (Section 3.1), and then describe the capture systems in detail, covering the home environments in which they are deployed and the sensors placed within them (Section 3.2).
3.1 System Overview
In Fig. 2, we provide the overall architecture of ACE. Recording everyday interaction holistically places three obligations on the capture system: it must capture every important signal an interaction may produce, keep those signals synchronized and registered in a common frame, and do all of these in environments that remain believably domestic. No single installation, however, can satisfy the first obligation at every spatial scale: resolving finger-object contact calls for close-range, densely placed sensors, whereas following locomotion across a scene calls for wide baselines and full-room coverage. Rather than compromise between the two, we build ACE at two spatial scales, table-scale and room-scale, and currently deploy it across two sites. The table-scale configuration concentrates on close-range cameras, optical motion capture, and tactile gloves, targeting fine-grained dexterous hand-object manipulation. The room-scale configuration spreads the same sensing suite over a fully furnished apartment, with wide-baseline cameras and optical motion capture covering the full activity area, targeting global motion and interactions distributed across the scene. Both systems share a common acquisition and annotation pipeline: all sensor streams are temporally synchronized, in hardware where sensors permit and in software otherwise (Section 4.1.1), registered into a shared world frame (Section 4.1.2), and enriched with annotations (Section 4.3). The result is a unified corpus in which every frame of every camera can be related to the ground-truth state of the human body, every tracked object, and the accompanying audio and tactile signals; this property underpins the benchmark in Section 5.
3.2 Capture System
Having sketched the architecture, we now zoom in on its physical form: first the environments that host the interactions, and then the sensors that observe them.
3.2.1 Environment Setup
A central design principle of ACE is to preserve the visual and physical realism of a lived-in home while permitting dense sensor coverage. The clutter, furniture, and spatial constraints that laboratory capture removes are precisely what make home-scene interaction difficult, so our environments deliberately keep them. The table-scale configuration is built around a work desk. It covers 30 square meters and is populated with over 25 interactable object instances from more than 8 categories, such as knives, bowls, and food containers. The workspace is instrumented with exocentric RGB cameras rigidly mounted on stands at close range (0.3–0.5 m), providing full coverage of the manipulation area and yielding multiple synchronized views at sub-millimeter effective resolution. An optical motion capture system with 16 OptiTrack cameras, mounted on a shared truss, spans a tracking volume that covers the entire workspace. An overview of the table-scale configuration is presented in Fig. 2(b). On the other hand, the room-scale configuration comprises a fully furnished apartment of approximately square meters, including a kitchen with a kitchen island, a dining area, a living room, and a bedroom, populated with over 25 object instances. A truss suspended from the ceiling and spanning the apartment carries both the RGB cameras and the optical motion-capture system: exocentric RGB cameras hang from the truss, each on an adjustable extension pole that fine-tunes its height and viewing angle, so that any point in the activity area remains visible from at least 4 views even under furniture occlusion; OptiTrack cameras are mounted on the same truss, with their positions and orientations tuned so that the full apartment lies within a single tracking volume. An overview of the room-scale configuration is presented in Fig. 2(a).
| Device | Role | Qty | Resolution / rate | Notes |
| OptiTrack PrimeX 22 | Optical motion capture | 28 | , 60 Hz | IR tracking: 41 body markers, objects, ego rig |
| ZED One | Exocentric RGB capture | 8 | , 30 FPS | GMSL2 to a single Jetson Orin host; shared frame trigger |
| GoPro | Exocentric RGB capture | 8 | , 30 FPS | Rigidly mounted on stands; audio-triggered recording |
| ACE-Ego-Head-V02 Lite | Egocentric capture | 4 cameras | , 20 FPS | One headset: front/back fisheye pairs, IMU, 5 markers |
| Manus | Hand pose | 2 gloves | 60 Hz | Per-finger articulation, both hands |
| ACE-Sense-Glove Lite | Contact pressure | 2 gloves | – | Full-palm pressure map, both hands |
| Jetson Orin | Recording host | 1 | – | Ingests all ZED One streams; NatNet-coordinated |
| Motive host (PC) | Mocap host, sync reference | 2 | – | Renders the optical clock for cross-system sync |
3.2.2 Hardware Setup
Within these environments, the sensor suite is assembled so that every facet of an interaction, from what the human sees, to how the body and objects move, to what the interaction sounds and feels like, has a dedicated sensor. Our two systems share this suite and differ only in sensor selection and placement, as summarized in Table 2.
Egocentric camera.
Participants in both systems wear an ACE-Ego-Head-V02 Lite by ACE Robotics, a head-mounted egocentric capture device with four fisheye cameras facing front-left, front-right, rear-left, and rear-right, recording at FPS with an onboard IMU. Five OptiTrack markers are attached to the device to provide its initial placement and its real-time 6-DoF pose throughout the capture. This helps register the egocentric streams into the same world frame as every other sensor.
Exocentric camera.
The table-scale configuration surrounds each workspace with close-range GoPro RGB cameras ( FPS), rigidly mounted on stands, while the room-scale configuration covers the apartment with wide-baseline ZED One RGB cameras ( FPS).
Human motion.
Each participant wears a motion-capture suit with markers attached at the joints across the whole body (including markers around the head), tracked by the OptiTrack system (PrimeX 22; 12 cameras in the room-scale configuration, 16 in the table-scale configuration) at Hz to yield 41-joint skeletons, complemented by articulated hand poses, acquired via Manus motion-capture gloves at Hz in the room-scale configuration, and via RANSAC triangulation of 2D hand keypoints from the exocentric cameras followed by manual refinement in the table-scale configuration.
Object motion.
Every interactable object is first digitized into a mesh by 3D scanning or 2D Gaussian Splatting reconstruction [40], and fitted with OptiTrack markers that are bound to this mesh in the tracking system, so that the resulting 6-DoF trajectories at 60 Hz place the exact object geometry in the world frame.
Audio.
We capture the audio via the GoPro exocentric cameras and the ACE-Ego-Head-V02 Lite to record the contact events, appliance operation, and ambient scene sound.
Tactile.
Participants wear full-palm tactile gloves, which record contact pressure maps across the palm and fingers area.
4 ACE-Data-0
With the capture systems in place, this section presents ACE-Data-0. Specifically, Section 4.1 presents the data acquisition pipeline, covering the capture workflow, multi-modal recording, sensor synchronization, and calibration. In Section 4.2, we illustrate the data collection process, including task design, capture settings, dataset statistics, and modalities. Finally, Section 4.3 introduces the annotations provided with the dataset and the pipeline used to produce them.
4.1 Data Acquisition Pipeline
The sensors above produce a dozen heterogeneous streams, and these streams are useful only when they can be precisely related to one another in both time and space. This subsection first describes how the streams are aligned, temporally (Section 4.1.1) and spatially (Section 4.1.2), and then how capture sessions are operated in practice (Section 4.1.3 and Section 4.1.4). An overview of the workflow for ACE is presented in Fig. 3.
4.1.1 Sensor Synchronization
Temporal alignment is the foundation of multi-modal capture: cameras, motion capture, object trackers, and tactile sensors run at different rates on independent clocks, and even a few milliseconds of drift are enough to corrupt contact-level annotation, making a hand appear to close on a cup several frames before the tactile stream registers the touch. We take the OptiTrack clock as the reference, since the tracking system internally synchronizes its own cameras and delivers marker positions as strictly simultaneous Hz frames, and align each remaining device to it in turn.
Exocentric cameras.
(1) For the room-scale configuration, the ZED One cameras are ingested by a single NVIDIA Jetson Orin host over GMSL2, whose capture cards drive all cameras from a common frame trigger, so their captures are aligned by construction. Alignment to the OptiTrack clock then proceeds in two steps. First, recording is coordinated over the local network: the camera host subscribes to the OptiTrack data stream via NatNet and starts and stops in lockstep with the motion-capture recording. This brackets the takes, but a network protocol without a shared clock cannot align individual frames. Unfortunately, the recorded timestamps cannot close this gap: they lag the true exposure moment by roughly s, and this lag also drifts slowly over time. To address this problem, we instead let the exocentric camera photograph a clock, in the form of QR codes. Specifically, the motion-capture host computer displays its own time on the monitor at nanosecond resolution. Reading this clock off a recorded frame gives the exact capture time of that frame, in motion-capture time. We sample such readings across a take and fit a line through them (a constant offset plus a slow drift). As a result, the final residuals after alignment are at the millisecond level, within a single motion-capture frame. (2) For the table-scale configuration, the GoPro cameras are synchronized with one another by aligning their audio streams.
Egocentric cameras.
The four fisheye views of ACE-Ego-Head are driven by the device’s onboard controller, with a measured mutual misalignment of under ms. (1) For the room-scale configuration, the alignment to the OptiTrack clock reuses the optical clock above, with one complication: the egocentric cameras face the hands and the wearer’s back, and never see the monitor during a take. Therefore, we start each take with a deliberate glance, in which the wearer points one camera of ACE-Ego-Head at the monitor for around seconds. Reading the clock from these frames gives the offset between the clock of the egocentric camera and the OptiTrack clock. The device’s shared clock then carries this offset to the other three views for overall synchronization. (2) For the table-scale configuration, the GoPro cameras and the egocentric headset read the same QR-code clock displayed on the monitor of the motion-capture host, as in the room-scale setup; this aligns the GoPro streams, already mutually synchronized via audio, to the headset clock. The egocentric headset is then aligned to the OptiTrack system by registering its internally estimated poses (from AprilTag-based tracking) to the corresponding poses tracked via optical markers, completing the chain from the GoPro cameras, through the headset, to the OptiTrack timeline. Together, these procedures register all devices in both configurations to the OptiTrack timeline.
Other sensors.
The tactile gloves are synchronized with ACE-Ego-Head using their onboard IMU signals. At the beginning of each take, the operator performs a short motion pattern while keeping the hand and head approximately rigid, producing correlated IMU signals between the glove and the egocentric headset. We align the two streams by matching this motion template, establishing temporal synchronization.
Verification and output.
As an independent check, we find that the per-camera time offsets re-estimated during calibration (Section 4.1.2) are under ms, confirming the alignment. All fitted offsets are folded into a per-take table that maps every camera frame, exocentric and egocentric alike, to its Hz motion-capture frame. Downstream processing consumes only this table. See Fig. 4 for an example of synchronized capturing.
4.1.2 Calibration
Whereas the synchronization aligns the streams in time, calibration registers them in space, and the two camera systems pose opposite challenges: the exocentric cameras never move but barely share a view, while the egocentric cameras move every frame. The two calibration procedures are illustrated in Fig. 5 and Fig. 6.
Exocentric cameras.
Standard multi-camera calibration assumes co-visibility: two cameras must observe the same target at the same time. Our sparse placement of the exocentric cameras breaks this assumption, as most camera pairs share no common area. We therefore bridge the cameras through the motion-capture system instead. Specifically, we utilize an ArUco board [29] with a retroreflective marker at each corner as the calibration target. This one physical object is visible to every sensing system at once: the RGB cameras see the printed pattern of the ArUco board, while the OptiTrack infrared cameras see both that pattern and the retroreflective corner markers. The board therefore ties the RGB cameras and the motion-capture volume to a single reference without requiring any two cameras to share a view. Then, the calibration proceeds in three steps: (1) First, the infrared cameras estimate where the corner markers actually sit on the board. It is important to note that these markers are attached by hand, so their exact positions are not known in advance. Aggregating the estimates over thousands of board poses reduces the error in this step to the millimeter level. (2) Second, with their positions known, the markers alone give the board’s pose in every frame, at sub-millimeter consistency. (3) Third, each RGB camera is fitted against this marker-anchored board trajectory. Not every observation contributes in this step: (i) frames in which the board was moving are discarded, because residual timing error grows with the motion, and (ii) the detections near the image edges are discarded, where the distortion model is least reliable. A final joint refinement then re-estimates the board pose from all cameras together and refits each camera in turn. Throughout, the tracked markers remain the dominant reference, keeping the solution anchored in the motion-capture world. On held-out frames, every camera reaches a median reprojection error below px, roughly a centimeter of 3D error at typical distances.
Egocentric cameras.
A moving rig is not described by one pose but by a pose for every frame. Considering our situation where the egocentric views are dominated by the wearer’s own arms and torso and the front and back camera pairs share no field of view, recovering this trajectory from vision alone, as SLAM would, is unreliable. We therefore do not estimate where the cameras are; we measure them. The five markers on the ACE-Ego-Head chassis form a rigid body that OptiTrack tracks at Hz. The only unknown left is the fixed transformation from each fisheye camera to this body, which is a classical hand-eye calibration problem [92]. We first solve it independently for each camera. A joint bundle adjustment [91] then refines the four transformations together, treating the board pose, shared by all cameras, and the per-camera time offsets as additional free variables. We also use only low-angular-velocity frames for the fitting, as the timing error grows with the motion. The final median reprojection error is about px. Once calibrated, producing camera poses for a take requires no images at all. OptiTrack tracks the egocentric camera’s rig body at Hz, while the cameras record at FPS; for each egocentric frame, we interpolate the rig’s tracked pose to the frame’s timestamp and apply that camera’s hand-eye transformation. Every pose is thus measured rather than estimated, and does not drift.
Other calibrations.
The two procedures above answer where the cameras are. The remaining calibrations describe each sensor itself. For the cameras, this means intrinsics: the exocentric ones use their factory parameters, and the egocentric fisheyes are calibrated with Kalibr [28]. For the objects, this means geometry: each marker set is registered once to its scanned or 2DGS-reconstructed mesh, so that tracking the markers places the full object in the world frame. All of these are re-verified twice a day. Together with the synchronization above, we complete the unified spatio-temporal frame on which all subsequent annotation and benchmarking rest.
4.1.3 Capture Workflow
With alignment in place, every capture session follows the same five-step protocol at both sites. (1) Scene preparation: objects are placed at randomized yet plausible initial positions. (2) Participant setup: the participant puts on the motion-capture suit, the ACE-Ego-Head headset, and the gloves, then performs a short T-pose routine that registers their skeleton with the tracking system. (3) Task briefing: the participant receives a goal-level instruction verbally; how to achieve the goal is left entirely to them. (4) Recording: the take opens with the clock glance described in Section 4.1.1, after which all sensors record continuously while an operator monitors stream health on a live dashboard. (5) Post-checks: synchronization and tracking quality are verified after each take, and failed takes are flagged for re-capture.
4.1.4 Multi-modal Recording
During step (4) of the capture workflow, each system simultaneously acquires: (i) the exocentric RGB streams; (ii) egocentric fisheye streams with IMU; (iii) full-body motion and hand poses at Hz; (iv) -DoF poses of all tracked objects at Hz; and (v) tactile signals. All streams are timestamped against the common clock established in Section 4.1.1. A one-hour session produces approximately TB of raw data.
4.2 Data Collection
The preceding subsections established the recording machinery; the question that remains is what to record with it. We describe the task design (Section 4.2.1), the capture settings (Section 4.2.2), the resulting dataset statistics (Section 4.2.3), and the modalities included in each released sequence (Section 4.2.4).
4.2.1 Task Design
The long-horizon character of ACE-Data-0 originates in its task design. By a long-horizon task we mean an activity that unfolds over minutes rather than seconds and decomposes into a sequence of sub-actions whose order is constrained by the goal rather than fixed in advance. For example, consider the task of “prepare a cup of tea and serve it at the table”. It involves a chain of human-object interactions: the participant walks to the cupboard, opens it, takes out a cup, fills the kettle, waits for the water to boil, retrieves a tea bag, pours, stirs, carries the cup across the room, and clears space before setting it down. This single take may thus contain planning, locomotion, and fine-grained manipulations.
However, most HOI datasets instead prescribe atomic actions [7, 15, 52, 22], capturing a single grasp or handover in isolation. To address this gap, we prescribe household goals and let the actions emerge: how to reach the goal is left to the participant. Different participants order the sub-tasks differently, grasp differently, and reach for different objects, so the variability of real behavior enters the data by itself.
Specifically, we record three types of takes:
- •
Atomic HOI tasks contain one to three household tasks each, drawn from more than 15 types of household activities, such as pouring water, drinking, making tea, watering plants, chopping vegetables, cooking, and tidying up; each take lasts roughly three minutes. Each of them provides clean, self-contained instances of everyday manipulation, the unit from which skills are most readily learned.
- •
Chains of HOI tasks combine the full range of short tasks into one continuous activity of roughly twenty to thirty minutes. Sub-tasks interleave freely, yet every take ends with the scene tidied back into order, tracing a full cycle of household activity. Compared to the atomic HOI tasks, they exercise long-horizon planning, state tracking, and memory at a further level of complexity.
- •
HSI tasks involve almost no objects, focusing instead on interactions with scene components such as tables, chairs, and sofas. Such takes record whole-body motion, such as walking and exercising, and human-scene contact, such as sitting, lying, and leaning, in about five minutes per take. These movements complete the range of behaviors a humanoid must master.
To decide what these recordings should contain, we surveyed the tasks that arise in everyday home life and designed the atomic HOI, chain-of-HOI, and HSI tasks accordingly. We present examples of different task categories in Fig. 7.
4.2.2 Capture Settings
To perform these tasks, we recruited 50 participants. Each participant completes all the designed tasks across a -day session. Diversity enters the data from two directions. On the human side, the goal-level instructions leave the behavior open: participants differ in where they start and end, the route they take in between, the interactions they choose, and how long each step takes. On the object side, takes vary in which categories appear and in what number, where the objects are placed, where they begin and end, and along what trajectories and in what manner they are moved.
4.2.3 Dataset Statistics
This collection effort yields more than 150 hours of synchronized multi-modal capture, comprising over 17M frames across more than 75,000 episodes. We define an episode as a contiguous segment of interaction that realizes one meaningful sub-goal, the smallest unit that remains semantically self-contained when used as a training example. Episodes are counted within takes rather than recorded in isolation, so the surrounding context, namely the actions that precede and follow each segment, is preserved in the same stream. In ACE-Data-0, even the shortest takes run minutes rather than seconds, roughly an order of magnitude longer than typical HOI clips [7, 15, 52, 22]. Fig. 8 reports the distributions of task categories, take lengths, object categories, etc.
4.2.4 Modalities
A released take contains more than the raw streams of Section 4.1.4: it also carries everything the pipeline of Section 4.1 derives from them, in a common frame and on a common timeline. Concretely, each take bundles:
- •
Egocentric video: the four fisheye views, their IMU readings, and per-frame 6-DoF headset poses from the tracked rig;
- •
Exocentric video: all views, each with intrinsics and its pose in the world frame;
- •
Human motion: -joint body skeletons with articulated hand poses and converted SMPL-X motion parameters [71];
- •
Object motion: per-object 6-DoF trajectories, with scanned or 2DGS-reconstructed meshes for more than 50 instances;
- •
Audio: synchronized multi-source audio recorded from GoPro exocentric cameras and the egocentric headset;
- •
Tactile: hand-shaped pressure grids remapped from raw glove sensors, with calibrated normalization and baseline correction.
4.3 Annotation
The raw signals described so far become substantially more useful once enriched with annotations. We describe what annotations ACE-Data-0 provides (Section 4.3.1) and how they are produced at scale (Section 4.3.2).
4.3.1 Annotation Types
ACE-Data-0 provides five types of annotation, which jointly characterize every take: the objects involved in the interaction, the configuration of the actor’s body, the articulation of the hands during manipulation, the physical contact established between hand and object, and the acoustic and semantic description of the activity itself. Specifically, the object annotations (Fig. 9) comprise a category label, a bounding box, and a per-frame 6-DoF pose for every tracked object, together with the motion trail traced by the object over the course of the take. Human annotations (Fig. 10) provide full body pose throughout each take, and hand annotations (Fig. 11) refine this to articulated finger configurations during periods of dexterous manipulation. Tactile annotations (Fig. 12) register the timing and spatial distribution of contact, resolving interaction events that remain ambiguous under visual occlusion. Audio and language annotations (Fig. 13) pair the recorded audio streams with natural-language descriptions that identify the sound events produced by the interaction, such as an object set down on a surface or a container being opened, and state the goal of each take alongside the sequence of sub-goals through which it is achieved.
Calibration and synchronization.
Every take ships with per-camera intrinsics, camera poses in the shared world frame, and the timeline that aligns all sensor streams. These are the outputs of Section 4.1, released as data: with them, any tracked 3D point can be projected onto any pixel of any view, and any two streams can be paired at any instant. Users can therefore combine the modalities freely for training, pairing any subset of them as inputs and supervision, without rerunning any part of our pipeline.
Human poses.
The captured body and hand poses are reprojected onto every frame of every camera, giving per-view 2D mesh overlays consistent with the 3D capture. Because these projections come from measured 3D states rather than image-based detectors, they remain correct where detectors typically fail: under furniture occlusion, extreme viewpoints, and motion blur. Each frame thus carries pixel-aligned 2D poses in every view and metric 3D poses in the world frame.
Tactile labels.
Each frame carries the tactile reading from the tactile glove, temporally aligned with the visual streams and the object label currently in use. Contact thus can be detected rather than inferred: the signal comes directly from the sensor surface, not from appearance or from proximity between reconstructed geometry. These readings mark when and with what an interaction happens, the anchor from which most HOI tasks start.
Object annotations.
Every object instance carries its scanned or 2DGS-reconstructed mesh, per-frame 6-DoF poses, 2D/3D bounding boxes in all views, and its motion trail over the take. Mesh and pose together place the exact object geometry in the scene at every instant; the boxes and trails are their projections into each camera. The full history of an object, where it sat, when it moved, and where it ended, can therefore be queried at any point of a take.
Audio descriptions.
Each take carries multi-channel audio recorded by the head-mounted egocentric capture rig and the GoPro cameras, sharing the synchronization clock of the visual streams and therefore aligned to the same timeline as the pose, object, and tactile annotations. Two classes of sound are present: those produced by the interaction itself, such as an object set down on a surface or liquid poured into a cup, and the ambient acoustics of the domestic environment, including appliance noise, footsteps, and room reverberation. Because the audio is captured from the actor’s own viewpoint, the acoustic perspective moves with the participant, so a sound event varies in loudness and spatial character with proximity.
Textual descriptions.
Gemini-3.1-pro-preview [89] watches the ego-view video and describes each time span in natural language: what the person is doing, and what happens in the scene. The descriptions give each take a searchable storyline and connect the physical record to language, in the form that vision-language and vision-language-action models consume directly.
4.3.2 Annotation Pipeline
Producing annotations of this breadth would ordinarily demand prohibitive manual effort. ACE avoids most of it, because among our five annotation types, all but the textual descriptions are measured rather than estimated. Human and object states are metrically tracked, every camera is calibrated, and contact is directly sensed. The pose reprojections, bounding boxes, motion trails, and contact events therefore follow from the recorded states by projection and tactile sensing, with no estimation model in the loop. The textual descriptions are the one generated type: Gemini produces them from the ego-view videos, segment by segment. Finally, human annotators check the auto-generated labels and correct the descriptions by hand.
5 Benchmark
Having described how ACE-Data-0 is captured and annotated, we now use it to evaluate existing methods, with a diagnostic goal: to expose where current approaches break on long-horizon home-scene data, and to indicate where solutions might lie. The benchmark comprises three levels, advancing from signals, to components, to interactions, with each level building upon the one beneath it: low-level signal inference operates directly on the raw sensory streams, i.e., predicting tactile signals from video (Section 5.1); scene component recovery assembles these signals into the 3D human components (Section 5.2); and interaction estimation recovers the hand motion through which the human engages objects, from egocentric and exocentric views (Section 5.3). This progression mirrors the perceptual capabilities an embodied agent must chain together when acting in a home: sensing contact, estimating scene state, and mastering hand-object coordination (see Fig. 14). We hold out 10 hours of capture as a test set for all benchmark evaluations in this report. Unless otherwise noted, baselines are evaluated with their officially released pre-trained checkpoints, for fair comparison.
5.1 Low-level Signals
We begin with the raw sensory signals that interactions produce, and among them, touch holds a special position. Vision tells an observer where things are; touch tells the actor whether a grasp has succeeded, how firmly to hold, and when an object begins to slip. Robots need this signal as much as humans do, yet contact sensing remains scarce in practice: tactile hardware is expensive and fragile, and absent from most deployed platforms, while cameras are everywhere. Learning to read touch out of video would therefore turn the most abundant sensor into a substitute for the scarcest one. ACE-Data-0 makes this mapping learnable: through the synchronization of Section 4.1.1, every video frame is paired with the tactile reading recorded at the same instant.
| Method | Temp Acc. | C-IoU | V-IoU | CoP |
|---|---|---|---|---|
| PressureVision [32] | 0.0093 | 0.0007 | 0.0000 | 10.9807 |
| EgoPressureDiff [109] | 0.2912 | 0.0197 | 0.0025 | 8.5152 |
| TouchAnything [115] | 0.7095 | 0.1646 | 0.1357 | 6.5846 |
5.1.1 Tactile from Vision
Visual observations provide only indirect evidence of physical contact. This ambiguity becomes particularly severe in egocentric views, where the manipulating hand frequently occludes the fingertips and the contacted object regions precisely when pressure is applied.
Task.
Given an egocentric video of a human-object interaction, the goal is to predict full-hand grasp pressure at each moment. We evaluate the predicted pressure distributions against synchronized measurements captured by the tactile glove.
Metrics.
We report four complementary metrics. Temporal accuracy measures frame-wise contact-state prediction over the interaction sequence. Contact IoU (C-IoU) evaluates the spatial overlap between thresholded predicted and ground-truth contact regions, while volumetric IoU (V-IoU) additionally accounts for pressure magnitude through min-max aggregation. Center-of-Pressure (CoP) error measures the spatial distance between predicted and ground-truth pressure centroids over the fingertips and palm, with lower values indicating more accurate contact localization.
Baselines.
We evaluate three baseline methods using their released pre-trained weights. PressureVision [32] directly regresses pixel-aligned pressure maps from RGB observations using a convolutional encoder-decoder. EgoPressureDiff [109] formulates pressure estimation as conditional diffusion and adapts a pre-trained video diffusion backbone. TouchAnything [115] learns a general vision-to-touch representation using cross-view fusion and view-dropout training.
Results.
Table 3 reports the results on the close-range table-scale recordings. PressureVision provides little meaningful contact prediction, producing nearly zero overlap under both C-IoU and V-IoU and the largest CoP error. EgoPressureDiff substantially improves temporal contact recognition and pressure localization, but its spatial overlap remains limited, indicating that detecting when contact occurs is considerably easier than recovering where pressure is distributed across the hand. In contrast, TouchAnything consistently achieves the strongest performance across all four metrics. Its advantage is especially pronounced in temporal accuracy and spatial overlap, suggesting that broader visual-tactile training and cross-view modeling improve transfer to the diverse interactions in ACE-Data-0. Nevertheless, its absolute C-IoU and V-IoU remain modest, and the remaining CoP error indicates that accurate pressure localization is still challenging. In particular, a model may correctly identify the presence of contact while failing to recover its precise position and intensity under hand-object occlusion.
Overall, the results reveal a substantial generalization gap for existing tactile-estimation models and establish ego-view pressure reconstruction as a challenging benchmark. Robust performance requires not only recognizing contact events, but also resolving fine-grained pressure distributions from visually ambiguous and frequently occluded interactions.
5.2 Scene Components
Where the previous track asks what an interaction feels like, this one asks where the components of the scene are and how they move over time. Between raw sensor signals and any understanding of an interaction lie the 3D states of the scene: the pose of the body, the articulation of the hands, and the 6-DoF pose of every object. These states are what most vision pipelines estimate first and what every downstream module consumes.
In home scenes, however, their reliability remains largely unquantified. In-the-wild footage offers no ground truth to measure against, datasets with metric ground truth are confined to laboratory settings, and datasets recorded in real environments are pseudo-labeled by the very methods under evaluation. ACE-Data-0 removes this obstacle by providing metric ground truth in furnished domestic scenes, allowing existing estimators to be evaluated under the conditions in which they are actually deployed. We therefore benchmark human motion estimation in this track. Hand articulation is treated separately in the following track, where dexterous manipulation provides the setting in which it matters most. Object pose estimation is left to future work: too few applicable methods [96] exist to support a meaningful comparison, though the released ground truth supports this evaluation directly.
| Method | PA-MPJPE | PA-PVE | MPJPE | PVE | WA-MPJPE |
|---|---|---|---|---|---|
| — Single-view exocentric, per-frame — | |||||
| Multi-HMR [4] | 110.5 | – | 115.3 | – | – |
| Multi-HMR2 [24] | 85.5 | – | 92.8 | – | – |
| SAM-3D-Body [103] | 71.0 | – | 76.4 | – | – |
| PyMAF-X [111] | 94.7 | 177.5 | 99.9 | 177.0 | – |
| PARE [51] | 98.7 | 132.3 | 102.0 | 134.8 | – |
| CameraHMR [70] | 78.7 | 98.7 | 85.7 | 105.2 | – |
| OSX [59] | 59.2 | 73.7 | 62.5 | 76.6 | – |
| — Single-view exocentric, temporal — | |||||
| SMPLer-X [11] | 57.0 | 70.2 | 59.8 | 71.7 | – |
| SMPLest-X [107] | 55.7 | 65.8 | 58.8 | 67.6 | – |
| Humans-in-4D [31] | 77.1 | 94.4 | 79.6 | 96.4 | – |
| GVHMR [84] | 88.0 | 104.1 | 96.3 | 112.0 | 217.1 |
| WHAM [85] | 131.1 | 160.2 | 144.0 | 166.4 | 243.1 |
| EasyMoCap [21, 86] | 103.4 | 155.8 | 106.0 | 156.1 | 256.2 |
| — Single-view exocentric, scene-aware — | |||||
| Phy-SIC [68] | 64.3 | 78.9 | 71.3 | 84.6 | – |
| UniSH [55] | 80.0 | 111.2 | 82.8 | 114.0 | 196.3 |
| JOSH [62] | 64.0 | 89.2 | 69.9 | 95.0 | 245.1 |
| Human3R [16] | 60.1 | 75.4 | 63.5 | 80.6 | 180.2 |
| — Multi-view exocentric — | |||||
| MAMMA [18] | 70.1 | 86.8 | 74.8 | 89.6 | 230.8 |
| U-HMR [57] | 134.4 | 176.0 | 150.2 | 179.2 | – |
| HSfM [67] | 92.8 | 133.8 | 93.8 | 134.8 | – |
| — Egocentric — | |||||
| EgoEgo [53] | 159.6 | 245.3 | 163.4 | 263.9 | 306.2 |
| EgoAllo [106] | 131.7 | 196.8 | 147.9 | 220.3 | 252.2 |
5.2.1 Human Motion Estimation
Human motion estimation in home environments remains challenging for several reasons. Furniture often blocks the lower body for long parts of a sequence. Common household actions, such as crouching in front of a cabinet, reaching above the head, or lying on a sofa, also differ from the mostly upright poses found in many training datasets. In addition, each take can last several minutes, allowing small frame-level errors to build up and cause large drift in the estimated global trajectory [113, 85, 84, 68]. Our dataset captures these challenges together and provides metric ground-truth motion for the full sequence.
Task.
Given the visual observations from a take, the goal is to estimate the articulated body pose over time. Per-frame methods process individual images, while temporal methods use the full video sequence. Methods with world-frame outputs are also expected to recover the person’s global trajectory through the scene.
We evaluate all predictions against motion-captured ground truth. The evaluation covers three input settings: multi-view exocentric video, single-view exocentric video, and egocentric video. It also includes three method families: per-frame, temporal, and scene-aware methods. Since all settings use the same ground truth, the results allow us to compare the effects of viewpoint, temporal information, and scene context directly.
Metric.
We report five metrics, ordered from local to global. PA-MPJPE and PA-PVE [42] evaluate the articulated pose and the recovered body surface after Procrustes alignment, independent of global position and orientation. MPJPE and PVE [63, 71] evaluate the same quantities with the global orientation retained. WA-MPJPE [85, 84] evaluates the global trajectory in the world frame, after a single alignment of the whole path.
Baselines.
We evaluate 22 methods, covering every major family of human motion estimation. For single-view exocentric input, these include per-frame approaches (Multi-HMR [4], Multi-HMR2 [24], SAM-3D-Body [103], PyMAF-X [111], PARE [51], CameraHMR [70], and OSX [59]), video-based approaches (SMPLer-X [11], SMPLest-X [107], Humans-in-4D [31], GVHMR [84], WHAM [85], and EasyMoCap [86]), and their scene-aware counterparts, Phy-SIC [68] for the former and UniSH [55], JOSH [62], and Human3R [16] for the latter. For multi-view exocentric input, we evaluate MAMMA [18], U-HMR [57], and HSfM [67]. For egocentric input, we evaluate EgoEgo [53] and EgoAllo [106]. All methods run with their released pre-trained weights. To our knowledge, this is the broadest evaluation of human motion estimation conducted in real home scenes to date.
Results.
We evaluate primarily on the room-scale recordings, which combine locomotion, furniture occlusion, and minutes of continuous motion. Quantitative results are reported in Table 4, and three findings stand out.
First, the results show a clear gap between local pose estimation and global trajectory recovery. Several methods achieve strong results on the Procrustes-aligned metrics, with accuracy close to that reported on standard benchmarks. This suggests that they can still recover articulated body poses under household occlusion and uncommon postures. However, their world-frame trajectory errors remain much higher. The ranking on local pose metrics also differs from that on global trajectory metrics. In other words, estimating the body pose correctly in each frame does not guarantee an accurate motion path over the full sequence.
Second, the comparison between scene-aware and temporal methods further shows where scene information is most useful. Scene-aware methods achieve lower trajectory errors, while their Procrustes-aligned results remain similar to those of temporal methods. This suggests that scene context mainly helps determine the body’s position in the room, rather than improving the relative joint configuration. Scene geometry provides useful constraints on where a person can stand or move, but gives less direct information about the exact body pose.
Third, the results also show a strong effect of viewpoint. Egocentric methods perform worse across most metrics because much of the body is outside the camera’s field of view and must be inferred from head motion. Multi-view methods, however, do not outperform the strongest single-view methods in this evaluation. We do not view this as evidence that multiple views are less useful. Instead, it likely reflects the strong performance of recent single-view whole-body and scene-aware methods, as well as the limited range of available multi-view baselines.
5.3 Embodied Interaction
Knowing the position of the person is not enough to describe how an interaction is carried out. Much of this information is contained in the hand motion, including how the hand approaches an object, forms a grasp, adjusts its position, and releases it. Hand motion is also especially relevant to robot learning because a hand trajectory can be retargeted to a robotic gripper more directly than raw visual observations.
We therefore evaluate how accurately current methods [78, 47, 77] recover hand motion from video. In ACE-Data-0, each interaction is recorded at the same time from both egocentric and exocentric viewpoints. This allows us to compare the two settings under the same conditions, using the same takes and the same ground truth while changing only the viewpoint.
| Method | PA-MPJPE | MPJPE | F@5 | F@15 | AUCJ | Traj. err. |
|---|---|---|---|---|---|---|
| WildHands [76] | 11.2 | 12.6 | 0.175 | 0.774 | 0.776 | – |
| Dyn-HaMR [108] | 18.9 | 21.1 | 0.042 | 0.413 | 0.624 | 98.2 |
| HaWoR [112] | 13.8 | 17.4 | 0.130 | 0.666 | 0.729 | 102.1 |
5.3.1 HOI from Ego-View
Egocentric video provides the viewpoint that is most similar to what a future robot may observe. The camera stays close to the hands and often captures fine details of the interaction. However, this viewpoint also introduces several practical challenges. The camera moves with the head, the wide-angle lens distorts the image near the boundary, and the hands often appear in these distorted regions. Large head or body movements may also move one or both hands outside the field of view.
Task.
Given the egocentric video of a take, a method estimates the 3D motion of both hands as MANO [81] sequences. This includes the articulated finger pose in each frame and, when supported by the method, the hand trajectory in a world coordinate frame. We obtain the ground truth from the Manus motion-capture gloves for the room-scale recordings; for the table-scale recordings, we apply RANSAC triangulation to 2D hand keypoints from the eight synchronized exocentric cameras, followed by manual refinement. Joints that are visible in too few views are masked during evaluation. The method is therefore not penalized on frames for which reliable ground truth cannot be obtained.
Metrics.
We report five metrics for hand pose and one metric for the global trajectory. The same definitions and units are used for both egocentric and exocentric evaluation. PA-MPJPE applies a separate Procrustes alignment to each predicted hand using rotation, translation, and scale. It therefore mainly measures errors in finger articulation. MPJPE uses only rotation and translation while keeping the scale fixed, and thus also reflects errors in metric hand size. F@5 and F@15 report the fractions of joints with errors below 5 and 15 mm, respectively. AUCJ is the normalized area under the PCK curve over thresholds from 0 to 50 mm. For methods that predict a persistent world coordinate frame, we also report the world-frame trajectory error. We align the predicted and ground-truth wrist trajectories using one similarity transform for the entire clip and compute the mean residual error. Unlike per-frame alignment, this metric captures drift and inconsistent scale across the sequence.
Baselines.
We evaluate three methods using their released pre-trained weights. WildHands [76] is a per-frame regressor and is applied independently to each egocentric frame. Dyn-HaMR [108] and HaWoR [112] process videos and recover the hands in a world coordinate frame by combining hand estimation with camera-motion estimation.
Results.
Table 5 shows a clear difference between frame-level hand pose estimation and world-space hand motion recovery. The per-frame method achieves the strongest articulation accuracy, while the two video-based methods perform less well on the local pose metrics. One possible reason is that the video-based methods estimate camera motion and hand pose together. Errors in camera motion can therefore affect the recovered hand pose, especially in our recordings where head motion is often large. The difference becomes more evident for global trajectory estimation. Both world-space methods produce trajectory errors that are much larger than their local joint errors. This suggests that the main challenge in egocentric hand reconstruction is not estimating finger articulation in individual frames, but maintaining a stable hand trajectory in the world coordinate system. Together with the exocentric results below, this points to egomotion estimation as a major source of error.
| Method | PA-MPJPE | MPJPE | F@5 | F@15 | AUCJ | Traj. err. |
|---|---|---|---|---|---|---|
| HORT [17] | 10.8 | 12.3 | 0.276 | 0.776 | 0.784 | – |
| HaMeR [72] | 9.6 | 10.4 | 0.281 | 0.848 | 0.812 | – |
| HaPTIC [105] | 10.0 | 10.7 | 0.247 | 0.838 | 0.804 | 63.0 |
| WiLoR [75] | 9.1 | 9.9 | 0.313 | 0.854 | 0.819 | – |
| OmniHands [58] | 10.7 | 11.5 | 0.211 | 0.814 | 0.791 | – |
5.3.2 HOI from Exo-View
Exocentric video presents a different set of advantages and challenges. The camera remains fixed, and the person and surrounding scene usually stay within the image. This provides a stable reference frame for tracking motion over time. However, the hands occupy only a small part of the image and are often occluded by the manipulated object or by the person’s body [26, 22]. Since the egocentric and exocentric cameras record the same interactions, we can compare the two viewpoints using the same takes and ground truth.
Task.
Given either one exocentric stream or all synchronized exocentric streams, a method estimates the articulated pose of both hands in each frame. Methods with world-frame outputs additionally recover the hand trajectories across the sequence. The ground truth and evaluation procedure are the same as those used in Section 5.3.1, allowing direct comparison between the two viewpoints.
Metrics.
We use the same five hand-pose metrics defined in Section 5.3.1. For video methods, we also report the world-frame trajectory error. Although the exocentric camera is static, a method must still recover hand positions consistently and at the correct metric scale throughout the clip.
Baselines.
We evaluate five single-view methods using their released pre-trained weights. HaMeR [72], WiLoR [75], and OmniHands [58] estimate the hands independently in each frame. HaPTIC [105] processes video and produces hand motion in a world coordinate frame. HORT [17] jointly reconstructs the right hand and the manipulated object, and therefore produces predictions only for frames containing object manipulation.
Results.
Table 6 summarizes the results. The five methods achieve similar articulation accuracy, with PA-MPJPE values ranging from 9.1 to 10.8 mm. WiLoR performs best, reaching 9.1 mm PA-MPJPE, an F@5 score of 0.313, and an AUCJ of 0.819. HaMeR follows closely with a PA-MPJPE of 9.6 mm. These results indicate that current crop-based hand regressors can recover detailed hand poses even when the hands occupy a small region of a 1080p image.
The fixed camera also makes global trajectory estimation more reliable. HaPTIC obtains a trajectory error of 63 mm while maintaining a PA-MPJPE of 10.0 mm, which is comparable to the per-frame methods. Its trajectory error is substantially lower than the 98–102 mm obtained by the egocentric world-space methods. This difference suggests that a stable camera coordinate system removes much of the uncertainty caused by egomotion.
HORT achieves a PA-MPJPE of 10.8 mm on the manipulation frames for which it produces predictions. This result is close to those of the other methods. Under this evaluation, jointly modeling the held object does not provide a clear improvement in articulated hand-pose accuracy, but it also does not noticeably reduce it.
5.3.3 Cross-View Analysis
Because the two viewpoints observe the same interactions, their results can be compared directly. Although the egocentric camera provides a closer view of the hands, the exocentric methods achieve better results for both articulation and global trajectory estimation. Among these results, the trajectory analysis shows the largest difference. The exocentric video method obtains an error of 63 mm, while the egocentric methods remain close to 100 mm. With a fixed exocentric camera, the camera coordinate system already provides a stable reference frame. Egocentric methods must estimate this reference frame from head motion, and errors in this step become a major part of the final trajectory error.
Overall, the two viewpoints still provide complementary observations. Egocentric video is more affected by truncation, distortion, and motion blur, while exocentric video is more affected by occlusion from the body and manipulated objects. Combining both streams could therefore reduce their individual failure cases. Providing measured headset motion to egocentric methods would also help separate errors caused by hand reconstruction from those caused by egomotion estimation. In addition, supplying object pose as an auxiliary input could test whether scene and object information improve the recovery of grasps and hand motion. We believe that ACE-Data-0 will support further research on combining egocentric and exocentric views for more accurate pose estimation.
6 Conclusion
We presented ACE, an ambient capture methodology for recording everyday household behavior in real homes. ACE uses two complementary configurations at different spatial scales: a table-scale setup for fine-grained hand-object manipulation and a room-scale setup for whole-body activity across a fully furnished apartment. Both capture synchronized egocentric and exocentric video, body, hand, and object motion, audio, and tactile signals. An optical-clock procedure aligns all streams to the motion-capture clock at millisecond precision, while marker-bridged calibration registers static and wearable cameras in a common world frame.
Using ACE, we collected ACE-Data-0, comprising over 150 hours of recording, 75,000 interaction episodes, 17M frames, and per-frame annotations. We benchmarked existing methods on touch prediction from video, full-body motion recovery, and hand-motion estimation from egocentric and exocentric views. Across these tasks, existing methods degrade under contact, occlusion, and long-duration activity—conditions that distinguish real homes from controlled laboratories. These results motivate models that fuse views and modalities, enforce physical constraints, and learn from measured contact and motion supervision.
By combining egocentric observations, demonstration trajectories, and contact-level supervision in a single time-aligned stream, ACE-Data-0 is designed to support research on manipulation policies, world models, and vision-language-action systems that connect perception, action, and physical state in real homes.
Limitations.
ACE has several limitations. First, it covers only two sites and therefore captures limited variation in layouts, furnishings, and lighting. Second, ground truth is restricted to instrumented entities: tracked objects must be scanned and equipped with markers in advance, and the dataset does not annotate state changes of articulated mechanisms, fluids, or deformable materials. Third, the suit, gloves, headset, and markers remain visible in the recordings and may introduce dataset-specific visual cues.
Ethics statement.
All participants volunteered and provided informed consent for both recording and public data release.
References
- [1] AgiBot-World-Contributors, Q. Bu, J. Cai, L. Chen, X. Cui, Y. Ding, et al. (2025) AgiBot World Colosseo: A large-scale manipulation platform for scalable and intelligent embodied systems. In RSS, Cited by: Table 1, §2.3.
- [2] M. Ahn, A. Brohan, N. Brown, Y. Chebotar, O. Cortes, et al. (2022) Do as I can, not as I say: grounding language in robotic affordances. In CoRL, Cited by: §2.3, §2.3.
- [3] P. Banerjee, S. Shkodrani, P. Moulon, S. Hampali, S. Han, F. Zhang, L. Zhang, J. Fountain, E. Miller, S. Basol, R. Newcombe, R. Wang, J. J. Engel, and T. Hodan (2025) HOT3D: hand and object tracking in 3D from egocentric multi-view videos. In CVPR, Cited by: Table 1, §1, §2.1, §2.2.
- [4] F. Baradel, M. Armando, S. Galaaoui, R. Brégier, P. Weinzaepfel, G. Rogez, and T. Lucas (2024) Multi-HMR: multi-person whole-body human mesh recovery in a single shot. In ECCV, Cited by: §5.2.1, Table 4.
- [5] B. L. Bhatnagar, X. Xie, I. A. Petrov, C. Sminchisescu, C. Theobalt, and G. Pons-Moll (2022) BEHAVE: dataset and method for tracking human object interactions. In CVPR, Cited by: Table 1, §1, §2.1.
- [6] K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. R. Equi, C. Finn, N. Fusai, M. Y. Galliker, D. Ghosh, L. Groom, K. Hausman, brian ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A. Z. Ren, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Walke, A. Walling, H. Wang, L. Yu, and U. Zhilinsky (2025) : A vision-language-action model with open-world generalization. In CoRL, Cited by: §1.
- [7] S. Brahmbhatt, C. Tang, C. D. Twigg, C. C. Kemp, and J. Hays (2020) ContactPose: A dataset of grasps with object contact and hand pose. In ECCV, Cited by: Table 1, §1, §2.1, §2.1, §4.2.1, §4.2.3.
- [8] T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei (2020) Language models are few-shot learners. In NeurIPS, Cited by: §1.
- [9] Q. Bu, Y. Yang, J. Cai, S. Gao, G. Ren, M. Yao, P. Luo, and H. Li (2025) UniVLA: learning to act anywhere with task-centric latent actions. In RSS, Cited by: §2.2.
- [10] X. Cai, R. Qiu, G. Chen, L. Wei, I. Liu, T. Huang, X. Cheng, and X. Wang (2025) In-N-On: scaling egocentric manipulation with in-the-wild and on-task data. arXiv 2511.15704. Cited by: §2.2.
- [11] Z. Cai, W. Yin, A. Zeng, C. Wei, Q. Sun, W. Yanjun, H. E. Pang, H. Mei, M. Zhang, L. Zhang, C. C. Loy, L. Yang, and Z. Liu (2023) SMPLer-X: scaling up expressive human pose and shape estimation. In NeurIPS, Cited by: §5.2.1, Table 4.
- [12] Y. Cao, J. Lu, Z. Huang, Z. Shen, C. Zhao, F. Hong, Z. Chen, X. Li, W. Wang, Y. Liu, and Z. Liu (2025) Reconstructing 4D spatial intelligence: A survey. arXiv 2507.21045. Cited by: §2.1.
- [13] Y. Cao, L. Pan, K. Han, K. K. Wong, and Z. Liu (2025) AvatarGO: zero-shot 4D human-object interaction generation and animation. In ICLR, Cited by: §2.3.
- [14] Y. Cao, H. Xie, F. Hong, L. Zhuo, Z. Chen, L. Pan, and Z. Liu (2026) HSImul3R: physics-in-the-loop reconstruction of simulation-ready human-scene interactions. In ECCV, Cited by: §2.3.
- [15] Y. Chao, W. Yang, Y. Xiang, P. Molchanov, A. Handa, J. Tremblay, Y. S. Narang, K. V. Wyk, U. Iqbal, S. Birchfield, J. Kautz, and D. Fox (2021) DexYCB: A benchmark for capturing hand grasping of objects. In CVPR, Cited by: Table 1, §1, §2.1, §4.2.1, §4.2.3.
- [16] Y. Chen, X. Chen, Y. Xue, A. Chen, Y. Xiu, and G. Pons-Moll (2026) Human3R: everyone everywhere all at once. In ICLR, Cited by: §5.2.1, Table 4.
- [17] Z. Chen, R. A. Potamias, S. Chen, and C. Schmid (2025) HORT: monocular hand-held objects reconstruction with transformers. In ICCV, Cited by: §2.1, §5.3.2, Table 6.
- [18] H. Cuevas-Velasquez, A. Yiannakidis, S. Shin, G. Becherini, M. Höschle, J. Tesch, T. Obersat, T. Alexiadis, and M. J. Black (2026) MAMMA: markerless and automatic multi-person motion action capture. In CVPR, Cited by: §5.2.1, Table 4.
- [19] D. Damen, H. Doughty, G. M. Farinella, A. Furnari, E. Kazakos, J. Ma, D. Moltisanti, J. Munro, T. Perrett, W. Price, and M. Wray (2022) Rescaling egocentric vision: collection, pipeline and challenges for EPIC-KITCHENS-100. IJCV 130 (1), pp. 33–55. Cited by: Table 1, §1.
- [20] J. Deng, W. Dong, R. Socher, L. Li, K. Li, and L. Fei-Fei (2009) ImageNet: A large-scale hierarchical image database. In CVPR, Cited by: §1.
- [21] (2021) EasyMoCap: make human motion capture easier. Note: GitHub External Links: Link Cited by: Table 4.
- [22] Z. Fan, O. Taheri, D. Tzionas, M. Kocabas, M. Kaufmann, M. J. Black, and O. Hilliges (2023) ARCTIC: A dataset for dexterous bimanual hand-object manipulation. In CVPR, Cited by: Table 1, §1, §2.1, §4.2.1, §4.2.3, §5.3.2.
- [23] H. Fang, H. Fang, Z. Tang, J. Liu, C. Wang, J. Wang, H. Zhu, and C. Lu (2024) RH20T: A comprehensive robotic dataset for learning diverse skills in one-shot. In ICRA, Cited by: Table 1.
- [24] G. Fiche, P. Weinzaepfel, R. Brégier, and F. Baradel (2026) Multi-HMR 2: multi-person camera-centric human detection, mesh recovery and tracking. arXiv 2606.14841. Cited by: §5.2.1, Table 4.
- [25] Fourier ActionNet Team and Y. Mu (2025) ActionNet: A dataset for dexterous bimanual manipulation. Note: Dataset website External Links: Link Cited by: §2.2.
- [26] R. Fu, D. Zhang, A. Jiang, W. Fu, A. Funk, D. Ritchie, and S. Sridhar (2025) GigaHands: A massive annotated dataset of bimanual hand activities. In CVPR, Cited by: Table 1, §1, §2.1, §5.3.2.
- [27] Z. Fu, T. Z. Zhao, and C. Finn (2024) Mobile ALOHA: learning bimanual mobile manipulation with low-cost whole-body teleoperation. In CoRL, Cited by: §2.3.
- [28] P. Furgale, J. Rehder, and R. Siegwart (2013) Unified temporal and spatial calibration for multi-sensor systems. In IROS, Cited by: §4.1.2.
- [29] S. Garrido-Jurado, R. Muñoz-Salinas, F.J. Madrid-Cuevas, and M.J. Marín-Jiménez (2014) Automatic generation and detection of highly reliable fiducial markers under occlusion. PR 47 (6), pp. 2280–2292. Cited by: §4.1.2.
- [30] J. J. Gibson (2014) The ecological approach to visual perception: classic edition. Psychology press. Cited by: §1.
- [31] S. Goel, G. Pavlakos, J. Rajasegaran, A. Kanazawa, and J. Malik (2023) Humans in 4D: reconstructing and tracking humans with transformers. In ICCV, Cited by: §5.2.1, Table 4.
- [32] P. Grady, C. Tang, S. Brahmbhatt, C. D. Twigg, C. Wan, J. Hays, and C. C. Kemp (2022) PressureVision: estimating hand pressure from a single RGB image. In ECCV, Cited by: §5.1.1, Table 3.
- [33] K. Grauman, A. Westbury, E. Byrne, et al. (2022) Ego4D: around the world in 3,000 hours of egocentric video. In CVPR, Cited by: Table 1, §2.2.
- [34] K. Grauman, A. Westbury, L. Torresani, K. Kitani, J. Malik, T. Afouras, et al. (2024) Ego-Exo4D: understanding skilled human activity from first- and third-person perspectives. In CVPR, Cited by: Table 1, §1, §2.2.
- [35] C. Gu, M. Zhang, H. Xie, Z. Cai, L. Yang, and Z. Liu (2026) Bridging semantic and kinematic conditions with diffusion-based discrete motion tokenizer. arXiv 2603.19227. Cited by: §2.3.
- [36] S. Hampali, M. Rad, M. Oberweger, and V. Lepetit (2020) HOnnotate: A method for 3D annotation of hand and object poses. In CVPR, Cited by: Table 1.
- [37] S. Hampali, S. D. Sarkar, M. Rad, and V. Lepetit (2022) Keypoint Transformer: solving joint identification in challenging hands and object interactions for accurate 3D pose estimation. In CVPR, Cited by: §2.1.
- [38] R. Hoque, P. Huang, D. J. Yoon, M. Sivapurapu, and J. Zhang (2026) EgoDex: learning dexterous manipulation from large-scale egocentric video. In ICLR, Cited by: Table 1, §2.2.
- [39] C. Hou, K. Wu, J. Liu, Z. Che, D. Wu, et al. (2025) RoboMIND 2.0: A multimodal, bimanual mobile manipulation dataset for generalizable embodied intelligence. arXiv 2512.24653. Cited by: §2.3.
- [40] B. Huang, Z. Yu, A. Chen, A. Geiger, and S. Gao (2024) 2D Gaussian splatting for geometrically accurate radiance fields. In SIGGRAPH, Cited by: §3.2.2.
- [41] Y. Huang, O. Taheri, M. J. Black, and D. Tzionas (2024) InterCap: joint markerless 3D tracking of humans and objects in interaction. IJCV 132 (7), pp. 2551–2566. Cited by: Table 1, §2.1, §2.1.
- [42] C. Ionescu, D. Papava, V. Olaru, and C. Sminchisescu (2014) Human3.6M: large scale datasets and predictive methods for 3D human sensing in natural environments. IEEE TPAMI 36 (7), pp. 1325–1339. Cited by: §5.2.1.
- [43] H. Jiang, J. Chen, Q. Bu, L. Chen, M. Shi, Y. Zhang, D. Li, C. Suo, C. Wang, Z. Peng, and H. Li (2026) WholeBodyVLA: towards unified latent VLA for whole-body loco-manipulation control. In ICLR, Cited by: §2.3.
- [44] N. Jiang, T. Liu, Z. Cao, J. Cui, Z. Zhang, Y. Chen, H. Wang, Y. Zhu, and S. Huang (2023) Full-body articulated human-object interaction. In ICCV, Cited by: Table 1.
- [45] N. Jiang, Z. Zhang, H. Li, X. Ma, Z. Wang, Y. Chen, T. Liu, Y. Zhu, and S. Huang (2024) Scaling up dynamic human-scene interaction modeling. In CVPR, Cited by: Table 1.
- [46] T. Jiang, T. Yuan, Y. Liu, C. Lu, J. Cui, X. Liu, S. Cheng, J. Gao, H. Xu, and H. Zhao (2025) Galaxea open-world dataset and G0 dual-system VLA model. arXiv 2509.00576. Cited by: Table 1, §2.3.
- [47] S. Kareer, D. Patel, R. Punamiya, P. Mathur, S. Cheng, C. Wang, J. Hoffman, and D. Xu (2025) EgoMimic: scaling imitation learning via egocentric video. In ICRA, Cited by: §2.2, §5.3.
- [48] S. Kareer, K. Pertsch, J. Darpinian, J. Hoffman, D. Xu, S. Levine, C. Finn, and S. Nair (2025) Emergence of human to robot transfer in vision-language-action models. arXiv 2512.22414. Cited by: §2.2.
- [49] A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, et al. (2024) DROID: A large-scale in-the-wild robot manipulation dataset. In RSS, Cited by: Table 1.
- [50] J. Kim, J. Kim, J. Na, and H. Joo (2024) ParaHome: parameterizing everyday home activities towards 3D generative modeling of human-object interactions. In CVPR, Cited by: Table 1, §2.3, §2.3.
- [51] M. Kocabas, C. P. Huang, O. Hilliges, and M. J. Black (2021) PARE: part attention regressor for 3D human body estimation. In ICCV, Cited by: §5.2.1, Table 4.
- [52] T. Kwon, B. Tekin, J. Stuhmer, F. Bogo, and M. Pollefeys (2021) H2O: two hands manipulating objects for first person interaction recognition. In ICCV, Cited by: Table 1, §1, §2.1, §2.2, §4.2.1, §4.2.3.
- [53] J. Li, C. K. Liu, and J. Wu (2023) Ego-body pose estimation via ego-head pose estimation. In CVPR, Cited by: §5.2.1, Table 4.
- [54] J. Li, J. Wu, and C. K. Liu (2023) Object motion guided human motion synthesis. ACM TOG 42 (6), pp. 197:1–197:11. Cited by: Table 1, §2.3.
- [55] M. Li, P. Li, Z. Zhang, J. Lu, C. Zhao, W. Xue, Q. Liu, S. Peng, W. Zhang, W. Luo, Y. Liu, and Y. Guo (2026) UniSH: unifying scene and human reconstruction in a feed-forward pass. arXiv 2601.01222. Cited by: §5.2.1, Table 4.
- [56] R. Li, Z. Hu, L. Siyao, Y. Zhang, H. Xie, M. Zhang, J. Guo, X. Li, and Z. Liu (2026) InfiniteDance: scalable 3D dance generation towards in-the-wild generalization. In ECCV, Cited by: §2.3.
- [57] X. Li, M. Meng, Z. Wu, T. Chen, F. Yang, and D. Shen (2024) Human mesh recovery from arbitrary multi-view images. arXiv 2403.12434. Cited by: §5.2.1, Table 4.
- [58] D. Lin, Y. Zhang, M. Li, Y. Liu, W. Jing, Q. Yan, Q. Wang, and H. Zhang (2026) OmniHands: robust motion capture of interactive hands via a versatile transformer. ACM TOG 42 (6), pp. 197:1–197:11. Cited by: §5.3.2, Table 6.
- [59] J. Lin, A. Zeng, H. Wang, L. Zhang, and Y. Li (2023) One-stage 3D whole-body mesh recovery with component aware transformer. In CVPR, Cited by: §5.2.1, Table 4.
- [60] Y. Liu, H. Yang, X. Si, L. Liu, Z. Li, Y. Zhang, Y. Liu, and L. Yi (2024) TACO: benchmarking generalizable bimanual Tool-ACtion-Object understanding. In CVPR, Cited by: Table 1, §2.1.
- [61] Y. Liu, Y. Liu, C. Jiang, K. Lyu, W. Wan, H. Shen, B. Liang, Z. Fu, H. Wang, and L. Yi (2022) HOI4D: A 4D egocentric dataset for category-level human-object interaction. In CVPR, Cited by: Table 1, §2.1, §2.2.
- [62] Z. Liu, J. Lin, W. Wu, and B. Zhou (2025) Joint optimization for 4D human-scene reconstruction in the wild. arXiv 2501.02158. Cited by: §5.2.1, Table 4.
- [63] M. Loper, N. Mahmood, J. Romero, G. Pons-Moll, and M. J. Black (2015) SMPL: a skinned multi-person linear model. ACM TOG 34 (6), pp. 248:1–248:16. Cited by: §5.2.1.
- [64] J. Lu, C. P. Huang, U. Bhattacharya, Q. Huang, and Y. Zhou (2025) HUMOTO: A 4D dataset of mocap human object interactions. In ICCV, Cited by: Table 1, §2.3.
- [65] X. Lv, L. Xu, Y. Yan, X. Jin, C. Xu, S. Wu, Y. Liu, L. Li, M. Bi, W. Zeng, and X. Yang (2024) HIMO: A new benchmark for full-body human interacting with multiple objects. In ECCV, Cited by: Table 1.
- [66] L. Ma, Y. Ye, F. Hong, V. Guzov, Y. Jiang, R. Postyeni, L. Pesqueira, A. Gamino, V. Baiyya, H. J. Kim, K. Bailey, D. S. Fosas, C. K. Liu, Z. Liu, J. J. Engel, R. De Nardi, and R. A. Newcombe (2024) Nymeria: A massive collection of multimodal egocentric daily motion in the wild. In ECCV, Cited by: Table 1.
- [67] L. Müller, H. Choi, A. Zhang, B. Yi, J. Malik, and A. Kanazawa (2025) Reconstructing people, places, and cameras. In CVPR, Cited by: §5.2.1, Table 4.
- [68] P. Y. Muralidhar, Y. Xue, X. Xie, M. Kostyrko, and G. Pons-Moll (2025) PhySIC: physically plausible 3D human-scene interaction and contact from a single image. In SIGGRAPH Asia, Cited by: §2.1, §5.2.1, §5.2.1, Table 4.
- [69] X. Pan, N. Charron, Y. Yang, S. Peters, T. Whelan, C. Kong, O. M. Parkhi, R. A. Newcombe, and C. Y. Ren (2023) Aria Digital Twin: A new benchmark dataset for egocentric 3D machine perception. In ICCV, Cited by: Table 1.
- [70] P. Patel and M. J. Black (2025) CameraHMR: aligning people with perspective. In 3DV, Cited by: §5.2.1, Table 4.
- [71] G. Pavlakos, V. Choutas, N. Ghorbani, T. Bolkart, A. A. A. Osman, D. Tzionas, and M. J. Black (2019) Expressive body capture: 3D hands, face, and body from a single image. In CVPR, Cited by: 3rd item, §5.2.1.
- [72] G. Pavlakos, D. Shan, I. Radosavovic, A. Kanazawa, D. Fouhey, and J. Malik (2024) Reconstructing hands in 3D with transformers. In CVPR, Cited by: §2.1, §5.3.2, Table 6.
- [73] T. Perrett, A. Darkhalil, S. Sinha, O. Emara, S. Pollard, K. Parida, K. Liu, P. Gatti, S. Bansal, K. Flanagan, J. Chalk, Z. Zhu, R. Guerrier, F. Abdelazim, B. Zhu, D. Moltisanti, M. Wray, H. Doughty, and D. Damen (2025) HD-EPIC: A highly-detailed egocentric video dataset. In CVPR, Cited by: Table 1.
- [74] R. Pfeifer and J. Bongard (2006) How the body shapes the way we think: a new view of intelligence. MIT press. Cited by: §1.
- [75] R. A. Potamias, J. Zhang, J. Deng, and S. Zafeiriou (2025) WiLoR: end-to-end 3D hand localization and reconstruction in-the-wild. In CVPR, Cited by: §5.3.2, Table 6.
- [76] A. Prakash, R. Tu, M. Chang, and S. Gupta (2024) 3D hand pose estimation in everyday egocentric images. In ECCV, Cited by: §5.3.1, Table 5.
- [77] R. Punamiya, D. Patel, P. Aphiwetsa, P. Kuppili, L. Y. Zhu, S. Kareer, J. Hoffman, and D. Xu (2025) EgoBridge: domain adaptation for generalizable imitation from egocentric human data. In NeurIPS, Cited by: §2.2, §5.3.
- [78] Y. Qin, Y. Wu, S. Liu, H. Jiang, R. Yang, Y. Fu, and X. Wang (2022) DexMV: imitation learning for dexterous manipulation from human videos. In ECCV, Cited by: §2.2, §5.3.
- [79] R. Qiu, S. Yang, X. Cheng, C. Chawla, J. Li, T. He, G. Yan, L. Paulsen, G. Yang, S. Yi, G. Shi, and X. Wang (2025) Humanoid policy human policy. In CoRL, Cited by: §2.2.
- [80] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. (2021) Learning transferable visual models from natural language supervision. In ICML, Cited by: §1.
- [81] J. Romero, D. Tzionas, and M. J. Black (2017) Embodied hands: modeling and capturing hands and bodies together. ACM TOG 36 (6), pp. 245:1–245:17. Cited by: §2.1, §5.3.1.
- [82] Ropedia (2026) Xperience-10M: A large-scale egocentric multimodal dataset with structured 3D/4D annotations. Note: Hugging Face dataset External Links: Link Cited by: Table 1, §1.
- [83] C. Schuhmann, R. Beaumont, R. Vencu, C. Gordon, R. Wightman, M. Cherti, T. Coombes, A. Katta, C. Mullis, M. Wortsman, et al. (2022) LAION-5B: an open large-scale dataset for training next generation image-text models. In NeurIPS, Cited by: §1.
- [84] Z. Shen, H. Pi, Y. Xia, Z. Cen, S. Peng, Z. Hu, H. Bao, R. Hu, and X. Zhou (2024) World-grounded human motion recovery via gravity-view coordinates. In SIGGRAPH Asia, Cited by: §5.2.1, §5.2.1, §5.2.1, Table 4.
- [85] S. Shin, J. Kim, E. Halilaj, and M. J. Black (2024) WHAM: reconstructing world-grounded humans with accurate 3D motion. In CVPR, Cited by: §5.2.1, §5.2.1, §5.2.1, Table 4.
- [86] Q. Shuai, C. Geng, Q. Fang, S. Peng, W. Shen, X. Zhou, and H. Bao (2022) Novel view synthesis of human interactions from sparse multi-view videos. In SIGGRAPH, Cited by: §5.2.1, Table 4.
- [87] Y. Su, S. Chen, H. Shi, M. Liu, Z. Zhang, N. Huang, W. Zhong, Z. Zhu, Y. Liu, and X. Liu (2026) World Guidance: world modeling in condition space for action generation. arXiv 2602.22010. Cited by: §1.
- [88] O. Taheri, N. Ghorbani, M. J. Black, and D. Tzionas (2020) GRAB: A dataset of whole-body human grasping of objects. In ECCV, Cited by: Table 1, §1, §2.1.
- [89] G. Team (2025) Gemini: A family of highly capable multimodal models. arXiv 2312.11805. Cited by: §4.3.1.
- [90] M. Torne, K. Pertsch, H. Walke, K. Vedder, S. Nair, B. Ichter, A. Z. Ren, H. Wang, J. Tang, K. Stachowicz, K. Dhabalia, M. Equi, Q. Vuong, J. T. Springenberg, S. Levine, C. Finn, and D. Driess (2026) MEM: multi-scale embodied memory for vision language action models. arXiv 2603.03596. Cited by: §2.3, §2.3.
- [91] B. Triggs, P. F. McLauchlan, R. I. Hartley, and A. W. Fitzgibbon (2000) Bundle adjustment — A modern synthesis. In Vision Algorithms: Theory and Practice, Cited by: §4.1.2.
- [92] R. Y. Tsai and R. Lenz (1989) A new technique for fully autonomous and efficient 3D robotics hand/eye calibration. IEEE T-RA 5 (3), pp. 345–358. Cited by: §4.1.2.
- [93] H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V. Myers, M. J. Kim, M. Du, A. Lee, K. Fang, C. Finn, and S. Levine (2023) BridgeData V2: A dataset for robot learning at scale. In CoRL, Cited by: Table 1.
- [94] C. Wang, H. Shi, W. Wang, R. Zhang, L. Fei-Fei, and C. K. Liu (2024) DexCap: scalable and portable mocap data collection system for dexterous manipulation. In RSS, Cited by: Table 1.
- [95] X. Wang, T. Kwon, M. Rad, B. Pan, I. Chakraborty, S. Andrist, D. Bohus, A. Feniello, B. Tekin, F. V. Frujeri, N. Joshi, and M. Pollefeys (2023) HoloAssist: an egocentric human interaction dataset for interactive AI assistants in the real world. In ICCV, Cited by: Table 1, §2.2.
- [96] B. Wen, W. Yang, J. Kautz, and S. Birchfield (2024) FoundationPose: unified 6D pose estimation and tracking of novel objects. In CVPR, Cited by: §5.2.
- [97] S. Wu, X. Liu, S. Xie, P. Wang, X. Li, et al. (2025) RoboCOIN: an open-sourced bimanual robotic data COllection for INtegrated manipulation. arXiv 2511.17441. Cited by: §2.3.
- [98] H. Xie, B. Wen, J. Zheng, Z. Chen, F. Hong, H. Diao, and Z. Liu (2026) DynamicVLA: A vision-language-action model for dynamic object manipulation. arXiv 2601.22153. Cited by: §2.3.
- [99] X. Xu, J. Park, H. Zhang, E. Cousineau, A. Bhat, J. Barreiros, D. Wang, and S. Song (2026) HoMMI: learning whole-body mobile manipulation from human demonstrations. arXiv 2603.03243. Cited by: §2.3.
- [100] J. Yang, S. Liu, H. Guo, Y. Dong, X. Zhang, S. Zhang, P. Wang, Z. Zhou, B. Xie, Z. Wang, B. Ouyang, Z. Lin, M. Cominelli, Z. Cai, B. Li, Y. Zhang, P. Zhang, F. Hong, J. Widmer, F. Gringoli, L. Yang, and Z. Liu (2025) EgoLife: towards egocentric life assistant. In CVPR, Cited by: Table 1.
- [101] L. Yang, K. Li, X. Zhan, F. Wu, A. Xu, L. Liu, and C. Lu (2022) OakInk: A large-scale knowledge repository for understanding hand-object interaction. In CVPR, Cited by: Table 1.
- [102] R. Yang, Q. Yu, Y. Wu, R. Yan, B. Li, A. Cheng, X. Zou, Y. Fang, X. Cheng, R. Qiu, H. Yin, S. Liu, S. Han, Y. Lu, and X. Wang (2025) EgoVLA: learning vision-language-action models from egocentric human videos. In CoRL, Cited by: §1, §2.2.
- [103] X. Yang, D. Kukreja, D. Pinkus, A. Sagar, T. Fan, J. Park, S. Shin, J. Cao, J. Liu, N. Ugrinovic, M. Feiszli, J. Malik, P. Dollár, and K. Kitani (2026) SAM 3D Body: robust full-body human mesh recovery. arXiv 2602.15989. Cited by: §5.2.1, Table 4.
- [104] S. Ye, Y. Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y. L. Tan, C. Zhu, J. Xiang, A. Malik, K. Lee, W. Liang, N. Ranawaka, J. Gu, Y. Xu, G. Wang, F. Hu, A. Narayan, J. Bjorck, J. Wang, G. Kim, D. Niu, R. Zheng, Y. Xie, J. Wu, Q. Wang, R. Julian, D. Xu, Y. Du, Y. Chebotar, S. Reed, J. Kautz, Y. Zhu, L. ”. Fan, and J. Jang (2026) World action models are zero-shot policies. arXiv 2602.15922. Cited by: §1.
- [105] Y. Ye, Y. Feng, O. Taheri, H. Feng, S. Tulsiani, and M. J. Black (2026) Predicting 4D hand trajectory from monocular videos. In 3DV, Cited by: §2.1, §5.3.2, Table 6.
- [106] B. Yi, V. Ye, M. Zheng, Y. Li, L. Müller, G. Pavlakos, Y. Ma, J. Malik, and A. Kanazawa (2025) Estimating body and hand motion in an ego-sensed world. In CVPR, Cited by: §5.2.1, Table 4.
- [107] W. Yin, Z. Cai, R. Wang, A. Zeng, C. Wei, Q. Sun, H. Mei, Y. Wang, H. E. Pang, M. Zhang, L. Zhang, C. C. Loy, A. Yamashita, L. Yang, and Z. Liu (2026) SMPLest-X: ultimate scaling for expressive human pose and shape estimation. IEEE TPAMI 48 (2), pp. 1778–1794. Cited by: §5.2.1, Table 4.
- [108] Z. Yu, S. Zafeiriou, and T. Birdal (2025) Dyn-HaMR: recovering 4D interacting hand motion from a dynamic camera. In CVPR, Cited by: §5.3.1, Table 5.
- [109] Y. Zeng, Y. Shi, T. Tan, X. Li, Y. Qin, Z. Lu, W. Yang, J. Xue, and Q. Liao (2026) EgoTactile: learning grasp pressure for everyday objects from egocentric video. In ICML, Cited by: §5.1.1, Table 3.
- [110] X. Zhan, L. Yang, Y. Zhao, K. Mao, H. Xu, Z. Lin, K. Li, and C. Lu (2024) OAKINK2: A dataset of bimanual hands-object manipulation in complex task completion. In CVPR, Cited by: Table 1, §1, §2.1, §2.3, §2.3.
- [111] H. Zhang, Y. Tian, Y. Zhang, M. Li, L. An, Z. Sun, and Y. Liu (2023) PyMAF-X: towards well-aligned full-body model regression from monocular images. IEEE TPAMI 45 (10), pp. 12287–12303. Cited by: §5.2.1, Table 4.
- [112] J. Zhang, J. Deng, C. Ma, and R. A. Potamias (2025) HaWoR: world-space hand motion reconstruction from egocentric videos. In CVPR, Cited by: §2.1, §5.3.1, Table 5.
- [113] S. Zhang, Q. Ma, Y. Zhang, Z. Qian, T. Kwon, M. Pollefeys, F. Bogo, and S. Tang (2022) EgoBody: human body shape and motion of interacting people from head-mounted devices. In ECCV, Cited by: Table 1, §5.2.1.
- [114] R. Zheng, D. Niu, Y. Xie, J. Wang, M. Xu, Y. Jiang, F. Castañeda, F. Hu, Y. L. Tan, L. Fu, T. Darrell, F. Huang, Y. Zhu, D. Xu, and L. Fan (2026) EgoScale: scaling dexterous manipulation with diverse egocentric human data. arXiv 2602.16710. Cited by: Table 1, §2.2.
- [115] J. Zhou, Z. Gao, F. Hong, Z. Liu, G. Zhang, W. Dai, R. Zhen, C. Lyu, H. Wu, Y. Mao, X. Wang, Y. Jiang, W. Ding, and S. Yang (2026) TouchAnything: A dataset and framework for bimanual tactile estimation from egocentric video. arXiv 2605.13083. Cited by: §5.1.1, Table 3.
- [116] L. Y. Zhu, P. Kuppili, R. Punamiya, P. Aphiwetsa, D. Patel, S. Kareer, S. Ha, and D. Xu (2025) EMMA: scaling mobile manipulation via egocentric human data. IEEE RA-L 11 (3), pp. 3087–3094. Cited by: §2.3.