{"work":{"id":"92a2f39f-9dd2-401e-a854-aab6064581f0","openalex_id":null,"doi":null,"arxiv_id":"2410.24221","raw_key":null,"title":"EgoMimic: Scaling Imitation Learning via Egocentric Video","authors":null,"authors_text":"Simar Kareer, Dhruv Patel, Ryan Punamiya, Pranay Mathur, Shuo Cheng, Chen Wang, Judy Hoffman, and Danfei Xu","year":2024,"venue":"cs.RO","abstract":"The scale and diversity of demonstration data required for imitation learning is a significant challenge. We present EgoMimic, a full-stack framework which scales manipulation via human embodiment data, specifically egocentric human videos paired with 3D hand tracking. EgoMimic achieves this through: (1) a system to capture human embodiment data using the ergonomic Project Aria glasses, (2) a low-cost bimanual manipulator that minimizes the kinematic gap to human data, (3) cross-domain data alignment techniques, and (4) an imitation learning architecture that co-trains on human and robot data. Compared to prior works that only extract high-level intent from human videos, our approach treats human and robot data equally as embodied demonstration data and learns a unified policy from both data sources. EgoMimic achieves significant improvement on a diverse set of long-horizon, single-arm and bimanual manipulation tasks over state-of-the-art imitation learning methods and enables generalization to entirely new scenes. Finally, we show a favorable scaling trend for EgoMimic, where adding 1 hour of additional hand data is significantly more valuable than 1 hour of additional robot data. Videos and additional information can be found at https://egomimic.github.io/","external_url":"https://arxiv.org/abs/2410.24221","cited_by_count":null,"metadata_source":"pith","metadata_fetched_at":"2026-07-10T18:47:31.552848+00:00","pith_arxiv_id":"2410.24221","created_at":"2026-05-11T05:10:54.135798+00:00","updated_at":"2026-07-10T18:47:31.552848+00:00","title_quality_ok":true,"display_title":"Egomimic: Scaling imitation learning via egocentric video","render_title":"Egomimic: Scaling imitation learning via egocentric video"},"hub":{"state":{"work_id":"92a2f39f-9dd2-401e-a854-aab6064581f0","tier":"hub","tier_reason":"10+ Pith inbound or 1,000+ external citations","pith_inbound_count":23,"external_cited_by_count":null,"distinct_field_count":2,"first_pith_cited_at":"2025-07-01T17:39:59+00:00","last_pith_cited_at":"2026-07-09T16:15:43+00:00","author_build_status":"not_needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"not_needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-22T07:09:39.110957+00:00","tier_text":"hub"},"tier":"hub","role_counts":[{"context_role":"dataset","n":4},{"context_role":"background","n":3}],"polarity_counts":[{"context_polarity":"background","n":5},{"context_polarity":"use_dataset","n":2}],"runs":{},"summary":{},"graph":{},"authors":[]}}