Pith. sign in

REVIEW 25 cited by

Ego4D: Around the World in 3,000 Hours of Egocentric Video

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.07058 v3 pith:WZOPS6BK submitted 2021-10-13 cs.CV cs.AI

Ego4D: Around the World in 3,000 Hours of Egocentric Video

classification cs.CV cs.AI
keywords videoegocentricbenchmarkego4darounddatasetfirst-personhours
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

We introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Project page: https://ego4d-data.org/

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 25 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution

    cs.CL 2023-09 unverdicted novelty 8.0

    Promptbreeder evolves both task prompts and the mutation prompts that improve them using LLMs, outperforming Chain-of-Thought and Plan-and-Solve on arithmetic and commonsense reasoning benchmarks.

  2. Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion

    cs.CV 2026-05 unverdicted novelty 7.0

    Introduces pause-and-think-T dataset and pause-and-think-B benchmark; fine-tunes 4B VLM to 58% accuracy matching 235B model while generalizing out-of-distribution.

  3. HRDexDB: A Paired Human-Robot Dataset for Cross-Embodiment Dexterous Grasping

    cs.RO 2026-04 unverdicted novelty 7.0

    HRDexDB is a multi-modal dataset of 1.4K human and robotic dexterous grasps across 100 objects, providing aligned 3D kinematics, high-resolution tactile data, and video streams.

  4. HRDexDB: A Paired Human-Robot Dataset for Cross-Embodiment Dexterous Grasping

    cs.RO 2026-04 conditional novelty 7.0

    HRDexDB provides 2.1K paired markerless human and multi-robot dexterous grasp sequences on 100 objects with multi-view 3D ground truth and tactile signals.

  5. How Can AI Augment Access to Justice? Public Defenders' Perspectives on AI Adoption

    cs.CY 2025-10 accept novelty 7.0

    Public defenders view AI as most useful for evidence investigation but limited in courtroom work and strategy, with adoption blocked by costs, confidentiality risks, and norms, requiring human oversight and open development.

  6. Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

    cs.CV 2022-04 unverdicted novelty 7.0

    Socratic Models compose zero-shot multimodal reasoning by prompting pretrained language and vision models to exchange information and enable new capabilities without finetuning.

  7. EgoSim: Egocentric World Simulator for Embodied Interaction Generation

    cs.CV 2026-04 conditional novelty 6.5

    EgoSim generates spatially consistent egocentric interaction videos by conditioning a video diffusion model on updatable 3D point-cloud states and action keypoints extracted at scale from monocular videos.

  8. HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

    cs.RO 2026-07 conditional novelty 6.0

    Robot-free HiFi-UMI demonstrations can replace teleoperated real-robot data in post-training: three policy backbones matched in-domain teleoperation within 3.1 percentage points, including 85% success on a precision i...

  9. MEMORA: Embodied Action Memory from Egocentric Videos for Reasoning and Planning

    cs.RO 2026-07 conditional novelty 6.0

    A typed, editable memory built from egocentric video improves memory-grounded question answering and out-of-distribution robot planning over flat-text and graph baselines.

  10. Towards Real-World Wearable Motion Reconstruction

    cs.CV 2026-07 conditional novelty 6.0

    A consumer-wearable MoCap dataset plus WHIP, a flow-matching model that reconstructs full-body motion from arbitrary sensor subsets and quantifies sensor complementarity.

  11. PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning

    cs.RO 2026-06 unverdicted novelty 6.0

    PoLAR imposes radial structure on latent actions in hyperbolic space to factorize extent and mode, improving robot policy performance over baselines.

  12. Contrastive Action-Image Pre-training for Visuomotor Control

    cs.RO 2026-06 unverdicted novelty 6.0

    CAIP learns action-aligned visual representations via contrastive pre-training on human hand keypoints from egocentric video, outperforming DINOv2, SigLIP, MVP, and R3M with >30% gains on real dexterous manipulation tasks.

  13. Pause and Think: A Dataset and Benchmark for Video-Grounded Assistive Action Suggestion

    cs.CV 2026-05 unverdicted novelty 6.0

    Introduces pause-and-think-T dataset and pause-and-think-B benchmark for video-grounded assistive action suggestion, enabling a 4B VLM to match larger models on reasoning tasks and generalize to EgoThink and TempCompass.

  14. HumanNet: Scaling Human-centric Video Learning to One Million Hours

    cs.CV 2026-05 unverdicted novelty 6.0

    HumanNet is a 1M-hour human-centric video dataset with interaction annotations that enables better vision-language-action model performance than equivalent robot data in a controlled test.

  15. WARPED: Wrist-Aligned Rendering for Robot Policy Learning from Egocentric Human Demonstrations

    cs.RO 2026-04 unverdicted novelty 6.0

    WARPED synthesizes realistic wrist-view observations from monocular egocentric human videos via foundation models, hand-object tracking, retargeting, and Gaussian Splatting to train visuomotor policies that match tele...

  16. RoSHI: A Versatile Robot-oriented Suit for Human Data In-the-Wild

    cs.RO 2026-04 unverdicted novelty 6.0

    RoSHI is a hybrid wearable that combines sparse IMUs and egocentric SLAM to capture accurate full-body 3D pose and shape data in natural environments for robot learning.

  17. Spotlighting Task-Relevant Features: Object-Centric Representations for Better Generalization in Robotic Manipulation

    cs.RO 2026-01 conditional novelty 6.0

    Slot-based object-centric visual representations, especially with robot-video pretraining, improve out-of-distribution generalization of robotic manipulation policies compared to global and dense pre-trained features.

  18. HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding

    cs.CV 2026-01 unverdicted novelty 6.0

    HERMES organizes the KV cache into a hierarchical memory to enable real-time streaming video understanding in MLLMs, achieving 10x faster TTFT and up to 11.4% accuracy gains on streaming benchmarks with 68% fewer tokens.

  19. How Can AI Augment Access to Justice? Public Defenders' Perspectives on AI Adoption

    cs.CY 2025-10 conditional novelty 6.0

    In 14 interviews, U.S. public defense professionals said AI is most helpful for evidence investigation and least helpful for courtroom advocacy and strategy.

  20. DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory

    cs.CV 2023-08 unverdicted novelty 6.0

    DragNUWA integrates text, image, and trajectory controls into a diffusion video model using a Trajectory Sampler, Multiscale Fusion, and Adaptive Training to enable fine-grained open-domain video generation.

  21. iFLYTEK-Embodied-Omni Technical Report

    cs.AI 2026-06 conditional novelty 5.0

    A three-branch Omni model (VLM+VGM brain, AGM cerebellum) with shared multimodal attention and four-stage training reaches 89.6% zero-shot on LIBERO-Plus and ~93% on RoboTwin 2.0.

  22. The Conversations Beneath the Code: Triadic Data for Long-Horizon Software Engineering Agents

    cs.SE 2026-05 unverdicted novelty 5.0

    Triadic data—synchronized human-human conversations, human-AI sessions, and cross-functional team work—is the essential substrate for training long-horizon software engineering agents.

  23. World Action Models: The Next Frontier in Embodied AI

    cs.RO 2026-05 unverdicted novelty 4.0

    The paper introduces World Action Models as a new paradigm unifying predictive world modeling with action generation in embodied foundation models and provides a taxonomy of existing approaches.

  24. Data Pyramid for Embodied Manipulation

    cs.RO 2026-07 conditional novelty 3.0

    Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.

  25. World Action Models: A Survey

    cs.RO 2026-06 unverdicted novelty 3.0

    A survey that clarifies boundaries and organizes World Action Models by generation requirements and predictive substrates, identifying a trend toward generating less of the future.