Pith. sign in

REVIEW 18 cited by

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.18259 v4 pith:6PPBQCZ5 submitted 2023-11-30 cs.CV cs.AI

classification cs.CVcs.AI
keywords videoactivityego-exo4dhumanskilledunderstandingactivitiesbenchmark
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present Ego-Exo4D, a diverse, large-scale multimodal multiview video dataset and benchmark challenge. Ego-Exo4D centers around simultaneously-captured egocentric and exocentric video of skilled human activities (e.g., sports, music, dance, bike repair). 740 participants from 13 cities worldwide performed these activities in 123 different natural scene contexts, yielding long-form captures from 1 to 42 minutes each and 1,286 hours of video combined. The multimodal nature of the dataset is unprecedented: the video is accompanied by multichannel audio, eye gaze, 3D point clouds, camera poses, IMU, and multiple paired language descriptions -- including a novel "expert commentary" done by coaches and teachers and tailored to the skilled-activity domain. To push the frontier of first-person video understanding of skilled human activity, we also present a suite of benchmark tasks and their annotations, including fine-grained activity understanding, proficiency estimation, cross-view translation, and 3D hand/body pose. All resources are open sourced to fuel new research in the community. Project page: http://ego-exo4d-data.org/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. The TIME Machine: On The Power of Motion for Efficient Perception

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    TIME is a motion-based embedding from point tracks, trained only on synthetic data via masked autoencoding, that matches state-of-the-art video model performance with up to 10,000x less training data.

  2. EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models

    cs.CV 2025-06 conditional novelty 7.0 of 10

    EPFL-Smart-Kitchen-30 is a 29.7-hour multimodal cooking dataset with 60k action segments and four benchmarks, including a kinematic-focused VQA benchmark that shows current video-language models struggle with hand and...

  3. EgoWAM: World Action Models Beyond Pixels with In-the-Wild Egocentric Human Data

    cs.RO 2026-07 accept novelty 6.5 of 10

    World Action Model co-training with DINO or 3D-flow targets scales human-to-robot transfer on bimanual tasks far better than behavior cloning, while pixel prediction transfers weakly.

  4. VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances

    cs.RO 2026-08 conditional novelty 6.0 of 10

    From 204K egocentric human videos, the authors automatically extract visual, grasp, and trajectory affordances and train one vision-language model, VLAff, that predicts all three for robot manipulation.

  5. HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Robot-free HiFi-UMI demonstrations can replace teleoperated real-robot data in post-training: three policy backbones matched in-domain teleoperation within 3.1 percentage points, including 85% success on a precision i...

  6. EgoHTR: Egocentric 4D Demonstrations of Human Terrain Traversal

    cs.RO 2026-07 conditional novelty 6.0 of 10

    EgoHTR is a 55-sequence, 150k-frame egocentric 4D human-terrain dataset with a reconstruction pipeline, MoCap-validated benchmark, and perceptive locomotion policies deployed on a Unitree G1.

  7. RGC-VQA: An Exploration Database for Robotic-Generated Video Quality Assessment

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A 2,100-video database with human opinions shows that current video quality models underperform on robot-generated content, motivating a new VQA subfield.

  8. CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CausalVQA provides 793 paired real-video causal reasoning questions on which the best multimodal model scores 61.66% versus 84.78% for humans, with the largest gaps on anticipation and hypothetical questions.

  9. Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new 28,497-sample dataset of 6DoF object manipulation trajectories is automatically extracted from egocentric video, and vision-language models are trained to generate these trajectories from action descriptions.

  10. Hierarchical and Multimodal Data for Daily Activity Understanding

    cs.CV 2025-04 conditional novelty 6.0 of 10

    DARai provides over 200 hours of multi-sensor, three-level annotated daily activity recordings, with benchmarks showing how accuracy varies across sensors, hierarchy levels, and camera or body placement.

  11. Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos

    cs.CV 2024-12 conditional novelty 6.0 of 10

    The authors train a view-switch predictor on pseudo-labeled web videos and show it transfers to selecting ego/exo views in new multi-view videos with limited labels.

  12. EgoCast: Forecasting Egocentric Human Pose in the Wild

    cs.CV 2024-12 conditional novelty 6.0 of 10

    EgoCast forecasts future 3D body poses from egocentric video and headset motion, using a pseudo-groundtruth module so no past body poses are required.

  13. Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

    cs.CV 2024-11 conditional novelty 6.0 of 10

    LangView uses per-view caption accuracy against view-agnostic narrations as pseudo-labels to train a view selector that outperforms heuristics and prior baselines on Ego-Exo4D and LEMMA.

  14. Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

    cs.CV 2024-12 conditional novelty 5.0 of 10

    A real-time egocentric assistant combines a vision-language model, memory, video retrieval, and video generation to answer questions and show how-to guidance from live wearable camera streams.

  15. CuriosAI Submission to the EgoExo4D Proficiency Estimation Challenge 2025

    cs.CV 2025-07 conditional novelty 4.0 of 10

    On EgoExo4D proficiency estimation, a two-stage pipeline with zero-shot scenario recognition and per-scenario, per-view VideoMAE classifiers (47.8% validation) outperforms a Sapiens-2B multi-task model (43.6%).

  16. Comparing Learning Paradigms for Egocentric Video Summarization

    cs.CV 2025-06 reject novelty 4.0 of 10

    A prompt-engineered GPT-4o (quality score 64.95) outperformed Shotluck Holmes (61.19) and TAC-SUM (58.43) on a 21-video egocentric summary evaluation, though all scores were modest.

  17. Data Pyramid for Embodied Manipulation: A Survey

    cs.RO 2026-07 conditional novelty 3.0 of 10

    Embodied training data form a five-layer pyramid—real-robot, UMI, ego/exo, simulation, general V–L—ordered by the trade-off between scale and robot alignment, and model capabilities track how those layers are mixed.

  18. PCIE_Pose Solution for EgoExo4D Pose and Proficiency Estimation Challenge

    cs.CV 2025-05 conditional novelty 3.0 of 10

    The paper reports first-place results on the CVPR2025 EgoExo4D hand pose, body pose, and proficiency estimation challenges using hybrid transformer-CNN and multimodal fusion architectures.

Pith tools