Pith. sign in

REVIEW 14 cited by

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.08380 v1 pith:MYBQHHR7 submitted 2024-11-13 cs.CV

EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation

classification cs.CV
keywords egocentricgenerationvideoactiondatasetegovid-5mdataannotations
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Video generation has emerged as a promising tool for world simulation, leveraging visual data to replicate real-world environments. Within this context, egocentric video generation, which centers on the human perspective, holds significant potential for enhancing applications in virtual reality, augmented reality, and gaming. However, the generation of egocentric videos presents substantial challenges due to the dynamic nature of egocentric viewpoints, the intricate diversity of actions, and the complex variety of scenes encountered. Existing datasets are inadequate for addressing these challenges effectively. To bridge this gap, we present EgoVid-5M, the first high-quality dataset specifically curated for egocentric video generation. EgoVid-5M encompasses 5 million egocentric video clips and is enriched with detailed action annotations, including fine-grained kinematic control and high-level textual descriptions. To ensure the integrity and usability of the dataset, we implement a sophisticated data cleaning pipeline designed to maintain frame consistency, action coherence, and motion smoothness under egocentric conditions. Furthermore, we introduce EgoDreamer, which is capable of generating egocentric videos driven simultaneously by action descriptions and kinematic control signals. The EgoVid-5M dataset, associated action annotations, and all data cleansing metadata will be released for the advancement of research in egocentric video generation.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks

    cs.CV 2026-04 unverdicted novelty 7.0

    EgoTL provides a new egocentric dataset with think-aloud chains and metric labels that benchmarks VLMs on long-horizon tasks and improves their planning, reasoning, and spatial grounding after finetuning.

  2. SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting

    cs.CV 2025-11 unverdicted novelty 7.0

    SFHand presents the first streaming language-guided autoregressive framework for 3D hand forecasting, achieving up to 35.8% gains over prior methods and 13.4% better downstream embodied task performance.

  3. EgoSim: Egocentric World Simulator for Embodied Interaction Generation

    cs.CV 2026-04 conditional novelty 6.5

    EgoSim generates spatially consistent egocentric interaction videos by conditioning a video diffusion model on updatable 3D point-cloud states and action keypoints extracted at scale from monocular videos.

  4. HandsOnWorld: Unconstrained Egocentric Video Generation with Camera-Disentangled Hand Control

    cs.CV 2026-07 unverdicted novelty 6.0

    HandsOnWorld creates a hand-controlled egocentric video generator from unconstrained monocular video via a new EgoVid-Pro dataset from monocular reconstruction and a Plücker Hand Map that disentangles camera and hand motion.

  5. E$^3$C: Video Generation with 3D Environmental Memory and Ego-Exo Human Pose Control

    cs.CV 2026-05 unverdicted novelty 6.0

    E³C is a video diffusion model that disentangles persistent 3D scene structure via point-cloud memory from human dynamics via ego-exo pose controls for improved egocentric video generation on the Nymeria dataset.

  6. Precise Action-to-Video Generation Through Visual Action Prompts

    cs.CV 2025-08 conditional novelty 6.0

    Skeleton-based visual action prompts give precise, cross-domain action control for video generation of human and robot interactions.

  7. EgoSafe: A First-Person Mobile-Captured Benchmark for Visual Safety Understanding

    cs.CV 2026-07 conditional novelty 5.0

    A new egocentric safety benchmark shows current video-language models can describe scenes well but fail at multi-step causal reasoning about blind spots and covert actions.

  8. iFLYTEK-Embodied-Omni Technical Report

    cs.AI 2026-06 conditional novelty 5.0

    A three-branch Omni model (VLM+VGM brain, AGM cerebellum) with shared multimodal attention and four-stage training reaches 89.6% zero-shot on LIBERO-Plus and ~93% on RoboTwin 2.0.

  9. VISTA: A Controllable Platform for Generating and Auditing Egocentric Assistance Scenarios

    cs.CL 2026-05 reject novelty 5.0

    A staged, inspectable pipeline converts scenario descriptions into egocentric videos for assistant-AI training, but its claimed edge over one-pass baselines is an estimate, not a measured result.

  10. VISTA: A Controllable Platform for Generating and Auditing Egocentric Assistance Scenarios

    cs.CL 2026-05 unverdicted novelty 5.0

    VISTA is a video synthesis framework that creates controllable egocentric videos of daily tasks with reactive and proactive agent intervention modes via causal reverse reasoning scripts.

  11. Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

    cs.RO 2025-08 unverdicted novelty 5.0

    This survey organizes large VLM-based VLA models for robotic manipulation into monolithic and hierarchical paradigms, reviews their integrations and datasets, and outlines future directions.

  12. World Action Models: The Next Frontier in Embodied AI

    cs.RO 2026-05 unverdicted novelty 4.0

    The paper introduces World Action Models as a new paradigm unifying predictive world modeling with action generation in embodied foundation models and provides a taxonomy of existing approaches.

  13. World Action Models: A Survey

    cs.RO 2026-06 unverdicted novelty 3.0

    A survey that clarifies boundaries and organizes World Action Models by generation requirements and predictive substrates, identifying a trend toward generating less of the future.

  14. Robot Learning from Human Videos: A Survey

    cs.RO 2026-04 unverdicted novelty 2.0

    The survey organizes human-video-based robot learning into task-, observation-, and action-oriented transfer pathways, reviews associated datasets, and outlines challenges for scalable embodied AI.