Pith. sign in

REVIEW 9 cited by

EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.18070 v4 pith:TVLZPACY submitted 2024-06-26 cs.CV

classification cs.CV
keywords egocentricegovideomodelactionchallengefoundationtracksadaptation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this report, we present our solutions to the EgoVis Challenges in CVPR 2024, including five tracks in the Ego4D challenge and three tracks in the EPIC-Kitchens challenge. Building upon the video-language two-tower model and leveraging our meticulously organized egocentric video data, we introduce a novel foundation model called EgoVideo. This model is specifically designed to cater to the unique characteristics of egocentric videos and provides strong support for our competition submissions. In the Ego4D challenges, we tackle various tasks including Natural Language Queries, Step Grounding, Moment Queries, Short-term Object Interaction Anticipation, and Long-term Action Anticipation. In addition, we also participate in the EPIC-Kitchens challenge, where we engage in the Action Recognition, Multiple Instance Retrieval, and Domain Adaptation for Action Recognition tracks. By adapting EgoVideo to these diverse tasks, we showcase its versatility and effectiveness in different egocentric video analysis scenarios, demonstrating the powerful representation ability of EgoVideo as an egocentric foundation model. Our codebase and pretrained models are publicly available at https://github.com/OpenGVLab/EgoVideo.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction

    cs.CV 2025-07 conditional novelty 6.0 of 10

    An MLLM trained with auxiliary goal-prediction tasks and multi-token prediction achieves SOTA on COIN and CrossTask visual planning and matches SOTA on Ego4D LTA.

  2. THOR: Thermal-guided Hand-Object Reasoning via Adaptive Vision Sampling

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A thermal-guided adaptive sampling system cuts RGB video data by roughly 97% while keeping hand-activity recognition F1 around 95%, comparable to processing all frames.

  3. Seeing Clearly, Forgetting Deeply: Revisiting Fine-Tuned Video Generators for Driving Simulation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Fine-tuning video generators on driving data can improve visual fidelity while degrading how accurately the model predicts the movement of cars and pedestrians.

  4. The Repeated-Stimulus Confound in Electroencephalography

    q-bio.NC 2025-08 unverdicted novelty 5.0 of 10

    The repeated-stimulus confound, where models are trained and tested on repeated presentations of identical stimuli, inflates reported EEG decoding accuracies by an estimated 4.46 to 7.42 percent.

  5. EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization

    cs.CV 2025-06 conditional novelty 5.0 of 10

    EVA02-AT combines full-dimension spatial and temporal rotary position embeddings with a symmetric multi-similarity loss to improve egocentric video-text retrieval.

  6. GazeNLQ @ Ego4D Natural Language Queries Challenge 2025

    cs.CV 2025-06 conditional novelty 5.0 of 10

    GazeNLQ adds contrastively pretrained gaze embeddings to a GroundNLQ-style grounding model, reporting 27.82 R1@0.3 on the Ego4D NLQ test split only when ensembled with GroundVQA.

  7. OSGNet @ Ego4D Episodic Memory Challenge 2025

    cs.CV 2025-06 conditional novelty 4.0 of 10

    OSGNet, an early-fusion grounding model, wins all three Ego4D Episodic Memory Challenge tracks by converting localization tasks into retrieval problems.

  8. Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A three-stage pipeline using the EgoVideo-V encoder, a verb-noun co-occurrence reranker, SAM2 hand-object features, and a fine-tuned Llama 2 model took first place in the Ego4D 2025 long-term action anticipation challenge.

  9. Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision

    cs.CV 2025-06 accept novelty 3.0 of 10

    A comprehensive review of cross-view video understanding that uses both first-person and third-person cameras, organized into a three-direction taxonomy with a dataset catalog and future research gaps.

Pith tools