Pith. sign in

REVIEW 3 cited by

InternVideo-Ego4D: A Pack of Champion Solutions to Ego4D Challenges

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2211.09529 v1 pith:YR3YA4UK submitted 2022-11-17 cs.CV

classification cs.CV
keywords ego4dfivefoundationinternvideo-ego4dmodeltasksvideochampion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this report, we present our champion solutions to five tracks at Ego4D challenge. We leverage our developed InternVideo, a video foundation model, for five Ego4D tasks, including Moment Queries, Natural Language Queries, Future Hand Prediction, State Change Object Detection, and Short-term Object Interaction Anticipation. InternVideo-Ego4D is an effective paradigm to adapt the strong foundation model to the downstream ego-centric video understanding tasks with simple head designs. In these five tasks, the performance of InternVideo-Ego4D comprehensively surpasses the baseline methods and the champions of CVPR2022, demonstrating the powerful representation ability of InternVideo as a video foundation model. Our code will be released at https://github.com/OpenGVLab/ego4d-eccv2022-solutions

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GazeNLQ @ Ego4D Natural Language Queries Challenge 2025

    cs.CV 2025-06 conditional novelty 5.0 of 10

    GazeNLQ adds contrastively pretrained gaze embeddings to a GroundNLQ-style grounding model, reporting 27.82 R1@0.3 on the Ego4D NLQ test split only when ensembled with GroundVQA.

  2. OSGNet @ Ego4D Episodic Memory Challenge 2025

    cs.CV 2025-06 conditional novelty 4.0 of 10

    OSGNet, an early-fusion grounding model, wins all three Ego4D Episodic Memory Challenge tracks by converting localization tasks into retrieval problems.

  3. Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision

    cs.CV 2025-06 accept novelty 3.0 of 10

    A comprehensive review of cross-view video understanding that uses both first-person and third-person cameras, organized into a three-direction taxonomy with a dataset catalog and future research gaps.

Pith tools