REVIEW 5 cited by
Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025
read the original abstract
In this report, we present a novel three-stage framework developed for the Ego4D Long-Term Action Anticipation (LTA) task. Inspired by recent advances in foundation models, our method consists of three stages: feature extraction, action recognition, and long-term action anticipation. First, visual features are extracted using a high-performance visual encoder. The features are then fed into a Transformer to predict verbs and nouns, with a verb-noun co-occurrence matrix incorporated to enhance recognition accuracy. Finally, the predicted verb-noun pairs are formatted as textual prompts and input into a fine-tuned large language model (LLM) to anticipate future action sequences. Our framework achieves first place in this challenge at CVPR 2025, establishing a new state-of-the-art in long-term action prediction. Our code will be released at https://github.com/CorrineQiu/Ego4D-LTA-Challenge-2025.
Forward citations
Cited by 5 Pith papers
-
FROST-STA: Frozen Dense Features for the Ego4D Short-Term Object Interaction Anticipation
FROST-STA ranks second in the Ego4D Short-Term Object Interaction Anticipation challenge with 5.13 mAP by adapting frozen V-JEPA features with object-centric heads and ensembling.
-
TAP-JEPA: Frozen Future-Latent Probing and Two-Stage Score Fusion for EPIC-KITCHENS-100 Action Anticipation
TAP-JEPA applies frozen V-JEPA features, latent future prediction, and two-stage fusion of attentive probes to reach 27.91% MT5R and second place on the EK-100 action anticipation leaderboard.
-
JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026
JFAA freezes a JEPA future-prediction model, adds a lightweight probe and ensemble, and wins the 2026 EK-100 action anticipation challenge.
-
MARS: Technical Report for the CASTLE Challenge at EgoVis 2026
MARS converts long videos to captions and summaries, maintains modality-specific memories, and deploys an agent to select evidence or answer, placing second on the CASTLE Challenge leaderboard.
-
VISTA: Technical Report for the Ego4D Short-Term Object Interaction Anticipation at EgoVis 2026
VISTA wins first place on the Ego4D Short-Term Object Interaction Anticipation challenge by combining spatial object proposals with temporal context via feature modulation and ROI fusion, followed by ensembling.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.