A three-stage pipeline using the EgoVideo-V encoder, a verb-noun co-occurrence reranker, SAM2 hand-object features, and a fine-tuned Llama 2 model took first place in the Ego4D 2025 long-term action anticipation challenge.
QueryMamba: A Mamba-Based Encoder-Decoder Architecture with a Statistical Verb-Noun Interaction Module for Video Action Forecasting @ Ego4D Long-Term Action Anticipation Challenge 2024
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
This report presents a novel Mamba-based encoder-decoder architecture, QueryMamba, featuring an integrated verb-noun interaction module that utilizes a statistical verb-noun co-occurrence matrix to enhance video action forecasting. This architecture not only predicts verbs and nouns likely to occur based on historical data but also considers their joint occurrence to improve forecast accuracy. The efficacy of this approach is substantiated by experimental results, with the method achieving second place in the Ego4D LTA challenge and ranking first in noun prediction accuracy.
citation-role summary
citation-polarity summary
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025
A three-stage pipeline using the EgoVideo-V encoder, a verb-noun co-occurrence reranker, SAM2 hand-object features, and a fine-tuned Llama 2 model took first place in the Ego4D 2025 long-term action anticipation challenge.