REVIEW 4 cited by
Technical Report for Ego4D Long Term Action Anticipation Challenge 2023
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this report, we describe the technical details of our approach for the Ego4D Long-Term Action Anticipation Challenge 2023. The aim of this task is to predict a sequence of future actions that will take place at an arbitrary time or later, given an input video. To accomplish this task, we introduce three improvements to the baseline model, which consists of an encoder that generates clip-level features from the video, an aggregator that integrates multiple clip-level features, and a decoder that outputs Z future actions. 1) Model ensemble of SlowFast and SlowFast-CLIP; 2) Label smoothing to relax order constraints for future actions; 3) Constraining the prediction of the action class (verb, noun) based on word co-occurrence. Our method outperformed the baseline performance and recorded as second place solution on the public leaderboard.
Forward citations
Cited by 4 Pith papers
-
Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings
SL-ASD replaces audio-visual synchronization with face-voice identity matching plus quality-weighted face aggregation, reporting competitive Ego4D validation accuracy with 0.4M learnable parameters, though the compari...
-
Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
A new 28,497-sample dataset of 6DoF object manipulation trajectories is automatically extracted from egocentric video, and vision-language models are trained to generate these trajectories from action descriptions.
-
Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models
Adding a backward prediction task to LLM training improves long-term action anticipation on Ego4D, lowering edit distance for predicted action sequences.
-
Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings
SCAN adds framewise voice comparison between reference speech and candidate audio to active speaker detection, improving mAP on Ego4D over the TalkNet and Light-ASD baselines.
Discussion (0). Continue with ORCID to comment.