Pith. sign in

REVIEW 4 cited by

Technical Report for Ego4D Long Term Action Anticipation Challenge 2023

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.01467 v1 pith:VGZOM5CA submitted 2023-07-04 cs.CV

classification cs.CV
keywords actionactionsfutureanticipationbaselinechallengeclip-levelego4d
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this report, we describe the technical details of our approach for the Ego4D Long-Term Action Anticipation Challenge 2023. The aim of this task is to predict a sequence of future actions that will take place at an arbitrary time or later, given an input video. To accomplish this task, we introduce three improvements to the baseline model, which consists of an encoder that generates clip-level features from the video, an aggregator that integrates multiple clip-level features, and a decoder that outputs Z future actions. 1) Model ensemble of SlowFast and SlowFast-CLIP; 2) Label smoothing to relax order constraints for future actions; 3) Constraining the prediction of the action class (verb, noun) based on word co-occurrence. Our method outperformed the baseline performance and recorded as second place solution on the public leaderboard.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Face-Voice Association for Audiovisual Active Speaker Detection in Egocentric Recordings

    cs.MM 2025-06 conditional novelty 6.0 of 10

    SL-ASD replaces audio-visual synchronization with face-voice identity matching plus quality-weighted face aggregation, reporting competitive Ego4D validation accuracy with 0.4M learnable parameters, though the compari...

  2. Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new 28,497-sample dataset of 6DoF object manipulation trajectories is automatically extracted from egocentric video, and vision-language models are trained to generate these trajectories from action descriptions.

  3. Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Adding a backward prediction task to LLM training improves long-term action anticipation on Ego4D, lowering edit distance for predicted action sequences.

  4. Speaker Embedding Informed Audiovisual Active Speaker Detection for Egocentric Recordings

    cs.MM 2025-02 conditional novelty 5.0 of 10

    SCAN adds framewise voice comparison between reference speech and candidate audio to active speaker detection, improving mAP on Ego4D over the TalkNet and Light-ASD baselines.

Pith tools