Pith. sign in

REVIEW 1 cited by

Cross-Domain First Person Audio-Visual Action Recognition through Relative Norm Alignment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.01689 v1 pith:J5IXIVKL submitted 2021-06-03 cs.CV

classification cs.CV
keywords actionrecognitionaudio-visualcross-domainfirstpersondataduring
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

First person action recognition is an increasingly researched topic because of the growing popularity of wearable cameras. This is bringing to light cross-domain issues that are yet to be addressed in this context. Indeed, the information extracted from learned representations suffers from an intrinsic environmental bias. This strongly affects the ability to generalize to unseen scenarios, limiting the application of current methods in real settings where trimmed labeled data are not available during training. In this work, we propose to leverage over the intrinsic complementary nature of audio-visual signals to learn a representation that works well on data seen during training, while being able to generalize across different domains. To this end, we introduce an audio-visual loss that aligns the contributions from the two modalities by acting on the magnitude of their feature norm representations. This new loss, plugged into a minimal multi-modal action recognition architecture, leads to strong results in cross-domain first person action recognition, as demonstrated by extensive experiments on the popular EPIC-Kitchens dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring Audio Cues for Enhanced Test-Time Video Model Adaptation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Using audio-assisted pseudo-labels generated by a pretrained audio model and an LLM improves test-time adaptation of video classifiers on corrupted videos.

Pith tools