Pith. sign in

REVIEW 1 cited by

Efficient Multiscale Multimodal Bottleneck Transformer for Audio-Video Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.04023 v1 pith:GYYTGEBM submitted 2024-01-08 cs.CV cs.AIcs.LGcs.MMcs.SDeess.AS

classification cs.CVcs.AIcs.LGcs.MMcs.SDeess.AS
keywords multiscaletransformercontrastiveefficientmultimodalaudioaudio-videoloss
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, researchers combine both audio and video signals to deal with challenges where actions are not well represented or captured by visual cues. However, how to effectively leverage the two modalities is still under development. In this work, we develop a multiscale multimodal Transformer (MMT) that leverages hierarchical representation learning. Particularly, MMT is composed of a novel multiscale audio Transformer (MAT) and a multiscale video Transformer [43]. To learn a discriminative cross-modality fusion, we further design multimodal supervised contrastive objectives called audio-video contrastive loss (AVC) and intra-modal contrastive loss (IMC) that robustly align the two modalities. MMT surpasses previous state-of-the-art approaches by 7.3% and 2.1% on Kinetics-Sounds and VGGSound in terms of the top-1 accuracy without external training data. Moreover, the proposed MAT significantly outperforms AST [28] by 22.2%, 4.4% and 4.7% on three public benchmark datasets, and is about 3% more efficient based on the number of FLOPs and 9.8% more efficient based on GPU memory usage.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Video-Based MPAA Rating Prediction: An Attention-Driven Hybrid Architecture Using Contrastive Learning

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A CNN+LSTM+attention model with contrastive learning predicts MPAA ratings from short video clips with 88% accuracy on a custom 323-clip dataset.

Pith tools