Towards Good Practices for Multi-modal Fusion in Large-scale Video Classification

Changhu Wang; Jinlai Liu; Zehuan Yuan

arxiv: 1809.05848 · v4 · pith:E5C2JAJJnew · submitted 2018-09-16 · 💻 cs.CV

Towards Good Practices for Multi-modal Fusion in Large-scale Video Classification

Jinlai Liu , Zehuan Yuan , Changhu Wang This is my paper

classification 💻 cs.CV

keywords videoclassificationfusionvisuallarge-scaleaudiobilinearframes

0 comments

read the original abstract

Leveraging both visual frames and audio has been experimentally proven effective to improve large-scale video classification. Previous research on video classification mainly focuses on the analysis of visual content among extracted video frames and their temporal feature aggregation. In contrast, multimodal data fusion is achieved by simple operators like average and concatenation. Inspired by the success of bilinear pooling in the visual and language fusion, we introduce multi-modal factorized bilinear pooling (MFB) to fuse visual and audio representations. We combine MFB with different video-level features and explore its effectiveness in video classification. Experimental results on the challenging Youtube-8M v2 dataset demonstrate that MFB significantly outperforms simple fusion methods in large-scale video classification.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Audio-Visual Kinship Verification
cs.CV 2019-06 unverdicted novelty 6.0

Introduces TALKIN dataset and deep Siamese fusion network showing audio-visual combination outperforms uni-modal baselines for kinship verification.