Cross and Learn: Cross-Modal Self-Supervision

Biagio Brattoli; Bj\"orn Ommer; Nawid Sayed

arxiv: 1811.03879 · v3 · pith:WVIXTORWnew · submitted 2018-11-09 · 💻 cs.CV

Cross and Learn: Cross-Modal Self-Supervision

Nawid Sayed , Biagio Brattoli , Bj\"orn Ommer This is my paper

classification 💻 cs.CV

keywords cross-modallearningmethodmodalitiesrepresentationself-supervisedablationaccessible

0 comments

read the original abstract

In this paper we present a self-supervised method for representation learning utilizing two different modalities. Based on the observation that cross-modal information has a high semantic meaning we propose a method to effectively exploit this signal. For our approach we utilize video data since it is available on a large scale and provides easily accessible modalities given by RGB and optical flow. We demonstrate state-of-the-art performance on highly contested action recognition datasets in the context of self-supervised learning. We show that our feature representation also transfers to other tasks and conduct extensive ablation studies to validate our core contributions. Code and model can be found at https://github.com/nawidsayed/Cross-and-Learn.

This paper has not been read by Pith yet.

Cross and Learn: Cross-Modal Self-Supervision

discussion (0)