Pith. sign in

REVIEW 6 cited by

data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.03555 v3 pith:KGO5X5TI submitted 2022-02-07 cs.LG

classification cs.LG
keywords learningspeechdata2vecgeneralinputself-supervisedframeworkidea
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While the general idea of self-supervised learning is identical across modalities, the actual algorithms and objectives differ widely because they were developed with a single modality in mind. To get us closer to general self-supervised learning, we present data2vec, a framework that uses the same learning method for either speech, NLP or computer vision. The core idea is to predict latent representations of the full input data based on a masked view of the input in a self-distillation setup using a standard Transformer architecture. Instead of predicting modality-specific targets such as words, visual tokens or units of human speech which are local in nature, data2vec predicts contextualized latent representations that contain information from the entire input. Experiments on the major benchmarks of speech recognition, image classification, and natural language understanding demonstrate a new state of the art or competitive performance to predominant approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 241 citations worldwide. Full citation record

  1. The Importance of Encoder Choice:A Tabular-Image Study

    cs.LG 2026-07 conditional novelty 6.5 of 10

    Tabular encoder choice reorders multimodal rankings, can erase apparent fusion gains, and requires non-vanilla extraction for in-context learning models to avoid train-test representation shift.

  2. Kepler-Encoder-v0.1: Towards a Multimodal Embedding Model for Robots

    cs.RO 2026-07 conditional novelty 6.0 of 10

    A self-supervised multimodal encoder trained with vision, proprioception, and force yields a vision-only latent that recovers end-effector state and force above vision baselines on RH20T, with modest absolute force accuracy.

  3. ZeroSyl: Simple Zero-Resource Syllable Tokenization for Spoken Language Modeling

    cs.CL 2026-02 conditional novelty 6.0 of 10

    Syllable boundaries can be derived directly from L2 norms of frozen WavLM features, yielding competitive spoken-language-model tokens without any training.

  4. Toward Annotation-Efficient Continuous Emotion Arousal Quantification via Group-Level EEG Dynamic Neural Synchrony

    cs.HC 2026-07 conditional novelty 5.0 of 10

    Group-level EEG dynamic neural synchrony (CorrCA) preferentially tracks the rate of change of continuous arousal and shows valence-dependent structure across four datasets.

  5. STST-JEPA: Shallow-Target Spatio-Temporal Joint Embedding Prediction Architecture For EEG Self-Supervised Learning

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A JEPA-style EEG foundation model with shallow EMA targets plus light reconstruction reaches strong multi-task transfer and 3.06-year validation age MAE on a large multi-site corpus.

  6. Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning

    cs.SD 2025-06 conditional novelty 3.0 of 10

    Transferring I-JEPA's masked latent prediction to mel-spectrograms yields competitive audio representations on music and environmental sound tasks with a small fraction of the training data.

Pith tools