Pith. sign in

REVIEW 22 cited by

wav2vec: Unsupervised Pre-training for Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.05862 v4 pith:M3GDESGW submitted 2019-04-11 cs.CL

wav2vec: Unsupervised Pre-training for Speech Recognition

classification cs.CL
keywords dataspeechaudiocharacter-basedpre-trainingrecognitionrepresentationstraining
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We explore unsupervised pre-training for speech recognition by learning representations of raw audio. wav2vec is trained on large amounts of unlabeled audio data and the resulting representations are then used to improve acoustic model training. We pre-train a simple multi-layer convolutional neural network optimized via a noise contrastive binary classification task. Our experiments on WSJ reduce WER of a strong character-based log-mel filterbank baseline by up to 36% when only a few hours of transcribed data is available. Our approach achieves 2.43% WER on the nov92 test set. This outperforms Deep Speech 2, the best reported character-based system in the literature while using two orders of magnitude less labeled training data.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Flexformer: Flexible Linear Transformer with Learnable Attention Kernel

    cs.LG 2026-06 unverdicted novelty 7.0

    Flexformer learns data-driven attention kernels for linear Transformers by optimizing spectral frequencies in random Fourier features, with stationary and nonstationary variants that outperform fixed-kernel baselines ...

  2. SpurAudio: A Benchmark for Studying Shortcut Learning in Few-Shot Audio Classification

    cs.CV 2026-05 unverdicted novelty 7.0

    SpurAudio benchmark shows state-of-the-art few-shot audio classifiers suffer large performance drops when background correlations are disrupted, even in large pretrained models.

  3. The Indra Representation Hypothesis for Multimodal Alignment

    cs.CV 2026-04 unverdicted novelty 7.0

    Unimodal model representations converge to a relational structure captured by the Indra representation via V-enriched Yoneda embedding, which is unique and structure-preserving and improves cross-model and cross-modal...

  4. A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection

    eess.AS 2026-03 unverdicted novelty 7.0

    Spoof-SUPERB benchmark shows large-scale discriminative SSL models such as XLS-R, UniSpeech-SAT, and WavLM Large outperform others in audio deepfake detection and maintain robustness under acoustic degradations.

  5. Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation

    cs.CV 2026-08 conditional novelty 6.0

    A diffusion-based talking-face generator uses 3D blendshape coefficients to continuously control the emotion intensity of generated facial expressions.

  6. Should Missing Modalities Always Be Necessary to Repair for Multi-modal Sentiment Analysis?

    cs.CL 2026-07 conditional novelty 6.0

    SIEVE learns a per-sample 'should we repair?' decision from the loss gap between direct and repair branches, improving three missing-modality MSA backbones on CMU-MOSI and IEMOCAP.

  7. Flexformer: Flexible Linear Transformer with Learnable Attention Kernel

    cs.LG 2026-06 unverdicted novelty 6.0

    Flexformer learns attention kernels by treating spectral frequencies as trainable parameters in random Fourier feature-based linear attention, with stationary and nonstationary variants that outperform fixed-kernel baselines.

  8. Extracting Governing Equations from Latent Dynamics via Multi-View Contrastive Learning

    cs.LG 2026-06 unverdicted novelty 6.0

    DYSCO jointly recovers latent trajectories and governing equations from noisy observations via multi-view contrastive learning, with theoretical guarantees up to affine indeterminacy.

  9. Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing

    cs.SD 2026-06 unverdicted novelty 6.0

    A gated fusion of XLSR-53 and CORES features with energy margin and diversity losses reaches 97.6% ID accuracy and reduces FPR95 by 83.5% relative to the Interspeech 2025 baseline on MLAAD.

  10. Evaluating Speech Articulation Synthesis with Articulatory Phoneme Recognition

    cs.CL 2026-05 unverdicted novelty 6.0

    The paper introduces phoneme recognition using articulatory features as a proxy metric for evaluating articulatory speech synthesis quality from phonetic sequences.

  11. STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

    cs.CV 2025-12 conditional novelty 6.0

    STARCaster is a 2D spatio-temporal video diffusion model that unifies identity-conditioned audio-driven portrait animation and novel-view synthesis without explicit 3D reconstruction.

  12. Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation

    cs.CV 2024-11 unverdicted novelty 6.0

    LetsTalk combines a multimodal diffusion transformer, noise-regularized memory bank, deep compression autoencoder, and symbiotic/direct fusion schemes to achieve state-of-the-art quality and efficiency in long-duratio...

  13. Vision Transformers Need Registers

    cs.CV 2023-09 unverdicted novelty 6.0

    Adding register tokens to Vision Transformers eliminates high-norm background artifacts and raises state-of-the-art performance on dense visual prediction tasks.

  14. wav2VOT: Automatic estimation of voice onset time, closure duration, and burst realisation with wav2vec2

    cs.SD 2026-06 unverdicted novelty 5.0

    wav2VOT shows wav2vec2 can estimate voice onset time and related stop consonant features with accuracy comparable to existing tools on unseen data and higher accuracy after fine-tuning.

  15. Fully Differentiable Neural Forced Alignment via Soft Dynamic Programming

    eess.AS 2026-06 unverdicted novelty 5.0

    Presents an end-to-end differentiable neural model for phoneme forced alignment that claims to outperform prior methods on English benchmarks and generalize to unseen languages.

  16. HighSync: High-Quality Lip Synchronization via Latent Diffusion Models

    cs.CV 2026-05 unverdicted novelty 5.0

    HighSync is a diffusion-based lip synchronization system that operates natively at 512x512 resolution by eliminating data leakage to enforce genuine audio dependence and reports state-of-the-art results on quality and...

  17. Clustering Unsupervised Representations as Defense against Poisoning Attacks on Speech Commands Classification System

    cs.SD 2026-06 unverdicted novelty 4.0

    Clustering DINO representations via K-means and LDA filters poisoned speech samples, reducing attack success rate from 99.75% to 0.25% at 10% poisoning level.

  18. Learning to Attend to Depression-Related Patterns: An Adaptive Cross-Modal Gating Network for Depression Detection

    cs.SD 2026-04 unverdicted novelty 4.0

    An adaptive cross-modal gating network improves depression detection from speech by selectively weighting sparse relevant segments across acoustic and textual modalities.

  19. Contextualized Token Discrimination for Speech Search Query Correction

    cs.SD 2025-09 reject novelty 4.0

    CTD uses BERT token representations plus a composition layer to correct Chinese spelling errors in ASR queries, but the reported gains lack matched baselines and released data.

  20. From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning

    eess.AS 2026-07 unverdicted novelty 3.0

    A survey that organizes audio SSL into five objective paradigms, relates their demands to architectural biases, and interprets downstream applications as tests of generalization.

  21. EGI: A Multimodal Emotional AI Framework for Enhancing Scrum Master Real-time Self-Awareness

    cs.AI 2026-05 unverdicted novelty 3.0

    EGI integrates four existing AI components for real-time multimodal emotion monitoring and feedback in simulated agile meetings, reporting 10% WER and improved self-awareness for Scrum Masters.

  22. Meta-Learning and Meta-Reinforcement Learning -- Tracing the Path towards DeepMind's Adaptive Agent

    cs.AI 2026-02 unverdicted novelty 2.0

    A survey provides a task-based formalization of meta-learning and meta-RL while chronicling algorithms that lead to DeepMind's Adaptive Agent.