Pith. sign in

REVIEW 12 cited by

vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1910.05453 v3 pith:KZT3ZSOY submitted 2019-10-12 cs.CL cs.LG

classification cs.CLcs.LG
keywords discreterepresentationsself-supervisedspeechvq-wav2vecachievesalgorithmalgorithms
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We propose vq-wav2vec to learn discrete representations of audio segments through a wav2vec-style self-supervised context prediction task. The algorithm uses either a gumbel softmax or online k-means clustering to quantize the dense representations. Discretization enables the direct application of algorithms from the NLP community which require discrete inputs. Experiments show that BERT pre-training achieves a new state of the art on TIMIT phoneme classification and WSJ speech recognition.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR

    cs.CL 2026-08 conditional novelty 6.0 of 10

    On a five-seed Garhwali ASR benchmark using the official VAANI splits, standard CTC beats Focal CTC, a matra-weighted objective, and Hindi transfer, while speed augmentation and encoder choice provide the only robust gains.

  2. Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A self-supervised diffusion framework with a low-bitrate vector-quantization bottleneck learns disentangled motion and content latents supporting motion transfer and auto-regressive generation.

  3. Multimodal Medical Code Tokenizer

    cs.CL 2025-02 conditional novelty 6.0 of 10

    MedTok encodes medical codes with text and graph information into a shared vector-quantized token space, improving downstream EHR prediction and medical QA when swapped in for standard tokenizers.

  4. Efficient Video Face Enhancement with Enhanced Spatial-Temporal Consistency

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A 3D-VQGAN with spatial-temporal codebooks and code-lookup transformers restores compressed face videos and removes flicker in about 3 seconds per 24-frame clip.

  5. ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram

    cs.SD 2024-11 conditional novelty 6.0 of 10

    ESTVocoder synthesizes speech by transforming the amplitude and phase spectra of an F0-derived harmonic excitation into speech spectra with a ConvNeXt v2 neural filter, improving several objective metrics over HiFi-GA...

  6. Bilevel Joint Unsupervised and Supervised Training for Automatic Speech Recognition

    cs.CL 2024-12 conditional novelty 5.0 of 10

    A bilevel training method that jointly optimizes supervised and unsupervised losses outperforms pretraining-then-finetuning for ASR on LibriSpeech, Switchboard, and an in-house dataset.

  7. Scalable Image Tokenization with Index Backpropagation Quantization

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Index Backpropagation Quantization updates all codebook embeddings via a straight-through softmax gradient, enabling a high-utilization 2^18-codebook image tokenizer with state-of-the-art reconstruction (rFID 1.00).

  8. VPBSD:Vessel-Pattern-Based Semi-Supervised Distillation for Efficient 3D Microscopic Cerebrovascular Segmentation

    cs.CV 2024-11 conditional novelty 5.0 of 10

    VpbSD uses a vessel-pattern codebook trained on unlabeled microscopy data to distill knowledge from a large teacher into a 0.12M-parameter student, reaching DSC 0.852 on VesSep2020.

  9. Heterogeneous Self-Supervised Acoustic Pre-Training with Local Constraints

    cs.LG 2025-08 conditional novelty 4.0 of 10

    PTLOC adds per-source local constraints to self-supervised acoustic pre-training via first-order MAML-style updates, reporting improved downstream ASR word error rates over a baseline that is not compute-matched.

  10. Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A comprehensive survey that frames multimodal understanding and generation as next token prediction and proposes a five-part taxonomy.

  11. Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis

    cs.SD 2024-11 conditional novelty 4.0 of 10

    Whisper-Tiny used as the audio feature extractor for RAD-NeRF and ER-NeRF talking portraits reduces AFE latency and yields modestly better SyncNet lip-sync scores than DeepSpeech, Wav2Vec 2.0, or HuBERT on three short...

  12. The SVASR System for Text-dependent Speaker Verification (TdSV) AAIC Challenge 2024

    cs.SD 2024-11 conditional novelty 3.0 of 10

    An ASR content gate plus concatenated wav2vec-BERT and ReDimNet speaker embeddings achieved normalized min-DCF 0.0452 and rank 2 on the TDSV 2024 text-dependent speaker verification challenge.

Pith tools