REVIEW 12 cited by
vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We propose vq-wav2vec to learn discrete representations of audio segments through a wav2vec-style self-supervised context prediction task. The algorithm uses either a gumbel softmax or online k-means clustering to quantize the dense representations. Discretization enables the direct application of algorithms from the NLP community which require discrete inputs. Experiments show that BERT pre-training achieves a new state of the art on TIMIT phoneme classification and WSJ speech recognition.
Forward citations
Cited by 12 Pith papers
-
Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR
On a five-seed Garhwali ASR benchmark using the official VAANI splits, standard CTC beats Focal CTC, a matra-weighted objective, and Hindi transfer, while speed augmentation and encoder choice provide the only robust gains.
-
Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video
A self-supervised diffusion framework with a low-bitrate vector-quantization bottleneck learns disentangled motion and content latents supporting motion transfer and auto-regressive generation.
-
Multimodal Medical Code Tokenizer
MedTok encodes medical codes with text and graph information into a shared vector-quantized token space, improving downstream EHR prediction and medical QA when swapped in for standard tokenizers.
-
Efficient Video Face Enhancement with Enhanced Spatial-Temporal Consistency
A 3D-VQGAN with spatial-temporal codebooks and code-lookup transformers restores compressed face videos and removes flicker in about 3 seconds per 24-frame clip.
-
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram
ESTVocoder synthesizes speech by transforming the amplitude and phase spectra of an F0-derived harmonic excitation into speech spectra with a ConvNeXt v2 neural filter, improving several objective metrics over HiFi-GA...
-
Bilevel Joint Unsupervised and Supervised Training for Automatic Speech Recognition
A bilevel training method that jointly optimizes supervised and unsupervised losses outperforms pretraining-then-finetuning for ASR on LibriSpeech, Switchboard, and an in-house dataset.
-
Scalable Image Tokenization with Index Backpropagation Quantization
Index Backpropagation Quantization updates all codebook embeddings via a straight-through softmax gradient, enabling a high-utilization 2^18-codebook image tokenizer with state-of-the-art reconstruction (rFID 1.00).
-
VPBSD:Vessel-Pattern-Based Semi-Supervised Distillation for Efficient 3D Microscopic Cerebrovascular Segmentation
VpbSD uses a vessel-pattern codebook trained on unlabeled microscopy data to distill knowledge from a large teacher into a 0.12M-parameter student, reaching DSC 0.852 on VesSep2020.
-
Heterogeneous Self-Supervised Acoustic Pre-Training with Local Constraints
PTLOC adds per-source local constraints to self-supervised acoustic pre-training via first-order MAML-style updates, reporting improved downstream ASR word error rates over a baseline that is not compute-matched.
-
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey
A comprehensive survey that frames multimodal understanding and generation as next token prediction and proposes a five-part taxonomy.
-
Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis
Whisper-Tiny used as the audio feature extractor for RAD-NeRF and ER-NeRF talking portraits reduces AFE latency and yields modestly better SyncNet lip-sync scores than DeepSpeech, Wav2Vec 2.0, or HuBERT on three short...
-
The SVASR System for Text-dependent Speaker Verification (TdSV) AAIC Challenge 2024
An ASR content gate plus concatenated wav2vec-BERT and ReDimNet speaker embeddings achieved normalized min-DCF 0.0452 and rank 2 on the TDSV 2024 text-dependent speaker verification challenge.
Discussion (0). Continue with ORCID to comment.