Pith. sign in

REVIEW 2 cited by

SGEM: Test-Time Adaptation for Automatic Speech Recognition via Sequential-Level Generalized Entropy Minimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.01981 v4 pith:K76TT3HG submitted 2023-06-03 eess.AS cs.AIcs.LG

classification eess.AScs.AIcs.LG
keywords sgemadaptationmodelmodelsoutputadaptautomaticdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Automatic speech recognition (ASR) models are frequently exposed to data distribution shifts in many real-world scenarios, leading to erroneous predictions. To tackle this issue, an existing test-time adaptation (TTA) method has recently been proposed to adapt the pre-trained ASR model on unlabeled test instances without source data. Despite decent performance gain, this work relies solely on naive greedy decoding and performs adaptation across timesteps at a frame level, which may not be optimal given the sequential nature of the model output. Motivated by this, we propose a novel TTA framework, dubbed SGEM, for general ASR models. To treat the sequential output, SGEM first exploits beam search to explore candidate output logits and selects the most plausible one. Then, it utilizes generalized entropy minimization and negative sampling as unsupervised objectives to adapt the model. SGEM achieves state-of-the-art performance for three mainstream ASR models under various domain shifts.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DHAuDS: A Dynamic and Heterogeneous Audio Benchmark for Test-Time Adaptation

    cs.SD 2025-11 conditional novelty 6.0 of 10

    DHAuDS is a new audio benchmark that corrupts four existing datasets with dynamically varying and diverse acoustic noise, and evaluates three classifiers under test-time adaptation.

  2. An Investigation of Test-time Adaptation for Audio Classification under Background Noise

    cs.LG 2025-07 reject novelty 5.0 of 10

    A modified CoNMix method achieved the lowest error rates for audio classification under background noise, but the comparison is confounded and the method was tuned on the test set.

Pith tools