Pith. sign in

REVIEW 3 cited by

Exploration of Adapter for Noise Robust Automatic Speech Recognition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.18275 v3 pith:R57F5FLF submitted 2024-02-28 cs.SD cs.CLeess.AS

classification cs.SDcs.CLeess.AS
keywords dataadapteradaptingnoisespeechadaptersautomaticenvironments
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Adapting an automatic speech recognition (ASR) system to unseen noise environments is crucial. Integrating adapters into neural networks has emerged as a potent technique for transfer learning. This study thoroughly investigates adapter-based ASR adaptation in noisy environments. We conducted experiments using the CHiME--4 dataset. The results show that inserting the adapter in the shallow layer yields superior effectiveness, and there is no significant difference between adapting solely within the shallow layer and adapting across all layers. The simulated data helps the system to improve its performance under real noise conditions. Nonetheless, when the amount of data is the same, the real data is more effective than the simulated data. Multi-condition training is still useful for adapter training. Furthermore, integrating adapters into speech enhancement-based ASR systems yields substantial improvements.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Beyond Acoustic Prefixes: Persistent Grounding in Serialized Acoustic Memory for LLM-Based Multi-Talker Speech Recognition

    cs.SD 2026-03 accept novelty 6.0 of 10

    Persistent gated residual cross-attention over onset-ordered talker acoustic memory, refined with LoRA, substantially improves LLM-SOT multi-talker ASR especially on three-talker mixtures.

  2. Serialized Output Prompting for Large Language Model-based Multi-Talker Speech Recognition

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Serialized CTC outputs from auxiliary branches, used as LLM prompts, improve LLM-based multi-talker ASR WER on Libri2Mix and Libri3Mix.

  3. Adapting Whisper for Code-Switching through Encoding Refining and Language-Aware Decoding

    cs.CL 2024-12 conditional novelty 5.0 of 10

    An LSTM-based encoder refiner plus language-aware dual adapters with a fusion module cuts Mandarin-English code-switching ASR errors on SEAME.

Pith tools