InterAligner adds progressive intermediate alignment objectives plus InterCTC to Aligner-Encoder ASR, cutting WER from 5.0/7.8 to 3.1/5.6 on LibriSpeech test-clean/other with largest gains on long utterances.
Progressive Alignment Objectives for Aligner-Encoder based ASR
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Aligner-Encoders are recently proposed seq2seq end-to-end ASR models that replace decoder attention by predicting the uth token directly from the u-th encoder position, so the encoder must learn the alignment internally without cross-attention or a transducer lattice. In practice, this alignment often forms abruptly in the upper layers, making training sensitive and brittle on long utterances. We propose InterAligner, which adds an intermediate Aligner objective so alignment can form progressively across depth, together with an intermediate CTC loss (InterCTC) to stabilize optimization. On LibriSpeech with a 17-layer Conformer, a final-only Aligner reaches 5.0/7.8 WER (test-clean/other). InterCTC improves to 3.4/6.0, and InterAligner further reduces WER to 3.1/5.6 with the largest gains on long utterances.
fields
eess.AS 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Progressive Alignment Objectives for Aligner-Encoder based ASR
InterAligner adds progressive intermediate alignment objectives plus InterCTC to Aligner-Encoder ASR, cutting WER from 5.0/7.8 to 3.1/5.6 on LibriSpeech test-clean/other with largest gains on long utterances.