REVIEW 1 cited by
Alternate Intermediate Conditioning with Syllable-level and Character-level Targets for Japanese ASR
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
End-to-end automatic speech recognition directly maps input speech to characters. However, the mapping can be problematic when several different pronunciations should be mapped into one character or when one pronunciation is shared among many different characters. Japanese ASR suffers the most from such many-to-one and one-to-many mapping problems due to Japanese kanji characters. To alleviate the problems, we introduce explicit interaction between characters and syllables using Self-conditioned connectionist temporal classification (CTC), in which the upper layers are ``self-conditioned'' on the intermediate predictions from the lower layers. The proposed method utilizes character-level and syllable-level intermediate predictions as conditioning features to deal with mutual dependency between characters and syllables. Experimental results on Corpus of Spontaneous Japanese show that the proposed method outperformed the conventional multi-task and Self-conditioned CTC methods.
Forward citations
Cited by 1 Pith paper
-
Cross-modal Knowledge Transfer Learning as Graph Matching Based on Optimal Transport for ASR
GM-OT, a fused Wasserstein and Gromov-Wasserstein graph-matching alignment for BERT-to-acoustic knowledge transfer, reports 3.98% CER on AISHELL-1 test versus 5.76% for a Conformer+CTC baseline.
Discussion (0). Continue with ORCID to comment.