Pith. sign in

REVIEW 1 cited by

Alternate Intermediate Conditioning with Syllable-level and Character-level Targets for Japanese ASR

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.00175 v2 pith:OHH2XYWR submitted 2022-04-01 cs.CL cs.SDeess.AS

classification cs.CLcs.SDeess.AS
keywords charactersjapaneseintermediateself-conditionedcharacter-levelconditioningdifferentlayers
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

End-to-end automatic speech recognition directly maps input speech to characters. However, the mapping can be problematic when several different pronunciations should be mapped into one character or when one pronunciation is shared among many different characters. Japanese ASR suffers the most from such many-to-one and one-to-many mapping problems due to Japanese kanji characters. To alleviate the problems, we introduce explicit interaction between characters and syllables using Self-conditioned connectionist temporal classification (CTC), in which the upper layers are ``self-conditioned'' on the intermediate predictions from the lower layers. The proposed method utilizes character-level and syllable-level intermediate predictions as conditioning features to deal with mutual dependency between characters and syllables. Experimental results on Corpus of Spontaneous Japanese show that the proposed method outperformed the conventional multi-task and Self-conditioned CTC methods.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cross-modal Knowledge Transfer Learning as Graph Matching Based on Optimal Transport for ASR

    eess.AS 2025-05 conditional novelty 5.0 of 10

    GM-OT, a fused Wasserstein and Gromov-Wasserstein graph-matching alignment for BERT-to-acoustic knowledge transfer, reports 3.98% CER on AISHELL-1 test versus 5.76% for a Conformer+CTC baseline.

Pith tools