Pith. sign in

REVIEW 3 cited by

Lenient Evaluation of Japanese Speech Recognition: Modeling Naturally Occurring Spelling Inconsistency

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.04530 v1 pith:F5JXMRJX submitted 2023-06-07 cs.CL

Lenient Evaluation of Japanese Speech Recognition: Modeling Naturally Occurring Spelling Inconsistency

classification cs.CL
keywords evaluationjapaneseerrorspellingsystemwordlenientplausible
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Word error rate (WER) and character error rate (CER) are standard metrics in Speech Recognition (ASR), but one problem has always been alternative spellings: If one's system transcribes adviser whereas the ground truth has advisor, this will count as an error even though the two spellings really represent the same word. Japanese is notorious for ``lacking orthography'': most words can be spelled in multiple ways, presenting a problem for accurate ASR evaluation. In this paper we propose a new lenient evaluation metric as a more defensible CER measure for Japanese ASR. We create a lattice of plausible respellings of the reference transcription, using a combination of lexical resources, a Japanese text-processing system, and a neural machine translation model for reconstructing kanji from hiragana or katakana. In a manual evaluation, raters rated 95.4% of the proposed spelling variants as plausible. ASR results show that our method, which does not penalize the system for choosing a valid alternate spelling of a word, affords a 2.4%-3.1% absolute reduction in CER depending on the task.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India

    cs.CL 2026-04 unverdicted novelty 7.0

    Voice of India is a new 536-hour benchmark of real telephonic conversations in 15 Indian languages with variant-aware transcripts for more realistic ASR evaluation.

  2. Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India

    cs.CL 2026-04 conditional novelty 6.0

    A 536-hour unscripted telephonic ASR benchmark covering 15 Indian languages and 139 regional clusters, with multi-reference transcripts for spelling variation and district-level performance analysis.

  3. Voice of India: A Large-Scale Benchmark for Real-World Speech Recognition in India

    cs.CL 2026-04 unverdicted novelty 5.0

    A 536-hour, 15-language, 139-cluster telephonic ASR benchmark for Indian languages with spelling-variation-aware transcripts and geographic performance analysis.