REVIEW 4 cited by
JTubeSpeech: corpus of Japanese speech collected from YouTube for speech recognition and speaker verification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we construct a new Japanese speech corpus called "JTubeSpeech." Although recent end-to-end learning requires large-size speech corpora, open-sourced such corpora for languages other than English have not yet been established. In this paper, we describe the construction of a corpus from YouTube videos and subtitles for speech recognition and speaker verification. Our method can automatically filter the videos and subtitles with almost no language-dependent processes. We consistently employ Connectionist Temporal Classification (CTC)-based techniques for automatic speech recognition (ASR) and a speaker variation-based method for automatic speaker verification (ASV). We build 1) a large-scale Japanese ASR benchmark with more than 1,300 hours of data and 2) 900 hours of data for Japanese ASV.
Forward citations
Cited by 4 Pith papers
-
A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors
On 1,168 professional voice actors, a misidentification floor in speaker embeddings survives calibration, normalization, and discriminative re-ranking, and the same floor makes fixed-threshold voice-clone attribution ...
-
Active Learning for Text-to-Speech Synthesis with Informative Sample Collection
An iterative active learning pipeline that filters web speech data by predicted quality and redundancy produces a TTS corpus with better speaker coverage at the same size.
-
Mitigating Language Mismatch in SSL-Based Speaker Anonymization
Fine-tuning an SSL content encoder on Japanese, especially when the encoder is pre-trained multilingually, makes anonymized Japanese and Mandarin speech much more intelligible while keeping speaker privacy at usable levels.
-
Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data
Whale, a 1.87B-parameter ASR model combining w2v-BERT and E-Branchformer, reports 2.4% WER on Librispeech test-clean and 3.4% CER on CSJ eval3, beating Whisper large-v3 and OWSM v3.1 on those benchmarks.
Discussion (0). Continue with ORCID to comment.