Pith. sign in

REVIEW 4 cited by

JTubeSpeech: corpus of Japanese speech collected from YouTube for speech recognition and speaker verification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.09323 v1 pith:6F4W3PUC submitted 2021-12-17 cs.SD eess.AS

classification cs.SDeess.AS
keywords speechjapanesespeakercorpusrecognitionverificationautomaticcorpora
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we construct a new Japanese speech corpus called "JTubeSpeech." Although recent end-to-end learning requires large-size speech corpora, open-sourced such corpora for languages other than English have not yet been established. In this paper, we describe the construction of a corpus from YouTube videos and subtitles for speech recognition and speaker verification. Our method can automatically filter the videos and subtitles with almost no language-dependent processes. We consistently employ Connectionist Temporal Classification (CTC)-based techniques for automatic speech recognition (ASR) and a speaker variation-based method for automatic speaker verification (ASV). We build 1) a large-scale Japanese ASR benchmark with more than 1,300 hours of data and 2) 900 hours of data for Japanese ASV.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors

    eess.AS 2026-07 conditional novelty 6.0 of 10

    On 1,168 professional voice actors, a misidentification floor in speaker embeddings survives calibration, normalization, and discriminative re-ranking, and the same floor makes fixed-threshold voice-clone attribution ...

  2. Active Learning for Text-to-Speech Synthesis with Informative Sample Collection

    cs.SD 2025-07 conditional novelty 6.0 of 10

    An iterative active learning pipeline that filters web speech data by predicted quality and redundancy produces a TTS corpus with better speaker coverage at the same size.

  3. Mitigating Language Mismatch in SSL-Based Speaker Anonymization

    eess.AS 2025-07 conditional novelty 5.0 of 10

    Fine-tuning an SSL content encoder on Japanese, especially when the encoder is pre-trained multilingually, makes anonymized Japanese and Mandarin speech much more intelligible while keeping speaker privacy at usable levels.

  4. Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Whale, a 1.87B-parameter ASR model combining w2v-BERT and E-Branchformer, reports 2.4% WER on Librispeech test-clean and 3.4% CER on CSJ eval3, beating Whisper large-v3 and OWSM v3.1 on those benchmarks.

Pith tools