Pith. sign in

REVIEW 9 cited by

JVS corpus: free Japanese multi-speaker voice corpus

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1908.06248 v1 pith:WC3PZTIH submitted 2019-08-17 cs.SD eess.AS

classification cs.SDeess.AS
keywords corpusvoicespeechsynthesiscontainshourslearningdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Thanks to improvements in machine learning techniques, including deep learning, speech synthesis is becoming a machine learning task. To accelerate speech synthesis research, we are developing Japanese voice corpora reasonably accessible from not only academic institutions but also commercial companies. In 2017, we released the JSUT corpus, which contains 10 hours of reading-style speech uttered by a single speaker, for end-to-end text-to-speech synthesis. For more general use in speech synthesis research, e.g., voice conversion and multi-speaker modeling, in this paper, we construct the JVS corpus, which contains voice data of 100 speakers in three styles (normal, whisper, and falsetto). The corpus contains 30 hours of voice data including 22 hours of parallel normal voices. This paper describes how we designed the corpus and summarizes the specifications. The corpus is available at our project page.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors

    eess.AS 2026-07 conditional novelty 6.0 of 10

    On 1,168 professional voice actors, a misidentification floor in speaker embeddings survives calibration, normalization, and discriminative re-ranking, and the same floor makes fixed-threshold voice-clone attribution ...

  2. Active Learning for Text-to-Speech Synthesis with Informative Sample Collection

    cs.SD 2025-07 conditional novelty 6.0 of 10

    An iterative active learning pipeline that filters web speech data by predicted quality and redundancy produces a TTS corpus with better speaker coverage at the same size.

  3. Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora

    eess.AS 2025-07 conditional novelty 6.0 of 10

    A voice conversion pipeline conditions FastSpeech 2 on predicted likability ratings, achieving partial subjective control but with degraded identity and content at strong settings.

  4. Transcript-Prompted Whisper with Dictionary-Enhanced Decoding for Japanese Speech Annotation

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A transcript-prompted Whisper model with dictionary-based decoding automatically produces phonemic and prosodic annotations for Japanese audio-transcript pairs, improving Japanese TTS naturalness.

  5. MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model

    eess.AS 2025-09 conditional novelty 5.0 of 10

    A T5 model predicts SSL-derived discrete speech tokens directly from mixed-script Japanese text, letting a FastSpeech 2 synthesizer produce speech without a grapheme-to-phoneme module.

  6. QHARMA-GAN: Quasi-Harmonic Neural Vocoder based on Autoregressive Moving Average Model

    eess.AS 2025-07 conditional novelty 5.0 of 10

    QHARMA-GAN combines quasi-harmonic speech modeling with a neural-network-estimated ARMA filter to synthesize and modify speech.

  7. SpeechAccentLLM: A Unified Framework for Foreign Accent Conversion and Text to Speech

    eess.AS 2025-07 conditional novelty 5.0 of 10

    SpeechAccentLLM jointly trains foreign accent conversion and text-to-speech on CTC-regularized discrete speech tokens, with a BERT-style restorer, and reports improved accent reduction and intelligibility over one baseline.

  8. Mitigating Language Mismatch in SSL-Based Speaker Anonymization

    eess.AS 2025-07 conditional novelty 5.0 of 10

    Fine-tuning an SSL content encoder on Japanese, especially when the encoder is pre-trained multilingually, makes anonymized Japanese and Mandarin speech much more intelligible while keeping speaker privacy at usable levels.

  9. Comparative Analysis of Fast and High-Fidelity Neural Vocoders for Low-Latency Streaming Synthesis in Resource-Constrained Environments

    cs.SD 2025-06 conditional novelty 5.0 of 10

    MS-Wavehax, a multi-stream extension of the Wavehax vocoder, achieves the best throughput in low-latency CPU streaming and near non-causal quality with one frame of lookahead.

Pith tools