Pith. sign in

REVIEW 2 major objections 5 minor 43 references

JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles

T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read The paper constructs JIS, a 169-speaker Japanese live-idol corpus meant to tighten speaker-similarity evaluation in TTS and VC.

desk verdict A genuinely new Japanese multi-speaker corpus with real potential for familiarity-based evaluation, but the central claim that familiarity transfers to read studio speech is unvalidated. read the letter →

arxiv 2506.18296 v2 pith:OO67EL2O submitted 2025-06-23 cs.SD eess.AS

classification cs.SDeess.AS
keywords speechcorpusJapaneseliveidolstext-to-speechsynthesisvoiceconversionspeakersimilarityspeakingstylespreferencenon-anonymousdataset
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper reports the construction of JIS, a speech corpus containing 169 young female Japanese live idols from 33 groups, to be distributed free for non-commercial basic research in text-to-speech and voice conversion. The central claim is that because every speaker is identified by a stage name and belongs to one highly specific category, researchers can recruit listeners who already know these voices, leading to stricter speaker-similarity evaluations than anonymous multi-speaker corpora allow. The corpus includes studio recordings from 61 speakers and quiet-room recordings from 108 speakers, covering phoneme-balanced read sentences, spontaneous greetings, personality phrases, and a short sung piece. The paper also supplies an overview of live idol culture, questionnaire-based voice-impression descriptions, and basic acoustic analyses to guide use of the data.

What carries the argument

The load-bearing object is the corpus itself, specifically its pairing of voice recordings with stage names and group identities. That pairing converts an ordinary multi-speaker corpus into a familiarity-enabled evaluation resource: the same brain processes that make familiar-voice recognition more accurate than unfamiliar-voice recognition can be recruited in listening tests. Around this pairing, the corpus offers style-specific utterances (an energetic post-performance farewell, an intimate photo-session greeting, a self-introduction), phoneme-balanced read sentences, a questionnaire in which each speaker describes her own and her groupmates' voice impressions, and embedding analyses showing that speakers remain distinguishable across styles.

What would settle it

A recognition test in which self-identified fans of specific JIS idols listen to short clips from Speech A and Speech B mixed with voices of same-age women and try to name the idol; if identification accuracy is near chance even for idols the listener claims to support, the stage-name advantage collapses and JIS becomes an ordinary anonymous corpus.

Watch

Extended reading notes

Core claim

JIS is, to the authors' knowledge, the first non-anonymous multi-speaker speech corpus built for speech generation AI research. All 169 speakers are young female Japanese live idols, and each is tied to a publicly used stage name and group affiliation, so an experiment planner can recruit fans who are familiar with the actual people behind the voices. The corpus is split into Speech A (61 speakers, professional studios, the full 100-sentence phoneme-balanced set plus spontaneous speech, greetings, and singing) and Speech B (108 speakers, quiet rooms, a partial set of the same content). Embedded speaker vectors cluster separately from those of an existing Japanese corpus and, within JIS, tend to form per-speaker clusters across speaking styles, which the authors read as evidence that the persona-matching recording instruction was at least partly effective.

Load-bearing premise

The load-bearing premise is that fans' familiarity with an idol from live performances, photo sessions, and social media transfers to recognizing her voice in JIS's studio recordings of read sentences and scripted phrases; the paper does not test this transfer with any listening experiment.

Editorial extensions

If this is right

  • TTS and VC systems can be evaluated by listeners who actually know the target speaker, so similarity scores reflect recognition of a real individual rather than matching of coarse attributes like age and gender.
  • Training and evaluation can be confined to a homogeneous speaker category, removing attribute mismatch as an accidental driver of high similarity ratings.
  • The style-conditioned recordings support research on separately controlling who is speaking and how they are speaking.
  • The questionnaire descriptions of voice impressions provide labels for work on listener-preference voice generation.
  • The documentation of recording instructions and rights transfer can guide other teams that want to collect additional live-idol speech under similar ethical terms.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If familiarity transfers from live performances and social media to these studio recordings, JIS could serve as a benchmark in which failing to recognize an idol's voice is an unambiguous system error; no anonymous corpus offers that oracle.
  • The voice-impression questionnaire could be used to learn a mapping from impression phrases to acoustic speaker embeddings, enabling text-described voice generation aimed at individual listener preferences, a direction the paper mentions but does not implement.
  • Because each speaker appears in several styles, JIS also enables within-speaker style disentanglement experiments without collecting new data.
  • A direct check of the corpus's core premise would be a fan-recognition test: if accuracy is high only for performance-style phrases and not for read sentences, evaluation protocols should weight the speaking styles that carry recognizability.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper introduces JIS, a freely distributed Japanese speech corpus of 169 young female live-idol speakers from 33 groups, recorded in two tiers: Speech A (studio, 61 speakers) and Speech B (quiet rooms, 108 speakers), with a total of 17.0 hours after VAD. The recordings include phoneme-balanced read sentences from the voiceactress100 text set, scripted spontaneous phrases, everyday greetings, and a royalty-free song, stored as 48 kHz/24-bit WAV. The authors argue that because speakers are identified by stage names rather than anonymized, researchers can recruit listeners familiar with these idols, enabling more rigorous speaker-similarity evaluation for TTS and VC than is possible with conventional anonymized corpora. The paper also provides an overview of Japanese live-idol culture, reports supplementary metadata (group, stage name, prefecture, questionnaire responses), and gives basic analyses: region histograms, BERT-based voice-impression embeddings from peer descriptions, UTMOS pseudo-MOS, and x-vector/ECAPA-TDNN speaker embeddings of JIS versus JVS.

Significance. If the corpus works as advertised, it is a valuable community resource: it is one of the first non-anonymous multi-speaker Japanese corpora for speech-generation research, it is free for non-commercial research, and it combines studio-quality and budget recordings with rich metadata that could support new research on familiarity-aware evaluation and listener-preference-driven voice generation. The construction details are concrete and reproducible (text sets, recording instructions, VAD, sampling/quantization), and the authors credit existing corpora and tools such as voiceactress100, JVS, UTMOS, x-vector, and ECAPA-TDNN. The main intellectual contribution is the hypothesis that familiar listeners will evaluate speaker similarity more discriminatorily; however, this hypothesis is not tested in the paper.

major comments (2)
  1. [Abstract, 3.1] The central usefulness claim is that stage names will let researchers recruit listeners familiar with JIS speakers, yielding 'more rigorous evaluations' of speaker similarity. This requires that fans can actually identify the recorded corpus voices as belonging to the named idols, but no listening experiment verifies that familiarity acquired from live performances, photo sessions, and social media transfers to read sentences, scripted spontaneous phrases, and studio speech recorded under an instruction to sound like the idol's persona (Section 3.2). The cited works [12, 13] support the general point that familiar voices are recognized differently and more accurately, but they do not establish the transfer step that is specific to JIS. The paper itself hedges in Section 3.1 ('may have the possibility to recruit listeners') while the abstract asserts the corpus 'will facilitate more rigorous evaluations.' This is load-bearing: without the transfer step, JIS is functionally anonymized for listeners and the 'first non-anonymous corpus' claim reduces to a metadata difference. I recommend adding a recognition experiment (e.g., familiar listeners match held-out JIS samples to stage names) or rewriting the abstract and Section 3.1 to frame familiarity-based evaluation as an untested potential benefit that JIS enables.
  2. [Section 4.2.1, Fig. 3] The pseudo-MOS analysis selects the normalization constant that maximizes the average pseudo-MOS ('We tested various normalization constants, choosing the one with the highest average pseudo MOS'). This procedure makes the reported mean values (Speech A 3.4, Speech B 2.8, JVS 3.7) non-reproducible and systematically optimistic, and it invalidates the intended comparison between recording environments and against JVS, since the constant is fit to the JIS data being compared. Use a pre-specified normalization rule (e.g., fixed target RMS or peak level) or report results for all tested constants. This issue is specific and affects the quantitative analysis that is offered as guidance for using the corpus.
minor comments (5)
  1. [4.1.2] The conclusion that 'overlapping yet offset regions' indicate differences in voice impressions between JIS speakers is based only on visual inspection of a t-SNE plot of BERT embeddings; a quantitative separability measure (e.g., silhouette score or classification accuracy) would make the claim more robust.
  2. [Table 1] The 'Recording environment' row is formatted ambiguously for JIS; the reader must infer that A refers to the studio condition and B to the quiet-room condition. Please spell out 'Speech A: Studio; Speech B: Quiet but unspecified rooms.'
  3. [3.1, 3.3] There is some tension between saying stage names allow fans to identify speakers while 'protecting the speakers' privacy' and the later statement that users must avoid harming idol reputations. Since stage names are public persona identities, the corpus is not anonymized; the privacy framing should be clarified.
  4. [4.2.2] The claim that the instruction to produce persona-consistent speech was 'somewhat effective' (Figure 5) is supported only by visual cluster inspection of t-SNE embeddings; please report a quantitative metric such as within-speaker versus between-speaker cosine similarity.
  5. [Throughout] Minor typographical and formatting issues: 'V AD' should be 'VAD', and the reference URLs are not consistently formatted. These do not affect the substance.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the corpus analyses are descriptive and the evaluation-advantage claim is forward-looking, not derived from fitted inputs.

full rationale

This paper is a corpus construction and analysis report rather than a mathematical or empirical derivation. The central claims are that JIS contains 169 voices of Japanese live idols identified by stage names; that this will enable recruiting listeners familiar with the idols and thereby support more rigorous speaker-similarity evaluation; and that several basic analyses guide future use of the corpus. None of these claims reduce to a fitted parameter renamed as a prediction. The pseudo-MOS analysis states that We tested various normalization constants, choosing the one with the highest average pseudo MOS; this is a disclosed calibration choice for a descriptive statistic, not a prediction evaluated on held-out data. The speaker-embedding and questionnaire analyses are descriptive visualizations using externally trained models, and they do not feed back into the corpus-value claim. The only author self-citation, VoiceGrad, appears in the introduction as a general reference on voice conversion and is not load-bearing. The evaluation-advantage claim rests on external citations about voice familiarity and on the explicit hedged statement that experiment planners may have the possibility to recruit listeners who are familiar with JIS speakers. Whether fans can recognize studio or read recordings of idols is an empirical assumption about familiarity transfer; it is a correctness risk, not a circular step. No equation in the paper defines a derived quantity in terms of the target claim, and no uniqueness theorem or self-citation chain forces the conclusion. Therefore no circularity step meets the evidentiary bar; score 0.

Assumptions & free parameters 1 free parameters · 3 assumptions · 0 invented entities

The corpus introduces no new physical or mathematical entities. The listed assumptions are the ones the paper's usefulness argument rests on; the first is supported by external citations, the second is an unvalidated bridge between stage persona and studio recordings.

free parameters (1)
  • Pseudo-MOS normalization constant = not reported
    Section 4.2.1: 'We tested various normalization constants, choosing the one with the highest average pseudo MOS.' This is a parameter fitted to the scored data, affecting the reported quality comparison.
assumptions (3)
  • domain assumption Familiarity with a speaker improves voice identification and discrimination.
    Section 1 cites Van Lancker and Kreiman (1987) and Schmidt-Nielsen and Stern (1985) to motivate the corpus; the central evaluation benefit depends on this.
  • ad hoc to paper The recorded speech reflects the voice characteristics fans recognize from idol activities.
    Speakers were instructed 'to produce speech that reflected the character they typically express in their idol activities' (Section 3.2), but no evidence links this to externally familiar voices.
  • domain assumption UTMOS pseudo-MOS approximates subjective quality for Japanese idol speech.
    Section 4.2.1 uses UTMOS, trained on TTS/VC MOS data; authors acknowledge linguistic mismatch with training data, so the comparison is approximate.

how reviews work

0 comments
Cite this review

Pith. "Pith review of JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles." pith.science (2026). https://pith.science/paper/OO67EL2O

@misc{pith2026250618296,
  author       = {Pith},
  title        = {Pith review of: JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OO67EL2O}},
  note         = {Machine review of arXiv:2506.18296}
}
read the original abstract

We construct Japanese Idol Speech Corpus (JIS) to advance research in speech generation AI, including text-to-speech synthesis (TTS) and voice conversion (VC). JIS will facilitate more rigorous evaluations of speaker similarity in TTS and VC systems since all speakers in JIS belong to a highly specific category: "young female live idols" in Japan, and each speaker is identified by a stage name, enabling researchers to recruit listeners familiar with these idols for listening experiments. With its unique speaker attributes, JIS will foster compelling research, including generating voices tailored to listener preferences-an area not yet widely studied. JIS will be distributed free of charge to promote research in speech generation AI, with usage restricted to non-commercial, basic research. We describe the construction of JIS, provide an overview of Japanese live idol culture to support effective and ethical use of JIS, and offer a basic analysis to guide application of JIS.

Figures

Figures reproduced from arXiv: 2506.18296 by the authors.

Figure 1
Figure 1. Histogram of JIS Speaker Counts by Group. Idol groups, typically with 3 to 10 members, perform live more than twice a week. Each group has a unique group name, such as “FRUITS ZIPPER” and “iLiFE!.” Their songs fea￾ture divided singing parts, inspiring datasets for singer diariza￾tion [20]. Idols use stage names, allowing experiment planners to engage fans familiar with JIS speakers while protecting the speakers’ pri… view at source ↗
Figure 2
Figure 2. Embedding plots of free descriptions given for JIS speakers of two idol groups ((I) and (II)). The same color rep￾resents the free descriptions given for the same JIS speaker, but the corresponding speakers differ between (I) and (II). 4.1. Analysis of supplementary information The analysis results of supplementary information will provide useful guidelines for designing listening experiments and con￾ditioning speec… view at source ↗
Figure 1
Figure 1. Note that the same color represents the free descriptions [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Embedding plots of JIS and JVS speech data [PITH_FULL_IMAGE:figures/full_fig_p004_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

43 extracted references · 40 canonical work pages

  1. [1]

    JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles

    Introduction We construct a new speech corpus, Japanese Idol Speech Cor- pus (JIS), to advance research in speech generation AI, includ- ing text-to-speech synthesis (TTS) [1, 2, 3, 4] and voice conver- sion (VC) [5, 6, 7]. TTS is a technology that generates speech from text, while VC is a technology that transforms a speaker’s voice into another person’s...

  2. [2]

    The initial challenge in TTS and VC research is to generate voices of speakers included in training data [1]

    Text-to-speech & voice conversion Neural model-based TTS and VC systems have advanced, achieving high speech quality. The initial challenge in TTS and VC research is to generate voices of speakers included in training data [1]. More practical TTS and VC systems have emerged, aiming to generate voices not included in the training data [4, 5]. Aiming for a ...

  3. [3]

    FRUITS ZIPPER

    JIS: Japanese Idol Speech Corpus 3.1. Japanese live idol This section provides an explanation of Japanese live idols [14] from the late 2010s to the early 2020s to promote the effective and ethical use of JIS, as well as to guide readers who may someday record additional live idol voices. Live idols are a key part of Japanese pop culture, enter- taining f...

  4. [4]

    energetic,

    Analysis of JIS This section presents the analysis of JIS. In Sec. 4.1, analysis of supplementary information is discussed, and in Sec. 4.2, analy- sis of the speech data is conducted. 1https://hiroshiba.github.io/voiceactress100_ ruby/ Figure 2: Embedding plots of free descriptions given for JIS speakers of two idol groups ( (I) and (II)). The same color...

  5. [5]

    young male live idols

    Conclusion We constructed a new voice corpus, JIS, composed of Japanese live idol voices. JIS will be distributed free of charge to other research institutions. To encourage the effective and ethical use of JIS and provide guidance for future efforts to collect live idol voices, we provided an overview of Japanese live idol culture. We analyzed JIS to sup...

  6. [6]

    Acknowledgement This work was supported by JST CREST Grant Number JP- MJCR19A3

  7. [7]

    WaveNet: A Generative Model for Raw Audio,

    A. Van Den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “WaveNet: A Generative Model for Raw Audio,” in Proc. ISCA SSW, 2016

  8. [8]

    Conditional Variational Autoen- coder with Adversarial Learning for End-to-End Text-to-Speech,

    J. Kim, J. Kong, and J. Son, “Conditional Variational Autoen- coder with Adversarial Learning for End-to-End Text-to-Speech,” in Proc. PMLR ICML, vol. 139, 2021, pp. 5530–5540

Show all 43 references
  1. [9]

    Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech,

    V . Popov, I. V ovk, V . Gogoryan, T. Sadekova, and M. Kudinov, “Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech,” in Proc. PMLR ICML, vol. 139, 2021, pp. 8599–8608

  2. [10]

    YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot V oice Conversion for Everyone,

    E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. G ¨olge, and M. A. Ponti, “YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot V oice Conversion for Everyone,” in Proc. PMLR ICML, vol. 162, 2022, pp. 2709–2720

  3. [11]

    AutoVC: Zero-Shot V oice Style Transfer with Only Au- toencoder Loss,

    K. Qian, Y . Zhang, S. Chang, X. Yang, and M. Hasegawa- Johnson, “AutoVC: Zero-Shot V oice Style Transfer with Only Au- toencoder Loss,” in Proc. PMLR ICML, vol. 97, 2019, pp. 5210– 5219

  4. [12]

    Diffusion-Based V oice Conversion with Fast Maximum Likelihood Sampling Scheme,

    V . Popov, I. V ovk, V . Gogoryan, T. Sadekova, M. Kudinov, and J. Wei, “Diffusion-Based V oice Conversion with Fast Maximum Likelihood Sampling Scheme,” Proc. PMLR ICLR, 2022

  5. [13]

    V oice- Grad: Non-Parallel Any-to-Many V oice Conversion With An- nealed Langevin Dynamics,

    H. Kameoka, T. Kaneko, K. Tanaka, N. Hojo, and S. Seki, “V oice- Grad: Non-Parallel Any-to-Many V oice Conversion With An- nealed Langevin Dynamics,” in IEEE/ACM Trans. ASLP, vol. 32, pp. 2213–2226, 2024

  6. [14]

    The CMU Arctic Speech Databases,

    J. Kominek and A. W. Black, “The CMU Arctic Speech Databases,” in Proc. ISCA SSW, 2004, pp. 223–224

  7. [15]

    Lib- riSpeech: An ASR Corpus Based on Public Domain Audio Books,

    V . Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Lib- riSpeech: An ASR Corpus Based on Public Domain Audio Books,” in Proc. IEEE ICASSP, 2015, pp. 5206–5210

  8. [16]

    CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR V oice Cloning Toolkit (version 0.92),

    J. Yamagishi, C. Veaux, and K. MacDonald, “CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR V oice Cloning Toolkit (version 0.92),” 2021, accessed: 2025-02-04. [Online]. Available: https://datashare.ed.ac.uk/handle/10283/3443

  9. [17]

    JSUT and JVS: Free Japanese V oice Corpora for Accelerating Speech Synthesis Research,

    S. Takamichi, R. Sonobe, K. Mitsui, Y . Saito, T. Koriyama, N. Tanji, and H. Saruwatari, “JSUT and JVS: Free Japanese V oice Corpora for Accelerating Speech Synthesis Research,” in J. ASJ AST, vol. 41, no. 5, pp. 761–768, 2020

  10. [18]

    V oice Discrimination and Recognition are Separate Abilities,

    D. Van Lancker and J. Kreiman, “V oice Discrimination and Recognition are Separate Abilities,” inNeuropsychologia, vol. 25, no. 5, pp. 829–834, 1987

  11. [19]

    Identification of Known V oices as a Function of Familiarity and Narrow-band Coding,

    A. Schmidt-Nielsen and K. R. Stern, “Identification of Known V oices as a Function of Familiarity and Narrow-band Coding,” in J. ASA, vol. 77, no. 2, pp. 658–663, 1985

  12. [20]

    Seeking‘Mutuality of Eyes’: Understanding Live Idol Experiences from the Perspective of Georg Simmel’s Sociology of the Senses,

    T. Ikeda, “Seeking‘Mutuality of Eyes’: Understanding Live Idol Experiences from the Perspective of Georg Simmel’s Sociology of the Senses,” in Konan Women’s College Researches , no. 57, pp. 163–170, 2021, in Japanese

  13. [21]

    V oice Activity Detection in the Wild: A Data-Driven Approach Using Teacher- Student Training,

    H. Dinkel, S. Wang, X. Xu, M. Wu, and K. Yu, “V oice Activity Detection in the Wild: A Data-Driven Approach Using Teacher- Student Training,” in IEEE/ACM Trans. ASLP, vol. 29, pp. 1542– 1555, 2021

  14. [22]

    Who Finds This V oice Attractive? A Large-Scale Experiment Using In-the-Wild Data,

    H. Suda, A. Watanabe, and S. Takamichi, “Who Finds This V oice Attractive? A Large-Scale Experiment Using In-the-Wild Data,” in Proc. ISCA Interspeech, 2024, pp. 3165–3169

  15. [23]

    DreamV oice: Text-Guided V oice Conversion,

    J. Hai, K. Thakkar, H. Wang, Z. Qin, and M. Elhilali, “DreamV oice: Text-Guided V oice Conversion,” inProc. ISCA In- terspeech, 2024, pp. 4373–4377

  16. [24]

    Generat- ing Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems,

    Z. Chen, X. Liu, E. Cooper, J. Yamagishi, and Y . Qian, “Generat- ing Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems,” in Proc. ISCA Inter- speech, 2024, pp. 4428–4432

  17. [25]

    Fle- Speech: Flexibly Controllable Speech Generation with Various Prompts,

    H. Li, Y . Li, X. Wang, J. Hu, Q. Xie, S. Yang, and L. Xie, “Fle- Speech: Flexibly Controllable Speech Generation with Various Prompts,” arXiv preprint arXiv:2501.04644, 2025

  18. [26]

    FruitsMusic: A Real-World Corpus of Japanese Idol-Group Songs,

    H. Suda, S. Yoshida, T. Nakamura, S. Fukayama, and J. Ogata, “FruitsMusic: A Real-World Corpus of Japanese Idol-Group Songs,” in Proc. ISMIR, 2024

  19. [27]

    Oishi, “Econometric Sociology of ‘Oshi’ that have become Popularized and Polysemous

    M. Oishi, “Econometric Sociology of ‘Oshi’ that have become Popularized and Polysemous. -Multivariate Analysis Focusing on Discrepancies in Perceptions Caused by the Presence or Absence of ‘Oshi’,” in the journal of sociology of education, Kyushu Uni- versity, vol. 27, pp. 47–...

  20. [28]

    YODAS: YouTube-Oriented Dataset for Audio and Speech,

    X. Li, S. Takamichi, T. Saeki, W. Chen, S. Shiota, and S. Watan- abe, “YODAS: YouTube-Oriented Dataset for Audio and Speech,” in Proc. IEEE ASRU, 2023, pp. 1–8

  21. [29]

    COCO-NUT: Corpus of Japanese Utterance and V oice Characteristics Description for Prompt-based Control,

    A. Watanabe, S. Takamichi, Y . Saito, W. Nakata, D. Xin, and H. Saruwatari, “COCO-NUT: Corpus of Japanese Utterance and V oice Characteristics Description for Prompt-based Control,” in Proc. IEEE ASRU, 2023, pp. 1–8

  22. [30]

    V oice-actress corpus,

    y benjo and MagnesiumRibbon, “V oice-actress corpus,” ac- cessed: Feb. 6, 2025. [Online]. Available: http: //voice-statistics.github.io/

  23. [31]

    JVS- MuSiC: Japanese Multispeaker Singing-V oice Corpus,

    H. Tamaru, S. Takamichi, N. Tanji, and H. Saruwatari, “JVS- MuSiC: Japanese Multispeaker Singing-V oice Corpus,” arXiv preprint arXiv:2001.07044, 2020

  24. [32]

    Introduction: The Mirror of Idols and Celebrity,

    P. W. Galbraith and J. G. Karlin, “Introduction: The Mirror of Idols and Celebrity,” inIdols and celebrity in Japanese media cul- ture. Springer, 2012, pp. 1–32

  25. [33]

    Multi-Dialect Speech Synthesis with Interpretable Accent latent Variable based on VQ- V AE,

    K. Yamauchi, Y . Saito, and H. Saruwatari, “Multi-Dialect Speech Synthesis with Interpretable Accent latent Variable based on VQ- V AE,” in IEICE Tech. Rep. , vol. 123, pp. 220–225, 2024, in Japanese

  26. [34]

    Regions of Japan,

    Japan Rail Pass, “Regions of Japan,” accessed: Feb. 6, 2025. [On- line]. Available: https://www.jrailpass.com/blog/regions-of-japan

  27. [35]

    The Classification and Division of Japanese Dialects,

    S. Abe, “The Classification and Division of Japanese Dialects,” Ph.D. dissertation, Gakushuin University, 2015, in Japanese

  28. [36]

    tohoku-nlp/bert-base-japanese,

    tohoku-nlp, “tohoku-nlp/bert-base-japanese,” accessed: Feb. 6,

  29. [38]

    UTMOS: UTokyo-Sarulab System for V oicemos Challenge 2022,

    T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, “UTMOS: UTokyo-Sarulab System for V oicemos Challenge 2022,” in Proc. ISCA Interspeech , 2022, pp. 4521– 4525

  30. [39]

    The V oiceMOS Challenge 2022,

    W.-C. Huang, E. Cooper, Y . Tsao, H.-M. Wang, T. Toda, and J. Ya- magishi, “The V oiceMOS Challenge 2022,” inProc. ISCA Inter- speech, 2022, pp. 4536–4540

  31. [40]

    X-Vectors: Robust DNN Embeddings for Speaker Recogni- tion,

    D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudan- pur, “X-Vectors: Robust DNN Embeddings for Speaker Recogni- tion,” in Proc. IEEE ICASSP, 2018, pp. 5329–5333

  32. [41]

    ECAPA- TDNN: Emphasized Channel Attention, Propagation and Aggre- gation in TDNN Based Speaker Verification,

    B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA- TDNN: Emphasized Channel Attention, Propagation and Aggre- gation in TDNN Based Speaker Verification,” in Proc. ISCA In- terspeech, 2020, pp. 3830–3834

  33. [42]

    sarulab-speech/xvector jtubespeech,

    S. Takamichi, “sarulab-speech/xvector jtubespeech,” accessed: Feb. 6, 2025. [Online]. Available: https://github.com/ sarulab-speech/xvector jtubespeech

  34. [43]

    washi/speaker-emb-ja-ecapa-tdnn,

    k-washi, “washi/speaker-emb-ja-ecapa-tdnn,” accessed: Feb. 6, 2025. [Online]. Available: https://github.com/k-washi/ speaker-emb-ja-ecapa-tdnn

  35. [2025]

    Available: https://huggingface.co/tohoku-nlp/ bert-base-japanese

    [Online]. Available: https://huggingface.co/tohoku-nlp/ bert-base-japanese

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.