REVIEW 2 major objections 5 minor 43 references
JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles
T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper constructs JIS, a 169-speaker Japanese live-idol corpus meant to tighten speaker-similarity evaluation in TTS and VC.
desk verdict A genuinely new Japanese multi-speaker corpus with real potential for familiarity-based evaluation, but the central claim that familiarity transfers to read studio speech is unvalidated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the corpus itself, specifically its pairing of voice recordings with stage names and group identities. That pairing converts an ordinary multi-speaker corpus into a familiarity-enabled evaluation resource: the same brain processes that make familiar-voice recognition more accurate than unfamiliar-voice recognition can be recruited in listening tests. Around this pairing, the corpus offers style-specific utterances (an energetic post-performance farewell, an intimate photo-session greeting, a self-introduction), phoneme-balanced read sentences, a questionnaire in which each speaker describes her own and her groupmates' voice impressions, and embedding analyses showing that speakers remain distinguishable across styles.
What would settle it
A recognition test in which self-identified fans of specific JIS idols listen to short clips from Speech A and Speech B mixed with voices of same-age women and try to name the idol; if identification accuracy is near chance even for idols the listener claims to support, the stage-name advantage collapses and JIS becomes an ordinary anonymous corpus.
Extended reading notes
Core claim
JIS is, to the authors' knowledge, the first non-anonymous multi-speaker speech corpus built for speech generation AI research. All 169 speakers are young female Japanese live idols, and each is tied to a publicly used stage name and group affiliation, so an experiment planner can recruit fans who are familiar with the actual people behind the voices. The corpus is split into Speech A (61 speakers, professional studios, the full 100-sentence phoneme-balanced set plus spontaneous speech, greetings, and singing) and Speech B (108 speakers, quiet rooms, a partial set of the same content). Embedded speaker vectors cluster separately from those of an existing Japanese corpus and, within JIS, tend to form per-speaker clusters across speaking styles, which the authors read as evidence that the persona-matching recording instruction was at least partly effective.
Load-bearing premise
The load-bearing premise is that fans' familiarity with an idol from live performances, photo sessions, and social media transfers to recognizing her voice in JIS's studio recordings of read sentences and scripted phrases; the paper does not test this transfer with any listening experiment.
Editorial extensions
If this is right
- TTS and VC systems can be evaluated by listeners who actually know the target speaker, so similarity scores reflect recognition of a real individual rather than matching of coarse attributes like age and gender.
- Training and evaluation can be confined to a homogeneous speaker category, removing attribute mismatch as an accidental driver of high similarity ratings.
- The style-conditioned recordings support research on separately controlling who is speaking and how they are speaking.
- The questionnaire descriptions of voice impressions provide labels for work on listener-preference voice generation.
- The documentation of recording instructions and rights transfer can guide other teams that want to collect additional live-idol speech under similar ethical terms.
Reading between the lines
- If familiarity transfers from live performances and social media to these studio recordings, JIS could serve as a benchmark in which failing to recognize an idol's voice is an unambiguous system error; no anonymous corpus offers that oracle.
- The voice-impression questionnaire could be used to learn a mapping from impression phrases to acoustic speaker embeddings, enabling text-described voice generation aimed at individual listener preferences, a direction the paper mentions but does not implement.
- Because each speaker appears in several styles, JIS also enables within-speaker style disentanglement experiments without collecting new data.
- A direct check of the corpus's core premise would be a fan-recognition test: if accuracy is high only for performance-style phrases and not for read sentences, evaluation protocols should weight the speaking styles that carry recognizability.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces JIS, a freely distributed Japanese speech corpus of 169 young female live-idol speakers from 33 groups, recorded in two tiers: Speech A (studio, 61 speakers) and Speech B (quiet rooms, 108 speakers), with a total of 17.0 hours after VAD. The recordings include phoneme-balanced read sentences from the voiceactress100 text set, scripted spontaneous phrases, everyday greetings, and a royalty-free song, stored as 48 kHz/24-bit WAV. The authors argue that because speakers are identified by stage names rather than anonymized, researchers can recruit listeners familiar with these idols, enabling more rigorous speaker-similarity evaluation for TTS and VC than is possible with conventional anonymized corpora. The paper also provides an overview of Japanese live-idol culture, reports supplementary metadata (group, stage name, prefecture, questionnaire responses), and gives basic analyses: region histograms, BERT-based voice-impression embeddings from peer descriptions, UTMOS pseudo-MOS, and x-vector/ECAPA-TDNN speaker embeddings of JIS versus JVS.
Significance. If the corpus works as advertised, it is a valuable community resource: it is one of the first non-anonymous multi-speaker Japanese corpora for speech-generation research, it is free for non-commercial research, and it combines studio-quality and budget recordings with rich metadata that could support new research on familiarity-aware evaluation and listener-preference-driven voice generation. The construction details are concrete and reproducible (text sets, recording instructions, VAD, sampling/quantization), and the authors credit existing corpora and tools such as voiceactress100, JVS, UTMOS, x-vector, and ECAPA-TDNN. The main intellectual contribution is the hypothesis that familiar listeners will evaluate speaker similarity more discriminatorily; however, this hypothesis is not tested in the paper.
major comments (2)
- [Abstract, 3.1] The central usefulness claim is that stage names will let researchers recruit listeners familiar with JIS speakers, yielding 'more rigorous evaluations' of speaker similarity. This requires that fans can actually identify the recorded corpus voices as belonging to the named idols, but no listening experiment verifies that familiarity acquired from live performances, photo sessions, and social media transfers to read sentences, scripted spontaneous phrases, and studio speech recorded under an instruction to sound like the idol's persona (Section 3.2). The cited works [12, 13] support the general point that familiar voices are recognized differently and more accurately, but they do not establish the transfer step that is specific to JIS. The paper itself hedges in Section 3.1 ('may have the possibility to recruit listeners') while the abstract asserts the corpus 'will facilitate more rigorous evaluations.' This is load-bearing: without the transfer step, JIS is functionally anonymized for listeners and the 'first non-anonymous corpus' claim reduces to a metadata difference. I recommend adding a recognition experiment (e.g., familiar listeners match held-out JIS samples to stage names) or rewriting the abstract and Section 3.1 to frame familiarity-based evaluation as an untested potential benefit that JIS enables.
- [Section 4.2.1, Fig. 3] The pseudo-MOS analysis selects the normalization constant that maximizes the average pseudo-MOS ('We tested various normalization constants, choosing the one with the highest average pseudo MOS'). This procedure makes the reported mean values (Speech A 3.4, Speech B 2.8, JVS 3.7) non-reproducible and systematically optimistic, and it invalidates the intended comparison between recording environments and against JVS, since the constant is fit to the JIS data being compared. Use a pre-specified normalization rule (e.g., fixed target RMS or peak level) or report results for all tested constants. This issue is specific and affects the quantitative analysis that is offered as guidance for using the corpus.
minor comments (5)
- [4.1.2] The conclusion that 'overlapping yet offset regions' indicate differences in voice impressions between JIS speakers is based only on visual inspection of a t-SNE plot of BERT embeddings; a quantitative separability measure (e.g., silhouette score or classification accuracy) would make the claim more robust.
- [Table 1] The 'Recording environment' row is formatted ambiguously for JIS; the reader must infer that A refers to the studio condition and B to the quiet-room condition. Please spell out 'Speech A: Studio; Speech B: Quiet but unspecified rooms.'
- [3.1, 3.3] There is some tension between saying stage names allow fans to identify speakers while 'protecting the speakers' privacy' and the later statement that users must avoid harming idol reputations. Since stage names are public persona identities, the corpus is not anonymized; the privacy framing should be clarified.
- [4.2.2] The claim that the instruction to produce persona-consistent speech was 'somewhat effective' (Figure 5) is supported only by visual cluster inspection of t-SNE embeddings; please report a quantitative metric such as within-speaker versus between-speaker cosine similarity.
- [Throughout] Minor typographical and formatting issues: 'V AD' should be 'VAD', and the reference URLs are not consistently formatted. These do not affect the substance.
Circularity Check
No significant circularity: the corpus analyses are descriptive and the evaluation-advantage claim is forward-looking, not derived from fitted inputs.
full rationale
This paper is a corpus construction and analysis report rather than a mathematical or empirical derivation. The central claims are that JIS contains 169 voices of Japanese live idols identified by stage names; that this will enable recruiting listeners familiar with the idols and thereby support more rigorous speaker-similarity evaluation; and that several basic analyses guide future use of the corpus. None of these claims reduce to a fitted parameter renamed as a prediction. The pseudo-MOS analysis states that We tested various normalization constants, choosing the one with the highest average pseudo MOS; this is a disclosed calibration choice for a descriptive statistic, not a prediction evaluated on held-out data. The speaker-embedding and questionnaire analyses are descriptive visualizations using externally trained models, and they do not feed back into the corpus-value claim. The only author self-citation, VoiceGrad, appears in the introduction as a general reference on voice conversion and is not load-bearing. The evaluation-advantage claim rests on external citations about voice familiarity and on the explicit hedged statement that experiment planners may have the possibility to recruit listeners who are familiar with JIS speakers. Whether fans can recognize studio or read recordings of idols is an empirical assumption about familiarity transfer; it is a correctness risk, not a circular step. No equation in the paper defines a derived quantity in terms of the target claim, and no uniqueness theorem or self-citation chain forces the conclusion. Therefore no circularity step meets the evidentiary bar; score 0.
Assumptions & free parameters
free parameters (1)
- Pseudo-MOS normalization constant =
not reported
assumptions (3)
- domain assumption Familiarity with a speaker improves voice identification and discrimination.
- ad hoc to paper The recorded speech reflects the voice characteristics fans recognize from idol activities.
- domain assumption UTMOS pseudo-MOS approximates subjective quality for Japanese idol speech.
Cite this review
Pith. "Pith review of JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles." pith.science (2026). https://pith.science/paper/OO67EL2O
@misc{pith2026250618296,
author = {Pith},
title = {Pith review of: JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles},
year = {2026},
howpublished = {\url{https://pith.science/paper/OO67EL2O}},
note = {Machine review of arXiv:2506.18296}
}
read the original abstract
We construct Japanese Idol Speech Corpus (JIS) to advance research in speech generation AI, including text-to-speech synthesis (TTS) and voice conversion (VC). JIS will facilitate more rigorous evaluations of speaker similarity in TTS and VC systems since all speakers in JIS belong to a highly specific category: "young female live idols" in Japan, and each speaker is identified by a stage name, enabling researchers to recruit listeners familiar with these idols for listening experiments. With its unique speaker attributes, JIS will foster compelling research, including generating voices tailored to listener preferences-an area not yet widely studied. JIS will be distributed free of charge to promote research in speech generation AI, with usage restricted to non-commercial, basic research. We describe the construction of JIS, provide an overview of Japanese live idol culture to support effective and ethical use of JIS, and offer a basic analysis to guide application of JIS.
Figures
Reference graph
Works this paper leans on
-
[1]
JIS: A Speech Corpus of Japanese Idol Speakers with Various Speaking Styles
Introduction We construct a new speech corpus, Japanese Idol Speech Cor- pus (JIS), to advance research in speech generation AI, includ- ing text-to-speech synthesis (TTS) [1, 2, 3, 4] and voice conver- sion (VC) [5, 6, 7]. TTS is a technology that generates speech from text, while VC is a technology that transforms a speaker’s voice into another person’s...
work page Pith review arXiv 2025
-
[2]
Text-to-speech & voice conversion Neural model-based TTS and VC systems have advanced, achieving high speech quality. The initial challenge in TTS and VC research is to generate voices of speakers included in training data [1]. More practical TTS and VC systems have emerged, aiming to generate voices not included in the training data [4, 5]. Aiming for a ...
-
[3]
JIS: Japanese Idol Speech Corpus 3.1. Japanese live idol This section provides an explanation of Japanese live idols [14] from the late 2010s to the early 2020s to promote the effective and ethical use of JIS, as well as to guide readers who may someday record additional live idol voices. Live idols are a key part of Japanese pop culture, enter- taining f...
-
[4]
Analysis of JIS This section presents the analysis of JIS. In Sec. 4.1, analysis of supplementary information is discussed, and in Sec. 4.2, analy- sis of the speech data is conducted. 1https://hiroshiba.github.io/voiceactress100_ ruby/ Figure 2: Embedding plots of free descriptions given for JIS speakers of two idol groups ( (I) and (II)). The same color...
work page 2022
-
[5]
Conclusion We constructed a new voice corpus, JIS, composed of Japanese live idol voices. JIS will be distributed free of charge to other research institutions. To encourage the effective and ethical use of JIS and provide guidance for future efforts to collect live idol voices, we provided an overview of Japanese live idol culture. We analyzed JIS to sup...
-
[6]
Acknowledgement This work was supported by JST CREST Grant Number JP- MJCR19A3
-
[7]
WaveNet: A Generative Model for Raw Audio,
A. Van Den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, and K. Kavukcuoglu, “WaveNet: A Generative Model for Raw Audio,” in Proc. ISCA SSW, 2016
work page 2016
-
[8]
Conditional Variational Autoen- coder with Adversarial Learning for End-to-End Text-to-Speech,
J. Kim, J. Kong, and J. Son, “Conditional Variational Autoen- coder with Adversarial Learning for End-to-End Text-to-Speech,” in Proc. PMLR ICML, vol. 139, 2021, pp. 5530–5540
work page 2021
Show all 43 references
-
[9]
Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech,
V . Popov, I. V ovk, V . Gogoryan, T. Sadekova, and M. Kudinov, “Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech,” in Proc. PMLR ICML, vol. 139, 2021, pp. 8599–8608
2021
-
[10]
YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot V oice Conversion for Everyone,
E. Casanova, J. Weber, C. D. Shulby, A. C. Junior, E. G ¨olge, and M. A. Ponti, “YourTTS: Towards Zero-Shot Multi-Speaker TTS and Zero-Shot V oice Conversion for Everyone,” in Proc. PMLR ICML, vol. 162, 2022, pp. 2709–2720
2022
-
[11]
AutoVC: Zero-Shot V oice Style Transfer with Only Au- toencoder Loss,
K. Qian, Y . Zhang, S. Chang, X. Yang, and M. Hasegawa- Johnson, “AutoVC: Zero-Shot V oice Style Transfer with Only Au- toencoder Loss,” in Proc. PMLR ICML, vol. 97, 2019, pp. 5210– 5219
2019
-
[12]
Diffusion-Based V oice Conversion with Fast Maximum Likelihood Sampling Scheme,
V . Popov, I. V ovk, V . Gogoryan, T. Sadekova, M. Kudinov, and J. Wei, “Diffusion-Based V oice Conversion with Fast Maximum Likelihood Sampling Scheme,” Proc. PMLR ICLR, 2022
2022
-
[13]
V oice- Grad: Non-Parallel Any-to-Many V oice Conversion With An- nealed Langevin Dynamics,
H. Kameoka, T. Kaneko, K. Tanaka, N. Hojo, and S. Seki, “V oice- Grad: Non-Parallel Any-to-Many V oice Conversion With An- nealed Langevin Dynamics,” in IEEE/ACM Trans. ASLP, vol. 32, pp. 2213–2226, 2024
2024
-
[14]
The CMU Arctic Speech Databases,
J. Kominek and A. W. Black, “The CMU Arctic Speech Databases,” in Proc. ISCA SSW, 2004, pp. 223–224
2004
-
[15]
Lib- riSpeech: An ASR Corpus Based on Public Domain Audio Books,
V . Panayotov, G. Chen, D. Povey, and S. Khudanpur, “Lib- riSpeech: An ASR Corpus Based on Public Domain Audio Books,” in Proc. IEEE ICASSP, 2015, pp. 5206–5210
2015
-
[16]
CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR V oice Cloning Toolkit (version 0.92),
J. Yamagishi, C. Veaux, and K. MacDonald, “CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR V oice Cloning Toolkit (version 0.92),” 2021, accessed: 2025-02-04. [Online]. Available: https://datashare.ed.ac.uk/handle/10283/3443
2021
-
[17]
JSUT and JVS: Free Japanese V oice Corpora for Accelerating Speech Synthesis Research,
S. Takamichi, R. Sonobe, K. Mitsui, Y . Saito, T. Koriyama, N. Tanji, and H. Saruwatari, “JSUT and JVS: Free Japanese V oice Corpora for Accelerating Speech Synthesis Research,” in J. ASJ AST, vol. 41, no. 5, pp. 761–768, 2020
2020
-
[18]
V oice Discrimination and Recognition are Separate Abilities,
D. Van Lancker and J. Kreiman, “V oice Discrimination and Recognition are Separate Abilities,” inNeuropsychologia, vol. 25, no. 5, pp. 829–834, 1987
1987
-
[19]
Identification of Known V oices as a Function of Familiarity and Narrow-band Coding,
A. Schmidt-Nielsen and K. R. Stern, “Identification of Known V oices as a Function of Familiarity and Narrow-band Coding,” in J. ASA, vol. 77, no. 2, pp. 658–663, 1985
1985
-
[20]
Seeking‘Mutuality of Eyes’: Understanding Live Idol Experiences from the Perspective of Georg Simmel’s Sociology of the Senses,
T. Ikeda, “Seeking‘Mutuality of Eyes’: Understanding Live Idol Experiences from the Perspective of Georg Simmel’s Sociology of the Senses,” in Konan Women’s College Researches , no. 57, pp. 163–170, 2021, in Japanese
2021
-
[21]
V oice Activity Detection in the Wild: A Data-Driven Approach Using Teacher- Student Training,
H. Dinkel, S. Wang, X. Xu, M. Wu, and K. Yu, “V oice Activity Detection in the Wild: A Data-Driven Approach Using Teacher- Student Training,” in IEEE/ACM Trans. ASLP, vol. 29, pp. 1542– 1555, 2021
2021
-
[22]
Who Finds This V oice Attractive? A Large-Scale Experiment Using In-the-Wild Data,
H. Suda, A. Watanabe, and S. Takamichi, “Who Finds This V oice Attractive? A Large-Scale Experiment Using In-the-Wild Data,” in Proc. ISCA Interspeech, 2024, pp. 3165–3169
2024
-
[23]
DreamV oice: Text-Guided V oice Conversion,
J. Hai, K. Thakkar, H. Wang, Z. Qin, and M. Elhilali, “DreamV oice: Text-Guided V oice Conversion,” inProc. ISCA In- terspeech, 2024, pp. 4373–4377
2024
-
[24]
Generat- ing Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems,
Z. Chen, X. Liu, E. Cooper, J. Yamagishi, and Y . Qian, “Generat- ing Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems,” in Proc. ISCA Inter- speech, 2024, pp. 4428–4432
2024
-
[25]
Fle- Speech: Flexibly Controllable Speech Generation with Various Prompts,
H. Li, Y . Li, X. Wang, J. Hu, Q. Xie, S. Yang, and L. Xie, “Fle- Speech: Flexibly Controllable Speech Generation with Various Prompts,” arXiv preprint arXiv:2501.04644, 2025
2025 arXiv
-
[26]
FruitsMusic: A Real-World Corpus of Japanese Idol-Group Songs,
H. Suda, S. Yoshida, T. Nakamura, S. Fukayama, and J. Ogata, “FruitsMusic: A Real-World Corpus of Japanese Idol-Group Songs,” in Proc. ISMIR, 2024
2024
-
[27]
Oishi, “Econometric Sociology of ‘Oshi’ that have become Popularized and Polysemous
M. Oishi, “Econometric Sociology of ‘Oshi’ that have become Popularized and Polysemous. -Multivariate Analysis Focusing on Discrepancies in Perceptions Caused by the Presence or Absence of ‘Oshi’,” in the journal of sociology of education, Kyushu Uni- versity, vol. 27, pp. 47–...
2024
-
[28]
YODAS: YouTube-Oriented Dataset for Audio and Speech,
X. Li, S. Takamichi, T. Saeki, W. Chen, S. Shiota, and S. Watan- abe, “YODAS: YouTube-Oriented Dataset for Audio and Speech,” in Proc. IEEE ASRU, 2023, pp. 1–8
2023
-
[29]
COCO-NUT: Corpus of Japanese Utterance and V oice Characteristics Description for Prompt-based Control,
A. Watanabe, S. Takamichi, Y . Saito, W. Nakata, D. Xin, and H. Saruwatari, “COCO-NUT: Corpus of Japanese Utterance and V oice Characteristics Description for Prompt-based Control,” in Proc. IEEE ASRU, 2023, pp. 1–8
2023
-
[30]
V oice-actress corpus,
y benjo and MagnesiumRibbon, “V oice-actress corpus,” ac- cessed: Feb. 6, 2025. [Online]. Available: http: //voice-statistics.github.io/
2025
-
[31]
JVS- MuSiC: Japanese Multispeaker Singing-V oice Corpus,
H. Tamaru, S. Takamichi, N. Tanji, and H. Saruwatari, “JVS- MuSiC: Japanese Multispeaker Singing-V oice Corpus,” arXiv preprint arXiv:2001.07044, 2020
2001 arXiv
-
[32]
Introduction: The Mirror of Idols and Celebrity,
P. W. Galbraith and J. G. Karlin, “Introduction: The Mirror of Idols and Celebrity,” inIdols and celebrity in Japanese media cul- ture. Springer, 2012, pp. 1–32
2012
-
[33]
Multi-Dialect Speech Synthesis with Interpretable Accent latent Variable based on VQ- V AE,
K. Yamauchi, Y . Saito, and H. Saruwatari, “Multi-Dialect Speech Synthesis with Interpretable Accent latent Variable based on VQ- V AE,” in IEICE Tech. Rep. , vol. 123, pp. 220–225, 2024, in Japanese
2024
-
[34]
Regions of Japan,
Japan Rail Pass, “Regions of Japan,” accessed: Feb. 6, 2025. [On- line]. Available: https://www.jrailpass.com/blog/regions-of-japan
2025
-
[35]
The Classification and Division of Japanese Dialects,
S. Abe, “The Classification and Division of Japanese Dialects,” Ph.D. dissertation, Gakushuin University, 2015, in Japanese
2015
-
[36]
tohoku-nlp/bert-base-japanese,
tohoku-nlp, “tohoku-nlp/bert-base-japanese,” accessed: Feb. 6,
-
[38]
UTMOS: UTokyo-Sarulab System for V oicemos Challenge 2022,
T. Saeki, D. Xin, W. Nakata, T. Koriyama, S. Takamichi, and H. Saruwatari, “UTMOS: UTokyo-Sarulab System for V oicemos Challenge 2022,” in Proc. ISCA Interspeech , 2022, pp. 4521– 4525
2022
-
[39]
The V oiceMOS Challenge 2022,
W.-C. Huang, E. Cooper, Y . Tsao, H.-M. Wang, T. Toda, and J. Ya- magishi, “The V oiceMOS Challenge 2022,” inProc. ISCA Inter- speech, 2022, pp. 4536–4540
2022
-
[40]
X-Vectors: Robust DNN Embeddings for Speaker Recogni- tion,
D. Snyder, D. Garcia-Romero, G. Sell, D. Povey, and S. Khudan- pur, “X-Vectors: Robust DNN Embeddings for Speaker Recogni- tion,” in Proc. IEEE ICASSP, 2018, pp. 5329–5333
2018
-
[41]
ECAPA- TDNN: Emphasized Channel Attention, Propagation and Aggre- gation in TDNN Based Speaker Verification,
B. Desplanques, J. Thienpondt, and K. Demuynck, “ECAPA- TDNN: Emphasized Channel Attention, Propagation and Aggre- gation in TDNN Based Speaker Verification,” in Proc. ISCA In- terspeech, 2020, pp. 3830–3834
2020
-
[42]
sarulab-speech/xvector jtubespeech,
S. Takamichi, “sarulab-speech/xvector jtubespeech,” accessed: Feb. 6, 2025. [Online]. Available: https://github.com/ sarulab-speech/xvector jtubespeech
2025
-
[43]
washi/speaker-emb-ja-ecapa-tdnn,
k-washi, “washi/speaker-emb-ja-ecapa-tdnn,” accessed: Feb. 6, 2025. [Online]. Available: https://github.com/k-washi/ speaker-emb-ja-ecapa-tdnn
2025
-
[2025]
Available: https://huggingface.co/tohoku-nlp/ bert-base-japanese
[Online]. Available: https://huggingface.co/tohoku-nlp/ bert-base-japanese
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.