{"id":"b0245581-bd29-4d04-b166-3bfb88989ec4","arxiv_id":"2506.18296","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"JIS is a new 169-speaker Japanese speech corpus of live idols, built to support listener-familiarity-based evaluation of TTS and VC speaker similarity.","lead":"A new Japanese speech corpus records 169 young female live idols, identified by stage name, across studio and quiet-room settings for text-to-speech and voice conversion research. The corpus aims to make speaker-similarity tests more rigorous by letting researchers recruit listeners who already know these performers.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The corpus's core evaluation advantage depends on fans recognizing studio/read recordings of idols by voice, but no recognition test is reported.","rationale":"The paper's value proposition rests on the claim that JIS enables more rigorous speaker-similarity evaluation through familiarity: because speakers are identified by stage names, researchers can recruit listeners who know these idols, and those listeners will make finer distinctions. That argument has two parts: (1) fans are familiar with the idols' voices, and (2) this familiarity transfers to the specific recordings in the corpus. The paper provides cultural background for part (1) but never tests part (2). The recordings are read sentences, scripted spontaneous phrases, and studio speech produced under a persona instruction (Sec. 3.2), which may be acoustically quite different from live performances and social media clips. The cited familiarity literature (Schmidt-Nielsen & Stern 1985) concerns recognition of known voices in general, not transfer across speaking styles and recording conditions for these particular voices. If listeners cannot identify the speakers from JIS clips, the stage names are inert: the corpus becomes an ordinary multi-speaker dataset with no evaluation advantage beyond what anonymized corpora provide. This is exactly the reader's weakest assumption, and I agree it is the load-bearing one. Other concerns are secondary: the pseudo-MOS analysis in Sec. 4.2.1 selects the normalization constant that maximizes the reported score, which undermines any quantitative quality comparison; the 'first non-anonymous' claim is under-scoped relative to VoxCeleb; and the data are not yet released. But the familiarity-transfer question is the one on which the scientific contribution of the corpus hinges. The corpus construction itself appears solid, with 169 speakers, two recording tiers, varied speech styles, and supplementary metadata, so the appropriate verdict remains CONDITIONAL: conditional on a positive familiarity-recognition result and on actual release.","tokens_in":8406,"tokens_out":2352,"duration_ms":24829,"concrete_test":"Recruit self-identified fans of two idol groups represented in JIS, and a control group unfamiliar with JIS. Play short clips (2–3 s) of read, spontaneous, and singing utterances from 10 speakers per group, and ask participants to select which stage name produced each clip from a list of the group's members. Compare accuracy against chance (1/N) and against the control group. The premise is supported only if familiar listeners score significantly above chance and above controls; near-chance performance would invalidate the central evaluation advantage.","verdict_should_be":"UNCHANGED","load_bearing_attack":"JIS's central claim is that stage names let researchers recruit fans familiar with the idols, enabling more rigorous speaker-similarity evaluation (Abstract, Sec. 3.1). For this to hold, fans must recognize the corpus voices as belonging to the named idols. The corpus contains read phoneme-balanced sentences, scripted spontaneous phrases, greetings, and singing (Sec. 3.2), recorded in studios or quiet rooms under an instruction to sound like the idol's persona. No listening experiment verifies that this transfers from live performances, photo sessions, or social media to these recordings. The cited evidence ([12], [13]) shows familiar voices are recognized better in general, but it does not establish that familiarity with an idol's public persona makes their read studio speech identifiable. Without such transfer, JIS is not functionally non-anonymous for listeners: stage names provide no evaluation advantage, and the 'first non-anonymous' claim reduces to a metadata difference. The paper even hedges in Sec. 3.1 ('may have the possibility'), yet the abstract asserts 'will facilitate more rigorous evaluations.'","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces JIS, a freely distributed Japanese speech corpus of 169 young female live-idol speakers from 33 groups, recorded in two tiers: Speech A (studio, 61 speakers) and Speech B (quiet rooms, 108 speakers), with a total of 17.0 hours after VAD. The recordings include phoneme-balanced read sentences from the voiceactress100 text set, scripted spontaneous phrases, everyday greetings, and a royalty-free song, stored as 48 kHz/24-bit WAV. The authors argue that because speakers are identified by stage names rather than anonymized, researchers can recruit listeners familiar with these idols, enabling more rigorous speaker-similarity evaluation for TTS and VC than is possible with conventional anonymized corpora. The paper also provides an overview of Japanese live-idol culture, reports supplementary metadata (group, stage name, prefecture, questionnaire responses), and gives basic analyses: region histograms, BERT-based voice-impression embeddings from peer descriptions, UTMOS pseudo-MOS, and x-vector/ECAPA-TDNN speaker embeddings of JIS versus JVS.","tokens_in":8728,"tokens_out":3243,"duration_ms":36204,"significance":"If the corpus works as advertised, it is a valuable community resource: it is one of the first non-anonymous multi-speaker Japanese corpora for speech-generation research, it is free for non-commercial research, and it combines studio-quality and budget recordings with rich metadata that could support new research on familiarity-aware evaluation and listener-preference-driven voice generation. The construction details are concrete and reproducible (text sets, recording instructions, VAD, sampling/quantization), and the authors credit existing corpora and tools such as voiceactress100, JVS, UTMOS, x-vector, and ECAPA-TDNN. The main intellectual contribution is the hypothesis that familiar listeners will evaluate speaker similarity more discriminatorily; however, this hypothesis is not tested in the paper.","major_comments":[{"comment":"The central usefulness claim is that stage names will let researchers recruit listeners familiar with JIS speakers, yielding 'more rigorous evaluations' of speaker similarity. This requires that fans can actually identify the recorded corpus voices as belonging to the named idols, but no listening experiment verifies that familiarity acquired from live performances, photo sessions, and social media transfers to read sentences, scripted spontaneous phrases, and studio speech recorded under an instruction to sound like the idol's persona (Section 3.2). The cited works [12, 13] support the general point that familiar voices are recognized differently and more accurately, but they do not establish the transfer step that is specific to JIS. The paper itself hedges in Section 3.1 ('may have the possibility to recruit listeners') while the abstract asserts the corpus 'will facilitate more rigorous evaluations.' This is load-bearing: without the transfer step, JIS is functionally anonymized for listeners and the 'first non-anonymous corpus' claim reduces to a metadata difference. I recommend adding a recognition experiment (e.g., familiar listeners match held-out JIS samples to stage names) or rewriting the abstract and Section 3.1 to frame familiarity-based evaluation as an untested potential benefit that JIS enables.","section":"Abstract, 3.1"},{"comment":"The pseudo-MOS analysis selects the normalization constant that maximizes the average pseudo-MOS ('We tested various normalization constants, choosing the one with the highest average pseudo MOS'). This procedure makes the reported mean values (Speech A 3.4, Speech B 2.8, JVS 3.7) non-reproducible and systematically optimistic, and it invalidates the intended comparison between recording environments and against JVS, since the constant is fit to the JIS data being compared. Use a pre-specified normalization rule (e.g., fixed target RMS or peak level) or report results for all tested constants. This issue is specific and affects the quantitative analysis that is offered as guidance for using the corpus.","section":"Section 4.2.1, Fig. 3"}],"minor_comments":[{"comment":"The conclusion that 'overlapping yet offset regions' indicate differences in voice impressions between JIS speakers is based only on visual inspection of a t-SNE plot of BERT embeddings; a quantitative separability measure (e.g., silhouette score or classification accuracy) would make the claim more robust.","section":"4.1.2"},{"comment":"The 'Recording environment' row is formatted ambiguously for JIS; the reader must infer that A refers to the studio condition and B to the quiet-room condition. Please spell out 'Speech A: Studio; Speech B: Quiet but unspecified rooms.'","section":"Table 1"},{"comment":"There is some tension between saying stage names allow fans to identify speakers while 'protecting the speakers' privacy' and the later statement that users must avoid harming idol reputations. Since stage names are public persona identities, the corpus is not anonymized; the privacy framing should be clarified.","section":"3.1, 3.3"},{"comment":"The claim that the instruction to produce persona-consistent speech was 'somewhat effective' (Figure 5) is supported only by visual cluster inspection of t-SNE embeddings; please report a quantitative metric such as within-speaker versus between-speaker cosine similarity.","section":"4.2.2"},{"comment":"Minor typographical and formatting issues: 'V AD' should be 'VAD', and the reference URLs are not consistently formatted. These do not affect the substance.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"This is a solid and useful resource paper whose main risk is over-claiming the evaluation advantage of non-anonymous stage names without a recognition experiment. The corpus construction itself is clear and well documented. I would encourage the editor to ask for either a small listening/recognition validation or a calibrated rewrite of the central claim, and for a fix of the pseudo-MOS normalization procedure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"JIS is a real contribution: a 169-speaker Japanese corpus of live idols with studio and quiet-room recordings, covering read, spontaneous, greeting, and sung speech, plus stage names and metadata. Construction is described concretely, the copyright and ethics discussion is serious, and free distribution is a plus. The novelty is genuine—nothing else targets this speaker category or explicitly aims at familiarity-based speaker similarity evaluation.\n\nThe analyses are modest and mostly well-hedged: embedding plots show JIS clustering apart from JVS, speaking styles cluster by speaker, and the questionnaire data give a glimpse of perceived voice impressions. The paper is also honest that the corpus is being released only after submission.\n\nThe soft spot is the load-bearing argument. The abstract says JIS 'will facilitate more rigorous evaluations' because fans can be recruited, but no listening test shows that fans can recognize these idols from read phoneme-balanced sentences or scripted spontaneous speech recorded in a studio. Familiarity with a public persona doesn't automatically transfer to a different speaking style and recording context. The authors themselves hedge in Sec. 3.1 ('may have the possibility'), so the abstract overstates. This is fixable with a simple recognition test, but until then the central claim is speculative.\n\nTwo smaller issues: the pseudo-MOS analysis picks the normalization constant that maximizes the average score, which is fitting a descriptive statistic and weakens the Speech A/B comparison, though it isn't a headline result. And the 'first non-anonymous' claim should at least acknowledge VoxCeleb, which is non-anonymous, even if geared to speaker recognition rather than speech generation.\n\nThis paper is for researchers in Japanese TTS/VC who need a multi-speaker corpus with a narrow speaker category and public identities, and for anyone working on listener-preference voice generation. The corpus deserves a serious referee. I'd accept it for review with the familiarity claim softened and a recognition experiment recommended as future work—and once the data are actually out, I'd cite it.","headline":"A genuinely new Japanese multi-speaker corpus with real potential for familiarity-based evaluation, but the central claim that familiarity transfers to read studio speech is unvalidated.","tokens_in":9113,"tokens_out":2682,"would_cite":true,"duration_ms":24861,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper constructs JIS, a 169-speaker Japanese live-idol corpus meant to tighten speaker-similarity evaluation in TTS and VC.","keywords":["speech corpus","Japanese live idols","text-to-speech synthesis","voice conversion","speaker similarity","speaking styles","voice preference","non-anonymous dataset"],"falsifier":"A recognition test in which self-identified fans of specific JIS idols listen to short clips from Speech A and Speech B mixed with voices of same-age women and try to name the idol; if identification accuracy is near chance even for idols the listener claims to support, the stage-name advantage collapses and JIS becomes an ordinary anonymous corpus.","tokens_in":8219,"feed_emoji":"🎤","tokens_out":8141,"duration_ms":77763,"temperature":0.7,"pith_summary":"This paper reports the construction of JIS, a speech corpus containing 169 young female Japanese live idols from 33 groups, to be distributed free for non-commercial basic research in text-to-speech and voice conversion. The central claim is that because every speaker is identified by a stage name and belongs to one highly specific category, researchers can recruit listeners who already know these voices, leading to stricter speaker-similarity evaluations than anonymous multi-speaker corpora allow. The corpus includes studio recordings from 61 speakers and quiet-room recordings from 108 speakers, covering phoneme-balanced read sentences, spontaneous greetings, personality phrases, and a short sung piece. The paper also supplies an overview of live idol culture, questionnaire-based voice-impression descriptions, and basic acoustic analyses to guide use of the data.","feed_headline":"169 idol voices, one corpus, stricter similarity tests","feed_subtitle":"Live-idol stage names let researchers recruit fans who know the speakers, making familiar-voice evaluation possible.","key_machinery":"The load-bearing object is the corpus itself, specifically its pairing of voice recordings with stage names and group identities. That pairing converts an ordinary multi-speaker corpus into a familiarity-enabled evaluation resource: the same brain processes that make familiar-voice recognition more accurate than unfamiliar-voice recognition can be recruited in listening tests. Around this pairing, the corpus offers style-specific utterances (an energetic post-performance farewell, an intimate photo-session greeting, a self-introduction), phoneme-balanced read sentences, a questionnaire in which each speaker describes her own and her groupmates' voice impressions, and embedding analyses showing that speakers remain distinguishable across styles.","core_discovery":"JIS is, to the authors' knowledge, the first non-anonymous multi-speaker speech corpus built for speech generation AI research. All 169 speakers are young female Japanese live idols, and each is tied to a publicly used stage name and group affiliation, so an experiment planner can recruit fans who are familiar with the actual people behind the voices. The corpus is split into Speech A (61 speakers, professional studios, the full 100-sentence phoneme-balanced set plus spontaneous speech, greetings, and singing) and Speech B (108 speakers, quiet rooms, a partial set of the same content). Embedded speaker vectors cluster separately from those of an existing Japanese corpus and, within JIS, tend to form per-speaker clusters across speaking styles, which the authors read as evidence that the persona-matching recording instruction was at least partly effective.","pith_inferences":["If familiarity transfers from live performances and social media to these studio recordings, JIS could serve as a benchmark in which failing to recognize an idol's voice is an unambiguous system error; no anonymous corpus offers that oracle.","The voice-impression questionnaire could be used to learn a mapping from impression phrases to acoustic speaker embeddings, enabling text-described voice generation aimed at individual listener preferences, a direction the paper mentions but does not implement.","Because each speaker appears in several styles, JIS also enables within-speaker style disentanglement experiments without collecting new data.","A direct check of the corpus's core premise would be a fan-recognition test: if accuracy is high only for performance-style phrases and not for read sentences, evaluation protocols should weight the speaking styles that carry recognizability."],"forward_implications":["TTS and VC systems can be evaluated by listeners who actually know the target speaker, so similarity scores reflect recognition of a real individual rather than matching of coarse attributes like age and gender.","Training and evaluation can be confined to a homogeneous speaker category, removing attribute mismatch as an accidental driver of high similarity ratings.","The style-conditioned recordings support research on separately controlling who is speaking and how they are speaking.","The questionnaire descriptions of voice impressions provide labels for work on listener-preference voice generation.","The documentation of recording instructions and rights transfer can guide other teams that want to collect additional live-idol speech under similar ethical terms."],"supporting_citations":[{"why":"Supplies evidence that voice discrimination and voice recognition are separate abilities, which motivates familiarity-based similarity testing.","marker":"[12]"},{"why":"Shows listeners identify familiar voices more accurately, the premise for recruiting fans as evaluators.","marker":"[13]"},{"why":"Describes JVS, the Japanese multi-speaker corpus against which JIS is compared and whose parallel100 texts are reused.","marker":"[11]"},{"why":"Provides the phoneme-balanced read sentences recorded by JIS speakers.","marker":"[24]"},{"why":"Provides background on Japanese live idol culture and photo-session interaction, motivating the spontaneous speech content and ethical guidance.","marker":"[14]"},{"why":"Defines oshi fan attachment and supports the claim that fans focus on individual idols rather than groups, justifying stage-name-linked evaluation.","marker":"[21]"}],"fun_headline_variants":["First non-anonymous speech corpus: 169 Japanese idols","Idol speech corpus lets fans judge voice similarity","169 idol voices with known identities for speech AI","JIS: non-anonymous voices for tougher speaker tests"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that fans' familiarity with an idol from live performances, photo sessions, and social media transfers to recognizing her voice in JIS's studio recordings of read sentences and scripted phrases; the paper does not test this transfer with any listening experiment.","fun_headline_variants_meta":{"raw":{"variants":["First non-anonymous speech corpus: 169 Japanese idols","Idol speech corpus lets fans judge voice similarity","169 idol voices with known identities for speech AI","JIS: non-anonymous voices for tougher speaker tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000198,"raw_usage":{"total_tokens":1329,"prompt_tokens":867,"completion_tokens":462,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":483,"completion_tokens_details":{"reasoning_tokens":399}},"tokens_in":483,"tokens_out":462,"duration_ms":4826,"temperature":1.0,"reasoning_tokens":399,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T18:51:33.503383+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A recognition test in which self-identified fans of specific JIS idols listen to short clips from Speech A and Speech B mixed with voices of same-age women and try to name the idol; if identification accuracy is near chance even for idols the listener claims to support, the stage-name advantage collapses and JIS becomes an ordinary anonymous corpus.","supporting_citations":[{"cited_title":"Diffusion-Based V oice Conversion with Fast Maximum Likelihood Sampling Scheme,","cited_arxiv_id":null,"evidence_quote":"Supplies evidence that voice discrimination and voice recognition are separate abilities, which motivates familiarity-based similarity testing."},{"cited_title":"V oice- Grad: Non-Parallel Any-to-Many V oice Conversion With An- nealed Langevin Dynamics,","cited_arxiv_id":null,"evidence_quote":"Shows listeners identify familiar voices more accurately, the premise for recruiting fans as evaluators."},{"cited_title":"AutoVC: Zero-Shot V oice Style Transfer with Only Au- toencoder Loss,","cited_arxiv_id":null,"evidence_quote":"Describes JVS, the Japanese multi-speaker corpus against which JIS is compared and whose parallel100 texts are reused."},{"cited_title":"Generat- ing Speakers by Prompting Listener Impressions for Pre-trained Multi-Speaker Text-to-Speech Systems,","cited_arxiv_id":null,"evidence_quote":"Provides the phoneme-balanced read sentences recorded by JIS speakers."},{"cited_title":"The CMU Arctic Speech Databases,","cited_arxiv_id":null,"evidence_quote":"Provides background on Japanese live idol culture and photo-session interaction, motivating the spontaneous speech content and ethical guidance."},{"cited_title":"V oice Activity Detection in the Wild: A Data-Driven Approach Using Teacher- Student Training,","cited_arxiv_id":null,"evidence_quote":"Defines oshi fan attachment and supports the claim that fans focus on individual idols rather than groups, justifying stage-name-linked evaluation."}],"review_version":2}