{"id":"2a02b3c1-813b-4493-9597-bec8c5797af0","arxiv_id":"2507.01349","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A newly built, freely distributable 15-song corpus of Japanese idol-style music with stems, dry vocals, and chord annotations for benchmarking music information processing.","lead":"The authors commissioned professional composers and singers to create 15 Japanese idol-style songs, recording clean stems, dry vocals, and chord annotations alongside mastered mixes. The resulting open corpus lets music software researchers test source separation, chord detection, and lyric transcription on realistic, high-loudness, multi-singer material.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"License clause forbidding ML training on instrumental stems conflicts with the corpus's stated MSS training purpose and needs clarification.","rationale":"The corpus construction itself is credible: the paper gives a detailed production pipeline, includes stems, dry vocals, solo versions, and expert chord annotations, and is transparent about the nonlinearity of mastering effects. The reader's weakest_assumption—that commissioned songs genuinely resemble real idol-group songs—is a real but softer concern: Section 3 offers only a UMAP plot with no statistical test, and the real-world comparison set contains only female-group tracks. Even if style match were imperfect, the corpus could still be useful as a controlled benchmark. The license contradiction is more decisive because it bears directly on whether the corpus can be used for its stated purpose of training and evaluating source separation. As written, the license simultaneously invites machine-learning use and prohibits training on the instrumental stems, making the distribution terms self-contradictory. This is an internal inconsistency, not a disagreement with external consensus. It can be fixed by clarifying the license text, so the appropriate disposition is the same CONDITIONAL verdict the reader reached: the paper should be accepted once the authors clarify the scope of the training prohibition and reconcile it with the MSS use case.","tokens_in":10741,"tokens_out":6092,"duration_ms":74238,"concrete_test":"Read the license file in the HuggingFace repository (https://huggingface.co/datasets/imprt/idol-songs-jp) and check whether the clause 'sampling the instrumental tracks ... to train machine learning models is prohibited' is present and unqualified. If it is, attempt a minimal supervised training run of a source-separation model on the four-stem subset (e.g., one song's stems) and verify whether the license permits it. If the license forbids training, the paper's claim that the corpus supports MSS training must be revised to 'evaluation-only', or the license must be amended to permit training while still prohibiting unauthorized sample-reuse of the instrument sounds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing problem is an internal contradiction in the corpus license, stated in Section 2's final paragraph. The text says 'Users may rearrange, parody, and apply machine learning techniques to the corpus,' but immediately adds: 'To protect the rights associated with software instruments and commercial sample libraries, sampling the instrumental tracks to create unrelated content or to train machine learning models is prohibited.' Section 1 motivates the corpus by noting that supervised MSS 'requires a corpus consisting of ground-truth stems,' and the corpus is presented as a resource for MSS. If the instrumental stems (drums, bass, other, and processed vocal stems) cannot be used to train machine-learning models, then a primary advertised use—training source-separation systems on the four-stem data—is blocked for the very files that make the corpus novel. This is not a style-realism question; it is a concrete licensing restriction that determines whether the benchmark can actually be used as advertised. The paper's own MSS experiments only evaluate a pretrained model and do not train on IdolSongsJp, so the conflict is not exercised by the authors' evaluation.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces IdolSongsJp, a corpus of 15 commissioned songs in the style of Japanese idol groups, produced by professional creators. The corpus provides mastered audio at −7 and −9 LUFS, instrument stems, four-stem versions for music source separation, 414 dry vocal tracks, 95 solo versions, off-vocal/minus-one tracks, and expert chord annotations. The authors report a qualitative embedding-space comparison with 4,483 real idol-group tracks and present benchmark evaluations for music source separation, automatic chord estimation, and automatic lyrics transcription. The central claim is that the corpus is a realistic, distributable resource for tasks such as source separation, singer diarization, and chord estimation under commercial-like loudness and arrangement conditions.","tokens_in":10839,"tokens_out":4365,"duration_ms":49701,"significance":"If the licensing contradictions are resolved, IdolSongsJp could be a valuable addition to MIR resources: it is, to my knowledge, a rare corpus with exact stems, dry vocal tracks, and expert chord annotations for multi-singer songs at commercial loudness, and the construction details (loudness targets, mixing paths, consensus annotations) are concrete and checkable. The authors are also transparent about the limitations of their MSS evaluation (nonlinear mastering effects) and about the fact that their system evaluations use pretrained models rather than training on the corpus. The main weaknesses are that the advertised usability for MSS training is currently undercut by the license text, and the realism/diversity evidence is only qualitative.","major_comments":[{"comment":"There is an internal contradiction in the license. The text states that 'Users may rearrange, parody, and apply machine learning techniques to the corpus,' but immediately adds that 'sampling the instrumental tracks to create unrelated content or to train machine learning models is prohibited.' Because Section 1 motivates the corpus by noting that supervised MSS 'requires a corpus consisting of ground-truth stems,' and Section 2 lists 'Stems for MSS' as a core data type, this prohibition removes the central advertised use of the instrumental stems. The phrase 'apply machine learning techniques' is unqualified and appears to allow what the next sentence forbids. Please clarify which data types (vocal vs instrumental; stems vs mastered tracks) may be used for training, and align the Hugging Face license with the text; the current wording makes the resource unusable for the main task it is designed to support.","section":"Section 2, final paragraph; Section 9"},{"comment":"The claim that the corpus songs are 'broadly distributed across the embedding space' and include both distinctive and typical idol-style songs is based solely on visual inspection of a UMAP plot. The comparison against 4,483 real-world tracks lacks any quantitative support, such as a coverage statistic, nearest-neighbor distances between corpus and real tracks, or a statistical test of distribution overlap. Since the conclusion that the corpus is a 'realistic resource' for real-world conditions depends in part on this comparison, please add a quantitative diversity/coverage analysis or explicitly weaken the claim to avoid overstating what Figure 2 demonstrates.","section":"Section 3, Figure 2"}],"minor_comments":[{"comment":"In the MSS evaluation, the reference stems for the mastered conditions are produced by applying nonlinear mastering effects to individual stems, and the paper acknowledges that simply summing the processed stems does not reproduce the final mastered tracks. This means the SDR values in Figure 3 are computed against approximate references whose relation to the true sources is not quantified; please state whether the approximation affects all stems equally or report the reconstruction error of the mastered mix from the processed stems.","section":"Section 4, conditions 2 and 3"},{"comment":"The statement that the proportion of major and minor chords is 'significantly lower' than in the McGill Billboard corpus would benefit from a statistical test or at least a confidence interval; as written, the comparison is qualitative and the denominator (after excluding non-chorded sections and tensions) is not fully specified.","section":"Section 5, Figure 4"},{"comment":"The terms 'calls and mixes' and 'utawari' are used without a brief definition for non-Japanese readers; a one-sentence explanation in Section 2 would improve accessibility.","section":"Section 2, bullet list"},{"comment":"The first-page footnote states that the paper is 'Licensed under a Creative Commons Attribution 4.0 International License (CC BY 4.0),' while Section 2 says the corpus is 'available free of charge for non-commercial research and entertainment purposes.' Please clarify which license applies to the paper and which to the corpus, because CC BY 4.0 permits commercial use and is at odds with the corpus's non-commercial restriction.","section":"First-page footnote vs Section 2"},{"comment":"The stem counts in Table 1 (ranging from 8 to 17) refer to the instrument stems, whereas the MSS data use four categories; please make this distinction explicit in the table caption or in Section 2 to avoid confusion.","section":"Table 1, column 'No. of stems'"}],"recommendation":"major_revision","confidential_remarks":"The license contradiction in Section 2 is the main blocker: it directly undermines the corpus's stated purpose for MSS training. The diversity comparison in Section 3 also needs quantitative grounding before the 'realistic resource' claim can be accepted. Both are fixable within the manuscript's scope, so major revision seems appropriate rather than rejection. The corpus itself appears genuinely useful if the licensing is clarified."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, the corpus is real and fills a concrete gap: commissioned Japanese idol-style songs with four-stem and finer stems, dry vocal tracks, utawari division, expert chord annotations, and multiple loudness masters, all publicly distributable for non-commercial use. The construction details are concrete and unusually honest. Second, there is a licensing contradiction the paper never resolves, and it blocks the corpus's primary advertised purpose. Section 2 says users may \"apply machine learning techniques to the corpus,\" then immediately prohibits \"sampling the instrumental tracks to create unrelated content or to train machine learning models.\" Since Section 1 motivates the corpus by the need for ground-truth stems to train source separation, the instrumental stems—drums, bass, other, and the processed vocal stems—cannot be used to train the very models the corpus is meant to support. The authors' own MSS experiments only evaluate a pretrained model, so the conflict is never exercised. This is not a style-realism quibble; it is a concrete restriction that determines whether the benchmark can be used as advertised. It needs clarification: either training on the stems is allowed under the non-commercial umbrella, or the paper should stop presenting MSS training as a use case.\n\nWhat is genuinely new is the combination: commissioned tracks with exact stems, dry vocals, utawari annotations, chord annotations, and multiple loudness masters. The corpus-building practice is sound—loudness targets are specified, stems are generated via explicit mixing paths, chord annotations require consensus of at least two experts, and the evaluation limitations are disclosed, including the nonlinearity of mastering effects. That transparency earns credit.\n\nThe soft spots are in the supporting evidence, not the construction. The style-diversity check is the weakest empirical section: a UMAP plot of 15 tracks against 4,483 real tracks, with no statistical test, and the real-world set is female groups only while the corpus includes male-singer tracks. The evaluation results have no error bars, and the MUSDB18 comparison leans on prior reported scores. These are minor-to-moderate issues; they don't undermine the corpus itself.\n\nBottom line: this deserves a serious referee, but the license clause must be resolved before release. As is, I wouldn't cite it as a training resource; once clarified, it becomes a useful evaluation benchmark for separation, diarization, and chord estimation under high-loudness, mastered conditions. Send it to peer review with a request to fix the license and tighten the diversity assessment.","headline":"A genuinely useful corpus that ships with a license clause contradicting its own advertised MSS training use—worth refereeing, but the license needs to be fixed before anyone can build on it.","tokens_in":11442,"tokens_out":2053,"would_cite":false,"duration_ms":23499,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"IdolSongsJp is a new, freely distributable corpus of 15 commissioned idol-style songs with exact stems, dry vocals, and expert chord annotations.","keywords":["Japanese idol group music","music source separation","chord estimation","lyrics transcription","singer diarization","multi-singer corpus","utawari","loudness"],"falsifier":"A quantitative distributional comparison between the 15 corpus tracks and the 4,483 real idol tracks (e.g., a two-sample test on CLAP embeddings or on loudness and instrument-count statistics) that shows the corpus tracks come from a different distribution, or a listener study in which participants do not identify the commissioned tracks as stylistically similar to idol songs, would falsify the realistic-resource claim.","tokens_in":10504,"feed_emoji":"🎤","tokens_out":8282,"duration_ms":83691,"temperature":0.7,"pith_summary":"The paper introduces IdolSongsJp, a new corpus of 15 multi-singer songs commissioned from professional composers to resemble Japanese idol group music. Each song ships with mastered audio at two commercial loudness levels, four-stem source-separation signals, 414 dry vocal tracks, 95 solo versions, and expert chord annotations, making it freely distributable for non-commercial research. The authors argue the corpus closes a gap: existing separation corpora are quieter and less dense than contemporary commercial releases, and many research corpora cannot be shared online without per-user permission. They show the corpus spans a wide stylistic range by comparing it against 4,483 real idol songs, and they demonstrate its use by benchmarking music source separation, chord estimation, and lyrics transcription systems.","feed_headline":"Idol-style song corpus offers clean ground truth for music AI","feed_subtitle":"Professional commissions with commercial loudness and full ground truth stress-test separation and chord systems.","key_machinery":"The load-bearing object is the corpus's multi-layer production design: every song exists as individual instrument stems, linear stem sums, stems processed with mastering effects, and fully mastered versions at two loudness targets, so that ground-truth audio can be compared against realistic commercial-style mixes. The utawari structure, in which singers alternate solo lines and sing together in sections, is the organizing musical feature that makes the corpus suited to multi-singer tasks such as singer diarization. Expert chord annotations in Harte's shorthand supply symbolic ground truth, and the 414 dry vocal tracks provide clean sources for singing voice synthesis and vocal processing.","core_discovery":"The central claim is that IdolSongsJp provides a realistic, distributable benchmark for music information processing on multi-singer idol-style songs. The 15 songs were created by professional or semi-professional musicians, with each song assigned a unique combination of singers among 10 female and 8 male vocalists, designed with utawari song-division structures, varied genres and lyrical themes, and mastered to a loudness of -7 LUFS (with a -9 LUFS alternative). The corpus bundles linear stem sums, stems for four-category source separation, 414 individually recorded dry vocal tracks, unmastered master-bus signals, off-vocal and minus-one versions, 95 solo-song mixes, and chord annotations in Harte's shorthand agreed by at least two expert annotators. Application experiments show the corpus is demanding for current systems: HT Demucs separation accuracy drops on mastered tracks compared with unmastered sums, chord estimation exceeds 80 percent on roots and major/minor quality but falls below 60 percent on tetrads and below 30 percent on four-note MIREX chords, and lyrics transcription with Whisper is more robust than with a HuBERT-Conformer system. These results support the claim that the corpus can serve as a challenging resource for general and song-specific music information processing tasks.","pith_inferences":["The paired -7 and -9 LUFS masters could be used as a controlled experiment to isolate how limiter gain alone changes separation and chord estimation performance, something the paper does not investigate.","Adding a quantitative distributional test to the UMAP visualization would strengthen the claim that the corpus matches real idol-song style; the current evidence is purely qualitative.","The license structure, which permits research use but forbids sampling instrumental stems to train generative models, offers a model for sharing music data while protecting sample-library rights.","The unusually poor separation result on the UK Garage track suggests genre-specific bass and sound design may need specialized models, pointing toward genre-conditioned source separation."],"forward_implications":["HT Demucs source separation performs noticeably worse on the mastered -7 LUFS tracks than on unmastered stem sums, especially for drums and vocals, showing that mastering effects remain a real obstacle for separation systems.","Chord estimation systems exceed 80 percent accuracy for roots and major/minor chords but stay below 60 percent for tetrads and below 30 percent for MIREX4 chords, so extended chord vocabulary is still an open problem.","Whisper-based lyrics transcription keeps similar character error rates across mastered tracks, Demucs-separated vocals, and dry lead vocals, while a HuBERT-Conformer system degrades on mastered audio and produces different error patterns.","Because the corpus includes utawari structures and dry vocals for each singer, it is directly reusable for singer diarization, multi-pitch detection, and singing voice synthesis, beyond the three tasks the paper evaluates."],"supporting_citations":[{"why":"Supplies the standard 4-stem source separation corpus whose performance level HT Demucs matches on unmastered sums.","marker":"[14]"},{"why":"Provides the instrument category definitions on which the corpus's stems are based.","marker":"[16]"},{"why":"Documents that existing separation corpora have lower loudness than commercial tracks, motivating the high-loudness design.","marker":"[18]"},{"why":"Exemplifies existing multi-singer corpora that the new corpus complements for diarization research.","marker":"[20]"},{"why":"Prior real-world idol song corpus by the same group, giving context for the commissioned-style approach.","marker":"[24]"},{"why":"A distributed research corpus that requires per-user permission, motivating the need for a freely distributable resource.","marker":"[27]"},{"why":"CLAP embeddings used to compare corpus songs to real idol songs in the diversity analysis.","marker":"[31]"},{"why":"One of the two chord estimation methods evaluated on the corpus.","marker":"[35]"}],"fun_headline_variants":["Idol songs get a benchmark for AI music tasks","Clean idol-style tracks stress-test music AI systems","Multi-singer idol corpus with full annotations for AI","Challenging idol-song dataset for music separation and chords","New idol-style corpus pushes music AI limits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the commissioned songs genuinely resemble real Japanese idol songs in the properties that matter — loudness, arrangement density, utawari structure, and stylistic coverage — which the paper supports only with a qualitative UMAP visualization and no statistical test.","fun_headline_variants_meta":{"raw":{"variants":["Idol songs get a benchmark for AI music tasks","Clean idol-style tracks stress-test music AI systems","Multi-singer idol corpus with full annotations for AI","Challenging idol-song dataset for music separation and chords","New idol-style corpus pushes music AI limits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000698,"raw_usage":{"total_tokens":3193,"prompt_tokens":1022,"completion_tokens":2171,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":638,"completion_tokens_details":{"reasoning_tokens":2111}},"tokens_in":638,"tokens_out":2171,"duration_ms":16040,"temperature":1.0,"reasoning_tokens":2111,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:54:02.631671+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A quantitative distributional comparison between the 15 corpus tracks and the 4,483 real idol tracks (e.g., a two-sample test on CLAP embeddings or on loudness and instrument-count statistics) that shows the corpus tracks come from a different distribution, or a listener study in which participants do not identify the commissioned tracks as stylistically similar to idol songs, would falsify the realistic-resource claim.","supporting_citations":[{"cited_title":"Deep learning for audio signal pro- cessing,","cited_arxiv_id":null,"evidence_quote":"Supplies the standard 4-stem source separation corpus whose performance level HT Demucs matches on unmastered sums."},{"cited_title":"Efficient tempo and beat tracking in audio recordings,","cited_arxiv_id":null,"evidence_quote":"Provides the instrument category definitions on which the corpus's stems are based."},{"cited_title":"Melody extraction from polyphonic music signals: Approaches, applications, and challenges,","cited_arxiv_id":null,"evidence_quote":"Documents that existing separation corpora have lower loudness than commercial tracks, motivating the high-loudness design."},{"cited_title":"Audio chord recognition with recurrent neural networks,","cited_arxiv_id":null,"evidence_quote":"Exemplifies existing multi-singer corpora that the new corpus complements for diarization research."},{"cited_title":"Automatic recognition of lyrics in singing,","cited_arxiv_id":null,"evidence_quote":"Prior real-world idol song corpus by the same group, giving context for the commissioned-style approach."},{"cited_title":"Deep learning based source separation ap- plied to choir ensembles,","cited_arxiv_id":null,"evidence_quote":"CLAP embeddings used to compare corpus songs to real idol songs in the diversity analysis."},{"cited_title":"Singer diarization for polyphonic music with unison singing,","cited_arxiv_id":null,"evidence_quote":"One of the two chord estimation methods evaluated on the corpus."}],"review_version":1}