{"id":"eabffc74-d2e6-494d-bc5f-b186347701b4","arxiv_id":"2507.01172","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"GuitarDuets provides about three hours of real and synthetic classical guitar duet audio, and experiments show that combining both data types improves Demucs-based separation of similar-timbre guitars.","lead":"This paper introduces GuitarDuets, a dataset of real and synthesized classical guitar duet recordings, and shows how a state-of-the-art separation model can separate two similar-timbre guitars. The dataset and experiments give researchers a new benchmark for monotimbral music source separation, a relatively underexplored task.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's blanket claim is contradicted by Table 3 for guitar 2: R+S vs R-only decreases G2 SDR and SAR, so the central claim needs qualification or statistical support.","rationale":"The reader's verdict is CONDITIONAL and the weakest assumption is that the seven-track bleeding-free test set is representative. I agree that test-set representativeness is a real external-validity concern, but the more immediate, internal problem is that the abstract's central claim is not supported by the paper's own Table 3 for the second guitar. The R+S vs R-only comparison shows a clear decrease in G2 SDR, SAR, and SIR, so the unqualified statement 'improved separation performance' is contradicted by the presented numbers. The reader's rationale did note that the headline claim is 'only partially supported by the tables,' which is close, but the weakest-assumption field points to test-set generalizability rather than this direct inconsistency. Because the dataset contribution remains credible and the issue is fixable by qualifying the claim and adding statistical robustness, the existing CONDITIONAL verdict should stand rather than escalate to REJECT. The proposed rerun with multiple seeds and per-track paired tests would settle whether the G1 improvement is real and whether G2 behaves as the table suggests.","tokens_in":11118,"tokens_out":4288,"duration_ms":47524,"concrete_test":"Rerun the two relevant training conditions (GuitarDuets(R)+GuitarDuets(S) and GuitarDuets(R) alone) with at least three random seeds and report per-track, per-guitar SDR and SI-SDR on the seven-track bleeding-free test set, with paired significance tests across tracks (e.g., Wilcoxon signed-rank). If G2 SDR is not consistently better under the combined condition, the headline claim must be narrowed to G1; if the G1 gain disappears within seed variance, the claim lacks statistical support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract states that using both the real and synthesized subsets of GuitarDuets leads to improved separation performance on the independent test set compared to using either subset alone. The paper's own Table 3 does not support this claim for the second guitar. For the direct comparison R+S vs R-only (rows 6 and 7), G2 SDR decreases from 1.014 to 0.920 dB and G2 SAR decreases from 1.424 to 0.896 dB; G2 SIR also decreases from 4.873 to 4.104 dB. Only G1 SDR improves (4.952 to 5.882 dB). Against S-only, G2 does improve (0.200 to 0.920 dB), but the unqualified phrase 'improved separation performance' is false for the R-only vs R+S comparison on G2. Since the task is duet separation, both sources matter, and the headline claim should be qualified to G1 or to an explicitly defined averaged metric. Moreover, all numbers are single runs without error bars or seed variance, so even the G1 gain of 0.93 dB may not be statistically reliable.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces GuitarDuets, a dataset of roughly 58.6 minutes of real and 106 minutes of synthesized classical guitar duet recordings, with MIDI note annotations for the synthetic portion. It adapts Hybrid Transformer Demucs to two-guitar separation and reports cross-dataset experiments on a bleeding-free held-out test set, comparing training on real-only, synthetic-only, and combined data. It also proposes score-informed conditioning with ground-truth labels and a two-stage transcription-separation pipeline with predicted labels, and analyzes SDR versus SI-SDR behavior on simulated monotimbral and multitimbral mixtures. The headline claim is that combining real and synthesized subsets improves separation on the independent test set compared with using either subset alone.","tokens_in":11344,"tokens_out":6049,"duration_ms":61624,"significance":"The main contribution is a publicly released dataset that fills a real gap: monotimbral source separation with polyphonic co-playing instruments. The design has several strengths: a dedicated bleeding-free test set, synthetic data with note-level annotations, use of an established strong baseline (Demucs), and a systematic cross-dataset training comparison. If the performance claims survive a more rigorous evaluation, the finding that synthetic data can complement a small real corpus in same-timbre separation would be practically useful, and the dataset could become a standard benchmark. The paper is honest about the difficulty of the task, noting the large SDR gap between the two guitars.","major_comments":[{"comment":"The abstract's claim that using both the real and synthesized subsets 'leads to improved separation performance' is not supported for the second guitar. In Table 3, comparing the R-only row with the R+S row, G2 SDR decreases from 1.014 to 0.920 dB, G2 SAR from 1.424 to 0.896 dB, and G2 SIR from 4.873 to 4.104 dB; G2 SI-SDR improves from -3.536 to -3.133 dB, while G1 improves on all four metrics (e.g., G1 SDR from 4.952 to 5.882 dB). Since a duet separation system must provide both sources, the headline claim should either be qualified to the first guitar or be based on a clearly defined aggregate metric; as written it is contradicted by part of the paper's own results.","section":"Abstract; §4.2, Table 3"},{"comment":"All separation results are reported as single numbers from a single training run, without seeds, error bars, or significance tests, and the test set consists of only 7 tracks. The key head-to-head gains (e.g., +0.93 dB G1 SDR for R+S versus R-only) are therefore not shown to be statistically reliable. Please report results over multiple random seeds, or at least per-track score distributions with paired tests such as Wilcoxon or bootstrap, for the central comparisons, and state the variance explicitly. This is necessary to support the comparison-based conclusions of the paper.","section":"§4.2, Tables 3 and 4"},{"comment":"The metric-behavior analysis is based on a single track (Track 29 of GuitarDuets(S)) and on synthetic additive mixtures m = αx1 + (1−α)x2, yet the text concludes that 'both metrics for the guitar mixtures are consistently higher' than for multitimbral mixtures. With no error bars or a range of tracks, 'consistently' is not established. Please extend this analysis to multiple tracks and report the variability; if the illustrative example is intended only as a motivating demonstration, say so explicitly and soften the claim.","section":"§4.4, Eq. (2), Figure 4"},{"comment":"The evaluation rests entirely on the 7-track bleeding-free test set. This is a well-motivated design, but the paper does not discuss whether those 7 tracks are representative of the real-recording conditions used in training (instruments, microphones, room, genre, tempo, or balance between guitars). A 7-track test set can support the paper's central claim only if the authors either provide evidence of representativeness or acknowledge the limitation and avoid generalizing beyond the test conditions. Please add a description of how the test tracks were selected and a discussion of this limitation.","section":"§2.2"}],"minor_comments":[{"comment":"The phrase 'have focused in the multi-timbral case' should be 'have focused on the multi-timbral case'; similar grammatical issues appear elsewhere in the introduction.","section":"Abstract"},{"comment":"There are minor typographical issues, including 'avalaible' in the introduction and the unusual spacing in 'W A V' in Section 2.2; these should be corrected.","section":"§1, §2.2"},{"comment":"The loss weights α and β are said to be set after preliminary experiments, but no sensitivity analysis or ablation is reported; please state the range explored or provide a brief justification.","section":"§3.1, Eq. (1)"},{"comment":"The header layout is confusing: 'Source Datasets' and 'Metrics' are on the same row, making it easy to misassign the checkmarks; use separate header rows or subheadings to clarify which datasets are included in each condition.","section":"Table 3"},{"comment":"The test sets differ between the GuitarDuets(S) and GuitarDuets(R) panels, so comparisons across the two panels should be explicitly flagged as not directly comparable.","section":"Table 4"},{"comment":"The training/validation split is described as 80-20, but it is not explicitly stated that the bleeding-free test set is excluded from training; please make this explicit.","section":"§4.2"},{"comment":"The qualitative claim that predicted note labels 'enable the separation model to more accurately sustain notes' is not supported by any quantitative measure; consider adding an objective analysis or softening the interpretation.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"To the editor: The dataset is a potentially valuable community resource and the paper is likely to be of interest to the MSS/MIR community. My main concern is not the dataset construction but the gap between the abstract's broad claim and the evidence in Table 3; with qualification and proper statistics, the paper could be suitable for publication. I would also encourage the authors to release code and model checkpoints, as the paper does not currently state their availability."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The main thing you should know: GuitarDuets itself is a solid, genuinely new resource. A public dataset of ~3 hours of real and synthesized classical guitar duets, with a bleeding-free test set and note annotations for the synthetic half, is exactly what the monotimbral separation community needs. The paper also does an honest job of benchmarking Demucs across different training pools and openly discusses the poor second-guitar results. That part deserves credit.\n\nBut the headline claim is overstretched. The abstract says using both real and synthesized subsets leads to improved separation performance on the independent test set. Their own Table 3 contradicts that for guitar 2. Comparing the R-only row with the R+S row: G2 SDR drops from 1.014 to 0.920 dB, SAR from 1.424 to 0.896, SIR from 4.873 to 4.104. Only G1 SDR improves, from 4.952 to 5.882. For a duet task, both sources matter, so the abstract needs to say \"for the first guitar\" or switch to an explicitly defined aggregate metric. The body text is actually more careful—it only claims the highest SDR for G1—so this is a fixable wording problem, not a deep methodological flaw.\n\nThe bigger issue is that every number in Tables 3 and 4 is a single run. No seeds, no error bars, no significance tests. The G1 SDR gain of 0.93 dB could easily be noise. The metric analysis in Figure 4 is also built on one track (Track 29 of the synthetic set), so the SDR vs SI-SDR discussion is suggestive, not robust. No code is released, only the dataset, which limits reproducibility of the training recipes but does not undermine the dataset itself.\n\nIs this paper worth refereeing? Yes. The dataset alone justifies it, and the empirical study, once qualified, is a useful benchmark. I would not desk-reject it. But I would send it back with a clear request: qualify the abstract, add variance information across multiple runs, and either expand the metric analysis to more tracks or explicitly label it as a case study. The citation pattern looks fine — relevant prior work is covered, including GuitarSet, EnsembleSet, and the joint training literature. No circularity in the evaluation.\n\nMy recommendation: engage with this. It is a legitimate dataset contribution with an overconfident abstract and thin statistical support, both of which are repairable.","headline":"The GuitarDuets dataset is a real and useful contribution, but the abstract's claim that combining real and synthetic data improves separation is only supported for guitar 1's SDR, and the lack of error bars means even that gain is shaky.","tokens_in":11841,"tokens_out":2077,"would_cite":true,"duration_ms":25713,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Training on both real and synthesized classical guitar duets improves same-timbre separation more than training on either data type alone.","keywords":["music source separation","monotimbral separation","classical guitar duet","GuitarDuets dataset","Demucs","score-informed separation","synthetic audio","permutation-invariant training"],"falsifier":"Run the same training setups on a larger held-out set of real guitar duets recorded with separate microphones or direct pickups; if real-only training matches or beats combined training on that set, the claimed benefit of mixing synthetic data would disappear.","tokens_in":10955,"feed_emoji":"🎸","tokens_out":6117,"duration_ms":176043,"temperature":0.7,"pith_summary":"Classical guitar duets are hard to separate because both instruments share a similar timbre, so standard source separation models struggle to tell them apart. The paper introduces GuitarDuets, roughly three hours of real and synthesized classical guitar duet recordings, with note-level annotations for the synthesized portion. Using an adapted Demucs separation architecture, it benchmarks monotimbral separation and a joint transcription-and-separation pipeline. Its central finding is that training on both the real and synthesized subsets improves separation on a bleeding-free real test set compared to training on either subset alone. The paper also reports that ground-truth note labels substantially help separation, whereas predicted labels yield only marginal gains.","feed_headline":"Guitar duets: real plus synthetic audio sharpens separation","feed_subtitle":"New 3-hour dataset and benchmarks show training on both recorded and synthesized duets beats either alone.","key_machinery":"The load-bearing mechanism is the adapted Hybrid Transformer Demucs separator, a dual U-Net operating in both waveform and spectrogram domains, trained with a permutation-invariant loss plus a mixture-consistency term: $\\alpha \\min(|\\hat{g}_1-g_1|+|\\hat{g}_2-g_2|,|\\hat{g}_2-g_1|+|\\hat{g}_1-g_2|) + \\beta |(\\hat{g}_1+\\hat{g}_2)-(g_1+g_2)|$, with $\\alpha=0.8$ and $\\beta=0.2$. This loss allows the two output channels to swap guitars without penalty and keeps their sum close to the input mixture. The dataset is the second piece of machinery: 58.6 minutes of real duets recorded with two microphones plus 106 minutes of synthesized duets generated with a virtual nylon-guitar instrument, the synthesized half carrying MIDI note annotations usable as score conditioning. Ground-truth labels are injected into both the temporal and spectral branches of Demucs, and when labels are unavailable a Residual Shuffle-Exchange transcription network generates them.","core_discovery":"The paper's central finding is that combining real recordings of simultaneously playing classical guitarists with synthesized duets generated from MIDI scores improves monotimbral separation over using either data type alone, when evaluated on a specially recorded seven-track test set free of cross-microphone bleed. Cross-dataset experiments with the adapted Demucs model show that training on the complete GuitarDuets dataset gives the highest SDR for the first guitar, while the second guitar is separated more consistently when GuitarSet is also included in training. A second result is that conditioning the separator on ground-truth note activity labels improves separation, especially SIR, while labels predicted by a separate Residual Shuffle-Exchange transcription network help only marginally on real duets and slightly hurt on synthetic duets. Finally, the paper argues that SDR and SI-SDR behave differently for monotimbral mixtures than for multitimbral ones, so direct SDR comparisons with multitimbral benchmarks are not appropriate.","pith_inferences":["The paper leaves implicit that the synthetic subset's MIDI annotations could be used to train far more polyphonic guitar transcribers, which might turn the marginal predicted-label gains into substantial ones once the transcriber generalizes better.","The mixture-consistency term in the loss is a general regularization that could be tested on other monotimbral tasks such as violin duets or choir separation.","The consistently weaker performance on the second guitar suggests the model treats it as residual noise; this points toward source-identity embeddings, similar to speaker embeddings, as a promising direction the paper does not explore.","A practical extension would be listening tests to check whether the measured SDR gains from combined training correspond to perceptually cleaner separation, a question the paper flags but does not answer."],"forward_implications":["The GuitarDuets dataset provides a benchmark for monotimbral separation where no suitable polyphonic same-instrument dataset previously existed.","Synthesized data can usefully supplement scarce real recordings for same-timbre separation, at least when the real test conditions are represented in training.","Score-informed separation helps most when note labels are exact, so systems that rely on predicted transcriptions must close the gap between transcription and separation.","SDR and SI-SDR scores for monotimbral separation should not be compared directly with multitimbral separation benchmarks.","The joint transcription-and-separation framework gives a template for using note predictions as auxiliary information in other monotimbral tasks."],"supporting_citations":[{"why":"Supplies the Hybrid Transformer Demucs architecture that the paper adapts as its separation backbone.","marker":"[5]"},{"why":"GuitarSet, the existing real acoustic-guitar dataset used in cross-dataset training and as a source of note labels for transcription.","marker":"[30]"},{"why":"The Residual Shuffle-Exchange Network used as the transcription model that predicts note-level labels for the joint framework.","marker":"[37]"},{"why":"Introduces permutation-invariant training, the basis of the loss used to resolve which output corresponds to which guitar.","marker":"[32]"},{"why":"Defines SDR, the primary separation metric whose behavior on monotimbral mixtures the paper analyzes.","marker":"[45]"},{"why":"Defines SI-SDR, the scale-invariant metric compared with SDR in the metric analysis.","marker":"[46]"}],"fun_headline_variants":["Real and synthesized guitar duets improve separation","GuitarDuets dataset merges real and synthesized audio","Monotimbral separation benefits from mixed training data","Combining recorded and synthetic duets beats single-source data","Guitar duet separation: real plus synthetic recordings perform best"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation rests on a seven-track test set recorded without microphone bleed, and if that test set is not representative of ordinary real duet recordings, the measured gains from adding synthetic data may not transfer to realistic conditions.","fun_headline_variants_meta":{"raw":{"variants":["Real and synthesized guitar duets improve separation","GuitarDuets dataset merges real and synthesized audio","Monotimbral separation benefits from mixed training data","Combining recorded and synthetic duets beats single-source data","Guitar duet separation: real plus synthetic recordings perform best"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000265,"raw_usage":{"total_tokens":1618,"prompt_tokens":966,"completion_tokens":652,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":582,"completion_tokens_details":{"reasoning_tokens":572}},"tokens_in":582,"tokens_out":652,"duration_ms":7649,"temperature":1.0,"reasoning_tokens":572,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T20:57:26.093493+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same training setups on a larger held-out set of real guitar duets recorded with separate microphones or direct pickups; if real-only training matches or beats combined training on that set, the claimed benefit of mixing synthetic data would disappear.","supporting_citations":[{"cited_title":"3rd Call for H.F.R.I. Research Projects to support Post-Doctoral Researchers","cited_arxiv_id":null,"evidence_quote":"Supplies the Hybrid Transformer Demucs architecture that the paper adapts as its separation backbone."},{"cited_title":"A unified model for zero-shot music source separation, transcrip- tion and synthesis,","cited_arxiv_id":null,"evidence_quote":"GuitarSet, the existing real acoustic-guitar dataset used in cross-dataset training and as a source of note labels for transcription."},{"cited_title":"Guitarset: A dataset for guitar transcription,","cited_arxiv_id":null,"evidence_quote":"The Residual Shuffle-Exchange Network used as the transcription model that predicts note-level labels for the joint framework."},{"cited_title":"Train from scratch: Single-stage joint training of speech separa- tion and recognition,","cited_arxiv_id":null,"evidence_quote":"Introduces permutation-invariant training, the basis of the loss used to resolve which output corresponds to which guitar."},{"cited_title":"Residual shuffle-exchange networks for fast processing of long sequences,","cited_arxiv_id":null,"evidence_quote":"Defines SDR, the primary separation metric whose behavior on monotimbral mixtures the paper analyzes."},{"cited_title":"Syn- thtab: Leveraging synthesized data for guitar tabla- ture transcription,","cited_arxiv_id":null,"evidence_quote":"Defines SI-SDR, the scale-invariant metric compared with SDR in the metric analysis."}],"review_version":1}