{"id":"7e6b7a79-02e1-4069-9d6f-aa2a8e350dfb","arxiv_id":"2411.18075","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Deliberately bad recorder playing is treated as a learnable style, and VAE-GAN outperforms StarGAN and DDSP at transferring normal instruments to that style on the new FR109 dataset.","lead":"This paper turns normal recordings of violin, clarinet, and saxophone into recordings of a deliberately off-pitch, poorly played recorder, and compares three existing style transfer models on that task. It also releases FR109, a five-hour dataset of intentionally failed recorder performances, and finds that a VAE-GAN model gets closest to the target sound.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"FR109's failed-recorder target comes from one professional player, so the demonstrated learnability may be an artifact of that player's idiosyncratic errors rather than a general failed-recorder style.","rationale":"The paper's central contribution is a new dataset and a first benchmark for failed-music style transfer. Taking the strongest claim at face value: if a model can convert clean instrumental audio into audio judged as a failed recorder, the task is learnable. For that claim to be meaningful beyond the specific recording session, the target distribution must be stable across recorder players who produce 'failed' performances. FR109 is the only target domain, and all target examples come from one professional. That is precisely the weakest spot: no amount of FAD or MOS computed against FR109 test clips can distinguish a well-defined style from a single-player confound. I checked whether internal reasoning breaks: the spectrogram and Wiener entropy analyses in Section V do show that generated audio is more noise-like than clean instruments, consistent with the target, but the same evidence would appear if the model merely imitated one player's noise characteristics. No independent code or implementation is provided in the paper, so there is no implementation evidence to lean on beyond the reported experiments. The concrete check is therefore an external-validity check: add at least two more players recording the same error taxonomy and re-run conversion. If cross-player FAD/MOS stays in the same range, the concern is resolved; if not, the conclusions should be scoped to one player's failed recorder. This does not require rejecting the dataset or the benchmark; it means current evidence is insufficient to establish the general claim. The reader's verdict of CONDITIONAL already captures this, so I recommend UNCHANGED.","tokens_in":7150,"tokens_out":3978,"duration_ms":39303,"concrete_test":"Recruit at least two additional recorder players to record a subset of the same musical pieces with the same instructed error taxonomy (cracked voice, weird dynamics, failed tonguing, overblowing, underblowing). Compute pairwise FAD (VGGish embeddings, as in the paper) across players and within-player split halves. Then train the same VAE-GAN on FR109 only and convert held-out Bach10 sources, measuring FAD/MOS against each new player's real recordings. If cross-player FAD and MOS are in the same range as the within-player numbers in Tables II–III, the failed-recorder style generalizes. If FAD degrades substantially, or listeners cannot match outputs to the intended player/error set, the benchmark is performer-specific and the paper's conclusions should be scoped to one player's failed recorder.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim is that failed-music style transfer is feasible and VAE-GAN is the best of the tested models. The load-bearing premise, explicit in Section II-C, is that FR109's 109 recordings by a single professional with intentionally introduced errors define a coherent target style. All target-domain training data, the FAD reference set (Section IV-B1), and the style reference for listening come from that same player. With only one performer, there is no way to distinguish properties of 'failed recorder playing' from artifacts of one individual's technique, instrument, microphone, or error habits. The paper gives no inter-performer consistency check, no per-piece error-type annotation, and no analysis of which learned spectrogram changes are due to the player's specific articulations. Consequently, FAD 7.27 and MOS SS 2.98 may measure fidelity to a single idiosyncratic target, and the conclusion that the models transfer to 'failed recorder' in general is not yet supported. Secondary weaknesses — no FAD error bars, only SQ reaching p<0.05 between VAE-GAN and StarGAN, and one randomly chosen clip per source-target pair in listening — make the model-ranking claim brittle but do not determine the benchmark's external validity. This is not an internal inconsistency; the paper is honest that FR109 is one player, but it is the main unresolved load-bearing assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a new task, “failed-music style transfer,” where normally played audio is converted to the style of a deliberately poor soprano recorder performance. The authors contribute the FR109 dataset: 109 recordings (5.05 hours) by one professional recorder player who intentionally injects errors such as cracked voice, overblowing, and failed tonguing. They compare three existing methods (StarGAN, VAE-GAN, and DDSP) for converting violin, clarinet, and saxophone clips from the Bach10 dataset into the failed-recorder target, using FAD, MOS listening tests, Mel-spectrogram inspection, and Wiener entropy. Their empirical claim is that VAE-GAN is the best of the tested models, with FAD 7.27 and MOS scores of SS 2.98, MS 3.56, SQ 3.00, beating StarGAN and DDSP. The paper also analyzes spectrographic differences, such as inharmonic partials, between well-played and failed-recorder audio.","tokens_in":7378,"tokens_out":4511,"duration_ms":41740,"significance":"If the FR109 dataset is accepted as representing a coherent failed-recorder style, the paper opens a new benchmark scenario for style transfer and provides a publicly available dataset to support it. The central premise, however, is that one professional player’s intentionally introduced errors define a generalizable style, and this is not validated by any inter-performer consistency check. The paper also includes no derivation; its value rests on the empirical comparison and dataset release. The comparison is useful preliminary evidence that adversarial spectrogram-based models can capture noisy, inharmonic target styles, while a DSP-based synthesizer (DDSP) fails under its pitch-invariance assumption. Overall, the contribution is more a dataset-and-task proposal than a conclusive model ranking, and the load-bearing assumption about the target style’s representativeness must be addressed before the broader claims can be accepted.","major_comments":[{"comment":"The entire target domain is a single professional player’s intentionally failed recordings. The paper repeatedly refers to “failed recorder style” in general, but no evidence is given that FR109 is representative of failed recorder playing beyond this one individual. There is no inter-performer consistency check, no per-piece annotation of error types, and no analysis of how much the learned spectrogram changes are player-specific. Consequently, the FAD and MOS results may measure fidelity to one person’s idiosyncratic artifacts rather than to a generalizable failed-recorder style. The paper should either add a second performer (even a small validation set) or explicitly restrict all claims to “the failed-recorder style of the FR109 player” and discuss the limitation.","section":"Section II-C (FR109 dataset)"},{"comment":"The analysis in Section V-A states that, based on Mel spectrograms and informal listening, StarGAN achieves better style similarity to a failed recorder than VAE-GAN (e.g., inharmonic partials are “not as clearly” present in VAE-GAN). This directly contradicts the formal MOS SS scores in Table III, which show VAE-GAN (2.98) ahead of StarGAN (2.54). The paper must reconcile this inconsistency: either the spectrogram inspection and informal listening are unreliable, or the MOS result is unexpected and needs explanation. As written, the contradiction undermines the credibility of both the objective analysis and the subjective evaluation.","section":"Section V-A vs. Table III"},{"comment":"The FAD scores are reported as single point estimates with no variance, confidence interval, or bootstrap uncertainty. With only 10 Bach10 pieces as the evaluation set, the difference between StarGAN (13.87) and VAE-GAN (7.27) cannot be assessed for significance. The text “StarGAN performs slightly worse than VAE-GAN” is therefore unsupported by the reported numbers. Please provide uncertainty estimates (e.g., bootstrap over clips or multiple FAD computations) or soften the claim accordingly.","section":"Section IV-B1, Table II (FAD)"},{"comment":"The listening test uses one randomly chosen clip per source-target pair, i.e., three clips per model in total (violin-to-recorder, clarinet-to-recorder, saxophone-to-recorder), with 16 responses. The presented p-values (SS p=0.09, MS p=0.06, SQ p=0.02) are only significant for SQ. The conclusion that “StarGAN’s overall performance falls behind VAE-GAN’s on all metrics” is thus not supported by conventional significance thresholds for two of the three metrics. Additionally, the paper does not state whether the 16 responses were pooled across the three clips per model or analyzed per clip; this should be clarified, and the statistical power of using a single clip per condition should be discussed.","section":"Section IV-B2 (Subjective evaluation)"}],"minor_comments":[{"comment":"There is a typo: “V AR-GAN” should be “VAE-GAN.”","section":"Section IV-B1"},{"comment":"The text says FAD is a “reference-free metric,” but FAD compares an evaluation set against a reference embedding set, so it is reference-based. Please correct the wording to avoid ambiguity.","section":"Section IV-B1"},{"comment":"The sentence “StarGAN performs slightly worse than VAE-GAN on both datasets” is misleading because only the Bach10 dataset is used for FAD evaluation; no FAD on URMP or FR109 test clips is reported.","section":"Section IV-B1"},{"comment":"Training details are sparse: the paper does not report hyperparameters (e.g., learning rate, number of steps, adversarial loss weights) for the three methods or how the 90/10 split of FR109 was performed. Adding these details would improve reproducibility.","section":"Section IV-A"},{"comment":"The STFT parameters used to compute Wiener entropy (window size, hop, FFT size) are not specified, making it difficult to reproduce the entropy values in Tables IV and V.","section":"Section V-B"}],"recommendation":"major_revision","confidential_remarks":"The main unresolved issue is the external validity of the single-player FR109 dataset as a representation of “failed recorder” style. This is not an internal inconsistency, but it is load-bearing for the paper’s central claim. The contradiction between Section V-A and Table III should be fixed before resubmission; it suggests the authors’ own informal judgments may not align with the formal MOS results. If the authors can add at least a small validation set with a different performer, or reframe all claims as single-player style transfer, the paper would be a solid contribution. Otherwise, the model-ranking conclusion remains brittle, especially given the lack of FAD uncertainty and the marginal p-values for SS and MS."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a modest but honest dataset-and-benchmark paper. The new thing is the task of transferring music to a deliberately failed-recorder target, and the FR109 dataset—109 pieces, 5.05 hours, recorded by one professional with intentional errors. That dataset is a real, reusable resource, and the paper makes a reasonable first attempt at benchmarking StarGAN, VAE-GAN, and DDSP on it. I agree with the reader that the central claim should be conditional, but I'd put more weight on the single-performer issue than on the statistics.\n\nWhat the paper does well: it frames a genuinely new scenario and releases the data (though no code or eval scripts, so 'reproducibility' is only partial). The analysis with Wiener entropy and Mel spectrograms is a nice touch—it gives a concrete acoustic correlate of the target style and shows the models do move in that direction. The write-up is candid; Section II-C openly states the recordings are by one professional, and the paper does not pretend otherwise.\n\nSoft spots, in order of importance. First, the target domain is one person's failed recorder. The paper treats 'failed recorder' as an instrument style, but with a single performer there is no way to separate the intended errors from that player's technique, microphone, or error habits. No inter-performer consistency check is done or even discussed. So the FAD 7.27 and MOS scores may measure fidelity to this one player rather than to a general failed-recorder style. Second, the statistics are thinner than the conclusion suggests. FAD has no variance estimate; the listening test uses one clip per source-target pair and 16 responses, and only sound quality reaches p<0.05 (SS p=0.09, MS p=0.06). The paper still says 'we can still conclude'—that is an overreach. Third, there is an internal tension: the qualitative spectrogram analysis says StarGAN captures the failed-recorder characteristic better, while the MOS trend favors VAE-GAN on style similarity. That inconsistency is worth resolving.\n\nCitation pattern is fine, and there is no circular derivation to worry about. If anything, the paper is under-ambitious in not trying a simple baseline like a pitch-shift plus noise augmentation to test whether the 'failed' part is even being modeled.\n\nRecommendation: send it to peer review. It is a legitimate empirical contribution with a new dataset and an honest write-up, and it deserves referee time. But the decision should be a conditional accept: accept the dataset and task framing, and ask for either an inter-performer check or a clear rewording that the target is 'one player's failed recorder.' My own verdict would be: valuable dataset, provisional model ranking.\n\nWho profits: researchers working on timbre transfer or expressive synthesis who want a harder, off-pitch target domain. I'd cite the dataset.","headline":"New failed-recorder dataset and task framing are the real contributions; the claimed VAE-GAN superiority is provisional given single-performer target and weak significance.","tokens_in":7940,"tokens_out":2643,"would_cite":true,"duration_ms":23218,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Failed recorder playing is a learnable AI music style","keywords":["failed music style transfer","FR109 dataset","recorder","off-pitch performance","VAE-GAN","StarGAN","DDSP","Wiener entropy"],"falsifier":"Record a second failed-recorder dataset from several different performers, retrain VAE-GAN on FR109 as before, and evaluate on this new set; if FAD jumps substantially or listeners no longer identify the output as a recorder, the style is not general, whereas comparable scores would confirm the style is a stable target.","tokens_in":6940,"feed_emoji":"🎵","tokens_out":3957,"duration_ms":34050,"temperature":0.7,"pith_summary":"This paper tries to establish that deliberately failed recorder performance, which most music style transfer work ignores, can be a legitimate target domain for style transfer. The authors introduce FR109, a dataset of 109 pieces of intentionally off-pitch recorder playing by a professional, and use it to convert clean violin, clarinet, and saxophone recordings into audio that listeners still recognize as a flawed recorder. They find that VAE-GAN outperforms StarGAN and DDSP on this task, with the best objective and subjective scores. The paper argues that learning the inharmonic, noise-heavy character of failed playing is a meaningful stress test for style transfer models that normally assume both source and target are well played.","feed_headline":"Failed recorder playing is a learnable AI style","feed_subtitle":"New FR109 dataset of intentional errors makes off-pitch recorder the target; VAE-GAN converts clean instruments best.","key_machinery":"The central object is the FR109 dataset: 109 pieces, 5.05 hours of soprano recorder performed by a professional with deliberate errors such as cracked voice, weird dynamics, failed tonguing, overblowing, and underblowing. The transfer pipeline converts source audio to Mel spectrograms, passes them through a style-transfer generator (StarGAN or VAE-GAN), and reconstructs the waveform with the BigVSAN vocoder. The paper quantifies the style signal using Wiener entropy, which measures how noise-like a signal is, and uses that metric to show that failed recorder is clearly noisier than well-played instruments and that current models still fall short of reproducing the full noise level.","core_discovery":"The central claim is that failed-music style transfer is feasible: a model trained on the FR109 failed-recorder dataset can turn well-played source audio into audio that sounds like an off-pitch recorder while preserving the melody, and VAE-GAN does this better than StarGAN or DDSP on the Bach10 test set. The paper reports FAD 7.27, MOS style similarity 2.98, melody similarity 3.56, and sound quality 3.00 for VAE-GAN, and it uses Wiener entropy to show that the failed-recorder style is characterized by much more noise (entropy 0.0345) than well-played instruments (0.0005), so the model must learn to generate deliberately noisy output.","pith_inferences":["Because FR109 was recorded by only one professional, the current findings may capture that player's individual 'failed' style rather than a general failed-recorder category; a multi-player version of the dataset would test whether the style is transferable.","The same approach could apply to other deliberately degraded or expressive instrumental styles, such as vocal fry, jazz inflections, or distorted electric guitar, turning performance 'mistakes' into a controllable style dimension.","Wiener entropy could be used as a training-time regularizer to push generated audio closer to the failed-recorder noise profile, since the paper only uses it for analysis."],"forward_implications":["VAE-GAN's per-domain decoders are better suited than StarGAN's single unified decoder for capturing the noisy, inharmonic character of failed playing.","DDSP's pitch-invariance assumption breaks down when the target is intentionally off-pitch, so differentiable-synthesizer pipelines need modification for this scenario.","The FR109 dataset gives style transfer researchers a reusable benchmark for testing models on expressive or intentionally flawed instrumental styles.","The measured gap between FR109's Wiener entropy and that of converted audio indicates that even the best tested model only partially reproduces the failed-recorder noise profile."],"supporting_citations":[{"why":"Supplies the clean multi-track violin, clarinet, and saxophone recordings used as training data for the source domains.","marker":"[12]"},{"why":"Provides the held-out Bach chorale recordings used to evaluate conversion to failed recorder.","marker":"[13]"},{"why":"StarGAN is the unified single-decoder multi-domain baseline that the paper compares against.","marker":"[15]"},{"why":"VAE-GAN is the variational multi-decoder method that achieves the best failed-recorder conversion results.","marker":"[7]"},{"why":"DDSP is the differentiable synthesizer baseline whose pitch-invariance assumption is challenged by off-pitch targets.","marker":"[3]"},{"why":"BigVSAN is the neural vocoder that reconstructs the final waveform from transferred Mel spectrograms.","marker":"[16]"},{"why":"FAD is the objective reference-free metric the paper uses to compare generated and real audio distributions.","marker":"[18]"},{"why":"CREPE pitch extraction is used to report the pitch statistics of the FR109 dataset.","marker":"[14]"},{"why":"Provides the Wiener entropy definition used to quantify the noise-likeness of failed recorder audio.","marker":"[21]"}],"fun_headline_variants":["Off-pitch recorder style now AI-transferable","AI learns the art of playing recorder badly","From clean to failed: style transfer to off-pitch recorder","FR109 dataset teaches AI to mimic bad recorder","VAE-GAN best at converting music to failed recorder"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire failed-recorder style is captured by one professional player's intentionally flawed performances, so the model may be learning that player's idiosyncrasies rather than a general 'failed recorder' style.","fun_headline_variants_meta":{"raw":{"variants":["Off-pitch recorder style now AI-transferable","AI learns the art of playing recorder badly","From clean to failed: style transfer to off-pitch recorder","FR109 dataset teaches AI to mimic bad recorder","VAE-GAN best at converting music to failed recorder"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000162,"raw_usage":{"total_tokens":1188,"prompt_tokens":846,"completion_tokens":342,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":462,"completion_tokens_details":{"reasoning_tokens":267}},"tokens_in":462,"tokens_out":342,"duration_ms":3267,"temperature":1.0,"reasoning_tokens":267,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:31:32.578967+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record a second failed-recorder dataset from several different performers, retrain VAE-GAN on FR109 as before, and evaluate on this new set; if FAD jumps substantially or listeners no longer identify the output as a recorder, the style is not general, whereas comparable scores would confirm the style is a stable target.","supporting_citations":[{"cited_title":"Cre- ating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications,","cited_arxiv_id":null,"evidence_quote":"Supplies the clean multi-track violin, clarinet, and saxophone recordings used as training data for the source domains."},{"cited_title":"Multiple fundamen- tal frequency estimation by modeling spectral peaks and non-peak regions,","cited_arxiv_id":null,"evidence_quote":"Provides the held-out Bach chorale recordings used to evaluate conversion to failed recorder."},{"cited_title":"Stargan: Unified generative adversarial net- works for multi-domain image-to-image translation,","cited_arxiv_id":null,"evidence_quote":"StarGAN is the unified single-decoder multi-domain baseline that the paper compares against."},{"cited_title":"Timbre trans- fer with variational auto encoding and cycle-consistent adversarial networks,","cited_arxiv_id":null,"evidence_quote":"VAE-GAN is the variational multi-decoder method that achieves the best failed-recorder conversion results."},{"cited_title":"Bigvsan: En- hancing gan-based neural vocoders with slicing adver- sarial network,","cited_arxiv_id":null,"evidence_quote":"BigVSAN is the neural vocoder that reconstructs the final waveform from transferred Mel spectrograms."},{"cited_title":"Crepe: A convolutional representation for pitch estimation,","cited_arxiv_id":null,"evidence_quote":"CREPE pitch extraction is used to report the pitch statistics of the FR109 dataset."},{"cited_title":"Transform coding of audio signals using perceptual noise criteria,","cited_arxiv_id":null,"evidence_quote":"Provides the Wiener entropy definition used to quantify the noise-likeness of failed recorder audio."}],"review_version":1}