{"id":"2c108818-3b74-41d7-a422-6efbfb800603","arxiv_id":"2608.05513","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A variational autoencoder with content and tone embeddings plus a tone-masking operation unifies electric guitar tone transfer and tone removal in one model.","lead":"This paper builds a single neural model, EG-VAE, that can both change an electric guitar recording's tone to match a reference and remove the tone to recover the dry, unprocessed signal. This matters because previous systems handled the two tasks separately and struggled with wet, processed guitar input.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Plugin-only evaluation makes the general EGTT/EGTR claim depend on unvalidated transfer from two Neural DSP renderings to real guitar chains.","rationale":"The reader's weakest assumption identifies the same load-bearing point: the evaluation is confined to a plugin-rendered distribution, so the broad claim about solving EGTT and EGTR for electric guitar depends on an unvalidated leap to real signal chains. My independent reading of the paper finds no internal inconsistency or circular reasoning in the architecture, training objective, or ablation logic; the tone-masking mechanism is coherent and the two-stage training is plausible. The ablations are internally consistent, and giving the baselines access to EG-VAE's own removal output strengthens the transfer comparison. The remaining risk is external: every reported improvement is measured inside a two-plugin universe, and the 'unseen' tone set is only a second plugin from the same vendor family. This is not a disagreement with current consensus; it is a correctness risk about whether the learned factorization is a general guitar-tone model or a plugin renderer. The concrete test I propose would settle this by measuring performance on genuinely external wet/dry data. Given this is exactly the condition the reader already attached, my recommendation is unchanged: the paper should be accepted conditionally on demonstrating transfer beyond the plugin renderings and on releasing data/code for verification.","tokens_in":18709,"tokens_out":3010,"duration_ms":32350,"concrete_test":"Evaluate EG-VAE and the same baselines on real third-party wet/dry pairs: source or collect simultaneous DI and processed recordings through non-Neural-DSP analog or software chains (e.g., tube amp with cabinet/mic, analog pedalboard, or another manufacturer's plugin). Keep the current training data unchanged, then compute the Table I and Table II metrics on these real pairs. If EG-VAE's margin over the strongest baseline narrows or reverses, the plugin-only evaluation is the source of the claimed advantage; if the margin persists, the external-validity concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is that EG-VAE outperforms task-specific baselines on tone transfer and removal for electric guitar. All evidence for this claim comes from a single constructed dataset: EGDB dry DI signals rendered through two commercial Neural DSP plugins, with Archetype Cory Wong X for seen tones and Morgan Amps Suite for unseen tones (Section IV-A). The \"unseen\" condition is therefore a held-out plugin of the same brand and rendering paradigm, not a real-world signal chain. The model is trained and evaluated entirely on this plugin-generated distribution, and the stage-2 Wiener-Hammerstein augmentation (Section III-E) further shapes the tone manifold toward plugin-like chains. If real electric guitar recordings involve signal chains outside this manifold (analog pedals, tube amplifiers, cabinet/microphone coloration, recording and mixing effects), nothing in the paper shows the learned content-tone factorization or the masking-based removal will transfer. The headline numbers in Tables I and II could shrink or invert outside the rendering distribution. Because the claim is framed as solving EGTT/EGTR generally, and no real or third-party wet/dry pairs are used anywhere, external validity is the most load-bearing assumption; lack of released code or data compounds it by preventing independent checks.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes EG-VAE, a unified variational autoencoder for electric guitar tone transfer (EGTT) and tone removal (EGTR). A wet recording is factorized into a frame-level content embedding and a global tone embedding; transfer recombines a source's content with a reference's tone, while removal is performed by a 'tone masking' operation that overrides the conditional layer normalization to decode from content alone. A second training stage adds variational sampling and Wiener-Hammerstein audio-effects augmentation to improve handling of unseen tones. The model is evaluated on a custom dataset rendered from EGDB dry signals through two Neural DSP plugins, compared with task-specific baselines using spectral distances and a 14-participant listening test, and analyzed with ablations and a perceptual-path-length smoothness measure.","tokens_in":18900,"tokens_out":6527,"duration_ms":56229,"significance":"If the results hold beyond the rendering distribution, the work is a meaningful step in guitar tone modeling: it unifies two previously separate tasks in a single representation, introduces a simple tone-masking mechanism that serves both training and inference, and reports large objective and subjective gains over task-specific baselines. The paper is clearly written, includes a thorough ablation of each proposed component, evaluates tone-space smoothness, and provides demonstration audio. The central novelty—tone masking as a dual-purpose disentanglement/removal objective, together with the two-stage training for unseen-tone transfer—is conceptually attractive and well motivated. The main open question is external validity, because all training and test data are plugin renderings and no real-world wet/dry recordings are used.","major_comments":[{"comment":"The entire empirical evaluation is based on renderings of EGDB dry signals through two Neural DSP plugins (Archetype: Cory Wong X for seen tones, Morgan Amps Suite for unseen tones); no real-world wet/dry recordings or third-party data are used anywhere. The 'unseen' condition is thus a held-out plugin of the same brand and rendering paradigm, not a real signal chain. Since the abstract and introduction claim that EG-VAE solves EGTT and EGTR generally, this is a load-bearing limitation: the learned content–tone factorization and masking-based removal may not transfer to chains involving analog pedals, tube amplifiers, cabinet/microphone coloration, or recording/mixing effects. The authors should either temper the generalization claim to 'plugin-rendered tones' or, preferably, evaluate on a modest set of real wet/dry guitar recordings (or a third-party dataset) to test whether the reported gains persist.","section":"Section IV-A and Abstract"},{"comment":"The baselines are retrained with the optimization objective standardized to the multi-scale mel-spectrogram loss used for EG-VAE, and Distortion Recovery's HiFi-GAN vocoder is replaced with BigVGAN. For methods originally designed with different losses (e.g., HDemucs hybrid spectrogram/waveform loss, DPTNet SI-SNR), this standardization may arbitrarily disadvantage them, and no ablation shows the standardized loss is suitable for each architecture. In addition, the 'w/ EGTR' baselines are trained on clean dry signals but at inference receive EG-VAE's estimated dry signal, a train/test distribution shift that could handicap them; the paper should justify this setup or fine-tune the baselines on estimated dry inputs. Without these controls, the magnitude of the reported advantage over baselines is uncertain.","section":"Section IV-C, Tables I and II"},{"comment":"The introduction states that EG-VAE improves over Distortion Recovery '1.10 vs 1.21 Mel on seen tones, matching on unseen.' Table II shows unseen Mel 1.19 vs 1.22 (EG-VAE better) and unseen STFT 1.87 vs 1.86 (Distortion Recovery better). This discrepancy should be corrected, and the abstract's 'outperforms task-specific baselines on transfer and removal' should be qualified to note that on unseen-tone STFT for removal the two models are essentially tied.","section":"Section I and Table II"}],"minor_comments":[{"comment":"The captions of Tables III and IV do not state which evaluation split they report; Table IV's numbers match the unseen-tone results in Tables I and II, but the text never says so explicitly. Please state in the captions and the ablation section that Table III is on seen tones and Table IV is on unseen tones.","section":"Section IV-D and Tables III/IV"},{"comment":"The objective metrics are averaged over 2,000 ten-second test examples, but the test partitions are described in hours (142 and 29 hours); please clarify how the 2,000 examples are sampled from these partitions.","section":"Section IV-D"},{"comment":"The subjective ratings are reported with standard deviations but no statistical significance testing; with 14 participants, a paired test (e.g., Wilcoxon signed-rank) would help confirm whether the AQ and Dryness differences between EG-VAE and the baselines are reliable.","section":"Section V-B"},{"comment":"The claim that tone masking 'directly' produces the dry signal is somewhat strong: it is only through training with L_mask that the decoder learns to map the no-modulation regime to the dry target, so the inference operation is a learned readout rather than an exact inverse. Consider softening the wording in Eq. (11) and the surrounding text.","section":"Section III-D"},{"comment":"The sentence 'owing to the highly nonlinear distortion and the wide diversity of effects ... its signal chain' is missing a preposition and is ungrammatical; please revise.","section":"Section I"}],"recommendation":"major_revision","confidential_remarks":"The paper is within the scope of the journal and the core idea is promising. The main risks are the absence of any real-world evaluation and the baseline-training choices described in Section IV-C; both directly affect the validity of the headline comparisons. I would be willing to review a revised version that adds a real-data test (even a small one) and justifies or corrects the baseline protocol."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core idea here is good, and it's new. Tying tone removal to a masking operation that also enforces disentanglement during training is a clever move -- one mechanism, two jobs -- and I don't know of prior work that does exactly this. The paper also does the right thing by giving baselines the best possible setup, including an EGTR-enhanced variant that uses the model's own removed signal as input; that makes the head-to-head comparison fairer than usual. The ablations are clean and each component (perturbation, MIDI supervision, tone masking, variational sampling, augmentation, pvpGD) shows a measurable contribution. The writing is clear and the limitations section is honest, noting the delay ambiguity as an open problem. This is a serious piece of work, not a desk reject.\n\nThe soft spot is the one the stress-test flagged, and it is real. Every piece of evaluation data comes from two commercial plugins (Cory Wong X and Morgan Amps Suite). The \"unseen\" tones are a held-out plugin of the same brand and rendering paradigm, not a real signal chain with analog pedals, tube amps, cabinets, or microphone coloration. The augmentation in stage 2 also stays within that plugin-like manifold. So the abstract's claim that EG-VAE outperforms task-specific baselines on transfer and removal is only demonstrated inside a plugin-rendered distribution. It may well transfer to real recordings -- the representation is learned from dry DI plus a broad chain, which is not obviously plugin-specific -- but the paper gives no evidence for that, and the claim is sweeping. This is the main reason I'd want external validation before believing the headline.\n\nMinor issues: no code, data, or checkpoints are released, which blocks independent checks. A few numbers are statistically tied (unseen STFT 1.87 vs 1.86) but discussed as \"comparable,\" and no significance tests are reported for subjective ratings or PPL. These are fixable in revision.\n\nWho is this for? Anyone working on guitar tone modeling, audio effects style transfer, or disentangled audio representations. It deserves a serious referee. I'd recommend conditional acceptance: require the authors to validate on real or third-party wet/dry recordings and release at least checkpoints or a demo instrument. If they do that, this becomes a strong reference point for the field.","headline":"A genuinely novel unified tone transfer/removal framework with convincing internal ablations, but the plugin-only evaluation makes the general guitar claim unproven; still deserving of peer review.","tokens_in":19538,"tokens_out":1484,"would_cite":true,"duration_ms":16221,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"EG-VAE claims that a single variational autoencoder, factorizing a wet electric guitar recording into frame-level content and global tone, can perform both tone transfer and tone removal better than task-specific baselines.","keywords":["electric guitar tone transfer","tone removal","dry DI recovery","content-tone disentanglement","variational autoencoder","tone masking","unseen tone generalization","audio effects modeling"],"falsifier":"Record wet/dry pairs from physical amplifier-and-pedal rigs or from many non-plugin renderers and compare EG-VAE against the same baselines; if its Mel-distance and listening-test advantages shrink or disappear outside the two-plugin distribution, the generalization claim fails.","tokens_in":18419,"feed_emoji":"🎸","tokens_out":6835,"duration_ms":60206,"temperature":0.7,"pith_summary":"EG-VAE claims that electric guitar tone transfer and tone removal are two directions of the same operation and can be solved by one model. The model factorizes a wet recording into a frame-level content embedding (what is played) and a global tone embedding (the signal chain), then realizes transfer by swapping tone embeddings and removal by a tone-masking forward pass. The paper reports that this single unified model beats task-specific baselines on both tasks, cutting the Mel distance of tone transfer from 1.53 to 0.86 on seen tones while also improving removal from 1.21 to 1.10. A second training stage with variational sampling and audio-effects augmentation is what makes unseen reference tones work. If true, this means removal and transfer can share one learned representation instead of separate task-specific systems.","feed_headline":"One VAE swaps or strips guitar tone, beating task-specific models","feed_subtitle":"The same content/tone split cuts transfer Mel distance to 0.86 and removes effects better than dedicated removal systems.","key_machinery":"The central mechanism is tone masking, which operates through conditional layer normalization (CLN), a layer that injects a global attribute via scale and shift parameters. Tone masking overrides those parameters to the identity ($\\gamma=1$, $\\beta=0$), so the content embedding alone reaches the decoder; supervised against the dry signal, this masked pass becomes the tone-removal procedure at inference and simultaneously penalizes tone leakage into content. The supporting machinery includes an information bottleneck on the content embedding, content-tone perturbation, pitch and preset supervision, posterior variance-parameterised Gaussian dropout to prevent tone-latent collapse, and a second training stage with variational sampling and audio-effects augmentation that smooths the tone space for unseen tones.","core_discovery":"On its own terms, the discovery is that a wet guitar recording can be split into a frame-level content embedding and a global tone embedding by a variational autoencoder, and that the same split serves both tasks. Tone transfer recombines one recording's content with another's tone embedding. Tone removal is the same masked forward pass used during training, where the conditional layer normalization's modulation is overridden to $\\gamma=1,\\beta=0$, leaving the content embedding to reconstruct the dry signal alone; this single operation both defines removal and strengthens disentanglement by forcing tone information out of the content path. The paper reports that EG-VAE reduces the Mel distance of transfer by 44% over the strongest baseline on seen tones (0.86 vs 1.53) and improves removal over the best specialized baseline (1.10 vs 1.21), with subjective ratings approaching the ground-truth anchor, and that a two-stage training procedure with variational tone sampling and effects augmentation preserves quality on unseen tones.","pith_inferences":["The plugin-based evaluation bounds the claim: real-world rigs add room acoustics, microphones, and physical amplifier behavior absent from the two commercial plugin chains, so the seen/unseen gains should be re-measured on recordings of physical chains.","Tone masking is a general recipe for conditional-normalization architectures: any attribute injected through scale/shift parameters can be removed by the identity modulation ($\\gamma=1$, $\\beta=0$), turning the masked pass into a self-supervised disentanglement objective for other attributes or instruments.","The paper's own conclusion flags delay as an open ambiguity: a delayed repetition is an effect acoustically but resembles performed content, so the current content/tone assignment has no consistent cue; resolving this could extend the framework to spatial and time-based effects.","Perceptual path length on unseen tones could serve as a tuning signal for audio style-transfer models generally, since it measures exactly the smoothness that makes zero-shot transfer work."],"forward_implications":["Tone removal stops being a separate module: the same masked forward pass used in training outputs the dry signal at inference, so a single checkpoint serves both tasks.","Because transfer operates on a representation already stripped of the input's tone, it sidesteps the error accumulation that a two-stage remove-then-model pipeline suffers.","Unseen reference tones remain usable: variational sampling plus effects augmentation makes the tone space smooth enough that held-out chains decode coherently, with perceptual path length dropping from 2.00 to 0.90.","Minimizing spectral distance is not enough for removal quality: one baseline with competitive spectral scores was rated lowest in audio quality, while EG-VAE led both audio quality and dryness in the listening test.","The ablation results indicate the disentanglement mechanisms are complementary; removing content-tone perturbation hurts all three modes most, and removing tone masking also degrades reconstruction and transfer, not just removal."],"supporting_citations":[{"why":"Supplies the dry DI guitar recordings that are rendered through effects chains to build the training and evaluation data.","marker":"[1]"},{"why":"Documents that guitar tone matching is a hard case for inference-time optimization and motivates the wet-input formulation.","marker":"[12]"},{"why":"The differentiable audio-effects style-transfer baseline; represents fixed-chain approaches that EG-VAE compares against.","marker":"[14]"},{"why":"The one-to-many amplifier-modeling baseline; in its EGTR-enhanced form it is the strongest EGTT competitor.","marker":"[15]"},{"why":"The specialized guitar effect-removal baseline that EG-VAE must beat on tone removal.","marker":"[17]"},{"why":"Supplies posterior variance-parameterised Gaussian dropout (pvpGD), which prevents collapse of the tone latent.","marker":"[19]"},{"why":"Supplies the disentangled sequential autoencoder factorization into frame-level and global latents that EG-VAE adapts.","marker":"[30]"},{"why":"The variational autoencoder framework and ELBO that ground the reconstruction and KL objectives.","marker":"[31]"},{"why":"The high-fidelity audio VAE whose encoder-decoder and adversarial setup EG-VAE adapts for content/tone factorization.","marker":"[33]"},{"why":"Conditional layer normalization is the combination mechanism that tone masking overrides to realize removal.","marker":"[35]"}],"fun_headline_variants":["Single VAE splits guitar tone and content, excels at both transfer and removal","Unified VAE disentangles guitar tone and content for better transfer and removal","One VAE does tone transfer and removal, beating specialized models","Guitar tone split into content and style: one VAE handles both tasks better","VAE framework unifies guitar tone swap and dry-signal recovery with better scores"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes that wet recordings rendered by two commercial guitar plugins are representative of real electric-guitar tones, so that a model trained and tested on these renderings will transfer to real amplifiers, pedals, and recording chains.","fun_headline_variants_meta":{"raw":{"variants":["Single VAE splits guitar tone and content, excels at both transfer and removal","Unified VAE disentangles guitar tone and content for better transfer and removal","One VAE does tone transfer and removal, beating specialized models","Guitar tone split into content and style: one VAE handles both tasks better","VAE framework unifies guitar tone swap and dry-signal recovery with better scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000719,"raw_usage":{"total_tokens":3231,"prompt_tokens":953,"completion_tokens":2278,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":569,"completion_tokens_details":{"reasoning_tokens":2177}},"tokens_in":569,"tokens_out":2278,"duration_ms":14250,"temperature":1.0,"reasoning_tokens":2177,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T11:42:05.741411+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record wet/dry pairs from physical amplifier-and-pedal rigs or from many non-plugin renderers and compare EG-VAE against the same baselines; if its Mel-distance and listening-test advantages shrink or disappear outside the two-plugin distribution, the generalization claim fails.","supporting_citations":[{"cited_title":"Towards automatic transcription of polyphonic electric guitar music: A new dataset and a multi-loss transformer model,","cited_arxiv_id":null,"evidence_quote":"Supplies the dry DI guitar recordings that are rendered through effects chains to build the training and evaluation data."},{"cited_title":"ST-ITO: Controlling audio effects for style transfer with inference-time optimization,","cited_arxiv_id":null,"evidence_quote":"Documents that guitar tone matching is a hard case for inference-time optimization and motivates the wet-input formulation."},{"cited_title":"Style transfer of audio effects with differentiable signal processing,","cited_arxiv_id":null,"evidence_quote":"The differentiable audio-effects style-transfer baseline; represents fixed-chain approaches that EG-VAE compares against."},{"cited_title":"Towards zero-shot amplifier modeling: One-to-many amplifier modeling via tone embedding control,","cited_arxiv_id":null,"evidence_quote":"The one-to-many amplifier-modeling baseline; in its EGTR-enhanced form it is the strongest EGTT competitor."},{"cited_title":"Distortion recovery: A two-stage method for guitar effect removal,","cited_arxiv_id":null,"evidence_quote":"The specialized guitar effect-removal baseline that EG-VAE must beat on tone removal."},{"cited_title":"Posterior variance-parameterised gaussian dropout: Improving disentangled sequential autoencoders for zero-shot voice conversion,","cited_arxiv_id":null,"evidence_quote":"Supplies posterior variance-parameterised Gaussian dropout (pvpGD), which prevents collapse of the tone latent."},{"cited_title":"Disentangled sequential autoencoder,","cited_arxiv_id":null,"evidence_quote":"Supplies the disentangled sequential autoencoder factorization into frame-level and global latents that EG-VAE adapts."},{"cited_title":"Adaspeech: Adaptive text to speech for custom voice,","cited_arxiv_id":null,"evidence_quote":"Conditional layer normalization is the combination mechanism that tone masking overrides to realize removal."}],"review_version":1}