{"id":"04efb736-c659-45f7-8b2c-6b7ade8d234f","arxiv_id":"2606.24912","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Transfer learning from synthetic velocity-labeled guitar data enables velocity prediction in automatic guitar transcription while maintaining competitive note transcription performance.","lead":"The paper presents a transfer learning method that pretrains a guitar transcription model on synthetic audio with velocity labels from virtual instruments, then fine-tunes on real recordings. This approach could allow velocity prediction in guitar AMT without requiring large labeled real-world datasets.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Velocity prediction accuracy demonstrated only on synthetic data; transfer effectiveness to real guitar audio unmeasured directly","rationale":"The identified concern is identical to the reader's weakest assumption. The abstract-only review already flags the synthetic-to-real transfer as unverified for velocity itself; the reported results do not alter that assessment.","tokens_in":1696,"tokens_out":263,"duration_ms":16998,"concrete_test":"Annotate velocity labels on a held-out set of 50–100 real guitar recordings (expert or MIDI-controller capture) and compute velocity MAE or correlation for the transferred model versus the non-pretrained baseline; if the gap disappears or reverses, the transfer claim does not hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that pretrained velocity weights from synthetic virtual-instrument data remain functional after fine-tuning on unlabeled real recordings. The paper reports velocity outperformance solely versus a non-pretrained baseline on synthetic test data; real-audio evaluation is limited to note-level transcription metrics that show only small, sometimes non-significant gains. No proxy (e.g., correlation with loudness, human ratings, or cross-instrument consistency) for velocity quality on real guitar is described, so the transfer of velocity statistical properties is assumed rather than tested.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes a method for velocity prediction in automatic guitar transcription by generating synthetic data with velocity labels from virtual instruments, pretraining a model on this data, transferring the weights to a model trained on real unlabeled guitar audio, and claiming that the resulting model outperforms a non-pretrained baseline on synthetic velocity prediction while yielding a small (sometimes non-significant) gain in note transcription and results comparable to SOTA guitar transcription.","tokens_in":1820,"tokens_out":411,"duration_ms":15311,"significance":"If the transfer of velocity prediction from synthetic to real guitar audio holds under direct testing, the work would address a notable gap in AMT for non-piano instruments by enabling velocity-aware models without requiring real velocity labels; the use of synthetic pretraining and transfer learning is a practical strength.","major_comments":[{"comment":"Abstract and evaluation description: outperformance on velocity is reported solely versus a non-pretrained baseline on synthetic test data, but no quantitative metrics, error bars, statistical tests, or definition of how velocity is computed or scored for guitar are provided.","section":"Abstract"},{"comment":"The central transfer claim (pretrained velocity weights remain functional after fine-tuning on real recordings) is supported only by indirect note-transcription metrics showing small and sometimes non-significant gains; no direct evaluation or proxy (loudness correlation, human ratings, or cross-instrument consistency) of velocity quality on real guitar audio is described.","section":"Evaluation / Results"},{"comment":"The weakest assumption—that statistical properties of synthetic velocity labels transfer meaningfully to real guitar—is stated but not tested, leaving generalization from virtual-instrument data unverified.","section":"Methodology / Discussion"}],"minor_comments":[{"comment":"The manuscript would benefit from explicit statements of the velocity definition used for guitar, details on the precise model architectures, and any ablation studies isolating the contribution of the velocity pretraining.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive comments, which highlight important aspects of our evaluation and assumptions. We address each major comment below, proposing revisions to strengthen the manuscript.","responses":[{"response":"We agree with this observation. The current abstract summarizes the results qualitatively. In the revised version, we will expand the abstract and evaluation description to include quantitative metrics (e.g., velocity prediction error), error bars from multiple runs, results of statistical tests, and a clear definition of how velocity is computed and scored for guitar, including the mapping from MIDI velocity values to synthesized audio intensity.","revision_made":"yes","referee_comment":"[Abstract] Abstract and evaluation description: outperformance on velocity is reported solely versus a non-pretrained baseline on synthetic test data, but no quantitative metrics, error bars, statistical tests, or definition of how velocity is computed or scored for guitar are provided."},{"response":"We acknowledge that our evaluation of the transferred velocity prediction relies on indirect evidence from note transcription improvements. Direct evaluation is challenging due to the absence of velocity annotations in real guitar datasets. We will revise the paper to explicitly state this limitation and discuss potential future proxies such as loudness correlations, while noting that the observed note transcription gains provide supporting evidence for retention of velocity information.","revision_made":"partial","referee_comment":"[Evaluation / Results] The central transfer claim (pretrained velocity weights remain functional after fine-tuning on real recordings) is supported only by indirect note-transcription metrics showing small and sometimes non-significant gains; no direct evaluation or proxy (loudness correlation, human ratings, or cross-instrument consistency) of velocity quality on real guitar audio is described."},{"response":"The assumption regarding the transfer of velocity statistics from synthetic to real data is indeed central and not directly tested in the current experiments. We will enhance the discussion section to better articulate the rationale behind this assumption with references to analogous transfer learning results in audio domains, and we will clearly label it as a limitation requiring further validation.","revision_made":"yes","referee_comment":"[Methodology / Discussion] The weakest assumption—that statistical properties of synthetic velocity labels transfer meaningfully to real guitar—is stated but not tested, leaving generalization from virtual-instrument data unverified."}],"tokens_in":1317,"tokens_out":521,"duration_ms":36516,"standing_objections":["Direct quantitative evaluation of velocity prediction quality on real guitar audio, as no ground-truth velocity labels exist for real recordings."]},"desk_editor":{"model":"grok-4.3","letter":"The paper introduces a straightforward transfer setup: pretrain on synthetic guitar audio that carries velocity labels from virtual instruments, then move those weights into a model fine-tuned on real recordings. This produces velocity output that beats a non-pretrained baseline on synthetic test data and yields a modest lift in note transcription accuracy.\n\nThe approach is new in the guitar AMT literature they cite. It directly tackles the lack of velocity labels without needing new real-data annotation, and the results stay comparable to existing guitar transcription systems. The method itself is clean and uses standard transfer techniques without obvious circular fitting.\n\nThe main limitation is the evaluation scope. Velocity gains are measured only on synthetic data; on real audio the paper reports only small note-level improvements that are not always statistically significant. No separate check on velocity quality for actual guitar (loudness correlation, listener ratings, or similar) is described, so the transfer of velocity statistics remains an assumption rather than a measured outcome. The abstract also omits concrete metrics and variance numbers.\n\nThis work is aimed at researchers extending AMT to additional attributes or instruments where labels are scarce. It is incremental but addresses a real gap with a reproducible pipeline. The reasoning is clear and the citations track prior synthetic-data and guitar work without overclaiming.\n\nI would send it to peer review. The contribution is concrete enough that referees can usefully comment on the transfer details and evaluation design.","headline":"Synthetic pretraining gives velocity prediction for guitar transcription with clear gains on synthetic tests but only small, indirect benefits shown on real audio.","tokens_in":2303,"tokens_out":355,"would_cite":false,"duration_ms":14988,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Pretraining on synthetic guitar data lets a transcription model predict note velocity on real audio.","keywords":["automatic guitar transcription","velocity prediction","synthetic data","transfer learning","polyphonic transcription","note intensity","virtual instruments"],"falsifier":"Both the transferred model and the baseline are evaluated on held-out synthetic guitar audio with known velocities; if the transferred model shows no lower velocity prediction error, the claim fails.","tokens_in":2607,"feed_emoji":"🎸","tokens_out":586,"duration_ms":23907,"temperature":0.7,"pith_summary":"The paper introduces a method to add velocity prediction to automatic guitar transcription despite the lack of intensity labels in real recordings. Virtual instruments generate synthetic audio with explicit velocity labels for pretraining. The resulting weights transfer to a second model trained on unlabeled real guitar audio, preserving the velocity head while adapting to actual instrument sounds. This transferred model beats a non-pretrained baseline at velocity estimation when both are tested on synthetic data and yields modest gains in note detection on some real test sets. The combined system reaches transcription accuracy on par with existing guitar-specific models while outputting velocity values.","feed_headline":"Synthetic pretraining adds velocity prediction to guitar transcription","feed_subtitle":"Weights learned from virtual instruments transfer to real audio, improving intensity estimates without hurting note accuracy.","key_machinery":"Weight transfer of a velocity prediction head pretrained on synthetic virtual-instrument data to a transcription network trained on real guitar audio.","core_discovery":"A model first trained on synthetic guitar data that includes velocity labels can have those weights transferred to a new network trained on real unlabeled guitar audio; the transferred model then predicts velocity more accurately than a baseline without pretraining when evaluated on synthetic test data, delivers small improvements to note transcription on some real test sets, and maintains overall performance comparable to the state of the art.","pith_inferences":["The same pretrain-and-transfer pattern could be applied to other instruments that lack velocity annotations.","If the acoustic mismatch between virtual and real instruments is reduced, velocity prediction accuracy on real audio may increase.","Downstream tools such as performance feedback systems could directly use the velocity output for expressive analysis."],"forward_implications":["Velocity prediction becomes possible in guitar transcription without requiring velocity-labeled real recordings.","The transferred model outperforms a non-pretrained baseline at velocity estimation on synthetic evaluation data.","Note transcription accuracy receives a small boost on certain real test sets when pretrained velocity weights are used.","Overall transcription performance stays comparable to existing state-of-the-art guitar transcription systems."],"fun_headline_variants":["Synthetic pretraining transfers velocity prediction to guitar transcription","Pretraining on virtual instruments adds guitar velocity prediction","Velocity prediction from synthetic guitar data transfers to real audio","Synthetic guitar pretraining enables velocity output on real data"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Statistical patterns of velocity learned from virtual instrument sounds remain useful when the model processes real guitar recordings.","fun_headline_variants_meta":{"raw":{"variants":["Synthetic pretraining transfers velocity prediction to guitar transcription","Pretraining on virtual instruments adds guitar velocity prediction","Velocity prediction from synthetic guitar data transfers to real audio","Synthetic guitar pretraining enables velocity output on real data"]},"model":"grok-4.3","cost_usd":0.011706,"raw_usage":{"total_tokens":5105,"prompt_tokens":630,"num_sources_used":0,"completion_tokens":59,"cost_in_usd_ticks":117062000,"prompt_tokens_details":{"text_tokens":630,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":4416,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":630,"tokens_out":59,"duration_ms":50445,"temperature":1.0,"reasoning_tokens":4416,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T12:59:03.712620+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Both the transferred model and the baseline are evaluated on held-out synthetic guitar audio with known velocities; if the transferred model shows no lower velocity prediction error, the claim fails.","supporting_citations":[],"review_version":1}