{"id":"b172cd84-8291-45d9-84fb-c0d1f2150f3f","arxiv_id":"2508.09790","paper_version":2,"verdict":"REJECT","confidence":"LOW","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"The abstract claims BeatFM achieves state-of-the-art beat tracking, but the body describes a different model, HingeNet, so the BeatFM claim is unsupported.","lead":"The abstract promises BeatFM, a beat tracking method that attaches a multi-domain aggregation module to a frozen music foundation model and claims state-of-the-art results. The full text instead presents HingeNet, a different architecture, so the submission's core claim has no supporting evidence in the body.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"BeatFM abstract and full text are different papers: body describes HingeNet (arXiv 2508.09788), so the claimed BeatFM paradigm, aggregation module, and SOTA results have no supporting derivation or experiments.","rationale":"The reader's weakest_assumption—that the body text must actually describe BeatFM for the abstract's claims to be evaluable—is exactly the load-bearing issue. The provided materials show a clear internal inconsistency: the abstract describes BeatFM, while the full text is a different paper, HingeNet, with a different title, a different arXiv identifier, and a different architecture. This is not a matter of consensus or interpretation; it is a factual mismatch that prevents any assessment of BeatFM's correctness, novelty, or experimental support. There is no derivation of the multi-dimensional semantic aggregation module, no ablation isolating its three sub-modules, and no benchmark comparison attributed to BeatFM. The HingeNet body may itself be a coherent paper, but it does not support the claims in the BeatFM abstract. Therefore the submitted manuscript cannot be accepted or conditionally accepted; rejection is appropriate, with the possibility of resubmission after correcting the title/abstract/body mismatch. I agree with the reader's identification of this problem and see no additional independent concern that would change the verdict.","tokens_in":7458,"tokens_out":2436,"duration_ms":28424,"concrete_test":"Extract the actual submitted PDF for arXiv:2508.09790 and search the body after the abstract for the strings 'BeatFM' and 'multi-dimensional semantic aggregation', specifically for three parallel sub-modules addressing temporal, frequency, and channel domains. If the body contains neither the name BeatFM nor any such module, and the title page identifies the paper as HingeNet, then the BeatFM claim is unsupported. Additionally, verify whether any experiment in Tables I–III reports results for BeatFM; if all rows correspond to HingeNet, the abstract's SOTA claim is unsubstantiated.","verdict_should_be":"REJECT","load_bearing_attack":"The central claim is that BeatFM—a pre-trained music foundation model plus a plug-and-play multi-dimensional semantic aggregation module with parallel temporal, frequency, and channel sub-modules—achieves state-of-the-art beat and downbeat tracking. The abstract alone asserts this. The full text is titled 'HingeNet: A Harmonic-Aware Fine-Tuning Approach for Beat Tracking', carries arXiv number 2508.09788, and describes a different architecture: projection layers, a learnable gating mechanism, and harmonic-aware modules built from dilated 1D convolutions. Nowhere in the body is 'BeatFM' defined, and the described three parallel semantic aggregation sub-modules do not appear. Consequently, the abstract's method has no architectural specification, no training details, and no experimental tables. The numerical results in Tables I–III are attributed to HingeNet, not BeatFM, so they cannot substantiate BeatFM's SOTA claim. Even if HingeNet is a valid contribution, this submission as written does not contain the paper promised by its abstract.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript as submitted consists of an abstract that announces a novel beat-tracking paradigm called BeatFM, which introduces a pre-trained music foundation model together with a plug-and-play multi-dimensional semantic aggregation module (temporal, frequency, and channel sub-modules), and claims state-of-the-art results on multiple benchmark datasets. The body text, however, is a different paper titled 'HingeNet: A Harmonic-Aware Fine-Tuning Approach for Beat Tracking' and carries arXiv number 2508.09788 rather than 2508.09790. The body defines HingeNet with projection layers, a learnable gating mechanism, and harmonic-aware modules in Eqs. (1)–(4), and reports experiments only for MERT+HingeNet and MusicFM+HingeNet in Tables I–III. The terms 'BeatFM' and 'multi-dimensional semantic aggregation' do not appear anywhere in the body. Consequently, the method promised in the abstract has no architectural specification, no training details, and no experimental evaluation in this document.","tokens_in":7704,"tokens_out":3208,"duration_ms":33865,"significance":"The claimed contribution—BeatFM, combining a frozen music foundation model with a multi-domain semantic aggregation module—would be a relevant and potentially valuable contribution to beat and downbeat tracking if the architecture and experiments were actually presented. The full-text HingeNet experiments (Tables I–III) may indicate a promising fine-tuning approach in their own right, but they do not support the BeatFM paradigm described in the abstract. The manuscript contains no machine-checked proofs, released code, or parameter-free derivations that could offset the absence of the claimed method. As submitted, the central claims of the abstract are not assessable because the body describes a different method and reports results for a different model. The significance of the submission is therefore not supported by its content.","major_comments":[{"comment":"The abstract proposes 'BeatFM' with a 'plug-and-play multi-dimensional semantic aggregation module' composed of temporal, frequency, and channel sub-modules. The body is titled 'HingeNet: A Harmonic-Aware Fine-Tuning Approach for Beat Tracking' and carries arXiv number 2508.09788. The method section (Sec. III) defines HingeNet via projection layers, a learnable gating mechanism, and harmonic-aware modules in Eqs. (1)–(4). The terms 'BeatFM', 'multi-dimensional semantic aggregation', and the three parallel sub-modules do not appear in the body. Thus the claimed method has no architectural definition, no training details, and no analysis in this manuscript.","section":"Abstract; Full Text Title"},{"comment":"The experimental section reports results for 'MERT+HingeNet' and 'MusicFM+HingeNet' on GTZAN, Ballroom, Hainsworth, and SMC. These tables cannot substantiate the abstract's state-of-the-art claim for BeatFM, because the evaluated models are HingeNet variants. No table or figure reports results for a 'BeatFM' model, and there is no ablation of the multi-dimensional semantic aggregation module. The central empirical claim of the abstract is therefore unsupported by the evidence contained in the document.","section":"Tables I–III; Section IV"},{"comment":"The Introduction, Method, Experiments, and Conclusion consistently discuss only HingeNet and its harmonic-aware mechanism. There is no passage connecting the abstract's BeatFM paradigm to the body. The discrepancy is not a terminological inconsistency: the two descriptions differ in architectural components, evaluation objects, and paper title. Even if HingeNet is a valid contribution, this submission as written does not contain the paper promised by its abstract.","section":"Abstract vs. Full Text (global)"}],"minor_comments":[{"comment":"Equation (4) has an extra closing parenthesis: '... H^M_i(\\tilde{h}_i)])))' should be '... H^M_i(\\tilde{h}_i)])' or similar. The notation for the dimension of h_i also relies on the implicit compatibility of h_{i-1} and h^P_i; making this explicit would improve readability.","section":"Section III.A, Eq. (4)"},{"comment":"Table I uses 'Downbeat Accuracy' with F-Measure, CMLt, and AMLt subcolumns, while Table II uses 'Downbeat F1'. For consistency, the metric naming should be unified across tables.","section":"Tables I–II"},{"comment":"The text refers to 'Music2Vce [25]' but the cited reference [25] is titled 'Map-Music2Vec'. Please correct the name to match the reference.","section":"Related Work, Section II.B"}],"recommendation":"reject","confidential_remarks":"The mismatch between the abstract and the full text is a fundamental problem: the submission does not describe the work claimed in its abstract. This is not a matter of local revision; the entire method and experimental sections would need to be replaced with the actual BeatFM content. If the authors intended to submit the HingeNet paper, they should do so under its own title and identifier. I recommend rejection of the current submission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe short version: the abstract and the body are two different papers. The abstract proposes BeatFM, a frozen foundation model plus a plug-and-play multi-dimensional semantic aggregation module with parallel temporal, frequency, and channel sub-modules, and claims state-of-the-art results. The body is titled “HingeNet: A Harmonic-Aware Fine-Tuning Approach for Beat Tracking,” carries arXiv 2508.09788, and never mentions BeatFM. The architecture described—projection layers, learnable gating, dilated 1D harmonic-aware modules—does not match the abstract. Tables I–III report HingeNet results, so the abstract’s SOTA claim has zero support in the text.\n\nWhat’s good: the HingeNet paper, taken on its own, looks like a credible contribution. It proposes a lightweight, separable fine-tuning network that takes intermediate features from a frozen MERT or MusicFM and applies harmonic-aware modules with a learnable gating mechanism. The ablation against Adapter and LoRA shows clear gains, and the results on GTZAN, Ballroom, Hainsworth, and SMC are competitive. The related work is properly cited. If I received that manuscript under its own title, I’d be inclined to send it to review.\n\nThe soft spot is fatal as submitted. This is not a weak argument; it’s a case where the delivered manuscript does not contain the claimed paper. The BeatFM method is only named in the abstract—no equations, no architecture diagram, no training details, no experiments. Even if HingeNet is solid, it is not what is being advertised, and there is no explanation for the mismatch.\n\nThe bottom line: this deserves a desk reject, not because the idea is weak, but because the submission is internally inconsistent. The authors should resubmit the correct full text. If that is the HingeNet paper, it deserves a serious referee. If they can actually describe BeatFM and show its results, that also deserves a look. But this version is not reviewable.\n\nMy recommendation: return it without peer review and ask for a matching manuscript. I would not cite this version.\n\nBest,\n\n[You]","headline":"Abstract and full text are two different papers—BeatFM is claimed, HingeNet is written; not reviewable as is.","tokens_in":8119,"tokens_out":2751,"would_cite":false,"duration_ms":27301,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims a frozen music foundation model plus a lightweight side network reaches state-of-the-art beat and downbeat tracking, but the abstract and body describe different architectures.","keywords":["beat tracking","downbeat tracking","music foundation model","parameter-efficient fine-tuning","harmonic-aware module","learnable gating","frozen encoder","music information retrieval"],"falsifier":"Open the body's model description and the released code: the body implements projection layers, a learnable gate, and harmonic-aware dilated convolutions, and every reported number comes from that design; the abstract promises parallel temporal, frequency, and channel aggregation sub-modules. If the code and checkpoints match the body rather than the abstract, the abstract's architecture claim has no experimental support. Separately, re-running the GTZAN fine-tuning ablation on identical MERT features would verify whether the roughly eight-point beat-F1 gap over LoRA is real.","tokens_in":7400,"feed_emoji":"🥁","tokens_out":15879,"duration_ms":161202,"temperature":0.7,"pith_summary":"The paper sets out to show that beat tracking can escape its label-scarcity bottleneck without inserting extra modules into a pre-trained model. Its proposal is to keep a music foundation model completely frozen and to attach a small, separately trainable network that reads the model's intermediate representations, converting them into beat and downbeat predictions; the abstract calls this BeatFM, with a plug-and-play module aggregating semantic features in the temporal, frequency, and channel domains, while the full text describes HingeNet, whose core layers use dimensionality-reducing projections, learnable gating, and harmonic-aware dilated convolutions. The paper claims this recipe reaches state-of-the-art beat and downbeat accuracy on GTZAN, Ballroom, Hainsworth, and SMC, and that it beats both adapter- and LoRA-style fine-tuning, which the authors report overfit on scarce beat labels. The experiments and ablations in the body belong to the HingeNet description; the abstract's architecture is not the one tested, so a careful reader must treat the abstract's specific claims as unverified by the supplied evidence.","feed_headline":"Bolt a side net onto a frozen music model to set beat-tracking records","feed_subtitle":"Keeping the pre-trained encoder frozen avoids overfitting on scarce beat labels, beating adapter- and LoRA-style fine-tuning.","key_machinery":"The load-bearing object is the HingeNet core layer, a lightweight side network named for its hinge-like silhouette, attached to every encoder of a frozen music foundation model (MERT or MusicFM). Each core layer combines a projection layer that shrinks the encoder output by a factor $r$ to cap the trainable parameter count; a learnable gate $\\mu_i=\\mathrm{sigmoid}(\\alpha_i)$ that fuses the previous core layer's output with the newly projected features; and a harmonic-aware module (HAM) of $M$ parallel 1D convolutional layers with differing dilation rates, concatenated and mapped by an MLP back to the original dimensionality. The design's bet is that the frozen model's intermediate features c","core_discovery":"On the paper's own terms, the discovery is that a frozen pre-trained music foundation model can be adapted to beat and downbeat tracking by a lightweight, separable network that consumes each encoder layer's intermediate features. Each core layer projects those features down by a factor of $r$, blends them with the previous layer's output through a learnable gate, and runs them through a harmonic-aware module of parallel 1D dilated convolutions whose concatenated outputs return through an MLP; a linear layer plus dynamic Bayesian network post-processing yields the beats. The paper reports that this recipe, with MERT or MusicFM as the frozen encoder, beats Adapter and LoRA fine-tuning and rea","pith_inferences":["A natural next experiment the paper leaves implicit: unfreeze part of the encoder and compare — if performance drops, the frozen-representation design, not extra capacity, carries the gains.","The harmonic-shift assumption is independently testable: check whether harmonic-change onsets in the audio align with predicted beats better than generic energy onsets do, rather than judging only end-to-end F1.","Because the abstract's multi-domain aggregation module is never implemented in the body's experiments, the benchmark numbers should be attributed to the projection-gating-HAM design only; a reader cannot infer how the temporal-frequency-channel module would score from this submission.","If the recipe generalizes, the same side-network pattern should transfer to other label-scarce frame-level MIR tasks, such as tempo estimation or structural segmentation, which share the mismatch between abundant pre-training data and scarce annotation."],"forward_implications":["The frozen-encoder recipe transfers across backbones: pairing HingeNet with either MERT or MusicFM improves beat and downbeat accuracy over each backbone alone and over the compared state-of-the-art models on GTZAN, Ballroom, Hainsworth, and SMC.","Insertion-style fine-tuning is the loser on scarce data: on GTZAN, MERT+HingeNet reaches 89.7 beat F1 against 73.4 for Adapter and 81.2 for LoRA, which the paper attributes to overfitting.","Harmonic structure is doing real work: adding the harmonic-aware module improves beat F1 at every tested projection factor, with the best trade-off at $r = 6$.","Beat tracking need not modify the foundation model's weights or architecture, keeping the pre-trained representation space intact while adding only the side network's parameters."],"supporting_citations":[{"why":"Side-tuning precedent: an independent additive side network that reads intermediate features instead of inserted modules, the design basis for HingeNet.","marker":"[16]"},{"why":"Evidence that insertion-based fine-tuning loses effectiveness when annotated data is scarce, the paper's core motivation for a separate fine-tuning network.","marker":"[19]"},{"why":"MERT, the frozen music foundation model whose intermediate features drive the main experiments.","marker":"[26]"},{"why":"MusicFM, the second frozen foundation model used to show HingeNet generalizes across backbones.","marker":"[27]"},{"why":"The dynamic Bayesian network state-space model used as post-processing to turn network outputs into final beat and downbeat predictions.","marker":"[29]"},{"why":"Beat transformer, the primary state-of-the-art baseline that all datasets compare against.","marker":"[14]"},{"why":"Beat This!, a recent beat-tracking baseline that removes DBN post-processing and is a head-to-head comparison on GTZAN.","marker":"[24]"},{"why":"Closest prior use of pre-trained representation models as feature extractors for beat tracking; the comparison point the paper extends.","marker":"[21]"}],"fun_headline_variants":["Frozen music model plus skinny net out-tracks fine-tuned rivals","BeatFM: freeze the music AI, add a small net, beat-tracking SOTA","Frozen audio foundation beats fine-tuning for beat tracking","Lightweight net on frozen music model outperforms adapters in beat tracking","Frozen backbone + side net: beat tracking SOTA without fine-tuning"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The whole argument rests on the submission's abstract and body describing the same method: the abstract promises a multi-domain aggregation module, while the body implements a harmonic-aware projection-and-gating side network, and only the latter is exercised by the reported experiments.","fun_headline_variants_meta":{"raw":{"variants":["Frozen music model plus skinny net out-tracks fine-tuned rivals","BeatFM: freeze the music AI, add a small net, beat-tracking SOTA","Frozen audio foundation beats fine-tuning for beat tracking","Lightweight net on frozen music model outperforms adapters in beat tracking","Frozen backbone + side net: beat tracking SOTA without fine-tuning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000595,"raw_usage":{"total_tokens":2589,"prompt_tokens":675,"completion_tokens":1914,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":419,"completion_tokens_details":{"reasoning_tokens":1818}},"tokens_in":419,"tokens_out":1914,"duration_ms":15317,"temperature":1.0,"reasoning_tokens":1818,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:47:40.376604+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Open the body's model description and the released code: the body implements projection layers, a learnable gate, and harmonic-aware dilated convolutions, and every reported number comes from that design; the abstract promises parallel temporal, frequency, and channel aggregation sub-modules. If the code and checkpoints match the body rather than the abstract, the abstract's architecture claim has no experimental support. Separately, re-running the GTZAN fine-tuning ablation on identical MERT features would verify whether the roughly eight-point beat-F1 gap over LoRA is real.","supporting_citations":[{"cited_title":"Side-tuning: a baseline for network adaptation via additive side networks,","cited_arxiv_id":null,"evidence_quote":"Side-tuning precedent: an independent additive side network that reads intermediate features instead of inserted modules, the design basis for HingeNet."},{"cited_title":"Mert: Acoustic music understanding model with large-scale self- supervised training,","cited_arxiv_id":null,"evidence_quote":"MERT, the frozen music foundation model whose intermediate features drive the main experiments."},{"cited_title":"A foundation model for music informatics,","cited_arxiv_id":null,"evidence_quote":"MusicFM, the second frozen foundation model used to show HingeNet generalizes across backbones."},{"cited_title":"An efficient state- space model for joint tempo and meter tracking.,","cited_arxiv_id":null,"evidence_quote":"The dynamic Bayesian network state-space model used as post-processing to turn network outputs into final beat and downbeat predictions."}],"review_version":1}