{"id":"c7e91749-3075-491f-ab9f-ec9495efeb0b","arxiv_id":"2608.03999","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Controlled experiments show that a 10ms performance-timed token stream lowers Frechet Music Distance roughly twofold versus beat-grid tokens, across model sizes from 0.8B to 27B.","lead":"A music-generation study trained the same language model at several sizes with different ways of encoding music and found that the encoding matters more than model size: a small model with fine-grained performance tokens matched or beat a large model with beat-grid tokens. The authors release the tokenizer, benchmark, and datasets so future claims about music representations can be tested directly.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Matched-budget scaling under-trains the 27B arms, so the flat FMD-vs-scale slopes and '0.8B beats 27B' headline may reflect convergence rather than representation; the 26M from-scratch control does not test a converged large beat-grid model.","rationale":"The reader's CONDITIONAL verdict is appropriate and should stand. I agree that the matched-budget scaling design is the load-bearing weakness. The paper's controls (from-scratch backbone, second performance-resolution tokenizer, embedder swap, source-stratified reference, onset snapping, model-free ceiling) substantially support a real representation effect, and the 0.8B FMD gap is not in doubt. But the paper's strongest headline extends a measured point into a scaling law. The scale curves show every arm's slope within noise of flat over 0.8B-27B, yet all arms share 10k steps; large models are systematically less converged at that budget, so the curve shape is confounded with convergence. The 26M from-scratch control, while well-converged, does not vary scale, so it cannot rule out that a converged 27B beat-grid model would close the gap. This is exactly the condition under which the central claim would fail, and it is testable with the released checkpoints by training 27B beat-grid arms longer. No internal inconsistency is alleged; the concern is that the evidence underdetermines the headline. Hence the verdict is unchanged: conditional until the convergence-matched large-model baseline is added or the abstract claim is softened.","tokens_in":62489,"tokens_out":5312,"duration_ms":49756,"concrete_test":"Train the 27B REMI (or Beat-TSD) arm with the same Qwen3.5 recipe and data for at least 50k steps, or until validation NLL plateaus, and evaluate FMD under the identical 100-caption protocol. If FMD stays above the PMT band (152-159), the representation claim survives; if it falls below ~159, the matched-budget design was the binding confound. A cheaper intermediate check is to report FMD at 20k/30k/50k checkpoints for the 27B beat-grid arm to show whether the flat slope is a plateau or a slow descent. Because the checkpoints are released, this can be run without retraining the 0.8B cells.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central Pareto claim ('representation, not model size, is the binding variable') rests on the FMD-vs-scale curves in Scaling Behavior, which are measured at a single matched budget: 10k steps, about 1.9 epochs, effective batch 16 (Implementation details). At 0.8B that budget may be near-converged, but a 27B model is far from converged at 1.9 epochs, so the flat slope could reflect undertraining at the top of the range rather than representation dominance. The paper's own from-scratch 26M control is well-converged, but at a much smaller capacity it cannot test whether a converged 27B beat-grid model would close the ~125-point FMD gap (PMT 159 vs REMI 272). The body carefully limits the claim to 'no closing trend within range'; the abstract's '34x parameter increase does not overturn the ordering' and 'representation, not model size' go beyond that evidence. The effect is not internally inconsistent, and the controls for embedder, reference, and onset snapping are strong; the weakest load-bearing premise is that the scale comparison isolates representation. A longer-trained large beat-grid arm could falsify the headline.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces PMT, a performance-timed symbolic music tokenizer (609 symbols; 10 ms timing, 32-level velocity, multi-track program structure), and evaluates it through a controlled representation swap over Qwen3.5 backbones from 0.8B to 27B, holding data, training budget, and decoding fixed across seven tokenization families. The central claim is that representation, not model size, is the binding variable for distributional fidelity as measured by Frechet Music Distance (FMD): PMT reaches FMD 152-159 across scales while beat-grid tokenizers sit at 272-286, with non-overlapping bootstrap confidence intervals. The paper includes extensive controls, including an embedder swap to CLaMP-3, coarse-quantization of onsets, a from-scratch 26M backbone, a second performance-resolution tokenizer (PerTok), ceiling-anchored evaluation, and an imprinting diagnostic showing that published text-to-MIDI systems are near-invariant to captions. It also releases two corpora, checkpoints, and a benchmark harness, and explicitly limits perceptual claims pending a pre-registered human study.","tokens_in":62864,"tokens_out":5327,"duration_ms":54688,"significance":"If the central claim holds, this is a strong and valuable result: it would establish an isolated representation effect in a domain where representation choices are usually entangled with backbone, data, and recipe, and it provides reusable datasets and a controlled benchmark for future work. The paper's strengths include unusually careful controls: velocity is held at 32 levels so the swap isolates timing quantization; the CLaMP-3 embedder swap and the reference swap test the headline FMD; the coarse-quantization control (Table 17) separates representability from learned timing; and the from-scratch 26M MAESTRO reproduction reduces pretraining-prior concerns. The round-trip exactness proof (Proposition 1) and the bootstrap intervals are also good. The main weakness is that the scaling argument underpinning the 'representation, not model size' claim is measured at a single matched budget that likely under-trains the large arms.","major_comments":[{"comment":"The central Pareto claim ('representation, not model size, is the binding variable for distributional fidelity') is measured at a single matched budget of 10k steps, about 1.9 epochs, with effective batch 16. At 0.8B this budget may be near convergence, but a 27B model is far from converged at 1.9 epochs, so the flat FMD-versus-scale slopes and the '0.8B beats 27B' headline cannot distinguish representation dominance from undertraining of the large beat-grid arms. The paper's own 26M from-scratch MAESTRO control is well-converged and does reproduce the ordering, but it is a 26M model; it does not test whether a converged 27B beat-grid model would close the roughly 125-point FMD gap (PMT 159 versus REMI 272). The body carefully restricts the claim to 'no closing trend within range' (Appendix K, Figure 13), but the abstract states the stronger '34x parameter increase does not overturn the ordering' and 'representation, not model size.' To support the headline claim, please train at least one large beat-grid arm to a convergence criterion (or a compute-matched budget with convergence diagnostics such as held-out loss plateaus for each arm), or explicitly soften the abstract and conclusion to the within-range statement. This is load-bearing because the scaling behavior is the main evidence for the central claim.","section":"Scaling Behavior / Implementation details"},{"comment":"There is a mismatch between the evidence and the abstract's causal phrasing. The manuscript's own scaling analysis concludes that the defensible statement is 'there is no closing trend within range' and explicitly declines to extrapolate a crossover from a flat fit over 1.5 decades. However, the abstract says 'representation, not model size, is the binding variable' and 'a 34x parameter increase does not overturn the ordering.' The latter is a statement about the measured points, while the former is a broader causal claim about what sets the ceiling. Since the matched-budget design cannot rule out convergence effects at the large end, the causal wording goes beyond what the controlled data establish. Please either add the convergence check requested above or rephrase the abstract and conclusion to match the within-range claim, which is already carefully worded in the body.","section":"Abstract / Conclusion"}],"minor_comments":[{"comment":"The abstract contains missing spaces in the extracted text (for example, 'Buildingatext-to-musiclanguagemodelbeginswithachoice usually made by default'); the camera-ready version should be checked for such spacing errors throughout.","section":"Abstract"},{"comment":"The naming of MIDI-Like, MIDI-LLM, and MIDILM is confusing because three similar names appear close together. The footnote helps, but consider renaming the MidiTok-based arm (for example, 'MIDI-TSD') to reduce reader load and avoid typographical mix-ups in later tables.","section":"Tables 1, 2, 24"},{"comment":"Appendix E's analysis of the decode-budget and length bias of FMD is important and clearly reported; a one- or two-sentence version of the 'longer is worse under FMD only because it is longer' finding should be moved into the main text where FMD is first introduced, since readers may otherwise interpret the 900-token budget as a neutral protocol choice.","section":"Evaluations / FMD caveats"}],"recommendation":"major_revision","confidential_remarks":"This is a carefully executed empirical paper and I found no evidence of circularity: the PMT design parameters are ablation-justified rather than fitted to the target result. My recommendation hinges on the matched-budget scaling issue; if the authors add a convergence check for a large beat-grid arm or align the abstract with the within-range evidence, I would support acceptance. The paper is well within the journal's scope and the released artifacts are a significant contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The controlled representation swap is the real contribution. They fix backbone, data, budget, and decoding, and show PMT (performance-resolution tokens) beats beat-grids on FMD across 0.8B–27B, with controls that survive embedder swaps, reference swaps, coarse-quantization, a from-scratch 26M backbone, and a second performance tokenizer. That is the first cross-family comparison I've seen that isolates representation this cleanly, and the paper is admirably honest about its own caveats: chord-time is mixture-confounded, the perceptual claim is explicitly left open, and caption adherence is weak. The released harness, aligned 86.6k corpus, 6.25M captioned corpus, and 25+ checkpoints are genuine artifacts the field can use, and the 'imprinting diagnostic' is a nice tool for exposing caption-invariant generation. The PMT tokenizer itself is a clean extension of Oore et al., with a formal round-trip guarantee and about 4 tokens/note.\n\nThe weakest link is the scale headline. The FMD-vs-scale curves are measured at a single matched budget: 10k steps, about 1.9 epochs, effective batch 16. At 0.8B that budget is close to convergence; a 27B model is nowhere near it. The body's careful language is 'no closing trend within range,' but the abstract says 'representation, not model size' and 'a 0.8B performance-resolution model beats a 27B beat grid.' Those are two different claims. The 26M from-scratch control rules out pretraining-prior confounds, but it is a small converged model and does not test whether a converged 27B beat-grid model would erode the gap. I don't think the central representation finding is in doubt—the gap is consistent across many controls—but the 'representation beats scale' framing is stronger than the evidence supports. Fixing it is straightforward: either add a convergence-matched large-model run or soften the abstract to match the body.\n\nMinor points: the FMD metric is protocol-sensitive (the paper admits ratios range 1.6–2.8x depending on reference), and fine-grained ordering among beat grids is within noise. Both are handled honestly in the text. The chord-time cancellation is openly acknowledged, and the per-source FMD breakdown shows the advantage is not just a folk artifact.\n\nBottom line: this paper deserves a serious referee. The empirical design is the strongest I've seen in this area, the artifacts are reusable, and the main finding is plausible and well-supported. The scale interpretation needs revision, but that's a revision, not a rejection. I'd bring it to a reading group and would cite the harness and PMT tokenizer if I worked in symbolic music generation.","headline":"Strong controlled study, but the scale headline is over-reached; the representation effect is real, the 'beats 27B' claim needs a convergence-matched baseline.","tokens_in":63322,"tokens_out":1639,"would_cite":true,"duration_ms":17890,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Holding backbone, data, budget, and decoding fixed, this paper shows that the token representation, not model scale, sets the distributional fidelity of text-to-music generation: a 0.8B performance-resolution model beats a 27B beat-grid…","keywords":["text-to-music generation","symbolic music","tokenization","performance timing","Fréchet Music Distance","controlled evaluation","multi-track MIDI","LLM vocabulary extension"],"falsifier":"Train the 27B beat-grid arm (or any large beat-grid model) to convergence under the same protocol and recompute FMD on the same 100 frozen captions; if a converged large beat-grid model reaches FMD at or below PMT's 159 at 0.8B, the claim that representation Pareto-dominates scale for distributional fidelity is overturned. A listener-level falsifier is the pre-registered forced-choice human study: a null result on block A (PMT vs. Beat-TSD at matched backbone) would bound the claim to the distributional level, which the paper already accepts.","tokens_in":62275,"feed_emoji":"🎵","tokens_out":4596,"duration_ms":35173,"temperature":0.7,"pith_summary":"This paper isolates one variable that text-to-music systems normally change together with everything else: how a score is turned into tokens. The authors hold the backbone, data, training budget, and decoding fixed, swap only the representation across seven tokenizations, and measure each arm's distributional distance to real music, normalized by what each format can express at all. They find that representation, not model scale, is the binding variable: a performance-resolution tokenizer with 10 ms timing reaches FMD 159 at 0.8B, while beat-grid tokenizers sit at 272–286 even at 27B, so a 0.8B model using the better representation beats a 27B model using a beat grid. The ordering reappears on a from-scratch 26M backbone and with a second performance-resolution tokenizer, which the paper reads as evidence that the effect belongs to the representation class. The paper is explicit that the advantage is distributional; whether listeners hear it is left to a pre-registered human study that is still collecting data.","feed_headline":"Representation, not scale, sets music-generation quality","feed_subtitle":"A controlled swap shows the token representation, not model size, decides distributional fidelity to real music.","key_machinery":"The load-bearing instrument is the controlled swap combined with ceiling-anchored evaluation. Seven tokenizations—performance-resolution (PMT and PerTok), beat-grid (Beat-TSD, REMI, MIDI-Like), grouped (Structured), and text (ABC)—are trained under identical backbone, data, budget, and decoding, with every texture metric divided by each representation's own model-free round-trip ceiling so that what a format can express is separated from what the model learns to use. The representation itself, PMT, serializes each note as an interpretable parameter tuple: pitch, duration, velocity on 32 levels, and inter-onset time-shift on a fixed 10 ms lattice, with track and program tokens emitted only at instrument changes, yielding about 4 tokens per note. That fixed-lattice time-shift is the single variable that the swap isolates, because beat-grid tokenizers place onsets on a tempo-relative grid whose step is coarser and tempo-dependent.","core_discovery":"The paper's central claim is that when backbone, data, budget, and decoding are fixed, the symbolic representation chosen for music determines distributional fidelity more than model size does. Its evidence is a controlled swap in which only the tokenization changes. A performance-resolution representation—PMT, a 609-symbol vocabulary of pitch, duration, velocity, 10 ms time-shift, track, and program tokens at about four tokens per note—produces Fréchet Music Distance 159 at 0.8B, versus 272–286 for REMI, Beat-TSD, and MIDI-Like beat grids, with non-overlapping bootstrap intervals; scaling the backbone 34× leaves the gap roughly constant. The claim is deliberately scoped: the gap is distributional, survives coarse-quantizing onsets to the beat grids' own resolution, and is reproduced on a 26M from-scratch backbone, so the authors attribute it to the representation class, not to pretraining or vocabulary luck. They do not claim it is audible; the human listening study is pre-registered and still running.","pith_inferences":["If representation is the binding variable in music, the same controlled-swap lens should be applied to other structured artifacts an LLM is trained to emit (vector graphics, 3D scenes, character rigs), where tokenization choices are currently entangled with the model recipe.","One testable extension is whether the 10 ms timing resolution itself is the active ingredient or whether coarser sub-grid lattices (20 ms, 50 ms) retain most of the benefit; the paper's own ablation predicts a sharp drop at 20 ms, which could be probed directly.","The imprinting diagnostic suggests that evaluations should report distance to each domain's own real reference rather than absolute texture statistics; otherwise a model that ignores the caption can look good on its home distribution.","A human-study null result would not touch the distributional claim but would usefully bound when micro-timing matters, while a positive result would extend the claim from distributional statistics to perception."],"forward_implications":["A 0.8B performance-resolution model can replace a 27B beat-grid model for distributional fidelity, cutting compute 34× at the same quality level on this axis.","The representation gap is a property of the performance-resolution class, not one vocabulary: PerTok shows the same low-FMD corner and the 26M from-scratch backbone reproduces the ordering.","The effect is not a finer-lattice artifact: snapping PMT onsets to a 60 ms grid still leaves it about 67 FMD points ahead of both beat grids.","Caption adherence is weak but separable: a decode-time constraint doubles instrument-F1 and correct-key rates with no distributional cost, suggesting conditioning and fidelity are orthogonal axes.","ABC's compactness does not convert into generation quality: its low token count comes with the lowest round-trip ceiling and the worst FMD in the grid."],"supporting_citations":[{"why":"Supplies the REMI beat-grid tokenization that serves as a main baseline arm in the controlled swap.","marker":"Huang and Yang 2020"},{"why":"Provides the MidiTok TSD/Beat-TSD tokenizer used for the beat-grid arms and the within-family comparison basis.","marker":"Fradet et al. 2021"},{"why":"Establishes the performance-event lineage that PMT extends with per-note velocity and multi-track program structure.","marker":"Oore et al. 2020"},{"why":"Contributes PerTok, the second performance-resolution tokenizer that corroborates the class-level claim.","marker":"Lenz and Mani 2024"},{"why":"Defines Fréchet Music Distance over CLaMP-2 embeddings, the primary distributional-fidelity metric.","marker":"Retkowski, Stępniak, and Modrzejewski 2024"},{"why":"Supplies CLaMP-3, used for caption alignment and as an independent embedder in the robustness check.","marker":"Wu et al. 2025"},{"why":"Provides MAESTRO, the from-scratch backbone dataset used to close the pretraining-prior confound.","marker":"Hawthorne et al. 2019"},{"why":"Contributes MidiCaps, the public benchmark used for zero-shot transfer and cross-system comparison.","marker":"Melechovsky, Roy, and Herremans 2024"},{"why":"Supplies MIDI-LLM, an external LLM-vocabulary baseline that the imprinting diagnostic exposes as caption-invariant.","marker":"Wu, Kim, and Huang 2025"},{"why":"Supplies text2midi, a published text-to-MIDI system used as an external baseline with its own training-domain imprint.","marker":"Bhandari et al. 2025"}],"fun_headline_variants":["Representation, not scale, dictates music fidelity","Token choice beats model size for symbolic music","Small model, right tokens: 34x scale can't match","PMT tokens halve music distance vs beat grids","Representation rules: 0.8B beats 27B on FMD"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison treats 10,000 training steps at effective batch 16 as a fair shared budget for a 0.8B and a 27B model; if the large beat-grid model is simply far from converged at that budget, the headline '0.8B beats 27B' could be an artifact of under-training the large arm, and the paper's well-converged 26M from-scratch control does not directly test a fully trained large beat-grid model.","fun_headline_variants_meta":{"raw":{"variants":["Representation, not scale, dictates music fidelity","Token choice beats model size for symbolic music","Small model, right tokens: 34x scale can't match","PMT tokens halve music distance vs beat grids","Representation rules: 0.8B beats 27B on FMD"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000205,"raw_usage":{"total_tokens":1518,"prompt_tokens":1195,"completion_tokens":323,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":811,"completion_tokens_details":{"reasoning_tokens":241}},"tokens_in":811,"tokens_out":323,"duration_ms":3105,"temperature":1.0,"reasoning_tokens":241,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:44:13.052556+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the 27B beat-grid arm (or any large beat-grid model) to convergence under the same protocol and recompute FMD on the same 100 frozen captions; if a converged large beat-grid model reaches FMD at or below PMT's 159 at 0.8B, the claim that representation Pareto-dominates scale for distributional fidelity is overturned. A listener-level falsifier is the pre-registered forced-choice human study: a null result on block A (PMT vs. Beat-TSD at matched backbone) would bound the claim to the distributional level, which the paper already accepts.","supporting_citations":[{"cited_title":"2020 , organization=","cited_arxiv_id":null,"evidence_quote":"Supplies the REMI beat-grid tokenization that serves as a main baseline arm in the controlled swap."}],"review_version":1}