Pith. sign in

REVIEW 2 major objections 2 minor 39 references

Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation

T0 review · 2 major / 2 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper claims that generating songs from human-editable bar-level symbolic scores, rather than from raw audio, makes a small model surpass commercial systems like Suno.

desk verdict The abstract makes an important claim about symbolic-score song generation, but the corrupted full text leaves the evidence unverifiable; ask for a clean copy before sending it out. read the letter →

arxiv 2508.01394 v1 pith:WXCE4BRX submitted 2025-08-02 cs.SD cs.AIeess.AS

classification cs.SDcs.AIeess.AS
keywords symbolicmusicgenerationsongbar-leveltokenizationhierarchicalhuman-editablescorescontrollabilityAIGCscore-based
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper argues that the four persistent problems of song generation—controllability, generalizability, perceptual quality, and long duration—persist because models try to learn musical structure directly from raw audio. It proposes BACH, a system that generates songs from human-editable bar-level symbolic scores, using a tokenization strategy and a hierarchical generative procedure built for song structure. The paper reports that BACH, at small model size, establishes a new state of the art among publicly reported song generation systems and even surpasses the commercial system Suno. A sympathetic reader should care because the claim is causal: the paradigm, not model scale, is what limits today's song generators, and the proposed alternative is testable and human-editable.

What carries the argument

The load-bearing machinery is BACH's bar-level symbolic tokenization combined with a hierarchical generative procedure. The tokenizer turns a song into discrete bar-level score tokens that preserve hierarchical structure—sections, bars, chords, and notes—so the generative model operates on explicit musical symbols rather than raw audio. This carries the argument because it makes the score both human-editable and structurally transparent, which the paper identifies as the reason controllability and long-form coherence improve.

What would settle it

A preregistered listening and editing study comparing BACH and Suno on identical prompts, with blinded raters and matched rendering, would settle the performance claim; if BACH's preference advantage disappears when renderings are matched, the superiority cannot be attributed to the symbolic-score paradigm.

Watch

Extended reading notes

Core claim

The central discovery the paper claims is that bar-level symbolic score representation is enough to make a small song generator exceed the quality, duration, efficiency, and controllability of much larger raw-audio generators. BACH tokenizes music into bar-level units and generates songs in a hierarchy aligned with song structure, so the model does not have to infer tonality, rhythm, and form from audio. Under the paper's evaluations, this design outperforms all publicly reported song generation systems, including Suno, and human raters prefer it across multiple subjective metrics.

Load-bearing premise

The load-bearing assumption is that BACH's measured advantage over raw-audio systems is caused by the symbolic-score paradigm itself, not by the specifics of its evaluation, prompts, rendering, or raters.

Editorial extensions

If this is right

  • Song generation can be made human-controllable by editing symbolic scores rather than prompting audio models.
  • Long-form structure can be generated without massively larger models, since the score compresses the musical decisions the model must learn.
  • Smaller models with symbolic input can produce perceptually competitive songs, challenging the assumption that raw-audio scale is necessary.
  • The symbolic-score representation makes the generative process inspectable and editable, opening a practical path for iterative human-AI composition.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct extension would be to measure how much of BACH's quality survives the rendering step by comparing its symbolic outputs rendered through different synthesizers against a raw-audio baseline.
  • If the causal claim generalizes, combining symbolic score priors with neural audio rendering could improve controllability of existing raw-audio systems without abandoning their acoustic richness.
  • A testable prediction is that a raw-audio model of the same size, trained with explicit bar-level structural conditioning, would close part of the gap, which would locate the benefit in structure rather than tokenization.
  • The human-editable interface suggests an evaluation metric based on edit distance or user task success, not just listening preference, to capture the controllability advantage.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper proposes BACH (Bar-level AI Composing Helper), a symbolic-score-based song generation model that tokenizes bar-level notation and generates hierarchically structured scores, which are then rendered to audio. The abstract claims that BACH, despite a small model size, achieves state-of-the-art results across all publicly reported song-generation systems, surpassing commercial systems such as Suno, and that human evaluations confirm its superiority on multiple subjective metrics. The paper attributes existing limitations in controllability, generalizability, perceptual quality, and duration to the dominant paradigm of learning music theory directly from raw audio. The full text of the submission, however, is almost entirely corrupted mojibake, so the technical and experimental contents cannot be examined.

Significance. If the central claims were fully supported, the work would be significant: it would demonstrate that a small symbolic-score model can outperform much larger raw-audio generative models on perceptual quality and controllability, and it would offer a practical, human-editable approach to long song generation. The paper's framing that the raw-audio paradigm itself, rather than model scale, is the bottleneck is a falsifiable and potentially impactful thesis. However, the current manuscript provides no inspectable evidence: the abstract contains no quantitative results or protocol details, and the body is unreadable. Consequently, the significance cannot be assessed independently, and no reproducible code, data, or demos are visible from the artifact.

major comments (2)
  1. [Full text (all sections after the Abstract)] The entire body of the manuscript, from the first page after the abstract onward, is rendered as mojibake (e.g., sequences of '��������' characters), with no legible equations, tables, or prose. As a result, the tokenization strategy, generative procedure, model architecture, datasets, training details, and experiments cannot be checked. The central empirical claim that BACH establishes a new SOTA and surpasses Suno therefore rests entirely on the abstract's assertions, with no inspectable evidence. A clean, complete manuscript with the full experimental protocol is required before this claim can be evaluated.
  2. [Abstract, sentences 4–6] The abstract states that 'Experiments demonstrate that BACH, with a small model size, establishes a new SOTA among all publicly reported song generation systems, even surpassing commercial solutions such as Suno,' and that 'Human evaluations further confirm its superiority across multiple subjective metrics.' However, no quantitative scores, effect sizes, error bars, p-values, number of raters, or details of the comparison protocol are given, even in the abstract. Since Suno is a black box, the comparison's fairness depends on matched prompts, blinded listening, and a comparable rendering pipeline; without this information, the superiority claim and the causal attribution of the four limitations to the raw-audio paradigm are unsupported. The resubmission must include a detailed experimental section with these specifics.
minor comments (2)
  1. [Header] The embedded header 'arXiv:2508.01392v2 [cs.LG] 26 Jun 2026' does not match the paper's arXiv ID 2508.01394 and subject class cs.SD; please correct the header to avoid provenance confusion.
  2. [Formatting] The title, abstract, and body are not cleanly separated in the corrupted text; the resubmission should use a properly encoded PDF so that the reviewers can read the technical content.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity is demonstrable from the readable abstract; the full text is corrupted, so no derivation step can be shown to reduce to its own inputs.

full rationale

Only the abstract is readable; the body, equations, tables, and experimental protocol are corrupted. No load-bearing derivation chain can be inspected, and therefore no specific reduction of a prediction to a fitted parameter, no self-definitional equivalence, and no load-bearing self-citation chain can be exhibited. The abstract's claims are evaluated against external benchmarks and human raters, which is structurally an external test rather than a tautology. The residual concerns about the unverifiable Suno comparison and the unsupported causal attribution to the raw-audio paradigm are correctness and evidence-quality issues, not circularity. Under the hard rule that circularity may be flagged only when the paper's own equations or citations show the reduction, the honest finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

Built from the abstract only because the full text is unreadable (mojibake) and its header references a different arXiv ID. No free parameters, hyperparameters, or fitted constants could be audited, so the ledger lists only the abstract-level premises the headline claim leans on. Scores throughout this report are correspondingly low-confidence.

assumptions (3)
  • domain assumption The four persistent limitations (controllability, generalizability, perceptual quality, duration) stem primarily from learning music theory from raw audio.
    Stated in the abstract as the motivating diagnosis. It justifies the design choice to generate from symbolic scores, but it is asserted, not demonstrated, and rival explanations (model scale, data, evaluation design) are not controlled for in the abstract.
  • domain assumption Human subjective evaluation is a valid and sufficient ground for the SOTA claim.
    The superiority claim over Suno and all published systems rests on subjective ratings in the abstract; no rater count, blinding, prompt matching, or inter-rater agreement is reported in the abstract, and the evaluation section is unreadable.
  • domain assumption Symbolic-score-then-render preserves enough musical information for perceptual quality competitive with audio-trained models.
    The paper's whole bet is that a bar-level symbolic representation plus rendering can match or beat end-to-end audio models on perceived quality. The abstract asserts the outcome but does not show the rendering pipeline; unverifiable here.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation." pith.science (2026). https://pith.science/paper/WXCE4BRX

@misc{pith2026250801394,
  author       = {Pith},
  title        = {Pith review of: Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/WXCE4BRX}},
  note         = {Machine review of arXiv:2508.01394}
}
read the original abstract

Song generation is regarded as the most challenging problem in music AIGC; nonetheless, existing approaches have yet to fully overcome four persistent limitations: controllability, generalizability, perceptual quality, and duration. We argue that these shortcomings stem primarily from the prevailing paradigm of attempting to learn music theory directly from raw audio, a task that remains prohibitively difficult for current models. To address this, we present Bar-level AI Composing Helper (BACH), the first model explicitly designed for song generation through human-editable symbolic scores. BACH introduces a tokenization strategy and a symbolic generative procedure tailored to hierarchical song structure. Consequently, it achieves substantial gains in the efficiency, duration, and perceptual quality of song generation. Experiments demonstrate that BACH, with a small model size, establishes a new SOTA among all publicly reported song generation systems, even surpassing commercial solutions such as Suno. Human evaluations further confirm its superiority across multiple subjective metrics.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

39 extracted references · 17 canonical work pages

  1. [1]

    , " * write output.state after.block = add.period write newline

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    I.; Borsos, Z.; Engel, J.; Verzetti, M.; Caillon, A.; Huang, Q.; Jansen, A.; Roberts, A.; Tagliasacchi, M.; et al

    Agostinelli, A.; Denk, T. I.; Borsos, Z.; Engel, J.; Verzetti, M.; Caillon, A.; Huang, Q.; Jansen, A.; Roberts, A.; Tagliasacchi, M.; et al. 2023. Musiclm: Generating music from text. arXiv preprint arXiv:2301.11325

  4. [4]

    Bai, Y.; et al. 2024. Seed Music: ByteDance's neural music generation system. https://bytedance.com/research/seed-music

  5. [5]

    Borsos, Z.; Marinier, R.; Vincent, D.; Kharitonov, E.; Pietquin, O.; Sharifi, M.; Roblek, D.; Teboul, O.; Grangier, D.; Tagliasacchi, M.; and Zeghidour, N. 2023. AudioLM: A Language Modeling Approach to Audio Generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31: 2523--2533

  6. [6]

    Chen, Y.; Huang, L.; and Gou, T. 2024. Applications and Advances of Artificial Intelligence in Music Generation:A Review. ArXiv, abs/2409.03715

  7. [7]

    Civit, M.; Civit-Masot, J.; Cuadrado, F.; and Cuaresma, M. J. E. 2022. A systematic review of artificial intelligence-based music generation: Scope, applications, and future trends. Expert Syst. Appl., 209: 118190

  8. [8]

    Copet, J.; Kreuk, F.; Gat, I.; Remez, T.; Kant, D.; Synnaeve, G.; Adi, Y.; and D \'e fossez, A. 2023. Simple and controllable music generation. Advances in Neural Information Processing Systems, 36: 47704--47720

Show all 39 references
  1. [9]

    W.; Radford, A.; and Sutskever, I

    Dhariwal, P.; Jun, H.; Payne, C.; Kim, J. W.; Radford, A.; and Sutskever, I. 2020. Jukebox: A generative model for music. arXiv preprint arXiv:2005.00341

  2. [10]

    Ding, S.; Liu, Z.; wen Dong, X.; Zhang, P.; Qian, R.; He, C.; Lin, D.; and Wang, J. 2024. SongComposer: A Large Language Model for Lyric and Melody Composition in Song Generation. ArXiv, abs/2402.17645

  3. [11]

    Donahue, C.; Caillon, A.; Roberts, A.; Manilow, E.; Esling, P.; Agostinelli, A.; Verzetti, M.; Simon, I.; Pietquin, O.; Zeghidour, N.; et al. 2023. Singsong: Generating musical accompaniments from singing. arXiv preprint arXiv:2301.12662

  4. [12]

    Du, X.; Wang, Z.; Liang, X.; Liang, H.; Zhu, B.; and Ma, Z. 2023. Bytecover3: Accurate cover song identification on short queries. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1--5. IEEE

  5. [13]

    Défossez, A.; Copet, J.; Synnaeve, G.; and Adi, Y. 2022. High Fidelity Neural Audio Compression. arXiv:2210.13438

  6. [14]

    Fang, R.; Duan, C.; Wang, K.; Huang, L.; Li, H.; Yan, S.; Tian, H.; Zeng, X.; Zhao, R.; Dai, J.; et al. 2025. Got: Unleashing reasoning capability of multimodal large language model for visual generation and editing. arXiv preprint arXiv:2503.10639

  7. [15]

    Hong, Z.; Cui, C.; Huang, R.; Zhang, L.; Liu, J.; He, J.; and Zhao, Z. 2023. Unisinger: Unified end-to-end singing voice synthesis with cross-modality information matching. In Proceedings of the 31st ACM International Conference on Multimedia, 7569--7579

  8. [16]

    A.; Vaswani, A.; Uszkoreit, J.; Shazeer, N.; Simon, I.; Hawthorne, C.; Dai, A

    Huang, C.-Z. A.; Vaswani, A.; Uszkoreit, J.; Shazeer, N.; Simon, I.; Hawthorne, C.; Dai, A. M.; Hoffman, M. D.; Dinculescu, M.; and Eck, D. 2018. Music transformer. arXiv preprint arXiv:1809.04281

  9. [17]

    Kim, T.; and Nam, J. 2023. All-in-One Metrical and Functional Structure Analysis with Neighborhood Attentions on Demixed Audio. In 2023 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 1--5

  10. [18]

    Latif, S.; Shoukat, M.; Shamshad, F.; Usama, M.; Cuay'ahuitl, H.; and Schuller, B. 2023. Sparks of Large Audio Models: A Survey and Outlook. ArXiv, abs/2308.12792

  11. [19]

    Lei, J.; et al. 2025. SongCreator: Transformer-based text-to-song generation. In Proceedings of the International Conference on Machine Learning

  12. [20]

    Li, R.; Hong, Z.; Wang, Y.; Zhang, L.; Huang, R.; Zheng, S.; and Zhao, Z. 2024. Accompanied Singing Voice Synthesis with Fully Text-controlled Melody. CoRR

  13. [21]

    Liang, X.; Du, X.; Lin, J.; Zou, P.; Wan, Y.; and Zhu, B. 2024. ByteComposer: a Human-like Melody Composition Method based on Language Model Agent. ArXiv, abs/2402.17785

  14. [22]

    Liu, H.; Chen, Z.; Yuan, Y.; Mei, X.; Liu, X.; Mandic, D.; Wang, W.; and Plumbley, M. D. 2023. Audioldm: Text-to-audio generation with latent diffusion models. arXiv preprint arXiv:2301.12503

  15. [23]

    Liu, Z.; Ding, S.; Zhang, Z.; Dong, X.; Zhang, P.; Zang, Y.; Cao, Y.; Lin, D.; and Wang, J. 2025. SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation. arXiv:2502.13128

  16. [24]

    Lv, A.; Tan, X.; Lu, P.; Ye, W.; Zhang, S.; Bian, J.; and Yan, R. 2023. GETMusic: Generating Any Music Tracks with a Unified Representation and Diffusion Framework. ArXiv, abs/2305.10841

  17. [25]

    Mesaros, A.; Heittola, T.; Virtanen, T.; and Plumbley, M. D. 2021. Sound event detection: A tutorial. IEEE Signal Processing Magazine, 38(5): 67--83

  18. [26]

    Qian, J.; Zhu, Z.; Zhou, H.; Feng, Z.; Zhai, Z.; and Mao, K. 2025. Beyond the Next Token: Towards Prompt-Robust Zero-Shot Classification via Efficient Multi-Token Prediction. In Chiruzzo, L.; Ritter, A.; and Wang, L., eds., Proceedings of the 2025 Conference of the Nations of ...

  19. [27]

    M.; Liu, J.; Yuan, R.; Min, L.; Liu, X.; Zhang, T.; Du, X.; Guo, S.; Liang, Y.; Li, Y.; Wu, S.; Zhou, J.; Zheng, T.; Ma, Z.; Han, F.; Xue, W.; Xia, G

    Qu, X.; Bai, Y.; Ma, Y.; Zhou, Z.; Lo, K. M.; Liu, J.; Yuan, R.; Min, L.; Liu, X.; Zhang, T.; Du, X.; Guo, S.; Liang, Y.; Li, Y.; Wu, S.; Zhou, J.; Zheng, T.; Ma, Z.; Han, F.; Xue, W.; Xia, G. G.; Benetos, E.; Yue, X.; Lin, C.; Tan, X.; Huang, S. W.; Chen, W.; Fu, J.; and Zhan...

  20. [28]

    A.; and Agchar, I

    Riedhammer, K.; Trump, S.; Ullrich, M.; Braun, F.; Baumann, I.; P \'e rez-Toro, P. A.; and Agchar, I. 2024. A Survey of Music Generation in the Context of Interaction. ArXiv, abs/2402.15294

  21. [29]

    Walshaw, C. 2021. The abc music standard 2.1 (Dec 2011)

  22. [30]

    Wang, C.; Chen, S.; Wu, Y.; Zhang, Z.; Zhou, L.; Liu, S.; Chen, Z.; Liu, Y.; Wang, H.; Li, J.; et al. 2023. Neural codec language models are zero-shot text to speech synthesizers. arXiv preprint arXiv:2301.02111

  23. [31]

    Wang, Y.; Wu, S.; Du, X.; and Sun, M. 2024 a . Exploring Tokenization Methods for Multitrack Sheet Music Generation. arXiv:2410.17584

  24. [32]

    Wang, Y.; Wu, S.; Hu, J.; Du, X.; Peng, Y.; Huang, Y.; Fan, S.; Li, X.; Yu, F.; and Sun, M. 2025. NotaGen: Advancing Musicality in Symbolic Music Generation with Large Language Model Training Paradigms. arXiv:2502.18008

  25. [33]

    Wang, Y.; Yang, W.; Dai, Z.; Zhang, Y.; Zhao, K.; and Wang, H. 2024 b . MeloTrans: A Text to Symbolic Music Generation Model Following Human Composition Habit. ArXiv, abs/2410.13419

  26. [34]

    Wu, S.; Li, X.; Yu, F.; and Sun, M. 2023. TunesFormer: Forming Irish Tunes with Control Codes by Bar Patching. In HCMIR@ISMIR

  27. [35]

    Wu, S.; Wang, Y.; Li, X.; Yu, F.; and Sun, M. 2024. MelodyT5: A Unified Score-to-Score Transformer for Symbolic Music Processing. ArXiv, abs/2407.02277

  28. [36]

    Yang, A.; Li, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Gao, C.; Huang, C.; Lv, C.; Zheng, C.; Liu, D.; Zhou, F.; Huang, F.; Hu, F.; Ge, H.; Wei, H.; Lin, H.; Tang, J.; Yang, J.; Tu, J.; Zhang, J.; Yang, J.; Yang, J.; Zhou, J.; Zhou, J.; Lin, J.; Dang, K.; Bao, K.; ...

  29. [37]

    Yuan, R.; Lin, H.; Guo, S.; Zhang, G.; Pan, J.; Zang, Y.; Liu, H.; Liang, Y.; Ma, W.; Du, X.; Du, X.; Ye, Z.; Zheng, T.; Ma, Y.; Liu, M.; Tian, Z.; Zhou, Z.; Xue, L.; Qu, X.; Li, Y.; Wu, S.; Shen, T.; Ma, Z.; Zhan, J.; Wang, C.; Wang, Y.; Chi, X.; Zhang, X.; Yang, Z.; Wang, X....

  30. [38]

    Yuan, R.; Lin, H.; Wang, Y.; Tian, Z.; Wu, S.; Shen, T.; Zhang, G.; Wu, Y.; Liu, C.; Zhou, Z.; et al. 2024. Chatmusician: Understanding and generating music intrinsically with llm. arXiv preprint arXiv:2402.16153

  31. [39]

    G.; Murata, N.; Ram \'i rez, M

    Zhang, Y.; Ikemiya, Y.; Xia, G. G.; Murata, N.; Ram \'i rez, M. A. M.; Liao, W.-H.; Mitsufuji, Y.; and Dixon, S. 2024. MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models. ArXiv, abs/2402.06178

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.