REVIEW 2 major objections 2 minor 39 references
Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation
T0 review · 2 major / 2 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper claims that generating songs from human-editable bar-level symbolic scores, rather than from raw audio, makes a small model surpass commercial systems like Suno.
desk verdict The abstract makes an important claim about symbolic-score song generation, but the corrupted full text leaves the evidence unverifiable; ask for a clean copy before sending it out. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is BACH's bar-level symbolic tokenization combined with a hierarchical generative procedure. The tokenizer turns a song into discrete bar-level score tokens that preserve hierarchical structure—sections, bars, chords, and notes—so the generative model operates on explicit musical symbols rather than raw audio. This carries the argument because it makes the score both human-editable and structurally transparent, which the paper identifies as the reason controllability and long-form coherence improve.
What would settle it
A preregistered listening and editing study comparing BACH and Suno on identical prompts, with blinded raters and matched rendering, would settle the performance claim; if BACH's preference advantage disappears when renderings are matched, the superiority cannot be attributed to the symbolic-score paradigm.
Extended reading notes
Core claim
The central discovery the paper claims is that bar-level symbolic score representation is enough to make a small song generator exceed the quality, duration, efficiency, and controllability of much larger raw-audio generators. BACH tokenizes music into bar-level units and generates songs in a hierarchy aligned with song structure, so the model does not have to infer tonality, rhythm, and form from audio. Under the paper's evaluations, this design outperforms all publicly reported song generation systems, including Suno, and human raters prefer it across multiple subjective metrics.
Load-bearing premise
The load-bearing assumption is that BACH's measured advantage over raw-audio systems is caused by the symbolic-score paradigm itself, not by the specifics of its evaluation, prompts, rendering, or raters.
Editorial extensions
If this is right
- Song generation can be made human-controllable by editing symbolic scores rather than prompting audio models.
- Long-form structure can be generated without massively larger models, since the score compresses the musical decisions the model must learn.
- Smaller models with symbolic input can produce perceptually competitive songs, challenging the assumption that raw-audio scale is necessary.
- The symbolic-score representation makes the generative process inspectable and editable, opening a practical path for iterative human-AI composition.
Reading between the lines
- A direct extension would be to measure how much of BACH's quality survives the rendering step by comparing its symbolic outputs rendered through different synthesizers against a raw-audio baseline.
- If the causal claim generalizes, combining symbolic score priors with neural audio rendering could improve controllability of existing raw-audio systems without abandoning their acoustic richness.
- A testable prediction is that a raw-audio model of the same size, trained with explicit bar-level structural conditioning, would close part of the gap, which would locate the benefit in structure rather than tokenization.
- The human-editable interface suggests an evaluation metric based on edit distance or user task success, not just listening preference, to capture the controllability advantage.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes BACH (Bar-level AI Composing Helper), a symbolic-score-based song generation model that tokenizes bar-level notation and generates hierarchically structured scores, which are then rendered to audio. The abstract claims that BACH, despite a small model size, achieves state-of-the-art results across all publicly reported song-generation systems, surpassing commercial systems such as Suno, and that human evaluations confirm its superiority on multiple subjective metrics. The paper attributes existing limitations in controllability, generalizability, perceptual quality, and duration to the dominant paradigm of learning music theory directly from raw audio. The full text of the submission, however, is almost entirely corrupted mojibake, so the technical and experimental contents cannot be examined.
Significance. If the central claims were fully supported, the work would be significant: it would demonstrate that a small symbolic-score model can outperform much larger raw-audio generative models on perceptual quality and controllability, and it would offer a practical, human-editable approach to long song generation. The paper's framing that the raw-audio paradigm itself, rather than model scale, is the bottleneck is a falsifiable and potentially impactful thesis. However, the current manuscript provides no inspectable evidence: the abstract contains no quantitative results or protocol details, and the body is unreadable. Consequently, the significance cannot be assessed independently, and no reproducible code, data, or demos are visible from the artifact.
major comments (2)
- [Full text (all sections after the Abstract)] The entire body of the manuscript, from the first page after the abstract onward, is rendered as mojibake (e.g., sequences of '��������' characters), with no legible equations, tables, or prose. As a result, the tokenization strategy, generative procedure, model architecture, datasets, training details, and experiments cannot be checked. The central empirical claim that BACH establishes a new SOTA and surpasses Suno therefore rests entirely on the abstract's assertions, with no inspectable evidence. A clean, complete manuscript with the full experimental protocol is required before this claim can be evaluated.
- [Abstract, sentences 4–6] The abstract states that 'Experiments demonstrate that BACH, with a small model size, establishes a new SOTA among all publicly reported song generation systems, even surpassing commercial solutions such as Suno,' and that 'Human evaluations further confirm its superiority across multiple subjective metrics.' However, no quantitative scores, effect sizes, error bars, p-values, number of raters, or details of the comparison protocol are given, even in the abstract. Since Suno is a black box, the comparison's fairness depends on matched prompts, blinded listening, and a comparable rendering pipeline; without this information, the superiority claim and the causal attribution of the four limitations to the raw-audio paradigm are unsupported. The resubmission must include a detailed experimental section with these specifics.
minor comments (2)
- [Header] The embedded header 'arXiv:2508.01392v2 [cs.LG] 26 Jun 2026' does not match the paper's arXiv ID 2508.01394 and subject class cs.SD; please correct the header to avoid provenance confusion.
- [Formatting] The title, abstract, and body are not cleanly separated in the corrupted text; the resubmission should use a properly encoded PDF so that the reviewers can read the technical content.
Circularity Check
No circularity is demonstrable from the readable abstract; the full text is corrupted, so no derivation step can be shown to reduce to its own inputs.
full rationale
Only the abstract is readable; the body, equations, tables, and experimental protocol are corrupted. No load-bearing derivation chain can be inspected, and therefore no specific reduction of a prediction to a fitted parameter, no self-definitional equivalence, and no load-bearing self-citation chain can be exhibited. The abstract's claims are evaluated against external benchmarks and human raters, which is structurally an external test rather than a tautology. The residual concerns about the unverifiable Suno comparison and the unsupported causal attribution to the raw-audio paradigm are correctness and evidence-quality issues, not circularity. Under the hard rule that circularity may be flagged only when the paper's own equations or citations show the reduction, the honest finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption The four persistent limitations (controllability, generalizability, perceptual quality, duration) stem primarily from learning music theory from raw audio.
- domain assumption Human subjective evaluation is a valid and sufficient ground for the SOTA claim.
- domain assumption Symbolic-score-then-render preserves enough musical information for perceptual quality competitive with audio-trained models.
Cite this review
Pith. "Pith review of Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation." pith.science (2026). https://pith.science/paper/WXCE4BRX
@misc{pith2026250801394,
author = {Pith},
title = {Pith review of: Via Score to Performance: Efficient Human-Controllable Long Song Generation with Bar-Level Symbolic Notation},
year = {2026},
howpublished = {\url{https://pith.science/paper/WXCE4BRX}},
note = {Machine review of arXiv:2508.01394}
}
read the original abstract
Song generation is regarded as the most challenging problem in music AIGC; nonetheless, existing approaches have yet to fully overcome four persistent limitations: controllability, generalizability, perceptual quality, and duration. We argue that these shortcomings stem primarily from the prevailing paradigm of attempting to learn music theory directly from raw audio, a task that remains prohibitively difficult for current models. To address this, we present Bar-level AI Composing Helper (BACH), the first model explicitly designed for song generation through human-editable symbolic scores. BACH introduces a tokenization strategy and a symbolic generative procedure tailored to hierarchical song structure. Consequently, it achieves substantial gains in the efficiency, duration, and perceptual quality of song generation. Experiments demonstrate that BACH, with a small model size, establishes a new SOTA among all publicly reported song generation systems, even surpassing commercial solutions such as Suno. Human evaluations further confirm its superiority across multiple subjective metrics.
Reference graph
Works this paper leans on
-
[1]
, " * write output.state after.block = add.period write newline
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint howpublished institution isbn journal key month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.a...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Agostinelli, A.; Denk, T. I.; Borsos, Z.; Engel, J.; Verzetti, M.; Caillon, A.; Huang, Q.; Jansen, A.; Roberts, A.; Tagliasacchi, M.; et al. 2023. Musiclm: Generating music from text. arXiv preprint arXiv:2301.11325
arXiv 2023
-
[4]
Bai, Y.; et al. 2024. Seed Music: ByteDance's neural music generation system. https://bytedance.com/research/seed-music
work page 2024
-
[5]
Borsos, Z.; Marinier, R.; Vincent, D.; Kharitonov, E.; Pietquin, O.; Sharifi, M.; Roblek, D.; Teboul, O.; Grangier, D.; Tagliasacchi, M.; and Zeghidour, N. 2023. AudioLM: A Language Modeling Approach to Audio Generation. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31: 2523--2533
work page 2023
-
[6]
Chen, Y.; Huang, L.; and Gou, T. 2024. Applications and Advances of Artificial Intelligence in Music Generation:A Review. ArXiv, abs/2409.03715
work page Pith review arXiv 2024
-
[7]
Civit, M.; Civit-Masot, J.; Cuadrado, F.; and Cuaresma, M. J. E. 2022. A systematic review of artificial intelligence-based music generation: Scope, applications, and future trends. Expert Syst. Appl., 209: 118190
work page 2022
-
[8]
Copet, J.; Kreuk, F.; Gat, I.; Remez, T.; Kant, D.; Synnaeve, G.; Adi, Y.; and D \'e fossez, A. 2023. Simple and controllable music generation. Advances in Neural Information Processing Systems, 36: 47704--47720
work page 2023
Show all 39 references
-
[9]
W.; Radford, A.; and Sutskever, I
Dhariwal, P.; Jun, H.; Payne, C.; Kim, J. W.; Radford, A.; and Sutskever, I. 2020. Jukebox: A generative model for music. arXiv preprint arXiv:2005.00341
2020 arXiv
-
[10]
Ding, S.; Liu, Z.; wen Dong, X.; Zhang, P.; Qian, R.; He, C.; Lin, D.; and Wang, J. 2024. SongComposer: A Large Language Model for Lyric and Melody Composition in Song Generation. ArXiv, abs/2402.17645
2024 arXiv
-
[11]
Donahue, C.; Caillon, A.; Roberts, A.; Manilow, E.; Esling, P.; Agostinelli, A.; Verzetti, M.; Simon, I.; Pietquin, O.; Zeghidour, N.; et al. 2023. Singsong: Generating musical accompaniments from singing. arXiv preprint arXiv:2301.12662
2023 arXiv
-
[12]
Du, X.; Wang, Z.; Liang, X.; Liang, H.; Zhu, B.; and Ma, Z. 2023. Bytecover3: Accurate cover song identification on short queries. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 1--5. IEEE
2023
-
[13]
Défossez, A.; Copet, J.; Synnaeve, G.; and Adi, Y. 2022. High Fidelity Neural Audio Compression. arXiv:2210.13438
2022 arXiv
-
[14]
Fang, R.; Duan, C.; Wang, K.; Huang, L.; Li, H.; Yan, S.; Tian, H.; Zeng, X.; Zhao, R.; Dai, J.; et al. 2025. Got: Unleashing reasoning capability of multimodal large language model for visual generation and editing. arXiv preprint arXiv:2503.10639
2025 arXiv
-
[15]
Hong, Z.; Cui, C.; Huang, R.; Zhang, L.; Liu, J.; He, J.; and Zhao, Z. 2023. Unisinger: Unified end-to-end singing voice synthesis with cross-modality information matching. In Proceedings of the 31st ACM International Conference on Multimedia, 7569--7579
2023
-
[16]
A.; Vaswani, A.; Uszkoreit, J.; Shazeer, N.; Simon, I.; Hawthorne, C.; Dai, A
Huang, C.-Z. A.; Vaswani, A.; Uszkoreit, J.; Shazeer, N.; Simon, I.; Hawthorne, C.; Dai, A. M.; Hoffman, M. D.; Dinculescu, M.; and Eck, D. 2018. Music transformer. arXiv preprint arXiv:1809.04281
2018 arXiv
-
[17]
Kim, T.; and Nam, J. 2023. All-in-One Metrical and Functional Structure Analysis with Neighborhood Attentions on Demixed Audio. In 2023 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (WASPAA), 1--5
2023
-
[18]
Latif, S.; Shoukat, M.; Shamshad, F.; Usama, M.; Cuay'ahuitl, H.; and Schuller, B. 2023. Sparks of Large Audio Models: A Survey and Outlook. ArXiv, abs/2308.12792
2023 arXiv
-
[19]
Lei, J.; et al. 2025. SongCreator: Transformer-based text-to-song generation. In Proceedings of the International Conference on Machine Learning
2025
-
[20]
Li, R.; Hong, Z.; Wang, Y.; Zhang, L.; Huang, R.; Zheng, S.; and Zhao, Z. 2024. Accompanied Singing Voice Synthesis with Fully Text-controlled Melody. CoRR
2024
-
[21]
Liang, X.; Du, X.; Lin, J.; Zou, P.; Wan, Y.; and Zhu, B. 2024. ByteComposer: a Human-like Melody Composition Method based on Language Model Agent. ArXiv, abs/2402.17785
2024 arXiv
-
[22]
Liu, H.; Chen, Z.; Yuan, Y.; Mei, X.; Liu, X.; Mandic, D.; Wang, W.; and Plumbley, M. D. 2023. Audioldm: Text-to-audio generation with latent diffusion models. arXiv preprint arXiv:2301.12503
2023 arXiv
-
[23]
Liu, Z.; Ding, S.; Zhang, Z.; Dong, X.; Zhang, P.; Zang, Y.; Cao, Y.; Lin, D.; and Wang, J. 2025. SongGen: A Single Stage Auto-regressive Transformer for Text-to-Song Generation. arXiv:2502.13128
2025 arXiv
-
[24]
Lv, A.; Tan, X.; Lu, P.; Ye, W.; Zhang, S.; Bian, J.; and Yan, R. 2023. GETMusic: Generating Any Music Tracks with a Unified Representation and Diffusion Framework. ArXiv, abs/2305.10841
2023 arXiv
-
[25]
Mesaros, A.; Heittola, T.; Virtanen, T.; and Plumbley, M. D. 2021. Sound event detection: A tutorial. IEEE Signal Processing Magazine, 38(5): 67--83
2021
-
[26]
Qian, J.; Zhu, Z.; Zhou, H.; Feng, Z.; Zhai, Z.; and Mao, K. 2025. Beyond the Next Token: Towards Prompt-Robust Zero-Shot Classification via Efficient Multi-Token Prediction. In Chiruzzo, L.; Ritter, A.; and Wang, L., eds., Proceedings of the 2025 Conference of the Nations of ...
2025
-
[27]
M.; Liu, J.; Yuan, R.; Min, L.; Liu, X.; Zhang, T.; Du, X.; Guo, S.; Liang, Y.; Li, Y.; Wu, S.; Zhou, J.; Zheng, T.; Ma, Z.; Han, F.; Xue, W.; Xia, G
Qu, X.; Bai, Y.; Ma, Y.; Zhou, Z.; Lo, K. M.; Liu, J.; Yuan, R.; Min, L.; Liu, X.; Zhang, T.; Du, X.; Guo, S.; Liang, Y.; Li, Y.; Wu, S.; Zhou, J.; Zheng, T.; Ma, Z.; Han, F.; Xue, W.; Xia, G. G.; Benetos, E.; Yue, X.; Lin, C.; Tan, X.; Huang, S. W.; Chen, W.; Fu, J.; and Zhan...
2024 arXiv
-
[28]
A.; and Agchar, I
Riedhammer, K.; Trump, S.; Ullrich, M.; Braun, F.; Baumann, I.; P \'e rez-Toro, P. A.; and Agchar, I. 2024. A Survey of Music Generation in the Context of Interaction. ArXiv, abs/2402.15294
2024 arXiv
-
[29]
Walshaw, C. 2021. The abc music standard 2.1 (Dec 2011)
2021
-
[30]
Wang, C.; Chen, S.; Wu, Y.; Zhang, Z.; Zhou, L.; Liu, S.; Chen, Z.; Liu, Y.; Wang, H.; Li, J.; et al. 2023. Neural codec language models are zero-shot text to speech synthesizers. arXiv preprint arXiv:2301.02111
2023 arXiv
-
[31]
Wang, Y.; Wu, S.; Du, X.; and Sun, M. 2024 a . Exploring Tokenization Methods for Multitrack Sheet Music Generation. arXiv:2410.17584
2024 arXiv
-
[32]
Wang, Y.; Wu, S.; Hu, J.; Du, X.; Peng, Y.; Huang, Y.; Fan, S.; Li, X.; Yu, F.; and Sun, M. 2025. NotaGen: Advancing Musicality in Symbolic Music Generation with Large Language Model Training Paradigms. arXiv:2502.18008
2025 arXiv
-
[33]
Wang, Y.; Yang, W.; Dai, Z.; Zhang, Y.; Zhao, K.; and Wang, H. 2024 b . MeloTrans: A Text to Symbolic Music Generation Model Following Human Composition Habit. ArXiv, abs/2410.13419
2024 arXiv
-
[34]
Wu, S.; Li, X.; Yu, F.; and Sun, M. 2023. TunesFormer: Forming Irish Tunes with Control Codes by Bar Patching. In HCMIR@ISMIR
2023
-
[35]
Wu, S.; Wang, Y.; Li, X.; Yu, F.; and Sun, M. 2024. MelodyT5: A Unified Score-to-Score Transformer for Symbolic Music Processing. ArXiv, abs/2407.02277
2024 arXiv
-
[36]
Yang, A.; Li, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Gao, C.; Huang, C.; Lv, C.; Zheng, C.; Liu, D.; Zhou, F.; Huang, F.; Hu, F.; Ge, H.; Wei, H.; Lin, H.; Tang, J.; Yang, J.; Tu, J.; Zhang, J.; Yang, J.; Yang, J.; Zhou, J.; Zhou, J.; Lin, J.; Dang, K.; Bao, K.; ...
2025 arXiv
-
[37]
Yuan, R.; Lin, H.; Guo, S.; Zhang, G.; Pan, J.; Zang, Y.; Liu, H.; Liang, Y.; Ma, W.; Du, X.; Du, X.; Ye, Z.; Zheng, T.; Ma, Y.; Liu, M.; Tian, Z.; Zhou, Z.; Xue, L.; Qu, X.; Li, Y.; Wu, S.; Shen, T.; Ma, Z.; Zhan, J.; Wang, C.; Wang, Y.; Chi, X.; Zhang, X.; Yang, Z.; Wang, X....
2025
-
[38]
Yuan, R.; Lin, H.; Wang, Y.; Tian, Z.; Wu, S.; Shen, T.; Zhang, G.; Wu, Y.; Liu, C.; Zhou, Z.; et al. 2024. Chatmusician: Understanding and generating music intrinsically with llm. arXiv preprint arXiv:2402.16153
2024 arXiv
-
[39]
G.; Murata, N.; Ram \'i rez, M
Zhang, Y.; Ikemiya, Y.; Xia, G. G.; Murata, N.; Ram \'i rez, M. A. M.; Liao, W.-H.; Mitsufuji, Y.; and Dixon, S. 2024. MusicMagus: Zero-Shot Text-to-Music Editing via Diffusion Models. ArXiv, abs/2402.06178
2024 arXiv
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.