VocalParse applies interleaved and Chain-of-Thought prompting to a Large Audio Language Model to jointly transcribe lyrics, melody and word-note alignments, achieving state-of-the-art results on multiple singing datasets.
Songtrans: An unified song transcription and alignment method for lyrics and notes
2 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.SD 2years
2026 2verdicts
UNVERDICTED 2roles
background 1polarities
background 1representative citing papers
MusicJudge is a modality-guided framework that performs block-aligned multimodal analysis for singing quality assessment by coupling lyrics with pitch-rhythm fidelity via multi-signal matching and Modality-Guided LoRA fine-tuning.
citing papers explorer
-
VocalParse: Towards Unified and Scalable Singing Voice Transcription with Large Audio Language Models
VocalParse applies interleaved and Chain-of-Thought prompting to a Large Audio Language Model to jointly transcribe lyrics, melody and word-note alignments, achieving state-of-the-art results on multiple singing datasets.
-
Listening Like a Judge: A Music-Aware Framework for Automatic Singing Performance Evaluation
MusicJudge is a modality-guided framework that performs block-aligned multimodal analysis for singing quality assessment by coupling lyrics with pitch-rhythm fidelity via multi-signal matching and Modality-Guided LoRA fine-tuning.