PitchBench shows that frontier audio-language models have highly unreliable pitch perception across instruments, durations, noise levels, and formats.
Onsets and Frames: Dual-Objective Piano Transcription
5 Pith papers cite this work, alongside 393 external citations. Polarity classification is still indexing.
abstract
We advance the state of the art in polyphonic piano music transcription by using a deep convolutional and recurrent neural network which is trained to jointly predict onsets and frames. Our model predicts pitch onset events and then uses those predictions to condition framewise pitch predictions. During inference, we restrict the predictions from the framewise detector by not allowing a new note to start unless the onset detector also agrees that an onset for that pitch is present in the frame. We focus on improving onsets and offsets together instead of either in isolation as we believe this correlates better with human musical perception. Our approach results in over a 100% relative improvement in note F1 score (with offsets) on the MAPS dataset. Furthermore, we extend the model to predict relative velocities of normalized audio which results in more natural-sounding transcriptions.
citation-role summary
citation-polarity summary
fields
cs.SD 5years
2026 5roles
background 1polarities
background 1representative citing papers
ONOTE is a multi-format benchmark that applies a deterministic pipeline to expose a disconnect between perceptual accuracy and music-theoretic comprehension in leading omnimodal AI models.
Transfer learning from synthetic velocity-labeled guitar data enables velocity prediction in automatic guitar transcription while maintaining competitive note transcription performance.
Cycle-consistent translation enables competitive music transcription performance with mostly unpaired audio and scores plus minimal paired supervision.
citing papers explorer
-
PitchBench: Measuring Pitch Hearing in Audio-Language Models
PitchBench shows that frontier audio-language models have highly unreliable pitch perception across instruments, durations, noise levels, and formats.
-
ONOTE: Benchmarking Omnimodal Notation Processing for Expert-level Music Intelligence
ONOTE is a multi-format benchmark that applies a deterministic pipeline to expose a disconnect between perceptual accuracy and music-theoretic comprehension in leading omnimodal AI models.
-
Velocity Prediction in Automatic Guitar Transcription
Transfer learning from synthetic velocity-labeled guitar data enables velocity prediction in automatic guitar transcription while maintaining competitive note transcription performance.
-
Music Transcription with (Almost) No Supervision
Cycle-consistent translation enables competitive music transcription performance with mostly unpaired audio and scores plus minimal paired supervision.
- MulTTiPop: A Multitrack Transcription Dataset for Pop Music