V2N uses a one-second keyboard video window and four jointly trained heads to predict onset, offset, key hold, and velocity, achieving state-of-the-art visual piano transcription on PianoVAM and R3.
V2N is our default full-configuration model (onset, offset, key hold, velocity heads)
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.SD 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Multi-Task Multi-Frame Visual Piano Transcription
V2N uses a one-second keyboard video window and four jointly trained heads to predict onset, offset, key hold, and velocity, achieving state-of-the-art visual piano transcription on PianoVAM and R3.