Pith. sign in

REVIEW 3 major objections 5 minor 37 references

PianoVAM: A Multimodal Piano Performance Dataset

T0 review · 3 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read PianoVAM publishes 21 hours of synchronized piano performance data with fingering labels.

desk verdict Valuable new multimodal piano dataset, but the fingering 'complete and accurate' claim is contradicted by the algorithm's own description and is validated on only 150 notes per piece. read the letter →

arxiv 2509.08800 v1 pith:NNC44T3K submitted 2025-09-10 cs.SD cs.AIcs.CVcs.MMeess.AS

classification cs.SDcs.AIcs.CVcs.MMeess.AS
keywords pianoperformancedatasetmultimodalaudio-visualtranscriptionfingeringannotationhandlandmarksMediaPipeMIDIDisklavier
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces PianoVAM, a multimodal dataset of amateur piano practice sessions recorded on a Disklavier, with synchronized top-view video, audio, MIDI, hand landmarks, fingering pseudo-labels, and metadata. The central claim is that this dataset fills gaps left by prior collections, which either lack video, rely on synthetic audio, or have incomplete fingering annotations. The paper further claims that its hybrid fingering annotation algorithm, combining hand-pose landmarks with manual refinement of ambiguous cases, yields complete fingering labels for all notes with about 95% precision. A sympathetic reader would care because the dataset enables audio-visual piano transcription research and the study of fingering and hand movement in real practice conditions.

What carries the argument

The key mechanism is the hybrid fingering annotation pipeline. It first runs MediaPipe Hands on each video frame to extract hand landmarks, then estimates the relative z-depth of each hand from the projected 2D skeleton using a model skeleton defined by the wrist, index metacarpal, and ring metacarpal, with a heuristic 28-degree angle for the neutral hand and a 0.9 depth threshold to identify and discard a floating hand. For each MIDI note, a fingering score is computed from the number of frames in which each fingertip lies within the note's key area; fingers scoring above 50% and 80% of the maximum become normal and strong candidates. Notes with a single strong candidate are labeled automat

What would settle it

Take a random sample of, say, 500 notes from the remaining ~1.05 million annotated notes, have an expert manually label the fingering from the video, and compare; if the precision falls substantially below the reported ~95%, or if certain pieces with heavy pedal use or fast passagework consistently show errors, the claim of complete and accurate fingering annotations across the dataset is falsified.

Watch

Extended reading notes

Core claim

The core discovery is the dataset itself: 106 solo piano recordings from 10 amateur performers, roughly 21 hours, captured in realistic practice conditions with a Yamaha Disklavier. Unlike MAESTRO, which lacks video, or OMAPS2 and PianoYT, which have limited or pseudo MIDI, PianoVAM provides aligned top-view video, real audio, ground-truth MIDI from the Disklavier, hand landmark sequences, and fingering pseudo-labels. The paper also demonstrates that the fingering annotation method, a hybrid of automatic hand-landmark-based candidate scoring and human selection when multiple candidates exist, reaches an average precision above 95% on a manually checked subset, and that a simple audio-visual

Load-bearing premise

The fingering pseudo-labels for the entire dataset rest on the assumption that the hand-pose landmarks and the z-depth heuristics (28-degree angle and 0.9 threshold) correctly identify which hand and finger plays each note from top-view video, an assumption validated only on the first 150 notes of 10 pieces.

Editorial extensions

If this is right

  • If the fingering annotations are accepted, PianoVAM becomes the largest piano dataset with both real performance audio and complete fingering labels, enabling supervised learning of fingering prediction from video or score.
  • The synchronized top-view video and MIDI enable audio-visual transcription models that can be tested on realistic practice-room conditions, including noise and reverberation, as shown by the benchmark.
  • The dataset's hand landmarks and per-hand MIDI separation support research on left-right hand assignment, hand posture analysis, and practice behavior analysis.
  • The reported benchmark indicates that combining audio-only models with video-based candidate filtering can improve onset precision in degraded acoustic conditions, pointing to a practical path for robust transcription.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper's own benchmarks, the hand landmarks and fingering labels could be used to study the relationship between fingering choices and expressive timing or dynamics across the same pieces performed by different amateurs.
  • The validation of the fingering algorithm covers only the first 150 notes of 10 pieces; a natural next step is to extend the manual ground-truth set to a random sample across all 106 recordings to measure the accuracy on the remaining ~1.05 million notes and on rarer techniques.
  • Because the dataset is dominated by practice sessions with heavy pedal use and a specific recording studio, models trained on it may need domain adaptation before transferring to concert recordings or other pianos; this is a caveat rather than a defect, but testable.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. PianoVAM is a new multimodal piano performance dataset comprising 21 hours (106 recordings) of amateur practice sessions captured with a Yamaha Disklavier, including synchronized top-view video, audio, MIDI, hand landmarks (MediaPipe), fingering pseudo-labels, and metadata. The paper describes the acquisition workflow, audio-MIDI alignment via FluidSynth rendering and DTW, loudness normalization, dataset statistics, and a semi-automated fingering annotation pipeline based on hand landmarks and a rule-based candidate scoring method with GUI-assisted disambiguation. Benchmark experiments cover audio-only piano transcription (Onsets and Frames) and an audio-visual post-processing filter, with statistical tests showing benefits from PianoVAM training data and from visual filtering in noisy/reverberant conditions. The central contribution is the dataset itself, with the fingering labeling being its most novel but least validated component.

Significance. If released as described, PianoVAM would be a valuable community resource: it combines real performance audio/MIDI with top-view video and hand landmarks at a scale comparable to OMAPS2, adds fingering annotations absent from MAESTRO, and provides a reproducible data split and baseline results. The distributional analysis relative to MAESTRO and the use of standard metrics with significance testing are useful. The manuscript also gives a fairly detailed description of the difficult audio/MIDI/video alignment problem. However, the paper's headline claim of 'complete and accurate fingering annotations for the entire dataset' is not supported by the evidence presented; this issue is load-bearing because the fingering modality is a primary differentiator of the dataset.

major comments (3)
  1. [§5 and §5.2 / Table 3] The statement 'This approach ensures complete and accurate fingering annotations for the entire dataset' is directly contradicted by §5.2's 'there might be notes with either no candidate or multiple candidates, in which case the algorithm will leave these notes unlabeled' and by Table 3, where 3.7–35.1% of notes (weighted average 13.0%) have no candidate. Table 2 nonetheless reports a 100% labeled ratio. If the custom GUI is used to fill in all no-candidate and multiple-candidate notes, that manual step must be stated and quantified; if it is not, the dataset is incomplete. The text also claims multiple-candidate cases affect ~20% of notes, whereas Table 3 gives a 5.1% weighted average, so the description of manual workload is internally inconsistent.
  2. [§5.2.1 / Table 3] The reliability estimate is based on the first 150 notes of 10 pieces—about 1,500 notes total, roughly 0.14% of the 1,050,966-note dataset—and the 10 excerpts are contiguous openings, which are unlikely to be representative of dense or technically difficult passages later in pieces. Precision is reported without recall and without specifying whether no-candidate notes are excluded from the denominator. Under the most optimistic interpretation, this gives an upper bound of ~95% × (1−0.13−0.05) ≈ 78% correctly labeled notes on the validation pieces, and this is unmeasured for the remaining 96 pieces. The authors should either evaluate on a random or stratified sample covering the full dataset or explicitly demote the claim to 'best-effort pseudo-labels.'
  3. [§5.2 / §7] The z-depth and candidate heuristics (28° IWR angle, 0.9 floating-hand threshold, 50%/80% candidate thresholds, and the ±2-key video filtering threshold in §6.3) contain several free parameters, but no sensitivity analysis is provided. Section 7 itself acknowledges failures due to motion blur in Ravel and shadows in Schumann, and Table 3 shows Ravel has 35.1% no-candidate notes. The ~95% precision number should therefore be presented as condition-specific rather than as a dataset-wide reliability certificate; at minimum, per-piece confidence intervals and a breakdown of error types are needed.
minor comments (5)
  1. [Title/Abstract] The dataset name is typeset inconsistently ('PianoV AM' vs. 'PianoVAM'); please unify the spelling in the title, abstract, and body.
  2. [§5.2, Eq. (1)–(5)] The notation is confusing: Eq. (1) defines I0W0R0 as a median of |△IW R|, but Eq. (3)–(5) use ||I0W0||, ||W0R0||, ||R0I0|| as distances. Please clarify how the median triangle gives segment lengths and define all symbols. The symbol AR in Eq. (2) is overloaded as both a coordinate bound and the aspect ratio.
  3. [§2.3 vs. §5] Section 2.3 says the dataset is 'improved by manual annotation of incomplete fingering labels,' but the annotation pipeline in §5 never describes this manual step except for GUI-based disambiguation of multiple candidates. Add a cross-reference and a short description of how no-candidate notes are handled.
  4. [§5.2] 'Powell's dog leg algorithm' should be 'Powell's dogleg algorithm.'
  5. [§6.3, Table 5] The table's footer states 'Bold: highest; Underline: significantly higher over the preceding method,' but the table as printed uses no underlining. Please make the formatting consistent with the caption.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the dataset is a new resource and its claimed annotations and benchmarks are evaluated against manual labels and held-out splits, not derived from the claims themselves.

full rationale

The paper's central contribution is a new multimodal dataset, not a formal derivation. The fingering labels are produced by an explicitly described geometric heuristic (MediaPipe landmarks, 28-degree IWR angle, 0.9 z-depth threshold, 50%/80% candidate-score thresholds) and are validated against manual ground truth on the first 150 notes of 10 pieces (Section 5.2.1, Table 3). This is an empirical measurement, not a fitted parameter renamed as a prediction; even if thresholds were tuned on the same corpus, the manuscript provides no equation that forces the evaluation result by construction. The phrase 'ensures complete and accurate fingering annotations for the entire dataset' is in tension with Section 5.2's statement that 'there might be notes with either no candidate or multiple candidates, in which case the algorithm will leave these notes unlabeled,' but that is a completeness/accuracy limitation, not a circular step. The benchmark results use the standard Onsets and Frames model on provided 80/10/10 splits (Sections 6.1-6.3), so no reported score is an input to the dataset construction. The only overlapping-author citation is [25], used for a noise-augmentation detail ('SNR randomly sampled from 0 to 24dB (cf. [25])'); the method is fully described inline and the citation is not load-bearing. Thus there is no self-definitional, fitted-input, or self-citation-based circularity.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The paper's central product is a dataset, so the main 'free parameters' are hand-tuned thresholds in the fingering and visual-filtering pipelines. These are not formally fit to a target output, but they were likely chosen on the same corpus used for the precision evaluation, so they affect the reported numbers. The axioms are standard assumptions about camera geometry, pretrained models, and evaluation conventions, plus the domain assumption that MediaPipe landmarks are reliable on piano video.

free parameters (6)
  • IWR angle heuristic = 28 degrees
    Chosen as a heuristic estimate for the average human hand in neutral position (Section 5.2); used to select the model skeleton for z-depth estimation.
  • floating-hand depth threshold = 0.9
    z-depth less than 0.9 indicates a hand floating more than 10% of the camera-keyboard distance; hand-tuned threshold for excluding non-playing hands (Section 5.2).
  • normal candidate threshold = 50% of note duration
    Fingering score threshold for adding a finger as a normal candidate (Section 5.2).
  • strong candidate threshold = 80% of note duration
    Threshold for a strong candidate that becomes the single candidate when unique (Section 5.2).
  • key candidate range = +/- 2 white keys
    Tunable threshold for candidate keys per fingertip in the audio-visual post-processing step (Section 6.3).
  • loudness normalization target = -23 LUFS
    Desired global average loudness, used to scale real recordings to match synthesized renderings (Section 3.2.2).
assumptions (5)
  • domain assumption MediaPipe Hands estimates hand landmarks accurately enough for piano fingering assignment in top-view video
    Used throughout Section 5; no independent validation on piano-specific hand poses or varied lighting.
  • standard math The three-point distance equations (Eqs. 3-5) with the median-based model skeleton recover the relative z-depth of the hand
    Assumes the hand skeleton is rigid and the projective camera model holds; the 28-degree angle heuristic fixes the skeleton.
  • domain assumption White keys are evenly spaced and the keyboard is a plane with known corners, so fingertip x-coordinates map linearly to key indices
    Used in Section 6.3 for visual filtering; ignores black-key geometry and perspective distortions beyond the affine transform.
  • domain assumption The sustain-pedal convention (offset set to pedal-release time) is valid for evaluating transcription
    Standard in MAESTRO-style evaluation, used in Section 6.2 without further justification.
  • domain assumption Audio-MIDI alignment via FluidSynth rendering plus Dynamic Time Warping within a ±2.5s band is accurate enough for frame-level synchronization
    Section 3.2.1; no explicit verification of the alignment error on the actual recordings.

how reviews work

0 comments
Cite this review

Pith. "Pith review of PianoVAM: A Multimodal Piano Performance Dataset." pith.science (2026). https://pith.science/paper/NNC44T3K

@misc{pith2026250908800,
  author       = {Pith},
  title        = {Pith review of: PianoVAM: A Multimodal Piano Performance Dataset},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NNC44T3K}},
  note         = {Machine review of arXiv:2509.08800}
}
read the original abstract

The multimodal nature of music performance has driven increasing interest in data beyond the audio domain within the music information retrieval (MIR) community. This paper introduces PianoVAM, a comprehensive piano performance dataset that includes videos, audio, MIDI, hand landmarks, fingering labels, and rich metadata. The dataset was recorded using a Disklavier piano, capturing audio and MIDI from amateur pianists during their daily practice sessions, alongside synchronized top-view videos in realistic and varied performance conditions. Hand landmarks and fingering labels were extracted using a pretrained hand pose estimation model and a semi-automated fingering annotation algorithm. We discuss the challenges encountered during data collection and the alignment process across different modalities. Additionally, we describe our fingering annotation method based on hand landmarks extracted from videos. Finally, we present benchmarking results for both audio-only and audio-visual piano transcription using the PianoVAM dataset and discuss additional potential applications.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

37 extracted references · 2 linked inside Pith

  1. [1]

    PianoV AM: A Multimodal Piano Perfor- mance Dataset

    INTRODUCTION Music performance is inherently multimodal, involving not only audio but also motion, posture, and other visual ele- ments as part of the expressive sound creation process [1,2]. In the field of Music Information Retrieval (MIR), there has been growing interest in collecting multimodal perfor- mance data to enhance the extraction of musical i...

  2. [2]

    RELA TED WORK 2.1 Audio-Visual Datasets The emergence of audio-visual datasets represents a promis- ing frontier in MIR, unlocking new research possibilities by providing visual information that complements audio signals. The URMP dataset [3] offers synchronized audio, video, and MIDI recordings of multi-instrument classical performances, supporting multi...

  3. [3]

    3.1.1 Acquisition Workflow The acquisition workflow comprises the following steps

    DA TASET ACQUISITION & PRE-PROCESSING 3.1 Acquisition We developed a data acquisition system to streamline the un- supervised recording of video, audio, MIDI, and associated metadata, such as performer and piece details. 3.1.1 Acquisition Workflow The acquisition workflow comprises the following steps. First, new users register by providing basic personal...

  4. [4]

    DA TASET STA TISTICS The dataset contains 106 solo piano recordings from 10 amateur performers, totaling approximately 21 hours. The repertoire is stylistically diverse, spanning works from 38 composers from the Baroque to modern eras (e.g., Bach, Chopin, Kapustin, Joe Hisaishi) and includes several im- provisations. Performers’ self-reported skill levels...

  5. [5]

    The algorithm first processes performance videos to map hand landmarks to potential finger candidates for each MIDI note

    ANNOTA TION OF FINGERING LABELS To generate fingering annotations, we developed the hybrid algorithm shown in Figure 3. The algorithm first processes performance videos to map hand landmarks to potential finger candidates for each MIDI note. For notes with a single, unambiguous candidate, the fingering is determined automatically, achieving a precision of...

  6. [6]

    The audio-visual experiments are specifically designed to assess the visual modality’s contribution to enhancing performance under challenging acoustic conditions

    BENCHMARK RESULTS To demonstrate its utility, we benchmark the PianoV AM dataset on the task of piano transcription under two settings: audio-only and audio-visual. The audio-visual experiments are specifically designed to assess the visual modality’s contribution to enhancing performance under challenging acoustic conditions. 6.1 Data Split To facilitate...

  7. [7]

    While this approach streamlines data acquisition, the dataset exhibits biases in performer identity, pedal usage, and composer representa- tion

    DISCUSSION The dataset was collected using a system designed to fa- cilitate unsupervised recording, allowing performers to play freely without on-site assistance. While this approach streamlines data acquisition, the dataset exhibits biases in performer identity, pedal usage, and composer representa- tion. In addition, since all recordings originate from...

  8. [8]

    CONCLUSION We presented PianoV AM, a comprehensive multimodal dataset of amateur piano practice sessions that captures synchronized top-view video, audio, MIDI, hand land- marks, fingering labels, and rich metadata. Recorded using a Yamaha Disklavier in natural, varied practice conditions, PianoV AM addresses key limitations of existing datasets that ofte...

Show all 37 references
  1. [9]

    KH2023-235)

    ETHICS STA TEMENT This study involved human participants for data collection, which was approved by the Institutional Review Board (IRB) at KAIST (Approval No. KH2023-235). All proce- dures strictly adhered to established ethical guidelines

  2. [10]

    This research was supported by the National Research Foundation of Ko- rea (NRF) funded by the Korea Government (MSIT) under Grant RS-2023-NR077289 and Grant RS-2024-00358448

    ACKNOWLEDGMENTS We sincerely appreciate the KAIST music and audio com- puting lab and PIAST (Piano club) members who partici- pated in the dataset acquisition as performers. This research was supported by the National Research Foundation of Ko- rea (NRF) funded by the Korea Go...

  3. [11]

    Hearing and seeing musical expression,

    V . Bergeron and D. M. Lopes, “Hearing and seeing musical expression,”Philosophy and Phenomenological Research, vol. 78, no. 1, pp. 1–16, 2009

  4. [12]

    When the eye listens: A meta- analysis of how audio-visual presentation enhances the appreciation of music performance,

    F. Platz and R. Kopiez, “When the eye listens: A meta- analysis of how audio-visual presentation enhances the appreciation of music performance,”Music Perception: An Interdisciplinary Journal, vol. 30, no. 1, pp. 71–83, 2012

  5. [13]

    Cre- ating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications,

    B. Li, X. Liu, K. Dinesh, Z. Duan, and G. Sharma, “Cre- ating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applications,”IEEE Transactions on Multimedia, vol. 21, no. 2, pp. 522–535, 2018

  6. [14]

    Audiovisual analysis of music perfor- mances: Overview of an emerging field,

    Z. Duan, S. Essid, C. C. S. Liem, G. Richard, and G. Sharma, “Audiovisual analysis of music perfor- mances: Overview of an emerging field,”IEEE Signal Processing Magazine, vol. 36, no. 1, pp. 63–73, 2019

  7. [15]

    Sight to sound: An end-to-end approach for visual piano transcription,

    A. S. Koepke, O. Wiles, Y . Moses, and A. Zisserman, “Sight to sound: An end-to-end approach for visual piano transcription,” inProceedings of the IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2020, pp. 1838–1842

  8. [16]

    A crnn- gcn piano transcription model based on audio and skele- ton features,

    Y . Li, X. Wang, R. Wu, W. Xu, and W. Chen, “A crnn- gcn piano transcription model based on audio and skele- ton features,” inIEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW), 2023, pp. 1–5

  9. [17]

    Audiovisual source associa- tion for string ensembles through multi-modal vibrato analysis,

    B. Li, C. Xu, and Z. Duan, “Audiovisual source associa- tion for string ensembles through multi-modal vibrato analysis,” inProceedings of the Sound and Music Com- puting Conference (SMC), 2017

  10. [18]

    A cap- pella: Audio-visual singing voice separation,

    J. F. Montesinos, V . S. Kadandale, and G. Haro, “A cap- pella: Audio-visual singing voice separation,” inPro- ceedings of the 32nd British Machine Vision Conference (BMVC), 2021

  11. [19]

    Identifying melodic motifs and stable notes from gestural informa- tion in indian vocal performances,

    S. Nadkarni, P. Rao, and M. Clayton, “Identifying melodic motifs and stable notes from gestural informa- tion in indian vocal performances,”Transactions of the International Society for Music Information Retrieval, vol. 7, no. 1, 2024

  12. [20]

    Enabling factorized piano music modeling and generation with the MAESTRO dataset,

    C. Hawthorne, A. Stasyuk, A. Roberts, I. Simon, C.- Z. A. Huang, S. Dieleman, E. Elsen, J. Engel, and D. Eck, “Enabling factorized piano music modeling and generation with the MAESTRO dataset,” inPro- ceedings of the International Conference on Learning Representations (ICLR), 2019

  13. [21]

    Maps - a piano database for multipitch estimation and automatic transcription of music,

    V . Emiya, N. Bertin, B. David, and R. Badeau, “Maps - a piano database for multipitch estimation and automatic transcription of music,” Research Report, Tech. Rep., 2010. [Online]. Available: https: //hal.inria.fr/inria-00544155

  14. [22]

    Statistical learn- ing and estimation of piano fingering,

    E. Nakamura, Y . Saito, and K. Yoshii, “Statistical learn- ing and estimation of piano fingering,”Information Sci- ences, vol. 517, pp. 68–85, 2020

  15. [23]

    Automatic piano fingering from partially an- notated scores using autoregressive neural networks,

    P. Ramoneda, D. Jeong, E. Nakamura, X. Serra, and M. Miron, “Automatic piano fingering from partially an- notated scores using autoregressive neural networks,” in Proceedings of the 30th ACM International Conference on Multimedia (ACM MM), 2022, pp. 6502–6510

  16. [24]

    Onsets and frames: Dual-objective piano transcription,

    C. Hawthorne, E. Elsen, J. Song, A. Roberts, I. Simon, C. Raffel, J. Engel, S. Oore, and E. D, “Onsets and frames: Dual-objective piano transcription,” inProceed- ings of the 19th International Society for Music Infor- mation Retrieval Conference (ISMIR), 2018, pp. 50–57

  17. [25]

    High- resolution piano transcription with pedals by regressing onset and offset times,

    Q. Kong, B. Li, X. Song, Y . Wan, and Y . Wang, “High- resolution piano transcription with pedals by regressing onset and offset times,”IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 29, pp. 3707–3717, 2021

  18. [26]

    Hppnet: Modeling the harmonic structure and pitch invariance in piano transcription,

    W. Wei, P. Li, Y . Yu, and W. Li, “Hppnet: Modeling the harmonic structure and pitch invariance in piano transcription,” inProceedings of the 23rd International Society for Music Information Retrieval Conference (ISMIR), 2022, pp. 709–716

  19. [27]

    Scoring time intervals using non- hierarchical transformer for automatic piano transcrip- tion,

    Y . Yan and Z. Duan, “Scoring time intervals using non- hierarchical transformer for automatic piano transcrip- tion,” inProceedings of the 25th International Society for Music Information Retrieval Conference (ISMIR), 2024, pp. 973–980

  20. [28]

    A data-driven analysis of robust automatic piano transcription,

    D. Edwards, S. Dixon, E. Benetos, A. Maezawa, and Y . Kusaka, “A data-driven analysis of robust automatic piano transcription,”IEEE Signal Processing Letters, vol. 31, pp. 681–685, 2024

  21. [29]

    Automatic piano music transcription using audio-visual features,

    Y . Wan, X. Wang, R. Zhou, and Y . Yan, “Automatic piano music transcription using audio-visual features,” Chinese Journal of Electronics, vol. 24, no. 3, pp. 597– 603, 2015

  22. [30]

    An audio-visual fusion piano transcription approach based on strategy,

    X. Wang, W. Xu, J. Liu, W. Yang, and W. Cheng, “An audio-visual fusion piano transcription approach based on strategy,” inProceedings of the 24th International Conference on Digital Audio Effects (DAFx), 2021, pp. 308–315

  23. [31]

    A two-stage audio-visual fusion piano transcription model based on the attention mechanism,

    Y . Li, X. Wang, R. Wu, W. Xu, and W. Cheng, “A two-stage audio-visual fusion piano transcription model based on the attention mechanism,”IEEE/ACM Trans- actions on Audio, Speech, and Language Processing, vol. 32, pp. 3618–3630, 2024

  24. [32]

    pyloudnorm: A simple yet flexible loudness meter in python,

    C. J. Steinmetz and J. D. Reiss, “pyloudnorm: A simple yet flexible loudness meter in python,” in150th AES Convention, 2021

  25. [33]

    Mediapipe hands: On-device real-time hand tracking,

    F. Zhang, V . Bazarevsky, A. Vakunov, A. Tkachenka, G. Sung, C.-L. Chang, and M. Grundmann, “Mediapipe hands: On-device real-time hand tracking,” 2020. [Online]. Available: https://arxiv.org/abs/2006.10214

  26. [34]

    Mir_eval: A transparent implementation of common mir metrics

    C. Raffel, B. McFee, E. J. Humphrey, J. Salamon, O. Ni- eto, D. Liang, D. P. Ellis, and C. C. Raffel, “Mir_eval: A transparent implementation of common mir metrics.” in Proceedings of the 15th International Society for Music Information Retrieval Conference (ISMIR), 2014, pp. 367–372

  27. [35]

    Towards robust transcription: Exploring noise injection strategies for training data augmentation,

    Y . Kim and A. Lerch, “Towards robust transcription: Exploring noise injection strategies for training data augmentation,” inLate Breaking Demo of the 25th Inter- national Society for Music Information Retrieval Con- ference (ISMIR), 2024

  28. [36]

    ViTPose++: Vision Transformer for Generic Body Pose Estimation,

    Y . Xu, J. Zhang, Q. Zhang, and D. Tao, “ViTPose++: Vision Transformer for Generic Body Pose Estimation,” IEEE Transactions on Pattern Analysis & Machine In- telligence, vol. 46, no. 02, pp. 1212–1230, 2024

  29. [37]

    Hamba: Single-view 3d hand reconstruction with graph-guided bi-scanning mamba,

    H. Dong, A. Chharia, W. Gou, F. Vicente Carrasco, and F. D. De la Torre, “Hamba: Single-view 3d hand reconstruction with graph-guided bi-scanning mamba,” Advances in Neural Information Processing Systems, vol. 37, pp. 2127–2160, 2024

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.