REVIEW 2 major objections 4 minor 80 references
VioPose: Violin Performance 4D Pose Estimation by Hierarchical Audiovisual Inference
T0 review · 2 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read VioPose estimates 4D violin-playing poses by using the music itself as a prior, beating visual-only methods on subtle motions such as vibrato.
desk verdict VioDat is a genuinely useful resource and the architecture is a sensible new combination, but the headline velocity/acceleration gains are confounded because VioPose is trained with extra losses that the baselines don't get. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The core mechanism is a three-layer hierarchy that treats audio as a prior for acceleration, cascades high-level dynamics down to lower layers through a summation that mimics a log-Bayesian update ($\log p(x|y) \approx \log p(y|x) + \log p(x)$), and then reconciles the estimated acceleration, velocity, and pose through a bidirectional mixing module that alternates integration and differentiation. The network consumes off-the-shelf 2D keypoints and a 35-dim audio feature vector, encodes both with transformer blocks, and is trained with a position loss plus max-cosine-similarity losses on velocity and acceleration, which forces the model to match motion dynamics rather than only average joint positions.
What would settle it
Retrain or test VioPose on VioDat with the input audio shifted by a fixed delay of about 100 ms relative to the video. If MPJPE, MPJVE, and MPJAE stay close to the synchronized results, the causal-audio prior is not the source of the improvements; a substantial degradation would confirm that the audio is doing the claimed work.
Extended reading notes
Core claim
The paper claims that estimating motion dynamics hierarchically, with acceleration as the highest-level signal and pose as the lowest, and feeding audio into the highest level as a causal prior, yields 4D pose sequences that are both more accurate and smoother than visual-only methods. On VioDat, VioPose reports MPJPE of 43.60 mm versus 46.09 mm for the best retrained visual-only baseline, MPJVE of 1.57 versus 1.90, and MPJAE of 1.02 versus 1.69. The largest joint-level gains are in the upper limbs, such as the left index and right wrist, and the method tracks vibrato perturbations of roughly 10 mm that other methods flatten into straight lines.
Load-bearing premise
In violin playing, the audio is caused by the same body motion that the pose estimator is trying to recover, so the sound can be trusted as a prior; if the audio track is desynchronized, contains unrelated sounds, or the player is miming, the reported gains should disappear.
Editorial extensions
If this is right
- Monocular pose estimation for instrument playing can be improved by treating the instrument's sound as a causal signal rather than a merely correlated one.
- Velocity and acceleration errors, not just joint positions, are the metrics that expose whether a model captures subtle motion such as vibrato.
- The same causal-audio prior should transfer to other instruments or activities where sound is physically produced by the body motion being estimated.
- A calibrated audiovisual dataset with motion-capture ground truth is necessary to fairly train and compare such methods, because existing music datasets lack full 3D kinematic ground truth.
Reading between the lines
- Beyond the paper, the causal-audio prior should transfer to other motion-produced sounds such as piano, drums, or speech articulation whenever the audio track can be isolated from background noise.
- Beyond the paper, the hierarchy's success suggests a general recipe for lifting methods: predict acceleration and velocity before position, and use the highest-derivative signal as the prior, which could improve fine-grained movement estimation in domains like typing or surgical motion.
- Beyond the paper, a direct way to test the claimed mechanism is to shift the input audio by a known delay, such as 100 ms, and measure how much the pose, velocity, and acceleration errors degrade; the paper does not report this experiment.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes VioPose, a multimodal 3D pose estimation network for violin performance that takes 2D keypoints and raw audio as input and outputs 3D pose, velocity, and acceleration. It introduces VioDat, a synchronized violin dataset with 12 players, four cameras, four microphones, and MoCap ground truth. The method uses a hierarchical architecture with cascade summation and bidirectional mixing, trained with pose, velocity, and acceleration losses. Experiments on VioDat show VioPose outperforms several visual-only SoTA baselines on MPJPE and especially on MPJVE/MPJAE, and an ablation demonstrates gains from the audio module. A downstream violin-performance analysis shows improvements in bowing direction, straight bow, violin hold, and vibrato detection.
Significance. If the results hold, the dataset alone is a significant contribution, filling a gap in calibrated audiovisual performance data. The hierarchical fusion and bidirectional mixing are plausible design choices, and the ablations are extensive, particularly the audio on/off comparison that isolates the audio contribution. The paper also promises code and dataset release, which supports reproducibility. However, the reported gains over SoTA on velocity and acceleration metrics are confounded with the training loss, and the test set is very small, so the magnitude and generalizability of the claimed improvements are not yet established. The central MPJPE improvement is plausible, but the fine-motion claims need a fairer evaluation protocol.
major comments (2)
- [§3.5 and §4.3, Eq. (8)] The large MPJVE and MPJAE improvements over SoTA in Table 2 are confounded: VioPose is trained with velocity and acceleration losses L_v and L_a, while the SoTA baselines are retrained with their original loss functions, which typically only supervise pose. Because VioPose directly optimizes the metrics on which it is evaluated, the claimed 17.37% and 39.64% improvements do not isolate the proposed architecture or the audio fusion. This is supported by Table 2, where VioPose w/o audio, trained with the same additional losses, already beats every baseline on MPJVE and MPJAE. Please retrain the baselines with the same velocity and acceleration losses, or remove those losses from VioPose, and report the comparison; without this, the fine-motion and vibrato claims are not established.
- [§4.1] The test set contains only three participants (an advanced male adult, a novice female teenager, and a novice male child), yet the paper reports point estimates without error bars, confidence intervals, or per-subject breakdowns. With n=3, the 5.4% MPJPE improvement and the MPJVE/MPJAE margins may not be statistically reliable. Please report per-subject results and uncertainty estimates, and ideally include more test subjects or cross-validation, to support the generalization claims made in the abstract and Section 4.3.
minor comments (4)
- [§4.4, Table 5] The sentence 'the w/o Cascade model shows the worst result in MPJPE' is inconsistent with the table: w/o Mixing has MPJPE 47.87, which is the worst, while w/o Cascade has MPJPE 44.21. Please correct the description or the table.
- [§3.5, Eq. (8)] The definitions of L_v and L_a are incomplete: as written, they sum the dot products of normalized vectors and do not clearly show the max operation or the normalization denominator that would define the max-cosine similarity. Please clarify the formula and its relation to the cited reference.
- [§3.2] In the sentence about the bottleneck layer, 'block box in Fig 2' should read 'black box in Fig 2'.
- [§4.5 and Figures 4-5] The qualitative trajectory plots are helpful, but they show only a few examples; adding quantitative trajectory errors per motion type (e.g., bowing vs. vibrato) would strengthen the claim that VioPose captures fine motions.
Circularity Check
No significant circularity: the central claim is a held-out empirical benchmark and the audio prior is learned end-to-end, with no parameter fitted from the target result.
full rationale
VioPose is an empirical supervised-learning paper; its central claim is benchmark performance on a held-out test split of the newly collected VioDat dataset. The hierarchical audiovisual fusion (Eqs. 2-7) is an architectural choice, not a derivation of the result from the causal-audio premise; the audio prior is learned end-to-end from synchronized recordings and its contribution is tested by ablation (VioPose w/o audio in Tables 2, 5, and 7). No parameter is fitted from the test-set quantities, and no uniqueness theorem or prior result by the same authors is invoked to force the design. The only author-overlap citations ([10, 26, 46, 61]) appear in related-work or future-work contexts and are not load-bearing. The potential concern that VioPose is trained with velocity and acceleration losses (Eq. 8) that match the MPJVE/MPJAE evaluation metrics is a fairness-of-comparison issue for the benchmark rather than a circularity of the derivation; the reported numbers are measurements on unseen participants, not identities forced by the loss. Accordingly, no circular step is identified.
Assumptions & free parameters
free parameters (3)
- lambda_v =
not reported
- lambda_a =
not reported
- Input frame count f =
150
assumptions (4)
- domain assumption The audio signal is causally generated by the same human motion that produces the 3D pose, so it carries reliable information about pose dynamics.
- domain assumption MediaPipe 2D keypoints provide a sufficiently accurate initialization for 3D lifting.
- domain assumption The Vicon mocap ground truth (0.1 mm accuracy) and the 4-camera synchronization are correct.
- standard math Transformer self-attention and 1D CNN components work as standard deep learning modules.
Cite this review
Pith. "Pith review of VioPose: Violin Performance 4D Pose Estimation by Hierarchical Audiovisual Inference." pith.science (2026). https://pith.science/paper/GAJ54CYR
@misc{pith2026241113607,
author = {Pith},
title = {Pith review of: VioPose: Violin Performance 4D Pose Estimation by Hierarchical Audiovisual Inference},
year = {2026},
howpublished = {\url{https://pith.science/paper/GAJ54CYR}},
note = {Machine review of arXiv:2411.13607}
}
read the original abstract
Musicians delicately control their bodies to generate music. Sometimes, their motions are too subtle to be captured by the human eye. To analyze how they move to produce the music, we need to estimate precise 4D human pose (3D pose over time). However, current state-of-the-art (SoTA) visual pose estimation algorithms struggle to produce accurate monocular 4D poses because of occlusions, partial views, and human-object interactions. They are limited by the viewing angle, pixel density, and sampling rate of the cameras and fail to estimate fast and subtle movements, such as in the musical effect of vibrato. We leverage the direct causal relationship between the music produced and the human motions creating them to address these challenges. We propose VioPose: a novel multimodal network that hierarchically estimates dynamics. High-level features are cascaded to low-level features and integrated into Bayesian updates. Our architecture is shown to produce accurate pose sequences, facilitating precise motion analysis, and outperforms SoTA. As part of this work, we collected the largest and the most diverse calibrated violin-playing dataset, including video, sound, and 3D motion capture poses. Code and dataset can be found in our project page \url{https://sj-yoo.info/viopose/}.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
“A Lip Sync Expert Is All You Need for Speech to Lip Genera- tion In the Wild — Proceedings of the 28th ACM International Conference on Multimedia.” 2
-
[2]
Synthesizing Obama: Learning lip sync from audio: ACM Transactions on Graphics: V ol 36, No 4
“Synthesizing Obama: Learning lip sync from audio: ACM Transactions on Graphics: V ol 36, No 4.” 2
-
[3]
Self- supervised Learning of Audio-Visual Objects from Video,
T. Afouras, A. Owens, J. S. Chung, and A. Zisserman, “Self- supervised Learning of Audio-Visual Objects from Video,” in Computer Vision – ECCV 2020 , ser. Lecture Notes in Computer Science, A. Vedaldi, H. Bischof, T. Brox, and J.-M. Frahm, Eds. Cham: Springer International Publishing, 2020, pp. 208–224. 2
work page 2020
-
[4]
Groovenet: Real- time music-driven dance movement generation using artificial neural networks,
O. Alemi, J. Fran c ¸oise, and P. Pasquier, “Groovenet: Real- time music-driven dance movement generation using artificial neural networks,” networks, vol. 8, no. 17, p. 26, 2017. 3
work page 2017
-
[5]
R. Arandjelovic and A. Zisserman, “Look, Listen and Learn,” in Proceedings of the IEEE International Conference on Com- puter Vision, 2017, pp. 609–617. 1
work page 2017
-
[6]
A. D. Blanco, S. Tassani, and R. Ramirez, “Real-Time Sound and Motion Feedback for Violin Bow Technique Learning: A Controlled, Randomized Trial,” Frontiers in Psychology, vol. 12, 2021. 3
work page 2021
-
[7]
Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional Networks,
Y . Cai, L. Ge, J. Liu, J. Cai, T.-J. Cham, J. Yuan, and N. M. Thalmann, “Exploiting Spatial-Temporal Relationships for 3D Pose Estimation via Graph Convolutional Networks,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 2272–2281. 2
work page 2019
-
[8]
SoundSpaces: Audio-Visual Navigation in 3D Environments,
C. Chen, U. Jain, C. Schissler, S. V . A. Gari, Z. Al-Halah, V . K. Ithapu, P. Robinson, and K. Grauman, “SoundSpaces: Audio-Visual Navigation in 3D Environments,” inComputer Vision – ECCV 2020 , ser. Lecture Notes in Computer Sci- ence, A. Vedaldi, H. Bischof, T. Brox, and J.-M. Frahm, Eds. Cham: Springer International Publishing, 2020, pp. 17–36. 2
work page 2020
Show all 80 references
-
[9]
MM-ViT: Multi-Modal Video Trans- former for Compressed Video Action Recognition,
J. Chen and C. M. Ho, “MM-ViT: Multi-Modal Video Trans- former for Compressed Video Action Recognition,” in Pro- ceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, 2022, pp. 1910–1921. 2
2022
-
[10]
Active Human Pose Estimation via an Autonomous UA V Agent,
J. Chen, B. He, C. D. Singh, C. Ferm ¨uller, and Y . Aloimonos, “Active Human Pose Estimation via an Autonomous UA V Agent,” Jul. 2024. [Online]. Available: http://arxiv.org/abs/2407.01811 2
2024 arXiv
-
[11]
Anatomy-Aware 3D Human Pose Estimation With Bone- Based Pose Decomposition,
T. Chen, C. Fang, X. Shen, Y . Zhu, Z. Chen, and J. Luo, “Anatomy-Aware 3D Human Pose Estimation With Bone- Based Pose Decomposition,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 32, no. 1, pp. 198–209, Jan. 2022. 2
2022
-
[12]
DiffuPose: Monocular 3D Human Pose Estimation via Denoising Diffusion Probabilistic Model,
J. Choi, D. Shim, and H. J. Kim, “DiffuPose: Monocular 3D Human Pose Estimation via Denoising Diffusion Probabilistic Model,” Dec. 2022. 1
2022
-
[13]
Optimizing Network Structure for 3D Human Pose Estimation,
H. Ci, C. Wang, X. Ma, and Y . Wang, “Optimizing Network Structure for 3D Human Pose Estimation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 2262–2271. 2
2019
-
[14]
Bowing Gestures Classifica- tion in Violin Performance: A Machine Learning Approach,
D. Dalmazzo and R. Ram´ırez, “Bowing Gestures Classifica- tion in Violin Performance: A Machine Learning Approach,” Frontiers in Psychology, vol. 10, 2019. 3
2019
-
[15]
A Machine Learn- ing Approach to Violin Bow Technique Classification: A Comparison Between IMU and MOCAP systems,
D. Dalmazzo, S. Tassani, and R. Ram´ırez, “A Machine Learn- ing Approach to Violin Bow Technique Classification: A Comparison Between IMU and MOCAP systems,” in Pro- ceedings of the 5th International Workshop on Sensor-based Activity Recognition and Interaction, ser. iWOAR ’18...
2018
-
[16]
Audiovisual Analysis of Music Performances: Overview of an Emerging Field,
Z. Duan, S. Essid, C. C. Liem, G. Richard, and G. Sharma, “Audiovisual Analysis of Music Performances: Overview of an Emerging Field,” IEEE Signal Processing Magazine, vol. 36, no. 1, pp. 63–73, Jan. 2019. 1, 3
2019
-
[17]
Uplift and Up- sample: Efficient 3D Human Pose Estimation With Uplift- ing Transformers,
M. Einfalt, K. Ludwig, and R. Lienhart, “Uplift and Up- sample: Efficient 3D Human Pose Estimation With Uplift- ing Transformers,” in Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision , 2023, pp. 2903–2913. 5
2023
-
[18]
Fusion of Multimodal Informa- tion in Music Content Analysis,
S. Essid and G. Richard, “Fusion of Multimodal Informa- tion in Music Content Analysis,” in Multimodal Music Pro- cessing, ser. Dagstuhl Follow-Ups, M. M ¨uller, M. Goto, and M. Schedl, Eds. Dagstuhl, Germany: Schloss Dagstuhl–Leibniz-Zentrum fuer Informatik, 2012, vol. 3, pp....
2012
-
[19]
Foley Music: Learning to Generate Music from Videos,
C. Gan, D. Huang, P. Chen, J. B. Tenenbaum, and A. Torralba, “Foley Music: Learning to Generate Music from Videos,” in Computer Vision – ECCV 2020 , ser. Lecture Notes in Computer Science, A. Vedaldi, H. Bischof, T. Brox, and J.-M. Frahm, Eds. Cham: Springer International Publ...
2020
-
[20]
Look, Listen, and Act: Towards Audio-Visual Embodied Nav- igation,
C. Gan, Y . Zhang, J. Wu, B. Gong, and J. B. Tenenbaum, “Look, Listen, and Act: Towards Audio-Visual Embodied Nav- igation,” in 2020 IEEE International Conference on Robotics and Automation (ICRA), May 2020, pp. 9701–9707. 2
2020
-
[21]
Automated Violin Bowing Gesture Recog- nition Using FMCW-Radar and Machine Learning,
H. Gao and C. Li, “Automated Violin Bowing Gesture Recog- nition Using FMCW-Radar and Machine Learning,” IEEE Sensors Journal, vol. 23, no. 9, pp. 9262–9270, May 2023. 3
2023
-
[22]
VisualEchoes: Spatial Image Representation Learn- ing Through Echolocation,
R. Gao, C. Chen, Z. Al-Halah, C. Schissler, and K. Grau- man, “VisualEchoes: Spatial Image Representation Learn- ing Through Echolocation,” in Computer Vision – ECCV 2020, ser. Lecture Notes in Computer Science, A. Vedaldi, H. Bischof, T. Brox, and J.-M. Frahm, Eds. Cham: Spri...
2020
-
[23]
Learning Individual Styles of Conversational Gesture,
S. Ginosar, A. Bar, G. Kohavi, C. Chan, A. Owens, and J. Ma- lik, “Learning Individual Styles of Conversational Gesture,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 3497–3506. 2 9
2019
-
[24]
Body Movement in Music Information Retrieval,
R. I. Godøy and A. R. Jensenius, “Body Movement in Music Information Retrieval,” pp. 45–50, 2009. 1
2009
-
[25]
DiffPose: Toward More Reliable 3D Pose Estimation,
J. Gong, L. G. Foo, Z. Fan, Q. Ke, H. Rahmani, and J. Liu, “DiffPose: Toward More Reliable 3D Pose Estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, 2023, pp. 13 041–13 051. 1
2023
-
[26]
Discov- ering a language for human activity,
G. Guerra-Filho, C. Ferm¨uller, and Y . Aloimonos, “Discov- ering a language for human activity,” in Proceedings of the AAAI 2005 Fall Symposium on Anticipatory Cognitive Em- bodied Systems, 2005. 8
2005
-
[27]
Deep Hierarchical Planning from Pixels,
D. Hafner, K.-H. Lee, I. Fischer, and P. Abbeel, “Deep Hierarchical Planning from Pixels,” Advances in Neural Information Processing Systems , vol. 35, pp. 26 091–26 104, Dec. 2022. [Online]. Avail- able: https://proceedings.neurips.cc/paper files/paper/ 2022/hash/a766f56d2da4...
2022
-
[28]
Resolv- ing 3D Human Pose Ambiguities With 3D Scene Constraints,
M. Hassan, V . Choutas, D. Tzionas, and M. J. Black, “Resolv- ing 3D Human Pose Ambiguities With 3D Scene Constraints,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 2282–2292. 2
2019
-
[29]
Exploiting temporal infor- mation for 3D human pose estimation,
M. R. I. Hossain and J. J. Little, “Exploiting temporal infor- mation for 3D human pose estimation,” in Proceedings of the European Conference on Computer Vision (ECCV), 2018, pp. 68–84. 1
2018
-
[30]
Hu- man3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments,
C. Ionescu, D. Papava, V . Olaru, and C. Sminchisescu, “Hu- man3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 36, no. 7, pp. 1325–1339, 2014. 3
2014
-
[31]
End- to-End Recovery of Human Shape and Pose,
A. Kanazawa, M. J. Black, D. W. Jacobs, and J. Malik, “End- to-End Recovery of Human Shape and Pose,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 7122–7131. 2
2018
-
[32]
Temporally Guided Music-to-Body- Movement Generation,
H.-K. Kao and L. Su, “Temporally Guided Music-to-Body- Movement Generation,” in Proceedings of the 28th ACM International Conference on Multimedia, ser. MM ’20. New York, NY , USA: Association for Computing Machinery, 10, 12, 2020, pp. 147–155. 1
2020
-
[33]
EPIC- Fusion: Audio-Visual Temporal Binding for Egocentric Ac- tion Recognition,
E. Kazakos, A. Nagrani, A. Zisserman, and D. Damen, “EPIC- Fusion: Audio-Visual Temporal Binding for Egocentric Ac- tion Recognition,” in Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision, 2019, pp. 5492–5501. 2
2019
-
[34]
Dancing to Music,
H.-Y . Lee, X. Yang, M.-Y . Liu, T.-C. Wang, Y .-D. Lu, M.- H. Yang, and J. Kautz, “Dancing to Music,” inAdvances in Neural Information Processing Systems , vol. 32. Curran Associates, Inc., 2019. 1
2019
-
[35]
Propagating LSTM: 3D Pose Estimation based on Joint Interdependency,
K. Lee, I. Lee, and S. Lee, “Propagating LSTM: 3D Pose Estimation based on Joint Interdependency,” in Proceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 119–135. 1
2018
-
[36]
An Interdis- ciplinary Review of Music Performance Analysis,
A. Lerch, C. Arthur, A. Pati, and S. Gururani, “An Interdis- ciplinary Review of Music Performance Analysis,” Trans- actions of the International Society for Music Information Retrieval, vol. 3, no. 1, pp. 221–245, Nov. 2020. 1
2020
-
[37]
Cre- ating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applica- tions,
B. Li, X. Liu, K. Dinesh, Z. Duan, and G. Sharma, “Cre- ating a multitrack classical music performance dataset for multimodal music analysis: Challenges, insights, and applica- tions,” IEEE Transactions on Multimedia, vol. 21, no. 2, pp. 522–535, 2018. 3
2018
-
[38]
Audio-Visual End-to-End Multi- Channel Speech Separation, Dereverberation and Recogni- tion,
G. Li, J. Deng, M. Geng, Z. Jin, T. Wang, S. Hu, M. Cui, H. Meng, and X. Liu, “Audio-Visual End-to-End Multi- Channel Speech Separation, Dereverberation and Recogni- tion,” IEEE/ACM Transactions on Audio, Speech, and Lan- guage Processing, vol. 31, pp. 2707–2723, 2023. 2
2023
-
[39]
HybrIK: A Hybrid Analytical-Neural Inverse Kinematics Solution for 3D Human Pose and Shape Estimation,
J. Li, C. Xu, Z. Chen, S. Bian, L. Yang, and C. Lu, “HybrIK: A Hybrid Analytical-Neural Inverse Kinematics Solution for 3D Human Pose and Shape Estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 3383–3393. 2
2021
-
[40]
NIKI: Neural Inverse Kinematics With Invertible Neural Networks for 3D Human Pose and Shape Estimation,
J. Li, S. Bian, Q. Liu, J. Tang, F. Wang, and C. Lu, “NIKI: Neural Inverse Kinematics With Invertible Neural Networks for 3D Human Pose and Shape Estimation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 12 933–12 942. 2
2023
-
[41]
AI Chore- ographer: Music Conditioned 3D Dance Generation With AIST++,
R. Li, S. Yang, D. A. Ross, and A. Kanazawa, “AI Chore- ographer: Music Conditioned 3D Dance Generation With AIST++,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 13 401–13 412. 3, 4, 7
2021
-
[42]
Learning Visual Styles from Audio-Visual Associations,
T. Li, Y . Liu, A. Owens, and H. Zhao, “Learning Visual Styles from Audio-Visual Associations,” inComputer Vision – ECCV 2022, ser. Lecture Notes in Computer Science, S. Avi- dan, G. Brostow, M. Ciss´e, G. M. Farinella, and T. Hassner, Eds. Cham: Springer Nature Switzerland, 2...
2022
-
[43]
Ex- ploiting Temporal Contexts With Strided Transformer for 3D Human Pose Estimation,
W. Li, H. Liu, R. Ding, M. Liu, P. Wang, and W. Yang, “Ex- ploiting Temporal Contexts With Strided Transformer for 3D Human Pose Estimation,” IEEE Transactions on Multimedia, vol. 25, pp. 1282–1293, 2022. 1, 5, 6, 8
2022
-
[44]
MHFormer: Multi-Hypothesis Transformer for 3D Hu- man Pose Estimation,
W. Li, H. Liu, H. Tang, P. Wang, and L. Van Gool, “MHFormer: Multi-Hypothesis Transformer for 3D Hu- man Pose Estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog- nition, 2022, pp. 13 147–13 156. [Online]. Available: https://openaccess.t...
2022
-
[45]
TokenPose: Learning Keypoint Tokens for Human Pose Estimation,
Y . Li, S. Zhang, Z. Wang, S. Yang, W. Yang, S.-T. Xia, and E. Zhou, “TokenPose: Learning Keypoint Tokens for Human Pose Estimation,” in Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision , 2021, pp. 11 313– 11 322. 1 10
2021
-
[46]
Learning shift- invariant sparse representation of actions,
Y . Li, C. Ferm¨uller, Y . Aloimonos, and H. Ji, “Learning shift- invariant sparse representation of actions,” in2010 IEEE Com- puter Society Conference on Computer Vision and Pattern Recognition. IEEE, 2010, pp. 2630–2637. 8
2010
-
[47]
MediaPipe: A Framework for Building Perception Pipelines,
C. Lugaresi, J. Tang, H. Nash, C. McClanahan, E. Uboweja, M. Hays, F. Zhang, C.-L. Chang, M. G. Yong, J. Lee, W.-T. Chang, W. Hua, M. Georg, and M. Grundmann, “MediaPipe: A Framework for Building Perception Pipelines,” Jun. 2019. 4, 5
2019
-
[48]
M2C: Concise Music Repre- sentation for 3D Dance Generation,
M. Marchellus and I. K. Park, “M2C: Concise Music Repre- sentation for 3D Dance Generation,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3126–3135. 1, 4, 7
2023
-
[49]
Hearing lips and seeing voices,
H. Mcgurk and J. Macdonald, “Hearing lips and seeing voices,” Nature, vol. 264, no. 5588, pp. 746–748, Dec. 1976. 1
1976
-
[50]
Marker-less piano fingering recognition using sequential depth images,
A. Oka and M. Hashimoto, “Marker-less piano fingering recognition using sequential depth images,” in The 19th Korea-Japan Joint Workshop on Frontiers of Computer Vision, Jan. 2013, pp. 1–4. 3
2013
-
[51]
A computational approach to studying interde- pendence in string quartet performance,
P. Papiotis, “A computational approach to studying interde- pendence in string quartet performance,” phdphd, Univesitat Pompeu Fabra, 2016. 3
2016
-
[52]
SyncTalk- Face: Talking Face Generation with Precise Lip-Syncing via Audio-Lip Memory,
S. J. Park, M. Kim, J. Hong, J. Choi, and Y . M. Ro, “SyncTalk- Face: Talking Face Generation with Precise Lip-Syncing via Audio-Lip Memory,” Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, no. 2, pp. 2062–2070, Jun
-
[53]
Ordinal Depth Supervision for 3D Human Pose Estimation,
G. Pavlakos, X. Zhou, and K. Daniilidis, “Ordinal Depth Supervision for 3D Human Pose Estimation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 7307–7316. 1, 2
2018
-
[54]
3D Human Pose Estimation in Video With Temporal Convo- lutions and Semi-Supervised Training,
D. Pavllo, C. Feichtenhofer, D. Grangier, and M. Auli, “3D Human Pose Estimation in Video With Temporal Convo- lutions and Semi-Supervised Training,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019, pp. 7753–7762. 1, 2
2019
-
[55]
Effective use of multimedia for computer-assisted musical instrument tutor- ing,
G. Percival, Y . Wang, and G. Tzanetakis, “Effective use of multimedia for computer-assisted musical instrument tutor- ing,” in Proceedings of the International Workshop on Educa- tional Multimedia and Multimedia Education, ser. Emme ’07. New York, NY , USA: Association for Co...
2007
-
[56]
Music-driven Dance Regeneration with Controllable Key Pose Constraints,
J. Pu and Y . Shan, “Music-driven Dance Regeneration with Controllable Key Pose Constraints,” Jul. 2022. 1
2022
-
[57]
A Joint Cross-Attention Model for Audio-Visual Fusion in Dimensional Emotion Recognition,
G. P. Rajasekar, W. C. de Melo, N. Ullah, H. Aslam, O. Zee- shan, T. Denorme, M. Pedersoli, A. Koerich, S. Bacon, P. Car- dinal, and E. Granger, “A Joint Cross-Attention Model for Audio-Visual Fusion in Dimensional Emotion Recognition,” Apr. 2022. 2
2022
-
[58]
P-STMO: Pre-trained Spatial Temporal Many-to-One Model for 3D Human Pose Estimation,
W. Shan, Z. Liu, X. Zhang, S. Wang, S. Ma, and W. Gao, “P-STMO: Pre-trained Spatial Temporal Many-to-One Model for 3D Human Pose Estimation,” inComputer Vision – ECCV 2022, ser. Lecture Notes in Computer Science, S. Avidan, G. Brostow, M. Ciss´e, G. M. Farinella, and T. Hassne...
2022
-
[59]
PhysCap: Physically plausible monocular 3D motion capture in real time,
S. Shimada, V . Golyanik, W. Xu, and C. Theobalt, “PhysCap: Physically plausible monocular 3D motion capture in real time,” ACM Transactions on Graphics , vol. 39, no. 6, pp. 235:1–235:16, 11, 27, 2020. 2
2020
-
[60]
Audio to Body Dynamics,
E. Shlizerman, L. Dery, H. Schoen, and I. Kemelmacher- Shlizerman, “Audio to Body Dynamics,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recog- nition, 2018, pp. 7574–7583. 1
2018
-
[61]
AIMusicGuru: Music Assisted Human Pose Correction,
S. Shrestha, C. Ferm¨uller, T. Huang, P. T. Win, A. Zukerman, C. M. Parameshwara, and Y . Aloimonos, “AIMusicGuru: Music Assisted Human Pose Correction,” Mar. 2022. 1
2022
-
[62]
HumanEva: Synchro- nized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion,
L. Sigal, A. Balan, and M. J. Black, “HumanEva: Synchro- nized video and motion capture dataset and baseline algorithm for evaluation of articulated human motion,” International Journal of Computer Vision, vol. 87, no. 1, pp. 4–27, Mar
-
[63]
Audeo: Audio generation for a silent performance video,
K. Su, X. Liu, and E. Shlizerman, “Audeo: Audio generation for a silent performance video,” Advances in Neural Informa- tion Processing Systems, vol. 33, pp. 3325–3337, 2020. 1, 7
2020
-
[64]
Physics-Driven Diffusion Models for Impact Sound Synthesis From Videos,
K. Su, K. Qian, E. Shlizerman, A. Torralba, and C. Gan, “Physics-Driven Diffusion Models for Impact Sound Synthesis From Videos,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 9749–
2023
-
[65]
3D Human Pose Estimation With Spatio-Temporal Criss-Cross Attention,
Z. Tang, Z. Qiu, Y . Hao, R. Hong, and T. Yao, “3D Human Pose Estimation With Spatio-Temporal Criss-Cross Attention,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 4790–4799. 1, 2
2023
-
[66]
Sight over sound in the judgment of music per- formance,
C.-J. Tsay, “Sight over sound in the judgment of music per- formance,” Proceedings of the National Academy of Sciences, vol. 110, no. 36, pp. 14 580–14 585, Sep. 2013. 1, 3
2013
-
[67]
Attention is All you Need,
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is All you Need,” in Advances in Neural Information Processing Systems, vol. 30. Curran Associates, Inc., 2017. 7
2017
-
[68]
A multimodal corpus for technology-enhanced learning of vi- olin playing,
G. V olpe, K. Kolykhalova, E. V olta, S. Ghisio, G. Waddell, P. Alborno, S. Piana, C. Canepa, and R. Ramirez-Melendez, “A multimodal corpus for technology-enhanced learning of vi- olin playing,” inProceedings of the 12th Biannual Conference on Italian SIGCHI Chapter, 2017, pp. 1–5. 3
2017
-
[69]
Con- volutional Pose Machines,
S.-E. Wei, V . Ramakrishna, T. Kanade, and Y . Sheikh, “Con- volutional Pose Machines,” in Proceedings of the IEEE Con- ference on Computer Vision and Pattern Recognition, 2016, pp. 4724–4732. 2 11
2016
-
[70]
Deep Kinematics Analysis for Monocular 3D Human Pose Esti- mation,
J. Xu, Z. Yu, B. Ni, J. Yang, X. Yang, and W. Zhang, “Deep Kinematics Analysis for Monocular 3D Human Pose Esti- mation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 899–908. 2
2020
-
[71]
U-shaped spa- tial–temporal transformer network for 3D human pose estima- tion,
H. Yang, L. Guo, Y . Zhang, and X. Wu, “U-shaped spa- tial–temporal transformer network for 3D human pose estima- tion,” Machine Vision and Applications, vol. 33, no. 6, p. 82, Sep. 2022. 2
2022
-
[72]
Automatic Music Transcription using Audio-Visual Fusion for Violin Practice in Home Environ- ment,
B. Zhang and Y . Wang, “Automatic Music Transcription using Audio-Visual Fusion for Violin Practice in Home Environ- ment,” Jul. 2009. 3
2009
-
[73]
MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video,
J. Zhang, Z. Tu, J. Yang, Y . Chen, and J. Yuan, “MixSTE: Seq2seq Mixed Spatio-Temporal Encoder for 3D Human Pose Estimation in Video,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2022, pp. 13 232–13 242. 5, 6, 8
2022
-
[74]
The Sound of Pixels,
H. Zhao, C. Gan, A. Rouditchenko, C. V ondrick, J. McDer- mott, and A. Torralba, “The Sound of Pixels,” inProceedings of the European Conference on Computer Vision (ECCV) , 2018, pp. 570–586. 2
2018
-
[75]
The Sound of Motions,
H. Zhao, C. Gan, W.-C. Ma, and A. Torralba, “The Sound of Motions,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 1735–1744. 2
2019
-
[76]
Pose- FormerV2: Exploring Frequency Domain for Efficient and Robust 3D Human Pose Estimation,
Q. Zhao, C. Zheng, M. Liu, P. Wang, and C. Chen, “Pose- FormerV2: Exploring Frequency Domain for Efficient and Robust 3D Human Pose Estimation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 8877–8886. 1, 2, 5, 6, 8
2023
-
[77]
3D Human Pose Estimation With Spatial and Tem- poral Transformers,
C. Zheng, S. Zhu, M. Mendieta, T. Yang, C. Chen, and Z. Ding, “3D Human Pose Estimation With Spatial and Tem- poral Transformers,” in Proceedings of the IEEE/CVF Inter- national Conference on Computer Vision, 2021, pp. 11 656– 11 665. 2, 4
2021
-
[78]
Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation,
H. Zhou, Y . Sun, W. Wu, C. C. Loy, X. Wang, and Z. Liu, “Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 4176–4186. 2
2021
-
[79]
Quantized GAN for Complex Music Gen- eration from Dance Videos,
Y . Zhu, K. Olszewski, Y . Wu, P. Achlioptas, M. Chai, Y . Yan, and S. Tulyakov, “Quantized GAN for Complex Music Gen- eration from Dance Videos,” in Computer Vision – ECCV 2022, ser. Lecture Notes in Computer Science, S. Avidan, G. Brostow, M. Ciss´e, G. M. Farinella, and T. ...
2022
-
[80]
Mu- sic2dance: Dancenet for music-driven dance generation,
W. Zhuang, C. Wang, S. Xia, J. Chai, and Y . Wang, “Mu- sic2dance: Dancenet for music-driven dance generation,” arXiv preprint arXiv:2002.03761, 2020. 3 12
2002 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.