REVIEW 3 major objections 5 minor 1 cited by
BimArt: A Unified Approach for the Synthesis of 3D Bimanual Interaction with Articulated Objects
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read BimArt generates realistic two-handed animations for articulated objects from only the object's 7-degree-of-freedom trajectory, without needing a reference grasp, a coarse hand trajectory, or separate grasping and articulating stages.
desk verdict Genuinely new representation and solid ablations, but the SOTA claim outruns the numbers; worth reviewing with major revisions. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing objects are the articulation-aware part-based BPS features, the bimanual contact maps, and the two diffusion models. Basis Point Sets (BPS) encode a geometry by storing, for each fixed basis point in space, the vector to the nearest object vertex; BimArt's extension computes these vectors separately for the top and bottom articulated parts after normalizing the object's scale, so that a bottle lid is sampled as densely as the bottle body. The contact maps are per-hand per-frame arrays that store, on each sampled object vertex, the minimum distance to any hand vertex; because they are distance-based and spatially embedded on the object, they reveal rich bimanual grasping patterns while leaving hand-object correspondence unspecified, which preserves diversity. The motion model generates hand surface keypoints plus direction vectors to the nearest object vertex, and uses the predicted contact maps in two ways: as a conditioning token with random dropout (classifier-free guidance) and as a gradient-based discrepancy target during denoising, pulling the generated hand geometry toward the predicted contact regions. The final MANO fitting optimization removes residual penetration and jitter with projection, penetration, and acceleration energy terms.
What would settle it
Take an articulated object not in the training categories whose maximum-extent articulation angle differs from the heuristic choices (0, $\pi/2$, or $\pi$), such as a folding chair or compound hinge, and run BimArt with its true trajectory: if penetration and contact or articulation percentages degrade sharply relative to the same object canonicalized with its correct angle, the heuristic is the load-bearing bottleneck. A sharper test is to feed a prismatic-joint object such as a sliding drawer, which violates the paper's two-part rotational-joint canonical frame (articulation axis aligned with the negative z-axis); the model should fail to produce stable contact on the moving part, delimiting the method's scope to rotational two-part articulated objects.
Extended reading notes
Core claim
On its own terms, the paper establishes that bimanual manipulation with articulated objects can be decomposed into contact-map prediction followed by contact-conditioned motion synthesis, and that this decomposition is what removes the need for grasp references. The contact generation network, a transformer-based denoising diffusion model, predicts per-frame distance-based contact maps for the left and right hands directly from an articulation-aware object encoding; the motion network then generates hand surface keypoints and object-direction vectors, conditioned on those maps with classifier-free guidance and an extra gradient term that aligns the generated contact geometry with the predicted maps at every denoising step. The object encoding that makes both stages category-agnostic is a normalized, part-based Basis Point Set representation in which the same basis points are mapped separately to each articulated part after scale normalization, so small moving parts receive equal sampling density to large static parts. A final optimization-based MANO fitting step with projection, penetration, and acceleration energies turns the generated keypoints into clean hand meshes. The paper demonstrates, via quantitative metrics and a forced-choice user study, that this pipeline produces motions judged more natural than adapted versions of several prior methods.
Load-bearing premise
The method's object encoding relies on a per-category guess about which articulation angle makes each object span its largest extent when normalizing scale; if that guess is wrong for an unfamiliar object, every downstream quantity—the BPS features, the contact maps, and the hand guidance—is computed from a distorted geometry, and the model's generalization degrades.
Editorial extensions
If this is right
- Animators and VR developers can generate multiple plausible two-handed interactions from a single object trajectory, without authoring grasps by hand.
- A single model handles multiple object categories and can execute object translation, rotation, and articulation simultaneously, which prior articulated-object methods could not do together.
- Because contact maps are an explicit intermediate output, users or downstream systems could edit or re-target those maps to steer hand placement while keeping the rest of the motion generation intact.
- The low penetration rate (2.03% of frames at 1 cm on ARCTIC) suggests the output motions can serve as initialization for physics-based simulators or motion retargeting pipelines without heavy cleanup.
Reading between the lines
- My inference: the contact-map-first idea should transfer beyond hands to whole-body interaction with articulated furniture or tools, where a sparse intent map on the object surface could mediate between environment geometry and full-body motion.
- My inference: the scale-normalization heuristic (a per-object guess of the articulation angle that maximizes extent) is the most likely point of failure for genuinely new objects; replacing it with a data-driven canonicalization would probably improve zero-shot generalization more than any other component change.
- My inference: the paper's multi-modality metric actually shows lower raw diversity than one adapted baseline (CAMS-B) on ARCTIC; the meaningful claim is diversity combined with physical plausibility, since that baseline's higher diversity comes with far more penetration.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents BimArt, a three-stage generative pipeline for synthesizing bimanual hand motions interacting with articulated objects, given only the 7-DoF trajectory of the object. It introduces a canonicalized, part-based BPS object representation, a diffusion-based contact map generator, and a diffusion-based motion generator that uses contact maps as conditioning and as guidance, followed by an optimization-based refinement. The method is evaluated on ARCTIC and HOI4D, and the authors claim state-of-the-art motion quality and diversity, with an ablation study and a perceptual user study.
Significance. If the method's claims are substantiated, the contribution is meaningful: it removes the need for a reference grasp or coarse hand trajectory, supports simultaneous articulation and global motion, and uses a single cross-category model. The paper's strengths include a clear description of the representation and losses, a thorough ablation over representations and contact usage, and a sensitivity analysis in the supplement. However, the headline 'surpasses the state of the art in motion quality and diversity' is not fully supported by the reported tables, as discussed below. The method does show very low penetration on ARCTIC compared to adapted baselines, which is a valuable result.
major comments (3)
- [4.1, 4.2, Tables 1 and 2] The central claim that BimArt 'surpasses the state of the art in motion quality and diversity' is not supported by the reported evidence. On ARCTIC (Table 1), CAMS-B has higher multimodality (8.5602 vs 6.9093 cm) and lower acceleration (0.11959 vs 0.18846 cm/s^2), which are the paper's own diversity and smoothness metrics. On HOI4D scissors (Table 2), BimArt with and without optimization has higher penetration than CAMS (0.591% and 1.204% vs 0.080%). The perceptual user study (Sec. 4.2) excludes CAMS-B, the baseline that is strongest on these axes, so no head-to-head perceptual justification exists for weighting penetration/contact above diversity/smoothness. In addition, all quantitative results are single-seed point estimates without confidence intervals, so it is unclear whether any of the differences are statistically significant. Please either temper the claim to 'competitive or better on contact/articulation metrics with substantially lower penetration on ARCTIC' or add a comparison that includes CAMS-B in the user study and report variance across seeds.
- [3.1 (Eq. 1) and Supplementary B] The 'category-agnostic' and 'unified' representation claim is weakened by the per-object-type heuristic used for the articulation angle in scale normalization: the supplement states that the angle is set to pi/2 for mixer and capsule machine, 0 for scissors and espresso machine, and pi for all others. This means that for a new object category, the user must supply a canonicalization rule, which is not category-agnostic. Since every downstream quantity (BPS features, contact maps, hand guidance) depends on this scale, the method's generalization to unseen articulated objects is at risk. Please provide a sensitivity analysis showing how the choice of this angle affects the final metrics, or revise the claim to specify that the canonicalization rule is object-model-specific.
- [4.2, Fig. 7] The user study only compares BimArt against MDM-B and OMOMO-B, while the caption of Fig. 7 states that 'Our method outperforms the existing state of the art for all objects.' Because CAMS-B is excluded and MDM-B/OMOMO-B are adapted baselines, this statement overstates the scope of the perceptual evidence. Please either include CAMS-B in the user study or explicitly limit the claim to the compared baselines.
minor comments (5)
- [Section 4, baseline details] The sentence 'The single-hand variant for HOI4D dataset is denoted as MDM-U' appears to be a typo; Table 2 lists both MDM-U and OMOMO-U, so the text should refer to 'OMOMO-U' for the OMOMO variant.
- [Eq. (7)] The notation M(\hat{X}(t), t, Z_o∅) is confusing; please define the null contact token and clarify which prediction is conditional and which is unconditional in the classifier-free guidance equations.
- [Table 3 and Sec. 4.4] The CM metric is computed against the contact maps predicted by the contact model, which is also used to guide the motion model; please state explicitly in the table caption that this is an internal consistency metric and not comparable across methods that do not share the same contact predictor.
- [Section 3.1, Eq. (2)] The part index p in O^p_i is not defined before use; please state explicitly that p ∈ {top, bottom}.
- [Section 2, Related Work] The concurrent work ManiDext [80] is mentioned as an exception but never compared quantitatively; please add a sentence explaining why it is excluded (e.g., different input assumptions or different evaluation protocol).
Circularity Check
No load-bearing circularity: one self-referential ablation metric (CM) is forced by the guidance objective; the central claim rests on ground-truth-based metrics.
-
fitted input called prediction
[Sec. 3.3 Eq. (6); Sec. 4 'Evaluation Metrics' (CM definition); Tab. 3 'Contact' ablations; Sec. 4.4]
"CM measures the l1 distance of the contact map derived from the generated hand motions from the predicted contact map. This metric is only applicable to our ablations. ... Contact guidance helps the hand motions better align with the contact maps, evidenced by a lower contact map discrepancy."
Eq. (6) implements contact-map guidance as gradient descent on ||Ĉ − C̃||, the l1 distance between the contact model's predicted map Ĉ and the map C̃ derived from the denoised hands. The CM ablation column is defined as exactly this same l1 distance, with the same reference Ĉ. The reported improvement (CM 1.1505 → 1.1284 when adding CG) is therefore the objective function of the guidance being minimized at inference, so the finding that guidance improves alignment is forced by construction rather than being empirical validation. CM also measures agreement with the model's own prediction (the same Ĉ that is fed as conditioning), not with ground-truth contact; the paper restricts the metric to ablations, so the headline ground-truth metrics (Pen, Con, Art, Mul) are unaffected.
full rationale
The central derivation is self-contained. Contact maps are ground-truth-derived (Eq. 4, from dataset hand meshes) and train a diffusion contact model; the motion model is trained with classifier-free dropout over these maps (Sec. 3.3); and the headline metrics Pen/Con/Art/Mul are computed against ground-truth object/hand geometry, not against the model's own outputs. The single self-referential element is the CM ablation metric, defined as the l1 distance from the generated hands' derived contact map to the contact model's predicted map — precisely the quantity that Eq. (6) minimizes at inference, with the same reference. Consequently the Sec. 4.4 claim that contact guidance helps hands align with the contact maps, evidenced by lower CM, holds by construction; because CM is labeled 'only applicable to our ablations,' it does not contaminate the baseline comparisons. No load-bearing self-citation exists: the self-citations (ROAM [81], MoFusion [11], IMOS [16], REMOS [17], ConvoFusion [47], MACS [54], HMP [13]) appear only in contextual related-work enumerations, dataset conventions, and baseline derivations; no uniqueness theorem or ansatz is imported via citation. The Supp. B scale-normalization heuristic (Eq. 1) is a robustness assumption about object canonicalization, not a circular step, and the appended Limitations passage concerns category coverage, not circularity. The legitimate reviewer concerns that are not circularity: the user study excludes CAMS-B (Sec. 4.2), all metrics are single-seed point estimates without confidence intervals, ARCTIC baselines do not receive the same post-optimization as BimArt, and the abstract's 'surpasses the state of the art in motion quality and diversity' is undercut by CAMS-B's better Mul and Accel in Tab. 1. These are evidence-strength and claim-scope issues, which fall under correctness risk, not circularity. Score 2: one minor self-referential ablation metric with an otherwise independent central claim.
Assumptions & free parameters
free parameters (6)
- Scale normalization margin d_margin =
0.15
- Canonical articulation angle for scale estimation =
pi/2 (mixer, capsule), 0 (scissors, espresso), pi (others)
- Classifier-free guidance scale lambda_f =
0.5
- Contact guidance scale lambda_c =
1/||grad X(t)|| (adaptive)
- Post-processing weights wproj, wpen, wacc =
100, 10, 1000 (ARCTIC); wacc=10^4 (HOI4D)
- Number of basis points K and hand keypoints J =
not specified in paper
assumptions (5)
- domain assumption Distance-based contact maps on the BPS-projected object vertices are a sufficient intermediate representation to guide bimanual hand motion synthesis.
- domain assumption The 7-DoF object trajectory (6D pose + 1D articulation) determines a learnable distribution of plausible bimanual interactions.
- ad hoc to paper Normalizing object scale by a heuristic articulation angle preserves sufficient geometric fidelity for the part-based BPS encoding.
- domain assumption Ground-truth hand-object contact in ARCTIC and HOI4D is accurate enough to supervise contact-map learning.
- domain assumption A fixed set of basis points uniformly sampled from the unit ball can represent a wide range of object geometries at comparable fidelity.
Cite this review
Pith. "Pith review of BimArt: A Unified Approach for the Synthesis of 3D Bimanual Interaction with Articulated Objects." pith.science (2026). https://pith.science/paper/LHEQDTBI
@misc{pith2026241205066,
author = {Pith},
title = {Pith review of: BimArt: A Unified Approach for the Synthesis of 3D Bimanual Interaction with Articulated Objects},
year = {2026},
howpublished = {\url{https://pith.science/paper/LHEQDTBI}},
note = {Machine review of arXiv:2412.05066}
}
read the original abstract
We present BimArt, a novel generative approach for synthesizing 3D bimanual hand interactions with articulated objects. Unlike prior works, we do not rely on a reference grasp, a coarse hand trajectory, or separate modes for grasping and articulating. To achieve this, we first generate distance-based contact maps conditioned on the object trajectory with an articulation-aware feature representation, revealing rich bimanual patterns for manipulation. The learned contact prior is then used to guide our hand motion generator, producing diverse and realistic bimanual motions for object movement and articulation. Our work offers key insights into feature representation and contact prior for articulated objects, demonstrating their effectiveness in taming the complex, high-dimensional space of bimanual hand-object interactions. Through comprehensive quantitative experiments, we demonstrate a clear step towards simplified and high-quality hand-object animations that surpass the state of the art in motion quality and diversity. Project page: https://vcai.mpi-inf.mpg.de/projects/bimart/.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 1 Pith paper
-
SyncDiff: Synchronized Motion Diffusion for Multi-Body Human-Object Interaction Synthesis
SyncDiff synthesizes multi-body human-object interaction motions with one diffusion model plus explicit synchronization and frequency decomposition, improving contact and action-quality metrics over prior methods on f...
Reference graph
Works this paper leans on
-
[1]
Bilinear spatiotemporal basis models
Ijaz Akhter, Tomas Simon, Sohaib Khan, Iain Matthews, and Yaser Sheikh. Bilinear spatiotemporal basis models. ACM Transactions on Graphics (TOG), 2012. 2
2012
-
[2]
Kemp, and James Hays
Samarth Brahmbhatt, Cusuh Ham, Charles C. Kemp, and James Hays. ContactDB: Analyzing and predicting grasp contact via thermal imaging. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019. 2
2019
-
[3]
Style machines
Matthew Brand and Aaron Hertzmann. Style machines. In Proceedings of the 27th Annual Conference on Computer Graphics and Interactive Techniques, 2000. 2
2000
-
[4]
Physically plausible full-body hand-object interaction synthesis
Jona Braun, Sammy Christen, Muhammed Kocabas, Emre Aksan, and Otmar Hilliges. Physically plausible full-body hand-object interaction synthesis. In Proceedings of the In- ternational Conference on 3D Vision (3DV), 2024. 2
2024
-
[5]
Text2hoi: Text-guided 3d motion generation for hand- object interaction
Junuk Cha, Jihyeon Kim, Jae Shin Yoon, and Seungryul Baek. Text2hoi: Text-guided 3d motion generation for hand- object interaction. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 2
2024
-
[6]
Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, Jan Kautz, and Di- eter Fox
Yu-Wei Chao, Wei Yang, Yu Xiang, Pavlo Molchanov, Ankur Handa, Jonathan Tremblay, Yashraj S. Narang, Karl Van Wyk, Umar Iqbal, Stan Birchfield, Jan Kautz, and Di- eter Fox. Dexycb: A benchmark for capturing hand grasping of objects. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021. 2
2021
-
[7]
Executing your commands via motion diffusion in latent space
Xin Chen, Biao Jiang, Wen Liu, Zilong Huang, Bin Fu, Tao Chen, and Gang Yu. Executing your commands via motion diffusion in latent space. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 8
work page 2023
-
[8]
Diffu- sion policy: Visuomotor policy learning via action diffusion
Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin Burchfiel, and Shuran Song. Diffu- sion policy: Visuomotor policy learning via action diffusion. In Proceedings of Robotics: Science and Systems, 2023. 2
work page 2023
Show all 86 references
-
[9]
D-Grasp: Phys- ically Plausible Dynamic Grasp Synthesis for Hand-Object Interactions
Sammy Christen, Muhammed Kocabas, Emre Aksan, Jemin Hwangbo, Jie Song, and Otmar Hilliges. D-Grasp: Phys- ically Plausible Dynamic Grasp Synthesis for Hand-Object Interactions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 1, 13
2022
-
[10]
Diffh2o: Diffusion-based synthesis of hand-object interactions from textual descriptions
Sammy Christen, Shreyas Hampali, Fadime Sener, Edoardo Remelli, Tomas Hodan, Eric Sauser, Shugao Ma, and Bugra Tekin. Diffh2o: Diffusion-based synthesis of hand-object interactions from textual descriptions. In SIGGRAPH Asia 2024 Conference Papers, 2024. 2
2024
-
[11]
Mofusion: A framework for denoising-diffusion-based motion synthesis
Rishabh Dabral, Muhammad Hamza Mughal, Vladislav Golyanik, and Christian Theobalt. Mofusion: A framework for denoising-diffusion-based motion synthesis. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2
2023
-
[12]
Adam: A method for stochastic opti- mization
P Kingma Diederik. Adam: A method for stochastic opti- mization. Int. Conf. Learn. Represent., 2014. 6
2014
-
[13]
Enes Duran, Muhammed Kocabas, Vasileios Choutas, Zi- cong Fan, and Michael J. Black. Hmp: Hand motion priors for pose and shape estimation from video. IEEE/CVF Win- ter Conference on Applications of Computer Vision (WACV),
-
[14]
Black, and Ot- mar Hilliges
Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J. Black, and Ot- mar Hilliges. ARCTIC: A dataset for dexterous bimanual hand-object manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)...
2023
-
[15]
Recurrent network models for human dynam- ics
Katerina Fragkiadaki, Sergey Levine, Panna Felsen, and Ji- tendra Malik. Recurrent network models for human dynam- ics. In Proceedings of the IEEE/CVF International Confer- ence on Computer Vision (ICCV), 2015. 2
2015
-
[16]
Imos: Intent-driven full-body motion synthesis for human-object interactions
Anindita Ghosh, Rishabh Dabral, Vladislav Golyanik, Chris- tian Theobalt, and Philipp Slusallek. Imos: Intent-driven full-body motion synthesis for human-object interactions. In Computer Graphics Forum (Eurographics), 2023. 1, 2, 13
2023
-
[17]
Remos: 3d motion- conditioned reaction synthesis for two-person interactions
Anindita Ghosh, Rishabh Dabral, Vladislav Golyanik, Chris- tian Theobalt, and Philipp Slusallek. Remos: 3d motion- conditioned reaction synthesis for two-person interactions. In Proceedings of the European Conference on Computer Vi- sion (ECCV), 2024. 2
2024
-
[18]
Honnotate: A method for 3d annotation of hand and object poses
Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vin- cent Lepetit. Honnotate: A method for 3d annotation of hand and object poses. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,
-
[19]
Hand-centric motion refinement for 3d hand-object in- teraction via hierarchical spatial-temporal modeling
Yuze Hao, Jianrong Zhang, Tao Zhuo, Fuan Wen, and Hehe Fan. Hand-centric motion refinement for 3d hand-object in- teraction via hierarchical spatial-temporal modeling. Asso- ciation for the Advancement of Artificial Intelligence , 2024. 2
2024
-
[20]
Stochas- tic scene-aware motion prediction
Mohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito, Jimei Yang, Yi Zhou, and Michael J Black. Stochas- tic scene-aware motion prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021. 1
2021
-
[21]
Black, Ivan Laptev, and Cordelia Schmid
Yana Hasson, G ¨ul Varol, Dimitrios Tzionas, Igor Kale- vatykh, Michael J. Black, Ivan Laptev, and Cordelia Schmid. Learning joint reconstruction of hands and manipulated ob- jects. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), 2019. 5
2019
-
[22]
Momentum contrast for unsupervised visual repre- sentation learning
Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual repre- sentation learning. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,
-
[23]
Classifier-free diffusion guidance
Jonathan Ho and Tim Salimans. Classifier-free diffusion guidance. In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications, 2021. 4, 6
2021
-
[24]
Denoising dif- fusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. Adv. Neural Inform. Process. Syst., 2020. 6
2020
-
[25]
Phase- functioned neural networks for character control
Daniel Holden, Taku Komura, and Jun Saito. Phase- functioned neural networks for character control. ACM Transactions on Graphics (TOG), 2017. 2
2017
-
[26]
3d-llm: Inject- ing the 3d world into large language models
Yining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng, Yilun Du, Zhenfang Chen, and Chuang Gan. 3d-llm: Inject- ing the 3d world into large language models. Adv. Neural Inform. Process. Syst., 2023. 8
2023
-
[27]
Dy- namic handover: Throw and catch with bimanual hands
Binghao Huang, Yuanpei Chen, Tianyu Wang, Yuzhe Qin, Yaodong Yang, Nikolay Atanasov, and Xiaolong Wang. Dy- namic handover: Throw and catch with bimanual hands. Conference on Robot Learning, 2023. 2
2023
-
[28]
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3. 6m: Large scale datasets and pre- dictive methods for 3d human sensing in natural environ- ments. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2013. 2
2013
-
[29]
Hand-object contact consistency reasoning for human grasps generation
Hanwen Jiang, Shaowei Liu, Jiashun Wang, and Xiaolong Wang. Hand-object contact consistency reasoning for human grasps generation. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision (ICCV), 2021. 2, 6
2021
-
[30]
Grasp- ing field: Learning implicit representations for human grasps
Korrawe Karunratanakul, Jinlong Yang, Yan Zhang, Michael J Black, Krikamol Muandet, and Siyu Tang. Grasp- ing field: Learning implicit representations for human grasps. In Proceedings of the International Conference on 3D Vision (3DV), 2020. 2
2020
-
[31]
Opti- mizing diffusion noise can serve as universal motion priors
Korrawe Karunratanakul, Konpat Preechakul, Emre Aksan, Thabo Beeler, Supasorn Suwajanakorn, and Siyu Tang. Opti- mizing diffusion noise can serve as universal motion priors. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 2
2024
-
[32]
Interhandgen: Two-hand interaction generation via cascaded reverse diffusion
Jihyun Lee, Shunsuke Saito, Giljoo Nam, Minhyuk Sung, and Tae-Kyun Kim. Interhandgen: Two-hand interaction generation via cascaded reverse diffusion. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 2
2024
-
[33]
Dextouch: Learning to seek and manipulate objects with tactile dexterity
Kang-Won Lee, Yuzhe Qin, Xiaolong Wang, and Soo-Chul Lim. Dextouch: Learning to seek and manipulate objects with tactile dexterity. IEEE Robotics and Automation Let- ters, 2024. 2
2024
-
[34]
Efficient nonlinear markov models for human mo- tion
Andreas M Lehrmann, Peter V Gehler, and Sebastian Nowozin. Efficient nonlinear markov models for human mo- tion. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), 2014. 2
2014
-
[35]
Object motion guided human motion synthesis
Jiaman Li, Jiajun Wu, and C Karen Liu. Object motion guided human motion synthesis. ACM Transactions on Graphics (TOG), 2023. 1, 2, 6, 13
2023
-
[36]
Controllable human-object interaction synthesis
Jiaman Li, Alexander Clegg, Roozbeh Mottaghi, Jiajun Wu, Xavier Puig, and C Karen Liu. Controllable human-object interaction synthesis. Proceedings of the European Confer- ence on Computer Vision (ECCV), 2024. 2
2024
-
[37]
Genzi: Zero-shot 3d human-scene interaction generation
Lei Li and Angela Dai. Genzi: Zero-shot 3d human-scene interaction generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 8
2024
-
[38]
Contactgen: Generative contact model- ing for grasp generation
Shaowei Liu, Yang Zhou, Jimei Yang, Saurabh Gupta, and Shenlong Wang. Contactgen: Generative contact model- ing for grasp generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2
2023
-
[39]
Geneoh diffusion: Towards general- izable hand-object interaction denoising via denoising diffu- sion
Xueyi Liu and Li Yi. Geneoh diffusion: Towards general- izable hand-object interaction denoising via denoising diffu- sion. In Int. Conf. Learn. Represent., 2024. 1, 2
2024
-
[40]
Hoi4d: A 4d egocentric dataset for category-level human- object interaction
Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. Hoi4d: A 4d egocentric dataset for category-level human- object interaction. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,
-
[41]
Taco: Benchmarking general- izable bimanual tool-action-object understanding
Yun Liu, Haolin Yang, Xu Si, Ling Liu, Zipeng Li, Yuxiang Zhang, Yebin Liu, and Li Yi. Taco: Benchmarking general- izable bimanual tool-action-object understanding. Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 2
2024
-
[42]
Sgdr: Stochastic gradi- ent descent with warm restarts
Ilya Loshchilov and Frank Hutter. Sgdr: Stochastic gradi- ent descent with warm restarts. Int. Conf. Learn. Represent.,
-
[43]
Troje, Ger- ard Pons-Moll, and Michael J
Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Ger- ard Pons-Moll, and Michael J. Black. Amass: Archive of motion capture as surface shapes. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019. 2
2019
-
[44]
Black, and Javier Romero
Julieta Martinez, Michael J. Black, and Javier Romero. On human motion prediction using recurrent neural networks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 2
2017
-
[45]
Promptable game models: Text-guided game simulation via masked diffusion models
Willi Menapace, Aliaksandr Siarohin, St ´ephane Lathuili`ere, Panos Achlioptas, Vladislav Golyanik, Sergey Tulyakov, and Elisa Ricci. Promptable game models: Text-guided game simulation via masked diffusion models. ACM Transactions on Graphics (TOG), 2024. 2
2024
-
[46]
Interhand2
Gyeongsik Moon, Shoou-I Yu, He Wen, Takaaki Shiratori, and Kyoung Mu Lee. Interhand2. 6m: A dataset and base- line for 3d interacting hand pose estimation from a single rgb image. In Proceedings of the European Conference on Computer Vision (ECCV), 2020. 2
2020
-
[47]
Convofusion: Multi-modal conversational dif- fusion for co-speech gesture synthesis
Muhammad Hamza Mughal, Rishabh Dabral, Ikhsanul Habibie, Lucia Donatelli, Marc Habermann, and Christian Theobalt. Convofusion: Multi-modal conversational dif- fusion for co-speech gesture synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog...
2024
-
[48]
Hoi-diff: Text-driven synthe- sis of 3d human-object interactions using diffusion models
Xiaogang Peng, Yiming Xie, Zizhao Wu, Varun Jampani, Deqing Sun, and Huaizu Jiang. Hoi-diff: Text-driven synthe- sis of 3d human-object interactions using diffusion models. arXiv preprint arXiv:2312.06553, 2023. 2
2023 arXiv
-
[49]
Action- conditioned 3d human motion synthesis with transformer vae
Mathis Petrovich, Michael J Black, and G ¨ul Varol. Action- conditioned 3d human motion synthesis with transformer vae. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision (ICCV), 2021. 2 10
2021
-
[50]
Black, G¨ul Varol, Xue Bin Peng, and Davis Rempe
Mathis Petrovich, Or Litany, Umar Iqbal, Michael J. Black, G¨ul Varol, Xue Bin Peng, and Davis Rempe. STMC: Multi- track timeline control for text-driven 3d human motion gen- eration. CVPR Workshop on Human Motion Generation ,
-
[51]
Ef- ficient learning on point clouds with basis point sets
Sergey Prokudin, Christoph Lassner, and Javier Romero. Ef- ficient learning on point clouds with basis point sets. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019. 2, 3
2019
-
[52]
Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: Modeling and capturing hands and bod- ies together. ACM Transactions on Graphics (TOG) , 2017. 3
2017
-
[53]
The vector heat method
Nicholas Sharp, Yousuf Soliman, and Keenan Crane. The vector heat method. ACM Transactions on Graphics (TOG),
-
[54]
Macs: Mass conditioned 3d hand and object motion synthe- sis
Soshi Shimada, Franziska Mueller, Jan Bednarik, Bardia Doosti, Bernd Bickel, Danhang Tang, Vladislav Golyanik, Jonathan Taylor, Christian Theobalt, and Thabo Beeler. Macs: Mass conditioned 3d hand and object motion synthe- sis. In Proceedings of the International Conference on...
2024
-
[55]
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International Confer- ence on Machine Learning, 2015. 4
2015
-
[56]
Denois- ing diffusion implicit models
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denois- ing diffusion implicit models. In Int. Conf. Learn. Repre- sent., 2021. 8
2021
-
[57]
Neural state machine for character-scene interactions
Sebastian Starke, He Zhang, Taku Komura, and Jun Saito. Neural state machine for character-scene interactions. ACM Transactions on Graphics (TOG), 2019. 1
2019
-
[58]
Deepphase: periodic autoencoders for learning motion phase manifolds
Sebastian Starke, Ian Mason, and Taku Komura. Deepphase: periodic autoencoders for learning motion phase manifolds. ACM Transactions on Graphics (TOG), 2022. 2
2022
-
[59]
Black, and Dim- itrios Tzionas
Omid Taheri, Nima Ghorbani, Michael J. Black, and Dim- itrios Tzionas. GRAB: A dataset of whole-body human grasping of objects. In Proceedings of the European Con- ference on Computer Vision (ECCV), 2020. 2
2020
-
[60]
Black, and Dim- itrios Tzionas
Omid Taheri, Vasileios Choutas, Michael J. Black, and Dim- itrios Tzionas. GOAL: Generating 4D whole-body motion for hand-object grasping. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2, 13
2022
-
[61]
Flex: Full- body grasping without full-body grasps
Purva Tendulkar, D´ıdac Sur´ıs, and Carl V ondrick. Flex: Full- body grasping without full-body grasps. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2
2023
-
[62]
Human motion diffu- sion model
Guy Tevet, Sigal Raab, Brian Gordon, Yoni Shafir, Daniel Cohen-or, and Amit Haim Bermano. Human motion diffu- sion model. In Int. Conf. Learn. Represent., 2023. 2, 4, 6
2023
-
[63]
Grasp’d: Differentiable contact-rich grasp syn- thesis for multi-fingered hands
Dylan Turpin, Liquan Wang, Eric Heiden, Yun-Chun Chen, Miles Macklin, Stavros Tsogkas, Sven Dickinson, and Ani- mesh Garg. Grasp’d: Differentiable contact-rich grasp syn- thesis for multi-fingered hands. In Proceedings of the Euro- pean Conference on Computer Vision (ECCV), 2022. 2
2022
-
[64]
Fast-grasp’d: Dexterous multi- finger grasp generation through differentiable simulation
Dylan Turpin, Tao Zhong, Shutong Zhang, Guanglei Zhu, Eric Heiden, Miles Macklin, Stavros Tsogkas, Sven Dick- inson, and Animesh Garg. Fast-grasp’d: Dexterous multi- finger grasp generation through differentiable simulation. In International Conference on Robotics and Automation ,
-
[65]
Attention is all you need
A Vaswani. Attention is all you need. Adv. Neural Inform. Process. Syst., 2017. 4
2017
-
[66]
Unidexgrasp++: Im- proving dexterous grasping policy learning via geometry- aware curriculum and iterative generalist-specialist learning
Weikang Wan, Haoran Geng, Yun Liu, Zikang Shan, Yaodong Yang, Li Yi, and He Wang. Unidexgrasp++: Im- proving dexterous grasping policy learning via geometry- aware curriculum and iterative generalist-specialist learning. Proceedings of the IEEE/CVF International Conference on ...
2023
-
[67]
Cy- berDemo: Augmenting Simulated Human Demonstration for Real-World Dexterous Manipulation
Jun Wang, Yuzhe Qin, Kaiming Kuang, Yigit Korkmaz, Akhilan Gurumoorthy, Hao Su, and Xiaolong Wang. Cy- berDemo: Augmenting Simulated Human Demonstration for Real-World Dexterous Manipulation. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CV...
2024
-
[68]
Wang, David J
Jack M. Wang, David J. Fleet, and Aaron Hertzmann. Gaus- sian process dynamical models for human motion. IEEE Transactions on Pattern Analysis and Machine Intelligence,
-
[69]
Dexgraspnet: A large- scale robotic dexterous grasp dataset for general objects based on simulation
Ruicheng Wang, Jialiang Zhang, Jiayi Chen, Yinzhen Xu, Puhao Li, Tengyu Liu, and He Wang. Dexgraspnet: A large- scale robotic dexterous grasp dataset for general objects based on simulation. International Conference on Robotics and Automation, 2023. 2
2023
-
[70]
Interdiff: Generating 3d human-object interactions with physics-informed diffusion
Sirui Xu, Zhengyuan Li, Yu-Xiong Wang, and Liang-Yan Gui. Interdiff: Generating 3d human-object interactions with physics-informed diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023. 1, 2
2023
-
[71]
What’s in your hands? 3d reconstruction of generic objects in hands
Yufei Ye, Abhinav Gupta, and Shubham Tulsiani. What’s in your hands? 3d reconstruction of generic objects in hands. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022. 2
2022
-
[72]
Diffusion-guided reconstruction of everyday hand- object interaction clips
Yufei Ye, Poorvi Hebbar, Abhinav Gupta, and Shubham Tul- siani. Diffusion-guided reconstruction of everyday hand- object interaction clips. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023. 2
2023
-
[73]
G-hop: Generative hand-object prior for interac- tion reconstruction and grasp synthesis
Yufei Ye, Abhinav Gupta, Kris Kitani, and Shubham Tul- siani. G-hop: Generative hand-object prior for interac- tion reconstruction and grasp synthesis. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 2
2024
-
[74]
Black, Xue Bin Peng, and Davis Rempe
Hongwei Yi, Justus Thies, Michael J. Black, Xue Bin Peng, and Davis Rempe. Generating human interaction motions in scenes with text control. Proceedings of the European Conference on Computer Vision (ECCV), 2024. 2
2024
-
[75]
Robot synesthesia: In-hand manipu- lation with visuotactile sensing
Ying Yuan, Haichuan Che, Yuzhe Qin, Binghao Huang, Zhao-Heng Yin, Kang-Won Lee, Yi Wu, Soo-Chul Lim, and Xiaolong Wang. Robot synesthesia: In-hand manipu- lation with visuotactile sensing. International Conference on Robotics and Automation, 2024. 2 11
2024
-
[76]
Ddf-ho: Hand-held object reconstruction via conditional directed distance field
Chenyangguang Zhang, Yan Di, Ruida Zhang, Guangyao Zhai, Fabian Manhardt, Federico Tombari, and Xiangyang Ji. Ddf-ho: Hand-held object reconstruction via conditional directed distance field. Adv. Neural Inform. Process. Syst. ,
-
[77]
ManipNet: Neural manipulation synthesis with a Hand- Object spatial representation
He Zhang, Yuting Ye, Takaaki Shiratori, and Taku Komura. ManipNet: Neural manipulation synthesis with a Hand- Object spatial representation. ACM Transactions on Graph- ics (TOG), 2021. 2, 6, 13
2021
-
[78]
GraspXL: Generating grasping motions for di- verse objects at scale
Hui Zhang, Sammy Christen, Zicong Fan, Otmar Hilliges, and Jie Song. GraspXL: Generating grasping motions for di- verse objects at scale. In Proceedings of the European Con- ference on Computer Vision (ECCV), 2024. 1
2024
-
[79]
ArtiGrasp: Physically plausible synthesis of bi-manual dexterous grasp- ing and articulation
Hui Zhang, Sammy Christen, Zicong Fan, Luocheng Zheng, Jemin Hwangbo, Jie Song, and Otmar Hilliges. ArtiGrasp: Physically plausible synthesis of bi-manual dexterous grasp- ing and articulation. In Proceedings of the International Conference on 3D Vision (3DV), 2024. 1, 2, 13
2024
-
[80]
Manidext: Hand-object manipulation synthesis via continuous corre- spondence embeddings and residual-guided diffusion, 2024
Jiajun Zhang, Yuxiang Zhang, Liang An, Mengcheng Li, Hongwen Zhang, Zonghai Hu, and Yebin Liu. Manidext: Hand-object manipulation synthesis via continuous corre- spondence embeddings and residual-guided diffusion, 2024. 2, 6
2024
-
[81]
Roam: Robust and object-aware motion genera- tion using neural pose descriptors
Wanyue Zhang, Rishabh Dabral, Thomas Leimk ¨uhler, Vladislav Golyanik, Marc Habermann, and Christian Theobalt. Roam: Robust and object-aware motion genera- tion using neural pose descriptors. Proceedings of the Inter- national Conference on 3D Vision (3DV), 2024. 2
2024
-
[82]
Couch: Towards controllable human-chair interactions
Xiaohan Zhang, Bharat Lal Bhatnagar, Sebastian Starke, Vladimir Guzov, and Gerard Pons-Moll. Couch: Towards controllable human-chair interactions. In Proceedings of the European Conference on Computer Vision (ECCV), 2022. 1
2022
-
[83]
Cams: Canonicalized manipulation spaces for category-level functional hand-object manipulation synthe- sis
Juntian Zheng, Qingyuan Zheng, Lixing Fang, Yun Liu, and Li Yi. Cams: Canonicalized manipulation spaces for category-level functional hand-object manipulation synthe- sis. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), 2023. 1, 2, 6, 13
2023
-
[84]
Toch: Spatio-temporal object-to-hand correspondence for motion refinement
Keyang Zhou, Bharat Lal Bhatnagar, Jan Eric Lenssen, and Gerard Pons-Moll. Toch: Spatio-temporal object-to-hand correspondence for motion refinement. InProceedings of the European Conference on Computer Vision (ECCV), 2022. 2
2022
-
[85]
Gears: Local geometry-aware hand-object interaction synthesis
Keyang Zhou, Bharat Lal Bhatnagar, Jan Eric Lenssen, and Gerard Pons-Moll. Gears: Local geometry-aware hand-object interaction synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024. 2
2024
-
[86]
grab” scenario, where the object’s articulation remains unchanged, and an “articulate
Zehao Zhu, Jiashun Wang, Yuzhe Qin, Deqing Sun, Varun Jampani, and Xiaolong Wang. Contactart: Learning 3d in- teraction priors for category-level articulated object and hand poses estimation. Proceedings of the International Confer- ence on 3D Vision (3DV), 2024. 2 12 BimArt: ...
2024
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.