REVIEW 4 major objections 6 minor 110 references
ECHO: Ego-Centric modeling of Human-Object interactions
T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Three trackers recover full human and object motion together
desk verdict Useful new generative model for egocentric HOI with a genuinely novel tri-variate diffusion, but the 'solely from head and wrist tracking' claim only holds if you already know the object and have its canonical mesh, which needs fixing before the paper is adopted. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a tri-variate diffusion process: human motion, object trajectory, and contact sequence are noised and denoised jointly, each with its own timestep per frame, inside a diffusion transformer. This per-frame, per-modality timestamp scheme is what turns the model into a flexible conditioning machine—any known portion of any modality can be supplied at a low noise level, and the model can attend to it while predicting the rest. The second piece is a conveyor inference: the denoising step increases monotonically along the temporal window, so completed frames stream out the front while fully noised frames enter from the back, enabling arbitrary-length, temporally consistent inference. The third piece is a head-centric canonical frame, expressed relative to the head pose at the first frame with the vertical axis aligned to gravity, which removes global orientation as a nuisance and, per the ablation, sharply improves prediction quality.
What would settle it
Take an unseen manipulation sequence recorded with only head and wrist trackers, supply the object mesh, and run ECHO while a baseline simply keeps the object at its most likely class-conditioned pose; if ECHO does not beat that baseline by a clear margin on object position and rotation error, then the apparent success would be better attributed to the dataset prior than to the sparse tracking signal.
Extended reading notes
Core claim
On its own terms, the paper establishes that human pose, object motion, and contact can be generated jointly from head-and-wrist conditioning by diffusing the three modalities together inside one transformer. The model works in a head-centric canonical frame to remove global-orientation bias, assigns an independent denoising timestamp to every frame and modality, and denoises a sliding temporal window so sequences of arbitrary length can be produced in real time. The authors show that feeding in sparse observations of any one modality—a few frames of human tracking, a few frames of object tracking, or partial contact labels—improves the reconstruction of the other modalities, and that ablations remove most of the gains when the contact modality, the head-centric representation, or the auxiliary motion data is removed. The central quantitative claim is state-of-the-art performance on both human and object metrics against a motion-diffusion baseline extended to object modeling, on two interaction datasets.
Load-bearing premise
At test time the user must know the object's class and provide its 3D canonical mesh, and the body shape parameters are assumed known; without that geometry and identity there is no object representation to denoise, so the operation is not fully self-contained from tracking alone.
Editorial extensions
If this is right
- Wearable-only setups in AR and VR could animate full-body avatars and manipulated objects in real time; the paper reports about 13.7 ms of inference per frame on a consumer GPU.
- A few observed frames of human, object, or contact information can be folded into the reconstruction to constrain the other modalities, so intermittent tracking does not break the output.
- Because the three modalities are noised independently, training can mix large motion-only datasets with smaller interaction datasets, giving the model a strong human-motion prior without sacrificing interaction detail.
- Conveyor inference removes the sliding-window stitching problem and supports arbitrarily long sequences, which is what a continuous wearable system would need.
Reading between the lines
- As a practical consequence the paper does not develop, ECHO cannot handle an object whose canonical mesh and class label are unavailable at test time; an obvious extension is to estimate the template on the fly from an egocentric camera and to quantify how template error propagates into trajectory error.
- The same joint-distribution formulation could be inverted into an interaction simulator: condition on the object's trajectory and generate the human response, or condition on the human and generate plausible object behavior, which would be useful for robotics and content creation.
- The per-frame timestamp mechanism suggests a natural online-filtering reading: if observed streams are fed in at low noise levels, ECHO could act as a continuously correcting state estimator rather than a one-shot generator, though the paper does not evaluate this mode directly.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces ECHO, a diffusion-transformer framework that jointly reconstructs human body pose, object motion, and contact signals from sparse egocentric sensors (head and wrist tracking). The key technical proposals are a tri-variate diffusion process with independent per-frame, per-modality noise schedules; a head-centric canonical coordinate representation; a conveyor-based autoregressive inference scheme for arbitrarily long sequences; and training on a mixture of HOI datasets (BEHAVE, OMOMO) plus large-scale human motion data (AMASS). Experiments report quantitative results on BEHAVE and OMOMO against a self-constructed BoDiffusion+Obj baseline, motion-generation comparisons on AMASS, sparse-tracking evaluations, and ablations supporting the design choices. The central claims are that ECHO is the first method to recover human and object motion jointly from 3-point tracking and that it achieves state-of-the-art egocentric HOI reconstruction.
Significance. If the claims hold, ECHO would be a meaningful step toward practical egocentric HOI capture from commodity wearables. The tri-variate diffusion formulation with per-frame timestamps is an interesting generalization of prior diffusion models, and the head-centric representation with conveyor inference addresses real sequence-length and orientation issues. The paper's strengths include internally consistent ablations (e.g., the benefit of the contact modality, AMASS pretraining, and head-centric coordinates), a clear exposition of the training objective, and a stated commitment to release code and models. The main reservation is that the headline claim of operating 'solely from head and wrist tracking' is contradicted by the requirement that the object's canonical mesh and class label be provided as input; the state-of-the-art claim is also weakened by the absence of quantitative comparison with the closest existing egocentric HOI method (iReplica) and the prior trilateral diffusion method (TriDi).
major comments (4)
- [Abstract and Section 3.1] The abstract and contribution list claim that ECHO recovers human pose, object motion, and contact 'solely from head and wrist tracking' and is 'the first method to solve for human and object motion sequences jointly, relying only on 3-point tracking.' Section 3.1 states that 'For every object we assume that its canonical mesh is given to the model as an input,' and the object conditioning C_O = (y_O, f_O) is computed from a one-hot class label and PointNext features of the canonicalized object vertices. Without this mesh, the object modality has no representation to denoise, so object trajectory cannot be predicted at all. The method therefore solves a conditional generation problem (3-point tracking plus known object identity and shape), not the sensor-only problem advertised. This is a material limitation that should be disclosed in the abstract and contributions, or the claims should be reworded to state the actual input requirements.
- [Section 4.1, Tables 1-2] The only HOI baseline in the comparison is BoDiffusion+Obj, which is constructed by the authors on top of BoDiffusion. There is no quantitative comparison with iReplica, the only existing egocentric HOI method cited as 'the closest approach to ours,' nor with TriDi, the trilateral diffusion model on which the proposed formulation is directly built. The paper's claim of 'state-of-the-art, significantly outperforming existing methods' is therefore not supported for the egocentric HOI setting. The authors should either add experiments against these methods (where input requirements are compatible) or explicitly state and justify why a quantitative comparison is not possible, and temper the SOTA claim accordingly.
- [Section 4.1, Table 2] The text states that ECHO 'performs on par or better than BoDiffusion+Obj' on AMASS, but Table 2 shows ECHO's MPJPE (93.9±8.7) is worse than BoDiffusion+Obj (91.5±4.2), while MPJVE is better (109.7 vs 115.5). Given the large variances, 'on par' may be defensible, but the claim as written is imprecise and should be corrected to reflect the actual direction of the differences.
- [Section 4.3, Tables 4 and S3] The ablation of the inference-time guidance (NoGuide) shows negligible differences on BEHAVE (human MPJPE 61.4 vs 61.2, object Ev2v 29.5 vs 29.8) and on OMOMO (Table S3, human MPJPE 64.1 vs 64.5, object Ev2v 26.7 vs 26.5). The paper claims guidance is useful, but these differences are well within the reported variances. The guidance's contribution should be characterized more cautiously, or the experimental setup (e.g., which weights are used) should be clarified to show where the benefit actually appears.
minor comments (6)
- [Abstract] The phrase 'tri-variate diffusion process with independent noise schedules' is ambiguous: the independent schedules are per frame and per modality, not merely per modality. Consider rewording to make the granularity explicit.
- [Section 3.2, Eq. (7)] The shorthand Ep, Et, Eq introduced in the background is reused in Eq. (7), but the subscripts are not restated; adding a one-line reminder would improve readability.
- [Table 2] The FC column is labeled 'FC↑1.0' in the header, which is confusing. The metric should be described in the caption as a fraction where higher is better, and the '1.0' in the header should be removed.
- [Section 4.3] There is a typo 'in thew Sup.Mat.'; it should read 'in the Sup. Mat.'
- [Figure 5] The qualitative comparison shows 'BoDiffusion + Obj.' but not iReplica or TriDi. Adding qualitative results for at least one of those methods, or a sentence explaining their absence, would strengthen the comparison.
- [Section 3.1] The notation SMPL (T_H, θ_H) should refer to SMPL-X consistently, as the paragraph begins by naming SMPL-X as the body model.
Circularity Check
No circularity found: ECHO's diffusion objective and evaluations are self-contained; the object-mesh input assumption is an overclaim, not a circular reduction.
full rationale
ECHO is a learned generative system, not a symbolic derivation, and no step reduces by construction to its own inputs. The training objective in Eq. 7 is a standard conditional DDPM loss: the network is trained to recover clean (H0,O0,I0) from independently noised versions of the three modalities plus conditioning (E, C_O), and the output is not algebraically forced to equal any input. The object conditioning C_O=(y_O,f_O) is a test-time input disclosed in Sec. 3.1 ('For every object we assume that its canonical mesh is given to the model as an input'), so the advertised 'solely from head and wrist tracking' wording overstates the sensor requirements, but this is a scope/correctness issue rather than a circular reduction: the object pose O is still a generated SE(3) sequence and is not obtained by refitting C_O. Citations to TriDi and UniDiffuser are to published methods that ECHO explicitly extends, the tri-variate objective is written out in Eqs. 6-7, and no uniqueness theorem or ansatz is imported from self-citations. Contact labels are defined from H/O geometry, but they are used as an auxiliary training signal and as an inference-time consistency constraint, not as the ground-truth target that defines H/O; the ablation shows empirical value. Evaluation is performed against held-out BEHAVE/OMOMO/AMASS splits with standard metrics, so the central performance claims are externally falsifiable. The only flagged limitation (canonical mesh and class label required at test time) undercuts the abstract's 'solely' wording but does not make the derivation circular.
Assumptions & free parameters
free parameters (6)
- Model weights (47.3M parameters) =
trained on AMASS + BEHAVE + OMOMO
- Loss weighting coefficients =
lambda_Hn = lambda_On = lambda_In = 1, lambda_Ov = lambda_Hj = 0.1, lambda_Hs = 0.05
- Contact distance threshold tau_c =
not stated in main text or provided supplementary
- Temporal window size W =
60 frames
- Inference denoising steps =
100
- Contact point set P_c =
64 points uniformly sampled from SMPL-X vertices
assumptions (5)
- domain assumption The object's canonical mesh and one-hot class label are available at test time.
- domain assumption Ground-truth 3-point tracking of head and wrists is available for conditioning.
- domain assumption SMPL-X body model with known shape parameters beta is a faithful representation of the human body.
- standard math DDPM forward and reverse processes with a DiT denoiser provide a valid generative model.
- domain assumption Head-centric canonicalization with a gravity-aligned height axis removes global orientation bias.
Cite this review
Pith. "Pith review of ECHO: Ego-Centric modeling of Human-Object interactions." pith.science (2026). https://pith.science/paper/E2IVXB2W
@misc{pith2026250821556,
author = {Pith},
title = {Pith review of: ECHO: Ego-Centric modeling of Human-Object interactions},
year = {2026},
howpublished = {\url{https://pith.science/paper/E2IVXB2W}},
note = {Machine review of arXiv:2508.21556}
}
read the original abstract
Modeling human-object interactions (HOI) from an egocentric perspective is a critical yet challenging task, particularly when relying on sparse signals from wearable devices like smart glasses and watches. We present ECHO, the first unified framework to jointly recover human pose, object motion, and contact dynamics solely from head and wrist tracking. To tackle the underconstrained nature of this problem, we introduce a novel tri-variate diffusion process with independent noise schedules that models the mutual dependencies between the human, object, and interaction modalities. This formulation allows ECHO to operate with flexible input configurations, making it robust to intermittent tracking and capable of leveraging partial observations. Crucially, it enables training on a combination of large-scale human motion datasets and smaller HOI collections, learning strong priors while capturing interaction nuances. Furthermore, we employ a smooth inpainting inference mechanism that enables the generation of temporally consistent interactions for arbitrarily long sequences. Extensive evaluations demonstrate that ECHO achieves state-of-the-art performance, significantly outperforming existing methods lacking such flexibility. The project page is available at https://ptrvilya.github.io/echo/.
Figures
Figures from the paper (1 more)
Reference graph
Works this paper leans on
-
[1]
Un- realego: A new dataset for robust egocentric 3d human mo- tion capture
Hiroyasu Akada, Jian Wang, Soshi Shimada, Masaki Taka- hashi, Christian Theobalt, and Vladislav Golyanik. Un- realego: A new dataset for robust egocentric 3d human mo- tion capture. In European Conference on Computer Vision, pages 1–17. Springer, 2022. 2
2022
-
[2]
Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon, Shreyas Hampali, Fan Zhang, Jade Fountain, Edward Miller, Selen Basol, Richard Newcombe, Robert Wang, et al. Introducing hot3d: An egocentric dataset for 3d hand and object tracking.arXiv preprint arXiv:2406.09598,
-
[3]
One transformer fits all distributions in multi-modal diffu- sion at scale
Fan Bao, Shen Nie, Kaiwen Xue, Chongxuan Li, Shi Pu, Yaole Wang, Gang Yue, Yue Cao, Hang Su, and Jun Zhu. One transformer fits all distributions in multi-modal diffu- sion at scale. In Proceedings of the 40th International Con- ference on Machine Learning. JMLR.org, 2023. 4
2023
-
[4]
From sparse signal to smooth motion: Real-time motion generation with rolling prediction models
German Barquero, Nadine Bertsch, Manojkumar Marram- reddy, Carlos Chac ´on, Filippo Arcadu, Ferran Rigual, Nicky Sijia He, Cristina Palmero, Sergio Escalera, Yuting Ye, et al. From sparse signal to smooth motion: Real-time motion generation with rolling prediction models. In Pro- ceedings of the Computer Vision and Pattern Recognition Conference, pages 18...
2025
-
[5]
Bharat Lal Bhatnagar, Suriya Singh, Chetan Arora, and C.V . Jawahar. Unsupervised learning of deep feature repre- sentation for clustering egocentric actions. In Proceedings of the Twenty-Sixth International Joint Conference on Arti- ficial Intelligence, IJCAI-17, pages 1447–1453, 2017. 2
2017
-
[6]
Behave: Dataset and method for tracking human object in- teractions
Bharat Lal Bhatnagar, Xianghui Xie, Ilya A Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. Behave: Dataset and method for tracking human object in- teractions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 15935– 15946, 2022. 3, 6
2022
-
[7]
Physically plausible full-body hand-object interaction synthesis
Jona Braun, Sammy Christen, Muhammed Kocabas, Emre Aksan, and Otmar Hilliges. Physically plausible full-body hand-object interaction synthesis. In 2024 International Conference on 3D Vision (3DV) , pages 464–473. IEEE,
2024
-
[8]
Egocentric gesture recognition using recurrent 3d convolutional neural networks with spatiotemporal trans- former modules
Congqi Cao, Yifan Zhang, Yi Wu, Hanqing Lu, and Jian Cheng. Egocentric gesture recognition using recurrent 3d convolutional neural networks with spatiotemporal trans- former modules. 2017 IEEE International Conference on Computer Vision (ICCV), 2017. 2
2017
Show all 110 references
-
[9]
Avatargo: Zero-shot 4d human-object interaction generation and animation
Yukang Cao, Liang Pan, Kai Han, Kwan-Yee K Wong, and Ziwei Liu. Avatargo: Zero-shot 4d human-object interaction generation and animation. arXiv preprint arXiv:2410.07164, 2024. 3
2024 arXiv
-
[10]
Bodiffusion: Diffusing sparse observations for full-body human motion synthesis
Angela Castillo, Maria Escobar, Guillaume Jeanneret, Al- bert Pumarola, Pablo Arbel ´aez, Ali Thabet, and Artsiom Sanakoyeu. Bodiffusion: Diffusing sparse observations for full-body human motion synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vis...
2023
-
[11]
Diffusion forcing: Next-token prediction meets full-sequence diffu- sion
Boyuan Chen, Diego Mart ´ı Mons ´o, Yilun Du, Max Sim- chowitz, Russ Tedrake, and Vincent Sitzmann. Diffusion forcing: Next-token prediction meets full-sequence diffu- sion. Advances in Neural Information Processing Systems, 37:24081–24125, 2024. 5
2024
-
[12]
Muscles in action
Mia Chiquier and Carl V ondrick. Muscles in action. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 22091–22101, 2023. 2
2023
-
[13]
Hmd-poser: On-device real-time human motion tracking from scalable sparse observations
Peng Dai, Yang Zhang, Tao Liu, Zhen Fan, Tianyuan Du, Zhuo Su, Xiaozheng Zheng, and Zeming Li. Hmd-poser: On-device real-time human motion tracking from scalable sparse observations. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages ...
2024
-
[14]
In- terfusion: Text-driven generation of 3d human-object inter- action
Sisi Dai, Wenhao Li, Haowen Sun, Haibin Huang, Chongyang Ma, Hui Huang, Kai Xu, and Ruizhen Hu. In- terfusion: Text-driven generation of 3d human-object inter- action. In European Conference on Computer Vision, pages 18–35. Springer, 2024. 3
2024
-
[15]
Hsc4d: Human- centered 4d scene capture in large-scale indoor-outdoor space using wearable imus and lidar
Yudi Dai, Yitai Lin, Chenglu Wen, Siqi Shen, Lan Xu, Jingyi Yu, Yuexin Ma, and Cheng Wang. Hsc4d: Human- centered 4d scene capture in large-scale indoor-outdoor space using wearable imus and lidar. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogn...
2022
-
[16]
Diffusion models beat gans on image synthesis
Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in neural informa- tion processing systems, 34:8780–8794, 2021. 6
2021
-
[17]
Cg-hoi: Contact-guided 3d human-object interaction generation
Christian Diller and Angela Dai. Cg-hoi: Contact-guided 3d human-object interaction generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19888–19901, 2024. 3
2024
-
[18]
Avatars grow legs: Generating smooth human motion from sparse track- ing inputs with diffusion model
Yuming Du, Robin Kips, Albert Pumarola, Sebastian Starke, Ali Thabet, and Artsiom Sanakoyeu. Avatars grow legs: Generating smooth human motion from sparse track- ing inputs with diffusion model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogniti...
2023
-
[19]
9 Egocast: Forecasting egocentric human pose in the wild
Maria Escobar, Juanita Puentes, Cristhian Forigua, Jordi Pont-Tuset, Kevis-Kokitsi Maninis, and Pablo Arbelaez. 9 Egocast: Forecasting egocentric human pose in the wild. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 5831–5841. IEEE, 2025. 2
2025
-
[20]
Arctic: A dataset for dexterous bimanual hand- object manipulation
Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J Black, and Otmar Hilliges. Arctic: A dataset for dexterous bimanual hand- object manipulation. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages...
2023
-
[21]
Understand- ing egocentric activities
Alireza Fathi, Ali Farhadi, and James M Rehg. Understand- ing egocentric activities. In 2011 international conference on computer vision, pages 407–414. IEEE, 2011. 2
2011
-
[22]
The matrix: Infinite-horizon world generation with real-time moving control
Ruili Feng, Han Zhang, Zhantao Yang, Jie Xiao, Zhilei Shu, Zhiheng Liu, Andy Zheng, Yukun Huang, Yu Liu, and Hongyang Zhang. The matrix: Infinite-horizon world generation with real-time moving control. arXiv preprint arXiv:2412.03568, 2024. 5
2024 arXiv
-
[23]
Imos: Intent- driven full-body motion synthesis for human-object inter- actions
Anindita Ghosh, Rishabh Dabral, Vladislav Golyanik, Christian Theobalt, and Philipp Slusallek. Imos: Intent- driven full-body motion synthesis for human-object inter- actions. In Computer Graphics Forum, pages 1–12. Wiley Online Library, 2023. 3
2023
-
[24]
Human poseitioning system (hps): 3d human pose estimation and self-localization in large scenes from body-mounted sensors
Vladimir Guzov, Aymen Mir, Torsten Sattler, and Gerard Pons-Moll. Human poseitioning system (hps): 3d human pose estimation and self-localization in large scenes from body-mounted sensors. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2021. 2, 3
2021
-
[25]
Interaction replica: Tracking human–object interac- tion and scene changes from human motion
Vladimir Guzov, Julian Chibane, Riccardo Marin, Yannan He, Yunus Saracoglu, Torsten Sattler, and Gerard Pons- Moll. Interaction replica: Tracking human–object interac- tion and scene changes from human motion. In Interna- tional Conference on 3D Vision (3DV), 2024. 2, 3
2024
-
[26]
Blendify – python rendering framework for blender
Vladimir Guzov, Ilya A Petrov, and Gerard Pons-Moll. Blendify – python rendering framework for blender. arXiv preprint arXiv:2410.17858, 2024. 6
2024 arXiv
-
[27]
Karen Liu, Yuting Ye, and Lingni Ma
Vladimir Guzov, Yifeng Jiang, Fangzhou Hong, Gerard Pons-Moll, Richard Newcombe, C. Karen Liu, Yuting Ye, and Lingni Ma. Hmd2: Environment-aware motion genera- tion from single egocentric head-mounted device. In Inter- national Conference on 3D Vision (3DV), 2025. 2, 3, 5
2025
-
[28]
Resolving 3d human pose ambiguities with 3d scene constraints
Mohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, and Michael J Black. Resolving 3d human pose ambiguities with 3d scene constraints. In Proceedings of the IEEE/CVF international conference on computer vision , pages 2282– 2292, 2019. 3
2019
-
[29]
Populating 3d scenes by learning human-scene interaction
Mohamed Hassan, Partha Ghosh, Joachim Tesch, Dimitrios Tzionas, and Michael J Black. Populating 3d scenes by learning human-scene interaction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14708–14718, 2021. 3
2021
-
[30]
Denoising dif- fusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. Advances in neural informa- tion processing systems, 33:6840–6851, 2020. 4, 1
2020
-
[31]
Egosim: An egocentric multi-view simulator and real dataset for body-worn cameras during motion and activity
Dominik Hollidt, Paul Streli, Jiaxi Jiang, Yasaman Haghighi, Changlin Qian, Xintong Liu, and Christian Holz. Egosim: An egocentric multi-view simulator and real dataset for body-worn cameras during motion and activity. Advances in Neural Information Processing Systems , 37: 10...
2024
-
[32]
Microsoft HoloLens, accessed January 7, 2025
HoloLens. Microsoft HoloLens, accessed January 7, 2025. https://learn.microsoft.com/en-us/hololens/. 2
2025
-
[33]
Diffusion- based generation, optimization, and planning in 3d scenes
Siyuan Huang, Zan Wang, Puhao Li, Baoxiong Jia, Tengyu Liu, Yixin Zhu, Wei Liang, and Song-Chun Zhu. Diffusion- based generation, optimization, and planning in 3d scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16750–16761, 2023. 3
2023
-
[34]
Black, Otmar Hilliges, and Gerard Pons-Moll
Yinghao Huang, Manuel Kaufmann, Emre Aksan, Michael J. Black, Otmar Hilliges, and Gerard Pons-Moll. Deep inertial poser: Learning to reconstruct human pose from sparse inertial measurements in real time. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia), 37(6): 185:1–185:15, 2018. 2
2018
-
[35]
Black, and Dim- itrios Tzionas
Yinghao Huang, Omid Taheri, Michael J. Black, and Dim- itrios Tzionas. InterCap: Joint markerless 3D tracking of humans and objects in interaction from multi-view RGB-D images. International Journal of Computer Vision (IJCV),
-
[36]
Avatar- poser: Articulated full-body pose tracking from sparse mo- tion sensing
Jiaxi Jiang, Paul Streli, Huajian Qiu, Andreas Fender, Larissa Laich, Patrick Snape, and Christian Holz. Avatar- poser: Articulated full-body pose tracking from sparse mo- tion sensing. In European conference on computer vision, pages 443–460. Springer, 2022. 2
2022
-
[37]
Scaling up dynamic human-scene interaction mod- eling
Nan Jiang, Zhiyuan Zhang, Hongjie Li, Xiaoxuan Ma, Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, and Siyuan Huang. Scaling up dynamic human-scene interaction mod- eling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 1737– 1747, 2024. 3
2024
-
[38]
Neuralho- fusion: Neural volumetric rendering under human-object interactions
Yuheng Jiang, Suyi Jiang, Guoxing Sun, Zhuo Su, Kai- wen Guo, Minye Wu, Jingyi Yu, and Lan Xu. Neuralho- fusion: Neural volumetric rendering under human-object interactions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6155– 6165, 2022. 3
2022
-
[39]
Transformer in- ertial poser: Real-time human motion reconstruction from sparse imus with simultaneous terrain generation
Yifeng Jiang, Yuting Ye, Deepak Gopinath, Jungdam Won, Alexander W Winkler, and C Karen Liu. Transformer in- ertial poser: Real-time human motion reconstruction from sparse imus with simultaneous terrain generation. In SIG- GRAPH Asia 2022 Conference Papers, pages 1–9, 2022. 2
2022
-
[40]
Ego3dpose: Capturing 3d cues from binocular egocentric views
Taeho Kang, Kyungjin Lee, Jinrui Zhang, and Youngki Lee. Ego3dpose: Capturing 3d cues from binocular egocentric views. In SIGGRAPH Asia 2023 Conference Papers, pages 1–10, 2023. 2
2023
-
[41]
Em-pose: 3d human pose estimation from sparse electromagnetic trackers
Manuel Kaufmann, Yi Zhao, Chengcheng Tang, Lingling Tao, Christopher Twigg, Jie Song, Robert Wang, and Otmar Hilliges. Em-pose: 3d human pose estimation from sparse electromagnetic trackers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 1151...
2021
-
[42]
aitviewer, 2022
Manuel Kaufmann, Velko Vechev, and Dario Mylonopou- los. aitviewer, 2022. 6
2022
-
[43]
Nifty: Neural object interaction fields for guided human motion synthesis
Nilesh Kulkarni, Davis Rempe, Kyle Genova, Abhi- jit Kundu, Justin Johnson, David Fouhey, and Leonidas 10 Guibas. Nifty: Neural object interaction fields for guided human motion synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , p...
2024
-
[44]
Mocap everyone everywhere: Lightweight motion capture with smartwatches and a head- mounted camera
Jiye Lee and Hanbyul Joo. Mocap everyone everywhere: Lightweight motion capture with smartwatches and a head- mounted camera. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 1091–1100, 2024. 2, 3
2024
-
[45]
Rewind: Real-time egocentric whole- body motion diffusion with exemplar-based identity condi- tioning
Jihyun Lee, Weipeng Xu, Alexander Richard, Shih-En Wei, Shunsuke Saito, Shaojie Bai, Te-Li Wang, Minhyuk Sung, Jason Saragih, et al. Rewind: Real-time egocentric whole- body motion diffusion with exemplar-based identity condi- tioning. arXiv preprint arXiv:2504.04956, 2025. 2
2025 arXiv
-
[46]
Ego-body pose es- timation via ego-head pose estimation
Jiaman Li, Karen Liu, and Jiajun Wu. Ego-body pose es- timation via ego-head pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17142–17151, 2023. 2, 7
2023
-
[47]
Object motion guided human motion synthesis
Jiaman Li, Jiajun Wu, and C Karen Liu. Object motion guided human motion synthesis. ACM Transactions on Graphics (TOG), 42(6):1–11, 2023. 3, 6
2023
-
[48]
Controllable human-object interaction synthesis
Jiaman Li, Alexander Clegg, Roozbeh Mottaghi, Jiajun Wu, Xavier Puig, and C Karen Liu. Controllable human-object interaction synthesis. In European Conference on Com- puter Vision, pages 54–72. Springer, 2024. 3
2024
-
[49]
Task-oriented human-object interactions generation with implicit neural representations
Quanzhou Li, Jingbo Wang, Chen Change Loy, and Bo Dai. Task-oriented human-object interactions generation with implicit neural representations. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 3035–3044, 2024. 3
2024
-
[50]
Egohdm: An online egocentric-inertial human motion capture, lo- calization, and dense mapping system
Bonan Liu, Handi Yin, Manuel Kaufmann, Jinhao He, Sammy Christen, Jie Song, and Pan Hui. Egohdm: An online egocentric-inertial human motion capture, lo- calization, and dense mapping system. arXiv preprint arXiv:2409.00343, 2024. 2, 3
2024 arXiv
-
[51]
Egofish3d: Egocentric 3d pose es- timation from a fisheye camera via self-supervised learning
Yuxuan Liu, Jianxin Yang, Xiao Gu, Yijun Chen, Yao Guo, and Guang-Zhong Yang. Egofish3d: Egocentric 3d pose es- timation from a fisheye camera via self-supervised learning. IEEE Transactions on Multimedia, 25:8880–8891, 2023. 2
2023
-
[52]
Egohmr: Egocentric human mesh recovery via hierarchical latent diffusion model
Yuxuan Liu, Jianxin Yang, Xiao Gu, Yao Guo, and Guang- Zhong Yang. Egohmr: Egocentric human mesh recovery via hierarchical latent diffusion model. In 2023 IEEE Inter- national Conference on Robotics and Automation (ICRA) , pages 9807–9813. IEEE, 2023. 2
2023
-
[53]
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 6
2017 arXiv
-
[54]
Nymeria: A massive collection of multimodal egocentric daily motion in the wild
Lingni Ma, Yuting Ye, Fangzhou Hong, Vladimir Guzov, Yifeng Jiang, Rowan Postyeni, Luis Pesqueira, Alexander Gamino, Vijay Baiyya, Hyo Jin Kim, et al. Nymeria: A massive collection of multimodal egocentric daily motion in the wild. In European Conference on Computer Vision, pa...
2024
-
[55]
Going deeper into first-person activity recognition
Minghuang Ma, Haoqi Fan, and Kris M Kitani. Going deeper into first-person activity recognition. In Proceed- ings of the IEEE Conference on Computer Vision and Pat- tern Recognition, pages 1894–1903, 2016. 2
1903
-
[56]
Troje, Gerard Pons-Moll, and Michael J
Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. AMASS: Archive of motion capture as surface shapes. In International Con- ference on Computer Vision , pages 5442–5451, 2019. 5, 6
2019
-
[57]
Generating continual human motion in diverse 3d scenes
Aymen Mir, Xavier Puig, Angjoo Kanazawa, and Gerard Pons-Moll. Generating continual human motion in diverse 3d scenes. In 2024 International Conference on 3D Vision (3DV), pages 903–913. IEEE, 2024. 3
2024
-
[58]
Imuposer: Full-body pose estima- tion using imus in phones, watches, and earbuds
Vimal Mollyn, Riku Arakawa, Mayank Goel, Chris Harri- son, and Karan Ahuja. Imuposer: Full-body pose estima- tion using imus in phones, watches, and earbuds. In Pro- ceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–12, 2023. 2
2023
-
[59]
Joint reconstruction of 3d human and object via contact-based refinement transformer
Hyeongjin Nam, Daniel Sungho Jung, Gyeongsik Moon, and Kyoung Mu Lee. Joint reconstruction of 3d human and object via contact-based refinement transformer. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10218–10227, 2024. 3
2024
-
[60]
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3D hands, face, and body from a single image. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pa...
2019
-
[61]
Scalable diffusion mod- els with transformers
William Peebles and Saining Xie. Scalable diffusion mod- els with transformers. In Proceedings of the IEEE/CVF international conference on computer vision , pages 4195– 4205, 2023. 5, 7
2023
-
[62]
Hoi-diff: Text-driven syn- thesis of 3d human-object interactions using diffusion mod- els
Xiaogang Peng, Yiming Xie, Zizhao Wu, Varun Jampani, Deqing Sun, and Huaizu Jiang. Hoi-diff: Text-driven syn- thesis of 3d human-object interactions using diffusion mod- els. arXiv preprint arXiv:2312.06553, 2023. 3
2023 arXiv
-
[63]
Object pop-up: Can we infer 3d objects and their poses from human interactions alone? In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023
Ilya A Petrov, Riccardo Marin, Julian Chibane, and Gerard Pons-Moll. Object pop-up: Can we infer 3d objects and their poses from human interactions alone? In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023. 3, 7
2023
-
[64]
Tridi: Trilateral diffusion of 3d humans, objects, and interactions
Ilya A Petrov, Riccardo Marin, Julian Chibane, and Ger- ard Pons-Moll. Tridi: Trilateral diffusion of 3d humans, objects, and interactions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025. 3, 4
2025
-
[65]
Project Aria , accessed January 7, 2025
Project Aria. Project Aria , accessed January 7, 2025. https://www.projectaria.com/. 2
2025
-
[66]
Project aria ma- chine perception services, accessed January 7, 2025
Project Aria Machine Perception Services. Project aria ma- chine perception services, accessed January 7, 2025. 2
2025
-
[67]
Pointnext: Revisiting pointnet++ with improved training and scaling strategies
Guocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai, Hasan Hammoud, Mohamed Elhoseiny, and Bernard Ghanem. Pointnext: Revisiting pointnet++ with improved training and scaling strategies. Advances in neural infor- mation processing systems, 35:23192–23204, 2022. 4
2022
-
[68]
Hierarchical text-conditional image generation with clip latents
Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 1(2):3, 2022. 4, 1
2022 arXiv
-
[69]
Egocap: egocentric marker-less motion capture with two fisheye cameras.ACM Transactions on Graphics (TOG), 35(6):1–11, 2016
Helge Rhodin, Christian Richardt, Dan Casas, Eldar In- safutdinov, Mohammad Shafiei, Hans-Peter Seidel, Bernt Schiele, and Christian Theobalt. Egocap: egocentric marker-less motion capture with two fisheye cameras.ACM Transactions on Graphics (TOG), 35(6):1–11, 2016. 2 11
2016
-
[70]
First-person pose recognition using egocentric workspaces
Gr ´egory Rogez, James S Supancic, and Deva Ramanan. First-person pose recognition using egocentric workspaces. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4325–4333, 2015. 2
2015
-
[71]
Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: Modeling and capturing hands and bod- ies together. ACM Transactions on Graphics, (Proc. SIG- GRAPH Asia), 36(6), 2017. 6
2017
-
[72]
Deep unsupervised learning using nonequilibrium thermodynamics
Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International confer- ence on machine learning, pages 2256–2265. PMLR, 2015. 4
2015
-
[73]
Hoianimator: Generating text-prompt human-object anima- tions using novel perceptive diffusion models
Wenfeng Song, Xinyu Zhang, Shuai Li, Yang Gao, Aimin Hao, Xia Hou, Chenglizhao Chen, Ning Li, and Hong Qin. Hoianimator: Generating text-prompt human-object anima- tions using novel perceptive diffusion models. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and...
2024
-
[74]
Neural localizer fields for continuous 3d human pose and shape estima- tion
Istv ´an S ´ar´andi and Gerard Pons-Moll. Neural localizer fields for continuous 3d human pose and shape estima- tion. Advances in Neural Information Processing Systems (NeurIPS), 2024. 6
2024
-
[75]
A unified diffusion framework for scene- aware human motion estimation from sparse signals
Jiangnan Tang, Jingya Wang, Kaiyang Ji, Lan Xu, Jingyi Yu, and Ye Shi. A unified diffusion framework for scene- aware human motion estimation from sparse signals. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition , pages 21251–21262, 2024. 3
2024
-
[76]
Human motion diffusion model
Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H Bermano. Human motion diffusion model. arXiv preprint arXiv:2209.14916, 2022. 5
2022 arXiv
-
[77]
Selfpose: 3d egocentric pose estimation from a headset mounted camera.IEEE Transactions on Pat- tern Analysis and Machine Intelligence , 45(6):6794–6806,
Denis Tome, Thiemo Alldieck, Patrick Peluse, Gerard Pons-Moll, Lourdes Agapito, Hernan Badino, and Fer- nando De la Torre. Selfpose: 3d egocentric pose estimation from a headset mounted camera.IEEE Transactions on Pat- tern Analysis and Machine Intelligence , 45(6):6794–6806,
-
[78]
Deco: Dense estimation of 3d human-scene contact in the wild
Shashank Tripathi, Agniv Chatterjee, Jean-Claude Passy, Hongwei Yi, Dimitrios Tzionas, and Michael J Black. Deco: Dense estimation of 3d human-scene contact in the wild. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 8001–8013, 2023. 3
2023
-
[79]
Practical motion capture in everyday surround- ings
Daniel Vlasic, Rolf Adelsberger, Giovanni Vannucci, John Barnwell, Markus Gross, Wojciech Matusik, and Jovan Popovi´c. Practical motion capture in everyday surround- ings. ACM transactions on graphics (TOG) , 26(3):35–es,
-
[80]
Sparse inertial poser: Automatic 3d hu- man pose estimation from sparse imus
Timo V on Marcard, Bodo Rosenhahn, Michael J Black, and Gerard Pons-Moll. Sparse inertial poser: Automatic 3d hu- man pose estimation from sparse imus. InComputer graph- ics forum, pages 349–360. Wiley Online Library, 2017. 2
2017
-
[81]
Egocentric whole-body motion capture with fisheyevit and diffusion-based motion refinement
Jian Wang, Zhe Cao, Diogo Luvizon, Lingjie Liu, Kri- pasindhu Sarkar, Danhang Tang, Thabo Beeler, and Chris- tian Theobalt. Egocentric whole-body motion capture with fisheyevit and diffusion-based motion refinement. In Pro- ceedings of the IEEE/CVF Conference on Computer Visio...
2024
-
[82]
Quest- sim: Human motion tracking from sparse sensors with sim- ulated avatars
Alexander Winkler, Jungdam Won, and Yuting Ye. Quest- sim: Human motion tracking from sparse sensors with sim- ulated avatars. In SIGGRAPH Asia 2022 Conference Pa- pers, pages 1–8, 2022. 2
2022
-
[83]
Chore: Contact, human and object reconstruction from a single rgb image
Xianghui Xie, Bharat Lal Bhatnagar, and Gerard Pons- Moll. Chore: Contact, human and object reconstruction from a single rgb image. In European Conference on Com- puter Vision (ECCV). Springer, 2022. 3
2022
-
[84]
Visibility aware human-object interaction tracking from single rgb camera
Xianghui Xie, Bharat Lal Bhatnagar, and Gerard Pons- Moll. Visibility aware human-object interaction tracking from single rgb camera. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023
2023
-
[85]
Template free reconstruction of human- object interaction with procedural interaction generation
Xianghui Xie, Bharat Lal Bhatnagar, Jan Eric Lenssen, and Gerard Pons-Moll. Template free reconstruction of human- object interaction with procedural interaction generation. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 10003–10015, 2024
2024
-
[86]
In- tertrack: Tracking human object interaction without object templates
Xianghui Xie, Jan Eric Lenssen, and Gerard Pons-Moll. In- tertrack: Tracking human object interaction without object templates. In International Conference on 3D Vision 2025,
2025
-
[87]
Regen- net: Towards human action-reaction synthesis
Liang Xu, Yizhou Zhou, Yichao Yan, Xin Jin, Wenhan Zhu, Fengyun Rao, Xiaokang Yang, and Wenjun Zeng. Regen- net: Towards human action-reaction synthesis. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1759–1769, 2024. 3
2024
-
[88]
Interdiff: Generating 3d human-object interactions with physics-informed diffusion
Sirui Xu, Zhengyuan Li, Yu-Xiong Wang, and Liang-Yan Gui. Interdiff: Generating 3d human-object interactions with physics-informed diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 14928–14940, 2023. 3, 7
2023
-
[89]
Inter- dreamer: Zero-shot text to 3d dynamic human-object in- teraction
Sirui Xu, Yu-Xiong Wang, Liangyan Gui, et al. Inter- dreamer: Zero-shot text to 3d dynamic human-object in- teraction. Advances in Neural Information Processing Sys- tems, 37:52858–52890, 2024. 3
2024
-
[90]
Mobileposer: Real-time full-body pose estimation and 3d human translation from imus in mobile consumer devices
Vasco Xu, Chenfeng Gao, Henry Hoffmann, and Karan Ahuja. Mobileposer: Real-time full-body pose estimation and 3d human translation from imus in mobile consumer devices. In Proceedings of the 37th Annual ACM Sympo- sium on User Interface Software and Technology, pages 1– 11, 2024. 2
2024
-
[91]
Mo 2Cap2 : Real-time mobile 3d motion capture with a cap-mounted fisheye camera
Weipeng Xu, Avishek Chatterjee, Michael Zollhoefer, Helge Rhodin, Pascal Fua, Hans-Peter Seidel, and Christian Theobalt. Mo 2Cap2 : Real-time mobile 3d motion capture with a cap-mounted fisheye camera. IEEE Transactions on Visualization and Computer Graphics, pages 1–1, 2019. 2
2019
-
[92]
F-hoi: Toward fine-grained semantic- aligned 3d human-object interactions
Jie Yang, Xuesong Niu, Nan Jiang, Ruimao Zhang, and Siyuan Huang. F-hoi: Toward fine-grained semantic- aligned 3d human-object interactions. In European Con- ference on Computer Vision, pages 91–110. Springer, 2024. 3
2024
-
[93]
Egochoir: Capturing 3d human-object interaction regions from egocentric views
Yuhang Yang, Wei Zhai, Chengfeng Wang, Chengjun Yu, Yang Cao, and Zheng-Jun Zha. Egochoir: Capturing 3d human-object interaction regions from egocentric views. Advances in Neural Information Processing Systems , 37: 54529–54557, 2024. 2
2024
-
[94]
Estimating body and hand motion in an ego- sensed world
Brent Yi, Vickie Ye, Maya Zheng, Yunqi Li, Lea M ¨uller, Georgios Pavlakos, Yi Ma, Jitendra Malik, and Angjoo 12 Kanazawa. Estimating body and hand motion in an ego- sensed world. arXiv preprint arXiv:2410.03665, 2024. 2, 3, 5, 6, 7
2024 arXiv
-
[95]
Transpose: Real- time 3d human translation and pose estimation with six in- ertial sensors
Xinyu Yi, Yuxiao Zhou, and Feng Xu. Transpose: Real- time 3d human translation and pose estimation with six in- ertial sensors. ACM Transactions On Graphics (TOG), 40 (4):1–13, 2021. 2
2021
-
[96]
Physical inertial poser (pip): Physics-aware real-time human motion tracking from sparse inertial sensors
Xinyu Yi, Yuxiao Zhou, Marc Habermann, Soshi Shi- mada, Vladislav Golyanik, Christian Theobalt, and Feng Xu. Physical inertial poser (pip): Physics-aware real-time human motion tracking from sparse inertial sensors. InPro- ceedings of the IEEE/CVF conference on computer vision...
2022
-
[97]
Yonemoto, K
H. Yonemoto, K. Murasaki, T. Osawa, K. Sudo, J. Shima- mura, and Y . Taniguchi. Egocentric articulated pose track- ing for action recognition. In International Conference on Machine Vision Applications (MVA), 2015. 2
2015
-
[98]
Neuraldome: A neural modeling pipeline on multi- view human-object interactions
Juze Zhang, Haimin Luo, Hongdi Yang, Xinru Xu, Qianyang Wu, Ye Shi, Jingyi Yu, Lan Xu, and Jingya Wang. Neuraldome: A neural modeling pipeline on multi- view human-object interactions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages ...
2023
-
[99]
Hoi-mˆ 3: Capture multiple humans and objects in- teraction within contextual environment
Juze Zhang, Jingyan Zhang, Zining Song, Zhanhe Shi, Chengfeng Zhao, Ye Shi, Jingyi Yu, Lan Xu, and Jingya Wang. Hoi-mˆ 3: Capture multiple humans and objects in- teraction within contextual environment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern R...
2024
-
[100]
Egobody: Human body shape and motion of interacting people from head-mounted devices
Siwei Zhang, Qianli Ma, Yan Zhang, Zhiyin Qian, Taein Kwon, Marc Pollefeys, Federica Bogo, and Siyu Tang. Egobody: Human body shape and motion of interacting people from head-mounted devices. In European confer- ence on computer vision , pages 180–200. Springer, 2022. 2
2022
-
[101]
Force: Physics-aware human-object interaction
Xiaohan Zhang, Bharat Lal Bhatnagar, Sebastian Starke, Ilya Petrov, Vladimir Guzov, Helisa Dhamo, Ed- uardo P ´erez-Pellitero, and Gerard Pons-Moll. Force: Physics-aware human-object interaction. arXiv preprint arXiv:2403.11237, 2024. 3
2024 arXiv
-
[102]
Scenic: Scene-aware semantic navi- gation with instruction-guided control
Xiaohan Zhang, Sebastian Starke, Vladimir Guzov, Zhensong Zhang, Eduardo P ´erez Pellitero, and Ger- ard Pons-Moll. Scenic: Scene-aware semantic navi- gation with instruction-guided control. arXiv preprint arXiv:2412.15664, 2024. 3
2024 arXiv
-
[103]
Instance tracking in 3d scenes from egocentric videos
Yunhan Zhao, Haoyu Ma, Shu Kong, and Charless Fowlkes. Instance tracking in 3d scenes from egocentric videos. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition , pages 21933–21944, 2024. 8
2024
-
[104]
Realistic full-body tracking from sparse ob- servations via joint-level modeling
Xiaozheng Zheng, Zhuo Su, Chao Wen, Zhou Xue, and Xiaojie Jin. Realistic full-body tracking from sparse ob- servations via joint-level modeling. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 14678–14688, 2023. 2, 7
2023
-
[105]
On the continuity of rotation representations in neu- ral networks
Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. On the continuity of rotation representations in neu- ral networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 5745– 5753, 2019. 4
2019
-
[106]
Loose inertial poser: Motion capture with imu-attached loose-wear jacket
Chengxu Zuo, Yiming Wang, Lishuang Zhan, Shihui Guo, Xinyu Yi, Feng Xu, and Yipeng Qin. Loose inertial poser: Motion capture with imu-attached loose-wear jacket. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 2209–2219, 2024. 2 13...
2024
-
[107]
This research direction could enable new applications for study- ing human behavior and developing realistic virtual expe- riences
Broader impacts The ability of our model to capture and generate contin- uous human-object interactions offers significant value for fields such as digital content creation and ergonomics. This research direction could enable new applications for study- ing human behavior and ...
-
[108]
The forward diffusion process can be for- mulated as a Markov chain with T steps
Background and Notation Background. The forward diffusion process can be for- mulated as a Markov chain with T steps. Starting from a clean sample z0 it produces a series of distributions q(zt|zt−1): q(z1:T|z0) = QT t=1q(zt|zt−1). We add noise to the distribution for T steps, ...
-
[109]
Fol- lowing the evaluation of ECHO with sparseH orO track- ing, we test the model’s performance withI data provided as an additional conditioning
Additional evaluation Evaluating the model with contacts conditioning. Fol- lowing the evaluation of ECHO with sparseH orO track- ing, we test the model’s performance withI data provided as an additional conditioning. We report the results in Ta- ble S2. Providing contact info...
-
[110]
Losses and Metrics Losses. The objective function used to train our network is the weighted combination of the following losses: LH n =∥θH−bθH∥2 +∥TH−bTH∥2 LO n =∥TO−bTO∥2 LI n =∥cI−bcI∥2 LO v =∥VO− bVO∥2 LH j =∥JH−bJH∥2 LH s =∥cfeet I ∗ bU feet H ∥2 (12) where bU feet H is th...
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.