Pith. sign in

REVIEW 4 major objections 6 minor 110 references

ECHO: Ego-Centric modeling of Human-Object interactions

T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read Three trackers recover full human and object motion together

desk verdict Useful new generative model for egocentric HOI with a genuinely novel tri-variate diffusion, but the 'solely from head and wrist tracking' claim only holds if you already know the object and have its canonical mesh, which needs fixing before the paper is adopted. read the letter →

arxiv 2508.21556 v3 pith:E2IVXB2W submitted 2025-08-29 cs.CV

classification cs.CV
keywords egocentrichuman-objectinteractionsparsemotiontrackingdiffusiontransformertri-variatecontactmodelinghead-centricrepresentationhumanreconstructionobjectposeestimation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

ECHO claims that the joint motion of a person and the object they are manipulating can be reconstructed from just three sparse points of tracking: the head and the two wrists. The paper argues that a single generative model can recover body pose, object trajectory, and contact pattern at once by treating them as three linked diffusion processes with independent noise schedules. If ECHO is right, the sensors already embedded in smart glasses and watches could support plausible 3D reconstruction of everyday manipulation without cameras, suits, or scene scans. The authors report that ECHO outperforms adapted baselines on standard human-object interaction benchmarks and stays competitive on motion-only data.

What carries the argument

The carrying mechanism is a tri-variate diffusion process: human motion, object trajectory, and contact sequence are noised and denoised jointly, each with its own timestep per frame, inside a diffusion transformer. This per-frame, per-modality timestamp scheme is what turns the model into a flexible conditioning machine—any known portion of any modality can be supplied at a low noise level, and the model can attend to it while predicting the rest. The second piece is a conveyor inference: the denoising step increases monotonically along the temporal window, so completed frames stream out the front while fully noised frames enter from the back, enabling arbitrary-length, temporally consistent inference. The third piece is a head-centric canonical frame, expressed relative to the head pose at the first frame with the vertical axis aligned to gravity, which removes global orientation as a nuisance and, per the ablation, sharply improves prediction quality.

What would settle it

Take an unseen manipulation sequence recorded with only head and wrist trackers, supply the object mesh, and run ECHO while a baseline simply keeps the object at its most likely class-conditioned pose; if ECHO does not beat that baseline by a clear margin on object position and rotation error, then the apparent success would be better attributed to the dataset prior than to the sparse tracking signal.

Watch

Extended reading notes

Core claim

On its own terms, the paper establishes that human pose, object motion, and contact can be generated jointly from head-and-wrist conditioning by diffusing the three modalities together inside one transformer. The model works in a head-centric canonical frame to remove global-orientation bias, assigns an independent denoising timestamp to every frame and modality, and denoises a sliding temporal window so sequences of arbitrary length can be produced in real time. The authors show that feeding in sparse observations of any one modality—a few frames of human tracking, a few frames of object tracking, or partial contact labels—improves the reconstruction of the other modalities, and that ablations remove most of the gains when the contact modality, the head-centric representation, or the auxiliary motion data is removed. The central quantitative claim is state-of-the-art performance on both human and object metrics against a motion-diffusion baseline extended to object modeling, on two interaction datasets.

Load-bearing premise

At test time the user must know the object's class and provide its 3D canonical mesh, and the body shape parameters are assumed known; without that geometry and identity there is no object representation to denoise, so the operation is not fully self-contained from tracking alone.

Editorial extensions

If this is right

  • Wearable-only setups in AR and VR could animate full-body avatars and manipulated objects in real time; the paper reports about 13.7 ms of inference per frame on a consumer GPU.
  • A few observed frames of human, object, or contact information can be folded into the reconstruction to constrain the other modalities, so intermittent tracking does not break the output.
  • Because the three modalities are noised independently, training can mix large motion-only datasets with smaller interaction datasets, giving the model a strong human-motion prior without sacrificing interaction detail.
  • Conveyor inference removes the sliding-window stitching problem and supports arbitrarily long sequences, which is what a continuous wearable system would need.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • As a practical consequence the paper does not develop, ECHO cannot handle an object whose canonical mesh and class label are unavailable at test time; an obvious extension is to estimate the template on the fly from an egocentric camera and to quantify how template error propagates into trajectory error.
  • The same joint-distribution formulation could be inverted into an interaction simulator: condition on the object's trajectory and generate the human response, or condition on the human and generate plausible object behavior, which would be useful for robotics and content creation.
  • The per-frame timestamp mechanism suggests a natural online-filtering reading: if observed streams are fed in at low noise levels, ECHO could act as a continuously correcting state estimator rather than a one-shot generator, though the paper does not evaluate this mode directly.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The paper introduces ECHO, a diffusion-transformer framework that jointly reconstructs human body pose, object motion, and contact signals from sparse egocentric sensors (head and wrist tracking). The key technical proposals are a tri-variate diffusion process with independent per-frame, per-modality noise schedules; a head-centric canonical coordinate representation; a conveyor-based autoregressive inference scheme for arbitrarily long sequences; and training on a mixture of HOI datasets (BEHAVE, OMOMO) plus large-scale human motion data (AMASS). Experiments report quantitative results on BEHAVE and OMOMO against a self-constructed BoDiffusion+Obj baseline, motion-generation comparisons on AMASS, sparse-tracking evaluations, and ablations supporting the design choices. The central claims are that ECHO is the first method to recover human and object motion jointly from 3-point tracking and that it achieves state-of-the-art egocentric HOI reconstruction.

Significance. If the claims hold, ECHO would be a meaningful step toward practical egocentric HOI capture from commodity wearables. The tri-variate diffusion formulation with per-frame timestamps is an interesting generalization of prior diffusion models, and the head-centric representation with conveyor inference addresses real sequence-length and orientation issues. The paper's strengths include internally consistent ablations (e.g., the benefit of the contact modality, AMASS pretraining, and head-centric coordinates), a clear exposition of the training objective, and a stated commitment to release code and models. The main reservation is that the headline claim of operating 'solely from head and wrist tracking' is contradicted by the requirement that the object's canonical mesh and class label be provided as input; the state-of-the-art claim is also weakened by the absence of quantitative comparison with the closest existing egocentric HOI method (iReplica) and the prior trilateral diffusion method (TriDi).

major comments (4)
  1. [Abstract and Section 3.1] The abstract and contribution list claim that ECHO recovers human pose, object motion, and contact 'solely from head and wrist tracking' and is 'the first method to solve for human and object motion sequences jointly, relying only on 3-point tracking.' Section 3.1 states that 'For every object we assume that its canonical mesh is given to the model as an input,' and the object conditioning C_O = (y_O, f_O) is computed from a one-hot class label and PointNext features of the canonicalized object vertices. Without this mesh, the object modality has no representation to denoise, so object trajectory cannot be predicted at all. The method therefore solves a conditional generation problem (3-point tracking plus known object identity and shape), not the sensor-only problem advertised. This is a material limitation that should be disclosed in the abstract and contributions, or the claims should be reworded to state the actual input requirements.
  2. [Section 4.1, Tables 1-2] The only HOI baseline in the comparison is BoDiffusion+Obj, which is constructed by the authors on top of BoDiffusion. There is no quantitative comparison with iReplica, the only existing egocentric HOI method cited as 'the closest approach to ours,' nor with TriDi, the trilateral diffusion model on which the proposed formulation is directly built. The paper's claim of 'state-of-the-art, significantly outperforming existing methods' is therefore not supported for the egocentric HOI setting. The authors should either add experiments against these methods (where input requirements are compatible) or explicitly state and justify why a quantitative comparison is not possible, and temper the SOTA claim accordingly.
  3. [Section 4.1, Table 2] The text states that ECHO 'performs on par or better than BoDiffusion+Obj' on AMASS, but Table 2 shows ECHO's MPJPE (93.9±8.7) is worse than BoDiffusion+Obj (91.5±4.2), while MPJVE is better (109.7 vs 115.5). Given the large variances, 'on par' may be defensible, but the claim as written is imprecise and should be corrected to reflect the actual direction of the differences.
  4. [Section 4.3, Tables 4 and S3] The ablation of the inference-time guidance (NoGuide) shows negligible differences on BEHAVE (human MPJPE 61.4 vs 61.2, object Ev2v 29.5 vs 29.8) and on OMOMO (Table S3, human MPJPE 64.1 vs 64.5, object Ev2v 26.7 vs 26.5). The paper claims guidance is useful, but these differences are well within the reported variances. The guidance's contribution should be characterized more cautiously, or the experimental setup (e.g., which weights are used) should be clarified to show where the benefit actually appears.
minor comments (6)
  1. [Abstract] The phrase 'tri-variate diffusion process with independent noise schedules' is ambiguous: the independent schedules are per frame and per modality, not merely per modality. Consider rewording to make the granularity explicit.
  2. [Section 3.2, Eq. (7)] The shorthand Ep, Et, Eq introduced in the background is reused in Eq. (7), but the subscripts are not restated; adding a one-line reminder would improve readability.
  3. [Table 2] The FC column is labeled 'FC↑1.0' in the header, which is confusing. The metric should be described in the caption as a fraction where higher is better, and the '1.0' in the header should be removed.
  4. [Section 4.3] There is a typo 'in thew Sup.Mat.'; it should read 'in the Sup. Mat.'
  5. [Figure 5] The qualitative comparison shows 'BoDiffusion + Obj.' but not iReplica or TriDi. Adding qualitative results for at least one of those methods, or a sentence explaining their absence, would strengthen the comparison.
  6. [Section 3.1] The notation SMPL (T_H, θ_H) should refer to SMPL-X consistently, as the paragraph begins by naming SMPL-X as the body model.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: ECHO's diffusion objective and evaluations are self-contained; the object-mesh input assumption is an overclaim, not a circular reduction.

full rationale

ECHO is a learned generative system, not a symbolic derivation, and no step reduces by construction to its own inputs. The training objective in Eq. 7 is a standard conditional DDPM loss: the network is trained to recover clean (H0,O0,I0) from independently noised versions of the three modalities plus conditioning (E, C_O), and the output is not algebraically forced to equal any input. The object conditioning C_O=(y_O,f_O) is a test-time input disclosed in Sec. 3.1 ('For every object we assume that its canonical mesh is given to the model as an input'), so the advertised 'solely from head and wrist tracking' wording overstates the sensor requirements, but this is a scope/correctness issue rather than a circular reduction: the object pose O is still a generated SE(3) sequence and is not obtained by refitting C_O. Citations to TriDi and UniDiffuser are to published methods that ECHO explicitly extends, the tri-variate objective is written out in Eqs. 6-7, and no uniqueness theorem or ansatz is imported from self-citations. Contact labels are defined from H/O geometry, but they are used as an auxiliary training signal and as an inference-time consistency constraint, not as the ground-truth target that defines H/O; the ablation shows empirical value. Evaluation is performed against held-out BEHAVE/OMOMO/AMASS splits with standard metrics, so the central performance claims are externally falsifiable. The only flagged limitation (canonical mesh and class label required at test time) undercuts the abstract's 'solely' wording but does not make the derivation circular.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The central claim depends on the trained model and on several domain assumptions. The most consequential is the availability of the object's canonical mesh at test time, which the abstract does not disclose. Learned network weights are fitted to data; hand-selected hyperparameters such as loss weights, contact threshold, and window size are listed as free parameters.

free parameters (6)
  • Model weights (47.3M parameters) = trained on AMASS + BEHAVE + OMOMO
    The denoising network is fitted end-to-end; its weights are the primary fitted quantities of the system. They are not claim-level constants, but they are the empirical content of the method.
  • Loss weighting coefficients = lambda_Hn = lambda_On = lambda_In = 1, lambda_Ov = lambda_Hj = 0.1, lambda_Hs = 0.05
    Hand-chosen in Equation 13 of the supplementary material; they trade off human pose, object pose, contact, and foot-skating objectives.
  • Contact distance threshold tau_c = not stated in main text or provided supplementary
    Defines the binary contact map c_I in Equation 3; a chosen hyperparameter on which all contact labels depend.
  • Temporal window size W = 60 frames
    Chosen for training sequence windows and conveyor scheduling; affects context length and inference stitching.
  • Inference denoising steps = 100
    DDPM scheduler steps at inference, chosen as a quality and latency tradeoff.
  • Contact point set P_c = 64 points uniformly sampled from SMPL-X vertices
    Contact representation size depends on this sampling choice; reported in supplementary Table S1.
assumptions (5)
  • domain assumption The object's canonical mesh and one-hot class label are available at test time.
    Stated in Section 3.1 Object paragraph; PointNext features and class encodings are computed from canonicalized vertices, and object motion is output as SE(3) relative to that mesh. The abstract's 'solely from head and wrist tracking' omits this input.
  • domain assumption Ground-truth 3-point tracking of head and wrists is available for conditioning.
    Section 3.1 Egocentric conditioning constructs an input vector from head and hand joint rotations, positions, velocities; no noise model for the tracking signal is considered.
  • domain assumption SMPL-X body model with known shape parameters beta is a faithful representation of the human body.
    Section 3.1 Human; ECHO predicts only pose theta_H and root transform T_H, treating shape and face parameters as fixed and known.
  • standard math DDPM forward and reverse processes with a DiT denoiser provide a valid generative model.
    Section 3.2 Background; the training objective and reverse process follow Ho et al. [30] and DiT [61] without modification beyond the multi-modality timestamps.
  • domain assumption Head-centric canonicalization with a gravity-aligned height axis removes global orientation bias.
    Section 3.1 Head-centric modeling; used to define all coordinate frames. The paper provides an ablation (NoCan) showing improved performance but no theoretical justification for the representation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of ECHO: Ego-Centric modeling of Human-Object interactions." pith.science (2026). https://pith.science/paper/E2IVXB2W

@misc{pith2026250821556,
  author       = {Pith},
  title        = {Pith review of: ECHO: Ego-Centric modeling of Human-Object interactions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/E2IVXB2W}},
  note         = {Machine review of arXiv:2508.21556}
}
read the original abstract

Modeling human-object interactions (HOI) from an egocentric perspective is a critical yet challenging task, particularly when relying on sparse signals from wearable devices like smart glasses and watches. We present ECHO, the first unified framework to jointly recover human pose, object motion, and contact dynamics solely from head and wrist tracking. To tackle the underconstrained nature of this problem, we introduce a novel tri-variate diffusion process with independent noise schedules that models the mutual dependencies between the human, object, and interaction modalities. This formulation allows ECHO to operate with flexible input configurations, making it robust to intermittent tracking and capable of leveraging partial observations. Crucially, it enables training on a combination of large-scale human motion datasets and smaller HOI collections, learning strong priors while capturing interaction nuances. Furthermore, we employ a smooth inpainting inference mechanism that enables the generation of temporally consistent interactions for arbitrarily long sequences. Extensive evaluations demonstrate that ECHO achieves state-of-the-art performance, significantly outperforming existing methods lacking such flexibility. The project page is available at https://ptrvilya.github.io/echo/.

Figures

Figures reproduced from arXiv: 2508.21556 by the authors.

Figure 1
Figure 1. ECHO. Wearable devices like smart glasses, watches, and rings are becoming more and more affordable at the customer level. But what can be inferred from such a sparse set of sensors? ECHO is the first model to recover Human-Object Interaction sequences (top) from sparse 3-point tracking. Our method is flexible and can operate in different modes (bottom), for example, incorporating available information (shown in red… view at source ↗
Figure 2
Figure 2. Head-centric rep￾resentation. ECHO operates in a head-centric coordinate system. This choice removes spatial bias from the global configuration. Head-centric modeling. The key design choice for our method was selecting the coordinate system for the network’s output. Although a global coor￾dinate frame might seem like a natural choice for modeling human-object interactions, we found that it introduces significant bia… view at source ↗
Figure 4
Figure 4. Conveyor Inference in Head-centric coordinates. In the standard formulation, all the frames are denoised in parallel with the same timestamp (left). We propose a conveyor diffu￾sion (right), where the denoising timestamp gradually increases towards the end of the window. When the first frame is fully dif￾fused, a new fully noised frame gets appended at the tail, with the appropriate coordinate transform. from ego-ce… view at source ↗
Figures from the paper (1 more)
Figure 5
Figure 5. Figure 5: Qualitative results of ECHO. Our method (middle) can infer movements well-aligned with the ground truth (left) with a vast variety of objects and motions. Competitors (right) struggle especially with large movements, causing implausible predictions such as penetrations…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

110 extracted references · 64 canonical work pages

  1. [1]

    Un- realego: A new dataset for robust egocentric 3d human mo- tion capture

    Hiroyasu Akada, Jian Wang, Soshi Shimada, Masaki Taka- hashi, Christian Theobalt, and Vladislav Golyanik. Un- realego: A new dataset for robust egocentric 3d human mo- tion capture. In European Conference on Computer Vision, pages 1–17. Springer, 2022. 2

  2. [2]

    Introducing hot3d: An egocentric dataset for 3d hand and object tracking.arXiv preprint arXiv:2406.09598,

    Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon, Shreyas Hampali, Fan Zhang, Jade Fountain, Edward Miller, Selen Basol, Richard Newcombe, Robert Wang, et al. Introducing hot3d: An egocentric dataset for 3d hand and object tracking.arXiv preprint arXiv:2406.09598,

  3. [3]

    One transformer fits all distributions in multi-modal diffu- sion at scale

    Fan Bao, Shen Nie, Kaiwen Xue, Chongxuan Li, Shi Pu, Yaole Wang, Gang Yue, Yue Cao, Hang Su, and Jun Zhu. One transformer fits all distributions in multi-modal diffu- sion at scale. In Proceedings of the 40th International Con- ference on Machine Learning. JMLR.org, 2023. 4

  4. [4]

    From sparse signal to smooth motion: Real-time motion generation with rolling prediction models

    German Barquero, Nadine Bertsch, Manojkumar Marram- reddy, Carlos Chac ´on, Filippo Arcadu, Ferran Rigual, Nicky Sijia He, Cristina Palmero, Sergio Escalera, Yuting Ye, et al. From sparse signal to smooth motion: Real-time motion generation with rolling prediction models. In Pro- ceedings of the Computer Vision and Pattern Recognition Conference, pages 18...

  5. [5]

    Bharat Lal Bhatnagar, Suriya Singh, Chetan Arora, and C.V . Jawahar. Unsupervised learning of deep feature repre- sentation for clustering egocentric actions. In Proceedings of the Twenty-Sixth International Joint Conference on Arti- ficial Intelligence, IJCAI-17, pages 1447–1453, 2017. 2

  6. [6]

    Behave: Dataset and method for tracking human object in- teractions

    Bharat Lal Bhatnagar, Xianghui Xie, Ilya A Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. Behave: Dataset and method for tracking human object in- teractions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 15935– 15946, 2022. 3, 6

  7. [7]

    Physically plausible full-body hand-object interaction synthesis

    Jona Braun, Sammy Christen, Muhammed Kocabas, Emre Aksan, and Otmar Hilliges. Physically plausible full-body hand-object interaction synthesis. In 2024 International Conference on 3D Vision (3DV) , pages 464–473. IEEE,

  8. [8]

    Egocentric gesture recognition using recurrent 3d convolutional neural networks with spatiotemporal trans- former modules

    Congqi Cao, Yifan Zhang, Yi Wu, Hanqing Lu, and Jian Cheng. Egocentric gesture recognition using recurrent 3d convolutional neural networks with spatiotemporal trans- former modules. 2017 IEEE International Conference on Computer Vision (ICCV), 2017. 2

Show all 110 references
  1. [9]

    Avatargo: Zero-shot 4d human-object interaction generation and animation

    Yukang Cao, Liang Pan, Kai Han, Kwan-Yee K Wong, and Ziwei Liu. Avatargo: Zero-shot 4d human-object interaction generation and animation. arXiv preprint arXiv:2410.07164, 2024. 3

  2. [10]

    Bodiffusion: Diffusing sparse observations for full-body human motion synthesis

    Angela Castillo, Maria Escobar, Guillaume Jeanneret, Al- bert Pumarola, Pablo Arbel ´aez, Ali Thabet, and Artsiom Sanakoyeu. Bodiffusion: Diffusing sparse observations for full-body human motion synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vis...

  3. [11]

    Diffusion forcing: Next-token prediction meets full-sequence diffu- sion

    Boyuan Chen, Diego Mart ´ı Mons ´o, Yilun Du, Max Sim- chowitz, Russ Tedrake, and Vincent Sitzmann. Diffusion forcing: Next-token prediction meets full-sequence diffu- sion. Advances in Neural Information Processing Systems, 37:24081–24125, 2024. 5

  4. [12]

    Muscles in action

    Mia Chiquier and Carl V ondrick. Muscles in action. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 22091–22101, 2023. 2

  5. [13]

    Hmd-poser: On-device real-time human motion tracking from scalable sparse observations

    Peng Dai, Yang Zhang, Tao Liu, Zhen Fan, Tianyuan Du, Zhuo Su, Xiaozheng Zheng, and Zeming Li. Hmd-poser: On-device real-time human motion tracking from scalable sparse observations. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages ...

  6. [14]

    In- terfusion: Text-driven generation of 3d human-object inter- action

    Sisi Dai, Wenhao Li, Haowen Sun, Haibin Huang, Chongyang Ma, Hui Huang, Kai Xu, and Ruizhen Hu. In- terfusion: Text-driven generation of 3d human-object inter- action. In European Conference on Computer Vision, pages 18–35. Springer, 2024. 3

  7. [15]

    Hsc4d: Human- centered 4d scene capture in large-scale indoor-outdoor space using wearable imus and lidar

    Yudi Dai, Yitai Lin, Chenglu Wen, Siqi Shen, Lan Xu, Jingyi Yu, Yuexin Ma, and Cheng Wang. Hsc4d: Human- centered 4d scene capture in large-scale indoor-outdoor space using wearable imus and lidar. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogn...

  8. [16]

    Diffusion models beat gans on image synthesis

    Prafulla Dhariwal and Alexander Nichol. Diffusion models beat gans on image synthesis. Advances in neural informa- tion processing systems, 34:8780–8794, 2021. 6

  9. [17]

    Cg-hoi: Contact-guided 3d human-object interaction generation

    Christian Diller and Angela Dai. Cg-hoi: Contact-guided 3d human-object interaction generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19888–19901, 2024. 3

  10. [18]

    Avatars grow legs: Generating smooth human motion from sparse track- ing inputs with diffusion model

    Yuming Du, Robin Kips, Albert Pumarola, Sebastian Starke, Ali Thabet, and Artsiom Sanakoyeu. Avatars grow legs: Generating smooth human motion from sparse track- ing inputs with diffusion model. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogniti...

  11. [19]

    9 Egocast: Forecasting egocentric human pose in the wild

    Maria Escobar, Juanita Puentes, Cristhian Forigua, Jordi Pont-Tuset, Kevis-Kokitsi Maninis, and Pablo Arbelaez. 9 Egocast: Forecasting egocentric human pose in the wild. In 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), pages 5831–5841. IEEE, 2025. 2

  12. [20]

    Arctic: A dataset for dexterous bimanual hand- object manipulation

    Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J Black, and Otmar Hilliges. Arctic: A dataset for dexterous bimanual hand- object manipulation. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages...

  13. [21]

    Understand- ing egocentric activities

    Alireza Fathi, Ali Farhadi, and James M Rehg. Understand- ing egocentric activities. In 2011 international conference on computer vision, pages 407–414. IEEE, 2011. 2

  14. [22]

    The matrix: Infinite-horizon world generation with real-time moving control

    Ruili Feng, Han Zhang, Zhantao Yang, Jie Xiao, Zhilei Shu, Zhiheng Liu, Andy Zheng, Yukun Huang, Yu Liu, and Hongyang Zhang. The matrix: Infinite-horizon world generation with real-time moving control. arXiv preprint arXiv:2412.03568, 2024. 5

  15. [23]

    Imos: Intent- driven full-body motion synthesis for human-object inter- actions

    Anindita Ghosh, Rishabh Dabral, Vladislav Golyanik, Christian Theobalt, and Philipp Slusallek. Imos: Intent- driven full-body motion synthesis for human-object inter- actions. In Computer Graphics Forum, pages 1–12. Wiley Online Library, 2023. 3

  16. [24]

    Human poseitioning system (hps): 3d human pose estimation and self-localization in large scenes from body-mounted sensors

    Vladimir Guzov, Aymen Mir, Torsten Sattler, and Gerard Pons-Moll. Human poseitioning system (hps): 3d human pose estimation and self-localization in large scenes from body-mounted sensors. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2021. 2, 3

  17. [25]

    Interaction replica: Tracking human–object interac- tion and scene changes from human motion

    Vladimir Guzov, Julian Chibane, Riccardo Marin, Yannan He, Yunus Saracoglu, Torsten Sattler, and Gerard Pons- Moll. Interaction replica: Tracking human–object interac- tion and scene changes from human motion. In Interna- tional Conference on 3D Vision (3DV), 2024. 2, 3

  18. [26]

    Blendify – python rendering framework for blender

    Vladimir Guzov, Ilya A Petrov, and Gerard Pons-Moll. Blendify – python rendering framework for blender. arXiv preprint arXiv:2410.17858, 2024. 6

  19. [27]

    Karen Liu, Yuting Ye, and Lingni Ma

    Vladimir Guzov, Yifeng Jiang, Fangzhou Hong, Gerard Pons-Moll, Richard Newcombe, C. Karen Liu, Yuting Ye, and Lingni Ma. Hmd2: Environment-aware motion genera- tion from single egocentric head-mounted device. In Inter- national Conference on 3D Vision (3DV), 2025. 2, 3, 5

  20. [28]

    Resolving 3d human pose ambiguities with 3d scene constraints

    Mohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, and Michael J Black. Resolving 3d human pose ambiguities with 3d scene constraints. In Proceedings of the IEEE/CVF international conference on computer vision , pages 2282– 2292, 2019. 3

  21. [29]

    Populating 3d scenes by learning human-scene interaction

    Mohamed Hassan, Partha Ghosh, Joachim Tesch, Dimitrios Tzionas, and Michael J Black. Populating 3d scenes by learning human-scene interaction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 14708–14718, 2021. 3

  22. [30]

    Denoising dif- fusion probabilistic models

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. Advances in neural informa- tion processing systems, 33:6840–6851, 2020. 4, 1

  23. [31]

    Egosim: An egocentric multi-view simulator and real dataset for body-worn cameras during motion and activity

    Dominik Hollidt, Paul Streli, Jiaxi Jiang, Yasaman Haghighi, Changlin Qian, Xintong Liu, and Christian Holz. Egosim: An egocentric multi-view simulator and real dataset for body-worn cameras during motion and activity. Advances in Neural Information Processing Systems , 37: 10...

  24. [32]

    Microsoft HoloLens, accessed January 7, 2025

    HoloLens. Microsoft HoloLens, accessed January 7, 2025. https://learn.microsoft.com/en-us/hololens/. 2

  25. [33]

    Diffusion- based generation, optimization, and planning in 3d scenes

    Siyuan Huang, Zan Wang, Puhao Li, Baoxiong Jia, Tengyu Liu, Yixin Zhu, Wei Liang, and Song-Chun Zhu. Diffusion- based generation, optimization, and planning in 3d scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16750–16761, 2023. 3

  26. [34]

    Black, Otmar Hilliges, and Gerard Pons-Moll

    Yinghao Huang, Manuel Kaufmann, Emre Aksan, Michael J. Black, Otmar Hilliges, and Gerard Pons-Moll. Deep inertial poser: Learning to reconstruct human pose from sparse inertial measurements in real time. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia), 37(6): 185:1–185:15, 2018. 2

  27. [35]

    Black, and Dim- itrios Tzionas

    Yinghao Huang, Omid Taheri, Michael J. Black, and Dim- itrios Tzionas. InterCap: Joint markerless 3D tracking of humans and objects in interaction from multi-view RGB-D images. International Journal of Computer Vision (IJCV),

  28. [36]

    Avatar- poser: Articulated full-body pose tracking from sparse mo- tion sensing

    Jiaxi Jiang, Paul Streli, Huajian Qiu, Andreas Fender, Larissa Laich, Patrick Snape, and Christian Holz. Avatar- poser: Articulated full-body pose tracking from sparse mo- tion sensing. In European conference on computer vision, pages 443–460. Springer, 2022. 2

  29. [37]

    Scaling up dynamic human-scene interaction mod- eling

    Nan Jiang, Zhiyuan Zhang, Hongjie Li, Xiaoxuan Ma, Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, and Siyuan Huang. Scaling up dynamic human-scene interaction mod- eling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 1737– 1747, 2024. 3

  30. [38]

    Neuralho- fusion: Neural volumetric rendering under human-object interactions

    Yuheng Jiang, Suyi Jiang, Guoxing Sun, Zhuo Su, Kai- wen Guo, Minye Wu, Jingyi Yu, and Lan Xu. Neuralho- fusion: Neural volumetric rendering under human-object interactions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6155– 6165, 2022. 3

  31. [39]

    Transformer in- ertial poser: Real-time human motion reconstruction from sparse imus with simultaneous terrain generation

    Yifeng Jiang, Yuting Ye, Deepak Gopinath, Jungdam Won, Alexander W Winkler, and C Karen Liu. Transformer in- ertial poser: Real-time human motion reconstruction from sparse imus with simultaneous terrain generation. In SIG- GRAPH Asia 2022 Conference Papers, pages 1–9, 2022. 2

  32. [40]

    Ego3dpose: Capturing 3d cues from binocular egocentric views

    Taeho Kang, Kyungjin Lee, Jinrui Zhang, and Youngki Lee. Ego3dpose: Capturing 3d cues from binocular egocentric views. In SIGGRAPH Asia 2023 Conference Papers, pages 1–10, 2023. 2

  33. [41]

    Em-pose: 3d human pose estimation from sparse electromagnetic trackers

    Manuel Kaufmann, Yi Zhao, Chengcheng Tang, Lingling Tao, Christopher Twigg, Jie Song, Robert Wang, and Otmar Hilliges. Em-pose: 3d human pose estimation from sparse electromagnetic trackers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 1151...

  34. [42]

    aitviewer, 2022

    Manuel Kaufmann, Velko Vechev, and Dario Mylonopou- los. aitviewer, 2022. 6

  35. [43]

    Nifty: Neural object interaction fields for guided human motion synthesis

    Nilesh Kulkarni, Davis Rempe, Kyle Genova, Abhi- jit Kundu, Justin Johnson, David Fouhey, and Leonidas 10 Guibas. Nifty: Neural object interaction fields for guided human motion synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , p...

  36. [44]

    Mocap everyone everywhere: Lightweight motion capture with smartwatches and a head- mounted camera

    Jiye Lee and Hanbyul Joo. Mocap everyone everywhere: Lightweight motion capture with smartwatches and a head- mounted camera. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 1091–1100, 2024. 2, 3

  37. [45]

    Rewind: Real-time egocentric whole- body motion diffusion with exemplar-based identity condi- tioning

    Jihyun Lee, Weipeng Xu, Alexander Richard, Shih-En Wei, Shunsuke Saito, Shaojie Bai, Te-Li Wang, Minhyuk Sung, Jason Saragih, et al. Rewind: Real-time egocentric whole- body motion diffusion with exemplar-based identity condi- tioning. arXiv preprint arXiv:2504.04956, 2025. 2

  38. [46]

    Ego-body pose es- timation via ego-head pose estimation

    Jiaman Li, Karen Liu, and Jiajun Wu. Ego-body pose es- timation via ego-head pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17142–17151, 2023. 2, 7

  39. [47]

    Object motion guided human motion synthesis

    Jiaman Li, Jiajun Wu, and C Karen Liu. Object motion guided human motion synthesis. ACM Transactions on Graphics (TOG), 42(6):1–11, 2023. 3, 6

  40. [48]

    Controllable human-object interaction synthesis

    Jiaman Li, Alexander Clegg, Roozbeh Mottaghi, Jiajun Wu, Xavier Puig, and C Karen Liu. Controllable human-object interaction synthesis. In European Conference on Com- puter Vision, pages 54–72. Springer, 2024. 3

  41. [49]

    Task-oriented human-object interactions generation with implicit neural representations

    Quanzhou Li, Jingbo Wang, Chen Change Loy, and Bo Dai. Task-oriented human-object interactions generation with implicit neural representations. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 3035–3044, 2024. 3

  42. [50]

    Egohdm: An online egocentric-inertial human motion capture, lo- calization, and dense mapping system

    Bonan Liu, Handi Yin, Manuel Kaufmann, Jinhao He, Sammy Christen, Jie Song, and Pan Hui. Egohdm: An online egocentric-inertial human motion capture, lo- calization, and dense mapping system. arXiv preprint arXiv:2409.00343, 2024. 2, 3

  43. [51]

    Egofish3d: Egocentric 3d pose es- timation from a fisheye camera via self-supervised learning

    Yuxuan Liu, Jianxin Yang, Xiao Gu, Yijun Chen, Yao Guo, and Guang-Zhong Yang. Egofish3d: Egocentric 3d pose es- timation from a fisheye camera via self-supervised learning. IEEE Transactions on Multimedia, 25:8880–8891, 2023. 2

  44. [52]

    Egohmr: Egocentric human mesh recovery via hierarchical latent diffusion model

    Yuxuan Liu, Jianxin Yang, Xiao Gu, Yao Guo, and Guang- Zhong Yang. Egohmr: Egocentric human mesh recovery via hierarchical latent diffusion model. In 2023 IEEE Inter- national Conference on Robotics and Automation (ICRA) , pages 9807–9813. IEEE, 2023. 2

  45. [53]

    Decoupled weight decay regularization

    Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017. 6

  46. [54]

    Nymeria: A massive collection of multimodal egocentric daily motion in the wild

    Lingni Ma, Yuting Ye, Fangzhou Hong, Vladimir Guzov, Yifeng Jiang, Rowan Postyeni, Luis Pesqueira, Alexander Gamino, Vijay Baiyya, Hyo Jin Kim, et al. Nymeria: A massive collection of multimodal egocentric daily motion in the wild. In European Conference on Computer Vision, pa...

  47. [55]

    Going deeper into first-person activity recognition

    Minghuang Ma, Haoqi Fan, and Kris M Kitani. Going deeper into first-person activity recognition. In Proceed- ings of the IEEE Conference on Computer Vision and Pat- tern Recognition, pages 1894–1903, 2016. 2

  48. [56]

    Troje, Gerard Pons-Moll, and Michael J

    Naureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll, and Michael J. Black. AMASS: Archive of motion capture as surface shapes. In International Con- ference on Computer Vision , pages 5442–5451, 2019. 5, 6

  49. [57]

    Generating continual human motion in diverse 3d scenes

    Aymen Mir, Xavier Puig, Angjoo Kanazawa, and Gerard Pons-Moll. Generating continual human motion in diverse 3d scenes. In 2024 International Conference on 3D Vision (3DV), pages 903–913. IEEE, 2024. 3

  50. [58]

    Imuposer: Full-body pose estima- tion using imus in phones, watches, and earbuds

    Vimal Mollyn, Riku Arakawa, Mayank Goel, Chris Harri- son, and Karan Ahuja. Imuposer: Full-body pose estima- tion using imus in phones, watches, and earbuds. In Pro- ceedings of the 2023 CHI Conference on Human Factors in Computing Systems, pages 1–12, 2023. 2

  51. [59]

    Joint reconstruction of 3d human and object via contact-based refinement transformer

    Hyeongjin Nam, Daniel Sungho Jung, Gyeongsik Moon, and Kyoung Mu Lee. Joint reconstruction of 3d human and object via contact-based refinement transformer. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10218–10227, 2024. 3

  52. [60]

    Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3D hands, face, and body from a single image. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pa...

  53. [61]

    Scalable diffusion mod- els with transformers

    William Peebles and Saining Xie. Scalable diffusion mod- els with transformers. In Proceedings of the IEEE/CVF international conference on computer vision , pages 4195– 4205, 2023. 5, 7

  54. [62]

    Hoi-diff: Text-driven syn- thesis of 3d human-object interactions using diffusion mod- els

    Xiaogang Peng, Yiming Xie, Zizhao Wu, Varun Jampani, Deqing Sun, and Huaizu Jiang. Hoi-diff: Text-driven syn- thesis of 3d human-object interactions using diffusion mod- els. arXiv preprint arXiv:2312.06553, 2023. 3

  55. [63]

    Object pop-up: Can we infer 3d objects and their poses from human interactions alone? In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023

    Ilya A Petrov, Riccardo Marin, Julian Chibane, and Gerard Pons-Moll. Object pop-up: Can we infer 3d objects and their poses from human interactions alone? In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023. 3, 7

  56. [64]

    Tridi: Trilateral diffusion of 3d humans, objects, and interactions

    Ilya A Petrov, Riccardo Marin, Julian Chibane, and Ger- ard Pons-Moll. Tridi: Trilateral diffusion of 3d humans, objects, and interactions. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025. 3, 4

  57. [65]

    Project Aria , accessed January 7, 2025

    Project Aria. Project Aria , accessed January 7, 2025. https://www.projectaria.com/. 2

  58. [66]

    Project aria ma- chine perception services, accessed January 7, 2025

    Project Aria Machine Perception Services. Project aria ma- chine perception services, accessed January 7, 2025. 2

  59. [67]

    Pointnext: Revisiting pointnet++ with improved training and scaling strategies

    Guocheng Qian, Yuchen Li, Houwen Peng, Jinjie Mai, Hasan Hammoud, Mohamed Elhoseiny, and Bernard Ghanem. Pointnext: Revisiting pointnet++ with improved training and scaling strategies. Advances in neural infor- mation processing systems, 35:23192–23204, 2022. 4

  60. [68]

    Hierarchical text-conditional image generation with clip latents

    Aditya Ramesh, Prafulla Dhariwal, Alex Nichol, Casey Chu, and Mark Chen. Hierarchical text-conditional image generation with clip latents. arXiv preprint arXiv:2204.06125, 1(2):3, 2022. 4, 1

  61. [69]

    Egocap: egocentric marker-less motion capture with two fisheye cameras.ACM Transactions on Graphics (TOG), 35(6):1–11, 2016

    Helge Rhodin, Christian Richardt, Dan Casas, Eldar In- safutdinov, Mohammad Shafiei, Hans-Peter Seidel, Bernt Schiele, and Christian Theobalt. Egocap: egocentric marker-less motion capture with two fisheye cameras.ACM Transactions on Graphics (TOG), 35(6):1–11, 2016. 2 11

  62. [70]

    First-person pose recognition using egocentric workspaces

    Gr ´egory Rogez, James S Supancic, and Deva Ramanan. First-person pose recognition using egocentric workspaces. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4325–4333, 2015. 2

  63. [71]

    Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: Modeling and capturing hands and bod- ies together. ACM Transactions on Graphics, (Proc. SIG- GRAPH Asia), 36(6), 2017. 6

  64. [72]

    Deep unsupervised learning using nonequilibrium thermodynamics

    Jascha Sohl-Dickstein, Eric Weiss, Niru Maheswaranathan, and Surya Ganguli. Deep unsupervised learning using nonequilibrium thermodynamics. In International confer- ence on machine learning, pages 2256–2265. PMLR, 2015. 4

  65. [73]

    Hoianimator: Generating text-prompt human-object anima- tions using novel perceptive diffusion models

    Wenfeng Song, Xinyu Zhang, Shuai Li, Yang Gao, Aimin Hao, Xia Hou, Chenglizhao Chen, Ning Li, and Hong Qin. Hoianimator: Generating text-prompt human-object anima- tions using novel perceptive diffusion models. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and...

  66. [74]

    Neural localizer fields for continuous 3d human pose and shape estima- tion

    Istv ´an S ´ar´andi and Gerard Pons-Moll. Neural localizer fields for continuous 3d human pose and shape estima- tion. Advances in Neural Information Processing Systems (NeurIPS), 2024. 6

  67. [75]

    A unified diffusion framework for scene- aware human motion estimation from sparse signals

    Jiangnan Tang, Jingya Wang, Kaiyang Ji, Lan Xu, Jingyi Yu, and Ye Shi. A unified diffusion framework for scene- aware human motion estimation from sparse signals. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition , pages 21251–21262, 2024. 3

  68. [76]

    Human motion diffusion model

    Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H Bermano. Human motion diffusion model. arXiv preprint arXiv:2209.14916, 2022. 5

  69. [77]

    Selfpose: 3d egocentric pose estimation from a headset mounted camera.IEEE Transactions on Pat- tern Analysis and Machine Intelligence , 45(6):6794–6806,

    Denis Tome, Thiemo Alldieck, Patrick Peluse, Gerard Pons-Moll, Lourdes Agapito, Hernan Badino, and Fer- nando De la Torre. Selfpose: 3d egocentric pose estimation from a headset mounted camera.IEEE Transactions on Pat- tern Analysis and Machine Intelligence , 45(6):6794–6806,

  70. [78]

    Deco: Dense estimation of 3d human-scene contact in the wild

    Shashank Tripathi, Agniv Chatterjee, Jean-Claude Passy, Hongwei Yi, Dimitrios Tzionas, and Michael J Black. Deco: Dense estimation of 3d human-scene contact in the wild. In Proceedings of the IEEE/CVF International Con- ference on Computer Vision, pages 8001–8013, 2023. 3

  71. [79]

    Practical motion capture in everyday surround- ings

    Daniel Vlasic, Rolf Adelsberger, Giovanni Vannucci, John Barnwell, Markus Gross, Wojciech Matusik, and Jovan Popovi´c. Practical motion capture in everyday surround- ings. ACM transactions on graphics (TOG) , 26(3):35–es,

  72. [80]

    Sparse inertial poser: Automatic 3d hu- man pose estimation from sparse imus

    Timo V on Marcard, Bodo Rosenhahn, Michael J Black, and Gerard Pons-Moll. Sparse inertial poser: Automatic 3d hu- man pose estimation from sparse imus. InComputer graph- ics forum, pages 349–360. Wiley Online Library, 2017. 2

  73. [81]

    Egocentric whole-body motion capture with fisheyevit and diffusion-based motion refinement

    Jian Wang, Zhe Cao, Diogo Luvizon, Lingjie Liu, Kri- pasindhu Sarkar, Danhang Tang, Thabo Beeler, and Chris- tian Theobalt. Egocentric whole-body motion capture with fisheyevit and diffusion-based motion refinement. In Pro- ceedings of the IEEE/CVF Conference on Computer Visio...

  74. [82]

    Quest- sim: Human motion tracking from sparse sensors with sim- ulated avatars

    Alexander Winkler, Jungdam Won, and Yuting Ye. Quest- sim: Human motion tracking from sparse sensors with sim- ulated avatars. In SIGGRAPH Asia 2022 Conference Pa- pers, pages 1–8, 2022. 2

  75. [83]

    Chore: Contact, human and object reconstruction from a single rgb image

    Xianghui Xie, Bharat Lal Bhatnagar, and Gerard Pons- Moll. Chore: Contact, human and object reconstruction from a single rgb image. In European Conference on Com- puter Vision (ECCV). Springer, 2022. 3

  76. [84]

    Visibility aware human-object interaction tracking from single rgb camera

    Xianghui Xie, Bharat Lal Bhatnagar, and Gerard Pons- Moll. Visibility aware human-object interaction tracking from single rgb camera. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023

  77. [85]

    Template free reconstruction of human- object interaction with procedural interaction generation

    Xianghui Xie, Bharat Lal Bhatnagar, Jan Eric Lenssen, and Gerard Pons-Moll. Template free reconstruction of human- object interaction with procedural interaction generation. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 10003–10015, 2024

  78. [86]

    In- tertrack: Tracking human object interaction without object templates

    Xianghui Xie, Jan Eric Lenssen, and Gerard Pons-Moll. In- tertrack: Tracking human object interaction without object templates. In International Conference on 3D Vision 2025,

  79. [87]

    Regen- net: Towards human action-reaction synthesis

    Liang Xu, Yizhou Zhou, Yichao Yan, Xin Jin, Wenhan Zhu, Fengyun Rao, Xiaokang Yang, and Wenjun Zeng. Regen- net: Towards human action-reaction synthesis. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1759–1769, 2024. 3

  80. [88]

    Interdiff: Generating 3d human-object interactions with physics-informed diffusion

    Sirui Xu, Zhengyuan Li, Yu-Xiong Wang, and Liang-Yan Gui. Interdiff: Generating 3d human-object interactions with physics-informed diffusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 14928–14940, 2023. 3, 7

  81. [89]

    Inter- dreamer: Zero-shot text to 3d dynamic human-object in- teraction

    Sirui Xu, Yu-Xiong Wang, Liangyan Gui, et al. Inter- dreamer: Zero-shot text to 3d dynamic human-object in- teraction. Advances in Neural Information Processing Sys- tems, 37:52858–52890, 2024. 3

  82. [90]

    Mobileposer: Real-time full-body pose estimation and 3d human translation from imus in mobile consumer devices

    Vasco Xu, Chenfeng Gao, Henry Hoffmann, and Karan Ahuja. Mobileposer: Real-time full-body pose estimation and 3d human translation from imus in mobile consumer devices. In Proceedings of the 37th Annual ACM Sympo- sium on User Interface Software and Technology, pages 1– 11, 2024. 2

  83. [91]

    Mo 2Cap2 : Real-time mobile 3d motion capture with a cap-mounted fisheye camera

    Weipeng Xu, Avishek Chatterjee, Michael Zollhoefer, Helge Rhodin, Pascal Fua, Hans-Peter Seidel, and Christian Theobalt. Mo 2Cap2 : Real-time mobile 3d motion capture with a cap-mounted fisheye camera. IEEE Transactions on Visualization and Computer Graphics, pages 1–1, 2019. 2

  84. [92]

    F-hoi: Toward fine-grained semantic- aligned 3d human-object interactions

    Jie Yang, Xuesong Niu, Nan Jiang, Ruimao Zhang, and Siyuan Huang. F-hoi: Toward fine-grained semantic- aligned 3d human-object interactions. In European Con- ference on Computer Vision, pages 91–110. Springer, 2024. 3

  85. [93]

    Egochoir: Capturing 3d human-object interaction regions from egocentric views

    Yuhang Yang, Wei Zhai, Chengfeng Wang, Chengjun Yu, Yang Cao, and Zheng-Jun Zha. Egochoir: Capturing 3d human-object interaction regions from egocentric views. Advances in Neural Information Processing Systems , 37: 54529–54557, 2024. 2

  86. [94]

    Estimating body and hand motion in an ego- sensed world

    Brent Yi, Vickie Ye, Maya Zheng, Yunqi Li, Lea M ¨uller, Georgios Pavlakos, Yi Ma, Jitendra Malik, and Angjoo 12 Kanazawa. Estimating body and hand motion in an ego- sensed world. arXiv preprint arXiv:2410.03665, 2024. 2, 3, 5, 6, 7

  87. [95]

    Transpose: Real- time 3d human translation and pose estimation with six in- ertial sensors

    Xinyu Yi, Yuxiao Zhou, and Feng Xu. Transpose: Real- time 3d human translation and pose estimation with six in- ertial sensors. ACM Transactions On Graphics (TOG), 40 (4):1–13, 2021. 2

  88. [96]

    Physical inertial poser (pip): Physics-aware real-time human motion tracking from sparse inertial sensors

    Xinyu Yi, Yuxiao Zhou, Marc Habermann, Soshi Shi- mada, Vladislav Golyanik, Christian Theobalt, and Feng Xu. Physical inertial poser (pip): Physics-aware real-time human motion tracking from sparse inertial sensors. InPro- ceedings of the IEEE/CVF conference on computer vision...

  89. [97]

    Yonemoto, K

    H. Yonemoto, K. Murasaki, T. Osawa, K. Sudo, J. Shima- mura, and Y . Taniguchi. Egocentric articulated pose track- ing for action recognition. In International Conference on Machine Vision Applications (MVA), 2015. 2

  90. [98]

    Neuraldome: A neural modeling pipeline on multi- view human-object interactions

    Juze Zhang, Haimin Luo, Hongdi Yang, Xinru Xu, Qianyang Wu, Ye Shi, Jingyi Yu, Lan Xu, and Jingya Wang. Neuraldome: A neural modeling pipeline on multi- view human-object interactions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages ...

  91. [99]

    Hoi-mˆ 3: Capture multiple humans and objects in- teraction within contextual environment

    Juze Zhang, Jingyan Zhang, Zining Song, Zhanhe Shi, Chengfeng Zhao, Ye Shi, Jingyi Yu, Lan Xu, and Jingya Wang. Hoi-mˆ 3: Capture multiple humans and objects in- teraction within contextual environment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern R...

  92. [100]

    Egobody: Human body shape and motion of interacting people from head-mounted devices

    Siwei Zhang, Qianli Ma, Yan Zhang, Zhiyin Qian, Taein Kwon, Marc Pollefeys, Federica Bogo, and Siyu Tang. Egobody: Human body shape and motion of interacting people from head-mounted devices. In European confer- ence on computer vision , pages 180–200. Springer, 2022. 2

  93. [101]

    Force: Physics-aware human-object interaction

    Xiaohan Zhang, Bharat Lal Bhatnagar, Sebastian Starke, Ilya Petrov, Vladimir Guzov, Helisa Dhamo, Ed- uardo P ´erez-Pellitero, and Gerard Pons-Moll. Force: Physics-aware human-object interaction. arXiv preprint arXiv:2403.11237, 2024. 3

  94. [102]

    Scenic: Scene-aware semantic navi- gation with instruction-guided control

    Xiaohan Zhang, Sebastian Starke, Vladimir Guzov, Zhensong Zhang, Eduardo P ´erez Pellitero, and Ger- ard Pons-Moll. Scenic: Scene-aware semantic navi- gation with instruction-guided control. arXiv preprint arXiv:2412.15664, 2024. 3

  95. [103]

    Instance tracking in 3d scenes from egocentric videos

    Yunhan Zhao, Haoyu Ma, Shu Kong, and Charless Fowlkes. Instance tracking in 3d scenes from egocentric videos. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition , pages 21933–21944, 2024. 8

  96. [104]

    Realistic full-body tracking from sparse ob- servations via joint-level modeling

    Xiaozheng Zheng, Zhuo Su, Chao Wen, Zhou Xue, and Xiaojie Jin. Realistic full-body tracking from sparse ob- servations via joint-level modeling. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 14678–14688, 2023. 2, 7

  97. [105]

    On the continuity of rotation representations in neu- ral networks

    Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. On the continuity of rotation representations in neu- ral networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 5745– 5753, 2019. 4

  98. [106]

    Loose inertial poser: Motion capture with imu-attached loose-wear jacket

    Chengxu Zuo, Yiming Wang, Lishuang Zhan, Shihui Guo, Xinyu Yi, Feng Xu, and Yipeng Qin. Loose inertial poser: Motion capture with imu-attached loose-wear jacket. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 2209–2219, 2024. 2 13...

  99. [107]

    This research direction could enable new applications for study- ing human behavior and developing realistic virtual expe- riences

    Broader impacts The ability of our model to capture and generate contin- uous human-object interactions offers significant value for fields such as digital content creation and ergonomics. This research direction could enable new applications for study- ing human behavior and ...

  100. [108]

    The forward diffusion process can be for- mulated as a Markov chain with T steps

    Background and Notation Background. The forward diffusion process can be for- mulated as a Markov chain with T steps. Starting from a clean sample z0 it produces a series of distributions q(zt|zt−1): q(z1:T|z0) = QT t=1q(zt|zt−1). We add noise to the distribution for T steps, ...

  101. [109]

    Fol- lowing the evaluation of ECHO with sparseH orO track- ing, we test the model’s performance withI data provided as an additional conditioning

    Additional evaluation Evaluating the model with contacts conditioning. Fol- lowing the evaluation of ECHO with sparseH orO track- ing, we test the model’s performance withI data provided as an additional conditioning. We report the results in Ta- ble S2. Providing contact info...

  102. [110]

    Losses and Metrics Losses. The objective function used to train our network is the weighted combination of the following losses: LH n =∥θH−bθH∥2 +∥TH−bTH∥2 LO n =∥TO−bTO∥2 LI n =∥cI−bcI∥2 LO v =∥VO− bVO∥2 LH j =∥JH−bJH∥2 LH s =∥cfeet I ∗ bU feet H ∥2 (12) where bU feet H is th...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.