Pith. sign in

REVIEW 4 major objections 2 cited by

Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions

T0 review · 4 major / 0 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read The paper's central claim is that phase poles causing black holes in NLINV MRI reconstruction can be detected in the smooth coil sensitivity maps and corrected mid-iteration by a phase-vortex multiplication, yielding pole-free images and co

desk verdict The submission is two papers in one envelope: the body is an MRI method paper, the abstract is a dataset paper, and as received there is nothing to referee. read the letter →

arxiv 2508.04681 v1 pith:F4MOSKZ3 submitted 2025-08-06 cs.CV

classification cs.CV PACS 87.61.-c
keywords MRIreconstructionphasesingularitycoilsensitivityestimationNLINVparallelimagingreal-timenon-linearinverseproblemswindingnumber
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Complex-valued MR images and coil sensitivity maps are only determined up to a common phase function, and regularized nonlinear inversion (NLINV) can get stuck in local minima where the image contains a phase pole — a point around which the phase winds, with a conjugate pole hidden in the coil maps — producing a black hole. This paper argues that such poles should be detected in the coil sensitivities rather than the image, because coil maps are smooth while image phase is noisy, and that a signal-weighted vote across coils isolates the artificial pole from true coil singularities. Detected poles are corrected by multiplying the image by a unit phase vortex and the coil sensitivities by its inverse, after which the remaining Gauss-Newton iterations refine the result. The demonstration on accelerated brain MPRAGE and interactive radial real-time cardiac MRI shows that pole-free images and singularity-free coil estimates can be obtained from auto-calibration regions as small as 7×7, without a body-coil reference. This matters because phase poles corrupt coil estimates and cause artifacts in compressed-sensing and real-time reconstruction.

What carries the argument

The central objects are phase poles — points where the complex phase is undefined and winds around the point, with amplitude zero — and the unit phase vortex $\vartheta_{r_0}^{\pm}(r)=\frac{(r-r_0)_x \pm i(r-r_0)_y}{\|r-r_0\|}=e^{\pm i\phi(r-r_0)}$. The carried mechanism is detection via a signal-weighted average of winding numbers computed in the smooth coil sensitivity maps, followed by correction that multiplies the image by one vortex and the coils by the conjugate vortex, integrated inside the iteratively regularized Gauss-Newton iterations of NLINV.

What would settle it

On the same real-time radial FLASH setup, record a long session with many interactive slice changes and check every corrected frame for a persistent black hole with a nonzero weighted winding number at that location; any such frame would refute the claim of reliable phase-pole removal. Alternatively, take a fully sampled brain dataset and initialize NLINV with a deliberately planted conjugate vortex in the coil maps; if the corrected reconstruction does not converge to the same image and coils as a body-coil-referenced initialization, the escape-from-local-minimum claim is false.

Watch

Extended reading notes

Core claim

The discovery is a recipe for escaping the phase-pole local minima of NLINV. At the end of an intermediate Gauss-Newton step, the algorithm computes, for each coil sensitivity map $c_i$, the winding number of the phase around a small circle centered at each pixel; the squared magnitude $|c_i|^2$ weights each coil's vote, and pixels where the weighted winding number exceeds a threshold are marked as poles. The image is then multiplied by the phase vortex $\vartheta_{r_0}^{\pm}$ that removes the pole, and the preconditioned coil sensitivities are multiplied by the inverse vortex before the next iteration. Because NLINV's iteratively regularized Gauss-Newton scheme continues after the correctio

Load-bearing premise

The method stands on the assumption that every artificial phase pole causes the same winding pattern in most of the coil sensitivity maps at the same location, while true coil singularities are rare enough that a signal-weighted vote can tell them apart; if that agreement is absent, correction will miss the pole or remove a real one.

Editorial extensions

If this is right

  • NLINV with phase-pole correction yields singularity-free coil sensitivity profiles from auto-calibration regions as small as $7\times7$, where ESPIRiT still shows aliasing artifacts.
  • Interactive real-time MRI stays real-time: the correction adds less than 10% runtime and the full reconstruction remains faster than the acquisition, so slice changes that trigger phase poles no longer corrupt the live images.
  • The body-coil reference is no longer needed to initialize NLINV away from pole local minima, removing a constraint in applications where no body coil is available.
  • Because detection is done in the coil maps rather than the image, the same idea can be transplanted to other auto-calibrated coil estimation methods such as ESPIRiT, though without NLINV's iterative refinement the post-correction phase may not be perfectly smooth.
  • The method corrects multiple poles simultaneously by multiplying the corresponding vortices, so reconstructions are not limited to single-singularity cases.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The abstract and body text of this manuscript are different works: the body never mentions the InterVLA dataset, egocentric video, or human-object-human interaction, so the abstract's dataset claim has no support in the text that follows it.
  • A natural stress test would sweep coil count and signal-to-noise ratio: the detection's vote-based logic predicts that reliability degrades when true coil singularities coincide with artificial poles or when only one or two coils contribute significant signal.
  • The paper's stated 2D limitation suggests a direct extension: in 3D, phase poles become vortex lines, so the point-detection algorithm would need to be replaced by a line-tracking version to handle vortex lines lying parallel to the imaging slices.
  • The demonstrated transfer to ESPIRiT suggests a general principle — detect artificial poles where coils agree, not where the image is dark — which could be built into other auto-calibrated reconstruction pipelines as a post-processing cleanup step.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 0 minor

Summary. The abstract announces InterVLA, a claimed first large-scale egocentric human-object-human interaction dataset with 11.4 hours and 1.2M frames of multimodal data, plus benchmarks for egocentric human motion estimation, interaction synthesis, and interaction prediction. The full text of the submission, however, is a completely different paper titled 'Phase-Pole-Free Images and Smooth Coil Sensitivity Maps by Regularized Nonlinear Inversion,' an MRI reconstruction methodology paper by different authors under arXiv:2508.04685v2 [physics.med-ph]. The body contains no description of InterVLA, no data collection or annotation protocol, no benchmark definitions, no experimental results, and no discussion of egocentric interactions. Treating all manuscript text as in-scope evidence, the submission as received does not support any of the abstract's central claims.

Significance. If the InterVLA dataset and benchmarks were real and fully documented, the resource could be valuable for research on egocentric perception, human-object-human interaction modeling, and vision-language-action assistants. The claimed scale (11.4 hours, 1.2M frames, multi-view, MoCap-synchronized data) would be a useful contribution, and the three proposed tasks address relevant open problems. However, none of this is present in the submitted manuscript. The only verifiable content is an unrelated MRI reconstruction paper. Therefore the significance of the contribution cannot be assessed from this submission; the claims are unverifiable and the resource, as presented, does not exist in the document.

major comments (4)
  1. [Full text (title page; §1–§4)] The body of the submission is not the InterVLA paper announced in the abstract. The full text is titled 'Phase-Pole-Free Images and Smooth Coil Sensitivity Maps by Regularized Nonlinear Inversion,' lists different authors, carries arXiv:2508.04685v2 [physics.med-ph], and describes an MRI reconstruction method with no connection to egocentric interactions. There is no section, equation, table, or figure in the body that supports the dataset or benchmark claims. Treating all manuscript text as in-scope evidence, the body refutes rather than supports the assertion that this is the InterVLA paper.
  2. [Abstract] The load-bearing quantitative claims — 11.4 hours, 1.2M frames, 2 egocentric + 5 exocentric views, accurate human/object motions, and verbal commands — are stated with no supporting methodology. The manuscript provides no acquisition protocol, sensor calibration, synchronization procedure, annotation pipeline, or accuracy metric. A dataset/benchmark contribution cannot be evaluated, reproduced, or used when the data description and validation are entirely absent.
  3. [Abstract (GPT-generated scripts)] The premise that behavior produced by 'pairs of assistants and instructors engag[ing] with multiple objects and the scene following GPT-generated scripts' transfers to real manual assistance is asserted without evidence. No naturalness evaluation, task-coverage analysis, or comparison to unscripted assistance is provided. This is a structural validity concern for any downstream benchmark built on the dataset, and it is unaddressed in the submission.
  4. [Abstract (benchmarks)] The three named benchmarks — egocentric human motion estimation, interaction synthesis, and interaction prediction — are not defined. There are no task formulations, evaluation metrics, data splits, baseline methods, or results. The phrase 'with comprehensive analysis' cannot be checked because the analysis does not appear in the manuscript. These are central components of the claimed contribution, not minor omissions.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found; the abstract claims stand without a supporting derivation, and the full text is unrelated to them.

full rationale

The received manuscript presents an abstract claiming creation of the InterVLA egocentric human-object-human interaction dataset (11.4 hours, 1.2M frames, RGB-MoCap hybrid capture, GPT-generated scripts) but the only full text supplied is an unrelated MRI reconstruction paper, 'Phase-Pole-Free Images and Smooth Coil Sensitivity Maps by Regularized Nonlinear Inversion' (arXiv:2508.04685v2), with a different title, author list, and field. There is thus no actual derivation chain in the manuscript connecting the abstract's claims to methods, benchmark definitions, or evaluation results. A circularity finding requires exhibiting a specific reduction: a defined term that is equivalent to the target result, a fitted parameter renamed as a prediction, a load-bearing self-citation chain, an ansatz smuggled via citation, or a known result merely renamed. None of these can be identified because the supporting technical content for InterVLA is absent. The abstract's unsupported factual assertions are a completeness, reproducibility, and integrity problem, not a circularity problem. The MRI body text, taken on its own, is a self-contained method-and-evaluation paper compared against ESPIRiT and other baselines; even if it were the intended paper, no circular step is evident in its phase-pole detection and correction scheme. Accordingly, per the hard rule that circularity must be demonstrated by quote and specific reduction rather than inferred from absence or mismatch, the honest verdict is no significant circularity, score 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The abstract exposes three domain assumptions the dataset's value depends on (egocentric necessity, GPT-script validity, MoCap accuracy) and no free parameters, because all benchmark hyperparameters and any fitted distribution choices would live in the missing methods sections. No new physical or conceptual entities beyond the dataset artifact itself are introduced.

assumptions (3)
  • domain assumption Egocentric, first-person acquisition is indispensable for an AI assistant to perceive and act in the world.
    The abstract's motivation ('we urge that both the generalist interaction knowledge and egocentric modality are indispensable') asserts this necessity, but the supplied material provides no comparative evidence that first-person data is required for the claimed downstream tasks.
  • domain assumption Behavior generated from GPT scripts by assistant-instructor pairs is ecologically valid training signal for manual assistance.
    Abstract: 'pairs of assistants and instructors engage with multiple objects and the scene following GPT-generated scripts.' The transfer of scripted interactions to unscripted real-world assistance is assumed without evidence on naturalness or task coverage.
  • domain assumption The hybrid RGB-MoCap system yields accurate human and object motion ground truth.
    Abstract claims 'accurate human/object motions'; no calibration, error metrics, or validation are visible in the supplied material to support the accuracy claim.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions." pith.science (2026). https://pith.science/paper/F4MOSKZ3

@misc{pith2026250804681,
  author       = {Pith},
  title        = {Pith review of: Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/F4MOSKZ3}},
  note         = {Machine review of arXiv:2508.04681}
}
read the original abstract

Learning action models from real-world human-centric interaction datasets is important towards building general-purpose intelligent assistants with efficiency. However, most existing datasets only offer specialist interaction category and ignore that AI assistants perceive and act based on first-person acquisition. We urge that both the generalist interaction knowledge and egocentric modality are indispensable. In this paper, we embed the manual-assisted task into a vision-language-action framework, where the assistant provides services to the instructor following egocentric vision and commands. With our hybrid RGB-MoCap system, pairs of assistants and instructors engage with multiple objects and the scene following GPT-generated scripts. Under this setting, we accomplish InterVLA, the first large-scale human-object-human interaction dataset with 11.4 hours and 1.2M frames of multimodal data, spanning 2 egocentric and 5 exocentric videos, accurate human/object motions and verbal commands. Furthermore, we establish novel benchmarks on egocentric human motion estimation, interaction synthesis, and interaction prediction with comprehensive analysis. We believe that our InterVLA testbed and the benchmarks will foster future works on building AI agents in the physical world.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial Fusion

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    IMU-HOI jointly recovers human pose and object motion from body and object IMUs by inferring contacts to guide a three-stage fusion of kinematic and inertial data for coherent trajectories.

  2. The N-Body Problem: Parallel Execution from Single-Person Egocentric Video

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Structured prompting for a VLM, fed with ground-truth spatial zone schedules, can generate more feasible multi-agent parallel executions from single-person egocentric videos.

Reference graph

Works this paper leans on

150 extracted references · 51 canonical work pages · cited by 2 Pith papers

  1. [1]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...

  2. [2]

    https://eth-ait.github.io/aitviewer/

    Aitviewer. https://eth-ait.github.io/aitviewer/

  3. [3]

    https://github.com/zju3dv/EasyMocap

    Easymocap - make human motion capture easier. https://github.com/zju3dv/EasyMocap

  4. [4]

    https://github.com/opencv/opencv/blob/master/data/haarcascades/haarcascade_frontalface_default.xml

    Opencv: opencv. https://github.com/opencv/opencv/blob/master/data/haarcascades/haarcascade_frontalface_default.xml

  5. [5]

    3d human pose perception from egocentric stereo videos

    Hiroyasu Akada, Jian Wang, Vladislav Golyanik, and Christian Theobalt. 3d human pose perception from egocentric stereo videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 767--776, 2024

  6. [6]

    Rgbmanip: Monocular image-based robotic manipulation through active object pose estimation

    Boshi An, Yiran Geng, Kai Chen, Xiaoqi Li, Qi Dou, and Hao Dong. Rgbmanip: Monocular image-based robotic manipulation through active object pose estimation. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 7748--7755. IEEE, 2024

  7. [7]

    Circle: Capture in rich contextual environments

    Joao Pedro Ara \'u jo, Jiaman Li, Karthik Vetrivel, Rishi Agarwal, Jiajun Wu, Deepak Gopinath, Alexander William Clegg, and Karen Liu. Circle: Capture in rich contextual environments. In CVPR, pages 21211--21221, 2023

  8. [8]

    Introducing hot3d: An egocentric dataset for 3d hand and object tracking

    Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon, Shreyas Hampali, Fan Zhang, Jade Fountain, Edward Miller, Selen Basol, Richard Newcombe, Robert Wang, et al. Introducing hot3d: An egocentric dataset for 3d hand and object tracking. arXiv preprint arXiv:2406.09598, 2024

Show all 150 references
  1. [9]

    Uncertainty-aware state space transformer for egocentric 3d hand trajectory forecasting

    Wentao Bao, Lele Chen, Libing Zeng, Zhong Li, Yi Xu, Junsong Yuan, and Yu Kong. Uncertainty-aware state space transformer for egocentric 3d hand trajectory forecasting. In Proceedings of the IEEE/CVF international conference on computer vision, pages 13702--13711, 2023

  2. [10]

    Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation, 2024

    Homanga Bharadhwaj, Debidatta Dwibedi, Abhinav Gupta, Shubham Tulsiani, Carl Doersch, Ted Xiao, Dhruv Shah, Fei Xia, Dorsa Sadigh, and Sean Kirmani. Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation, 2024

  3. [11]

    Behave: Dataset and method for tracking human object interactions

    Bharat Lal Bhatnagar, Xianghui Xie, Ilya A Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. Behave: Dataset and method for tracking human object interactions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15...

  4. [12]

    Bedlam: A synthetic dataset of bodies exhibiting detailed lifelike animated motion

    Michael J Black, Priyanka Patel, Joachim Tesch, and Jinlong Yang. Bedlam: A synthetic dataset of bodies exhibiting detailed lifelike animated motion. In CVPR, pages 8726--8737, 2023

  5. [13]

    Contactpose: A dataset of grasps with object contact and hand pose

    Samarth Brahmbhatt, Chengcheng Tang, Christopher D Twigg, Charles C Kemp, and James Hays. Contactpose: A dataset of grasps with object contact and hand pose. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part XIII 16, p...

  6. [14]

    Rt-1: Robotics transformer for real-world control at scale

    Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al. Rt-1: Robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817, 2022

  7. [15]

    Rt-2: Vision-language-action models transfer web knowledge to robotic control

    Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818, 2023

  8. [16]

    u tepage, Ali Ghadirzadeh, \

    Judith B \"u tepage, Ali Ghadirzadeh, \"O zge \"O ztimur Karadaǧ, M rten Bj \"o rkman, and Danica Kragic. Imitating by generating: Deep generative models for imitation of interactive tasks. Frontiers in Robotics and AI, 7: 0 47, 2020

  9. [17]

    Understanding hand-object manipulation with grasp types and object attributes

    Minjie Cai, Kris M Kitani, and Yoichi Sato. Understanding hand-object manipulation with grasp types and object attributes. In Robotics: Science and Systems, 2016

  10. [18]

    Long-term human motion prediction with scene context

    Zhe Cao, Hang Gao, Karttikeya Mangalam, Qi-Zhi Cai, Minh Vo, and Jitendra Malik. Long-term human motion prediction with scene context. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part I 16, pages 387--404. Springer, 2020

  11. [19]

    A multi-sensor dataset of human-human handover

    Alessandro Carf \` , Francesco Foglino, Barbara Bruno, and Fulvio Mastrogiovanni. A multi-sensor dataset of human-human handover. Data in brief, 22: 0 109--117, 2019

  12. [20]

    On the choice of grasp type and location when handing over an object

    Francesca Cini, V Ortenzi, P Corke, and MJSR Controzzi. On the choice of grasp type and location when handing over an object. Science Robotics, 4 0 (27): 0 eaau9757, 2019

  13. [21]

    Context-aware human motion prediction, 2020

    Enric Corona, Albert Pumarola, Guillem Alenyà, and Francesc Moreno-Noguer. Context-aware human motion prediction, 2020

  14. [22]

    Towards collaborative robots as intelligent co-workers in human-robot joint tasks: what to do and who does it? In ISR 2020; 52th International Symposium on Robotics, pages 1--8

    Ana Cunha, Flora Ferreira, Emanuel Sousa, Luis Louro, Paulo Vicente, Sergio Monteiro, Wolfram Erlhagen, and Estela Bicho. Towards collaborative robots as intelligent co-workers in human-robot joint tasks: what to do and who does it? In ISR 2020; 52th International Symposium on...

  15. [23]

    You-do, i-learn: Egocentric unsupervised discovery of objects and their modes of interaction towards video-based guidance

    Dima Damen, Teesid Leelasawassuk, and Walterio Mayol-Cuevas. You-do, i-learn: Egocentric unsupervised discovery of objects and their modes of interaction towards video-based guidance. Computer Vision and Image Understanding, 149: 0 98--112, 2016

  16. [24]

    Scaling egocentric vision: The epic-kitchens dataset

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Scaling egocentric vision: The epic-kitchens dataset. In Proceedings of the European conference on comput...

  17. [25]

    Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100

    Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. International Journal ...

  18. [26]

    Cg-hoi: Contact-guided 3d human-object interaction generation

    Christian Diller and Angela Dai. Cg-hoi: Contact-guided 3d human-object interaction generation. arXiv preprint arXiv:2311.16097, 2023

  19. [27]

    Palm-e: An embodied multimodal language model

    Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023

  20. [28]

    Avatars grow legs: Generating smooth human motion from sparse tracking inputs with diffusion model

    Yuming Du, Robin Kips, Albert Pumarola, Sebastian Starke, Ali Thabet, and Artsiom Sanakoyeu. Avatars grow legs: Generating smooth human motion from sparse tracking inputs with diffusion model. In CVPR, 2023

  21. [29]

    Human preferences for robot eye gaze in human-to-robot handovers

    Tair Faibish, Alap Kshirsagar, Guy Hoffman, and Yael Edan. Human preferences for robot eye gaze in human-to-robot handovers. International Journal of Social Robotics, 14 0 (4): 0 995--1012, 2022

  22. [30]

    Go to zero: Towards zero-shot motion generation with million-scale data

    Ke Fan, Shunlin Lu, Minyue Dai, Runyi Yu, Lixing Xiao, Zhiyang Dou, Junting Dong, Lizhuang Ma, and Jingbo Wang. Go to zero: Towards zero-shot motion generation with million-scale data. arXiv preprint arXiv:2507.07095, 2025 a

  23. [31]

    Arctic: A dataset for dexterous bimanual hand-object manipulation

    Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J Black, and Otmar Hilliges. Arctic: A dataset for dexterous bimanual hand-object manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12...

  24. [32]

    Hold: Category-agnostic 3d reconstruction of interacting hands and objects from video

    Zicong Fan, Maria Parelli, Maria Eleni Kadoglou, Xu Chen, Muhammed Kocabas, Michael J Black, and Otmar Hilliges. Hold: Category-agnostic 3d reconstruction of interacting hands and objects from video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...

  25. [33]

    Benchmarks and challenges in pose estimation for egocentric hand interactions with objects

    Zicong Fan, Takehiko Ohkawa, Linlin Yang, Nie Lin, Zhishan Zhou, Shihao Zhou, Jiajun Liang, Zhong Gao, Xuanyang Zhang, Xue Zhang, et al. Benchmarks and challenges in pose estimation for egocentric hand interactions with objects. In European Conference on Computer Vision, pages...

  26. [34]

    Social interactions: A first-person perspective

    Alircza Fathi, Jessica K Hodgins, and James M Rehg. Social interactions: A first-person perspective. In 2012 IEEE Conference on Computer Vision and Pattern Recognition, pages 1226--1233. IEEE, 2012

  27. [35]

    Three-dimensional reconstruction of human interactions

    Mihai Fieraru, Mihai Zanfir, Elisabeta Oneata, Alin-Ionut Popa, Vlad Olaru, and Cristian Sminchisescu. Three-dimensional reconstruction of human interactions. In CVPR, pages 7214--7223, 2020

  28. [36]

    Ego4d: Around the world in 3,000 hours of egocentric video

    Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF Conference on Computer Vision ...

  29. [37]

    Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

    Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. In Proceedings...

  30. [38]

    Generating diverse and natural 3d human motions from text

    Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. Generating diverse and natural 3d human motions from text. In CVPR, pages 5152--5161, 2022 a

  31. [39]

    Multi-person extreme motion prediction

    Wen Guo, Xiaoyu Bie, Xavier Alameda-Pineda, and Francesc Moreno-Noguer. Multi-person extreme motion prediction. In CVPR, pages 13053--13064, 2022 b

  32. [40]

    Interaction replica: Tracking human--object interaction and scene changes from human motion

    Vladimir Guzov, Julian Chibane, Riccardo Marin, Yannan He, Yunus Saracoglu, Torsten Sattler, and Gerard Pons-Moll. Interaction replica: Tracking human--object interaction and scene changes from human motion. In 2024 International Conference on 3D Vision (3DV), pages 1006--1016...

  33. [41]

    Honnotate: A method for 3d annotation of hand and object poses

    Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. Honnotate: A method for 3d annotation of hand and object poses. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3196--3206, 2020

  34. [42]

    Keypoint transformer: Solving joint identification in challenging hands and object interactions for accurate 3d pose estimation

    Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit. Keypoint transformer: Solving joint identification in challenging hands and object interactions for accurate 3d pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit...

  35. [43]

    Ego3dt: Tracking every 3d object in ego-centric videos

    Shengyu Hao, Wenhao Chai, Zhonghan Zhao, Meiqi Sun, Wendi Hu, Jieyang Zhou, Yixian Zhao, Qi Li, Yizhou Wang, Xi Li, et al. Ego3dt: Tracking every 3d object in ego-centric videos. arXiv preprint arXiv:2410.08530, 2024

  36. [44]

    Resolving 3d human pose ambiguities with 3d scene constraints

    Mohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, and Michael J Black. Resolving 3d human pose ambiguities with 3d scene constraints. In ICCV, pages 2282--2292, 2019

  37. [45]

    Stochastic scene-aware motion prediction

    Mohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito, Jimei Yang, Yi Zhou, and Michael J Black. Stochastic scene-aware motion prediction. In ICCV, pages 11374--11384, 2021

  38. [46]

    Omnih2o: Universal and dexterous human-to-humanoid whole-body teleoperation and learning

    Tairan He, Zhengyi Luo, Xialin He, Wenli Xiao, Chong Zhang, Weinan Zhang, Kris Kitani, Changliu Liu, and Guanya Shi. Omnih2o: Universal and dexterous human-to-humanoid whole-body teleoperation and learning. arXiv preprint arXiv:2406.08858, 2024 a

  39. [47]

    Learning human-to-humanoid real-time whole-body teleoperation, 2024 b

    Tairan He, Zhengyi Luo, Wenli Xiao, Chong Zhang, Kris Kitani, Changliu Liu, and Guanya Shi. Learning human-to-humanoid real-time whole-body teleoperation, 2024 b

  40. [48]

    Gans trained by a two time-scale update rule converge to a local nash equilibrium

    Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017

  41. [49]

    Voxposer: Composable 3d value maps for robotic manipulation with language models

    Wenlong Huang, Chen Wang, Ruohan Zhang, Yunzhu Li, Jiajun Wu, and Li Fei-Fei. Voxposer: Composable 3d value maps for robotic manipulation with language models. arXiv preprint arXiv:2307.05973, 2023

  42. [50]

    Intercap: Joint markerless 3d tracking of humans and objects in interaction

    Yinghao Huang, Omid Taheri, Michael J Black, and Dimitrios Tzionas. Intercap: Joint markerless 3d tracking of humans and objects in interaction. In DAGM German Conference on Pattern Recognition, pages 281--299. Springer, 2022

  43. [51]

    Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments

    Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. IEEE transactions on pattern analysis and machine intelligence, 36 0 (7): 0 1325--1339, 2013

  44. [52]

    A large-scale rgb-d database for arbitrary-view human action recognition

    Yanli Ji, Feixiang Xu, Yang Yang, Fumin Shen, Heng Tao Shen, and Wei-Shi Zheng. A large-scale rgb-d database for arbitrary-view human action recognition. In ACMMM, page 1510–1518, 2018

  45. [53]

    Affordpose: A large-scale dataset of hand-object interactions with affordance-driven hand pose

    Juntao Jian, Xiuping Liu, Manyi Li, Ruizhen Hu, and Jian Liu. Affordpose: A large-scale dataset of hand-object interactions with affordance-driven hand pose. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14713--14724, 2023

  46. [54]

    Avatarposer: Articulated full-body pose tracking from sparse motion sensing

    Jiaxi Jiang, Paul Streli, Huajian Qiu, Andreas Fender, Larissa Laich, Patrick Snape, and Christian Holz. Avatarposer: Articulated full-body pose tracking from sparse motion sensing. In ECCV, pages 443--460. Springer, 2022

  47. [55]

    Egoposer: Robust real-time ego-body pose estimation in large scenes

    Jiaxi Jiang, Paul Streli, Manuel Meier, and Christian Holz. Egoposer: Robust real-time ego-body pose estimation in large scenes. arXiv preprint arXiv:2308.06493, 2023 a

  48. [56]

    Full-body articulated human-object interaction

    Nan Jiang, Tengyu Liu, Zhexuan Cao, Jieming Cui, Zhiyuan Zhang, Yixin Chen, He Wang, Yixin Zhu, and Siyuan Huang. Full-body articulated human-object interaction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9365--9376, 2023 b

  49. [57]

    Scaling up dynamic human-scene interaction modeling

    Nan Jiang, Zhiyuan Zhang, Hongjie Li, Xiaoxuan Ma, Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, and Siyuan Huang. Scaling up dynamic human-scene interaction modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1737--1747, 2024 a

  50. [58]

    A probabilistic attention model with occlusion-aware texture regression for 3d hand reconstruction from a single rgb image

    Zheheng Jiang, Hossein Rahmani, Sue Black, and Bryan M Williams. A probabilistic attention model with occlusion-aware texture regression for 3d hand reconstruction from a single rgb image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pa...

  51. [59]

    Harmon: Whole-body motion generation of humanoid robots from language descriptions

    Zhenyu Jiang, Yuqi Xie, Jinhan Li, Ye Yuan, Yifeng Zhu, and Yuke Zhu. Harmon: Whole-body motion generation of humanoid robots from language descriptions. arXiv preprint arXiv:2410.12773, 2024 b

  52. [60]

    Epic-fusion: Audio-visual temporal binding for egocentric action recognition

    Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen. Epic-fusion: Audio-visual temporal binding for egocentric action recognition. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5492--5501, 2019

  53. [61]

    Real-time active vision for a humanoid soccer robot using deep reinforcement learning

    Soheil Khatibi, Meisam Teimouri, and Mahdi Rezaei. Real-time active vision for a humanoid soccer robot using deep reinforcement learning. arXiv preprint arXiv:2011.13851, 2020

  54. [62]

    Openvla: An open-source vision-language-action model

    Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024

  55. [63]

    Dataset of bimanual human-to-human object handovers

    Alap Kshirsagar, Raphael Fortuna, Zhiming Xie, and Guy Hoffman. Dataset of bimanual human-to-human object handovers. Data in Brief, 48: 0 109277, 2023

  56. [64]

    H2o: Two hands manipulating objects for first person interaction recognition

    Taein Kwon, Bugra Tekin, Jan St \"u hmer, Federica Bogo, and Marc Pollefeys. H2o: Two hands manipulating objects for first person interaction recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10138--10148, 2021

  57. [65]

    Ego-body pose estimation via ego-head pose estimation

    Jiaman Li, Karen Liu, and Jiajun Wu. Ego-body pose estimation via ego-head pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17142--17151, 2023

  58. [66]

    Dngaussian: Optimizing sparse-view 3d gaussian radiance fields with global-local depth normalization

    Jiahe Li, Jiawei Zhang, Xiao Bai, Jin Zheng, Xin Ning, Jun Zhou, and Lin Gu. Dngaussian: Optimizing sparse-view 3d gaussian radiance fields with global-local depth normalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20775-...

  59. [67]

    In the eye of beholder: Joint learning of gaze and actions in first person video

    Yin Li, Miao Liu, and James M Rehg. In the eye of beholder: Joint learning of gaze and actions in first person video. In Proceedings of the European conference on computer vision (ECCV), pages 619--635, 2018

  60. [68]

    Ego-exo: Transferring visual representations from third-person to first-person videos

    Yanghao Li, Tushar Nagarajan, Bo Xiong, and Kristen Grauman. Ego-exo: Transferring visual representations from third-person to first-person videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6943--6953, 2021

  61. [69]

    Intergen: Diffusion-based multi-human motion generation under complex interactions

    Han Liang, Wenqian Zhang, Wenxuan Li, Jingyi Yu, and Lan Xu. Intergen: Diffusion-based multi-human motion generation under complex interactions. arXiv preprint arXiv:2304.05684, 2023

  62. [70]

    Motion-x: A large-scale 3d expressive whole-body human motion dataset

    Jing Lin, Ailing Zeng, Shunlin Lu, Yuanhao Cai, Ruimao Zhang, Haoqian Wang, and Lei Zhang. Motion-x: A large-scale 3d expressive whole-body human motion dataset. arXiv preprint arXiv:2307.00818, 2023

  63. [71]

    Gaussian-flow: 4d reconstruction with dynamic 3d gaussian particle

    Youtian Lin, Zuozhuo Dai, Siyu Zhu, and Yao Yao. Gaussian-flow: 4d reconstruction with dynamic 3d gaussian particle. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21136--21145, 2024

  64. [72]

    Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding

    Jun Liu, Amir Shahroudy, Mauricio Perez, Gang Wang, Ling-Yu Duan, and Alex C Kot. Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding. T-PAMI, 42 0 (10): 0 2684--2701, 2019

  65. [73]

    Forecasting human-object interaction: joint prediction of motor attention and actions in first person video

    Miao Liu, Siyu Tang, Yin Li, and James M Rehg. Forecasting human-object interaction: joint prediction of motor attention and actions in first person video. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part I 16, pages ...

  66. [74]

    Joint hand motion and interaction hotspots prediction from egocentric videos

    Shaowei Liu, Subarna Tripathi, Somdeb Majumdar, and Xiaolong Wang. Joint hand motion and interaction hotspots prediction from egocentric videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3282--3292, 2022 a

  67. [75]

    Hoi4d: A 4d egocentric dataset for category-level human-object interaction

    Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. Hoi4d: A 4d egocentric dataset for category-level human-object interaction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pa...

  68. [76]

    Taco: Benchmarking generalizable bimanual tool-action-object understanding

    Yun Liu, Haolin Yang, Xu Si, Ling Liu, Zipeng Li, Yuxiang Zhang, Yebin Liu, and Li Yi. Taco: Benchmarking generalizable bimanual tool-action-object understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21740--21751, 2024

  69. [77]

    Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons - Moll, and Michael J. Black. SMPL: a skinned multi-person linear model. ACM Trans. Graph. , 34 0 (6): 0 248:1--248:16, 2015

  70. [78]

    Dynamics-regulated kinematic policy for egocentric pose estimation

    Zhengyi Luo, Ryo Hachiuma, Ye Yuan, and Kris Kitani. Dynamics-regulated kinematic policy for egocentric pose estimation. Advances in Neural Information Processing Systems, 34: 0 25019--25032, 2021

  71. [79]

    Himo: A new benchmark for full-body human interacting with multiple objects

    Xintao Lv, Liang Xu, Yichao Yan, Xin Jin, Congsheng Xu, Shuwen Wu, Yifan Liu, Lincheng Li, Mengxiao Bi, Wenjun Zeng, et al. Himo: A new benchmark for full-body human interacting with multiple objects. In European Conference on Computer Vision, pages 300--318. Springer, 2025

  72. [80]

    Diff-ip2d: Diffusion-based hand-object interaction prediction on egocentric videos

    Junyi Ma, Jingyi Xu, Xieyuanli Chen, and Hesheng Wang. Diff-ip2d: Diffusion-based hand-object interaction prediction on egocentric videos. arXiv preprint arXiv:2405.04370, 2024 a

  73. [81]

    A survey on vision-language-action models for embodied ai

    Yueen Ma, Zixing Song, Yuzheng Zhuang, Jianye Hao, and Irwin King. A survey on vision-language-action models for embodied ai. arXiv preprint arXiv:2405.14093, 2024 b

  74. [82]

    Unifying representations and large-scale whole-body motion databases for studying human motion

    Christian Mandery, \"O mer Terlemez, Martin Do, Nikolaus Vahrenkamp, and Tamim Asfour. Unifying representations and large-scale whole-body motion databases for studying human motion. IEEE Transactions on Robotics, 32 0 (4): 0 796--809, 2016

  75. [83]

    Dexvip: Learning dexterous grasping with human hand pose priors from video

    Priyanka Mandikal and Kristen Grauman. Dexvip: Learning dexterous grasping with human hand pose priors from video. In Conference on Robot Learning, pages 651--661. PMLR, 2022

  76. [84]

    Hoi4abot: Human-object interaction anticipation for human intention reading collaborative robots, 2024

    Esteve Valls Mascaro, Daniel Sliwowski, and Dongheui Lee. Hoi4abot: Human-object interaction anticipation for human intention reading collaborative robots, 2024

  77. [85]

    Eventego3d: 3d human motion capture from egocentric event streams

    Christen Millerdurai, Hiroyasu Akada, Jian Wang, Diogo Luvizon, Christian Theobalt, and Vladislav Golyanik. Eventego3d: 3d human motion capture from egocentric event streams. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1186--1195, 2024

  78. [86]

    imapper: interaction-guided scene mapping from monocular videos

    Aron Monszpart, Paul Guerrero, Duygu Ceylan, Ersin Yumer, and Niloy J Mitra. imapper: interaction-guided scene mapping from monocular videos. ACM Transactions On Graphics (TOG), 38 0 (4): 0 1--15, 2019

  79. [87]

    Grounded human-object interaction hotspots from video

    Tushar Nagarajan, Christoph Feichtenhofer, and Kristen Grauman. Grounded human-object interaction hotspots from video. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8688--8697, 2019

  80. [88]

    Jointly learning energy expenditures and activities using egocentric multimodal signals

    Katsuyuki Nakamura, Serena Yeung, Alexandre Alahi, and Li Fei-Fei. Jointly learning energy expenditures and activities using egocentric multimodal signals. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1868--1877, 2017

  81. [89]

    You2me: Inferring body pose in egocentric video via first and second person interactions

    Evonne Ng, Donglai Xiang, Hanbyul Joo, and Kristen Grauman. You2me: Inferring body pose in egocentric video via first and second person interactions. In CVPR, pages 9890--9900, 2020

  82. [90]

    Open x-embodiment: Robotic learning datasets and rt-x models

    Abby O'Neill, Abdul Rehman, Abhinav Gupta, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, et al. Open x-embodiment: Robotic learning datasets and rt-x models. arXiv preprint arXiv:2310.08864, 2023

  83. [91]

    GPT -3.5 turbo fine-tuning and api updates

    OpenAI. GPT -3.5 turbo fine-tuning and api updates. https://openai.com/index/gpt-3-5-turbo-fine-tuning-and-api-updates/, 2023

  84. [92]

    Handoccnet: Occlusion-robust 3d hand mesh estimation network

    JoonKyu Park, Yeonguk Oh, Gyeongsik Moon, Hongsuk Choi, and Kyoung Mu Lee. Handoccnet: Occlusion-robust 3d hand mesh estimation network. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1496--1505, 2022

  85. [93]

    Reconstructing hands in 3 D with transformers

    Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Reconstructing hands in 3 D with transformers. In CVPR, 2024

  86. [94]

    Hoi-diff: Text-driven synthesis of 3d human-object interactions using diffusion models

    Xiaogang Peng, Yiming Xie, Zizhao Wu, Varun Jampani, Deqing Sun, and Huaizu Jiang. Hoi-diff: Text-driven synthesis of 3d human-object interactions using diffusion models. arXiv preprint arXiv:2312.06553, 2023

  87. [95]

    Action-conditioned 3d human motion synthesis with transformer vae

    Mathis Petrovich, Michael J Black, and G \"u l Varol. Action-conditioned 3d human motion synthesis with transformer vae. In CVPR, pages 10985--10995, 2021

  88. [96]

    The kit motion-language dataset

    Matthias Plappert, Christian Mandery, and Tamim Asfour. The kit motion-language dataset. Big data, 4 0 (4): 0 236--252, 2016

  89. [97]

    Wilor: End-to-end 3d hand localization and reconstruction in-the-wild

    Rolandos Alexandros Potamias, Jinglei Zhang, Jiankang Deng, and Stefanos Zafeiriou. Wilor: End-to-end 3d hand localization and reconstruction in-the-wild. arXiv preprint arXiv:2409.12259, 2024

  90. [98]

    Mild: multimodal interactive latent dynamics for learning human-robot interaction

    Vignesh Prasad, Dorothea Koert, Ruth Stock-Homburg, Jan Peters, and Georgia Chalvatzaki. Mild: multimodal interactive latent dynamics for learning human-robot interaction. In 2022 IEEE-RAS 21st International Conference on Humanoid Robots (Humanoids), pages 472--479. IEEE, 2022

  91. [99]

    Moveint: Mixture of variational experts for learning human-robot interactions from demonstrations

    Vignesh Prasad, Alap Kshirsagar, Dorothea Koert Ruth Stock-Homburg, Jan Peters, and Georgia Chalvatzaki. Moveint: Mixture of variational experts for learning human-robot interactions from demonstrations. IEEE Robotics and Automation Letters, 2024

  92. [100]

    The virtual caliper: Rapid creation of metrically accurate avatars from 3d measurements

    Sergi Pujades, Betty Mohler, Anne Thaler, Joachim Tesch, Naureen Mahmood, Nikolas Hesse, Heinrich H B \"u lthoff, and Michael J Black. The virtual caliper: Rapid creation of metrically accurate avatars from 3d measurements. IEEE transactions on visualization and computer graph...

  93. [101]

    D-nerf: Neural radiance fields for dynamic scenes

    Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-nerf: Neural radiance fields for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10318--10327, 2021

  94. [102]

    Babel: bodies, action and behavior with english labels

    Abhinanda R Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez, and Michael J Black. Babel: bodies, action and behavior with english labels. In CVPR, pages 722--731, 2021

  95. [103]

    Dexmv: Imitation learning for dexterous manipulation from human videos

    Yuzhe Qin, Yueh-Hua Wu, Shaowei Liu, Hanwen Jiang, Ruihan Yang, Yang Fu, and Xiaolong Wang. Dexmv: Imitation learning for dexterous manipulation from human videos. In European Conference on Computer Vision, pages 570--587. Springer, 2022

  96. [104]

    Imitation learning based on bilateral control for human--robot cooperation

    Ayumu Sasagawa, Kazuki Fujimoto, Sho Sakaino, and Toshiaki Tsuji. Imitation learning based on bilateral control for human--robot cooperation. IEEE Robotics and Automation Letters, 5 0 (4): 0 6169--6176, 2020

  97. [105]

    Pigraphs: learning interaction snapshots from observations

    Manolis Savva, Angel X Chang, Pat Hanrahan, Matthew Fisher, and Matthias Nie ner. Pigraphs: learning interaction snapshots from observations. ACM Transactions On Graphics (TOG), 35 0 (4): 0 1--12, 2016

  98. [106]

    Deep imitation learning for humanoid loco-manipulation through human teleoperation, 2023

    Mingyo Seo, Steve Han, Kyutae Sim, Seung Hyeon Bang, Carlos Gonzalez, Luis Sentis, and Yuke Zhu. Deep imitation learning for humanoid loco-manipulation through human teleoperation, 2023

  99. [107]

    Human motion diffusion as a generative prior

    Yonatan Shafir, Guy Tevet, Roy Kapon, and Amit H Bermano. Human motion diffusion as a generative prior. arXiv preprint arXiv:2303.01418, 2023

  100. [108]

    Modeling ambient scene dynamics for free-view synthesis

    Meng-Li Shih, Jia-Bin Huang, Changil Kim, Rajvi Shah, Johannes Kopf, and Chen Gao. Modeling ambient scene dynamics for free-view synthesis. In ACM SIGGRAPH 2024 Conference Papers, pages 1--11, 2024

  101. [109]

    Wham: Reconstructing world-grounded humans with accurate 3d motion

    Soyong Shin, Juyong Kim, Eni Halilaj, and Michael J Black. Wham: Reconstructing world-grounded humans with accurate 3d motion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2070--2080, 2024

  102. [110]

    Trace: 5d temporal regression of avatars with dynamic cameras in 3d environments

    Yu Sun, Qian Bao, Wu Liu, Tao Mei, and Michael J Black. Trace: 5d temporal regression of avatars with dynamic cameras in 3d environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8856--8866, 2023

  103. [111]

    Grab: A dataset of whole-body human grasping of objects

    Omid Taheri, Nima Ghorbani, Michael J Black, and Dimitrios Tzionas. Grab: A dataset of whole-body human grasping of objects. In ECCV, pages 581--600, 2020

  104. [112]

    Human motion diffusion model

    Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H Bermano. Human motion diffusion model. arXiv preprint arXiv:2209.14916, 2022

  105. [113]

    Learning robot soccer from egocentric vision with deep reinforcement learning

    Dhruva Tirumala, Markus Wulfmeier, Ben Moran, Sandy Huang, Jan Humplik, Guy Lever, Tuomas Haarnoja, Leonard Hasenclever, Arunkumar Byravan, Nathan Batchelor, et al. Learning robot soccer from egocentric vision with deep reinforcement learning. arXiv preprint arXiv:2405.02425, 2024

  106. [114]

    Umpm benchmark: A multi-person dataset with synchronized video and motion capture data for evaluation of articulated human motion and interaction

    NP Van der Aa, Xinghan Luo, Geert-Jan Giezeman, Robby T Tan, and Remco C Veltkamp. Umpm benchmark: A multi-person dataset with synchronized video and motion capture data for evaluation of articulated human motion and interaction. In ICCV Workshops, pages 1264--1269. IEEE, 2011

  107. [115]

    Sparsenerf: Distilling depth ranking for few-shot novel view synthesis

    Guangcong Wang, Zhaoxi Chen, Chen Change Loy, and Ziwei Liu. Sparsenerf: Distilling depth ranking for few-shot novel view synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9065--9076, 2023 a

  108. [116]

    Scene-aware egocentric 3d human pose estimation

    Jian Wang, Diogo Luvizon, Weipeng Xu, Lingjie Liu, Kripasindhu Sarkar, and Christian Theobalt. Scene-aware egocentric 3d human pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13031--13040, 2023 b

  109. [117]

    Egocentric whole-body motion capture with fisheyevit and diffusion-based motion refinement

    Jian Wang, Zhe Cao, Diogo Luvizon, Lingjie Liu, Kripasindhu Sarkar, Danhang Tang, Thabo Beeler, and Christian Theobalt. Egocentric whole-body motion capture with fisheyevit and diffusion-based motion refinement. In Proceedings of the IEEE/CVF Conference on Computer Vision and ...

  110. [118]

    Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world

    Xin Wang, Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bugra Tekin, Felipe Vieira Frujeri, et al. Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world. In Proceedings of the IEEE/...

  111. [119]

    Quo vadis, motion generation? from large language models to large motion models, 2024 b

    Ye Wang, Sipeng Zheng, Bin Cao, Qianshan Wei, Qin Jin, and Zongqing Lu. Quo vadis, motion generation? from large language models to large motion models, 2024 b

  112. [120]

    Tram: Global trajectory and motion of 3d humans from in-the-wild videos

    Yufu Wang, Ziyun Wang, Lingjie Liu, and Kostas Daniilidis. Tram: Global trajectory and motion of 3d humans from in-the-wild videos. In European Conference on Computer Vision, pages 467--487. Springer, 2025

  113. [121]

    Genh2r: Learning generalizable human-to-robot handover via scalable simulation demonstration and imitation

    Zifan Wang, Junyu Chen, Ziqing Chen, Pengwei Xie, Rui Chen, and Li Yi. Genh2r: Learning generalizable human-to-robot handover via scalable simulation demonstration and imitation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16362-...

  114. [122]

    Hoh: Markerless multimodal human-object-human handover dataset with large object count

    Noah Wiederhold, Ava Megyeri, DiMaggio Paris, Sean Banerjee, and Natasha Banerjee. Hoh: Markerless multimodal human-object-human handover dataset with large object count. Advances in Neural Information Processing Systems, 36, 2024

  115. [123]

    4d gaussian splatting for real-time dynamic scene rendering

    Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 4d gaussian splatting for real-time dynamic scene rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20310--20320, 2024

  116. [124]

    Unified human-scene interaction via prompted chain-of-contacts

    Zeqi Xiao, Tai Wang, Jingbo Wang, Jinkun Cao, Wenwei Zhang, Bo Dai, Dahua Lin, and Jiangmiao Pang. Unified human-scene interaction via prompted chain-of-contacts. arXiv preprint arXiv:2309.07918, 2023

  117. [125]

    Sparsegs: Real-time 360 \ deg \ sparse view synthesis using gaussian splatting

    Haolin Xiong, Sairisheek Muttukuru, Rishi Upadhyay, Pradyumna Chari, and Achuta Kadambi. Sparsegs: Real-time 360 \ deg \ sparse view synthesis using gaussian splatting. arXiv preprint arXiv:2312.00206, 2023

  118. [126]

    Actformer: A gan-based transformer towards general action-conditioned 3d human motion generation

    Liang Xu, Ziyang Song, Dongliang Wang, Jing Su, Zhicheng Fang, Chenjing Ding, Weihao Gan, Yichao Yan, Xin Jin, Xiaokang Yang, et al. Actformer: A gan-based transformer towards general action-conditioned 3d human motion generation. In ICCV, pages 2228--2238, 2023 a

  119. [127]

    Motionbank: A large-scale video motion benchmark with disentangled rule-based annotations, 2024 a

    Liang Xu, Shaoyang Hua, Zili Lin, Yifan Liu, Feipeng Ma, Yichao Yan, Xin Jin, Xiaokang Yang, and Wenjun Zeng. Motionbank: A large-scale video motion benchmark with disentangled rule-based annotations, 2024 a

  120. [128]

    Inter-x: Towards versatile human-human interaction analysis

    Liang Xu, Xintao Lv, Yichao Yan, Xin Jin, Shuwen Wu, Congsheng Xu, Yifan Liu, Yizhou Zhou, Fengyun Rao, Xingdong Sheng, et al. Inter-x: Towards versatile human-human interaction analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pag...

  121. [129]

    Regennet: Towards human action-reaction synthesis

    Liang Xu, Yizhou Zhou, Yichao Yan, Xin Jin, Wenhan Zhu, Fengyun Rao, Xiaokang Yang, and Wenjun Zeng. Regennet: Towards human action-reaction synthesis. In CVPR, pages 1759--1769, 2024 c

  122. [130]

    InterDiff : Generating 3d human-object interactions with physics-informed diffusion

    Sirui Xu, Zhengyuan Li, Yu-Xiong Wang, and Liang-Yan Gui. InterDiff : Generating 3d human-object interactions with physics-informed diffusion. In ICCV, 2023 b

  123. [131]

    D3d-hoi: Dynamic 3d human-object interactions from videos

    Xiang Xu, Hanbyul Joo, Greg Mori, and Manolis Savva. D3d-hoi: Dynamic 3d human-object interactions from videos. arXiv preprint arXiv:2108.08420, 2021

  124. [132]

    Humanvla: Towards vision-language directed object rearrangement by physical humanoid

    Xinyu Xu, Yizheng Zhang, Yong-Lu Li, Lei Han, and Cewu Lu. Humanvla: Towards vision-language directed object rearrangement by physical humanoid. arXiv preprint arXiv:2406.19972, 2024 d

  125. [133]

    What's in your hands? 3d reconstruction of generic objects in hands

    Yufei Ye, Abhinav Gupta, and Shubham Tulsiani. What's in your hands? 3d reconstruction of generic objects in hands. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3895--3905, 2022

  126. [134]

    Diffusion-guided reconstruction of everyday hand-object interaction clips

    Yufei Ye, Poorvi Hebbar, Abhinav Gupta, and Shubham Tulsiani. Diffusion-guided reconstruction of everyday hand-object interaction clips. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 19717--19728, 2023

  127. [135]

    Estimating body and hand motion in an ego-sensed world, 2024

    Brent Yi, Vickie Ye, Maya Zheng, Lea Müller, Georgios Pavlakos, Yi Ma, Jitendra Malik, and Angjoo Kanazawa. Estimating body and hand motion in an ego-sensed world, 2024

  128. [136]

    Hi4d: 4d instance segmentation of close human interaction

    Yifei Yin, Chen Guo, Manuel Kaufmann, Juan Jose Zarate, Jie Song, and Otmar Hilliges. Hi4d: 4d instance segmentation of close human interaction. In CVPR, pages 17016--17027, 2023

  129. [137]

    Lm-gaussian: Boost sparse-view 3d gaussian splatting with large model priors

    Hanyang Yu, Xiaoxiao Long, and Ping Tan. Lm-gaussian: Boost sparse-view 3d gaussian splatting with large model priors. arXiv preprint arXiv:2409.03456, 2024

  130. [138]

    Acr: Attention collaboration-based regressor for arbitrary two-hand reconstruction

    Zhengdi Yu, Shaoli Huang, Chen Fang, Toby P Breckon, and Jue Wang. Acr: Attention collaboration-based regressor for arbitrary two-hand reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12955--12964, 2023

  131. [139]

    Glamr: Global occlusion-aware human mesh recovery with dynamic cameras

    Ye Yuan, Umar Iqbal, Pavlo Molchanov, Kris Kitani, and Jan Kautz. Glamr: Global occlusion-aware human mesh recovery with dynamic cameras. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11038--11049, 2022

  132. [140]

    Two-person interaction detection using body-pose features and multiple instance learning

    Kiwon Yun, Jean Honorio, Debaleena Chattopadhyay, Tamara L Berg, and Dimitris Samaras. Two-person interaction detection using body-pose features and multiple instance learning. In CVPR, pages 28--35. IEEE, 2012

  133. [141]

    Core4d: A 4d human-object-human interaction dataset for collaborative object rearrangement

    Chengwen Zhang, Yun Liu, Ruofan Xing, Bingda Tang, and Li Yi. Core4d: A 4d human-object-human interaction dataset for collaborative object rearrangement. arXiv preprint arXiv:2406.19353, 2024 a

  134. [142]

    Transplat: Generalizable 3d gaussian splatting from sparse multi-view images with transformers

    Chuanrui Zhang, Yingshuang Zou, Zhuoling Li, Minmin Yi, and Haoqian Wang. Transplat: Generalizable 3d gaussian splatting from sparse multi-view images with transformers. arXiv preprint arXiv:2408.13770, 2024 b

  135. [143]

    Enhancing human-centered dynamic scene understanding via multiple llms collaborated reasoning

    Hang Zhang, Wenxiao Zhang, Haoxuan Qu, and Jun Liu. Enhancing human-centered dynamic scene understanding via multiple llms collaborated reasoning. Visual Intelligence, 3 0 (1): 0 3, 2025

  136. [144]

    Hoi-m\^ 3: Capture multiple humans and objects interaction within contextual environment

    Juze Zhang, Jingyan Zhang, Zining Song, Zhanhe Shi, Chengfeng Zhao, Ye Shi, Jingyi Yu, Lan Xu, and Jingya Wang. Hoi-m\^ 3: Capture multiple humans and objects interaction within contextual environment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re...

  137. [145]

    Couch: towards controllable human-chair interactions

    Xiaohan Zhang, Bharat Lal Bhatnagar, Sebastian Starke, Vladimir Guzov, and Gerard Pons-Moll. Couch: towards controllable human-chair interactions. In ECCV, pages 518--535, 2022

  138. [146]

    Scenic: Scene-aware semantic navigation with instruction-guided control

    Xiaohan Zhang, Sebastian Starke, Vladimir Guzov, Zhensong Zhang, Eduardo P \'e rez Pellitero, and Gerard Pons-Moll. Scenic: Scene-aware semantic navigation with instruction-guided control. arXiv preprint arXiv:2412.15664, 2024 d

  139. [147]

    3d-vla: A 3d vision-language-action generative world model

    Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, and Chuang Gan. 3d-vla: A 3d vision-language-action generative world model. arXiv preprint arXiv:2403.09631, 2024

  140. [148]

    Temporal perception and prediction in ego-centric video

    Yipin Zhou and Tamara L Berg. Temporal perception and prediction in ego-centric video. In Proceedings of the IEEE International Conference on Computer Vision, pages 4498--4506, 2015

  141. [149]

    On the continuity of rotation representations in neural networks

    Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. On the continuity of rotation representations in neural networks. In CVPR , pages 5745--5753. Computer Vision Foundation / IEEE , 2019

  142. [150]

    3d human shape reconstruction from a polarization image

    Shihao Zou, Xinxin Zuo, Yiming Qian, Sen Wang, Chi Xu, Minglun Gong, and Li Cheng. 3d human shape reconstruction from a polarization image. In ECCV, pages 351--368, 2020

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.