REVIEW 4 major objections 2 cited by
Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
T0 review · 4 major / 0 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read The paper's central claim is that phase poles causing black holes in NLINV MRI reconstruction can be detected in the smooth coil sensitivity maps and corrected mid-iteration by a phase-vortex multiplication, yielding pole-free images and co
desk verdict The submission is two papers in one envelope: the body is an MRI method paper, the abstract is a dataset paper, and as received there is nothing to referee. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central objects are phase poles — points where the complex phase is undefined and winds around the point, with amplitude zero — and the unit phase vortex $\vartheta_{r_0}^{\pm}(r)=\frac{(r-r_0)_x \pm i(r-r_0)_y}{\|r-r_0\|}=e^{\pm i\phi(r-r_0)}$. The carried mechanism is detection via a signal-weighted average of winding numbers computed in the smooth coil sensitivity maps, followed by correction that multiplies the image by one vortex and the coils by the conjugate vortex, integrated inside the iteratively regularized Gauss-Newton iterations of NLINV.
What would settle it
On the same real-time radial FLASH setup, record a long session with many interactive slice changes and check every corrected frame for a persistent black hole with a nonzero weighted winding number at that location; any such frame would refute the claim of reliable phase-pole removal. Alternatively, take a fully sampled brain dataset and initialize NLINV with a deliberately planted conjugate vortex in the coil maps; if the corrected reconstruction does not converge to the same image and coils as a body-coil-referenced initialization, the escape-from-local-minimum claim is false.
Extended reading notes
Core claim
The discovery is a recipe for escaping the phase-pole local minima of NLINV. At the end of an intermediate Gauss-Newton step, the algorithm computes, for each coil sensitivity map $c_i$, the winding number of the phase around a small circle centered at each pixel; the squared magnitude $|c_i|^2$ weights each coil's vote, and pixels where the weighted winding number exceeds a threshold are marked as poles. The image is then multiplied by the phase vortex $\vartheta_{r_0}^{\pm}$ that removes the pole, and the preconditioned coil sensitivities are multiplied by the inverse vortex before the next iteration. Because NLINV's iteratively regularized Gauss-Newton scheme continues after the correctio
Load-bearing premise
The method stands on the assumption that every artificial phase pole causes the same winding pattern in most of the coil sensitivity maps at the same location, while true coil singularities are rare enough that a signal-weighted vote can tell them apart; if that agreement is absent, correction will miss the pole or remove a real one.
Editorial extensions
If this is right
- NLINV with phase-pole correction yields singularity-free coil sensitivity profiles from auto-calibration regions as small as $7\times7$, where ESPIRiT still shows aliasing artifacts.
- Interactive real-time MRI stays real-time: the correction adds less than 10% runtime and the full reconstruction remains faster than the acquisition, so slice changes that trigger phase poles no longer corrupt the live images.
- The body-coil reference is no longer needed to initialize NLINV away from pole local minima, removing a constraint in applications where no body coil is available.
- Because detection is done in the coil maps rather than the image, the same idea can be transplanted to other auto-calibrated coil estimation methods such as ESPIRiT, though without NLINV's iterative refinement the post-correction phase may not be perfectly smooth.
- The method corrects multiple poles simultaneously by multiplying the corresponding vortices, so reconstructions are not limited to single-singularity cases.
Reading between the lines
- The abstract and body text of this manuscript are different works: the body never mentions the InterVLA dataset, egocentric video, or human-object-human interaction, so the abstract's dataset claim has no support in the text that follows it.
- A natural stress test would sweep coil count and signal-to-noise ratio: the detection's vote-based logic predicts that reliability degrades when true coil singularities coincide with artificial poles or when only one or two coils contribute significant signal.
- The paper's stated 2D limitation suggests a direct extension: in 3D, phase poles become vortex lines, so the point-detection algorithm would need to be replaced by a line-tracking version to handle vortex lines lying parallel to the imaging slices.
- The demonstrated transfer to ESPIRiT suggests a general principle — detect artificial poles where coils agree, not where the image is dark — which could be built into other auto-calibrated reconstruction pipelines as a post-processing cleanup step.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The abstract announces InterVLA, a claimed first large-scale egocentric human-object-human interaction dataset with 11.4 hours and 1.2M frames of multimodal data, plus benchmarks for egocentric human motion estimation, interaction synthesis, and interaction prediction. The full text of the submission, however, is a completely different paper titled 'Phase-Pole-Free Images and Smooth Coil Sensitivity Maps by Regularized Nonlinear Inversion,' an MRI reconstruction methodology paper by different authors under arXiv:2508.04685v2 [physics.med-ph]. The body contains no description of InterVLA, no data collection or annotation protocol, no benchmark definitions, no experimental results, and no discussion of egocentric interactions. Treating all manuscript text as in-scope evidence, the submission as received does not support any of the abstract's central claims.
Significance. If the InterVLA dataset and benchmarks were real and fully documented, the resource could be valuable for research on egocentric perception, human-object-human interaction modeling, and vision-language-action assistants. The claimed scale (11.4 hours, 1.2M frames, multi-view, MoCap-synchronized data) would be a useful contribution, and the three proposed tasks address relevant open problems. However, none of this is present in the submitted manuscript. The only verifiable content is an unrelated MRI reconstruction paper. Therefore the significance of the contribution cannot be assessed from this submission; the claims are unverifiable and the resource, as presented, does not exist in the document.
major comments (4)
- [Full text (title page; §1–§4)] The body of the submission is not the InterVLA paper announced in the abstract. The full text is titled 'Phase-Pole-Free Images and Smooth Coil Sensitivity Maps by Regularized Nonlinear Inversion,' lists different authors, carries arXiv:2508.04685v2 [physics.med-ph], and describes an MRI reconstruction method with no connection to egocentric interactions. There is no section, equation, table, or figure in the body that supports the dataset or benchmark claims. Treating all manuscript text as in-scope evidence, the body refutes rather than supports the assertion that this is the InterVLA paper.
- [Abstract] The load-bearing quantitative claims — 11.4 hours, 1.2M frames, 2 egocentric + 5 exocentric views, accurate human/object motions, and verbal commands — are stated with no supporting methodology. The manuscript provides no acquisition protocol, sensor calibration, synchronization procedure, annotation pipeline, or accuracy metric. A dataset/benchmark contribution cannot be evaluated, reproduced, or used when the data description and validation are entirely absent.
- [Abstract (GPT-generated scripts)] The premise that behavior produced by 'pairs of assistants and instructors engag[ing] with multiple objects and the scene following GPT-generated scripts' transfers to real manual assistance is asserted without evidence. No naturalness evaluation, task-coverage analysis, or comparison to unscripted assistance is provided. This is a structural validity concern for any downstream benchmark built on the dataset, and it is unaddressed in the submission.
- [Abstract (benchmarks)] The three named benchmarks — egocentric human motion estimation, interaction synthesis, and interaction prediction — are not defined. There are no task formulations, evaluation metrics, data splits, baseline methods, or results. The phrase 'with comprehensive analysis' cannot be checked because the analysis does not appear in the manuscript. These are central components of the claimed contribution, not minor omissions.
Circularity Check
No circularity found; the abstract claims stand without a supporting derivation, and the full text is unrelated to them.
full rationale
The received manuscript presents an abstract claiming creation of the InterVLA egocentric human-object-human interaction dataset (11.4 hours, 1.2M frames, RGB-MoCap hybrid capture, GPT-generated scripts) but the only full text supplied is an unrelated MRI reconstruction paper, 'Phase-Pole-Free Images and Smooth Coil Sensitivity Maps by Regularized Nonlinear Inversion' (arXiv:2508.04685v2), with a different title, author list, and field. There is thus no actual derivation chain in the manuscript connecting the abstract's claims to methods, benchmark definitions, or evaluation results. A circularity finding requires exhibiting a specific reduction: a defined term that is equivalent to the target result, a fitted parameter renamed as a prediction, a load-bearing self-citation chain, an ansatz smuggled via citation, or a known result merely renamed. None of these can be identified because the supporting technical content for InterVLA is absent. The abstract's unsupported factual assertions are a completeness, reproducibility, and integrity problem, not a circularity problem. The MRI body text, taken on its own, is a self-contained method-and-evaluation paper compared against ESPIRiT and other baselines; even if it were the intended paper, no circular step is evident in its phase-pole detection and correction scheme. Accordingly, per the hard rule that circularity must be demonstrated by quote and specific reduction rather than inferred from absence or mismatch, the honest verdict is no significant circularity, score 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Egocentric, first-person acquisition is indispensable for an AI assistant to perceive and act in the world.
- domain assumption Behavior generated from GPT scripts by assistant-instructor pairs is ecologically valid training signal for manual assistance.
- domain assumption The hybrid RGB-MoCap system yields accurate human and object motion ground truth.
Cite this review
Pith. "Pith review of Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions." pith.science (2026). https://pith.science/paper/F4MOSKZ3
@misc{pith2026250804681,
author = {Pith},
title = {Pith review of: Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions},
year = {2026},
howpublished = {\url{https://pith.science/paper/F4MOSKZ3}},
note = {Machine review of arXiv:2508.04681}
}
read the original abstract
Learning action models from real-world human-centric interaction datasets is important towards building general-purpose intelligent assistants with efficiency. However, most existing datasets only offer specialist interaction category and ignore that AI assistants perceive and act based on first-person acquisition. We urge that both the generalist interaction knowledge and egocentric modality are indispensable. In this paper, we embed the manual-assisted task into a vision-language-action framework, where the assistant provides services to the instructor following egocentric vision and commands. With our hybrid RGB-MoCap system, pairs of assistants and instructors engage with multiple objects and the scene following GPT-generated scripts. Under this setting, we accomplish InterVLA, the first large-scale human-object-human interaction dataset with 11.4 hours and 1.2M frames of multimodal data, spanning 2 egocentric and 5 exocentric videos, accurate human/object motions and verbal commands. Furthermore, we establish novel benchmarks on egocentric human motion estimation, interaction synthesis, and interaction prediction with comprehensive analysis. We believe that our InterVLA testbed and the benchmarks will foster future works on building AI agents in the physical world.
Forward citations
Cited by 2 Pith papers
-
IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial Fusion
IMU-HOI jointly recovers human pose and object motion from body and object IMUs by inferring contacts to guide a three-stage fusion of kinematic and inertial data for coherent trajectories.
-
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
Structured prompting for a VLM, fed with ground-truth spatial zone schedules, can generate more feasible multi-agent parallel executions from single-person egocentric videos.
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION format.date year duplicate empty "emp...
-
[2]
https://eth-ait.github.io/aitviewer/
Aitviewer. https://eth-ait.github.io/aitviewer/
-
[3]
https://github.com/zju3dv/EasyMocap
Easymocap - make human motion capture easier. https://github.com/zju3dv/EasyMocap
-
[4]
https://github.com/opencv/opencv/blob/master/data/haarcascades/haarcascade_frontalface_default.xml
Opencv: opencv. https://github.com/opencv/opencv/blob/master/data/haarcascades/haarcascade_frontalface_default.xml
-
[5]
3d human pose perception from egocentric stereo videos
Hiroyasu Akada, Jian Wang, Vladislav Golyanik, and Christian Theobalt. 3d human pose perception from egocentric stereo videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 767--776, 2024
2024
-
[6]
Rgbmanip: Monocular image-based robotic manipulation through active object pose estimation
Boshi An, Yiran Geng, Kai Chen, Xiaoqi Li, Qi Dou, and Hao Dong. Rgbmanip: Monocular image-based robotic manipulation through active object pose estimation. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages 7748--7755. IEEE, 2024
2024
-
[7]
Circle: Capture in rich contextual environments
Joao Pedro Ara \'u jo, Jiaman Li, Karthik Vetrivel, Rishi Agarwal, Jiajun Wu, Deepak Gopinath, Alexander William Clegg, and Karen Liu. Circle: Capture in rich contextual environments. In CVPR, pages 21211--21221, 2023
2023
-
[8]
Introducing hot3d: An egocentric dataset for 3d hand and object tracking
Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon, Shreyas Hampali, Fan Zhang, Jade Fountain, Edward Miller, Selen Basol, Richard Newcombe, Robert Wang, et al. Introducing hot3d: An egocentric dataset for 3d hand and object tracking. arXiv preprint arXiv:2406.09598, 2024
arXiv 2024
Show all 150 references
-
[9]
Uncertainty-aware state space transformer for egocentric 3d hand trajectory forecasting
Wentao Bao, Lele Chen, Libing Zeng, Zhong Li, Yi Xu, Junsong Yuan, and Yu Kong. Uncertainty-aware state space transformer for egocentric 3d hand trajectory forecasting. In Proceedings of the IEEE/CVF international conference on computer vision, pages 13702--13711, 2023
2023
-
[10]
Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation, 2024
Homanga Bharadhwaj, Debidatta Dwibedi, Abhinav Gupta, Shubham Tulsiani, Carl Doersch, Ted Xiao, Dhruv Shah, Fei Xia, Dorsa Sadigh, and Sean Kirmani. Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation, 2024
2024
-
[11]
Behave: Dataset and method for tracking human object interactions
Bharat Lal Bhatnagar, Xianghui Xie, Ilya A Petrov, Cristian Sminchisescu, Christian Theobalt, and Gerard Pons-Moll. Behave: Dataset and method for tracking human object interactions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15...
2022
-
[12]
Bedlam: A synthetic dataset of bodies exhibiting detailed lifelike animated motion
Michael J Black, Priyanka Patel, Joachim Tesch, and Jinlong Yang. Bedlam: A synthetic dataset of bodies exhibiting detailed lifelike animated motion. In CVPR, pages 8726--8737, 2023
2023
-
[13]
Contactpose: A dataset of grasps with object contact and hand pose
Samarth Brahmbhatt, Chengcheng Tang, Christopher D Twigg, Charles C Kemp, and James Hays. Contactpose: A dataset of grasps with object contact and hand pose. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part XIII 16, p...
2020
-
[14]
Rt-1: Robotics transformer for real-world control at scale
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Joseph Dabis, Chelsea Finn, Keerthana Gopalakrishnan, Karol Hausman, Alex Herzog, Jasmine Hsu, et al. Rt-1: Robotics transformer for real-world control at scale. arXiv preprint arXiv:2212.06817, 2022
2022 arXiv
-
[15]
Rt-2: Vision-language-action models transfer web knowledge to robotic control
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, Xi Chen, Krzysztof Choromanski, Tianli Ding, Danny Driess, Avinava Dubey, Chelsea Finn, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. arXiv preprint arXiv:2307.15818, 2023
2023 arXiv
-
[16]
u tepage, Ali Ghadirzadeh, \
Judith B \"u tepage, Ali Ghadirzadeh, \"O zge \"O ztimur Karadaǧ, M rten Bj \"o rkman, and Danica Kragic. Imitating by generating: Deep generative models for imitation of interactive tasks. Frontiers in Robotics and AI, 7: 0 47, 2020
2020
-
[17]
Understanding hand-object manipulation with grasp types and object attributes
Minjie Cai, Kris M Kitani, and Yoichi Sato. Understanding hand-object manipulation with grasp types and object attributes. In Robotics: Science and Systems, 2016
2016
-
[18]
Long-term human motion prediction with scene context
Zhe Cao, Hang Gao, Karttikeya Mangalam, Qi-Zhi Cai, Minh Vo, and Jitendra Malik. Long-term human motion prediction with scene context. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part I 16, pages 387--404. Springer, 2020
2020
-
[19]
A multi-sensor dataset of human-human handover
Alessandro Carf \` , Francesco Foglino, Barbara Bruno, and Fulvio Mastrogiovanni. A multi-sensor dataset of human-human handover. Data in brief, 22: 0 109--117, 2019
2019
-
[20]
On the choice of grasp type and location when handing over an object
Francesca Cini, V Ortenzi, P Corke, and MJSR Controzzi. On the choice of grasp type and location when handing over an object. Science Robotics, 4 0 (27): 0 eaau9757, 2019
2019
-
[21]
Context-aware human motion prediction, 2020
Enric Corona, Albert Pumarola, Guillem Alenyà, and Francesc Moreno-Noguer. Context-aware human motion prediction, 2020
2020
-
[22]
Towards collaborative robots as intelligent co-workers in human-robot joint tasks: what to do and who does it? In ISR 2020; 52th International Symposium on Robotics, pages 1--8
Ana Cunha, Flora Ferreira, Emanuel Sousa, Luis Louro, Paulo Vicente, Sergio Monteiro, Wolfram Erlhagen, and Estela Bicho. Towards collaborative robots as intelligent co-workers in human-robot joint tasks: what to do and who does it? In ISR 2020; 52th International Symposium on...
2020
-
[23]
You-do, i-learn: Egocentric unsupervised discovery of objects and their modes of interaction towards video-based guidance
Dima Damen, Teesid Leelasawassuk, and Walterio Mayol-Cuevas. You-do, i-learn: Egocentric unsupervised discovery of objects and their modes of interaction towards video-based guidance. Computer Vision and Image Understanding, 149: 0 98--112, 2016
2016
-
[24]
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Scaling egocentric vision: The epic-kitchens dataset. In Proceedings of the European conference on comput...
2018
-
[25]
Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens-100. International Journal ...
2022
-
[26]
Cg-hoi: Contact-guided 3d human-object interaction generation
Christian Diller and Angela Dai. Cg-hoi: Contact-guided 3d human-object interaction generation. arXiv preprint arXiv:2311.16097, 2023
2023 arXiv
-
[27]
Palm-e: An embodied multimodal language model
Danny Driess, Fei Xia, Mehdi SM Sajjadi, Corey Lynch, Aakanksha Chowdhery, Brian Ichter, Ayzaan Wahid, Jonathan Tompson, Quan Vuong, Tianhe Yu, et al. Palm-e: An embodied multimodal language model. arXiv preprint arXiv:2303.03378, 2023
2023 arXiv
-
[28]
Avatars grow legs: Generating smooth human motion from sparse tracking inputs with diffusion model
Yuming Du, Robin Kips, Albert Pumarola, Sebastian Starke, Ali Thabet, and Artsiom Sanakoyeu. Avatars grow legs: Generating smooth human motion from sparse tracking inputs with diffusion model. In CVPR, 2023
2023
-
[29]
Human preferences for robot eye gaze in human-to-robot handovers
Tair Faibish, Alap Kshirsagar, Guy Hoffman, and Yael Edan. Human preferences for robot eye gaze in human-to-robot handovers. International Journal of Social Robotics, 14 0 (4): 0 995--1012, 2022
2022
-
[30]
Go to zero: Towards zero-shot motion generation with million-scale data
Ke Fan, Shunlin Lu, Minyue Dai, Runyi Yu, Lixing Xiao, Zhiyang Dou, Junting Dong, Lizhuang Ma, and Jingbo Wang. Go to zero: Towards zero-shot motion generation with million-scale data. arXiv preprint arXiv:2507.07095, 2025 a
2025 arXiv
-
[31]
Arctic: A dataset for dexterous bimanual hand-object manipulation
Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J Black, and Otmar Hilliges. Arctic: A dataset for dexterous bimanual hand-object manipulation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12...
2023
-
[32]
Hold: Category-agnostic 3d reconstruction of interacting hands and objects from video
Zicong Fan, Maria Parelli, Maria Eleni Kadoglou, Xu Chen, Muhammed Kocabas, Michael J Black, and Otmar Hilliges. Hold: Category-agnostic 3d reconstruction of interacting hands and objects from video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco...
2024
-
[33]
Benchmarks and challenges in pose estimation for egocentric hand interactions with objects
Zicong Fan, Takehiko Ohkawa, Linlin Yang, Nie Lin, Zhishan Zhou, Shihao Zhou, Jiajun Liang, Zhong Gao, Xuanyang Zhang, Xue Zhang, et al. Benchmarks and challenges in pose estimation for egocentric hand interactions with objects. In European Conference on Computer Vision, pages...
2025
-
[34]
Social interactions: A first-person perspective
Alircza Fathi, Jessica K Hodgins, and James M Rehg. Social interactions: A first-person perspective. In 2012 IEEE Conference on Computer Vision and Pattern Recognition, pages 1226--1233. IEEE, 2012
2012
-
[35]
Three-dimensional reconstruction of human interactions
Mihai Fieraru, Mihai Zanfir, Elisabeta Oneata, Alin-Ionut Popa, Vlad Olaru, and Cristian Sminchisescu. Three-dimensional reconstruction of human interactions. In CVPR, pages 7214--7223, 2020
2020
-
[36]
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF Conference on Computer Vision ...
2022
-
[37]
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. In Proceedings...
2024
-
[38]
Generating diverse and natural 3d human motions from text
Chuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang, Wei Ji, Xingyu Li, and Li Cheng. Generating diverse and natural 3d human motions from text. In CVPR, pages 5152--5161, 2022 a
2022
-
[39]
Multi-person extreme motion prediction
Wen Guo, Xiaoyu Bie, Xavier Alameda-Pineda, and Francesc Moreno-Noguer. Multi-person extreme motion prediction. In CVPR, pages 13053--13064, 2022 b
2022
-
[40]
Interaction replica: Tracking human--object interaction and scene changes from human motion
Vladimir Guzov, Julian Chibane, Riccardo Marin, Yannan He, Yunus Saracoglu, Torsten Sattler, and Gerard Pons-Moll. Interaction replica: Tracking human--object interaction and scene changes from human motion. In 2024 International Conference on 3D Vision (3DV), pages 1006--1016...
2024
-
[41]
Honnotate: A method for 3d annotation of hand and object poses
Shreyas Hampali, Mahdi Rad, Markus Oberweger, and Vincent Lepetit. Honnotate: A method for 3d annotation of hand and object poses. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3196--3206, 2020
2020
-
[42]
Keypoint transformer: Solving joint identification in challenging hands and object interactions for accurate 3d pose estimation
Shreyas Hampali, Sayan Deb Sarkar, Mahdi Rad, and Vincent Lepetit. Keypoint transformer: Solving joint identification in challenging hands and object interactions for accurate 3d pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognit...
2022
-
[43]
Ego3dt: Tracking every 3d object in ego-centric videos
Shengyu Hao, Wenhao Chai, Zhonghan Zhao, Meiqi Sun, Wendi Hu, Jieyang Zhou, Yixian Zhao, Qi Li, Yizhou Wang, Xi Li, et al. Ego3dt: Tracking every 3d object in ego-centric videos. arXiv preprint arXiv:2410.08530, 2024
2024 arXiv
-
[44]
Resolving 3d human pose ambiguities with 3d scene constraints
Mohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, and Michael J Black. Resolving 3d human pose ambiguities with 3d scene constraints. In ICCV, pages 2282--2292, 2019
2019
-
[45]
Stochastic scene-aware motion prediction
Mohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito, Jimei Yang, Yi Zhou, and Michael J Black. Stochastic scene-aware motion prediction. In ICCV, pages 11374--11384, 2021
2021
-
[46]
Omnih2o: Universal and dexterous human-to-humanoid whole-body teleoperation and learning
Tairan He, Zhengyi Luo, Xialin He, Wenli Xiao, Chong Zhang, Weinan Zhang, Kris Kitani, Changliu Liu, and Guanya Shi. Omnih2o: Universal and dexterous human-to-humanoid whole-body teleoperation and learning. arXiv preprint arXiv:2406.08858, 2024 a
2024 arXiv
-
[47]
Learning human-to-humanoid real-time whole-body teleoperation, 2024 b
Tairan He, Zhengyi Luo, Wenli Xiao, Chong Zhang, Kris Kitani, Changliu Liu, and Guanya Shi. Learning human-to-humanoid real-time whole-body teleoperation, 2024 b
2024
-
[48]
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. Gans trained by a two time-scale update rule converge to a local nash equilibrium. Advances in neural information processing systems, 30, 2017
2017
-
[49]
Voxposer: Composable 3d value maps for robotic manipulation with language models
Wenlong Huang, Chen Wang, Ruohan Zhang, Yunzhu Li, Jiajun Wu, and Li Fei-Fei. Voxposer: Composable 3d value maps for robotic manipulation with language models. arXiv preprint arXiv:2307.05973, 2023
2023 arXiv
-
[50]
Intercap: Joint markerless 3d tracking of humans and objects in interaction
Yinghao Huang, Omid Taheri, Michael J Black, and Dimitrios Tzionas. Intercap: Joint markerless 3d tracking of humans and objects in interaction. In DAGM German Conference on Pattern Recognition, pages 281--299. Springer, 2022
2022
-
[51]
Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6m: Large scale datasets and predictive methods for 3d human sensing in natural environments. IEEE transactions on pattern analysis and machine intelligence, 36 0 (7): 0 1325--1339, 2013
2013
-
[52]
A large-scale rgb-d database for arbitrary-view human action recognition
Yanli Ji, Feixiang Xu, Yang Yang, Fumin Shen, Heng Tao Shen, and Wei-Shi Zheng. A large-scale rgb-d database for arbitrary-view human action recognition. In ACMMM, page 1510–1518, 2018
2018
-
[53]
Affordpose: A large-scale dataset of hand-object interactions with affordance-driven hand pose
Juntao Jian, Xiuping Liu, Manyi Li, Ruizhen Hu, and Jian Liu. Affordpose: A large-scale dataset of hand-object interactions with affordance-driven hand pose. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14713--14724, 2023
2023
-
[54]
Avatarposer: Articulated full-body pose tracking from sparse motion sensing
Jiaxi Jiang, Paul Streli, Huajian Qiu, Andreas Fender, Larissa Laich, Patrick Snape, and Christian Holz. Avatarposer: Articulated full-body pose tracking from sparse motion sensing. In ECCV, pages 443--460. Springer, 2022
2022
-
[55]
Egoposer: Robust real-time ego-body pose estimation in large scenes
Jiaxi Jiang, Paul Streli, Manuel Meier, and Christian Holz. Egoposer: Robust real-time ego-body pose estimation in large scenes. arXiv preprint arXiv:2308.06493, 2023 a
2023 arXiv
-
[56]
Full-body articulated human-object interaction
Nan Jiang, Tengyu Liu, Zhexuan Cao, Jieming Cui, Zhiyuan Zhang, Yixin Chen, He Wang, Yixin Zhu, and Siyuan Huang. Full-body articulated human-object interaction. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9365--9376, 2023 b
2023
-
[57]
Scaling up dynamic human-scene interaction modeling
Nan Jiang, Zhiyuan Zhang, Hongjie Li, Xiaoxuan Ma, Zan Wang, Yixin Chen, Tengyu Liu, Yixin Zhu, and Siyuan Huang. Scaling up dynamic human-scene interaction modeling. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1737--1747, 2024 a
2024
-
[58]
A probabilistic attention model with occlusion-aware texture regression for 3d hand reconstruction from a single rgb image
Zheheng Jiang, Hossein Rahmani, Sue Black, and Bryan M Williams. A probabilistic attention model with occlusion-aware texture regression for 3d hand reconstruction from a single rgb image. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pa...
2023
-
[59]
Harmon: Whole-body motion generation of humanoid robots from language descriptions
Zhenyu Jiang, Yuqi Xie, Jinhan Li, Ye Yuan, Yifeng Zhu, and Yuke Zhu. Harmon: Whole-body motion generation of humanoid robots from language descriptions. arXiv preprint arXiv:2410.12773, 2024 b
2024 arXiv
-
[60]
Epic-fusion: Audio-visual temporal binding for egocentric action recognition
Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen. Epic-fusion: Audio-visual temporal binding for egocentric action recognition. In Proceedings of the IEEE/CVF international conference on computer vision, pages 5492--5501, 2019
2019
-
[61]
Real-time active vision for a humanoid soccer robot using deep reinforcement learning
Soheil Khatibi, Meisam Teimouri, and Mahdi Rezaei. Real-time active vision for a humanoid soccer robot using deep reinforcement learning. arXiv preprint arXiv:2011.13851, 2020
2011 arXiv
-
[62]
Openvla: An open-source vision-language-action model
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan Foster, Grace Lam, Pannag Sanketi, et al. Openvla: An open-source vision-language-action model. arXiv preprint arXiv:2406.09246, 2024
2024 arXiv
-
[63]
Dataset of bimanual human-to-human object handovers
Alap Kshirsagar, Raphael Fortuna, Zhiming Xie, and Guy Hoffman. Dataset of bimanual human-to-human object handovers. Data in Brief, 48: 0 109277, 2023
2023
-
[64]
H2o: Two hands manipulating objects for first person interaction recognition
Taein Kwon, Bugra Tekin, Jan St \"u hmer, Federica Bogo, and Marc Pollefeys. H2o: Two hands manipulating objects for first person interaction recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 10138--10148, 2021
2021
-
[65]
Ego-body pose estimation via ego-head pose estimation
Jiaman Li, Karen Liu, and Jiajun Wu. Ego-body pose estimation via ego-head pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17142--17151, 2023
2023
-
[66]
Dngaussian: Optimizing sparse-view 3d gaussian radiance fields with global-local depth normalization
Jiahe Li, Jiawei Zhang, Xiao Bai, Jin Zheng, Xin Ning, Jun Zhou, and Lin Gu. Dngaussian: Optimizing sparse-view 3d gaussian radiance fields with global-local depth normalization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20775-...
2024
-
[67]
In the eye of beholder: Joint learning of gaze and actions in first person video
Yin Li, Miao Liu, and James M Rehg. In the eye of beholder: Joint learning of gaze and actions in first person video. In Proceedings of the European conference on computer vision (ECCV), pages 619--635, 2018
2018
-
[68]
Ego-exo: Transferring visual representations from third-person to first-person videos
Yanghao Li, Tushar Nagarajan, Bo Xiong, and Kristen Grauman. Ego-exo: Transferring visual representations from third-person to first-person videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6943--6953, 2021
2021
-
[69]
Intergen: Diffusion-based multi-human motion generation under complex interactions
Han Liang, Wenqian Zhang, Wenxuan Li, Jingyi Yu, and Lan Xu. Intergen: Diffusion-based multi-human motion generation under complex interactions. arXiv preprint arXiv:2304.05684, 2023
2023 arXiv
-
[70]
Motion-x: A large-scale 3d expressive whole-body human motion dataset
Jing Lin, Ailing Zeng, Shunlin Lu, Yuanhao Cai, Ruimao Zhang, Haoqian Wang, and Lei Zhang. Motion-x: A large-scale 3d expressive whole-body human motion dataset. arXiv preprint arXiv:2307.00818, 2023
2023 arXiv
-
[71]
Gaussian-flow: 4d reconstruction with dynamic 3d gaussian particle
Youtian Lin, Zuozhuo Dai, Siyu Zhu, and Yao Yao. Gaussian-flow: 4d reconstruction with dynamic 3d gaussian particle. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21136--21145, 2024
2024
-
[72]
Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding
Jun Liu, Amir Shahroudy, Mauricio Perez, Gang Wang, Ling-Yu Duan, and Alex C Kot. Ntu rgb+ d 120: A large-scale benchmark for 3d human activity understanding. T-PAMI, 42 0 (10): 0 2684--2701, 2019
2019
-
[73]
Forecasting human-object interaction: joint prediction of motor attention and actions in first person video
Miao Liu, Siyu Tang, Yin Li, and James M Rehg. Forecasting human-object interaction: joint prediction of motor attention and actions in first person video. In Computer Vision--ECCV 2020: 16th European Conference, Glasgow, UK, August 23--28, 2020, Proceedings, Part I 16, pages ...
2020
-
[74]
Joint hand motion and interaction hotspots prediction from egocentric videos
Shaowei Liu, Subarna Tripathi, Somdeb Majumdar, and Xiaolong Wang. Joint hand motion and interaction hotspots prediction from egocentric videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3282--3292, 2022 a
2022
-
[75]
Hoi4d: A 4d egocentric dataset for category-level human-object interaction
Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, and Li Yi. Hoi4d: A 4d egocentric dataset for category-level human-object interaction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pa...
2022
-
[76]
Taco: Benchmarking generalizable bimanual tool-action-object understanding
Yun Liu, Haolin Yang, Xu Si, Ling Liu, Zipeng Li, Yuxiang Zhang, Yebin Liu, and Li Yi. Taco: Benchmarking generalizable bimanual tool-action-object understanding. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21740--21751, 2024
2024
-
[77]
Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons - Moll, and Michael J. Black. SMPL: a skinned multi-person linear model. ACM Trans. Graph. , 34 0 (6): 0 248:1--248:16, 2015
2015
-
[78]
Dynamics-regulated kinematic policy for egocentric pose estimation
Zhengyi Luo, Ryo Hachiuma, Ye Yuan, and Kris Kitani. Dynamics-regulated kinematic policy for egocentric pose estimation. Advances in Neural Information Processing Systems, 34: 0 25019--25032, 2021
2021
-
[79]
Himo: A new benchmark for full-body human interacting with multiple objects
Xintao Lv, Liang Xu, Yichao Yan, Xin Jin, Congsheng Xu, Shuwen Wu, Yifan Liu, Lincheng Li, Mengxiao Bi, Wenjun Zeng, et al. Himo: A new benchmark for full-body human interacting with multiple objects. In European Conference on Computer Vision, pages 300--318. Springer, 2025
2025
-
[80]
Diff-ip2d: Diffusion-based hand-object interaction prediction on egocentric videos
Junyi Ma, Jingyi Xu, Xieyuanli Chen, and Hesheng Wang. Diff-ip2d: Diffusion-based hand-object interaction prediction on egocentric videos. arXiv preprint arXiv:2405.04370, 2024 a
2024
-
[81]
A survey on vision-language-action models for embodied ai
Yueen Ma, Zixing Song, Yuzheng Zhuang, Jianye Hao, and Irwin King. A survey on vision-language-action models for embodied ai. arXiv preprint arXiv:2405.14093, 2024 b
2024 arXiv
-
[82]
Unifying representations and large-scale whole-body motion databases for studying human motion
Christian Mandery, \"O mer Terlemez, Martin Do, Nikolaus Vahrenkamp, and Tamim Asfour. Unifying representations and large-scale whole-body motion databases for studying human motion. IEEE Transactions on Robotics, 32 0 (4): 0 796--809, 2016
2016
-
[83]
Dexvip: Learning dexterous grasping with human hand pose priors from video
Priyanka Mandikal and Kristen Grauman. Dexvip: Learning dexterous grasping with human hand pose priors from video. In Conference on Robot Learning, pages 651--661. PMLR, 2022
2022
-
[84]
Hoi4abot: Human-object interaction anticipation for human intention reading collaborative robots, 2024
Esteve Valls Mascaro, Daniel Sliwowski, and Dongheui Lee. Hoi4abot: Human-object interaction anticipation for human intention reading collaborative robots, 2024
2024
-
[85]
Eventego3d: 3d human motion capture from egocentric event streams
Christen Millerdurai, Hiroyasu Akada, Jian Wang, Diogo Luvizon, Christian Theobalt, and Vladislav Golyanik. Eventego3d: 3d human motion capture from egocentric event streams. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1186--1195, 2024
2024
-
[86]
imapper: interaction-guided scene mapping from monocular videos
Aron Monszpart, Paul Guerrero, Duygu Ceylan, Ersin Yumer, and Niloy J Mitra. imapper: interaction-guided scene mapping from monocular videos. ACM Transactions On Graphics (TOG), 38 0 (4): 0 1--15, 2019
2019
-
[87]
Grounded human-object interaction hotspots from video
Tushar Nagarajan, Christoph Feichtenhofer, and Kristen Grauman. Grounded human-object interaction hotspots from video. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8688--8697, 2019
2019
-
[88]
Jointly learning energy expenditures and activities using egocentric multimodal signals
Katsuyuki Nakamura, Serena Yeung, Alexandre Alahi, and Li Fei-Fei. Jointly learning energy expenditures and activities using egocentric multimodal signals. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 1868--1877, 2017
2017
-
[89]
You2me: Inferring body pose in egocentric video via first and second person interactions
Evonne Ng, Donglai Xiang, Hanbyul Joo, and Kristen Grauman. You2me: Inferring body pose in egocentric video via first and second person interactions. In CVPR, pages 9890--9900, 2020
2020
-
[90]
Open x-embodiment: Robotic learning datasets and rt-x models
Abby O'Neill, Abdul Rehman, Abhinav Gupta, Abhiram Maddukuri, Abhishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, et al. Open x-embodiment: Robotic learning datasets and rt-x models. arXiv preprint arXiv:2310.08864, 2023
2023 arXiv
-
[91]
GPT -3.5 turbo fine-tuning and api updates
OpenAI. GPT -3.5 turbo fine-tuning and api updates. https://openai.com/index/gpt-3-5-turbo-fine-tuning-and-api-updates/, 2023
2023
-
[92]
Handoccnet: Occlusion-robust 3d hand mesh estimation network
JoonKyu Park, Yeonguk Oh, Gyeongsik Moon, Hongsuk Choi, and Kyoung Mu Lee. Handoccnet: Occlusion-robust 3d hand mesh estimation network. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1496--1505, 2022
2022
-
[93]
Reconstructing hands in 3 D with transformers
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Reconstructing hands in 3 D with transformers. In CVPR, 2024
2024
-
[94]
Hoi-diff: Text-driven synthesis of 3d human-object interactions using diffusion models
Xiaogang Peng, Yiming Xie, Zizhao Wu, Varun Jampani, Deqing Sun, and Huaizu Jiang. Hoi-diff: Text-driven synthesis of 3d human-object interactions using diffusion models. arXiv preprint arXiv:2312.06553, 2023
2023 arXiv
-
[95]
Action-conditioned 3d human motion synthesis with transformer vae
Mathis Petrovich, Michael J Black, and G \"u l Varol. Action-conditioned 3d human motion synthesis with transformer vae. In CVPR, pages 10985--10995, 2021
2021
-
[96]
The kit motion-language dataset
Matthias Plappert, Christian Mandery, and Tamim Asfour. The kit motion-language dataset. Big data, 4 0 (4): 0 236--252, 2016
2016
-
[97]
Wilor: End-to-end 3d hand localization and reconstruction in-the-wild
Rolandos Alexandros Potamias, Jinglei Zhang, Jiankang Deng, and Stefanos Zafeiriou. Wilor: End-to-end 3d hand localization and reconstruction in-the-wild. arXiv preprint arXiv:2409.12259, 2024
2024 arXiv
-
[98]
Mild: multimodal interactive latent dynamics for learning human-robot interaction
Vignesh Prasad, Dorothea Koert, Ruth Stock-Homburg, Jan Peters, and Georgia Chalvatzaki. Mild: multimodal interactive latent dynamics for learning human-robot interaction. In 2022 IEEE-RAS 21st International Conference on Humanoid Robots (Humanoids), pages 472--479. IEEE, 2022
2022
-
[99]
Moveint: Mixture of variational experts for learning human-robot interactions from demonstrations
Vignesh Prasad, Alap Kshirsagar, Dorothea Koert Ruth Stock-Homburg, Jan Peters, and Georgia Chalvatzaki. Moveint: Mixture of variational experts for learning human-robot interactions from demonstrations. IEEE Robotics and Automation Letters, 2024
2024
-
[100]
The virtual caliper: Rapid creation of metrically accurate avatars from 3d measurements
Sergi Pujades, Betty Mohler, Anne Thaler, Joachim Tesch, Naureen Mahmood, Nikolas Hesse, Heinrich H B \"u lthoff, and Michael J Black. The virtual caliper: Rapid creation of metrically accurate avatars from 3d measurements. IEEE transactions on visualization and computer graph...
2019
-
[101]
D-nerf: Neural radiance fields for dynamic scenes
Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-nerf: Neural radiance fields for dynamic scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10318--10327, 2021
2021
-
[102]
Babel: bodies, action and behavior with english labels
Abhinanda R Punnakkal, Arjun Chandrasekaran, Nikos Athanasiou, Alejandra Quiros-Ramirez, and Michael J Black. Babel: bodies, action and behavior with english labels. In CVPR, pages 722--731, 2021
2021
-
[103]
Dexmv: Imitation learning for dexterous manipulation from human videos
Yuzhe Qin, Yueh-Hua Wu, Shaowei Liu, Hanwen Jiang, Ruihan Yang, Yang Fu, and Xiaolong Wang. Dexmv: Imitation learning for dexterous manipulation from human videos. In European Conference on Computer Vision, pages 570--587. Springer, 2022
2022
-
[104]
Imitation learning based on bilateral control for human--robot cooperation
Ayumu Sasagawa, Kazuki Fujimoto, Sho Sakaino, and Toshiaki Tsuji. Imitation learning based on bilateral control for human--robot cooperation. IEEE Robotics and Automation Letters, 5 0 (4): 0 6169--6176, 2020
2020
-
[105]
Pigraphs: learning interaction snapshots from observations
Manolis Savva, Angel X Chang, Pat Hanrahan, Matthew Fisher, and Matthias Nie ner. Pigraphs: learning interaction snapshots from observations. ACM Transactions On Graphics (TOG), 35 0 (4): 0 1--12, 2016
2016
-
[106]
Deep imitation learning for humanoid loco-manipulation through human teleoperation, 2023
Mingyo Seo, Steve Han, Kyutae Sim, Seung Hyeon Bang, Carlos Gonzalez, Luis Sentis, and Yuke Zhu. Deep imitation learning for humanoid loco-manipulation through human teleoperation, 2023
2023
-
[107]
Human motion diffusion as a generative prior
Yonatan Shafir, Guy Tevet, Roy Kapon, and Amit H Bermano. Human motion diffusion as a generative prior. arXiv preprint arXiv:2303.01418, 2023
2023 arXiv
-
[108]
Modeling ambient scene dynamics for free-view synthesis
Meng-Li Shih, Jia-Bin Huang, Changil Kim, Rajvi Shah, Johannes Kopf, and Chen Gao. Modeling ambient scene dynamics for free-view synthesis. In ACM SIGGRAPH 2024 Conference Papers, pages 1--11, 2024
2024
-
[109]
Wham: Reconstructing world-grounded humans with accurate 3d motion
Soyong Shin, Juyong Kim, Eni Halilaj, and Michael J Black. Wham: Reconstructing world-grounded humans with accurate 3d motion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2070--2080, 2024
-
[110]
Trace: 5d temporal regression of avatars with dynamic cameras in 3d environments
Yu Sun, Qian Bao, Wu Liu, Tao Mei, and Michael J Black. Trace: 5d temporal regression of avatars with dynamic cameras in 3d environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8856--8866, 2023
2023
-
[111]
Grab: A dataset of whole-body human grasping of objects
Omid Taheri, Nima Ghorbani, Michael J Black, and Dimitrios Tzionas. Grab: A dataset of whole-body human grasping of objects. In ECCV, pages 581--600, 2020
2020
-
[112]
Human motion diffusion model
Guy Tevet, Sigal Raab, Brian Gordon, Yonatan Shafir, Daniel Cohen-Or, and Amit H Bermano. Human motion diffusion model. arXiv preprint arXiv:2209.14916, 2022
2022 arXiv
-
[113]
Learning robot soccer from egocentric vision with deep reinforcement learning
Dhruva Tirumala, Markus Wulfmeier, Ben Moran, Sandy Huang, Jan Humplik, Guy Lever, Tuomas Haarnoja, Leonard Hasenclever, Arunkumar Byravan, Nathan Batchelor, et al. Learning robot soccer from egocentric vision with deep reinforcement learning. arXiv preprint arXiv:2405.02425, 2024
2024 arXiv
-
[114]
Umpm benchmark: A multi-person dataset with synchronized video and motion capture data for evaluation of articulated human motion and interaction
NP Van der Aa, Xinghan Luo, Geert-Jan Giezeman, Robby T Tan, and Remco C Veltkamp. Umpm benchmark: A multi-person dataset with synchronized video and motion capture data for evaluation of articulated human motion and interaction. In ICCV Workshops, pages 1264--1269. IEEE, 2011
2011
-
[115]
Sparsenerf: Distilling depth ranking for few-shot novel view synthesis
Guangcong Wang, Zhaoxi Chen, Chen Change Loy, and Ziwei Liu. Sparsenerf: Distilling depth ranking for few-shot novel view synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 9065--9076, 2023 a
2023
-
[116]
Scene-aware egocentric 3d human pose estimation
Jian Wang, Diogo Luvizon, Weipeng Xu, Lingjie Liu, Kripasindhu Sarkar, and Christian Theobalt. Scene-aware egocentric 3d human pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13031--13040, 2023 b
2023
-
[117]
Egocentric whole-body motion capture with fisheyevit and diffusion-based motion refinement
Jian Wang, Zhe Cao, Diogo Luvizon, Lingjie Liu, Kripasindhu Sarkar, Danhang Tang, Thabo Beeler, and Christian Theobalt. Egocentric whole-body motion capture with fisheyevit and diffusion-based motion refinement. In Proceedings of the IEEE/CVF Conference on Computer Vision and ...
2024
-
[118]
Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world
Xin Wang, Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bugra Tekin, Felipe Vieira Frujeri, et al. Holoassist: an egocentric human interaction dataset for interactive ai assistants in the real world. In Proceedings of the IEEE/...
2023
-
[119]
Quo vadis, motion generation? from large language models to large motion models, 2024 b
Ye Wang, Sipeng Zheng, Bin Cao, Qianshan Wei, Qin Jin, and Zongqing Lu. Quo vadis, motion generation? from large language models to large motion models, 2024 b
2024
-
[120]
Tram: Global trajectory and motion of 3d humans from in-the-wild videos
Yufu Wang, Ziyun Wang, Lingjie Liu, and Kostas Daniilidis. Tram: Global trajectory and motion of 3d humans from in-the-wild videos. In European Conference on Computer Vision, pages 467--487. Springer, 2025
2025
-
[121]
Genh2r: Learning generalizable human-to-robot handover via scalable simulation demonstration and imitation
Zifan Wang, Junyu Chen, Ziqing Chen, Pengwei Xie, Rui Chen, and Li Yi. Genh2r: Learning generalizable human-to-robot handover via scalable simulation demonstration and imitation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16362-...
2024
-
[122]
Hoh: Markerless multimodal human-object-human handover dataset with large object count
Noah Wiederhold, Ava Megyeri, DiMaggio Paris, Sean Banerjee, and Natasha Banerjee. Hoh: Markerless multimodal human-object-human handover dataset with large object count. Advances in Neural Information Processing Systems, 36, 2024
2024
-
[123]
4d gaussian splatting for real-time dynamic scene rendering
Guanjun Wu, Taoran Yi, Jiemin Fang, Lingxi Xie, Xiaopeng Zhang, Wei Wei, Wenyu Liu, Qi Tian, and Xinggang Wang. 4d gaussian splatting for real-time dynamic scene rendering. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20310--20320, 2024
2024
-
[124]
Unified human-scene interaction via prompted chain-of-contacts
Zeqi Xiao, Tai Wang, Jingbo Wang, Jinkun Cao, Wenwei Zhang, Bo Dai, Dahua Lin, and Jiangmiao Pang. Unified human-scene interaction via prompted chain-of-contacts. arXiv preprint arXiv:2309.07918, 2023
2023 arXiv
-
[125]
Sparsegs: Real-time 360 \ deg \ sparse view synthesis using gaussian splatting
Haolin Xiong, Sairisheek Muttukuru, Rishi Upadhyay, Pradyumna Chari, and Achuta Kadambi. Sparsegs: Real-time 360 \ deg \ sparse view synthesis using gaussian splatting. arXiv preprint arXiv:2312.00206, 2023
2023 arXiv
-
[126]
Actformer: A gan-based transformer towards general action-conditioned 3d human motion generation
Liang Xu, Ziyang Song, Dongliang Wang, Jing Su, Zhicheng Fang, Chenjing Ding, Weihao Gan, Yichao Yan, Xin Jin, Xiaokang Yang, et al. Actformer: A gan-based transformer towards general action-conditioned 3d human motion generation. In ICCV, pages 2228--2238, 2023 a
2023
-
[127]
Motionbank: A large-scale video motion benchmark with disentangled rule-based annotations, 2024 a
Liang Xu, Shaoyang Hua, Zili Lin, Yifan Liu, Feipeng Ma, Yichao Yan, Xin Jin, Xiaokang Yang, and Wenjun Zeng. Motionbank: A large-scale video motion benchmark with disentangled rule-based annotations, 2024 a
2024
-
[128]
Inter-x: Towards versatile human-human interaction analysis
Liang Xu, Xintao Lv, Yichao Yan, Xin Jin, Shuwen Wu, Congsheng Xu, Yifan Liu, Yizhou Zhou, Fengyun Rao, Xingdong Sheng, et al. Inter-x: Towards versatile human-human interaction analysis. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pag...
2024
-
[129]
Regennet: Towards human action-reaction synthesis
Liang Xu, Yizhou Zhou, Yichao Yan, Xin Jin, Wenhan Zhu, Fengyun Rao, Xiaokang Yang, and Wenjun Zeng. Regennet: Towards human action-reaction synthesis. In CVPR, pages 1759--1769, 2024 c
2024
-
[130]
InterDiff : Generating 3d human-object interactions with physics-informed diffusion
Sirui Xu, Zhengyuan Li, Yu-Xiong Wang, and Liang-Yan Gui. InterDiff : Generating 3d human-object interactions with physics-informed diffusion. In ICCV, 2023 b
2023
-
[131]
D3d-hoi: Dynamic 3d human-object interactions from videos
Xiang Xu, Hanbyul Joo, Greg Mori, and Manolis Savva. D3d-hoi: Dynamic 3d human-object interactions from videos. arXiv preprint arXiv:2108.08420, 2021
2021 arXiv
-
[132]
Humanvla: Towards vision-language directed object rearrangement by physical humanoid
Xinyu Xu, Yizheng Zhang, Yong-Lu Li, Lei Han, and Cewu Lu. Humanvla: Towards vision-language directed object rearrangement by physical humanoid. arXiv preprint arXiv:2406.19972, 2024 d
2024 arXiv
-
[133]
What's in your hands? 3d reconstruction of generic objects in hands
Yufei Ye, Abhinav Gupta, and Shubham Tulsiani. What's in your hands? 3d reconstruction of generic objects in hands. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 3895--3905, 2022
2022
-
[134]
Diffusion-guided reconstruction of everyday hand-object interaction clips
Yufei Ye, Poorvi Hebbar, Abhinav Gupta, and Shubham Tulsiani. Diffusion-guided reconstruction of everyday hand-object interaction clips. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 19717--19728, 2023
2023
-
[135]
Estimating body and hand motion in an ego-sensed world, 2024
Brent Yi, Vickie Ye, Maya Zheng, Lea Müller, Georgios Pavlakos, Yi Ma, Jitendra Malik, and Angjoo Kanazawa. Estimating body and hand motion in an ego-sensed world, 2024
2024
-
[136]
Hi4d: 4d instance segmentation of close human interaction
Yifei Yin, Chen Guo, Manuel Kaufmann, Juan Jose Zarate, Jie Song, and Otmar Hilliges. Hi4d: 4d instance segmentation of close human interaction. In CVPR, pages 17016--17027, 2023
2023
-
[137]
Lm-gaussian: Boost sparse-view 3d gaussian splatting with large model priors
Hanyang Yu, Xiaoxiao Long, and Ping Tan. Lm-gaussian: Boost sparse-view 3d gaussian splatting with large model priors. arXiv preprint arXiv:2409.03456, 2024
2024 arXiv
-
[138]
Acr: Attention collaboration-based regressor for arbitrary two-hand reconstruction
Zhengdi Yu, Shaoli Huang, Chen Fang, Toby P Breckon, and Jue Wang. Acr: Attention collaboration-based regressor for arbitrary two-hand reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 12955--12964, 2023
2023
-
[139]
Glamr: Global occlusion-aware human mesh recovery with dynamic cameras
Ye Yuan, Umar Iqbal, Pavlo Molchanov, Kris Kitani, and Jan Kautz. Glamr: Global occlusion-aware human mesh recovery with dynamic cameras. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 11038--11049, 2022
2022
-
[140]
Two-person interaction detection using body-pose features and multiple instance learning
Kiwon Yun, Jean Honorio, Debaleena Chattopadhyay, Tamara L Berg, and Dimitris Samaras. Two-person interaction detection using body-pose features and multiple instance learning. In CVPR, pages 28--35. IEEE, 2012
2012
-
[141]
Core4d: A 4d human-object-human interaction dataset for collaborative object rearrangement
Chengwen Zhang, Yun Liu, Ruofan Xing, Bingda Tang, and Li Yi. Core4d: A 4d human-object-human interaction dataset for collaborative object rearrangement. arXiv preprint arXiv:2406.19353, 2024 a
2024 arXiv
-
[142]
Transplat: Generalizable 3d gaussian splatting from sparse multi-view images with transformers
Chuanrui Zhang, Yingshuang Zou, Zhuoling Li, Minmin Yi, and Haoqian Wang. Transplat: Generalizable 3d gaussian splatting from sparse multi-view images with transformers. arXiv preprint arXiv:2408.13770, 2024 b
2024 arXiv
-
[143]
Enhancing human-centered dynamic scene understanding via multiple llms collaborated reasoning
Hang Zhang, Wenxiao Zhang, Haoxuan Qu, and Jun Liu. Enhancing human-centered dynamic scene understanding via multiple llms collaborated reasoning. Visual Intelligence, 3 0 (1): 0 3, 2025
2025
-
[144]
Hoi-m\^ 3: Capture multiple humans and objects interaction within contextual environment
Juze Zhang, Jingyan Zhang, Zining Song, Zhanhe Shi, Chengfeng Zhao, Ye Shi, Jingyi Yu, Lan Xu, and Jingya Wang. Hoi-m\^ 3: Capture multiple humans and objects interaction within contextual environment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Re...
2024
-
[145]
Couch: towards controllable human-chair interactions
Xiaohan Zhang, Bharat Lal Bhatnagar, Sebastian Starke, Vladimir Guzov, and Gerard Pons-Moll. Couch: towards controllable human-chair interactions. In ECCV, pages 518--535, 2022
2022
-
[146]
Scenic: Scene-aware semantic navigation with instruction-guided control
Xiaohan Zhang, Sebastian Starke, Vladimir Guzov, Zhensong Zhang, Eduardo P \'e rez Pellitero, and Gerard Pons-Moll. Scenic: Scene-aware semantic navigation with instruction-guided control. arXiv preprint arXiv:2412.15664, 2024 d
2024 arXiv
-
[147]
3d-vla: A 3d vision-language-action generative world model
Haoyu Zhen, Xiaowen Qiu, Peihao Chen, Jincheng Yang, Xin Yan, Yilun Du, Yining Hong, and Chuang Gan. 3d-vla: A 3d vision-language-action generative world model. arXiv preprint arXiv:2403.09631, 2024
2024 arXiv
-
[148]
Temporal perception and prediction in ego-centric video
Yipin Zhou and Tamara L Berg. Temporal perception and prediction in ego-centric video. In Proceedings of the IEEE International Conference on Computer Vision, pages 4498--4506, 2015
2015
-
[149]
On the continuity of rotation representations in neural networks
Yi Zhou, Connelly Barnes, Jingwan Lu, Jimei Yang, and Hao Li. On the continuity of rotation representations in neural networks. In CVPR , pages 5745--5753. Computer Vision Foundation / IEEE , 2019
2019
-
[150]
3d human shape reconstruction from a polarization image
Shihao Zou, Xinxin Zuo, Yiming Qian, Sen Wang, Chi Xu, Minglun Gong, and Li Cheng. 3d human shape reconstruction from a polarization image. In ECCV, pages 351--368, 2020
2020
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.