REVIEW 5 major objections 6 minor 119 references
FIction: 4D Future Interaction Prediction from Video
T0 review · 5 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read FICTION predicts the 3D locations and body poses of all future object interactions for three minutes ahead, from a 30-second video observation plus a voxel scene map, beating six adapted baselines by more than 30 percent.
desk verdict New task and benchmark are real, but the central comparison is undermined by an unstated full-take object map, missing error bars, and an overblown abstract. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the multimodal transformer encoder that fuses three streams: visual features from the observed video, the actor's SMPL pose parameters, and a $N \times N \times N$ voxelized scene map $S$ in which occupied cells carry object-class indices and the actor's location is marked with a reserved index. A linear decoder maps the fused representation onto a future-interaction voxel grid trained with binary cross-entropy, and a conditional VAE takes the same representation plus a query location and produces a Gaussian latent space from which SMPL body poses can be sampled. The voxel grid is what ties 'where' and 'how' together; in the ablations, removing the environment stream causes the largest performance drop.
What would settle it
Train and evaluate FICTION twice, once with the voxel scene map $S$ built only from frames at or before $\tau_o=30$ s and once with $S$ built from all frames, and count how many object boxes in $S$ come from objects whose first appearance in the video is after 30 seconds. If PR-AUC drops materially when $S$ is cropped to the observed period, the prediction is partly reading the future object inventory from the map rather than forecasting it.
Extended reading notes
Core claim
The paper's central discovery is that early fusion of three signals, past video, past SMPL body pose, and a persistent 3D voxel representation of object locations, lets a single transformer predict both components of future interaction in a shared 3D space. The location branch decodes the fused representation into a voxel grid marking every 3D cell that will be touched in the next 180 seconds; the pose branch is a conditional VAE that, given a query location, samples the body pose distribution likely to be executed there. On the EgoExo4D procedural benchmark, this beats six adapted baselines: autoregressive models drift and diverge over the three-minute horizon, while video-to-3D scene models without explicit activity grounding predict less accurately. Ablations support the same conclusion: removing the scene map drops PR-AUC from 21.0 to 9.9 on cooking, while removing video or pose degrades it less.
Load-bearing premise
The paper never states whether the 3D object map is built only from video frames up to the $\tau_o=30$ second observation time or from the whole video, so objects that appear only after the observation could be leaking future scene content into the model's input.
Editorial extensions
If this is right
- If FICTION is right, explicit 3D scene conditioning is the ingredient that makes minute-scale interaction anticipation work: removing the object map in the ablations causes a larger drop in PR-AUC than removing either the video or the pose stream.
- Autoregressive methods that generate the next token in a sequence are shown to accumulate errors and diverge over the three-minute horizon, which positions direct decoding into a voxel grid as the better architecture for long-horizon forecasts.
- The released dataset and fixed task protocol ($\tau_o = 30$ s observation, $\tau_a = 5$ s anticipation gap, $\tau_f = 180$ s future, with PR-AUC, Chamfer distance, MPJPE, and PA-MPJPE) give subsequent work a public benchmark for 4D future interaction prediction.
- Assistive agents could turn the forecast into concrete preparation: knowing the interaction location and likely pose distribution, a robot or AR coach can position itself and cue the right assistance before the contact happens.
Reading between the lines
- Our inference: the paper does not state whether the 3D object map $S$ is built only from frames up to $\tau_o = 30$ s or from the whole video, so if $S$ contains objects that appear after the observation, some of the reported prediction could be reading the future object inventory from the map.
- Our inference: if the method retains its gains when $S$ is strictly cropped to the observed period, the model is probably learning procedural scripts such as 'fridge, then faucet, then cabinet' rather than physical dynamics; swapping object layouts in otherwise identical procedures would separate those two mechanisms.
- Our inference: the pose evaluation scores the closest of five sampled poses to the ground truth, so a model that spreads its samples widely is rewarded; a stricter utility measure would use expected error or calibrated diversity over all samples.
- Our inference: the most natural deployment setting for this input design is a robot that already knows the room layout from a prior scan, where the object map is genuinely available before the activity starts; the paper's setup implicitly assumes such a map.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces FICTION, a method and benchmark for 4D future interaction prediction from egocentric video. Given an observation window up to time tau_o, the model predicts (i) all future 3D object interaction locations within a tau_f = 3-minute horizon, encoded as a voxel grid, and (ii) a distribution of SMPL body poses at each query interaction location, modeled with a CVAE. The model fuses three inputs: egocentric video features (EgoVLPv2), the actor's observed body pose (WHAM), and a voxelized 3D object placement map S built from Detic detections and DBSCAN clustering. The dataset is curated from Ego-Exo4D procedural videos (cooking, bike repair, health) with interaction instances defined by hand-in-3D-bounding-box plus Llama-3.1 narration matching. Experiments compare against six baselines (HierVL, OCT, OccFormer, VoxFormer, 4D-Humans, T2M-GPT, and a video-to-pose CVAE) and report PR-AUC and Chamfer distance for locations and MPJPE/PA-MPJPE for poses, showing consistent improvements on all three scenarios.
Significance. If the empirical claims hold, this is a novel and worthwhile task formulation that connects long-horizon activity anticipation with explicit 3D environment context, and the released dataset could enable follow-up work. The architecture is simple and interpretable, and the paper includes useful ablations (video, pose, environment) and hyperparameter studies in the supplementary material. However, the central comparison is compromised by an unspecified temporal cutoff in the construction of the input object map S, and the headline '>30% relative gains' is not supported by the paper's own tables. The significance of the contribution is therefore contingent on resolving these issues.
major comments (5)
- [Sec. 3.2, Sec. 3.3, Appendix E] The 3D object placement map S, a primary input, is not explicitly constrained to frames with timestamps up to tau_o. Section 3.2 describes building S as a voxel grid with object indices and the actor location, and Section 3.3 says object bounding boxes are computed by running Detic at 30 fps on video frames and clustering the resulting 3D points, without any temporal cutoff. Appendix E states: 'We also assume a static point cloud when creating the dataset... It is possible to use 3D information only from the last time segment for improving the spatial input to the model, we do not consider this case for the ease of the I/O.' This strongly suggests that S aggregates object placements across the full take, including frames after tau_o. If an object first appears after the observation window but is present in S, the model knows that the object exists and where it is, leaking future scene content into the 'observation.' The ablations show the model depends heavily on S (w/o env PR-AUC drops from 21.0 to 9.9 on cooking, 18.7 to 6.0 on bike repair, and 12.7 to 4.7 on health), and baselines do not receive S. The reported gains may therefore be inflated by this leakage. The manuscript must either state that S is built only from frames up to tau_o, or clarify that S comes from a genuinely prior environment scan independent of the current take; in the latter case, baselines should be given equivalent information.
- [Abstract, Sec. 1, Table 1] The abstract and introduction claim 'more than 30% relative gains' over the best baseline, but Table 1 does not support this as a summary statement. Location PR-AUC relative improvements are 24.3% on cooking (21.0 vs 16.9), 32.6% on bike repair (18.7 vs 14.1), and 13.4% on health (12.7 vs 11.2). Pose MPJPE improvements are 13.3% on cooking (229 vs 264), 7.5% on bike repair (372 vs 402), and 22.2% on health (172 vs 221). Only the bike-repair location setting exceeds 30%. The claim should be qualified, e.g., 'up to 32% relative gain,' or the precise settings should be named.
- [Sec. 4, Table 1] No error bars, standard deviations, confidence intervals, or significance tests are reported for any metric. The text repeatedly states that FICTION 'significantly outperforms' all baselines, but the margin on health location PR-AUC is only 1.5 points (12.7 vs 11.2), and on pose PA-MPJPE the differences are 4-6 mm. Given that the ground-truth interactions are produced by an automatic pipeline (Detic, WHAM, Llama-3.1), a few runs with different seeds or a bootstrap over test episodes would be needed to establish that these gaps are not noise.
- [Sec. 3.3, Sec. 4 (metrics)] The pose distribution is evaluated by sampling N=5 poses and selecting the one closest to ground truth (MPJPE and PA-MPJPE). This best-of-N protocol rewards diversity without penalizing implausible samples, so the reported numbers can be artificially low if the model produces a broad distribution. The paper should also report the average error over all samples, or a coverage/frequency metric over the interaction locations, so the reader can assess the quality of the full distribution rather than only the closest sample.
- [Sec. 3.3, Appendix B] The dataset curation pipeline defines an interaction as hands inside an object's 3D bounding box plus an LLM-based narration match. The paper does not report the precision/recall of this automatic pipeline against a human-annotated subset, nor the rate of agreement between the geometric cue and the narration cue. Since the benchmark is being released, a validation of the interaction annotation quality is important for its credibility and for interpreting the results.
minor comments (6)
- [Supplementary Appendix E] The sentence 'Note that this simplification does not affect the curated dataset quality, since we use narrations from Ego-Exo4D as an additional signal' is a non-sequitur: the static point-cloud assumption is about the model input, not about dataset quality, and the narrations do not compensate for future object locations being visible in S.
- [Supplementary Appendix B] There is a typo: 'we do not over-emphaisze on the hands' should be 'emphasize.'
- [Throughout] Inconsistent spacing appears in the model name ('FI CTION' in section headings and Figure 2) and in baseline names ('V oxFormer' in the baselines paragraph and 'CV AE' in Section 3.2); consider unifying these.
- [Sec. 3.1, Eq. (1)] Equation (1) defines Fo as returning a set of points x in R^3, but Section 3.2 and the loss describe a binary voxel grid output; please align the notation so it is clear whether the output is a set of points or a voxel map.
- [Sec. 3.2, Figure 2] Figure 2 does not indicate the temporal cutoff tau_o in the observation inputs, nor does it show which inputs are used for the location decoder versus the CVAE; adding these annotations would improve readability and prevent the temporal-cutoff ambiguity from persisting in the figure.
- [Sec. 3.4, Sec. 4] The paper does not specify how many random seeds were used for training or whether the reported results are averaged over seeds; please state this for all models.
Circularity Check
No circularity in the derivation chain; the Appendix E static-point-cloud statement is a data-leakage risk, not a circular reduction.
full rationale
The paper's derivation is a supervised learning benchmark, not a deductive chain. Equation (1) defines the prediction target as future interaction locations given the observation video V0:τo and the 3D locations P; the ground-truth voxel mask L is a future subset of the object placement map S, and the model is trained with binary cross-entropy against that mask and evaluated on held-out takes from EgoExo4D, so the reported gains (e.g., PR-AUC 21.0 vs 16.9 on cooking) are empirical and not forced by construction. The input object map is a legitimate conditioning modality because the task is to select which of the known objects will be touched, not to invent object positions from scratch. The baselines are external or properly adapted methods (OCT, 4D-Humans, T2M-GPT, VoxFormer, OccFormer, HierVL); self-citations to EgoExo4D [38] and HierVL [5] are data and baseline sources, not load-bearing theorems. The one in-scope limitation is Appendix E: "We also assume a static point cloud when creating the dataset, while in practice, the object location can change with time. It is possible to use 3D information only from the last time segment for improving the spatial input to the model, we do not consider this case for the ease of the I/O." This indicates S may aggregate object detections across the full take, which would let future-object knowledge enter the observation and inflate the reported advantage over baselines that lack S; however, that is a temporal leakage/validity concern about the benchmark, not a case where a prediction reduces to its inputs by definition or where a fitted parameter is renamed as a prediction. No circularity under the seven enumerated patterns is present.
Assumptions & free parameters
free parameters (6)
- Voxel grid resolution N =
16
- Observation time tau_o =
30s
- Future horizon tau_f =
180s
- Anticipation buffer tau_a =
5s
- DBSCAN distance threshold =
50cm
- DBSCAN min cluster points =
100
assumptions (5)
- domain assumption The 3D scene representation S (object placements) is available at test time and is constructed without a stated temporal restriction from the video.
- ad hoc to paper Interaction ground truth is defined by the automatic pipeline: hands inside detected object bounding boxes plus Llama-3.1 narrative matching.
- domain assumption The environment is static; object placements do not change over time.
- domain assumption One actor per video; multi-person scenarios are out of scope.
- domain assumption The off-the-shelf encoders EgoVLPv2 and WHAM provide reliable video and pose representations.
Cite this review
Pith. "Pith review of FIction: 4D Future Interaction Prediction from Video." pith.science (2026). https://pith.science/paper/GNJEASGU
@misc{pith2026241200932,
author = {Pith},
title = {Pith review of: FIction: 4D Future Interaction Prediction from Video},
year = {2026},
howpublished = {\url{https://pith.science/paper/GNJEASGU}},
note = {Machine review of arXiv:2412.00932}
}
read the original abstract
Anticipating how a person will interact with objects in an environment is essential for activity understanding, but existing methods are limited to the 2D space of video frames-capturing physically ungrounded predictions of "what" and ignoring the "where" and "how". We introduce FIction for 4D future interaction prediction from videos. Given an input video of a human activity, the goal is to predict which objects at what 3D locations the person will interact with in the next time period (e.g., cabinet, fridge), and how they will execute that interaction (e.g., poses for bending, reaching, pulling). Our novel model FIction fuses the past video observation of the person's actions and their environment to predict both the "where" and "how" of future interactions. Through comprehensive experiments on a variety of activities and real-world environments in EgoExo4D, we show that our proposed approach outperforms prior autoregressive and (lifted) 2D video models substantially, with more than 30% relative gains.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
When will you do what?-anticipating temporal occurrences of activities
Yazan Abu Farha, Alexander Richard, and Juergen Gall. When will you do what?-anticipating temporal occurrences of activities. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5343–5352,
-
[2]
A spatio-temporal transformer for 3d human mo- tion prediction
Emre Aksan, Manuel Kaufmann, Peng Cao, and Otmar Hilliges. A spatio-temporal transformer for 3d human mo- tion prediction. In 2021 International Conference on 3D Vision (3DV), pages 565–574. IEEE, 2021. 3
2021
-
[3]
Zero experience required: Plug & play modu- lar transfer learning for semantic visual navigation
Ziad Al-Halah, Santhosh Kumar Ramakrishnan, and Kris- ten Grauman. Zero experience required: Plug & play modu- lar transfer learning for semantic visual navigation. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17031–17041, 2022. 1
2022
-
[4]
On evaluation of embodied navigation agents
Peter Anderson, Angel Chang, Devendra Singh Chap- lot, Alexey Dosovitskiy, Saurabh Gupta, Vladlen Koltun, Jana Kosecka, Jitendra Malik, Roozbeh Mottaghi, Manolis Savva, et al. On evaluation of embodied navigation agents. arXiv preprint arXiv:1807.06757, 2018. 1
arXiv 2018
-
[5]
Hiervl: Learning hierarchical video- language embeddings
Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, and Kristen Grauman. Hiervl: Learning hierarchical video- language embeddings. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 23066–23078, 2023. 1, 2, 6, 8
2023
-
[6]
ExpertAF: Ex- pert actionable feedback from video
Kumar Ashutosh, Tushar Nagarajan, Georgios Pavlakos, Kris Kitani, and Kristen Grauman. ExpertAF: Ex- pert actionable feedback from video. arXiv preprint arXiv:2408.00672, 2024. 1
arXiv 2024
-
[7]
Video-mined task graphs for keystep recognition in instructional videos
Kumar Ashutosh, Santhosh Kumar Ramakrishnan, Tri- antafyllos Afouras, and Kristen Grauman. Video-mined task graphs for keystep recognition in instructional videos. Advances in Neural Information Processing Systems , 36,
-
[8]
Affordances from human videos as a versatile representation for robotics
Shikhar Bahl, Russell Mendonca, Lili Chen, Unnat Jain, and Deepak Pathak. Affordances from human videos as a versatile representation for robotics. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13778–13790, 2023. 1, 3
2023
Show all 119 references
-
[9]
Objectnav revisited: On evaluation of embodied agents navigating to objects
Dhruv Batra, Aaron Gokaslan, Aniruddha Kembhavi, Olek- sandr Maksymets, Roozbeh Mottaghi, Manolis Savva, Alexander Toshev, and Erik Wijmans. Objectnav revisited: On evaluation of embodied agents navigating to objects. arXiv preprint arXiv:2006.13171, 2020. 1
2006 arXiv
-
[10]
Is space-time attention all you need for video understanding? In ICML, page 4, 2021
Gedas Bertasius, Heng Wang, and Lorenzo Torresani. Is space-time attention all you need for video understanding? In ICML, page 4, 2021. 4
2021
-
[11]
Procedure planning in instructional videos via contextual modeling and model- based policy learning
Jing Bi, Jiebo Luo, and Chenliang Xu. Procedure planning in instructional videos via contextual modeling and model- based policy learning. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision , pages 15611– 15620, 2021. 2
2021
-
[12]
Long-term human mo- tion prediction with scene context
Zhe Cao, Hang Gao, Karttikeya Mangalam, Qi-Zhi Cai, Minh V o, and Jitendra Malik. Long-term human mo- tion prediction with scene context. In Computer Vision– ECCV 2020: 16th European Conference, Glasgow, UK, Au- gust 23–28, 2020, Proceedings, Part I 16 , pages 387–404. Springe...
2020
-
[13]
Quo vadis, action recognition? a new model and the kinetics dataset
Joao Carreira and Andrew Zisserman. Quo vadis, action recognition? a new model and the kinetics dataset. In pro- ceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 6299–6308, 2017. 2
2017
-
[14]
Procedure planning in instructional videos
Chien-Yi Chang, De-An Huang, Danfei Xu, Ehsan Adeli, Li Fei-Fei, and Juan Carlos Niebles. Procedure planning in instructional videos. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part XI, pages 334–350. Springer, 2020. 2
2020
-
[15]
Expressive whole-body con- trol for humanoid robots
Xuxin Cheng, Yandong Ji, Junming Chen, Ruihan Yang, Ge Yang, and Xiaolong Wang. Expressive whole-body con- trol for humanoid robots. arXiv preprint arXiv:2402.16796,
-
[16]
Context-aware human motion prediction
Enric Corona, Albert Pumarola, Guillem Alenya, and Francesc Moreno-Noguer. Context-aware human motion prediction. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 6992–7001,
-
[17]
Enrichme: Per- ception and interaction of an assistive robot for the elderly at home
Serhan Cos ¸ar, Manuel Fernandez-Carmona, Roxana Agrig- oroaie, Jordi Pages, Franc ¸ois Ferland, Feng Zhao, Shigang Yue, Nicola Bellotto, and Adriana Tapus. Enrichme: Per- ception and interaction of an assistive robot for the elderly at home. International Journal of Social Ro...
2020
-
[18]
Rescaling egocentric vision: collection, pipeline and chal- lenges for epic-kitchens-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Rescaling egocentric vision: collection, pipeline and chal- lenges for epic-kitchens-100. International Journa...
2022
-
[19]
3d affordancenet: A benchmark for visual object affordance understanding
Shengheng Deng, Xun Xu, Chaozheng Wu, Ke Chen, and Kui Jia. 3d affordancenet: A benchmark for visual object affordance understanding. In proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 1778–1787, 2021. 2
2021
-
[20]
Cg-hoi: Contact-guided 3d human-object interaction generation
Christian Diller and Angela Dai. Cg-hoi: Contact-guided 3d human-object interaction generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 19888–19901, 2024. 2
2024
-
[21]
The llama 3 herd of models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Ab- hishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al. The llama 3 herd of models. arXiv preprint arXiv:2407.21783,
-
[22]
Simultaneous local- ization and mapping: part i
Hugh Durrant-Whyte and Tim Bailey. Simultaneous local- ization and mapping: part i. IEEE robotics & automation magazine, 13(2):99–110, 2006. 1, 6
2006
-
[23]
Flow graph to video grounding for weakly-supervised multi-step localization
Nikita Dvornik, Isma Hadji, Hai Pham, Dhaivat Bhatt, Brais Martinez, Afsaneh Fazly, and Allan D Jepson. Flow graph to video grounding for weakly-supervised multi-step localization. In Computer Vision–ECCV 2022: 17th Eu- ropean Conference, Tel Aviv, Israel, October 23–27, 2022,...
2022
-
[24]
Tokenhmr: Advancing human mesh re- 9 covery with a tokenized pose representation
Sai Kumar Dwivedi, Yu Sun, Priyanka Patel, Yao Feng, and Michael J Black. Tokenhmr: Advancing human mesh re- 9 covery with a tokenized pose representation. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1323–1333, 2024. 3
2024
-
[25]
Project aria: A new tool for egocentric multi-modal ai research
Jakob Engel, Kiran Somasundaram, Michael Goesele, Al- bert Sun, Alexander Gamino, Andrew Turner, Arjang Talat- tof, Arnie Yuan, Bilal Souti, Brighid Meredith, et al. Project aria: A new tool for egocentric multi-modal ai research. arXiv preprint arXiv:2308.13561, 2023. 1, 5
2023 arXiv
-
[26]
A density-based algorithm for discovering clusters in large spatial databases with noise
Martin Ester, Hans-Peter Kriegel, J ¨org Sander, Xiaowei Xu, et al. A density-based algorithm for discovering clusters in large spatial databases with noise. In kdd, pages 226–231,
-
[27]
Slowfast networks for video recognition
Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. Slowfast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision, pages 6202–6211, 2019. 2
2019
-
[28]
Yao Feng, Jing Lin, Sai Kumar Dwivedi, Yu Sun, Priyanka Patel, and Michael J. Black. Chatpose: Chatting about 3d human pose, 2024. 1
2024
-
[29]
Rolling- unrolling lstms for action anticipation from first-person video
Antonino Furnari and Giovanni Maria Farinella. Rolling- unrolling lstms for action anticipation from first-person video. IEEE transactions on pattern analysis and machine intelligence, 43(11):4021–4036, 2020. 2
2020
-
[30]
Next-active-object predic- tion from egocentric videos
Antonino Furnari, Sebastiano Battiato, Kristen Grauman, and Giovanni Maria Farinella. Next-active-object predic- tion from egocentric videos. Journal of Visual Communi- cation and Image Representation, 49:401–411, 2017. 1
2017
-
[32]
Red: Re- inforced encoder-decoder networks for action anticipation
Jiyang Gao, Zhenheng Yang, and Ram Nevatia. Red: Re- inforced encoder-decoder networks for action anticipation. arXiv preprint arXiv:1707.04818, 2017
2017 arXiv
-
[33]
Anticipative video transformer
Rohit Girdhar and Kristen Grauman. Anticipative video transformer. In Proceedings of the IEEE/CVF Inter- national Conference on Computer Vision , pages 13505– 13515, 2021. 1, 2, 3
2021
-
[34]
Omni- vore: A single model for many visual modalities
Rohit Girdhar, Mannat Singh, Nikhila Ravi, Laurens van der Maaten, Armand Joulin, and Ishan Misra. Omni- vore: A single model for many visual modalities. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16102–16112, 2022. 2
2022
-
[35]
Humans in 4d: Re- constructing and tracking humans with transformers
Shubham Goel, Georgios Pavlakos, Jathushan Rajasegaran, Angjoo Kanazawa, and Jitendra Malik. Humans in 4d: Re- constructing and tracking humans with transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 14783–14794, 2023. 1, 2, 3, 5, 6, 8
2023
-
[36]
Con- tactopt: Optimizing contact to improve grasps
Patrick Grady, Chengcheng Tang, Christopher D Twigg, Minh V o, Samarth Brahmbhatt, and Charles C Kemp. Con- tactopt: Optimizing contact to improve grasps. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1471–1481, 2021. 2
2021
-
[37]
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. In Proceedings of the IEEE/CVF Conference on Computer Vi- sio...
2022
-
[38]
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. In CVPR, 2024....
2024
-
[39]
Lvis: A dataset for large vocabulary instance segmentation
Agrim Gupta, Piotr Dollar, and Ross Girshick. Lvis: A dataset for large vocabulary instance segmentation. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5356–5364, 2019. 4, 5, 6, 8, 1
2019
-
[40]
Resolving 3d human pose ambiguities with 3d scene constraints
Mohamed Hassan, Vasileios Choutas, Dimitrios Tzionas, and Michael J Black. Resolving 3d human pose ambiguities with 3d scene constraints. In Proceedings of the IEEE/CVF international conference on computer vision , pages 2282– 2292, 2019. 3
2019
-
[41]
Stochas- tic scene-aware motion prediction
Mohamed Hassan, Duygu Ceylan, Ruben Villegas, Jun Saito, Jimei Yang, Yi Zhou, and Michael J Black. Stochas- tic scene-aware motion prediction. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 11374–11384, 2021. 1, 3
2021
-
[42]
Synthesizing phys- ical character-scene interactions
Mohamed Hassan, Yunrong Guo, Tingwu Wang, Michael Black, Sanja Fidler, and Xue Bin Peng. Synthesizing phys- ical character-scene interactions. In ACM SIGGRAPH 2023 Conference Proceedings, pages 1–9, 2023. 3
2023
-
[43]
Learning human- to-humanoid real-time whole-body teleoperation
Tairan He, Zhengyi Luo, Wenli Xiao, Chong Zhang, Kris Kitani, Changliu Liu, and Guanya Shi. Learning human- to-humanoid real-time whole-body teleoperation. arXiv preprint arXiv:2403.04436, 2024. 1
2024 arXiv
-
[44]
Diffusion- based generation, optimization, and planning in 3d scenes
Siyuan Huang, Zan Wang, Puhao Li, Baoxiong Jia, Tengyu Liu, Yixin Zhu, Wei Liang, and Song-Chun Zhu. Diffusion- based generation, optimization, and planning in 3d scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16750–16761, 2023. 1
2023
-
[45]
Human3.6m: Large scale datasets and pre- dictive methods for 3d human sensing in natural environ- ments
Catalin Ionescu, Dragos Papava, Vlad Olaru, and Cristian Sminchisescu. Human3.6m: Large scale datasets and pre- dictive methods for 3d human sensing in natural environ- ments. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(7):1325–1339, 2014. 2
2014
-
[46]
Hand-object contact consistency reasoning for hu- man grasps generation
Hanwen Jiang, Shaowei Liu, Jiashun Wang, and Xiaolong Wang. Hand-object contact consistency reasoning for hu- man grasps generation. In Proceedings of the IEEE/CVF international conference on computer vision, pages 11107– 11116, 2021. 2
2021
-
[47]
Sym- phonize 3d semantic scene completion with contextual in- stance queries, 2023
Haoyi Jiang, Tianheng Cheng, Naiyu Gao, Haoyang Zhang, Tianwei Lin, Wenyu Liu, and Xinggang Wang. Sym- phonize 3d semantic scene completion with contextual in- stance queries, 2023. 1
2023
-
[48]
R2cnn: Rota- tional region cnn for orientation robust scene text detection
Yingying Jiang, Xiangyu Zhu, Xiaobing Wang, Shuli Yang, Wei Li, Hua Wang, Pei Fu, and Zhenbo Luo. R2cnn: Rota- tional region cnn for orientation robust scene text detection. arXiv preprint arXiv:1706.09579, 2017. 5
2017 arXiv
-
[49]
Segment anything
Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer 10 Whitehead, Alexander C Berg, Wan-Yen Lo, et al. Segment anything. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 4015–4026, 2023. 5
2023
-
[50]
Interactive object segmentation in 3d point clouds
Theodora Kontogianni, Ekin Celikkan, Siyu Tang, and Konrad Schindler. Interactive object segmentation in 3d point clouds. In 2023 IEEE International Conference on Robotics and Automation (ICRA), pages 2891–2897. IEEE,
2023
-
[51]
A hi- erarchical representation for future action prediction
Tian Lan, Tsung-Chuan Chen, and Silvio Savarese. A hi- erarchical representation for future action prediction. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part III 13, pages 689–704. Springer, 2014. 2
2014
-
[52]
Llava-med: Training a large language-and-vision assistant for biomedicine in one day
Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. Llava-med: Training a large language-and-vision assistant for biomedicine in one day. arXiv preprint arXiv:2306.00890, 2023. 4
2023 arXiv
-
[53]
Uniformer: Unifying convolution and self-attention for visual recogni- tion
Kunchang Li, Yali Wang, Junhao Zhang, Peng Gao, Guan- glu Song, Yu Liu, Hongsheng Li, and Yu Qiao. Uniformer: Unifying convolution and self-attention for visual recogni- tion. arXiv preprint arXiv:2201.09450, 2022. 2
2022 arXiv
-
[54]
Mvitv2: Improved multiscale vision transform- ers for classification and detection
Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Man- galam, Bo Xiong, Jitendra Malik, and Christoph Feicht- enhofer. Mvitv2: Improved multiscale vision transform- ers for classification and detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Rec...
2022
-
[55]
Alvarez, Sanja Fidler, Chen Feng, and Anima Anandkumar
Yiming Li, Zhiding Yu, Christopher Choy, Chaowei Xiao, Jose M. Alvarez, Sanja Fidler, Chen Feng, and Anima Anandkumar. V oxformer: Sparse voxel transformer for camera-based 3d semantic scene completion. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern...
2023
-
[56]
Egocen- tric video-language pretraining
Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Difei Gao, Rongcheng Tu, Wenzhe Zhao, Weijie Kong, et al. Egocen- tric video-language pretraining. In NeurIPS, 2022. 2
2022
-
[57]
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedin...
2014
-
[58]
Learn- ing to recognize procedural activities with distant supervi- sion
Xudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach, Shih-Fu Chang, and Lorenzo Torresani. Learn- ing to recognize procedural activities with distant supervi- sion. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, pages 13853–13863,
-
[59]
Deep patch visual slam
Lahav Lipson, Zachary Teed, and Jia Deng. Deep patch visual slam. In European Conference on Computer Vision , pages 424–440. Springer, 2025. 1, 6
2025
-
[60]
Visual instruction tuning
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning. arXiv preprint arXiv:2304.08485, 2023. 4
2023 arXiv
-
[61]
Fore- casting human-object interaction: joint prediction of motor attention and actions in first person video
Miao Liu, Siyu Tang, Yin Li, and James M Rehg. Fore- casting human-object interaction: joint prediction of motor attention and actions in first person video. In Computer Vision–ECCV 2020: 16th European Conference, Glasgow, UK, August 23–28, 2020, Proceedings, Part I 16 , pages...
2020
-
[62]
Joint hand motion and interaction hotspots prediction from egocentric videos
Shaowei Liu, Subarna Tripathi, Somdeb Majumdar, and Xi- aolong Wang. Joint hand motion and interaction hotspots prediction from egocentric videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3282–3292, 2022. 3
2022
-
[63]
Joint hand motion and interaction hotspots prediction from egocentric videos
Shaowei Liu, Subarna Tripathi, Somdeb Majumdar, and Xi- aolong Wang. Joint hand motion and interaction hotspots prediction from egocentric videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 3282–3292, 2022. 1, 2, 6, 8
2022
-
[64]
Matthew Loper, Naureen Mahmood, Javier Romero, Ger- ard Pons-Moll, and Michael J. Black. Smpl: a skinned multi-person linear model. ACM Trans. Graph. , 34(6),
-
[65]
Multimodal sense-informed forecasting of 3d human motions
Zhenyu Lou, Qiongjie Cui, Haofan Wang, Xu Tang, and Hong Zhou. Multimodal sense-informed forecasting of 3d human motions. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition , pages 2144–2154, 2024. 3
2024
-
[66]
Amass: Archive of motion capture as surface shapes
Naureen Mahmood, Nima Ghorbani, Nikolaus F Troje, Gerard Pons-Moll, and Michael J Black. Amass: Archive of motion capture as surface shapes. In Proceedings of the IEEE/CVF international conference on computer vi- sion, pages 5442–5451, 2019. 2
2019
-
[67]
Contact-aware human motion forecasting
Wei Mao, Richard I Hartley, Mathieu Salzmann, et al. Contact-aware human motion forecasting. Advances in Neural Information Processing Systems , 35:7356–7367,
-
[68]
On human motion prediction using recurrent neural networks
Julieta Martinez, Michael J Black, and Javier Romero. On human motion prediction using recurrent neural networks. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2891–2900, 2017. 3
2017
-
[69]
Intention-conditioned long-term human egocentric action forecasting@ ego4d challenge 2022
Esteve Valls Mascaro, Hyemin Ahn, and Dongheui Lee. Intention-conditioned long-term human egocentric action forecasting@ ego4d challenge 2022. arXiv preprint arXiv:2207.12080, 2022. 2
2022 arXiv
-
[70]
Howto100m: Learning a text-video embedding by watch- ing hundred million narrated video clips
Antoine Miech, Dimitri Zhukov, Jean-Baptiste Alayrac, Makarand Tapaswi, Ivan Laptev, and Josef Sivic. Howto100m: Learning a text-video embedding by watch- ing hundred million narrated video clips. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pa...
2019
-
[71]
End-to-end learning of visual representations from uncurated instruc- tional videos
Antoine Miech, Jean-Baptiste Alayrac, Lucas Smaira, Ivan Laptev, Josef Sivic, and Andrew Zisserman. End-to-end learning of visual representations from uncurated instruc- tional videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages...
2020
-
[72]
Nerf: Representing scenes as neural radiance fields for view syn- 11 thesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- 11 thesis. Communications of the ACM , 65(1):99–106, 2021. 3
2021
-
[73]
Graspit! a versatile simulator for robotic grasping
Andrew T Miller and Peter K Allen. Graspit! a versatile simulator for robotic grasping. IEEE Robotics & Automa- tion Magazine, 11(4):110–122, 2004. 2
2004
-
[74]
Where2act: From pixels to actions for articulated 3d objects
Kaichun Mo, Leonidas J Guibas, Mustafa Mukadam, Abhi- nav Gupta, and Shubham Tulsiani. Where2act: From pixels to actions for articulated 3d objects. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 6813–6823, 2021. 2
2021
-
[75]
Orb-slam: a versatile and accurate monocular slam system
Raul Mur-Artal, Jose Maria Martinez Montiel, and Juan D Tardos. Orb-slam: a versatile and accurate monocular slam system. IEEE transactions on robotics , 31(5):1147–1163,
-
[77]
Multi-label affordance mapping from egocentric vision
Lorenzo Mur-Labadia, Jose J Guerrero, and Ruben Martinez-Cantin. Multi-label affordance mapping from egocentric vision. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision , pages 5238–5249,
-
[78]
Aff-ttention! affordances and attention models for short- term object interaction anticipation
Lorenzo Mur-Labadia, Ruben Martinez-Cantin, Josechu Guerrero, Giovanni Maria Farinella, and Antonino Furnari. Aff-ttention! affordances and attention models for short- term object interaction anticipation. arXiv preprint arXiv:2406.01194, 2024. 1, 3
2024 arXiv
-
[79]
Nagarajan and K
T. Nagarajan and K. Grauman. Learning affordance land- scapes for interaction exploration in 3d environments. In Proceedings of the Advances on Neural Information Pro- cessing Systems (NeurIPS), 2020. 3
2020
-
[80]
Grounded human-object interaction hotspots from video
Tushar Nagarajan, Christoph Feichtenhofer, and Kristen Grauman. Grounded human-object interaction hotspots from video. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019. 1, 3
2019
-
[81]
Future event prediction: If and when
Lukas Neumann, Andrew Zisserman, and Andrea Vedaldi. Future event prediction: If and when. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pages 0–0, 2019. 2
2019
-
[82]
Using geometry to detect grasps in 3d point clouds
Andreas ten Pas and Robert Platt. Using geometry to detect grasps in 3d point clouds. arXiv preprint arXiv:1501.03100, 2015. 2
2015 arXiv
-
[83]
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3D hands, face, and body from a single image. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pa...
2019
-
[84]
Egovlpv2: Egocentric video-language pre-training with fusion in the backbone
Shraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin, Hardik Shah, Mike Zheng Shou, Rama Chellappa, and Pengchuan Zhang. Egovlpv2: Egocentric video-language pre-training with fusion in the backbone. In Proceedings of the IEEE/CVF International Conference on Computer Vis...
2023
-
[85]
Habitat 3.0: A co-habitat for humans, avatars and robots
Xavier Puig, Eric Undersander, Andrew Szot, Mikael Dal- laire Cote, Tsung-Yen Yang, Ruslan Partsey, Ruta Desai, Alexander William Clegg, Michal Hlavac, So Yeon Min, et al. Habitat 3.0: A co-habitat for humans, avatars and robots. arXiv preprint arXiv:2310.13724, 2023. 3
-
[86]
Vins-mono: A ro- bust and versatile monocular visual-inertial state estimator
Tong Qin, Peiliang Li, and Shaojie Shen. Vins-mono: A ro- bust and versatile monocular visual-inertial state estimator. IEEE transactions on robotics, 34(4):1004–1020, 2018. 5
2018
-
[87]
State-only imitation learning for dexterous manipulation
I Radosavovic, X Wang, L Pinto, and J Malik. State-only imitation learning for dexterous manipulation. InIEEE/RSJ International Conference on Intelligent Robots and Sys- tems, 2021. 1
2021
-
[88]
Poni: Potential functions for objectgoal navigation with interaction-free learning
Santhosh Kumar Ramakrishnan, Devendra Singh Chap- lot, Ziad Al-Halah, Jitendra Malik, and Kristen Grauman. Poni: Potential functions for objectgoal navigation with interaction-free learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition ,...
2022
-
[89]
Humor: 3d human motion model for robust pose estimation
Davis Rempe, Tolga Birdal, Aaron Hertzmann, Jimei Yang, Srinath Sridhar, and Leonidas J Guibas. Humor: 3d human motion model for robust pose estimation. In Proceedings of the IEEE/CVF international conference on computer vi- sion, pages 11488–11499, 2021. 3
2021
-
[90]
First-person activity forecasting with online inverse reinforcement learning
Nicholas Rhinehart and Kris M Kitani. First-person activity forecasting with online inverse reinforcement learning. In Proceedings of the IEEE International Conference on Com- puter Vision, pages 3696–3705, 2017. 2
2017
-
[91]
Structure-from-motion revisited
Johannes L Schonberger and Jan-Michael Frahm. Structure-from-motion revisited. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 4104–4113, 2016. 5
2016
-
[92]
Wham: Reconstructing world-grounded humans with accurate 3d motion
Soyong Shin, Juyong Kim, Eni Halilaj, and Michael J Black. Wham: Reconstructing world-grounded humans with accurate 3d motion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 2070–2080, 2024. 3, 5
2024
-
[93]
Learn- ing structured output representation using deep conditional generative models
Kihyuk Sohn, Honglak Lee, and Xinchen Yan. Learn- ing structured output representation using deep conditional generative models. Advances in neural information pro- cessing systems, 28, 2015. 3, 4, 6, 2
2015
-
[94]
Generating notifications for missing actions: Don’t forget to turn the lights off! In Proceedings of the IEEE International Con- ference on Computer Vision, pages 4669–4677, 2015
Bilge Soran, Ali Farhadi, and Linda Shapiro. Generating notifications for missing actions: Don’t forget to turn the lights off! In Proceedings of the IEEE International Con- ference on Computer Vision, pages 4669–4677, 2015. 1
2015
-
[95]
Segcloud: Semantic segmen- tation of 3d point clouds
Lyne Tchapmi, Christopher Choy, Iro Armeni, JunYoung Gwak, and Silvio Savarese. Segcloud: Semantic segmen- tation of 3d point clouds. In 2017 international conference on 3D vision (3DV), pages 537–547. IEEE, 2017. 5
2017
-
[96]
Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras
Zachary Teed and Jia Deng. Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras. Advances in neural information processing systems , 34:16558–16569,
-
[97]
EPIC Fields: Marrying 3D Geometry and Video Understanding
Vadim Tschernezki, Ahmad Darkhalil, Zhifan Zhu, David Fouhey, Iro Larina, Diane Larlus, Dima Damen, and An- drea Vedaldi. EPIC Fields: Marrying 3D Geometry and Video Understanding. In Proceedings of the Neural Infor- mation Processing Systems (NeurIPS), 2023. 5 12
2023
-
[98]
Transformation-based models of video sequences
Joost Van Amersfoort, Anitha Kannan, Marc’Aurelio Ran- zato, Arthur Szlam, Du Tran, and Soumith Chintala. Transformation-based models of video sequences. arXiv preprint arXiv:1701.08435, 2017. 2
2017 arXiv
-
[99]
Chaoqun Wang, Jiyu Cheng, Wenzheng Chi, Tingfang Yan, and Max Q.-H. Meng. Semantic-aware informative path planning for efficient object search using mobile robot. IEEE Transactions on Systems, Man, and Cybernetics: Sys- tems, 51(8):5230–5243, 2021. 1
2021
-
[100]
Synthesizing long-term 3d human mo- tion and interaction in 3d scenes
Jiashun Wang, Huazhe Xu, Jingwei Xu, Sifei Liu, and Xiaolong Wang. Synthesizing long-term 3d human mo- tion and interaction in 3d scenes. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 9401–9411, 2021. 1, 3
2021
-
[101]
Adaafford: Learn- ing to adapt manipulation affordance for 3d articulated ob- jects via few-shot interactions
Yian Wang, Ruihai Wu, Kaichun Mo, Jiaqi Ke, Qingnan Fan, Leonidas J Guibas, and Hao Dong. Adaafford: Learn- ing to adapt manipulation affordance for 3d articulated ob- jects via few-shot interactions. In European conference on computer vision, pages 90–107. Springer, 2022. 2
2022
-
[102]
Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition
Chao-Yuan Wu, Yanghao Li, Karttikeya Mangalam, Haoqi Fan, Bo Xiong, Jitendra Malik, and Christoph Feichten- hofer. Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion a...
2022
-
[103]
Videoclip: Contrastive pre- training for zero-shot video-text understanding
Hu Xu, Gargi Ghosh, Po-Yao Huang, Dmytro Okhonko, Armen Aghajanyan, Florian Metze, Luke Zettlemoyer, and Christoph Feichtenhofer. Videoclip: Contrastive pre- training for zero-shot video-text understanding. arXiv preprint arXiv:2109.14084, 2021. 2
2021 arXiv
-
[104]
Vitpose: Simple vision transformer baselines for human pose estimation
Yufei Xu, Jing Zhang, Qiming Zhang, and Dacheng Tao. Vitpose: Simple vision transformer baselines for human pose estimation. Advances in Neural Information Process- ing Systems, 35:38571–38584, 2022. 5
2022
-
[105]
Fore- casting of 3d whole-body human poses with grasping ob- jects
Haitao Yan, Qiongjie Cui, Jiexin Xie, and Shijie Guo. Fore- casting of 3d whole-body human poses with grasping ob- jects. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition , pages 1726–1736,
-
[106]
Taco: Token-aware cascade contrastive learning for video-text alignment
Jianwei Yang, Yonatan Bisk, and Jianfeng Gao. Taco: Token-aware cascade contrastive learning for video-text alignment. In Proceedings of the IEEE/CVF Interna- tional Conference on Computer Vision , pages 11562– 11572, 2021. 4
2021
-
[107]
Lemon: Learning 3d human-object in- teraction relation from 2d images
Yuhang Yang, Wei Zhai, Hongchen Luo, Yang Cao, and Zheng-Jun Zha. Lemon: Learning 3d human-object in- teraction relation from 2d images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16284–16295, 2024. 2
2024
-
[108]
Decoupling human and camera motion from videos in the wild
Vickie Ye, Georgios Pavlakos, Jitendra Malik, and Angjoo Kanazawa. Decoupling human and camera motion from videos in the wild. In Proceedings of the IEEE/CVF con- ference on computer vision and pattern recognition , pages 21222–21232, 2023. 3
2023
-
[109]
Affordance diffusion: Synthesizing hand-object inter- actions
Yufei Ye, Xueting Li, Abhinav Gupta, Shalini De Mello, Stan Birchfield, Jiaming Song, Shubham Tulsiani, and Sifei Liu. Affordance diffusion: Synthesizing hand-object inter- actions. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 22...
2023
-
[110]
Dlow: Diversifying latent flows for diverse human motion prediction
Ye Yuan and Kris Kitani. Dlow: Diversifying latent flows for diverse human motion prediction. In Computer Vision– ECCV 2020: 16th European Conference, Glasgow, UK, Au- gust 23–28, 2020, Proceedings, Part IX 16, pages 346–364. Springer, 2020. 3
2020
-
[111]
Generating human motion from textual descrip- tions with discrete representations
Jianrong Zhang, Yangsong Zhang, Xiaodong Cun, Yong Zhang, Hongwei Zhao, Hongtao Lu, Xi Shen, and Ying Shan. Generating human motion from textual descrip- tions with discrete representations. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, ...
2023
-
[112]
Predicting 3d human dynamics from video
Jason Y Zhang, Panna Felsen, Angjoo Kanazawa, and Ji- tendra Malik. Predicting 3d human dynamics from video. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 7114–7123, 2019. 3
2019
-
[113]
We are more than our joints: Predicting how 3d bodies move
Yan Zhang, Michael J Black, and Siyu Tang. We are more than our joints: Predicting how 3d bodies move. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3372–3382, 2021. 3
2021
-
[114]
Occformer: Dual-path transformer for vision-based 3d semantic occu- pancy prediction
Yunpeng Zhang, Zheng Zhu, and Dalong Du. Occformer: Dual-path transformer for vision-based 3d semantic occu- pancy prediction. In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision (ICCV) , pages 9433–9443, 2023. 2, 6, 8
2023
-
[115]
Gimo: Gaze-informed human motion prediction in context
Yang Zheng, Yanchao Yang, Kaichun Mo, Jiaman Li, Tao Yu, Yebin Liu, C Karen Liu, and Leonidas J Guibas. Gimo: Gaze-informed human motion prediction in context. In Eu- ropean Conference on Computer Vision , pages 676–694. Springer, 2022. 3
2022
-
[116]
Learning procedure-aware video represen- tation from instructional videos and their narrations
Yiwu Zhong, Licheng Yu, Yang Bai, Shangwen Li, Xueting Yan, and Yin Li. Learning procedure-aware video represen- tation from instructional videos and their narrations. arXiv preprint arXiv:2303.17839, 2023. 2
2023 arXiv
-
[117]
Procedure-aware pretraining for instructional video understanding
Honglu Zhou, Roberto Mart ´ın-Mart´ın, Mubbasir Kapadia, Silvio Savarese, and Juan Carlos Niebles. Procedure-aware pretraining for instructional video understanding. Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023. 2
2023
-
[118]
Detecting twenty-thousand classes using image-level supervision
Xingyi Zhou, Rohit Girdhar, Armand Joulin, Philipp Kr¨ahenb¨uhl, and Ishan Misra. Detecting twenty-thousand classes using image-level supervision. InEuropean Confer- ence on Computer Vision , pages 350–368. Springer, 2022. 5, 1, 3 13 FI CTION : 4D Future Interaction Prediction...
2022
-
[119]
{rewrite first narration } - answer: (object1, ob- ject2)
-
[120]
{rewrite second narration} - answer: NO INTER- ACTION
-
[121]
Use ‘NO INTERACTION’ and ‘NO MATCHING OBJECTS’ in cases with no in- teraction and matching objects, respectively
{rewrite third narration} - answer: NO MATCH- ING OBJECTS. Use ‘NO INTERACTION’ and ‘NO MATCHING OBJECTS’ in cases with no in- teraction and matching objects, respectively. Here are the numbered narrations: {narrations} C. Details of baseline implementation We introduce the ba...
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.