REVIEW 4 major objections 5 minor 4 cited by
EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World
T0 review · 4 major / 5 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read The paper introduces EgoMe, a large-scale dataset of 7,902 paired videos in which one person watches a demonstration from a third-person view and then imitates it from a first-person view, with gaze, IMU, and language annotations designed…
desk verdict A genuinely new paired observe-then-imitate egocentric dataset with gaze and IMU, but missing label-reliability evidence and split details; deserves conditional acceptance. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the paired observation–imitation video record: one head-mounted-glasses wearer first watches a demonstrator perform an activity (the exocentric video) and then performs the same activity themselves (the egocentric video), with the pair recorded in the same scenario against the same action. The pairing is what carries the argument, because it converts imitation learning into a supervised cross-view problem with a known ground-truth correspondence between seeing and doing. Supporting it are the eye-gaze stream recorded at 120 Hz, IMU signals, separate coarse and fine-grained language captions for each view, and True/False verification labels that mark wrong imitation steps; a pretrained cross-modal retrieval model with an EgoExoNCE contrastive loss supplies aligned exo-ego features used across the benchmarks.
What would settle it
Re-annotate a random sample of video pairs with new annotators and measure agreement on the True/False label and on the localization of wrong steps; if agreement is near chance, or if 'True' pairs show step timestamps that do not align between exo and ego videos, the pairing assumption fails.
Extended reading notes
Core claim
The central claim is that recording the imitator, not just the demonstrator, is what has been missing from ego-exo datasets, and that EgoMe fills that gap. Prior paired datasets capture the same person from two views, while asynchronous collections stitch together videos from different environments; EgoMe instead has one person watch from an exocentric view and then copy the same action from their own egocentric view in the same scene. The paper argues this observation-to-imitation pairing reflects the high-level process robots must reproduce, and it provides the paired videos, gaze and IMU streams, view-specific language annotations, and imitation-verification labels needed to train and test it. The accompanying benchmarks show measurable gaps between single-view, cross-view, and co-training settings, with cross-view transfer consistently harder, which the authors interpret as evidence that the dataset captures a genuine learning challenge rather than a saturated one.
Load-bearing premise
The True/False labels and paired step timestamps assume that the recruited imitators reproduce observed actions faithfully and that annotators agree on what counts as a correct copy, so the egocentric video can be treated as a behavioral replay of the exocentric demonstration.
Editorial extensions
If this is right
- A model can be trained to localize, within an egocentric imitation, exactly which step was performed incorrectly relative to the exocentric demonstration, turning imitation checking into a supervised temporal-localization task.
- Because each pair shares the same scene and action, the dataset supports view-invariant representation learning and cross-view retrieval without the noise of stitching together videos from different environments.
- The paired gaze streams make attention transfer a learnable task: gaze observed during watching can be used to predict where the imitator will look while replaying, and vice versa.
- Fine-grained step captions with timestamps for both views allow procedural understanding methods to be trained and evaluated on cross-view step alignment and reorganization.
- The multimodal IMU and gaze signals give embodied agents a route to correlate body motion and attention between the observing and following phases, beyond RGB alone.
Reading between the lines
- Because only 16.03% of pairs are labeled False, the benchmarks would benefit from class-balanced or per-class results; otherwise a model can score well by predicting 'True' almost everywhere, and the reported aggregate numbers may overstate cross-view skill.
- A natural extension the paper lists but does not run is cross-view video generation; the same-scene paired structure makes EgoMe a direct testbed for synthesizing an imitation view from an observation view.
- The gaze data could be used to test a causal hypothesis the paper leaves implicit: attention during observation predicts which steps the imitator will get wrong, which would connect gaze prediction to imitation-error assessment.
- The dataset's caregiving and assisting activities point toward a human-robot transfer use case where a robot watches a demonstration and then checks its own execution against the demonstration; the verification labels are a ready-made reward signal for that loop.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces EgoMe, a dataset of 7,902 paired videos (15,804 total) recorded with head-mounted glasses worn by an imitator, where each pair contains one video of the imitator observing a demonstrator and one video of the imitator subsequently mimicking the observed activity. The dataset includes eye gaze, IMU data, video-level action labels, coarse- and fine-level language annotations, and imitative verification labels (True/False). The authors also present six benchmarks: exo-ego cross-modal retrieval, exo-ego gaze prediction, imitative action assessment, imitative action reorganization, coarse-level video understanding, and fine-level procedural understanding, each with baseline results on a validation set and a test set reported in the supplementary material. The central claim is that EgoMe is the first large-scale dataset capturing the imitation-learning process from the imitator's egocentric perspective.
Significance. If the core claims are validated, EgoMe would fill a real gap: existing ego-exo datasets record the same person from different viewpoints, whereas EgoMe records the transition from observing another person to performing the observed actions, which is directly relevant for robot imitation learning and embodied AI. The multimodal signals (eye gaze at 120 Hz, IMU, pupil data) and the paired observation-following structure are genuinely useful resources, and the benchmark suite covers a broad set of tasks. The paper also ships a public dataset release link, which enables community use. However, as detailed below, the current manuscript lacks the reliability evidence needed to support the annotation-centric benchmarks, and the 'exocentric' labeling of one of the two views is problematic. The contribution is potentially significant but not yet fully supported.
major comments (4)
- [Section 3.1 and Table 1] The videos labeled 'exocentric' are recorded with the same head-mounted glasses worn by the imitator, so they are egocentric videos of the imitator observing the demonstrator rather than third-person/exocentric videos. This makes the 'Exo+Ego?' column in Table 1 and the cross-view benchmark framing in Sections 4.1, 4.2, 4.5, and 4.6 inaccurate: the two views differ in task phase (observation vs. imitation), not in camera perspective. Please either rename the views (e.g., 'observation view' and 'imitation view') and adapt the benchmark definitions accordingly, or collect genuine third-person exocentric videos. As written, the claim that EgoMe is a cross-view (exo-ego) dataset is not supported by the collection protocol.
- [Section 3.2, Imitative Verification Annotation] The True/False imitation labels are load-bearing for the imitative action assessment benchmark (Section 4.3) and the imitative action reorganization benchmark (Section 4.4), but the paper provides no inter-annotator agreement statistic, no adjudication protocol, and no temporal-alignment statistics between the fine-level step timestamps of the paired exo and ego videos. Section 3.2 asserts that for True pairs 'the number of action timestamps and contents of the exocentric-egocentric video pair should be consistent,' but this consistency is not quantified. Without reporting e.g., Cohen's kappa on True/False labels and agreement on step boundaries (with tolerances), the reader cannot judge whether the supervisory labels are reliable enough to make the reported baseline numbers meaningful measures of imitation ability. Please add these statistics and describe how disagreements were resolved.
- [Section 3.3 and benchmark sections] The train/val/test split is never described. The paper reports 'validate set' results in the main text and 'test set' results in the supplementary material, but gives no split sizes, no indication of whether the split is performed at the video level or at the pair level, and no per-category or per-scenario distribution of the split. Without this information, the baseline numbers cannot be reproduced, and there is a risk of information leakage between paired exo and ego videos if the split is not pair-disjoint. Please specify the exact split protocol and release the split indices with the dataset.
- [Section 3.2, annotation quality control] The claim that 'annotation quality is strictly controlled and checked by at least three individuals and checking code (including syntax checks using LLM models)' is not supported by quantitative evidence. Please report concrete quality-control metrics, such as the percentage of annotations that required correction after the three-person review, the inter-annotator agreement for the language annotations, and the outcome of the LLM-based syntax checks. Without such numbers, the 'strictly controlled' claim remains unverified.
minor comments (5)
- [Section 4.1] The text says 'Table 9 reports the experimental results of the baseline model,' but in the main text the corresponding results are in Table 2. The same mismatch appears for Tables 10-14 in Sections 4.2-4.6. Please correct the cross-references so the main text points to the main-text tables.
- [Section 4.4, Task definition] The 'Ranking Accuracy' metric is not defined; please specify the exact computation (e.g., top-1 accuracy over permuted fragment orders, or Kendall's tau) so that the reported values in Table 5 are interpretable.
- [Section 3.1] The device model is written as '7 invensun aSee Glasses'; please clarify the correct model name and provide the sampling rate and the synchronization method between the video stream and the eye-gaze/IMU streams, since temporal alignment across modalities is essential for the gaze and IMU benchmarks.
- [Table 1] There is a typo in the column header 'annottaions', and the entry for EgoExoLearn under 'Exo+Ego?' uses a checkmark while the text describes its videos as asynchronous; please make the table consistent with the main text description.
- [Supplementary, Figure 8] The caption says 'Distribution of sentence length of global and fine-grained descriptions,' but the main text (Section 3.3) reports average sentence lengths of 32.15 and 18.15 words; please clarify whether the figure shows the same statistics and define which descriptions correspond to each curve.
Circularity Check
No circular derivation: EgoMe is an empirically collected dataset with supervised benchmarks, and no prediction reduces by construction to its own inputs.
full rationale
The paper's central contribution is the EgoMe dataset: physically recorded paired exo-ego videos with eye gaze, IMU, language annotations, and True/False imitation labels. These annotations come from instructed human imitators and manual annotation, not from the benchmark models. The benchmark tasks (cross-modal retrieval, gaze prediction, imitative action assessment/reorganization, and video understanding) are standard supervised evaluations on held-out splits; even though the authors design the tasks and baselines, this is ordinary benchmark practice rather than circular derivation. The fine-level annotation rule that consistent exo-ego steps are expected for correct imitation is a labeling convention, not a fitted parameter that is later renamed as a prediction. The one self-citation ([57], used only as an example of prior single-view video captioning work) is not load-bearing. The absence of inter-annotator agreement or temporal alignment statistics is a data-quality and validity concern, not evidence that the dataset's claims are equivalent to their inputs. No quoted reduction of a stated result to the paper's own definitions or fitted values could be exhibited, so the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption The recorded pairs are behaviorally aligned: the imitator genuinely observes the exocentric demonstration and then attempts to follow it.
- domain assumption Human annotations, especially the imitative verification labels and step timestamps, are accurate and consistent across annotators.
- domain assumption Video encoders pretrained on Ego4D transfer well enough to EgoMe to serve as meaningful features for cross-view tasks.
Cite this review
Pith. "Pith review of EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World." pith.science (2026). https://pith.science/paper/LCYOVNLZ
@misc{pith2026250119061,
author = {Pith},
title = {Pith review of: EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World},
year = {2026},
howpublished = {\url{https://pith.science/paper/LCYOVNLZ}},
note = {Machine review of arXiv:2501.19061}
}
read the original abstract
In human imitation learning, the imitator typically take the egocentric view as a benchmark, naturally transferring behaviors observed from an exocentric view to their owns, which provides inspiration for researching how robots can more effectively imitate human behavior. However, current research primarily focuses on the basic alignment issues of ego-exo data from different cameras, rather than collecting data from the imitator's perspective, which is inconsistent with the high-level cognitive process. To advance this research, we introduce a novel large-scale egocentric dataset, called EgoMe, which towards following the process of human imitation learning via the imitator's egocentric view in the real world. Our dataset includes 7902 paired exo-ego videos (totaling15804 videos) spanning diverse daily behaviors in various real-world scenarios. For each video pair, one video captures an exocentric view of the imitator observing the demonstrator's actions, while the other captures an egocentric view of the imitator subsequently following those actions. Notably, EgoMe uniquely incorporates exo-ego eye gaze, other multi-modal sensor IMU data and different-level annotations for assisting in establishing correlations between observing and imitating process. We further provide a suit of challenging benchmarks for fully leveraging this data resource and promoting the robot imitation learning research. Extensive analysis demonstrates significant advantages over existing datasets. Our EgoMe dataset and benchmarks are available at https://huggingface.co/datasets/HeqianQiu/EgoMe.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 4 Pith papers
-
CoMind: Understanding Collaborative Human Activity from Multiple Minds and Views
CoMind releases 41 h of synchronized multi-view cooking collaboration with social-cue annotations and three ToM-oriented benchmarks on which current VLMs score poorly until fine-tuned.
-
Exo2EgoPose: Leveraging Exocentric Demonstrations for Vision-Language guided Egocentric 3D Hand Pose Forecasting
A transformer-based hand-pose forecaster that reconstructs exocentric video features at video and frame levels from egocentric inputs, then uses them to modulate egocentric pose queries, beating re-implemented baselin...
-
RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
A unified transformer model generates language-and-image planning sequences for embodied tasks, and shows real-robot manipulation without large-scale action pretraining.
-
Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
A comprehensive review of cross-view video understanding that uses both first-person and third-person cameras, organized into a three-direction taxonomy with a dataset catalog and future research gaps.
Reference graph
Works this paper leans on
-
[1]
Tsp: Temporally-sensitive pretraining of video encoders for localization tasks
Humam Alwassel, Silvio Giancola, and Bernard Ghanem. Tsp: Temporally-sensitive pretraining of video encoders for localization tasks. In Proceedings of the IEEE/CVF Inter- national Conference on Computer Vision, pages 3173–3183,
-
[2]
Hiervl: Learning hierarchical video- language embeddings
Kumar Ashutosh, Rohit Girdhar, Lorenzo Torresani, and Kristen Grauman. Hiervl: Learning hierarchical video- language embeddings. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 23066–23078, 2023. 1
work page 2023
-
[3]
My view is the best view: Procedure learning from egocentric videos
Siddhant Bansal, Chetan Arora, and CV Jawahar. My view is the best view: Procedure learning from egocentric videos. In European Conference on Computer Vision, pages 657–675. Springer, 2022. 2
work page 2022
-
[4]
Activitynet: A large-scale video benchmark for human activity understanding
Fabian Caba Heilbron, Victor Escorcia, Bernard Ghanem, and Juan Carlos Niebles. Activitynet: A large-scale video benchmark for human activity understanding. In Proceed- ings of the ieee conference on computer vision and pattern recognition, pages 961–970, 2015. 1, 3
work page 2015
-
[5]
Collecting highly paral- lel data for paraphrase evaluation
David Chen and William B Dolan. Collecting highly paral- lel data for paraphrase evaluation. InProceedings of the 49th annual meeting of the association for computational linguis- tics: human language technologies , pages 190–200, 2011. 3
2011
-
[6]
Egoplan-bench: Benchmarking egocentric embodied plan- ning with multimodal large language models
Yi Chen, Yuying Ge, Yixiao Ge, Mingyu Ding, Bohao Li, Rui Wang, Ruifeng Xu, Ying Shan, and Xihui Liu. Egoplan-bench: Benchmarking egocentric embodied plan- ning with multimodal large language models. arXiv preprint arXiv:2312.06722, 2023. 8
arXiv 2023
-
[7]
Sijie Cheng, Kechen Fang, Yangyang Yu, Sicheng Zhou, Bo- hao Li, Ye Tian, Tingguang Li, Lei Han, and Yang Liu. Vide- gothink: Assessing egocentric video understanding capabili- ties for embodied ai.arXiv preprint arXiv:2410.11623, 2024. 8
-
[8]
Egocap and egoformer: First-person image captioning with context fusion
Zhuangzhuang Dai, Vu Tran, Andrew Markham, Niki Trigoni, M Arif Rahman, Lahiru NS Wijayasingha, John Stankovic, and Chen Li. Egocap and egoformer: First-person image captioning with context fusion. Pattern Recognition Letters, 181:50–56, 2024. 8
work page 2024
Show all 68 references
-
[9]
Scaling egocentric vision: The epic-kitchens dataset
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Sanja Fidler, Antonino Furnari, Evangelos Kazakos, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Scaling egocentric vision: The epic-kitchens dataset. In Proceedings of the European conference on comput...
2018
-
[10]
Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100
Dima Damen, Hazel Doughty, Giovanni Maria Farinella, Antonino Furnari, Evangelos Kazakos, Jian Ma, Davide Moltisanti, Jonathan Munro, Toby Perrett, Will Price, et al. Rescaling egocentric vision: Collection, pipeline and chal- lenges for epic-kitchens-100. International Journa...
2022
-
[11]
Guide to the carnegie mellon university multimodal activity (cmu-mmac) database
Fernando De la Torre, Jessica Hodgins, Adam Bargteil, Xavier Martin, Justin Macey, Alex Collado, and Pep Beltran. Guide to the carnegie mellon university multimodal activity (cmu-mmac) database. 2009. 2, 3
2009
-
[12]
Text-guided graph temporal modeling for few-shot video classification
Fuqin Deng, Jiaming Zhong, Nannan Li, Lanhui Fu, Bingchun Jiang, Yi Ningbo, Feng Qi, He Xin, and Tin Lun Lam. Text-guided graph temporal modeling for few-shot video classification. Engineering Applications of Artificial Intelligence, 137:109076, 2024. 3
2024
-
[13]
A survey of embodied ai: From simulators to research tasks
Jiafei Duan, Samson Yu, Hui Li Tan, Hongyuan Zhu, and Cheston Tan. A survey of embodied ai: From simulators to research tasks. IEEE Transactions on Emerging Topics in Computational Intelligence, 6(2):230–244, 2022. 1
2022
-
[14]
The” something something” video database for learning and evaluating visual common sense
Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michal- ski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. The” something something” video database for learning and evaluating visual common sense. In ...
2017
-
[15]
Ego4d: Around the world in 3,000 hours of egocentric video
Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, et al. Ego4d: Around the world in 3,000 hours of egocentric video. InPro- ceedings of the IEEE/CVF Conference on Computer Vision...
2022
-
[16]
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, et al. Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives. In Proceedings...
2024
-
[17]
Egolifter: Open-world 3d seg- mentation for egocentric perception
Qiao Gu, Zhaoyang Lv, Duncan Frost, Simon Green, Julian Straub, and Chris Sweeney. Egolifter: Open-world 3d seg- mentation for egocentric perception. In European Confer- ence on Computer Vision , pages 382–400. Springer, 2025. 2
2025
-
[18]
Egoexolearn: A dataset for bridging asyn- chronous ego-and exo-centric view of procedural activities in real world
Yifei Huang, Guo Chen, Jilan Xu, Mingfang Zhang, Li- jin Yang, Baoqi Pei, Hongjie Zhang, Lu Dong, Yali Wang, Limin Wang, et al. Egoexolearn: A dataset for bridging asyn- chronous ego-and exo-centric view of procedural activities in real world. In Proceedings of the IEEE/CVF Co...
2024
-
[19]
Tgif-qa: Toward spatio-temporal reasoning in visual question answering
Yunseok Jang, Yale Song, Youngjae Yu, Youngjin Kim, and Gunhee Kim. Tgif-qa: Toward spatio-temporal reasoning in visual question answering. In Proceedings of the IEEE con- ference on computer vision and pattern recognition , pages 2758–2766, 2017. 3
2017
-
[20]
Epic- tent: An egocentric video dataset for camping tent assembly
Youngkyoon Jang, Brian Sullivan, Casimir Ludwig, Iain Gilchrist, Dima Damen, and Walterio Mayol-Cuevas. Epic- tent: An egocentric video dataset for camping tent assembly. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019. 2 9
2019
-
[21]
Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities
Baoxiong Jia, Yixin Chen, Siyuan Huang, Yixin Zhu, and Song-chun Zhu. Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities. In European Conference on Computer Vision, pages 767–786. Springer, 2020. 2, 3
2020
-
[22]
Large-scale video classification with convolutional neural networks
Andrej Karpathy, George Toderici, Sanketh Shetty, Thomas Leung, Rahul Sukthankar, and Li Fei-Fei. Large-scale video classification with convolutional neural networks. In Pro- ceedings of the IEEE conference on Computer Vision and Pattern Recognition, pages 1725–1732, 2014. 3
2014
-
[23]
Seed4d: A synthetic ego–exo dynamic 4d data generator, driving dataset and benchmark
Marius K ¨astingsch¨afer, Th ´eo Gieruc, Sebastian Bernhard, Dylan Campbell, Eldar Insafutdinov, Eyvaz Najafli, and Thomas Brox. Seed4d: A synthetic ego–exo dynamic 4d data generator, driving dataset and benchmark. arXiv preprint arXiv:2412.00730, 2024. 3
2024 arXiv
-
[24]
The kinetics hu- man action video dataset
Will Kay, Joao Carreira, Karen Simonyan, Brian Zhang, Chloe Hillier, Sudheendra Vijayanarasimhan, Fabio Viola, Tim Green, Trevor Back, Paul Natsev, et al. The kinetics hu- man action video dataset. arXiv preprint arXiv:1705.06950,
-
[25]
Hmdb: a large video database for human motion recognition
Hildegard Kuehne, Hueihan Jhuang, Est ´ıbaliz Garrote, Tomaso Poggio, and Thomas Serre. Hmdb: a large video database for human motion recognition. In 2011 Inter- national conference on computer vision , pages 2556–2563. IEEE, 2011. 1, 3
2011
-
[26]
H2o: Two hands manipulating objects for first person interaction recognition
Taein Kwon, Bugra Tekin, Jan St ¨uhmer, Federica Bogo, and Marc Pollefeys. H2o: Two hands manipulating objects for first person interaction recognition. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 10138–10148, 2021. 2, 3
2021
-
[27]
In the eye of transformer: Global-local correlation for egocentric gaze estimation
Bolin Lai, Miao Liu, Fiona Ryan, and James Rehg. In the eye of transformer: Global-local correlation for egocentric gaze estimation. British Machine Vision Conference, 2022. 6, 7
2022
-
[28]
In the eye of transformer: Global–local correlation for egocentric gaze estimation and beyond
Bolin Lai, Miao Liu, Fiona Ryan, and James M Rehg. In the eye of transformer: Global–local correlation for egocentric gaze estimation and beyond. International Journal of Com- puter Vision, pages 1–18, 2023. 7
2023
-
[29]
Error detection in egocentric procedural task videos
Shih-Po Lee, Zijia Lu, Zekun Zhang, Minh Hoai, and Ehsan Elhamifar. Error detection in egocentric procedural task videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 18655– 18666, 2024. 3, 7
2024
-
[30]
Dis- covering important people and objects for egocentric video summarization
Yong Jae Lee, Joydeep Ghosh, and Kristen Grauman. Dis- covering important people and objects for egocentric video summarization. In 2012 IEEE conference on computer vi- sion and pattern recognition, pages 1346–1353. IEEE, 2012. 2
2012
-
[31]
Uniformerv2: Unlocking the po- tential of image vits for video understanding
Kunchang Li, Yali Wang, Yinan He, Yizhuo Li, Yi Wang, Limin Wang, and Yu Qiao. Uniformerv2: Unlocking the po- tential of image vits for video understanding. InProceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 1632–1643, 2023. 6
2023
-
[32]
Zeroi2v: Zero-cost adaptation of pre-trained transformers from image to video
Xinhao Li, Yuhan Zhu, and Limin Wang. Zeroi2v: Zero-cost adaptation of pre-trained transformers from image to video. In European Conference on Computer Vision , pages 425–
-
[33]
In the eye of beholder: Joint learning of gaze and actions in first person video
Yin Li, Miao Liu, and James M Rehg. In the eye of beholder: Joint learning of gaze and actions in first person video. In Proceedings of the European conference on computer vision (ECCV), pages 619–635, 2018. 2, 3
2018
-
[34]
Egoexo-fitness: Towards egocentric and exocentric full-body action under- standing
Yuan-Ming Li, Wei-Jin Huang, An-Lan Wang, Ling-An Zeng, Jing-Ke Meng, and Wei-Shi Zheng. Egoexo-fitness: Towards egocentric and exocentric full-body action under- standing. arXiv preprint arXiv:2406.08877, 2024. 2, 3, 4
2024 arXiv
-
[35]
Joint hand motion and interaction hotspots prediction from egocentric videos
Shaowei Liu, Subarna Tripathi, Somdeb Majumdar, and Xi- aolong Wang. Joint hand motion and interaction hotspots prediction from egocentric videos. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 3282–3292, 2022. 1
2022
-
[36]
Hota: A higher order metric for evaluating multi-object tracking
Jonathon Luiten, Aljosa Osep, Patrick Dendorfer, Philip Torr, Andreas Geiger, Laura Leal-Taix´e, and Bastian Leibe. Hota: A higher order metric for evaluating multi-object tracking. International journal of computer vision, 129:548– 578, 2021. 1
2021
-
[37]
Aria everyday activities dataset
Zhaoyang Lv, Nicholas Charron, Pierre Moulon, Alexander Gamino, Cheng Peng, Chris Sweeney, Edward Miller, Huix- uan Tang, Jeff Meissner, Jing Dong, et al. Aria everyday activities dataset. arXiv preprint arXiv:2402.13349, 2024. 2
2024 arXiv
-
[38]
Nymeria: A mas- sive collection of multimodal egocentric daily motion in the wild
Lingni Ma, Yuting Ye, Fangzhou Hong, Vladimir Guzov, Yifeng Jiang, Rowan Postyeni, Luis Pesqueira, Alexander Gamino, Vijay Baiyya, Hyo Jin Kim, et al. Nymeria: A mas- sive collection of multimodal egocentric daily motion in the wild. In European Conference on Computer Vision ,...
2025
-
[39]
Mod- eling context between objects for referring expression under- standing
Varun K Nagaraja, Vlad I Morariu, and Larry S Davis. Mod- eling context between objects for referring expression under- standing. In Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14 , pages 792–807. Springer,
2016
-
[40]
Sensor-augmented egocentric-video captioning with dy- namic modal attention
Katsuyuki Nakamura, Hiroki Ohashi, and Mitsuhiro Okada. Sensor-augmented egocentric-video captioning with dy- namic modal attention. In Proceedings of the 29th ACM International Conference on Multimedia, pages 4220–4229,
-
[41]
Cap- taincook4d: A dataset for understanding errors in procedural activities
Rohith Peddi, Shivvrat Arya, Bharath Challa, Likhitha Pal- lapothula, Akshay Vyas, Jikai Wang, Qifan Zhang, Vasund- hara Komaragiri, Eric Ragan, Nicholas Ruozzi, et al. Cap- taincook4d: A dataset for understanding errors in procedural activities. arXiv preprint arXiv:2312.1455...
2023 arXiv
-
[42]
Detecting activi- ties of daily living in first-person camera views
Hamed Pirsiavash and Deva Ramanan. Detecting activi- ties of daily living in first-person camera views. In 2012 IEEE conference on computer vision and pattern recogni- tion, pages 2847–2854. IEEE, 2012. 2, 3
2012
-
[43]
Svip: Sequence verification for procedures in videos
Yicheng Qian, Weixin Luo, Dongze Lian, Xu Tang, Peilin Zhao, and Shenghua Gao. Svip: Sequence verification for procedures in videos. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 19890–19902, 2022. 7
2022
-
[44]
Egoplan-bench2: A benchmark for multimodal large language model planning in real-world scenarios.arXiv preprint arXiv:2412.04447, 2024
Lu Qiu, Yuying Ge, Yi Chen, Yixiao Ge, Ying Shan, and Xihui Liu. Egoplan-bench2: A benchmark for multimodal large language model planning in real-world scenarios.arXiv preprint arXiv:2412.04447, 2024. 8
2024 arXiv
-
[45]
Learning transferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, 10 Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. In International conference on machine learnin...
2021
-
[46]
The meccano dataset: Understanding human-object interactions from egocentric videos in an industrial-like domain
Francesco Ragusa, Antonino Furnari, Salvatore Livatino, and Giovanni Maria Farinella. The meccano dataset: Understanding human-object interactions from egocentric videos in an industrial-like domain. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer ...
2021
-
[47]
Home action genome: Cooperative compositional action understanding
Nishant Rai, Haofeng Chen, Jingwei Ji, Rishi Desai, Kazuki Kozuka, Shun Ishizaka, Ehsan Adeli, and Juan Carlos Niebles. Home action genome: Cooperative compositional action understanding. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, p...
2021
-
[48]
As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities
Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, and Angela Yao. As- sembly101: A large-scale multi-view video dataset for un- derstanding procedural activities. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern ...
2022
-
[49]
Charades-ego: A large-scale dataset of paired third and first person videos
Gunnar A Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari. Charades-ego: A large-scale dataset of paired third and first person videos. arXiv preprint arXiv:1804.09626, 2018. 2, 3
2018 arXiv
-
[50]
Krishnacam: Using a longitudinal, single-person, egocentric dataset for scene understanding tasks
Krishna Kumar Singh, Kayvon Fatahalian, and Alexei A Efros. Krishnacam: Using a longitudinal, single-person, egocentric dataset for scene understanding tasks. In 2016 IEEE Winter Conference on Applications of Computer Vision (WACV), pages 1–9. IEEE, 2016. 2
2016
-
[51]
Long-form video-language pre- training with multimodal temporal contrastive learning
Yuchong Sun, Hongwei Xue, Ruihua Song, Bei Liu, Huan Yang, and Jianlong Fu. Long-form video-language pre- training with multimodal temporal contrastive learning. Ad- vances in neural information processing systems, 35:38032– 38045, 2022. 1, 3
2022
-
[52]
Clip4caption: Clip for video caption
Mingkang Tang, Zhanyu Wang, Zhenhua Liu, Fengyun Rao, Dian Li, and Xiu Li. Clip4caption: Clip for video caption. In Proceedings of the 29th ACM International Conference on Multimedia, pages 4858–4862, 2021. 8
2021
-
[53]
Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training
Zhan Tong, Yibing Song, Jue Wang, and Limin Wang. Videomae: Masked autoencoders are data-efficient learners for self-supervised video pre-training. Advances in neural information processing systems, 35:10078–10093, 2022. 6
2022
-
[54]
Learning language-visual embedding for movie understanding with natural-language
Atousa Torabi, Niket Tandon, and Leonid Sigal. Learning language-visual embedding for movie understanding with natural-language. arXiv preprint arXiv:1609.08124 , 2016. 3
2016 arXiv
-
[55]
Epic fields: Marrying 3d geometry and video under- standing
Vadim Tschernezki, Ahmad Darkhalil, Zhifan Zhu, David Fouhey, Iro Laina, Diane Larlus, Dima Damen, and Andrea Vedaldi. Epic fields: Marrying 3d geometry and video under- standing. Advances in Neural Information Processing Sys- tems, 36, 2024. 2
2024
-
[56]
Scene-aware ego- centric 3d human pose estimation
Jian Wang, Diogo Luvizon, Weipeng Xu, Lingjie Liu, Kri- pasindhu Sarkar, and Christian Theobalt. Scene-aware ego- centric 3d human pose estimation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13031–13040, 2023. 1
2023
-
[57]
Pos-trends dynamic-aware model for video caption
Lanxiao Wang, Hongliang Li, Heqian Qiu, Qingbo Wu, Fan- man Meng, and King Ngi Ngan. Pos-trends dynamic-aware model for video caption. IEEE Transactions on Circuits and Systems for Video Technology, 32(7):4751–4764, 2021. 8
2021
-
[58]
End-to-end dense video captioning with parallel decoding
Teng Wang, Ruimao Zhang, Zhichao Lu, Feng Zheng, Ran Cheng, and Ping Luo. End-to-end dense video captioning with parallel decoding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 6847– 6857, 2021. 8
2021
-
[59]
Holoassist: an egocen- tric human interaction dataset for interactive ai assistants in the real world
Xin Wang, Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bu- gra Tekin, Felipe Vieira Frujeri, et al. Holoassist: an egocen- tric human interaction dataset for interactive ai assistants in the real world. In Proceedings of the I...
-
[60]
Assistq: Affordance-centric question-driven task completion for ego- centric assistant
Benita Wong, Joya Chen, You Wu, Stan Weixian Lei, Dongxing Mao, Difei Gao, and Mike Zheng Shou. Assistq: Affordance-centric question-driven task completion for ego- centric assistant. In European Conference on Computer Vi- sion, pages 485–501. Springer, 2022. 2
2022
-
[61]
Msr-vtt: A large video description dataset for bridging video and language
Jun Xu, Tao Mei, Ting Yao, and Yong Rui. Msr-vtt: A large video description dataset for bridging video and language. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 5288–5296, 2016. 3
2016
-
[62]
Retrieval-augmented egocentric video captioning
Jilan Xu, Yifei Huang, Junlin Hou, Guo Chen, Yuejie Zhang, Rui Feng, and Weidi Xie. Retrieval-augmented egocentric video captioning. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 13525–13536, 2024. 6
2024
-
[63]
Learning fine- grained view-invariant representations from unpaired ego- exo videos via temporal alignment
Zihui Sherry Xue and Kristen Grauman. Learning fine- grained view-invariant representations from unpaired ego- exo videos via temporal alignment. Advances in Neural In- formation Processing Systems, 36:53688–53710, 2023. 2, 3, 6
2023
-
[64]
Hitea: Hierarchical temporal- aware video-language pre-training
Qinghao Ye, Guohai Xu, Ming Yan, Haiyang Xu, Qi Qian, Ji Zhang, and Fei Huang. Hitea: Hierarchical temporal- aware video-language pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision , pages 15405–15416, 2023. 1, 3
2023
-
[65]
A joint se- quence fusion model for video question answering and re- trieval
Youngjae Yu, Jongseok Kim, and Gunhee Kim. A joint se- quence fusion model for video question answering and re- trieval. In Proceedings of the European conference on com- puter vision (ECCV), pages 471–487, 2018. 3, 6
2018
-
[66]
A survey on deep learning technique for video segmentation
Tianfei Zhou, Fatih Porikli, David J Crandall, Luc Van Gool, and Wenguan Wang. A survey on deep learning technique for video segmentation. IEEE transactions on pattern analysis and machine intelligence, 45(6):7099–7122, 2022. 3 11 EgoMe: A New Dataset and Challenge for Followi...
2022
-
[67]
In addition, each kind of scenario comprises a diverse range of layouts
Dataset Details In our EgoMe dataset, we consider daily human activities and their follow-up learning in real-world environments, therefore we collect data in up to over 41 scenarios includ- ing various classrooms, living rooms, supermarkets, offices, libraries, washrooms, gar...
-
[68]
A shown in Table ??, it can be observed that these results on the test set are a similar experimental results to validate set in our EgoMe dataset
Benchmark Details Similar to benchmark section in formal content paper, we also conduct all experiments on the test set of our EgoMe dataset, including Exo-Ego cross-modal retrieval, Exo-Ego gaze prediction, imitative action assessment or reorganiza- tion, Exo-Ego coarse-level...
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.