REVIEW 3 major objections 6 minor 217 references
Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A first survey unifies first- and third-person video AI.
desk verdict Useful organizing survey of ego-exo video understanding, solid dataset table, but 'comprehensive' claim needs either a search methodology or a softer claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is a three-way taxonomy of research directions—ego-for-exo (egocentric cues improve exocentric tasks), exo-for-ego (exocentric data improves egocentric analysis), and joint learning (both views are used at training and inference together). The survey maps each direction to concrete tasks such as video generation, action understanding, affordance grounding, and cross-view retrieval, and links those tasks back to eight application scenarios. This taxonomy does the work of a framework: it positions every reviewed method, exposes where the literature is thin, and drives the paper's dataset inventory and future-directions discussion.
What would settle it
An independent, systematic literature search with explicit inclusion criteria that surfaces a substantial body of ego-exo work omitted from this survey, or that reveals a major research direction (for instance multi-agent ego-exo collaboration) not covered by the three-way taxonomy, would falsify the survey's central claims of comprehensiveness and completeness.
Extended reading notes
Core claim
The paper's central claim is that cross-view collaboration between egocentric and exocentric video is a coherent field whose progress can be mapped by three research directions: egocentric-for-exocentric, exocentric-for-egocentric, and joint learning. On the paper's own terms, no prior survey has integrated both perspectives, so this review fills that gap by organizing existing methods, identifying eight application domains that would benefit from ego-exo collaboration, and cataloguing datasets that contain both viewpoints. The review concludes that current work clusters around daily-life activity understanding, while tasks such as ego-to-exo video generation, view birdification, skill assessment, and affordance grounding for industrial or surgical tools remain under-explored.
Load-bearing premise
The survey's value rests on its selection of papers and datasets being complete and representative enough that the 'first comprehensive review' claim and the resulting gap analysis are trustworthy.
Editorial extensions
If this is right
- A researcher can place any ego-exo method into one of three categories, making the field's structure explicit for newcomers.
- The gap analysis identifies ego-to-exo video generation, view birdification, and cross-view skill assessment as under-studied tasks worth pursuing.
- Specialized domains such as healthcare, education, and public service lack dedicated ego-exo datasets, so collecting them is a clear next step.
- Because synchronized paired data is expensive, aligning unpaired ego and exo videos via language or retrieval is a promising route to scale.
- Vision-language models that can take both perspectives as input are proposed as a path toward unified cross-view frameworks.
Reading between the lines
- The 'first comprehensive survey' claim is only as strong as the literature search behind it; a reader can probe it by checking whether recent ego-exo works beyond the cited set fit the taxonomy.
- The three-way taxonomy may under-represent multi-agent ego-exo collaboration, which appears in applications like rescue and public service but is not given its own research direction.
- A concrete testable extension of the gap analysis is to benchmark ego-to-exo video generation on the driving and robotics scenarios the paper names, where latency constraints are severe.
- The dataset table offers a way to quantify domain imbalance, for instance by counting ego-exo datasets per application area, which could prioritize future data collection.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a survey of video understanding research that integrates egocentric (first-person) and exocentric (third-person) perspectives. It opens with eight application domains, derives a set of research tasks from them, organizes the literature into three directions (ego-for-exo, exo-for-ego, and joint learning), and provides a table of ego-exo datasets. The paper claims to be the first comprehensive survey of this integration and concludes with a gap analysis and future-work suggestions. A companion GitHub repository is announced.
Significance. If its coverage is representative, the survey is timely and useful: it gives the area a coherent taxonomy, collects many relevant methods across disparate subfields, and the dataset table in Section V is a practical contribution. The paper also ships a companion GitHub repository, which helps researchers track new work. However, the central 'comprehensive' claim and the under-explored task analysis rest on an undocumented literature sample, and a few taxonomically misplaced sections weaken the review's internal consistency. The contribution is therefore not yet fully supported, although the issues appear fixable within the manuscript's scope.
major comments (3)
- [Section I, Section VI, Fig. 4] The paper's central claim that it is the first comprehensive survey of egocentric-exocentric video understanding (Section I) and the gap analysis in Section VI are not backed by a documented literature search. There is no search protocol, inclusion criterion, or comparison with prior or concurrent surveys, so the reader cannot judge whether the works sampled in Sections IV and V are representative. This matters because Section VI draws conclusions about which tasks are 'under-explored' and Fig. 4 is used to state that 'many tasks critical to applications remain under-investigated.' I request an explicit methodology paragraph (databases, query terms, inclusion/exclusion criteria, and screening flow) and, ideally, a recall-style check against a systematic ego-exo query to substantiate the comprehensiveness claim. Without this, the claim remains an unverified external completeness assertion.
- [Section IV.B, Fig. 5] The task 'Remote Drone Teleoperation' is placed under 'Exocentric for Egocentric' video understanding, but the works reviewed there ([132]–[135] and [133]) are human-computer interaction systems that use VR, additional cameras, or AR overlays to provide a pilot with a better exocentric view during teleoperation; they do not use exocentric data to improve egocentric video understanding, which is the definition given at the start of Section IV.B. This inclusion stretches the taxonomy and weakens the review's stated focus on video understanding. Either move these works to an application-oriented subsection or add an explicit justification of how they inform ego-exo video understanding.
- [Section III, Fig. 4] The 'from applications to research tasks' bridge is incomplete. Section IV reviews view birdification, cross-view retrieval, 3D camera localization, egocentric wearer identification, and cross-view human identification, but none of these tasks appears in Fig. 4, which is used in Section III to support the statement that many application-critical tasks are under-investigated. The figure should include all tasks discussed in Section IV, or the text should clearly state that Fig. 4 is a partial mapping; alternatively, the gap analysis should be supported by the full task list in Section IV.
minor comments (6)
- [Figure 2 caption] The caption says Section IV is divided into 'ego for exo, exo for exo, and joint learning'; 'exo for exo' should be 'exo for ego'.
- [Figure 4] There are typos in the figure: 'Camara Localization' should be 'Camera Localization' and 'Healthcar e' should be 'Healthcare'.
- [References] References [75] and [177] are the same paper (Seo et al., 'Multi-view masked world models for visual robotic manipulation'), and references [116] and [170] are the same paper (Jia et al., 'The audio-visual conversational graph'); these duplicates should be unified.
- [Section IV.B, Self-supervised methods] EgoFish3D [114] is described as using an exocentric pose estimator as supervision; this is more accurately characterized as weakly supervised than self-supervised, given the terminology used in the surrounding text.
- [Section V.D] The dataset entry 'ThirdtoFirst [51]' cites a method paper (Li et al., 'Ego-Exo: Transferring visual representations from third-person to first-person videos') rather than a dataset paper; please align the citation with the actual dataset source or rename the entry.
- [Fig. 1] The caption says the citation counts cover 'papers and datasets discussed in Sections IV and V,' but the collection and curation rules for the Google Scholar data are not described; please clarify the inclusion criteria so readers can interpret the growth curve.
Circularity Check
No circular derivation; survey is descriptive and self-contained, with only minor self-citation bias.
full rationale
This is a survey paper with no equations, no fitted parameters, and no predictive claim whose outcome could be forced by construction. The central organizational scheme (ego-for-exo, exo-for-ego, joint learning) is a taxonomy, not a derived result; it is defined by the direction of information flow and then applied to the literature. The novelty claim that 'no survey has yet addressed the integration of both perspectives' is an external literature-coverage assertion supported by references to existing surveys [28]–[32], none of which is authored by the present authors; it is not a self-citation chain used to forbid alternatives. The self-citations are numerous (e.g., [13], [14], [54], [60], [136], [142], [145], [158], [192], [194]), but they appear as entries in the reviewed literature, not as load-bearing premises that define the taxonomy or establish the gap. Fig. 1's citation curve is computed from papers the survey itself discusses, but the citation counts themselves are external Google Scholar data, and the figure is presented as a descriptive trend of that selected set, not as a prediction or proof. The absence of a search protocol or inclusion criteria is a legitimate completeness-risk concern, but it is an external validity issue, not circularity: the survey never claims to derive its completeness from its own contents. No passage satisfies the quoted-reduction threshold required to flag a circular step. A score of 1 reflects the elevated self-citation rate and the self-selected Fig. 1 sample, which create a mild self-referential flavor without making any central claim equivalent to its inputs.
Assumptions & free parameters
assumptions (3)
- domain assumption The three-way taxonomy (ego for exo, exo for ego, joint learning) is the natural and exhaustive organizing principle for cross-view ego-exo work.
- domain assumption The surveyed papers and datasets are accurately described and representative.
- domain assumption Citation counts in Figure 1 from Google Scholar are a meaningful proxy for field growth.
Cite this review
Pith. "Pith review of Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision." pith.science (2026). https://pith.science/paper/KJTJ3EHL
@misc{pith2026250606253,
author = {Pith},
title = {Pith review of: Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision},
year = {2026},
howpublished = {\url{https://pith.science/paper/KJTJ3EHL}},
note = {Machine review of arXiv:2506.06253}
}
read the original abstract
Perceiving the world from both egocentric (first-person) and exocentric (third-person) perspectives is fundamental to human cognition, enabling rich and complementary understanding of dynamic environments. In recent years, allowing the machines to leverage the synergistic potential of these dual perspectives has emerged as a compelling research direction in video understanding. In this survey, we provide a comprehensive review of video understanding from both exocentric and egocentric viewpoints. We begin by highlighting the practical applications of integrating egocentric and exocentric techniques, envisioning their potential collaboration across domains. We then identify key research tasks to realize these applications. Next, we systematically organize and review recent advancements into three main research directions: (1) leveraging egocentric data to enhance exocentric understanding, (2) utilizing exocentric data to improve egocentric analysis, and (3) joint learning frameworks that unify both perspectives. For each direction, we analyze a diverse set of tasks and relevant works. Additionally, we discuss benchmark datasets that support research in both perspectives, evaluating their scope, diversity, and applicability. Finally, we discuss limitations in current works and propose promising future research directions. By synthesizing insights from both perspectives, our goal is to inspire advancements in video understanding and artificial intelligence, bringing machines closer to perceiving the world in a human-like manner. A GitHub repo of related works can be found at https://github.com/ayiyayi/Awesome-Egocentric-and-Exocentric-Vision.
Figures
Figures from the paper (15 more)
Reference graph
Works this paper leans on
-
[132]
Starhopper: A touch interface for remote object-centric drone navigation,
J. Li, R. Balakrishnan, and T. Grossman, “Starhopper: A touch interface for remote object-centric drone navigation,” inProc Graphics Interface, ser. GI 2020, 2020, pp. 317 – 326
2020
-
[135]
Birdviewar: Surroundings-aware remote drone piloting using an augmented third- person perspective,
M. Inoue, K. Takashima, K. Fujita, and Y . Kitamura, “Birdviewar: Surroundings-aware remote drone piloting using an augmented third- person perspective,” inConf Hum Fact Comput Syst Proc, 2023
2023
-
[133]
Drone- augmented human vision: Exocentric control for drones exploring hidden areas,
O. Erat, W. A. Isop, D. Kalkofen, and D. Schmalstieg, “Drone- augmented human vision: Exocentric control for drones exploring hidden areas,”IEEE Trans. Vis. Comput. Graphics, vol. 24, pp. 1437– 1446, 04 2018
2018
-
[1]
The mirror-neuron system,
G. Rizzolatti and L. Craighero, “The mirror-neuron system,”Annual review of neuroscience, vol. 27, pp. 169–92, 02 2004
2004
-
[2]
Actor and observer: Joint modeling of first and third-person videos,
G. A. Sigurdsson, A. Gupta, C. Schmid, A. Farhadi, and K. Alahari, “Actor and observer: Joint modeling of first and third-person videos,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 7396–7404
2018
-
[3]
The epic-kitchens dataset: Collection, challenges and baselines,
D. Damenet al., “The epic-kitchens dataset: Collection, challenges and baselines,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 43, no. 11, pp. 4125–4141, 2021
2021
-
[4]
Ego4d: Around the world in 3,000 hours of egocentric video,
K. Graumanet al., “Ego4d: Around the world in 3,000 hours of egocentric video,”IEEE Trans. Pattern Anal. Mach. Intell., pp. 1–32, 2024
2024
-
[5]
Ego4d goal-step: Toward hierarchical understanding of procedural activities,
Y . Song, E. Byrne, T. Nagarajan, H. Wang, M. Martin, and L. Torresani, “Ego4d goal-step: Toward hierarchical understanding of procedural activities,”Adv. Neural Inform. Process. Syst., vol. 36, 2024
2024
Show all 217 references
-
[6]
Ego-humans: An ego-centric 3d multi-human benchmark,
R. Khirodkar, A. Bansal, L. Ma, R. Newcombe, M. V o, and K. Kitani, “Ego-humans: An ego-centric 3d multi-human benchmark,” inInt. Conf. Comput. Vis., 2023, pp. 19 807–19 819
2023
-
[7]
Identifying first-person camera wearers in third-person videos,
C. Fanet al., “Identifying first-person camera wearers in third-person videos,” inIEEE Conf. Comput. Vis. Pattern Recog., July 2017
2017
-
[8]
Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives,
K. Graumanet al., “Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 19 383–19 400
2024
-
[9]
Hd-epic: A highly-detailed egocentric video dataset,
T. Perrettet al., “Hd-epic: A highly-detailed egocentric video dataset,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2025
2025
-
[10]
Egocentric video-language pretraining,
K. Q. Linet al., “Egocentric video-language pretraining,” inAdv. Neural Inform. Process. Syst., vol. 35, 2022, pp. 7575–7586
2022
-
[11]
Egovlpv2: Egocentric video-language pre-training with fusion in the backbone,
S. Pramanicket al., “Egovlpv2: Egocentric video-language pre-training with fusion in the backbone,” inInt. Conf. Comput. Vis., 2023, pp. 5262–5274
2023
-
[12]
Helping hands: An object- aware ego-centric video recognition model,
C. Zhang, A. Gupta, and A. Zisserman, “Helping hands: An object- aware ego-centric video recognition model,” inInt. Conf. Comput. Vis., 2023, pp. 13 855–13 866
2023
-
[13]
Internvideo-ego4d: A pack of champion solutions to ego4d challenges,
G. Chenet al., “Internvideo-ego4d: A pack of champion solutions to ego4d challenges,” 2022,arXiv: 2211.09529
2022 arXiv
-
[14]
Egovideo: Exploring egocentric foundation model and downstream adaptation,
B. Peiet al., “Egovideo: Exploring egocentric foundation model and downstream adaptation,” 2024,arXiv:2406.18070
2024 arXiv
-
[15]
Modeling fine-grained hand-object dynamics for egocentric video representation learning,
——, “Modeling fine-grained hand-object dynamics for egocentric video representation learning,” 2025,arXiv:2503.00986
2025 arXiv
-
[16]
Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,
A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, “Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,”Int. Conf. Comput. Vis., pp. 2630–2640, 2019
2019
-
[17]
Ava: A video dataset of spatio-temporally localized atomic visual actions,
C. Guet al., “Ava: A video dataset of spatio-temporally localized atomic visual actions,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 6047–6056
2018
-
[18]
Ucf101: A dataset of 101 human actions classes from videos in the wild,
K. Soomro, A. R. Zamir, and M. Shah, “Ucf101: A dataset of 101 human actions classes from videos in the wild,” 2012,arXiv:1212.0402
2012 arXiv
-
[19]
Quo vadis, action recognition? a new model and the kinetics dataset,
J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,”IEEE Conf. Comput. Vis. Pattern Recog., pp. 4724–4733, 2017
2017
-
[20]
Internvid: A large-scale video-text dataset for multi- modal understanding and generation,
Y . Wanget al., “Internvid: A large-scale video-text dataset for multi- modal understanding and generation,” 2023,arXiv: 2307.06942
2023 arXiv
-
[21]
Cg-bench: Clue-grounded question answering bench- mark for long video understanding,
G. Chenet al., “Cg-bench: Clue-grounded question answering bench- mark for long video understanding,” 2024,arXiv: 2412.12075
2024 arXiv
-
[22]
Two-stream convolutional networks for action recognition in videos,
K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” inAdv. Neural Inform. Process. Syst. Cambridge, MA, USA: MIT Press, 2014, p. 568–576
2014
-
[23]
Vivit: A video vision transformer,
A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lu ˇci´c, and C. Schmid, “Vivit: A video vision transformer,” inInt. Conf. Comput. Vis., 2021, pp. 6816–6826
2021
-
[24]
Is space-time attention all you need for video understanding?
G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?” inInt. Conf. Mach. Learn., July 2021
2021
-
[25]
Dcan: improving temporal action detection via dual context aggregation,
G. Chen, Y .-D. Zheng, L. Wang, and T. Lu, “Dcan: improving temporal action detection via dual context aggregation,” inAAAI Conf. Artif. Intell., vol. 36, no. 1, 2022, pp. 248–257
2022
-
[26]
Video mamba suite: State space model as a versatile alternative for video understanding,
G. Chenet al., “Video mamba suite: State space model as a versatile alternative for video understanding,” 2024,arXiv: 2403.09626
2024 arXiv
-
[27]
Memory-and- anticipation transformer for online action understanding,
J. Wang, G. Chen, Y . Huang, L. Wang, and T. Lu, “Memory-and- anticipation transformer for online action understanding,” inInt. Conf. Comput. Vis., 2023, pp. 13 824–13 835
2023
-
[28]
Video super-resolution based on deep learning: a comprehensive survey,
H. Liuet al., “Video super-resolution based on deep learning: a comprehensive survey,”Artif Intell Rev, vol. 55, pp. 5981–6035, 2022
2022
-
[29]
About time: Advances, challenges, and outlooks of action understanding,
A. Stergiou and R. Poppe, “About time: Advances, challenges, and outlooks of action understanding,” 2024,arXiv:2411.15106
2024 arXiv
-
[30]
New generation deep learning for video object detection: A survey,
L. Jiaoet al., “New generation deep learning for video object detection: A survey,”IEEE Trans. Neural Netw. Learn. Syst., vol. 33, no. 8, pp. 3195–3215, 2022
2022
-
[31]
A survey of single- scene video anomaly detection,
B. Ramachandra, M. J. Jones, and R. R. Vatsavai, “A survey of single- scene video anomaly detection,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 5, pp. 2293–2312, 2022
2022
-
[32]
An outlook into the future of egocentric vision,
C. Plizzariet al., “An outlook into the future of egocentric vision,”Int. J. Comput. Vis., pp. 1–57, 2024
2024
-
[33]
Samsung’s new smart fridge lets you check in on its contents through internal cameras,
N. Lavars, “Samsung’s new smart fridge lets you check in on its contents through internal cameras,” Jan 2016. [Online]. Available: https://newatlas.com/samsung-family-hub-smart-fridge/41192/
2016
-
[34]
June oven,
J. Oven, “June oven,” Aug 2018. [Online]. Available: https: //firewireblog.com/2018/08/19/june-oven/
2018
-
[35]
Sportvu stats can be helpful, overwhelming,
J. McDonald, “Sportvu stats can be helpful, overwhelming,” Nov
-
[36]
Hawk-eye’s eagle eye on wimbledon tennis,
D. Winter, “Hawk-eye’s eagle eye on wimbledon tennis,” 2024. [Online]. Available: https://www.redsharknews.com/hawk-eyes-eye- on-wimbledon
2024
-
[37]
Super bowl li preview: Inside fox sports’ “be the player
J. Dachman, “Super bowl li preview: Inside fox sports’ “be the player” first-person pov replay tech,” Jan 2017. [Online]. Avail- able: https://www.sportsvideo.org/2017/01/13/super-bowl-li-preview- inside-fox-sports-be-the-player-360-pov-replay-technology/
2017
-
[38]
Smart glasses in hospitals: Viewing care delivery through a new lens,
H. M. Asia, “Smart glasses in hospitals: Viewing care delivery through a new lens,” HMA, 03 2022. [Online]. Available: https://www.hospitalmanagementasia.com/tech-innovation/ smart-glasses-in-hospitals-viewing-care-delivery-through-a-new-lens/
2022
-
[39]
Here’s how asean’s first 5g smart hospital is using ai to usher in a new era of healthcare,
——, “Here’s how asean’s first 5g smart hospital is using ai to usher in a new era of healthcare,” HMA, 02 2022. [Online]. Available: https://www.hospitalmanagementasia.com/tech-innovation/heres-how- aseans-first-5g-smart-hospital-is-using-ai-to-usher-in-a-new-era-of- healthcare/
2022
-
[40]
Cameras could be installed in classrooms in these states,
K. Rahman, “Cameras could be installed in classrooms in these states,” Jan 2024. [Online]. Available: https://www.newsweek.com/cameras- installed-classrooms-1859098
2024
-
[41]
Classroom camera: Transform education,
Reolink, “Classroom camera: Transform education,” Nov 2024. [Online]. Available: https://reolink.com/blog/classroom-camera
2024
-
[42]
3d surround view system,
E. News, “3d surround view system,” Jun 2018. [Online]. Available: https://www.ien.eu/article/3d-surround-view-system/
2018
-
[43]
Nauto launches real-time driver behavior learning platform for fleets,
FreightWaves, “Nauto launches real-time driver behavior learning platform for fleets,” Nov 2019. [Online]. Available: https://finance. yahoo.com/news/nauto-launches-real-time-driver-144742897.html
2019
-
[44]
Vision-centric semantic occupancy predic- tion for autonomous driving,
P. L. Liu, “Vision-centric semantic occupancy predic- tion for autonomous driving,” May 2023. [Online]. Available: https://towardsdatascience.com/vision-centric-semantic- occupancy-prediction-for-autonomous-driving-16a46dbd6f65
2023
-
[45]
Top 10 applications of robotics in 2024,
harkiran78, “Top 10 applications of robotics in 2024,” geeksforgeeks, Feb 2024. [Online]. Available: https://www.geeksforgeeks.org/ applications-of-robotics
2024
-
[46]
Research on body-worn cameras and law enforcement,
N. I. of Justice, “Research on body-worn cameras and law enforcement,” National Institute of Justice, Jan 2022. [Online]. Available: https://nij.ojp.gov/topics/articles/research-body- worn-cameras-and-law-enforcement
2022
-
[47]
Over 1,000 people saved with drone search and rescue: Dji,
I. Singh, “Over 1,000 people saved with drone search and rescue: Dji,” DroneDJ, Jul 2023. [Online]. Available: https://dronedj.com/ 2023/07/12/dji-drone-search-rescue-map/
2023
-
[48]
Empowering health & safety monitoring in manufacturing - fogsphere,
Fogsphere, “Empowering health & safety monitoring in manufacturing - fogsphere,” Sep 2023. [Online]. Available: https://fogsphere.com/ industries-served/manufacturing/
2023
-
[49]
A vision-guided robotic system designed to grab any object,
R. Begg, “A vision-guided robotic system designed to grab any object,” Aug 2024. [Online]. Available: https://www.machinedesign.com/markets/robotics/video/55131589/ cynlr-a-vision-guided-robotic-system-designed-to-grab-any-object
2024
-
[50]
How visual ai transforms assembly line operations in factories,
Jobit, “How visual ai transforms assembly line operations in factories,”
-
[51]
Ego-exo: Transferring visual representations from third-person to first-person videos,
Y . Li, T. Nagarajan, B. Xiong, and K. Grauman, “Ego-exo: Transferring visual representations from third-person to first-person videos,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 6943–6953
2021
-
[52]
Cross-view action recognition un- derstanding from exocentric to egocentric perspective,
T.-D. Truong and K. Luu, “Cross-view action recognition un- derstanding from exocentric to egocentric perspective,” 2023, arXiv:2305.15699
2023 arXiv
-
[53]
Unlocking exocentric video-language data for ego- centric video representation learning,
Z.-Y . Douet al., “Unlocking exocentric video-language data for ego- centric video representation learning,” 2024,arXiv:2408.03567
2024 arXiv
-
[54]
Holographic feature learning of egocentric-exocentric videos for multi-domain action recognition,
Y . Huang, X. Yang, J. Gao, and C. Xu, “Holographic feature learning of egocentric-exocentric videos for multi-domain action recognition,” IEEE Trans. Multimedia, vol. 24, pp. 2273–2286, 2022. 17
2022
-
[55]
Synchronization is all you need: Exocentric-to-egocentric transfer for temporal action segmentation with unlabeled synchronized video pairs,
C. Quattrocchi, A. Furnari, D. Di Mauro, M. V . Giuffrida, and G. M. Farinella, “Synchronization is all you need: Exocentric-to-egocentric transfer for temporal action segmentation with unlabeled synchronized video pairs,” 2023,arXiv:2312.02638
2023 arXiv
-
[56]
Action recognition in the presence of one egocentric and multiple static cameras,
B. Soran, A. Farhadi, and L. G. Shapiro, “Action recognition in the presence of one egocentric and multiple static cameras,” inLect. Notes Comput. Sci., 2014
2014
-
[57]
Cross-view exocentric to egocentric video synthesis,
G. Liu, H. Tang, H. Latapie, J. J. Corso, and Y . Yan, “Cross-view exocentric to egocentric video synthesis,”ACM Int. Conf. Multimedia, 2021
2021
-
[58]
4diff: 3d-aware diffusion model for third-to-first viewpoint translation,
F. Chenget al., “4diff: 3d-aware diffusion model for third-to-first viewpoint translation,” inEur. Conf. Comput. Vis., 2024
2024
-
[59]
Exo2egodvc: Dense video captioning of ego- centric procedural activities using web instructional videos,
T. Ohkawaet al., “Exo2egodvc: Dense video captioning of ego- centric procedural activities using web instructional videos,” 2023, arXiv:2311.16444
2023 arXiv
-
[60]
Retrieval-augmented egocentric video captioning,
J. Xuet al., “Retrieval-augmented egocentric video captioning,” in IEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 13 525–13 536
2024
-
[62]
Enhancing egocentric 3d pose estimation with third person views,
A. Dhamanaskar, M. Dimiccoli, E. Corona, A. Pumarola, and F. Moreno-Noguer, “Enhancing egocentric 3d pose estimation with third person views,”Pattern Recognition, vol. 138, p. 109358, 2023
2023
-
[63]
An exocentric look at egocentric actions and vice versa,
S. Ardeshir and A. Borji, “An exocentric look at egocentric actions and vice versa,”Comput Vision Image Understanding, pp. 61–68, 2018
2018
-
[64]
Learning fine-grained view-invariant representations from unpaired ego-exo videos via temporal alignment,
Z. S. Xue and K. Grauman, “Learning fine-grained view-invariant representations from unpaired ego-exo videos via temporal alignment,” inAdv. Neural Inform. Process. Syst., vol. 36, 2023, pp. 53 688–53 710
2023
-
[65]
Camera selection for occlusion-less surgery recording via training with an egocentric camera,
Y . Saito, R. Hachiuma, H. Saito, H. Kajita, Y . Takatsume, and T. Hayashida, “Camera selection for occlusion-less surgery recording via training with an egocentric camera,”IEEE Access, vol. 9, pp. 138 307–138 322, 2021
2021
-
[66]
Next-generation surgical navigation: Marker-less multi-view 6dof pose estimation of surgical instruments,
J. Heinet al., “Next-generation surgical navigation: Marker-less multi-view 6dof pose estimation of surgical instruments,” 2023, arXiv:2305.03535
2023 arXiv
-
[67]
Look both ways: Self-supervising driver gaze estimation and road scene saliency,
I. Kasahara, S. Stent, and H. S. Park, “Look both ways: Self-supervising driver gaze estimation and road scene saliency,” inEur. Conf. Comput. Vis.Springer, 2022, pp. 126–142
2022
-
[68]
Aide: A vision-driven multi-view, multi-modal, multi- tasking dataset for assistive driving perception,
D. Yanget al., “Aide: A vision-driven multi-view, multi-modal, multi- tasking dataset for assistive driving perception,” inInt. Conf. Comput. Vis., 2023, pp. 20 402–20 413
2023
-
[69]
Vision-based ma- nipulators need to also see from their hands,
K. Hsu, M. J. Kim, R. Rafailov, J. Wu, and C. Finn, “Vision-based ma- nipulators need to also see from their hands,” 2022,arXiv:2203.12677
2022 arXiv
-
[70]
Look closer: Bridging egocentric and third-person views with transformers for robotic manipulation,
R. Jangir, N. Hansen, S. Ghosal, M. Jain, and X. Wang, “Look closer: Bridging egocentric and third-person views with transformers for robotic manipulation,”IEEE Trans. Robot. Autom., vol. 7, no. 2, pp. 3046–3053, 2022
2022
-
[71]
Self-supervised disentangled representation learning for third-person imitation learning,
J. Shang and M. S. Ryoo, “Self-supervised disentangled representation learning for third-person imitation learning,” inIEEE Int. Conf. Intell. Rob. Syst., 2021, pp. 214–221
2021
-
[72]
Visual-policy learning through multi-camera view to single-camera view knowledge distilla- tion for robot manipulation tasks,
C. Acar, K. Binici, A. Tekirda ˘g, and Y . Wu, “Visual-policy learning through multi-camera view to single-camera view knowledge distilla- tion for robot manipulation tasks,”IEEE Trans. Robot. Autom., vol. 9, no. 1, pp. 691–698, 2024
2024
-
[73]
Third-person visual imitation learning via decoupled hierarchical controller,
P. Sharma, D. Pathak, and A. K. Gupta, “Third-person visual imitation learning via decoupled hierarchical controller,” inAdv. Neural Inform. Process. Syst., 2019
2019
-
[74]
Multi-view disentanglement for rein- forcement learning with multiple cameras,
M. Dunion and S. V . Albrecht, “Multi-view disentanglement for rein- forcement learning with multiple cameras,” 2024,arXiv:2404.14064
2024 arXiv
-
[75]
Multi-view masked world models for visual robotic manipulation,
Y . Seo, J. Kim, S. James, K. Lee, J. Shin, and P. Abbeel, “Multi-view masked world models for visual robotic manipulation,” inInt. Conf. Mach. Learn., 2023
2023
-
[76]
Learning fused state representations for control from multi-view observations,
Z. Wang, Y .-H. Li, X. Li, H. Zang, R. Laroche, and R. Islam, “Learning fused state representations for control from multi-view observations,” 2025,arXiv: 2502.01316
2025
-
[77]
Self-explainable affordance learning with embodied caption,
Z. Zhang, Z. Wei, G. Sun, P. Wang, and L. Van Gool, “Self-explainable affordance learning with embodied caption,” 2024,arXiv:2404.05603
2024 arXiv
-
[78]
Learning granularity-aware affordances from human- object interaction for tool-based functional grasping in dexterous robotics,
F. Yanget al., “Learning granularity-aware affordances from human- object interaction for tool-based functional grasping in dexterous robotics,” 2024,arXiv:2407.00614
2024 arXiv
-
[79]
Towards third-person visual imitation learning using generative adversarial networks,
L. Garello, F. Rea, N. Noceti, and A. Sciutti, “Towards third-person visual imitation learning using generative adversarial networks,” in IEEE Int. Conf. Dev. Learn.IEEE, 2022, pp. 121–126
2022
-
[80]
Diffusing in someone else’s shoes: Robotic perspective taking with diffusion,
J. Spisak, M. Kerzel, and S. Wermter, “Diffusing in someone else’s shoes: Robotic perspective taking with diffusion,” 2024, arXiv:2404.07735
2024 arXiv
-
[81]
Ego-to-exo: Interfacing third person visuals from egocentric views in real-time for improved rov teleoperation,
A. Abdullah, R. Chen, I. Rekleitis, and M. J. Islam, “Ego-to-exo: Interfacing third person visuals from egocentric views in real-time for improved rov teleoperation,” 2024,arXiv:2407.00848
2024 arXiv
-
[82]
Spatial assisted human-drone collabo- rative navigation and interaction through immersive mixed reality,
L. Morando and G. Loianno, “Spatial assisted human-drone collabo- rative navigation and interaction through immersive mixed reality,” in Int. Conf. Robot. Autom.IEEE, 2024, pp. 8707–8713
2024
-
[83]
Yowo: You only walk once to jointly map an indoor scene and register ceiling-mounted cameras,
F. Yang, S. Yamao, I. Kusajima, A. Moteki, S. Masui, and S. Jiang, “Yowo: You only walk once to jointly map an indoor scene and register ceiling-mounted cameras,”IEEE Trans. Circuit Syst. Video Technol., pp. 1–1, 2024
2024
-
[84]
Fourth- person captioning: Describing daily events by uni-supervised and tri- regularized training,
K. Nakashima, Y . Iwashita, A. Kawamura, and R. Kurazume, “Fourth- person captioning: Describing daily events by uni-supervised and tri- regularized training,” inIEEE Trans. Syst., Man, Cybern., 2018, pp. 2122–2127
2018
-
[85]
Put myself in your shoes: Lifting the egocentric perspective from exocentric videos,
M. Luo, Z. Xue, A. Dimakis, and K. Grauman, “Put myself in your shoes: Lifting the egocentric perspective from exocentric videos,” in Eur. Conf. Comput. Vis.Springer, 2025, pp. 407–425
2025
-
[86]
Animate any- one: Consistent and controllable image-to-video synthesis for character animation,
L. Hu, X. Gao, P. Zhang, K. Sun, B. Zhang, and L. Bo, “Animate any- one: Consistent and controllable image-to-video synthesis for character animation,”IEEE Conf. Comput. Vis. Pattern Recog., pp. 8153–8163, 2023
2023
-
[87]
Stable video diffusion: Scaling latent video diffusion models to large datasets,
A. Blattmannet al., “Stable video diffusion: Scaling latent video diffusion models to large datasets,” 2023,arXiv:2311.15127
2023 arXiv
-
[88]
Cogvideox: Text-to-video diffusion models with an expert transformer,
Z. Yanget al., “Cogvideox: Text-to-video diffusion models with an expert transformer,” 2024,arXiv:2408.06072
2024 arXiv
-
[89]
Structure and content-guided video synthesis with diffusion models,
P. Esser, J. Chiu, P. Atighehchian, J. Granskog, and A. Germanidis, “Structure and content-guided video synthesis with diffusion models,” inInt. Conf. Comput. Vis., 2023, pp. 7312–7322
2023
-
[90]
Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory,
S. Yinet al., “Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory,” 2023,arXiv:2308.08089
2023 arXiv
-
[91]
Intention-driven ego-to-exo video generation,
H. Luo, K. Zhu, W. Zhai, and Y . Cao, “Intention-driven ego-to-exo video generation,” 2024,arXiv:2403.09194
2024 arXiv
-
[92]
From my view to yours: Ego- augmented learning in large vision language models for understanding exocentric daily living activities,
D. Reilly, M. K. Govind, and S. Das, “From my view to yours: Ego- augmented learning in large vision language models for understanding exocentric daily living activities,” 2025,arXiv:2501.05711
2025 arXiv
-
[93]
Viewbirdiformer: Learn- ing to recover ground-plane crowd trajectories and ego-motion from a single ego-centric view,
M. Nishimura, S. Nobuhara, and K. Nishino, “Viewbirdiformer: Learn- ing to recover ground-plane crowd trajectories and ego-motion from a single ego-centric view,”IEEE Trans. Robot. Autom., vol. 8, no. 1, pp. 368–375, 2023
2023
-
[94]
View birdification in the crowd: Ground-plane localization from perceived movements,
——, “View birdification in the crowd: Ground-plane localization from perceived movements,”Int. J. Comput. Vis., vol. 131, no. 8, p. 2015–2031, May 2023
2015
-
[95]
Incrowdformer: On-ground pedestrian world model from ego- centric views,
——, “Incrowdformer: On-ground pedestrian world model from ego- centric views,” 2023,arXiv:2303.09534
2023 arXiv
-
[96]
Fastllve: Real-time low- light video enhancement with intensity-aware look-up table,
W. Li, G. Wu, W. Wang, P. Ren, and X. Liu, “Fastllve: Real-time low- light video enhancement with intensity-aware look-up table,” inACM Int. Conf. Multimedia, 2023, pp. 8134–8144
2023
-
[97]
Flow-guided sparse transformer for video deblurring,
J. Linet al., “Flow-guided sparse transformer for video deblurring,” in Int. Conf. Mach. Learn., vol. 162, 17–23 Jul 2022, pp. 13 334–13 343
2022
-
[98]
Blur interpolation transformer for real-world motion from blur,
Z. Zhong, M. Cao, X. Ji, Y . Zheng, and I. Sato, “Blur interpolation transformer for real-world motion from blur,” inInt. Conf. Comput. Vis., 2023, pp. 5713–5723
2023
-
[99]
Surgery recording without occlusions by multi-view surgical videos,
T. Shimizu, K. Oishi, R. Hachiuma, H. Kajita, Y . Takatsume, and H. Saito, “Surgery recording without occlusions by multi-view surgical videos,” in15th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications, 2020, pp. 837–844
2020
-
[100]
Deep selection: A fully supervised camera selection network for surgery recordings,
R. Hachiuma, T. Shimizu, H. Saito, H. Kajita, and Y . Takatsume, “Deep selection: A fully supervised camera selection network for surgery recordings,” inMedical Image Computing and Computer Assisted Intervention. Springer, 2020, pp. 419–428
2020
-
[101]
From third person to first person: Dataset and baselines for synthesis and retrieval,
M. Elfeki, K. Regmi, S. Ardeshir, and A. Borji, “From third person to first person: Dataset and baselines for synthesis and retrieval,” 2018, arXiv:1812.00104
2018 arXiv
-
[102]
Cross- view exocentric to egocentric video synthesis,
G. Liu, H. Tang, H. M. Latapie, J. J. Corso, and Y . Yan, “Cross- view exocentric to egocentric video synthesis,” inACM Int. Conf. Multimedia, 2021, pp. 974–982
2021
-
[103]
Parallel generative adversarial network for third-person to first-person image generation,
G. Liu, H. Latapie, O. Kilic, and A. Lawrence, “Parallel generative adversarial network for third-person to first-person image generation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 1917–1923
2022
-
[104]
Egoexo-gen: Ego-centric video prediction by watching exo-centric videos,
J. Xuet al., “Egoexo-gen: Ego-centric video prediction by watching exo-centric videos,” 2025,arXiv:2504.11732
2025 arXiv
-
[105]
Sibnet: Sibling convolutional encoder for video captioning,
S. Liu, Z. Ren, and J. Yuan, “Sibnet: Sibling convolutional encoder for video captioning,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 43, pp. 3259–3272, 2018. 18
2018
-
[106]
Spatio-temporal graph for video captioning with knowledge distillation,
B. Panet al., “Spatio-temporal graph for video captioning with knowledge distillation,”IEEE Conf. Comput. Vis. Pattern Recog., pp. 10 867–10 876, 2020
2020
-
[107]
Domain-specific semantics guided approach to video captioning,
H. Munusamy and C. C. Sekhar, “Domain-specific semantics guided approach to video captioning,”IEEE Winter Conf. Appl. Comput. Vis., pp. 1576–1585, 2020
2020
-
[108]
Worldscribe: Towards context-aware live visual descriptions,
R.-C. Chang, Y . Liu, and A. Guo, “Worldscribe: Towards context-aware live visual descriptions,” inACM Symp. User Interface Softw. Technol., 2024, pp. 1–18
2024
-
[109]
Wanderguide: Indoor map-less robotic guide for exploration by blind people,
M. Kuribayashi, K. Uehara, A. Wang, S. Morishima, and C. Asakawa, “Wanderguide: Indoor map-less robotic guide for exploration by blind people,” 2025,arXiv:2502.08906
2025 arXiv
-
[110]
Audio-adaptive activity recognition across video domains,
Y . Zhang, H. Doughty, L. Shao, and C. G. Snoek, “Audio-adaptive activity recognition across video domains,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 13 791–13 800
2022
-
[111]
Cross-view generalisation in action recognition: Feature design for transitioning from exocentric to egocentric views,
B. Rocha, P. Moreno, and A. Bernardino, “Cross-view generalisation in action recognition: Feature design for transitioning from exocentric to egocentric views,” inIberian Robotics Conference, 2023, pp. 155–166
2023
-
[112]
Learning from semantic alignment between unpaired multiviews for egocentric video recognition,
Q. Wang, L. Zhao, L. Yuan, T. Liu, and X. Peng, “Learning from semantic alignment between unpaired multiviews for egocentric video recognition,” inInt. Conf. Comput. Vis., 2023, pp. 3284–3294
2023
-
[113]
Unsupervised and semi-supervised domain adaptation for action recognition from drones,
J. Choi, G. Sharma, M. Chandraker, and J.-B. Huang, “Unsupervised and semi-supervised domain adaptation for action recognition from drones,” inIEEE Winter Conf. Appl. Comput. Vis., 2020, pp. 1717– 1726
2020
-
[114]
Egofish3d: Egocentric 3d pose estimation from a fisheye camera via self- supervised learning,
Y . Liu, J. Yang, X. Gu, Y . Chen, Y . Guo, and G.-Z. Yang, “Egofish3d: Egocentric 3d pose estimation from a fisheye camera via self- supervised learning,”IEEE Trans. Multimedia, pp. 8880–8891, 2023
2023
-
[115]
Ex2eg-mae: A framework for adaptation of exocentric video masked autoencoders for egocentric social role understanding,
M. Tran, Y . Kim, C.-C. Su, C.-H. Kuo, M. Sun, and M. Soleymani, “Ex2eg-mae: A framework for adaptation of exocentric video masked autoencoders for egocentric social role understanding,” inEur. Conf. Comput. Vis.Springer, 2025, pp. 1–19
2025
-
[116]
The audio-visual conversational graph: From an egocentric-exocentric perspective,
W. Jiaet al., “The audio-visual conversational graph: From an egocentric-exocentric perspective,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 26 396–26 405
2024
-
[117]
Learning to score figure skating sport videos,
C. Xu, Y . Fu, B. Zhang, Z. Chen, Y .-G. Jiang, and X. Xue, “Learning to score figure skating sport videos,”IEEE Trans. Circuit Syst. Video Technol., vol. 30, no. 12, pp. 4578–4590, 2019
2019
-
[118]
Expertaf: Expert actionable feedback from video,
K. Ashutosh, T. Nagarajan, G. Pavlakos, K. Kitani, and K. Grau- man, “Expertaf: Expert actionable feedback from video,” 2024, arXiv:2408.00672
2024 arXiv
-
[119]
Towards universal soccer video understanding,
J. Rao, H. Wu, H. Jiang, Y . Zhang, Y . Wang, and W. Xie, “Towards universal soccer video understanding,” 2024,arXiv:2412.01820
2024 arXiv
-
[120]
Learning 6-dof fine-grained grasp detection based on part affordance grounding,
Y . Songet al., “Learning 6-dof fine-grained grasp detection based on part affordance grounding,” 2024,arXiv: 2301.11564
2024 arXiv
-
[121]
Robo-abc: Affordance generalization beyond categories via semantic correspon- dence for robot manipulation,
Y . Ju, K. Hu, G. Zhang, G. Zhang, M. Jiang, and H. Xu, “Robo-abc: Affordance generalization beyond categories via semantic correspon- dence for robot manipulation,” inEur. Conf. Comput. Vis., 2025, pp. 222–239
2025
-
[122]
Learning affordance grounding from exocentric images,
H. Luo, W. Zhai, J. Zhang, Y . Cao, and D. Tao, “Learning affordance grounding from exocentric images,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2022, pp. 2252–2261
2022
-
[123]
Locate: Localize and transfer object parts for weakly supervised affordance grounding,
G. Li, V . Jampani, D. Sun, and L. Sevilla-Lara, “Locate: Localize and transfer object parts for weakly supervised affordance grounding,” in IEEE Conf. Comput. Vis. Pattern Recog., June 2023, pp. 10 922–10 931
2023
-
[124]
Weakly supervised multimodal affordance grounding for egocentric images,
L. Xu, Y . Gao, W. Song, and A. Hao, “Weakly supervised multimodal affordance grounding for egocentric images,” inAAAI Conf. Artif. Intell., vol. 38, no. 6, 2024, pp. 6324–6332
2024
-
[125]
Strategies to leverage foun- dational model knowledge in object affordance grounding,
A. Rai, K. Buettner, and A. Kovashka, “Strategies to leverage foun- dational model knowledge in object affordance grounding,” inIEEE. Conf. Comput. Vis. Pattern Recog. Workshops., June 2024, pp. 1714– 1723
2024
-
[126]
Learning deep features for discriminative localization,
B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” inIEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 2921–2929
2016
-
[127]
Learning transferable visual models from natural language supervision,
A. Radfordet al., “Learning transferable visual models from natural language supervision,” inInt. Conf. Mach. Learn., 2021
2021
-
[128]
Intra: Interaction relationship-aware weakly supervised affordance grounding,
J. H. Jang, H. Seo, and S. Y . Chun, “Intra: Interaction relationship-aware weakly supervised affordance grounding,” 2024, arXiv:2409.06210
2024 arXiv
-
[129]
Visual affordance and function understanding: A survey,
M. Hassanin, S. Khan, and M. Tahtali, “Visual affordance and function understanding: A survey,”ACM Comput. Surv., vol. 54, no. 3, Apr. 2021
2021
-
[130]
Emergencynet: Efficient aerial image classification for drone-based emergency monitoring using atrous con- volutional feature fusion,
C. Kyrkou and T. Theocharides, “Emergencynet: Efficient aerial image classification for drone-based emergency monitoring using atrous con- volutional feature fusion,”IEEE J. Sel. Topics Appl. Earth Observations Remote Sens., vol. 13, pp. 1687–1699, 2020
2020
-
[131]
Toward a roadmap for human-drone interaction,
J. R. Cauchard, M. Khamis, J. Garcia, M. Kljun, and A. M. Brock, “Toward a roadmap for human-drone interaction,”Interactions, vol. 28, pp. 76–81, 03 2021
2021
-
[134]
Third-person piloting: Increasing situational awareness using a spa- tially coupled second drone,
R. Temma, K. Takashima, K. Fujita, K. Sueda, and Y . Kitamura, “Third-person piloting: Increasing situational awareness using a spa- tially coupled second drone,” inACM Symp. User Interface Softw. Technol., 2019, pp. 507–519
2019
-
[136]
Vinci: A real-time embodied smart assistant based on egocentric vision-language model,
Y . Huanget al., “Vinci: A real-time embodied smart assistant based on egocentric vision-language model,” 2024,arXiv: 2412.21080
2024 arXiv
-
[137]
Lita: Language instructed temporal-localization assistant,
D.-A. Huanget al., “Lita: Language instructed temporal-localization assistant,” inEur. Conf. Comput. Vis., 2024
2024
-
[138]
An egocentric vision-language model based portable real-time smart assistant,
Y . Huanget al., “An egocentric vision-language model based portable real-time smart assistant,” 2025,arXiv:2503.04250
2025 arXiv
-
[139]
Lifelogging caption generation via fourth-person vision in a human–robot symbiotic envi- ronment,
K. Nakashima, Y . Iwashita, and R. Kurazume, “Lifelogging caption generation via fourth-person vision in a human–robot symbiotic envi- ronment,”Robomech J., vol. 7, 2020
2020
-
[140]
Faster r-cnn: Towards real-time object detection with region proposal networks,
S. Ren, K. He, R. B. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 39, pp. 1137–1149, 2015
2015
-
[141]
Temporally-weighted hierarchical clustering for un- supervised action segmentation,
M. S. Sarfraz, N. Murray, V . Sharma, A. Diba, L. V . Gool, and R. Stiefelhagen, “Temporally-weighted hierarchical clustering for un- supervised action segmentation,”IEEE Conf. Comput. Vis. Pattern Recog., pp. 11 220–11 229, 2021
2021
-
[142]
Improving action segmentation via graph-based temporal reasoning,
Y . Huang, Y . Sugano, and Y . Sato, “Improving action segmentation via graph-based temporal reasoning,” inIEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 14 024–14 034
2020
-
[143]
Collaborative weakly supervised video correlation learn- ing for procedure-aware instructional video analysis,
T. Heet al., “Collaborative weakly supervised video correlation learn- ing for procedure-aware instructional video analysis,”AAAI Conf. Artif. Intell., vol. 38, no. 3, pp. 2112–2120, Mar. 2024
2024
-
[144]
Svip: Sequence verification for procedures in videos,
Y . Qian, W. Luo, D. Lian, X. Tang, P. Zhao, and S. Gao, “Svip: Sequence verification for procedures in videos,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 19 858–19 870
2022
-
[145]
Mutual context network for jointly estimating egocentric gaze and action,
Y . Huang, M. Cai, Z. Li, F. Lu, and Y . Sato, “Mutual context network for jointly estimating egocentric gaze and action,”IEEE Transactions on Image Processing, vol. 29, pp. 7795–7806, 2020
2020
-
[146]
Egotransfer: Transferring motion across egocentric and exocentric domains using deep neural networks,
S. Ardeshir, K. Regmi, and A. Borji, “Egotransfer: Transferring motion across egocentric and exocentric domains using deep neural networks,” 2016,arXiv:1612.05836
2016 arXiv
-
[147]
First- and third-person video co-analysis by learning spatial-temporal joint attention,
H. Yu, M. Cai, Y . Liu, and F. Lu, “First- and third-person video co-analysis by learning spatial-temporal joint attention,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 6, pp. 6631–6646, 2023
2023
-
[148]
Psalm: Pixelwise segmentation with large multi-modal model,
Z. Zhang, Y . Ma, E. Zhang, and X. Bai, “Psalm: Pixelwise segmentation with large multi-modal model,” inEur. Conf. Comput. Vis., 2025, pp. 74–91
2025
-
[149]
Objectrelator: Enabling cross-view object relation understanding in ego-centric and exo-centric videos,
Y . Fu, R. Wang, Y . Fu, D. P. Paudel, X. Huang, and L. V . Gool, “Objectrelator: Enabling cross-view object relation understanding in ego-centric and exo-centric videos,” 2024,arXiv:2411.19083
2024 arXiv
-
[150]
Relating view directions of complementary-view mobile cameras via the human shadow,
R. Han, Y . Gan, L. Wang, N. Li, W. Feng, and S. Wang, “Relating view directions of complementary-view mobile cameras via the human shadow,”Int. J. Comput. Vis., vol. 131, no. 5, p. 1106–1121, Jan. 2023
2023
-
[151]
From a bird’s eye view to see: Joint camera and subject registration without the camera calibration,
Z. Qian, R. Han, W. Feng, F. F. Wang, and S. Wang, “From a bird’s eye view to see: Joint camera and subject registration without the camera calibration,”IEEE Conf. Comput. Vis. Pattern Recog., pp. 863–873, 2022
2022
-
[152]
Cal- ibration of non-overlapping cameras using an external slam system,
E. Ataer-Cansizoglu, Y . Taguchi, S. Ramalingam, and Y . Miki, “Cal- ibration of non-overlapping cameras using an external slam system,” in2014 2nd International Conference on 3D Vision, vol. 1, 2014, pp. 509–516
2014
-
[153]
Egolocate: Real-time motion capture, localization, and mapping with sparse body-mounted sensors,
X. Yiet al., “Egolocate: Real-time motion capture, localization, and mapping with sparse body-mounted sensors,”ACM Transactions on Graphics, vol. 42, no. 4, 2023
2023
-
[154]
Recognizing micro-actions and reactions from paired egocentric videos,
R. Yonetani, K. M. Kitani, and Y . Sato, “Recognizing micro-actions and reactions from paired egocentric videos,” inIEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 2629–2638
2016
-
[155]
Multi-view object pose estimation from correspondence distributions and epipolar geometry,
R. L. Haugaard and T. M. Iversen, “Multi-view object pose estimation from correspondence distributions and epipolar geometry,” inInt. Conf. Robot. Autom., 2023, pp. 1786–1792. 19
2023
-
[156]
Tell, don’t show: Language guidance eases transfer across domains in images and videos,
T. Kalluri, B. P. Majumder, and M. Chandraker, “Tell, don’t show: Language guidance eases transfer across domains in images and videos,” inInt. Conf. Mach. Learn., Jul 2024, pp. 22 879–22 894
2024
-
[157]
Pov: Prompt-oriented view-agnostic learning for egocentric hand-object interaction in the multi-view world,
B. Xu, S. Zheng, and Q. Jin, “Pov: Prompt-oriented view-agnostic learning for egocentric hand-object interaction in the multi-view world,” inACM Int. Conf. Multimedia, 2023, pp. 2807–2816
2023
-
[158]
Masked video and body- worn imu autoencoder for egocentric action recognition,
M. Zhang, Y . Huang, R. Liu, and Y . Sato, “Masked video and body- worn imu autoencoder for egocentric action recognition,” inEur. Conf. Comput. Vis., 2025, pp. 312–330
2025
-
[159]
Ego2top: Matching viewers in egocentric and top-view videos,
S. Ardeshir and A. Borji, “Ego2top: Matching viewers in egocentric and top-view videos,” inEur. Conf. Comput. Vis., 2016, pp. 253–268
2016
-
[160]
Egocentric meets top-view,
——, “Egocentric meets top-view,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 41, no. 6, pp. 1353–1366, 2019
2019
-
[161]
Integrating egocentric videos in top-view surveillance videos: Joint identification and temporal alignment,
——, “Integrating egocentric videos in top-view surveillance videos: Joint identification and temporal alignment,” inEur. Conf. Comput. Vis., Sep 2018
2018
-
[162]
Class segmentation and object localization with superpixel neighborhoods,
B. Fulkerson, A. Vedaldi, and S. Soatto, “Class segmentation and object localization with superpixel neighborhoods,”Int. Conf. Comput. Vis., pp. 670–677, 2009
2009
-
[163]
Joint person segmentation and identification in synchronized first- and third-person videos,
M. Xu, C. Fan, Y . Wang, M. S. Ryoo, and D. J. Crandall, “Joint person segmentation and identification in synchronized first- and third-person videos,” inEur. Conf. Comput. Vis., 2018, pp. 656–672
2018
-
[164]
Fusing personal and environmental cues for identification and segmentation of first-person camera wearers in third-person views,
Z. Zhao, Y . Wang, and C. Wang, “Fusing personal and environmental cues for identification and segmentation of first-person camera wearers in third-person views,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 16 477–16 487
2024
-
[165]
Visual-gps: Ego-downward and ambient video based person location association,
L. Yang, H. Jiang, Z. Huo, and J. Xiao, “Visual-gps: Ego-downward and ambient video based person location association,” inIEEE Conf. Comput. Vis. Pattern Recog Workshops., 2019, pp. 371–380
2019
-
[166]
Seeing the unseen: Predicting the first-person camera wearer’s location and pose in third-person scenes,
Y . Wen, K. K. Singh, M. Anderson, W.-P. Jan, and Y . J. Lee, “Seeing the unseen: Predicting the first-person camera wearer’s location and pose in third-person scenes,” inInt. Conf. Comput. Vis., 2021, pp. 3446–3455
2021
-
[167]
Egopca: A new framework for egocentric hand-object interaction understanding,
Y . Xuet al., “Egopca: A new framework for egocentric hand-object interaction understanding,” inInt. Conf. Comput. Vis., 2023, pp. 5273– 5284
2023
-
[168]
Ego- centric action recognition by capturing hand-object contact and object state,
T. Shiota, M. Takagi, K. Kumagai, H. Seshimo, and Y . Aono, “Ego- centric action recognition by capturing hand-object contact and object state,” inIEEE Winter Conf. Appl. Comput. Vis., 2024, pp. 6541–6551
2024
-
[169]
Egochoir: Capturing 3d human-object interaction regions from egocentric views,
Y . Yang, W. Zhai, C. Wang, C. Yu, Y . Cao, and Z.-J. Zha, “Egochoir: Capturing 3d human-object interaction regions from egocentric views,” Adv. Neural Inform. Process. Syst., vol. 37, pp. 54 529–54 557, 2025
2025
-
[170]
The audio-visual conversational graph: From an egocentric-exocentric perspective,
W. Jia, M. Liu, H. Jiang, I. Ananthabhotla, J. M. Rehg, V . K. Ithapu, and R. Gao, “The audio-visual conversational graph: From an egocentric-exocentric perspective,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 26 386–26 395
2024
-
[171]
Complementary-view co-interest person detection,
R. Han, J. Zhao, W. Feng, Y . Gan, L. Wan, and S. Wang, “Complementary-view co-interest person detection,” inACM Int. Conf. Multimedia, 2020, p. 2746–2754
2020
-
[172]
Complementary-view multiple human tracking,
R. Han, W. Feng, J. Zhao, Z. Niu, Y . Zhang, and L. Wan, “Complementary-view multiple human tracking,” inAAAI Conf. Artif. Intell., vol. 34, 02 2020
2020
-
[173]
Multiple human as- sociation and tracking from egocentric and complementary top views,
R. Han, W. Feng, Y . Zhang, J. Zhao, and S. Wang, “Multiple human as- sociation and tracking from egocentric and complementary top views,” IEEE Trans. Pattern Anal. Mach. Intell., pp. 5225–5242, 2022
2022
-
[174]
Connecting the complementary-view videos: Joint camera identification and subject association,
R. Han, Y . Gan, J. Li, F. Wang, W. Feng, and S. Wang, “Connecting the complementary-view videos: Joint camera identification and subject association,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 2406–2415
2022
-
[175]
You only look once: Unified, real-time object detection,
J. Redmon, S. K. Divvala, R. B. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,”IEEE Conf. Comput. Vis. Pattern Recog., pp. 779–788, 2015
2015
-
[176]
Learning visual robotic control efficiently with contrastive pre-training and data augmentation,
A. Zhan, R. Zhao, L. Pinto, P. Abbeel, and M. Laskin, “Learning visual robotic control efficiently with contrastive pre-training and data augmentation,” inIEEE Int. Conf. Intell. Rob. Syst., 2022, pp. 4040– 4047
2022
-
[177]
Multi-view masked world models for visual robotic manipulation,
Y . Seo, J. Kim, S. James, K. Lee, J. Shin, and P. Abbeel, “Multi-view masked world models for visual robotic manipulation,” inInt. Conf. Mach. Learn., vol. 202, 23–29 Jul 2023, pp. 30 613–30 632
2023
-
[178]
Rt-1: Robotics transformer for real-world control at scale,
A. Brohanet al., “Rt-1: Robotics transformer for real-world control at scale,” 2022,arXiv:2212.06817
2022 arXiv
-
[179]
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware,
T. Z. Zhao, V . Kumar, S. Levine, and C. Finn, “Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware,” inProc. Robot.: Sci. Syst., Daegu, Republic of Korea, July 2023
2023
-
[180]
Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking,
H. Bharadhwaj, J. Vakil, M. Sharma, A. Gupta, S. Tulsiani, and V . Ku- mar, “Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking,” inInt. Conf. Robot. Autom., 2024, pp. 4788–4795
2024
-
[181]
Diffusion policy: Visuomotor policy learning via action diffusion,
C. Chiet al., “Diffusion policy: Visuomotor policy learning via action diffusion,” inProc. Robot.: Sci. Syst., 2023
2023
-
[182]
Aloha unleashed: A simple recipe for robot dexterity,
T. Z. Zhaoet al., “Aloha unleashed: A simple recipe for robot dexterity,” 2024,arXiv:2410.13126
2024 arXiv
-
[183]
Deep 360 pilot: Learning a deep agent for piloting through 360 sports videos,
H.-N. Hu, Y .-C. Lin, M.-Y . Liu, H.-T. Cheng, Y .-J. Chang, and M. Sun, “Deep 360 pilot: Learning a deep agent for piloting through 360 sports videos,” inIEEE Conf. Comput. Vis. Pattern Recog., 2017, pp. 1396– 1405
2017
-
[184]
Making 360 video watchable in 2d: Learning videography for click free viewing,
Y .-C. Su and K. Grauman, “Making 360 video watchable in 2d: Learning videography for click free viewing,” inIEEE Conf. Comput. Vis. Pattern Recog., 2017, pp. 1368–1376
2017
-
[185]
Multicamba: a system for selecting camera views in live broadcasting of sport events using a dynamic 3d model,
R. Yus, E. Mena, S. Ilarri, A. Illarramendi, and J. Bernad, “Multicamba: a system for selecting camera views in live broadcasting of sport events using a dynamic 3d model,”Multimedia Tools and Applications, vol. 74, pp. 4059–4090, 2015
2015
-
[186]
Learning sports camera selection from internet videos,
J. Chen, K. Lu, S. Tian, and J. Little, “Learning sports camera selection from internet videos,” inIEEE Winter Conf. Appl. Comput. Vis.IEEE, 2019, pp. 1682–1691
2019
-
[187]
Which viewpoint shows it best? language for weakly supervising view selection in multi-view videos,
S. Majumder, T. Nagarajan, Z. Al-Halah, R. Pradhan, and K. Grauman, “Which viewpoint shows it best? language for weakly supervising view selection in multi-view videos,” 2024,arXiv: 2411.08753
2024 arXiv
-
[188]
Switch- a-view: Few-shot view selection learned from edited videos,
S. Majumder, T. Nagarajan, Z. Al-Halah, and K. Grauman, “Switch- a-view: Few-shot view selection learned from edited videos,” 2024, arXiv: 2412.18386
2024 arXiv
-
[189]
Lemma: A multi- view dataset for le arning m ulti-agent m ulti-task a ctivities,
B. Jia, Y . Chen, S. Huang, Y . Zhu, and S.-c. Zhu, “Lemma: A multi- view dataset for le arning m ulti-agent m ulti-task a ctivities,” inEur. Conf. Comput. Vis.Springer, 2020, pp. 767–786
2020
-
[190]
Ft-hid: a large-scale rgb-d dataset for first-and third-person human interaction analysis,
Z. Guo, Y . Hou, P. Wang, Z. Gao, M. Xu, and W. Li, “Ft-hid: a large-scale rgb-d dataset for first-and third-person human interaction analysis,”Neural Computing and Applications, pp. 2007–2024, 2023
2007
-
[191]
Wts: A pedestrian-centric traffic video dataset for fine-grained spatial-temporal understanding,
Q. Konget al., “Wts: A pedestrian-centric traffic video dataset for fine-grained spatial-temporal understanding,” 2024,arXiv:2407.15350
2024 arXiv
-
[192]
Egoexolearn: A dataset for bridging asynchronous ego-and exo-centric view of procedural activities in real world,
Y . Huanget al., “Egoexolearn: A dataset for bridging asynchronous ego-and exo-centric view of procedural activities in real world,” in IEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 22 072–22 086
2024
-
[193]
Gazevqa: A video question answering dataset for multiview eye-gaze task-oriented collaborations,
M. Ilaslanet al., “Gazevqa: A video question answering dataset for multiview eye-gaze task-oriented collaborations,” inConf. Empir. Methods Nat. Lang. Process., Proc., 2023, pp. 10 462–10 479
2023
-
[194]
Predicting gaze in egocentric video by learning task-dependent attention transition,
Y . Huang, M. Cai, Z. Li, and Y . Sato, “Predicting gaze in egocentric video by learning task-dependent attention transition,” inEur. Conf. Comput. Vis.Springer, 2018, pp. 789–804
2018
-
[195]
Guide to the carnegie mellon university multimodal activity (cmu-mmac) database,
F. D. la Torre Fradeet al., “Guide to the carnegie mellon university multimodal activity (cmu-mmac) database,” Carnegie Mellon Univer- sity, Pittsburgh, PA, Tech. Rep. CMU-RI-TR-08-22, April 2008
2008
-
[196]
H2o: Two hands manipulating objects for first person interaction recognition,
T. Kwon, B. Tekin, J. St ¨uhmer, F. Bogo, and M. Pollefeys, “H2o: Two hands manipulating objects for first person interaction recognition,”Int. Conf. Comput. Vis., pp. 10 118–10 128, 2021
2021
-
[197]
Assembly101: A large-scale multi-view video dataset for understanding procedural activities,
F. Seneret al., “Assembly101: A large-scale multi-view video dataset for understanding procedural activities,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 21 096–21 106
2022
-
[198]
ARCTIC: A dataset for dexterous bimanual hand-object manipulation,
Z. Fanet al., “ARCTIC: A dataset for dexterous bimanual hand-object manipulation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023
2023
-
[199]
Oakink2: A dataset of bimanual hands-object manipu- lation in complex task completion,
X. Zhanet al., “Oakink2: A dataset of bimanual hands-object manipu- lation in complex task completion,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 445–456
2024
-
[200]
Home action genome: Cooperative compositional action understanding,
N. Raiet al., “Home action genome: Cooperative compositional action understanding,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 11 179–11 188
2021
-
[201]
Egoexo-fitness: Towards egocentric and exocentric full-body action understanding,
Y .-M. Li, W.-J. Huang, A.-L. Wang, L.-A. Zeng, J.-K. Meng, and W.-S. Zheng, “Egoexo-fitness: Towards egocentric and exocentric full-body action understanding,” 2024,arXiv:2406.08877
2024 arXiv
-
[202]
Hollywood in homes: Crowdsourcing data collection for activity understanding,
G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. K. Gupta, “Hollywood in homes: Crowdsourcing data collection for activity understanding,” inEur. Conf. Comput. Vis., 2016
2016
-
[203]
Core4d: A 4d human- object-human interaction dataset for collaborative object rearrange- ment,
C. Zhang, Y . Liu, R. Xing, B. Tang, and L. Yi, “Core4d: A 4d human- object-human interaction dataset for collaborative object rearrange- ment,” 2024,arXiv:2406.19353
2024 arXiv
-
[204]
Egome: Follow me via egocentric view in real world,
H. Qiu, Z. Shi, L. Wang, H. Xiong, X. Li, and H. Li, “Egome: Follow me via egocentric view in real world,” 2025,arXiv: 2501.19061
2025 arXiv
-
[205]
Estimating egocentric 3d human pose in the wild with external weak supervision,
J. Wang, L. Liu, W. Xu, K. Sarkar, D. Luvizon, and C. Theobalt, “Estimating egocentric 3d human pose in the wild with external weak supervision,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 13 157–13 166
2022
-
[206]
Assemblyhands: Towards egocentric activity understanding via 3d hand pose estimation,
T. Ohkawa, K. He, F. Sener, T. Hodan, L. Tran, and C. Keskin, “Assemblyhands: Towards egocentric activity understanding via 3d hand pose estimation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 12 999–13 008. 20
2023
-
[207]
Thermohands: A benchmark for 3d hand pose estimation from egocentric thermal image,
F. Ding, Y . Zhu, X. Wen, and C. X. Lu, “Thermohands: A benchmark for 3d hand pose estimation from egocentric thermal image,” 2024, arXiv:2403.09871
2024 arXiv
-
[208]
Nymeria: A massive collection of multimodal egocentric daily motion in the wild,
L. Maet al., “Nymeria: A massive collection of multimodal egocentric daily motion in the wild,” inEur. Conf. Comput. Vis., 2024
2024
-
[209]
Ovr: A dataset for open vocabulary temporal repetition counting in videos,
D. Dwibedi, Y . Aytar, J. Tompson, and A. Zisserman, “Ovr: A dataset for open vocabulary temporal repetition counting in videos,” 2024, arXiv:2407.17085
2024 arXiv
-
[210]
One-shot affordance detection,
H. Luo, W. Zhai, J. Zhang, Y . Cao, and D. Tao, “One-shot affordance detection,” 2021,arXiv:2106.14747
2021 arXiv
-
[211]
360 +x: A panoptic multi-modal scene understanding dataset,
H. Chen, Y . Hou, C. Qu, I. Testini, X. Hong, and J. Jiao, “360 +x: A panoptic multi-modal scene understanding dataset,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 19 373–19 382
2024
-
[212]
Visual instruction tuning,
H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,”Adv. Neural Inform. Process. Syst., vol. 36, 2024
2024
-
[213]
Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,
Z. Chenet al., “Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 24 185–24 198
2024
-
[214]
Eagle 2: Building post-training data strategies from scratch for frontier vision-language models,
Z. Liet al., “Eagle 2: Building post-training data strategies from scratch for frontier vision-language models,” 2025,arXiv: 2501.14818
2025 arXiv
-
[215]
Videollm: Modeling video sequence with large language models,
G. Chenet al., “Videollm: Modeling video sequence with large language models,” 2023,arXiv: 2305.13292
2023 arXiv
-
[216]
Video-rag: Visually-aligned retrieval-augmented long video comprehension,
Y . Luoet al., “Video-rag: Visually-aligned retrieval-augmented long video comprehension,” 2024,arXiv:2411.13093
2024
-
[2013]
Available: https://www.expressnews.com/sports/spurs/ article/SportVU-stats-can-be-helpful-overwhelming-4993731.php
[Online]. Available: https://www.expressnews.com/sports/spurs/ article/SportVU-stats-can-be-helpful-overwhelming-4993731.php
-
[2024]
Available: https://randomwalk.ai/blog/how-visual-ai- transforms-assembly-line-operations-in-factories/
[Online]. Available: https://randomwalk.ai/blog/how-visual-ai- transforms-assembly-line-operations-in-factories/
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.