Pith. sign in

REVIEW 3 major objections 6 minor 217 references

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A first survey unifies first- and third-person video AI.

desk verdict Useful organizing survey of ego-exo video understanding, solid dataset table, but 'comprehensive' claim needs either a search methodology or a softer claim. read the letter →

arxiv 2506.06253 v1 pith:KJTJ3EHL submitted 2025-06-06 cs.CV

classification cs.CV
keywords videounderstandingegocentricexocentriccross-viewlearningfirst-personvisionsurveydatasetsandbenchmarks
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This survey argues that perceiving the world from both a first-person (egocentric) and a third-person (exocentric) perspective is a distinct and rapidly growing research direction in video understanding, and claims to be the first comprehensive review of that direction. It proposes that the field is naturally organized into three lines of work: using egocentric data to improve exocentric understanding, using exocentric data to improve egocentric analysis, and jointly learning from both views. The survey connects these directions to eight application areas, from cooking assistants to traffic and surgery, and to the datasets that support them. A sympathetic reader would care because the paper supplies a taxonomy and a gap analysis that can steer future work toward the tasks and data that are still missing.

What carries the argument

The organizing device is a three-way taxonomy of research directions—ego-for-exo (egocentric cues improve exocentric tasks), exo-for-ego (exocentric data improves egocentric analysis), and joint learning (both views are used at training and inference together). The survey maps each direction to concrete tasks such as video generation, action understanding, affordance grounding, and cross-view retrieval, and links those tasks back to eight application scenarios. This taxonomy does the work of a framework: it positions every reviewed method, exposes where the literature is thin, and drives the paper's dataset inventory and future-directions discussion.

What would settle it

An independent, systematic literature search with explicit inclusion criteria that surfaces a substantial body of ego-exo work omitted from this survey, or that reveals a major research direction (for instance multi-agent ego-exo collaboration) not covered by the three-way taxonomy, would falsify the survey's central claims of comprehensiveness and completeness.

Watch

Extended reading notes

Core claim

The paper's central claim is that cross-view collaboration between egocentric and exocentric video is a coherent field whose progress can be mapped by three research directions: egocentric-for-exocentric, exocentric-for-egocentric, and joint learning. On the paper's own terms, no prior survey has integrated both perspectives, so this review fills that gap by organizing existing methods, identifying eight application domains that would benefit from ego-exo collaboration, and cataloguing datasets that contain both viewpoints. The review concludes that current work clusters around daily-life activity understanding, while tasks such as ego-to-exo video generation, view birdification, skill assessment, and affordance grounding for industrial or surgical tools remain under-explored.

Load-bearing premise

The survey's value rests on its selection of papers and datasets being complete and representative enough that the 'first comprehensive review' claim and the resulting gap analysis are trustworthy.

Editorial extensions

If this is right

  • A researcher can place any ego-exo method into one of three categories, making the field's structure explicit for newcomers.
  • The gap analysis identifies ego-to-exo video generation, view birdification, and cross-view skill assessment as under-studied tasks worth pursuing.
  • Specialized domains such as healthcare, education, and public service lack dedicated ego-exo datasets, so collecting them is a clear next step.
  • Because synchronized paired data is expensive, aligning unpaired ego and exo videos via language or retrieval is a promising route to scale.
  • Vision-language models that can take both perspectives as input are proposed as a path toward unified cross-view frameworks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The 'first comprehensive survey' claim is only as strong as the literature search behind it; a reader can probe it by checking whether recent ego-exo works beyond the cited set fit the taxonomy.
  • The three-way taxonomy may under-represent multi-agent ego-exo collaboration, which appears in applications like rescue and public service but is not given its own research direction.
  • A concrete testable extension of the gap analysis is to benchmark ego-to-exo video generation on the driving and robotics scenarios the paper names, where latency constraints are severe.
  • The dataset table offers a way to quantify domain imbalance, for instance by counting ego-exo datasets per application area, which could prioritize future data collection.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper presents a survey of video understanding research that integrates egocentric (first-person) and exocentric (third-person) perspectives. It opens with eight application domains, derives a set of research tasks from them, organizes the literature into three directions (ego-for-exo, exo-for-ego, and joint learning), and provides a table of ego-exo datasets. The paper claims to be the first comprehensive survey of this integration and concludes with a gap analysis and future-work suggestions. A companion GitHub repository is announced.

Significance. If its coverage is representative, the survey is timely and useful: it gives the area a coherent taxonomy, collects many relevant methods across disparate subfields, and the dataset table in Section V is a practical contribution. The paper also ships a companion GitHub repository, which helps researchers track new work. However, the central 'comprehensive' claim and the under-explored task analysis rest on an undocumented literature sample, and a few taxonomically misplaced sections weaken the review's internal consistency. The contribution is therefore not yet fully supported, although the issues appear fixable within the manuscript's scope.

major comments (3)
  1. [Section I, Section VI, Fig. 4] The paper's central claim that it is the first comprehensive survey of egocentric-exocentric video understanding (Section I) and the gap analysis in Section VI are not backed by a documented literature search. There is no search protocol, inclusion criterion, or comparison with prior or concurrent surveys, so the reader cannot judge whether the works sampled in Sections IV and V are representative. This matters because Section VI draws conclusions about which tasks are 'under-explored' and Fig. 4 is used to state that 'many tasks critical to applications remain under-investigated.' I request an explicit methodology paragraph (databases, query terms, inclusion/exclusion criteria, and screening flow) and, ideally, a recall-style check against a systematic ego-exo query to substantiate the comprehensiveness claim. Without this, the claim remains an unverified external completeness assertion.
  2. [Section IV.B, Fig. 5] The task 'Remote Drone Teleoperation' is placed under 'Exocentric for Egocentric' video understanding, but the works reviewed there ([132]–[135] and [133]) are human-computer interaction systems that use VR, additional cameras, or AR overlays to provide a pilot with a better exocentric view during teleoperation; they do not use exocentric data to improve egocentric video understanding, which is the definition given at the start of Section IV.B. This inclusion stretches the taxonomy and weakens the review's stated focus on video understanding. Either move these works to an application-oriented subsection or add an explicit justification of how they inform ego-exo video understanding.
  3. [Section III, Fig. 4] The 'from applications to research tasks' bridge is incomplete. Section IV reviews view birdification, cross-view retrieval, 3D camera localization, egocentric wearer identification, and cross-view human identification, but none of these tasks appears in Fig. 4, which is used in Section III to support the statement that many application-critical tasks are under-investigated. The figure should include all tasks discussed in Section IV, or the text should clearly state that Fig. 4 is a partial mapping; alternatively, the gap analysis should be supported by the full task list in Section IV.
minor comments (6)
  1. [Figure 2 caption] The caption says Section IV is divided into 'ego for exo, exo for exo, and joint learning'; 'exo for exo' should be 'exo for ego'.
  2. [Figure 4] There are typos in the figure: 'Camara Localization' should be 'Camera Localization' and 'Healthcar e' should be 'Healthcare'.
  3. [References] References [75] and [177] are the same paper (Seo et al., 'Multi-view masked world models for visual robotic manipulation'), and references [116] and [170] are the same paper (Jia et al., 'The audio-visual conversational graph'); these duplicates should be unified.
  4. [Section IV.B, Self-supervised methods] EgoFish3D [114] is described as using an exocentric pose estimator as supervision; this is more accurately characterized as weakly supervised than self-supervised, given the terminology used in the surrounding text.
  5. [Section V.D] The dataset entry 'ThirdtoFirst [51]' cites a method paper (Li et al., 'Ego-Exo: Transferring visual representations from third-person to first-person videos') rather than a dataset paper; please align the citation with the actual dataset source or rename the entry.
  6. [Fig. 1] The caption says the citation counts cover 'papers and datasets discussed in Sections IV and V,' but the collection and curation rules for the Google Scholar data are not described; please clarify the inclusion criteria so readers can interpret the growth curve.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation; survey is descriptive and self-contained, with only minor self-citation bias.

full rationale

This is a survey paper with no equations, no fitted parameters, and no predictive claim whose outcome could be forced by construction. The central organizational scheme (ego-for-exo, exo-for-ego, joint learning) is a taxonomy, not a derived result; it is defined by the direction of information flow and then applied to the literature. The novelty claim that 'no survey has yet addressed the integration of both perspectives' is an external literature-coverage assertion supported by references to existing surveys [28]–[32], none of which is authored by the present authors; it is not a self-citation chain used to forbid alternatives. The self-citations are numerous (e.g., [13], [14], [54], [60], [136], [142], [145], [158], [192], [194]), but they appear as entries in the reviewed literature, not as load-bearing premises that define the taxonomy or establish the gap. Fig. 1's citation curve is computed from papers the survey itself discusses, but the citation counts themselves are external Google Scholar data, and the figure is presented as a descriptive trend of that selected set, not as a prediction or proof. The absence of a search protocol or inclusion criteria is a legitimate completeness-risk concern, but it is an external validity issue, not circularity: the survey never claims to derive its completeness from its own contents. No passage satisfies the quoted-reduction threshold required to flag a circular step. A score of 1 reflects the elevated self-citation rate and the self-selected Fig. 1 sample, which create a mild self-referential flavor without making any central claim equivalent to its inputs.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

This is a survey with no fitted parameters and no introduced entities. The central claim rests on the adequacy of the literature selection and the validity of its taxonomy.

assumptions (3)
  • domain assumption The three-way taxonomy (ego for exo, exo for ego, joint learning) is the natural and exhaustive organizing principle for cross-view ego-exo work.
    Section IV divides all reviewed work into these three categories; this partitioning is asserted, not derived, and some tasks (e.g., video captioning) appear under both unidirectional and joint headings.
  • domain assumption The surveyed papers and datasets are accurately described and representative.
    The paper relies on secondary descriptions of methods and datasets; factual errors in these summaries would propagate into the survey's claims.
  • domain assumption Citation counts in Figure 1 from Google Scholar are a meaningful proxy for field growth.
    Figure 1 uses Google Scholar citation data over a hand-picked paper set to claim rapid growth; the selection of papers for citation counting is not detailed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision." pith.science (2026). https://pith.science/paper/KJTJ3EHL

@misc{pith2026250606253,
  author       = {Pith},
  title        = {Pith review of: Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KJTJ3EHL}},
  note         = {Machine review of arXiv:2506.06253}
}
read the original abstract

Perceiving the world from both egocentric (first-person) and exocentric (third-person) perspectives is fundamental to human cognition, enabling rich and complementary understanding of dynamic environments. In recent years, allowing the machines to leverage the synergistic potential of these dual perspectives has emerged as a compelling research direction in video understanding. In this survey, we provide a comprehensive review of video understanding from both exocentric and egocentric viewpoints. We begin by highlighting the practical applications of integrating egocentric and exocentric techniques, envisioning their potential collaboration across domains. We then identify key research tasks to realize these applications. Next, we systematically organize and review recent advancements into three main research directions: (1) leveraging egocentric data to enhance exocentric understanding, (2) utilizing exocentric data to improve egocentric analysis, and (3) joint learning frameworks that unify both perspectives. For each direction, we analyze a diverse set of tasks and relevant works. Additionally, we discuss benchmark datasets that support research in both perspectives, evaluating their scope, diversity, and applicability. Finally, we discuss limitations in current works and propose promising future research directions. By synthesizing insights from both perspectives, our goal is to inspire advancements in video understanding and artificial intelligence, bringing machines closer to perceiving the world in a human-like manner. A GitHub repo of related works can be found at https://github.com/ayiyayi/Awesome-Egocentric-and-Exocentric-Vision.

Figures

Figures reproduced from arXiv: 2506.06253 by the authors.

Figure 1
Figure 1. Number of citations to egocentric-exocentric related papers from 2015 [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The overall structure of the survey. We first highlight the application value of egocentric and exocentric collaboration (Section [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Examples of the potential collaboration of egocentric and exocentric vision in diverse applications. We illustrate how integrating egocentric and [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (15 more)
Figure 4
Figure 4. Figure 4: Mapping relevant research works to applications and research tasks. [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Overall structure of Section Research Tasks. We discuss the research task from three aspects: Egocentric for Exocentric, Exocentric for Egocentric, and Joint Learning. Each subsection reviews a variety of tasks and their existing works. Action Understanding. Human acti…
Figure 7
Figure 7. Figure 7: Illustration of a typical method for ego-for-exo action understanding, [PITH_FULL_IMAGE:figures/full_fig_p006_7.png]
Figure 9
Figure 9. Figure 9: Illustration of a typical method for view selection in surgical recording, [PITH_FULL_IMAGE:figures/full_fig_p007_9.png]
Figure 10
Figure 10. Figure 10: Illustration of the general GAN-based framework for exo-to [PITH_FULL_IMAGE:figures/full_fig_p007_10.png]
Figure 12
Figure 12. Figure 12: Illustration of a general adversarial-based approach for exo-for [PITH_FULL_IMAGE:figures/full_fig_p008_12.png]
Figure 13
Figure 13. Figure 13: Illustration of the exo-for-ego affordance grounding framework. [PITH_FULL_IMAGE:figures/full_fig_p009_13.png]
Figure 16
Figure 16. Figure 16: Illustration of the general cross-view retrieval framework. Exocentric [PITH_FULL_IMAGE:figures/full_fig_p010_16.png]
Figure 17
Figure 17. Figure 17: Illustration of a typical method for egocentric camera localization, [PITH_FULL_IMAGE:figures/full_fig_p011_17.png]
Figure 18
Figure 18. Figure 18: Illustration of cross-view action understanding. This involves action [PITH_FULL_IMAGE:figures/full_fig_p011_18.png]
Figure 20
Figure 20. Figure 20: Illustration of a typical method for cross-view human tracking and [PITH_FULL_IMAGE:figures/full_fig_p012_20.png]
Figure 21
Figure 21. Figure 21: Illustration of a typical framework of multi-view robotic manipula [PITH_FULL_IMAGE:figures/full_fig_p012_21.png]
Figure 19
Figure 19. Figure 19: Illustration of a typical method for egocentric wearer identification, [PITH_FULL_IMAGE:figures/full_fig_p012_19.png]
Figure 23
Figure 23. Figure 23: Illustration of a typical ego-exo view selection method, adapted [PITH_FULL_IMAGE:figures/full_fig_p013_23.png]
Figure 22
Figure 22. Figure 22: Illustration of a typical method for remote drone teleoperation with [PITH_FULL_IMAGE:figures/full_fig_p013_22.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

217 extracted references · 47 canonical work pages

  1. [132]

    Starhopper: A touch interface for remote object-centric drone navigation,

    J. Li, R. Balakrishnan, and T. Grossman, “Starhopper: A touch interface for remote object-centric drone navigation,” inProc Graphics Interface, ser. GI 2020, 2020, pp. 317 – 326

  2. [135]

    Birdviewar: Surroundings-aware remote drone piloting using an augmented third- person perspective,

    M. Inoue, K. Takashima, K. Fujita, and Y . Kitamura, “Birdviewar: Surroundings-aware remote drone piloting using an augmented third- person perspective,” inConf Hum Fact Comput Syst Proc, 2023

  3. [133]

    Drone- augmented human vision: Exocentric control for drones exploring hidden areas,

    O. Erat, W. A. Isop, D. Kalkofen, and D. Schmalstieg, “Drone- augmented human vision: Exocentric control for drones exploring hidden areas,”IEEE Trans. Vis. Comput. Graphics, vol. 24, pp. 1437– 1446, 04 2018

  4. [1]

    The mirror-neuron system,

    G. Rizzolatti and L. Craighero, “The mirror-neuron system,”Annual review of neuroscience, vol. 27, pp. 169–92, 02 2004

  5. [2]

    Actor and observer: Joint modeling of first and third-person videos,

    G. A. Sigurdsson, A. Gupta, C. Schmid, A. Farhadi, and K. Alahari, “Actor and observer: Joint modeling of first and third-person videos,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 7396–7404

  6. [3]

    The epic-kitchens dataset: Collection, challenges and baselines,

    D. Damenet al., “The epic-kitchens dataset: Collection, challenges and baselines,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 43, no. 11, pp. 4125–4141, 2021

  7. [4]

    Ego4d: Around the world in 3,000 hours of egocentric video,

    K. Graumanet al., “Ego4d: Around the world in 3,000 hours of egocentric video,”IEEE Trans. Pattern Anal. Mach. Intell., pp. 1–32, 2024

  8. [5]

    Ego4d goal-step: Toward hierarchical understanding of procedural activities,

    Y . Song, E. Byrne, T. Nagarajan, H. Wang, M. Martin, and L. Torresani, “Ego4d goal-step: Toward hierarchical understanding of procedural activities,”Adv. Neural Inform. Process. Syst., vol. 36, 2024

Show all 217 references
  1. [6]

    Ego-humans: An ego-centric 3d multi-human benchmark,

    R. Khirodkar, A. Bansal, L. Ma, R. Newcombe, M. V o, and K. Kitani, “Ego-humans: An ego-centric 3d multi-human benchmark,” inInt. Conf. Comput. Vis., 2023, pp. 19 807–19 819

  2. [7]

    Identifying first-person camera wearers in third-person videos,

    C. Fanet al., “Identifying first-person camera wearers in third-person videos,” inIEEE Conf. Comput. Vis. Pattern Recog., July 2017

  3. [8]

    Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives,

    K. Graumanet al., “Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 19 383–19 400

  4. [9]

    Hd-epic: A highly-detailed egocentric video dataset,

    T. Perrettet al., “Hd-epic: A highly-detailed egocentric video dataset,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2025

  5. [10]

    Egocentric video-language pretraining,

    K. Q. Linet al., “Egocentric video-language pretraining,” inAdv. Neural Inform. Process. Syst., vol. 35, 2022, pp. 7575–7586

  6. [11]

    Egovlpv2: Egocentric video-language pre-training with fusion in the backbone,

    S. Pramanicket al., “Egovlpv2: Egocentric video-language pre-training with fusion in the backbone,” inInt. Conf. Comput. Vis., 2023, pp. 5262–5274

  7. [12]

    Helping hands: An object- aware ego-centric video recognition model,

    C. Zhang, A. Gupta, and A. Zisserman, “Helping hands: An object- aware ego-centric video recognition model,” inInt. Conf. Comput. Vis., 2023, pp. 13 855–13 866

  8. [13]

    Internvideo-ego4d: A pack of champion solutions to ego4d challenges,

    G. Chenet al., “Internvideo-ego4d: A pack of champion solutions to ego4d challenges,” 2022,arXiv: 2211.09529

  9. [14]

    Egovideo: Exploring egocentric foundation model and downstream adaptation,

    B. Peiet al., “Egovideo: Exploring egocentric foundation model and downstream adaptation,” 2024,arXiv:2406.18070

  10. [15]

    Modeling fine-grained hand-object dynamics for egocentric video representation learning,

    ——, “Modeling fine-grained hand-object dynamics for egocentric video representation learning,” 2025,arXiv:2503.00986

  11. [16]

    Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,

    A. Miech, D. Zhukov, J.-B. Alayrac, M. Tapaswi, I. Laptev, and J. Sivic, “Howto100m: Learning a text-video embedding by watching hundred million narrated video clips,”Int. Conf. Comput. Vis., pp. 2630–2640, 2019

  12. [17]

    Ava: A video dataset of spatio-temporally localized atomic visual actions,

    C. Guet al., “Ava: A video dataset of spatio-temporally localized atomic visual actions,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 6047–6056

  13. [18]

    Ucf101: A dataset of 101 human actions classes from videos in the wild,

    K. Soomro, A. R. Zamir, and M. Shah, “Ucf101: A dataset of 101 human actions classes from videos in the wild,” 2012,arXiv:1212.0402

  14. [19]

    Quo vadis, action recognition? a new model and the kinetics dataset,

    J. Carreira and A. Zisserman, “Quo vadis, action recognition? a new model and the kinetics dataset,”IEEE Conf. Comput. Vis. Pattern Recog., pp. 4724–4733, 2017

  15. [20]

    Internvid: A large-scale video-text dataset for multi- modal understanding and generation,

    Y . Wanget al., “Internvid: A large-scale video-text dataset for multi- modal understanding and generation,” 2023,arXiv: 2307.06942

  16. [21]

    Cg-bench: Clue-grounded question answering bench- mark for long video understanding,

    G. Chenet al., “Cg-bench: Clue-grounded question answering bench- mark for long video understanding,” 2024,arXiv: 2412.12075

  17. [22]

    Two-stream convolutional networks for action recognition in videos,

    K. Simonyan and A. Zisserman, “Two-stream convolutional networks for action recognition in videos,” inAdv. Neural Inform. Process. Syst. Cambridge, MA, USA: MIT Press, 2014, p. 568–576

  18. [23]

    Vivit: A video vision transformer,

    A. Arnab, M. Dehghani, G. Heigold, C. Sun, M. Lu ˇci´c, and C. Schmid, “Vivit: A video vision transformer,” inInt. Conf. Comput. Vis., 2021, pp. 6816–6826

  19. [24]

    Is space-time attention all you need for video understanding?

    G. Bertasius, H. Wang, and L. Torresani, “Is space-time attention all you need for video understanding?” inInt. Conf. Mach. Learn., July 2021

  20. [25]

    Dcan: improving temporal action detection via dual context aggregation,

    G. Chen, Y .-D. Zheng, L. Wang, and T. Lu, “Dcan: improving temporal action detection via dual context aggregation,” inAAAI Conf. Artif. Intell., vol. 36, no. 1, 2022, pp. 248–257

  21. [26]

    Video mamba suite: State space model as a versatile alternative for video understanding,

    G. Chenet al., “Video mamba suite: State space model as a versatile alternative for video understanding,” 2024,arXiv: 2403.09626

  22. [27]

    Memory-and- anticipation transformer for online action understanding,

    J. Wang, G. Chen, Y . Huang, L. Wang, and T. Lu, “Memory-and- anticipation transformer for online action understanding,” inInt. Conf. Comput. Vis., 2023, pp. 13 824–13 835

  23. [28]

    Video super-resolution based on deep learning: a comprehensive survey,

    H. Liuet al., “Video super-resolution based on deep learning: a comprehensive survey,”Artif Intell Rev, vol. 55, pp. 5981–6035, 2022

  24. [29]

    About time: Advances, challenges, and outlooks of action understanding,

    A. Stergiou and R. Poppe, “About time: Advances, challenges, and outlooks of action understanding,” 2024,arXiv:2411.15106

  25. [30]

    New generation deep learning for video object detection: A survey,

    L. Jiaoet al., “New generation deep learning for video object detection: A survey,”IEEE Trans. Neural Netw. Learn. Syst., vol. 33, no. 8, pp. 3195–3215, 2022

  26. [31]

    A survey of single- scene video anomaly detection,

    B. Ramachandra, M. J. Jones, and R. R. Vatsavai, “A survey of single- scene video anomaly detection,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 5, pp. 2293–2312, 2022

  27. [32]

    An outlook into the future of egocentric vision,

    C. Plizzariet al., “An outlook into the future of egocentric vision,”Int. J. Comput. Vis., pp. 1–57, 2024

  28. [33]

    Samsung’s new smart fridge lets you check in on its contents through internal cameras,

    N. Lavars, “Samsung’s new smart fridge lets you check in on its contents through internal cameras,” Jan 2016. [Online]. Available: https://newatlas.com/samsung-family-hub-smart-fridge/41192/

  29. [34]

    June oven,

    J. Oven, “June oven,” Aug 2018. [Online]. Available: https: //firewireblog.com/2018/08/19/june-oven/

  30. [35]

    Sportvu stats can be helpful, overwhelming,

    J. McDonald, “Sportvu stats can be helpful, overwhelming,” Nov

  31. [36]

    Hawk-eye’s eagle eye on wimbledon tennis,

    D. Winter, “Hawk-eye’s eagle eye on wimbledon tennis,” 2024. [Online]. Available: https://www.redsharknews.com/hawk-eyes-eye- on-wimbledon

  32. [37]

    Super bowl li preview: Inside fox sports’ “be the player

    J. Dachman, “Super bowl li preview: Inside fox sports’ “be the player” first-person pov replay tech,” Jan 2017. [Online]. Avail- able: https://www.sportsvideo.org/2017/01/13/super-bowl-li-preview- inside-fox-sports-be-the-player-360-pov-replay-technology/

  33. [38]

    Smart glasses in hospitals: Viewing care delivery through a new lens,

    H. M. Asia, “Smart glasses in hospitals: Viewing care delivery through a new lens,” HMA, 03 2022. [Online]. Available: https://www.hospitalmanagementasia.com/tech-innovation/ smart-glasses-in-hospitals-viewing-care-delivery-through-a-new-lens/

  34. [39]

    Here’s how asean’s first 5g smart hospital is using ai to usher in a new era of healthcare,

    ——, “Here’s how asean’s first 5g smart hospital is using ai to usher in a new era of healthcare,” HMA, 02 2022. [Online]. Available: https://www.hospitalmanagementasia.com/tech-innovation/heres-how- aseans-first-5g-smart-hospital-is-using-ai-to-usher-in-a-new-era-of- healthcare/

  35. [40]

    Cameras could be installed in classrooms in these states,

    K. Rahman, “Cameras could be installed in classrooms in these states,” Jan 2024. [Online]. Available: https://www.newsweek.com/cameras- installed-classrooms-1859098

  36. [41]

    Classroom camera: Transform education,

    Reolink, “Classroom camera: Transform education,” Nov 2024. [Online]. Available: https://reolink.com/blog/classroom-camera

  37. [42]

    3d surround view system,

    E. News, “3d surround view system,” Jun 2018. [Online]. Available: https://www.ien.eu/article/3d-surround-view-system/

  38. [43]

    Nauto launches real-time driver behavior learning platform for fleets,

    FreightWaves, “Nauto launches real-time driver behavior learning platform for fleets,” Nov 2019. [Online]. Available: https://finance. yahoo.com/news/nauto-launches-real-time-driver-144742897.html

  39. [44]

    Vision-centric semantic occupancy predic- tion for autonomous driving,

    P. L. Liu, “Vision-centric semantic occupancy predic- tion for autonomous driving,” May 2023. [Online]. Available: https://towardsdatascience.com/vision-centric-semantic- occupancy-prediction-for-autonomous-driving-16a46dbd6f65

  40. [45]

    Top 10 applications of robotics in 2024,

    harkiran78, “Top 10 applications of robotics in 2024,” geeksforgeeks, Feb 2024. [Online]. Available: https://www.geeksforgeeks.org/ applications-of-robotics

  41. [46]

    Research on body-worn cameras and law enforcement,

    N. I. of Justice, “Research on body-worn cameras and law enforcement,” National Institute of Justice, Jan 2022. [Online]. Available: https://nij.ojp.gov/topics/articles/research-body- worn-cameras-and-law-enforcement

  42. [47]

    Over 1,000 people saved with drone search and rescue: Dji,

    I. Singh, “Over 1,000 people saved with drone search and rescue: Dji,” DroneDJ, Jul 2023. [Online]. Available: https://dronedj.com/ 2023/07/12/dji-drone-search-rescue-map/

  43. [48]

    Empowering health & safety monitoring in manufacturing - fogsphere,

    Fogsphere, “Empowering health & safety monitoring in manufacturing - fogsphere,” Sep 2023. [Online]. Available: https://fogsphere.com/ industries-served/manufacturing/

  44. [49]

    A vision-guided robotic system designed to grab any object,

    R. Begg, “A vision-guided robotic system designed to grab any object,” Aug 2024. [Online]. Available: https://www.machinedesign.com/markets/robotics/video/55131589/ cynlr-a-vision-guided-robotic-system-designed-to-grab-any-object

  45. [50]

    How visual ai transforms assembly line operations in factories,

    Jobit, “How visual ai transforms assembly line operations in factories,”

  46. [51]

    Ego-exo: Transferring visual representations from third-person to first-person videos,

    Y . Li, T. Nagarajan, B. Xiong, and K. Grauman, “Ego-exo: Transferring visual representations from third-person to first-person videos,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 6943–6953

  47. [52]

    Cross-view action recognition un- derstanding from exocentric to egocentric perspective,

    T.-D. Truong and K. Luu, “Cross-view action recognition un- derstanding from exocentric to egocentric perspective,” 2023, arXiv:2305.15699

  48. [53]

    Unlocking exocentric video-language data for ego- centric video representation learning,

    Z.-Y . Douet al., “Unlocking exocentric video-language data for ego- centric video representation learning,” 2024,arXiv:2408.03567

  49. [54]

    Holographic feature learning of egocentric-exocentric videos for multi-domain action recognition,

    Y . Huang, X. Yang, J. Gao, and C. Xu, “Holographic feature learning of egocentric-exocentric videos for multi-domain action recognition,” IEEE Trans. Multimedia, vol. 24, pp. 2273–2286, 2022. 17

  50. [55]

    Synchronization is all you need: Exocentric-to-egocentric transfer for temporal action segmentation with unlabeled synchronized video pairs,

    C. Quattrocchi, A. Furnari, D. Di Mauro, M. V . Giuffrida, and G. M. Farinella, “Synchronization is all you need: Exocentric-to-egocentric transfer for temporal action segmentation with unlabeled synchronized video pairs,” 2023,arXiv:2312.02638

  51. [56]

    Action recognition in the presence of one egocentric and multiple static cameras,

    B. Soran, A. Farhadi, and L. G. Shapiro, “Action recognition in the presence of one egocentric and multiple static cameras,” inLect. Notes Comput. Sci., 2014

  52. [57]

    Cross-view exocentric to egocentric video synthesis,

    G. Liu, H. Tang, H. Latapie, J. J. Corso, and Y . Yan, “Cross-view exocentric to egocentric video synthesis,”ACM Int. Conf. Multimedia, 2021

  53. [58]

    4diff: 3d-aware diffusion model for third-to-first viewpoint translation,

    F. Chenget al., “4diff: 3d-aware diffusion model for third-to-first viewpoint translation,” inEur. Conf. Comput. Vis., 2024

  54. [59]

    Exo2egodvc: Dense video captioning of ego- centric procedural activities using web instructional videos,

    T. Ohkawaet al., “Exo2egodvc: Dense video captioning of ego- centric procedural activities using web instructional videos,” 2023, arXiv:2311.16444

  55. [60]

    Retrieval-augmented egocentric video captioning,

    J. Xuet al., “Retrieval-augmented egocentric video captioning,” in IEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 13 525–13 536

  56. [62]

    Enhancing egocentric 3d pose estimation with third person views,

    A. Dhamanaskar, M. Dimiccoli, E. Corona, A. Pumarola, and F. Moreno-Noguer, “Enhancing egocentric 3d pose estimation with third person views,”Pattern Recognition, vol. 138, p. 109358, 2023

  57. [63]

    An exocentric look at egocentric actions and vice versa,

    S. Ardeshir and A. Borji, “An exocentric look at egocentric actions and vice versa,”Comput Vision Image Understanding, pp. 61–68, 2018

  58. [64]

    Learning fine-grained view-invariant representations from unpaired ego-exo videos via temporal alignment,

    Z. S. Xue and K. Grauman, “Learning fine-grained view-invariant representations from unpaired ego-exo videos via temporal alignment,” inAdv. Neural Inform. Process. Syst., vol. 36, 2023, pp. 53 688–53 710

  59. [65]

    Camera selection for occlusion-less surgery recording via training with an egocentric camera,

    Y . Saito, R. Hachiuma, H. Saito, H. Kajita, Y . Takatsume, and T. Hayashida, “Camera selection for occlusion-less surgery recording via training with an egocentric camera,”IEEE Access, vol. 9, pp. 138 307–138 322, 2021

  60. [66]

    Next-generation surgical navigation: Marker-less multi-view 6dof pose estimation of surgical instruments,

    J. Heinet al., “Next-generation surgical navigation: Marker-less multi-view 6dof pose estimation of surgical instruments,” 2023, arXiv:2305.03535

  61. [67]

    Look both ways: Self-supervising driver gaze estimation and road scene saliency,

    I. Kasahara, S. Stent, and H. S. Park, “Look both ways: Self-supervising driver gaze estimation and road scene saliency,” inEur. Conf. Comput. Vis.Springer, 2022, pp. 126–142

  62. [68]

    Aide: A vision-driven multi-view, multi-modal, multi- tasking dataset for assistive driving perception,

    D. Yanget al., “Aide: A vision-driven multi-view, multi-modal, multi- tasking dataset for assistive driving perception,” inInt. Conf. Comput. Vis., 2023, pp. 20 402–20 413

  63. [69]

    Vision-based ma- nipulators need to also see from their hands,

    K. Hsu, M. J. Kim, R. Rafailov, J. Wu, and C. Finn, “Vision-based ma- nipulators need to also see from their hands,” 2022,arXiv:2203.12677

  64. [70]

    Look closer: Bridging egocentric and third-person views with transformers for robotic manipulation,

    R. Jangir, N. Hansen, S. Ghosal, M. Jain, and X. Wang, “Look closer: Bridging egocentric and third-person views with transformers for robotic manipulation,”IEEE Trans. Robot. Autom., vol. 7, no. 2, pp. 3046–3053, 2022

  65. [71]

    Self-supervised disentangled representation learning for third-person imitation learning,

    J. Shang and M. S. Ryoo, “Self-supervised disentangled representation learning for third-person imitation learning,” inIEEE Int. Conf. Intell. Rob. Syst., 2021, pp. 214–221

  66. [72]

    Visual-policy learning through multi-camera view to single-camera view knowledge distilla- tion for robot manipulation tasks,

    C. Acar, K. Binici, A. Tekirda ˘g, and Y . Wu, “Visual-policy learning through multi-camera view to single-camera view knowledge distilla- tion for robot manipulation tasks,”IEEE Trans. Robot. Autom., vol. 9, no. 1, pp. 691–698, 2024

  67. [73]

    Third-person visual imitation learning via decoupled hierarchical controller,

    P. Sharma, D. Pathak, and A. K. Gupta, “Third-person visual imitation learning via decoupled hierarchical controller,” inAdv. Neural Inform. Process. Syst., 2019

  68. [74]

    Multi-view disentanglement for rein- forcement learning with multiple cameras,

    M. Dunion and S. V . Albrecht, “Multi-view disentanglement for rein- forcement learning with multiple cameras,” 2024,arXiv:2404.14064

  69. [75]

    Multi-view masked world models for visual robotic manipulation,

    Y . Seo, J. Kim, S. James, K. Lee, J. Shin, and P. Abbeel, “Multi-view masked world models for visual robotic manipulation,” inInt. Conf. Mach. Learn., 2023

  70. [76]

    Learning fused state representations for control from multi-view observations,

    Z. Wang, Y .-H. Li, X. Li, H. Zang, R. Laroche, and R. Islam, “Learning fused state representations for control from multi-view observations,” 2025,arXiv: 2502.01316

  71. [77]

    Self-explainable affordance learning with embodied caption,

    Z. Zhang, Z. Wei, G. Sun, P. Wang, and L. Van Gool, “Self-explainable affordance learning with embodied caption,” 2024,arXiv:2404.05603

  72. [78]

    Learning granularity-aware affordances from human- object interaction for tool-based functional grasping in dexterous robotics,

    F. Yanget al., “Learning granularity-aware affordances from human- object interaction for tool-based functional grasping in dexterous robotics,” 2024,arXiv:2407.00614

  73. [79]

    Towards third-person visual imitation learning using generative adversarial networks,

    L. Garello, F. Rea, N. Noceti, and A. Sciutti, “Towards third-person visual imitation learning using generative adversarial networks,” in IEEE Int. Conf. Dev. Learn.IEEE, 2022, pp. 121–126

  74. [80]

    Diffusing in someone else’s shoes: Robotic perspective taking with diffusion,

    J. Spisak, M. Kerzel, and S. Wermter, “Diffusing in someone else’s shoes: Robotic perspective taking with diffusion,” 2024, arXiv:2404.07735

  75. [81]

    Ego-to-exo: Interfacing third person visuals from egocentric views in real-time for improved rov teleoperation,

    A. Abdullah, R. Chen, I. Rekleitis, and M. J. Islam, “Ego-to-exo: Interfacing third person visuals from egocentric views in real-time for improved rov teleoperation,” 2024,arXiv:2407.00848

  76. [82]

    Spatial assisted human-drone collabo- rative navigation and interaction through immersive mixed reality,

    L. Morando and G. Loianno, “Spatial assisted human-drone collabo- rative navigation and interaction through immersive mixed reality,” in Int. Conf. Robot. Autom.IEEE, 2024, pp. 8707–8713

  77. [83]

    Yowo: You only walk once to jointly map an indoor scene and register ceiling-mounted cameras,

    F. Yang, S. Yamao, I. Kusajima, A. Moteki, S. Masui, and S. Jiang, “Yowo: You only walk once to jointly map an indoor scene and register ceiling-mounted cameras,”IEEE Trans. Circuit Syst. Video Technol., pp. 1–1, 2024

  78. [84]

    Fourth- person captioning: Describing daily events by uni-supervised and tri- regularized training,

    K. Nakashima, Y . Iwashita, A. Kawamura, and R. Kurazume, “Fourth- person captioning: Describing daily events by uni-supervised and tri- regularized training,” inIEEE Trans. Syst., Man, Cybern., 2018, pp. 2122–2127

  79. [85]

    Put myself in your shoes: Lifting the egocentric perspective from exocentric videos,

    M. Luo, Z. Xue, A. Dimakis, and K. Grauman, “Put myself in your shoes: Lifting the egocentric perspective from exocentric videos,” in Eur. Conf. Comput. Vis.Springer, 2025, pp. 407–425

  80. [86]

    Animate any- one: Consistent and controllable image-to-video synthesis for character animation,

    L. Hu, X. Gao, P. Zhang, K. Sun, B. Zhang, and L. Bo, “Animate any- one: Consistent and controllable image-to-video synthesis for character animation,”IEEE Conf. Comput. Vis. Pattern Recog., pp. 8153–8163, 2023

  81. [87]

    Stable video diffusion: Scaling latent video diffusion models to large datasets,

    A. Blattmannet al., “Stable video diffusion: Scaling latent video diffusion models to large datasets,” 2023,arXiv:2311.15127

  82. [88]

    Cogvideox: Text-to-video diffusion models with an expert transformer,

    Z. Yanget al., “Cogvideox: Text-to-video diffusion models with an expert transformer,” 2024,arXiv:2408.06072

  83. [89]

    Structure and content-guided video synthesis with diffusion models,

    P. Esser, J. Chiu, P. Atighehchian, J. Granskog, and A. Germanidis, “Structure and content-guided video synthesis with diffusion models,” inInt. Conf. Comput. Vis., 2023, pp. 7312–7322

  84. [90]

    Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory,

    S. Yinet al., “Dragnuwa: Fine-grained control in video generation by integrating text, image, and trajectory,” 2023,arXiv:2308.08089

  85. [91]

    Intention-driven ego-to-exo video generation,

    H. Luo, K. Zhu, W. Zhai, and Y . Cao, “Intention-driven ego-to-exo video generation,” 2024,arXiv:2403.09194

  86. [92]

    From my view to yours: Ego- augmented learning in large vision language models for understanding exocentric daily living activities,

    D. Reilly, M. K. Govind, and S. Das, “From my view to yours: Ego- augmented learning in large vision language models for understanding exocentric daily living activities,” 2025,arXiv:2501.05711

  87. [93]

    Viewbirdiformer: Learn- ing to recover ground-plane crowd trajectories and ego-motion from a single ego-centric view,

    M. Nishimura, S. Nobuhara, and K. Nishino, “Viewbirdiformer: Learn- ing to recover ground-plane crowd trajectories and ego-motion from a single ego-centric view,”IEEE Trans. Robot. Autom., vol. 8, no. 1, pp. 368–375, 2023

  88. [94]

    View birdification in the crowd: Ground-plane localization from perceived movements,

    ——, “View birdification in the crowd: Ground-plane localization from perceived movements,”Int. J. Comput. Vis., vol. 131, no. 8, p. 2015–2031, May 2023

  89. [95]

    Incrowdformer: On-ground pedestrian world model from ego- centric views,

    ——, “Incrowdformer: On-ground pedestrian world model from ego- centric views,” 2023,arXiv:2303.09534

  90. [96]

    Fastllve: Real-time low- light video enhancement with intensity-aware look-up table,

    W. Li, G. Wu, W. Wang, P. Ren, and X. Liu, “Fastllve: Real-time low- light video enhancement with intensity-aware look-up table,” inACM Int. Conf. Multimedia, 2023, pp. 8134–8144

  91. [97]

    Flow-guided sparse transformer for video deblurring,

    J. Linet al., “Flow-guided sparse transformer for video deblurring,” in Int. Conf. Mach. Learn., vol. 162, 17–23 Jul 2022, pp. 13 334–13 343

  92. [98]

    Blur interpolation transformer for real-world motion from blur,

    Z. Zhong, M. Cao, X. Ji, Y . Zheng, and I. Sato, “Blur interpolation transformer for real-world motion from blur,” inInt. Conf. Comput. Vis., 2023, pp. 5713–5723

  93. [99]

    Surgery recording without occlusions by multi-view surgical videos,

    T. Shimizu, K. Oishi, R. Hachiuma, H. Kajita, Y . Takatsume, and H. Saito, “Surgery recording without occlusions by multi-view surgical videos,” in15th International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications, 2020, pp. 837–844

  94. [100]

    Deep selection: A fully supervised camera selection network for surgery recordings,

    R. Hachiuma, T. Shimizu, H. Saito, H. Kajita, and Y . Takatsume, “Deep selection: A fully supervised camera selection network for surgery recordings,” inMedical Image Computing and Computer Assisted Intervention. Springer, 2020, pp. 419–428

  95. [101]

    From third person to first person: Dataset and baselines for synthesis and retrieval,

    M. Elfeki, K. Regmi, S. Ardeshir, and A. Borji, “From third person to first person: Dataset and baselines for synthesis and retrieval,” 2018, arXiv:1812.00104

  96. [102]

    Cross- view exocentric to egocentric video synthesis,

    G. Liu, H. Tang, H. M. Latapie, J. J. Corso, and Y . Yan, “Cross- view exocentric to egocentric video synthesis,” inACM Int. Conf. Multimedia, 2021, pp. 974–982

  97. [103]

    Parallel generative adversarial network for third-person to first-person image generation,

    G. Liu, H. Latapie, O. Kilic, and A. Lawrence, “Parallel generative adversarial network for third-person to first-person image generation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 1917–1923

  98. [104]

    Egoexo-gen: Ego-centric video prediction by watching exo-centric videos,

    J. Xuet al., “Egoexo-gen: Ego-centric video prediction by watching exo-centric videos,” 2025,arXiv:2504.11732

  99. [105]

    Sibnet: Sibling convolutional encoder for video captioning,

    S. Liu, Z. Ren, and J. Yuan, “Sibnet: Sibling convolutional encoder for video captioning,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 43, pp. 3259–3272, 2018. 18

  100. [106]

    Spatio-temporal graph for video captioning with knowledge distillation,

    B. Panet al., “Spatio-temporal graph for video captioning with knowledge distillation,”IEEE Conf. Comput. Vis. Pattern Recog., pp. 10 867–10 876, 2020

  101. [107]

    Domain-specific semantics guided approach to video captioning,

    H. Munusamy and C. C. Sekhar, “Domain-specific semantics guided approach to video captioning,”IEEE Winter Conf. Appl. Comput. Vis., pp. 1576–1585, 2020

  102. [108]

    Worldscribe: Towards context-aware live visual descriptions,

    R.-C. Chang, Y . Liu, and A. Guo, “Worldscribe: Towards context-aware live visual descriptions,” inACM Symp. User Interface Softw. Technol., 2024, pp. 1–18

  103. [109]

    Wanderguide: Indoor map-less robotic guide for exploration by blind people,

    M. Kuribayashi, K. Uehara, A. Wang, S. Morishima, and C. Asakawa, “Wanderguide: Indoor map-less robotic guide for exploration by blind people,” 2025,arXiv:2502.08906

  104. [110]

    Audio-adaptive activity recognition across video domains,

    Y . Zhang, H. Doughty, L. Shao, and C. G. Snoek, “Audio-adaptive activity recognition across video domains,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 13 791–13 800

  105. [111]

    Cross-view generalisation in action recognition: Feature design for transitioning from exocentric to egocentric views,

    B. Rocha, P. Moreno, and A. Bernardino, “Cross-view generalisation in action recognition: Feature design for transitioning from exocentric to egocentric views,” inIberian Robotics Conference, 2023, pp. 155–166

  106. [112]

    Learning from semantic alignment between unpaired multiviews for egocentric video recognition,

    Q. Wang, L. Zhao, L. Yuan, T. Liu, and X. Peng, “Learning from semantic alignment between unpaired multiviews for egocentric video recognition,” inInt. Conf. Comput. Vis., 2023, pp. 3284–3294

  107. [113]

    Unsupervised and semi-supervised domain adaptation for action recognition from drones,

    J. Choi, G. Sharma, M. Chandraker, and J.-B. Huang, “Unsupervised and semi-supervised domain adaptation for action recognition from drones,” inIEEE Winter Conf. Appl. Comput. Vis., 2020, pp. 1717– 1726

  108. [114]

    Egofish3d: Egocentric 3d pose estimation from a fisheye camera via self- supervised learning,

    Y . Liu, J. Yang, X. Gu, Y . Chen, Y . Guo, and G.-Z. Yang, “Egofish3d: Egocentric 3d pose estimation from a fisheye camera via self- supervised learning,”IEEE Trans. Multimedia, pp. 8880–8891, 2023

  109. [115]

    Ex2eg-mae: A framework for adaptation of exocentric video masked autoencoders for egocentric social role understanding,

    M. Tran, Y . Kim, C.-C. Su, C.-H. Kuo, M. Sun, and M. Soleymani, “Ex2eg-mae: A framework for adaptation of exocentric video masked autoencoders for egocentric social role understanding,” inEur. Conf. Comput. Vis.Springer, 2025, pp. 1–19

  110. [116]

    The audio-visual conversational graph: From an egocentric-exocentric perspective,

    W. Jiaet al., “The audio-visual conversational graph: From an egocentric-exocentric perspective,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 26 396–26 405

  111. [117]

    Learning to score figure skating sport videos,

    C. Xu, Y . Fu, B. Zhang, Z. Chen, Y .-G. Jiang, and X. Xue, “Learning to score figure skating sport videos,”IEEE Trans. Circuit Syst. Video Technol., vol. 30, no. 12, pp. 4578–4590, 2019

  112. [118]

    Expertaf: Expert actionable feedback from video,

    K. Ashutosh, T. Nagarajan, G. Pavlakos, K. Kitani, and K. Grau- man, “Expertaf: Expert actionable feedback from video,” 2024, arXiv:2408.00672

  113. [119]

    Towards universal soccer video understanding,

    J. Rao, H. Wu, H. Jiang, Y . Zhang, Y . Wang, and W. Xie, “Towards universal soccer video understanding,” 2024,arXiv:2412.01820

  114. [120]

    Learning 6-dof fine-grained grasp detection based on part affordance grounding,

    Y . Songet al., “Learning 6-dof fine-grained grasp detection based on part affordance grounding,” 2024,arXiv: 2301.11564

  115. [121]

    Robo-abc: Affordance generalization beyond categories via semantic correspon- dence for robot manipulation,

    Y . Ju, K. Hu, G. Zhang, G. Zhang, M. Jiang, and H. Xu, “Robo-abc: Affordance generalization beyond categories via semantic correspon- dence for robot manipulation,” inEur. Conf. Comput. Vis., 2025, pp. 222–239

  116. [122]

    Learning affordance grounding from exocentric images,

    H. Luo, W. Zhai, J. Zhang, Y . Cao, and D. Tao, “Learning affordance grounding from exocentric images,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2022, pp. 2252–2261

  117. [123]

    Locate: Localize and transfer object parts for weakly supervised affordance grounding,

    G. Li, V . Jampani, D. Sun, and L. Sevilla-Lara, “Locate: Localize and transfer object parts for weakly supervised affordance grounding,” in IEEE Conf. Comput. Vis. Pattern Recog., June 2023, pp. 10 922–10 931

  118. [124]

    Weakly supervised multimodal affordance grounding for egocentric images,

    L. Xu, Y . Gao, W. Song, and A. Hao, “Weakly supervised multimodal affordance grounding for egocentric images,” inAAAI Conf. Artif. Intell., vol. 38, no. 6, 2024, pp. 6324–6332

  119. [125]

    Strategies to leverage foun- dational model knowledge in object affordance grounding,

    A. Rai, K. Buettner, and A. Kovashka, “Strategies to leverage foun- dational model knowledge in object affordance grounding,” inIEEE. Conf. Comput. Vis. Pattern Recog. Workshops., June 2024, pp. 1714– 1723

  120. [126]

    Learning deep features for discriminative localization,

    B. Zhou, A. Khosla, A. Lapedriza, A. Oliva, and A. Torralba, “Learning deep features for discriminative localization,” inIEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 2921–2929

  121. [127]

    Learning transferable visual models from natural language supervision,

    A. Radfordet al., “Learning transferable visual models from natural language supervision,” inInt. Conf. Mach. Learn., 2021

  122. [128]

    Intra: Interaction relationship-aware weakly supervised affordance grounding,

    J. H. Jang, H. Seo, and S. Y . Chun, “Intra: Interaction relationship-aware weakly supervised affordance grounding,” 2024, arXiv:2409.06210

  123. [129]

    Visual affordance and function understanding: A survey,

    M. Hassanin, S. Khan, and M. Tahtali, “Visual affordance and function understanding: A survey,”ACM Comput. Surv., vol. 54, no. 3, Apr. 2021

  124. [130]

    Emergencynet: Efficient aerial image classification for drone-based emergency monitoring using atrous con- volutional feature fusion,

    C. Kyrkou and T. Theocharides, “Emergencynet: Efficient aerial image classification for drone-based emergency monitoring using atrous con- volutional feature fusion,”IEEE J. Sel. Topics Appl. Earth Observations Remote Sens., vol. 13, pp. 1687–1699, 2020

  125. [131]

    Toward a roadmap for human-drone interaction,

    J. R. Cauchard, M. Khamis, J. Garcia, M. Kljun, and A. M. Brock, “Toward a roadmap for human-drone interaction,”Interactions, vol. 28, pp. 76–81, 03 2021

  126. [134]

    Third-person piloting: Increasing situational awareness using a spa- tially coupled second drone,

    R. Temma, K. Takashima, K. Fujita, K. Sueda, and Y . Kitamura, “Third-person piloting: Increasing situational awareness using a spa- tially coupled second drone,” inACM Symp. User Interface Softw. Technol., 2019, pp. 507–519

  127. [136]

    Vinci: A real-time embodied smart assistant based on egocentric vision-language model,

    Y . Huanget al., “Vinci: A real-time embodied smart assistant based on egocentric vision-language model,” 2024,arXiv: 2412.21080

  128. [137]

    Lita: Language instructed temporal-localization assistant,

    D.-A. Huanget al., “Lita: Language instructed temporal-localization assistant,” inEur. Conf. Comput. Vis., 2024

  129. [138]

    An egocentric vision-language model based portable real-time smart assistant,

    Y . Huanget al., “An egocentric vision-language model based portable real-time smart assistant,” 2025,arXiv:2503.04250

  130. [139]

    Lifelogging caption generation via fourth-person vision in a human–robot symbiotic envi- ronment,

    K. Nakashima, Y . Iwashita, and R. Kurazume, “Lifelogging caption generation via fourth-person vision in a human–robot symbiotic envi- ronment,”Robomech J., vol. 7, 2020

  131. [140]

    Faster r-cnn: Towards real-time object detection with region proposal networks,

    S. Ren, K. He, R. B. Girshick, and J. Sun, “Faster r-cnn: Towards real-time object detection with region proposal networks,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 39, pp. 1137–1149, 2015

  132. [141]

    Temporally-weighted hierarchical clustering for un- supervised action segmentation,

    M. S. Sarfraz, N. Murray, V . Sharma, A. Diba, L. V . Gool, and R. Stiefelhagen, “Temporally-weighted hierarchical clustering for un- supervised action segmentation,”IEEE Conf. Comput. Vis. Pattern Recog., pp. 11 220–11 229, 2021

  133. [142]

    Improving action segmentation via graph-based temporal reasoning,

    Y . Huang, Y . Sugano, and Y . Sato, “Improving action segmentation via graph-based temporal reasoning,” inIEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 14 024–14 034

  134. [143]

    Collaborative weakly supervised video correlation learn- ing for procedure-aware instructional video analysis,

    T. Heet al., “Collaborative weakly supervised video correlation learn- ing for procedure-aware instructional video analysis,”AAAI Conf. Artif. Intell., vol. 38, no. 3, pp. 2112–2120, Mar. 2024

  135. [144]

    Svip: Sequence verification for procedures in videos,

    Y . Qian, W. Luo, D. Lian, X. Tang, P. Zhao, and S. Gao, “Svip: Sequence verification for procedures in videos,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 19 858–19 870

  136. [145]

    Mutual context network for jointly estimating egocentric gaze and action,

    Y . Huang, M. Cai, Z. Li, F. Lu, and Y . Sato, “Mutual context network for jointly estimating egocentric gaze and action,”IEEE Transactions on Image Processing, vol. 29, pp. 7795–7806, 2020

  137. [146]

    Egotransfer: Transferring motion across egocentric and exocentric domains using deep neural networks,

    S. Ardeshir, K. Regmi, and A. Borji, “Egotransfer: Transferring motion across egocentric and exocentric domains using deep neural networks,” 2016,arXiv:1612.05836

  138. [147]

    First- and third-person video co-analysis by learning spatial-temporal joint attention,

    H. Yu, M. Cai, Y . Liu, and F. Lu, “First- and third-person video co-analysis by learning spatial-temporal joint attention,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 6, pp. 6631–6646, 2023

  139. [148]

    Psalm: Pixelwise segmentation with large multi-modal model,

    Z. Zhang, Y . Ma, E. Zhang, and X. Bai, “Psalm: Pixelwise segmentation with large multi-modal model,” inEur. Conf. Comput. Vis., 2025, pp. 74–91

  140. [149]

    Objectrelator: Enabling cross-view object relation understanding in ego-centric and exo-centric videos,

    Y . Fu, R. Wang, Y . Fu, D. P. Paudel, X. Huang, and L. V . Gool, “Objectrelator: Enabling cross-view object relation understanding in ego-centric and exo-centric videos,” 2024,arXiv:2411.19083

  141. [150]

    Relating view directions of complementary-view mobile cameras via the human shadow,

    R. Han, Y . Gan, L. Wang, N. Li, W. Feng, and S. Wang, “Relating view directions of complementary-view mobile cameras via the human shadow,”Int. J. Comput. Vis., vol. 131, no. 5, p. 1106–1121, Jan. 2023

  142. [151]

    From a bird’s eye view to see: Joint camera and subject registration without the camera calibration,

    Z. Qian, R. Han, W. Feng, F. F. Wang, and S. Wang, “From a bird’s eye view to see: Joint camera and subject registration without the camera calibration,”IEEE Conf. Comput. Vis. Pattern Recog., pp. 863–873, 2022

  143. [152]

    Cal- ibration of non-overlapping cameras using an external slam system,

    E. Ataer-Cansizoglu, Y . Taguchi, S. Ramalingam, and Y . Miki, “Cal- ibration of non-overlapping cameras using an external slam system,” in2014 2nd International Conference on 3D Vision, vol. 1, 2014, pp. 509–516

  144. [153]

    Egolocate: Real-time motion capture, localization, and mapping with sparse body-mounted sensors,

    X. Yiet al., “Egolocate: Real-time motion capture, localization, and mapping with sparse body-mounted sensors,”ACM Transactions on Graphics, vol. 42, no. 4, 2023

  145. [154]

    Recognizing micro-actions and reactions from paired egocentric videos,

    R. Yonetani, K. M. Kitani, and Y . Sato, “Recognizing micro-actions and reactions from paired egocentric videos,” inIEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 2629–2638

  146. [155]

    Multi-view object pose estimation from correspondence distributions and epipolar geometry,

    R. L. Haugaard and T. M. Iversen, “Multi-view object pose estimation from correspondence distributions and epipolar geometry,” inInt. Conf. Robot. Autom., 2023, pp. 1786–1792. 19

  147. [156]

    Tell, don’t show: Language guidance eases transfer across domains in images and videos,

    T. Kalluri, B. P. Majumder, and M. Chandraker, “Tell, don’t show: Language guidance eases transfer across domains in images and videos,” inInt. Conf. Mach. Learn., Jul 2024, pp. 22 879–22 894

  148. [157]

    Pov: Prompt-oriented view-agnostic learning for egocentric hand-object interaction in the multi-view world,

    B. Xu, S. Zheng, and Q. Jin, “Pov: Prompt-oriented view-agnostic learning for egocentric hand-object interaction in the multi-view world,” inACM Int. Conf. Multimedia, 2023, pp. 2807–2816

  149. [158]

    Masked video and body- worn imu autoencoder for egocentric action recognition,

    M. Zhang, Y . Huang, R. Liu, and Y . Sato, “Masked video and body- worn imu autoencoder for egocentric action recognition,” inEur. Conf. Comput. Vis., 2025, pp. 312–330

  150. [159]

    Ego2top: Matching viewers in egocentric and top-view videos,

    S. Ardeshir and A. Borji, “Ego2top: Matching viewers in egocentric and top-view videos,” inEur. Conf. Comput. Vis., 2016, pp. 253–268

  151. [160]

    Egocentric meets top-view,

    ——, “Egocentric meets top-view,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 41, no. 6, pp. 1353–1366, 2019

  152. [161]

    Integrating egocentric videos in top-view surveillance videos: Joint identification and temporal alignment,

    ——, “Integrating egocentric videos in top-view surveillance videos: Joint identification and temporal alignment,” inEur. Conf. Comput. Vis., Sep 2018

  153. [162]

    Class segmentation and object localization with superpixel neighborhoods,

    B. Fulkerson, A. Vedaldi, and S. Soatto, “Class segmentation and object localization with superpixel neighborhoods,”Int. Conf. Comput. Vis., pp. 670–677, 2009

  154. [163]

    Joint person segmentation and identification in synchronized first- and third-person videos,

    M. Xu, C. Fan, Y . Wang, M. S. Ryoo, and D. J. Crandall, “Joint person segmentation and identification in synchronized first- and third-person videos,” inEur. Conf. Comput. Vis., 2018, pp. 656–672

  155. [164]

    Fusing personal and environmental cues for identification and segmentation of first-person camera wearers in third-person views,

    Z. Zhao, Y . Wang, and C. Wang, “Fusing personal and environmental cues for identification and segmentation of first-person camera wearers in third-person views,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 16 477–16 487

  156. [165]

    Visual-gps: Ego-downward and ambient video based person location association,

    L. Yang, H. Jiang, Z. Huo, and J. Xiao, “Visual-gps: Ego-downward and ambient video based person location association,” inIEEE Conf. Comput. Vis. Pattern Recog Workshops., 2019, pp. 371–380

  157. [166]

    Seeing the unseen: Predicting the first-person camera wearer’s location and pose in third-person scenes,

    Y . Wen, K. K. Singh, M. Anderson, W.-P. Jan, and Y . J. Lee, “Seeing the unseen: Predicting the first-person camera wearer’s location and pose in third-person scenes,” inInt. Conf. Comput. Vis., 2021, pp. 3446–3455

  158. [167]

    Egopca: A new framework for egocentric hand-object interaction understanding,

    Y . Xuet al., “Egopca: A new framework for egocentric hand-object interaction understanding,” inInt. Conf. Comput. Vis., 2023, pp. 5273– 5284

  159. [168]

    Ego- centric action recognition by capturing hand-object contact and object state,

    T. Shiota, M. Takagi, K. Kumagai, H. Seshimo, and Y . Aono, “Ego- centric action recognition by capturing hand-object contact and object state,” inIEEE Winter Conf. Appl. Comput. Vis., 2024, pp. 6541–6551

  160. [169]

    Egochoir: Capturing 3d human-object interaction regions from egocentric views,

    Y . Yang, W. Zhai, C. Wang, C. Yu, Y . Cao, and Z.-J. Zha, “Egochoir: Capturing 3d human-object interaction regions from egocentric views,” Adv. Neural Inform. Process. Syst., vol. 37, pp. 54 529–54 557, 2025

  161. [170]

    The audio-visual conversational graph: From an egocentric-exocentric perspective,

    W. Jia, M. Liu, H. Jiang, I. Ananthabhotla, J. M. Rehg, V . K. Ithapu, and R. Gao, “The audio-visual conversational graph: From an egocentric-exocentric perspective,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 26 386–26 395

  162. [171]

    Complementary-view co-interest person detection,

    R. Han, J. Zhao, W. Feng, Y . Gan, L. Wan, and S. Wang, “Complementary-view co-interest person detection,” inACM Int. Conf. Multimedia, 2020, p. 2746–2754

  163. [172]

    Complementary-view multiple human tracking,

    R. Han, W. Feng, J. Zhao, Z. Niu, Y . Zhang, and L. Wan, “Complementary-view multiple human tracking,” inAAAI Conf. Artif. Intell., vol. 34, 02 2020

  164. [173]

    Multiple human as- sociation and tracking from egocentric and complementary top views,

    R. Han, W. Feng, Y . Zhang, J. Zhao, and S. Wang, “Multiple human as- sociation and tracking from egocentric and complementary top views,” IEEE Trans. Pattern Anal. Mach. Intell., pp. 5225–5242, 2022

  165. [174]

    Connecting the complementary-view videos: Joint camera identification and subject association,

    R. Han, Y . Gan, J. Li, F. Wang, W. Feng, and S. Wang, “Connecting the complementary-view videos: Joint camera identification and subject association,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 2406–2415

  166. [175]

    You only look once: Unified, real-time object detection,

    J. Redmon, S. K. Divvala, R. B. Girshick, and A. Farhadi, “You only look once: Unified, real-time object detection,”IEEE Conf. Comput. Vis. Pattern Recog., pp. 779–788, 2015

  167. [176]

    Learning visual robotic control efficiently with contrastive pre-training and data augmentation,

    A. Zhan, R. Zhao, L. Pinto, P. Abbeel, and M. Laskin, “Learning visual robotic control efficiently with contrastive pre-training and data augmentation,” inIEEE Int. Conf. Intell. Rob. Syst., 2022, pp. 4040– 4047

  168. [177]

    Multi-view masked world models for visual robotic manipulation,

    Y . Seo, J. Kim, S. James, K. Lee, J. Shin, and P. Abbeel, “Multi-view masked world models for visual robotic manipulation,” inInt. Conf. Mach. Learn., vol. 202, 23–29 Jul 2023, pp. 30 613–30 632

  169. [178]

    Rt-1: Robotics transformer for real-world control at scale,

    A. Brohanet al., “Rt-1: Robotics transformer for real-world control at scale,” 2022,arXiv:2212.06817

  170. [179]

    Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware,

    T. Z. Zhao, V . Kumar, S. Levine, and C. Finn, “Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware,” inProc. Robot.: Sci. Syst., Daegu, Republic of Korea, July 2023

  171. [180]

    Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking,

    H. Bharadhwaj, J. Vakil, M. Sharma, A. Gupta, S. Tulsiani, and V . Ku- mar, “Roboagent: Generalization and efficiency in robot manipulation via semantic augmentations and action chunking,” inInt. Conf. Robot. Autom., 2024, pp. 4788–4795

  172. [181]

    Diffusion policy: Visuomotor policy learning via action diffusion,

    C. Chiet al., “Diffusion policy: Visuomotor policy learning via action diffusion,” inProc. Robot.: Sci. Syst., 2023

  173. [182]

    Aloha unleashed: A simple recipe for robot dexterity,

    T. Z. Zhaoet al., “Aloha unleashed: A simple recipe for robot dexterity,” 2024,arXiv:2410.13126

  174. [183]

    Deep 360 pilot: Learning a deep agent for piloting through 360 sports videos,

    H.-N. Hu, Y .-C. Lin, M.-Y . Liu, H.-T. Cheng, Y .-J. Chang, and M. Sun, “Deep 360 pilot: Learning a deep agent for piloting through 360 sports videos,” inIEEE Conf. Comput. Vis. Pattern Recog., 2017, pp. 1396– 1405

  175. [184]

    Making 360 video watchable in 2d: Learning videography for click free viewing,

    Y .-C. Su and K. Grauman, “Making 360 video watchable in 2d: Learning videography for click free viewing,” inIEEE Conf. Comput. Vis. Pattern Recog., 2017, pp. 1368–1376

  176. [185]

    Multicamba: a system for selecting camera views in live broadcasting of sport events using a dynamic 3d model,

    R. Yus, E. Mena, S. Ilarri, A. Illarramendi, and J. Bernad, “Multicamba: a system for selecting camera views in live broadcasting of sport events using a dynamic 3d model,”Multimedia Tools and Applications, vol. 74, pp. 4059–4090, 2015

  177. [186]

    Learning sports camera selection from internet videos,

    J. Chen, K. Lu, S. Tian, and J. Little, “Learning sports camera selection from internet videos,” inIEEE Winter Conf. Appl. Comput. Vis.IEEE, 2019, pp. 1682–1691

  178. [187]

    Which viewpoint shows it best? language for weakly supervising view selection in multi-view videos,

    S. Majumder, T. Nagarajan, Z. Al-Halah, R. Pradhan, and K. Grauman, “Which viewpoint shows it best? language for weakly supervising view selection in multi-view videos,” 2024,arXiv: 2411.08753

  179. [188]

    Switch- a-view: Few-shot view selection learned from edited videos,

    S. Majumder, T. Nagarajan, Z. Al-Halah, and K. Grauman, “Switch- a-view: Few-shot view selection learned from edited videos,” 2024, arXiv: 2412.18386

  180. [189]

    Lemma: A multi- view dataset for le arning m ulti-agent m ulti-task a ctivities,

    B. Jia, Y . Chen, S. Huang, Y . Zhu, and S.-c. Zhu, “Lemma: A multi- view dataset for le arning m ulti-agent m ulti-task a ctivities,” inEur. Conf. Comput. Vis.Springer, 2020, pp. 767–786

  181. [190]

    Ft-hid: a large-scale rgb-d dataset for first-and third-person human interaction analysis,

    Z. Guo, Y . Hou, P. Wang, Z. Gao, M. Xu, and W. Li, “Ft-hid: a large-scale rgb-d dataset for first-and third-person human interaction analysis,”Neural Computing and Applications, pp. 2007–2024, 2023

  182. [191]

    Wts: A pedestrian-centric traffic video dataset for fine-grained spatial-temporal understanding,

    Q. Konget al., “Wts: A pedestrian-centric traffic video dataset for fine-grained spatial-temporal understanding,” 2024,arXiv:2407.15350

  183. [192]

    Egoexolearn: A dataset for bridging asynchronous ego-and exo-centric view of procedural activities in real world,

    Y . Huanget al., “Egoexolearn: A dataset for bridging asynchronous ego-and exo-centric view of procedural activities in real world,” in IEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 22 072–22 086

  184. [193]

    Gazevqa: A video question answering dataset for multiview eye-gaze task-oriented collaborations,

    M. Ilaslanet al., “Gazevqa: A video question answering dataset for multiview eye-gaze task-oriented collaborations,” inConf. Empir. Methods Nat. Lang. Process., Proc., 2023, pp. 10 462–10 479

  185. [194]

    Predicting gaze in egocentric video by learning task-dependent attention transition,

    Y . Huang, M. Cai, Z. Li, and Y . Sato, “Predicting gaze in egocentric video by learning task-dependent attention transition,” inEur. Conf. Comput. Vis.Springer, 2018, pp. 789–804

  186. [195]

    Guide to the carnegie mellon university multimodal activity (cmu-mmac) database,

    F. D. la Torre Fradeet al., “Guide to the carnegie mellon university multimodal activity (cmu-mmac) database,” Carnegie Mellon Univer- sity, Pittsburgh, PA, Tech. Rep. CMU-RI-TR-08-22, April 2008

  187. [196]

    H2o: Two hands manipulating objects for first person interaction recognition,

    T. Kwon, B. Tekin, J. St ¨uhmer, F. Bogo, and M. Pollefeys, “H2o: Two hands manipulating objects for first person interaction recognition,”Int. Conf. Comput. Vis., pp. 10 118–10 128, 2021

  188. [197]

    Assembly101: A large-scale multi-view video dataset for understanding procedural activities,

    F. Seneret al., “Assembly101: A large-scale multi-view video dataset for understanding procedural activities,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 21 096–21 106

  189. [198]

    ARCTIC: A dataset for dexterous bimanual hand-object manipulation,

    Z. Fanet al., “ARCTIC: A dataset for dexterous bimanual hand-object manipulation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023

  190. [199]

    Oakink2: A dataset of bimanual hands-object manipu- lation in complex task completion,

    X. Zhanet al., “Oakink2: A dataset of bimanual hands-object manipu- lation in complex task completion,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 445–456

  191. [200]

    Home action genome: Cooperative compositional action understanding,

    N. Raiet al., “Home action genome: Cooperative compositional action understanding,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 11 179–11 188

  192. [201]

    Egoexo-fitness: Towards egocentric and exocentric full-body action understanding,

    Y .-M. Li, W.-J. Huang, A.-L. Wang, L.-A. Zeng, J.-K. Meng, and W.-S. Zheng, “Egoexo-fitness: Towards egocentric and exocentric full-body action understanding,” 2024,arXiv:2406.08877

  193. [202]

    Hollywood in homes: Crowdsourcing data collection for activity understanding,

    G. A. Sigurdsson, G. Varol, X. Wang, A. Farhadi, I. Laptev, and A. K. Gupta, “Hollywood in homes: Crowdsourcing data collection for activity understanding,” inEur. Conf. Comput. Vis., 2016

  194. [203]

    Core4d: A 4d human- object-human interaction dataset for collaborative object rearrange- ment,

    C. Zhang, Y . Liu, R. Xing, B. Tang, and L. Yi, “Core4d: A 4d human- object-human interaction dataset for collaborative object rearrange- ment,” 2024,arXiv:2406.19353

  195. [204]

    Egome: Follow me via egocentric view in real world,

    H. Qiu, Z. Shi, L. Wang, H. Xiong, X. Li, and H. Li, “Egome: Follow me via egocentric view in real world,” 2025,arXiv: 2501.19061

  196. [205]

    Estimating egocentric 3d human pose in the wild with external weak supervision,

    J. Wang, L. Liu, W. Xu, K. Sarkar, D. Luvizon, and C. Theobalt, “Estimating egocentric 3d human pose in the wild with external weak supervision,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 13 157–13 166

  197. [206]

    Assemblyhands: Towards egocentric activity understanding via 3d hand pose estimation,

    T. Ohkawa, K. He, F. Sener, T. Hodan, L. Tran, and C. Keskin, “Assemblyhands: Towards egocentric activity understanding via 3d hand pose estimation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 12 999–13 008. 20

  198. [207]

    Thermohands: A benchmark for 3d hand pose estimation from egocentric thermal image,

    F. Ding, Y . Zhu, X. Wen, and C. X. Lu, “Thermohands: A benchmark for 3d hand pose estimation from egocentric thermal image,” 2024, arXiv:2403.09871

  199. [208]

    Nymeria: A massive collection of multimodal egocentric daily motion in the wild,

    L. Maet al., “Nymeria: A massive collection of multimodal egocentric daily motion in the wild,” inEur. Conf. Comput. Vis., 2024

  200. [209]

    Ovr: A dataset for open vocabulary temporal repetition counting in videos,

    D. Dwibedi, Y . Aytar, J. Tompson, and A. Zisserman, “Ovr: A dataset for open vocabulary temporal repetition counting in videos,” 2024, arXiv:2407.17085

  201. [210]

    One-shot affordance detection,

    H. Luo, W. Zhai, J. Zhang, Y . Cao, and D. Tao, “One-shot affordance detection,” 2021,arXiv:2106.14747

  202. [211]

    360 +x: A panoptic multi-modal scene understanding dataset,

    H. Chen, Y . Hou, C. Qu, I. Testini, X. Hong, and J. Jiao, “360 +x: A panoptic multi-modal scene understanding dataset,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 19 373–19 382

  203. [212]

    Visual instruction tuning,

    H. Liu, C. Li, Q. Wu, and Y . J. Lee, “Visual instruction tuning,”Adv. Neural Inform. Process. Syst., vol. 36, 2024

  204. [213]

    Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,

    Z. Chenet al., “Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks,” inIEEE Conf. Comput. Vis. Pattern Recog., 2024, pp. 24 185–24 198

  205. [214]

    Eagle 2: Building post-training data strategies from scratch for frontier vision-language models,

    Z. Liet al., “Eagle 2: Building post-training data strategies from scratch for frontier vision-language models,” 2025,arXiv: 2501.14818

  206. [215]

    Videollm: Modeling video sequence with large language models,

    G. Chenet al., “Videollm: Modeling video sequence with large language models,” 2023,arXiv: 2305.13292

  207. [216]

    Video-rag: Visually-aligned retrieval-augmented long video comprehension,

    Y . Luoet al., “Video-rag: Visually-aligned retrieval-augmented long video comprehension,” 2024,arXiv:2411.13093

  208. [2013]

    Available: https://www.expressnews.com/sports/spurs/ article/SportVU-stats-can-be-helpful-overwhelming-4993731.php

    [Online]. Available: https://www.expressnews.com/sports/spurs/ article/SportVU-stats-can-be-helpful-overwhelming-4993731.php

  209. [2024]

    Available: https://randomwalk.ai/blog/how-visual-ai- transforms-assembly-line-operations-in-factories/

    [Online]. Available: https://randomwalk.ai/blog/how-visual-ai- transforms-assembly-line-operations-in-factories/

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.