{"id":"e6a715f4-2206-4e4e-b3b4-d6bfcce8a055","arxiv_id":"2506.06253","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A comprehensive review of cross-view video understanding that uses both first-person and third-person cameras, organized into a three-direction taxonomy with a dataset catalog and future research gaps.","lead":"This survey collects and organizes research on video understanding that combines first-person (egocentric) and third-person (exocentric) views, grouping the field into three directions: egocentric-for-exocentric, exocentric-for-egocentric, and joint learning. It also catalogs benchmark datasets and lists under-explored tasks such as cross-view action assessment and affordance grounding for specialized domains.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's 'first comprehensive' claim rests on an undocumented sampling of the literature; without a recall check against a systematic ego-exo query, its taxonomy and gap analysis may be incomplete.","rationale":"The reader's weakest assumption is exactly that the paper's sample of papers and datasets is sufficiently complete to support the 'comprehensive' framing. That is the most load-bearing assumption: the paper's novel contribution is its three-way taxonomy and its gap analysis, both of which depend on seeing the field. The paper provides no search protocol, inclusion criteria, or comparison with competing surveys, so the claim is currently unverifiable from the text. I do not treat this as a demonstrated error: the taxonomy is internally coherent, the dataset table is informative, and I found no specific omission that clearly falsifies the comprehensiveness claim. The appropriate posture is therefore UNCHANGED, with a concrete recall check that would settle whether the concern actually lands. The reader's verdict already notes this issue, so agreement is 'agree.' No ad hominem is involved; the critique concerns the survey's evidentiary basis, not the authors' conduct.","tokens_in":31755,"tokens_out":4922,"duration_ms":54913,"concrete_test":"Run a systematic search on DBLP/OpenAlex/arXiv for 2015–2025 using a query such as (egocentric AND exocentric) AND (video OR action OR retrieval OR captioning OR localization OR generation); screen results with the survey's own definition of ego-exo video understanding; then compute recall of the survey's ego-exo reference list and count screened papers absent from the manuscript. Also inspect Plizzari et al. [32] and any other ego-vision surveys for an ego-exo integration section. If a distinct task cluster is absent or recall is below 80%, the 'first comprehensive' claim and the under-explored conclusions need qualification; if recall is high, the claim stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section I states 'no survey has yet addressed the integration of both perspectives,' and Sections IV and VI draw conclusions about which tasks are 'under-explored.' These conclusions are valid only if the referenced work is a representative sample of ego-exo video research. The paper gives no search protocol, inclusion criteria, or comparison with existing surveys; the companion GitHub repo cannot be validated from the manuscript. Selection appears curated rather than systematic: IV.C View Selection includes [183]–[188] even though the text later says those methods treat ego and exo views separately, and IV.B Remote Drone Teleoperation covers HCI teleoperation rather than video understanding. A curated sample can still yield a useful taxonomy, but 'comprehensive' and the under-explored gap analysis become unsubstantiated if a substantial body of ego-exo work is omitted. This is not an internal inconsistency; it is an unverified external completeness claim. The concern lands only if a systematic recall check shows materially low coverage.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a survey of video understanding research that integrates egocentric (first-person) and exocentric (third-person) perspectives. It opens with eight application domains, derives a set of research tasks from them, organizes the literature into three directions (ego-for-exo, exo-for-ego, and joint learning), and provides a table of ego-exo datasets. The paper claims to be the first comprehensive survey of this integration and concludes with a gap analysis and future-work suggestions. A companion GitHub repository is announced.","tokens_in":31891,"tokens_out":7835,"duration_ms":73112,"significance":"If its coverage is representative, the survey is timely and useful: it gives the area a coherent taxonomy, collects many relevant methods across disparate subfields, and the dataset table in Section V is a practical contribution. The paper also ships a companion GitHub repository, which helps researchers track new work. However, the central 'comprehensive' claim and the under-explored task analysis rest on an undocumented literature sample, and a few taxonomically misplaced sections weaken the review's internal consistency. The contribution is therefore not yet fully supported, although the issues appear fixable within the manuscript's scope.","major_comments":[{"comment":"The paper's central claim that it is the first comprehensive survey of egocentric-exocentric video understanding (Section I) and the gap analysis in Section VI are not backed by a documented literature search. There is no search protocol, inclusion criterion, or comparison with prior or concurrent surveys, so the reader cannot judge whether the works sampled in Sections IV and V are representative. This matters because Section VI draws conclusions about which tasks are 'under-explored' and Fig. 4 is used to state that 'many tasks critical to applications remain under-investigated.' I request an explicit methodology paragraph (databases, query terms, inclusion/exclusion criteria, and screening flow) and, ideally, a recall-style check against a systematic ego-exo query to substantiate the comprehensiveness claim. Without this, the claim remains an unverified external completeness assertion.","section":"Section I, Section VI, Fig. 4"},{"comment":"The task 'Remote Drone Teleoperation' is placed under 'Exocentric for Egocentric' video understanding, but the works reviewed there ([132]–[135] and [133]) are human-computer interaction systems that use VR, additional cameras, or AR overlays to provide a pilot with a better exocentric view during teleoperation; they do not use exocentric data to improve egocentric video understanding, which is the definition given at the start of Section IV.B. This inclusion stretches the taxonomy and weakens the review's stated focus on video understanding. Either move these works to an application-oriented subsection or add an explicit justification of how they inform ego-exo video understanding.","section":"Section IV.B, Fig. 5"},{"comment":"The 'from applications to research tasks' bridge is incomplete. Section IV reviews view birdification, cross-view retrieval, 3D camera localization, egocentric wearer identification, and cross-view human identification, but none of these tasks appears in Fig. 4, which is used in Section III to support the statement that many application-critical tasks are under-investigated. The figure should include all tasks discussed in Section IV, or the text should clearly state that Fig. 4 is a partial mapping; alternatively, the gap analysis should be supported by the full task list in Section IV.","section":"Section III, Fig. 4"}],"minor_comments":[{"comment":"The caption says Section IV is divided into 'ego for exo, exo for exo, and joint learning'; 'exo for exo' should be 'exo for ego'.","section":"Figure 2 caption"},{"comment":"There are typos in the figure: 'Camara Localization' should be 'Camera Localization' and 'Healthcar e' should be 'Healthcare'.","section":"Figure 4"},{"comment":"References [75] and [177] are the same paper (Seo et al., 'Multi-view masked world models for visual robotic manipulation'), and references [116] and [170] are the same paper (Jia et al., 'The audio-visual conversational graph'); these duplicates should be unified.","section":"References"},{"comment":"EgoFish3D [114] is described as using an exocentric pose estimator as supervision; this is more accurately characterized as weakly supervised than self-supervised, given the terminology used in the surrounding text.","section":"Section IV.B, Self-supervised methods"},{"comment":"The dataset entry 'ThirdtoFirst [51]' cites a method paper (Li et al., 'Ego-Exo: Transferring visual representations from third-person to first-person videos') rather than a dataset paper; please align the citation with the actual dataset source or rename the entry.","section":"Section V.D"},{"comment":"The caption says the citation counts cover 'papers and datasets discussed in Sections IV and V,' but the collection and curation rules for the Google Scholar data are not described; please clarify the inclusion criteria so readers can interpret the growth curve.","section":"Fig. 1"}],"recommendation":"major_revision","confidential_remarks":"The survey includes a notable number of works from the authors' own groups (e.g., [8], [54], [60], [83], [145], [192], and others), which, in the absence of a search protocol, creates a risk of coverage bias. This should be addressed in revision through a transparent methodology and, if possible, a comparison with other recent ego-exo efforts. The paper is best characterized at present as a curated topical review; with the requested methodological additions, it could become a properly comprehensive survey."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Josh,\n\nThis survey does something useful: it organizes the growing ego-exo video understanding literature into three directions (ego-for-exo, exo-for-ego, joint learning) and provides a genuinely handy dataset table. If you're new to the area, this is a good starting point to see the task landscape and which datasets exist. The figures are mostly adapted, but they're clear.\n\nThe soft spot is the 'first comprehensive survey' claim. The paper gives no search protocol, no inclusion criteria, no comparison with other surveys. The selection looks curated rather than systematic. For example, the View Selection subsection includes 360° video work that is not really ego-exo, and the drone teleoperation subsection is largely HCI rather than video understanding. That does not break the survey's value, but it does mean the under-explored task gap analysis is only as good as the sampling. A short methodology section or a comment on how the references were collected would fix this.\n\nSelf-citation is noticeable but not damning. The authors have worked on Ego-Exo4D, EgoExoLearn, and related topics, and they cite their own work heavily, but they also cover plenty of other groups. The main risk is that some areas may be underrepresented because they are less familiar with them.\n\nMinor issues: Figure 2 mislabels 'Exo for Exo' (should be 'Exo for Ego'), and a few dataset table entries are incomplete. These are easy fixes.\n\nOverall, this is a solid survey. It does not need to be the first comprehensive one to be useful. I'd accept it with minor revisions. A reader getting into the field will get a good map; a specialist will find the dataset table handy and the gap analysis a starting point.\n\nSend it to peer review. It is not a groundbreaking contribution, but it earns its place.","headline":"Useful organizing survey of ego-exo video understanding, solid dataset table, but 'comprehensive' claim needs either a search methodology or a softer claim.","tokens_in":32445,"tokens_out":2609,"would_cite":true,"duration_ms":22984,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A first survey unifies first- and third-person video AI.","keywords":["video understanding","egocentric video","exocentric video","cross-view learning","first-person vision","survey","datasets and benchmarks"],"falsifier":"An independent, systematic literature search with explicit inclusion criteria that surfaces a substantial body of ego-exo work omitted from this survey, or that reveals a major research direction (for instance multi-agent ego-exo collaboration) not covered by the three-way taxonomy, would falsify the survey's central claims of comprehensiveness and completeness.","tokens_in":31542,"feed_emoji":"🎥","tokens_out":4950,"duration_ms":46296,"temperature":0.7,"pith_summary":"This survey argues that perceiving the world from both a first-person (egocentric) and a third-person (exocentric) perspective is a distinct and rapidly growing research direction in video understanding, and claims to be the first comprehensive review of that direction. It proposes that the field is naturally organized into three lines of work: using egocentric data to improve exocentric understanding, using exocentric data to improve egocentric analysis, and jointly learning from both views. The survey connects these directions to eight application areas, from cooking assistants to traffic and surgery, and to the datasets that support them. A sympathetic reader would care because the paper supplies a taxonomy and a gap analysis that can steer future work toward the tasks and data that are still missing.","feed_headline":"A first survey unifies first- and third-person video AI","feed_subtitle":"It maps the field into three directions and pinpoints under-explored tasks and missing datasets.","key_machinery":"The organizing device is a three-way taxonomy of research directions—ego-for-exo (egocentric cues improve exocentric tasks), exo-for-ego (exocentric data improves egocentric analysis), and joint learning (both views are used at training and inference together). The survey maps each direction to concrete tasks such as video generation, action understanding, affordance grounding, and cross-view retrieval, and links those tasks back to eight application scenarios. This taxonomy does the work of a framework: it positions every reviewed method, exposes where the literature is thin, and drives the paper's dataset inventory and future-directions discussion.","core_discovery":"The paper's central claim is that cross-view collaboration between egocentric and exocentric video is a coherent field whose progress can be mapped by three research directions: egocentric-for-exocentric, exocentric-for-egocentric, and joint learning. On the paper's own terms, no prior survey has integrated both perspectives, so this review fills that gap by organizing existing methods, identifying eight application domains that would benefit from ego-exo collaboration, and cataloguing datasets that contain both viewpoints. The review concludes that current work clusters around daily-life activity understanding, while tasks such as ego-to-exo video generation, view birdification, skill assessment, and affordance grounding for industrial or surgical tools remain under-explored.","pith_inferences":["The 'first comprehensive survey' claim is only as strong as the literature search behind it; a reader can probe it by checking whether recent ego-exo works beyond the cited set fit the taxonomy.","The three-way taxonomy may under-represent multi-agent ego-exo collaboration, which appears in applications like rescue and public service but is not given its own research direction.","A concrete testable extension of the gap analysis is to benchmark ego-to-exo video generation on the driving and robotics scenarios the paper names, where latency constraints are severe.","The dataset table offers a way to quantify domain imbalance, for instance by counting ego-exo datasets per application area, which could prioritize future data collection."],"forward_implications":["A researcher can place any ego-exo method into one of three categories, making the field's structure explicit for newcomers.","The gap analysis identifies ego-to-exo video generation, view birdification, and cross-view skill assessment as under-studied tasks worth pursuing.","Specialized domains such as healthcare, education, and public service lack dedicated ego-exo datasets, so collecting them is a clear next step.","Because synchronized paired data is expensive, aligning unpaired ego and exo videos via language or retrieval is a promising route to scale.","Vision-language models that can take both perspectives as input are proposed as a path toward unified cross-view frameworks."],"supporting_citations":[{"why":"Supplies the mirror-neuron framing that motivates studying first- and third-person vision together.","marker":"[1]"},{"why":"Early joint model of first- and third-person video, a foundational cross-view method the survey builds its taxonomy around.","marker":"[2]"},{"why":"Ego-Exo4D, the large-scale benchmark that anchors the field's recent growth and many reviewed works.","marker":"[8]"},{"why":"The egocentric-vision survey whose 'future-to-present' structure this survey explicitly adopts while extending it to both perspectives.","marker":"[32]"},{"why":"Key exo-for-ego method transferring third-person representations to first-person video, a core thread in the review.","marker":"[51]"},{"why":"AE2's unpaired temporal alignment approach, central to the joint-learning section's discussion of data-efficient methods.","marker":"[64]"},{"why":"Exo2Ego, a representative diffusion-based method for exocentric-to-egocentric video generation reviewed in the exo-for-ego direction.","marker":"[85]"},{"why":"Foundational method for identifying the camera wearer in third-person video, a task the survey treats as a distinct research direction.","marker":"[7]"}],"fun_headline_variants":["Cross-view video survey maps ego-exo collaboration","First survey unifies ego-exo video into three directions","Ego-exo video: survey highlights untapped tasks","Bridging first-person and third-person video understanding"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's value rests on its selection of papers and datasets being complete and representative enough that the 'first comprehensive review' claim and the resulting gap analysis are trustworthy.","fun_headline_variants_meta":{"raw":{"variants":["Cross-view video survey maps ego-exo collaboration","First survey unifies ego-exo video into three directions","Ego-exo video: survey highlights untapped tasks","Bridging first-person and third-person video understanding"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00022,"raw_usage":{"total_tokens":1447,"prompt_tokens":945,"completion_tokens":502,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":561,"completion_tokens_details":{"reasoning_tokens":438}},"tokens_in":561,"tokens_out":502,"duration_ms":5010,"temperature":1.0,"reasoning_tokens":438,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:56:59.203020+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent, systematic literature search with explicit inclusion criteria that surfaces a substantial body of ego-exo work omitted from this survey, or that reveals a major research direction (for instance multi-agent ego-exo collaboration) not covered by the three-way taxonomy, would falsify the survey's central claims of comprehensiveness and completeness.","supporting_citations":[],"review_version":1}