Pith. sign in

REVIEW 14 cited by

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2507.18342 v1 pith:K6L7Y2LH submitted 2025-07-24 cs.CV

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

classification cs.CV
keywords egoexobenchmllmsreasoningacrossbenchmarkcross-viewintelligencemodels
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Transferring and integrating knowledge across first-person (egocentric) and third-person (exocentric) viewpoints is intrinsic to human intelligence, enabling humans to learn from others and convey insights from their own experiences. Despite rapid progress in multimodal large language models (MLLMs), their ability to perform such cross-view reasoning remains unexplored. To address this, we introduce EgoExoBench, the first benchmark for egocentric-exocentric video understanding and reasoning. Built from publicly available datasets, EgoExoBench comprises over 7,300 question-answer pairs spanning eleven sub-tasks organized into three core challenges: semantic alignment, viewpoint association, and temporal reasoning. We evaluate 13 state-of-the-art MLLMs and find that while these models excel on single-view tasks, they struggle to align semantics across perspectives, accurately associate views, and infer temporal dynamics in the ego-exo context. We hope EgoExoBench can serve as a valuable resource for research on embodied agents and intelligent assistants seeking human-like cross-view intelligence.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering

    cs.CV 2026-05 unverdicted novelty 7.0

    CaST-Bench creates a benchmark with causal-chain annotations and novel metrics showing that current VLMs struggle to construct precise grounded causal chains in video QA.

  2. CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering

    cs.CV 2026-05 unverdicted novelty 7.0

    Introduces CaST-Bench, a dataset of 2,066 causal questions on 1,015 videos with annotated causal chains and metrics to evaluate VLMs on spatio-temporal causal reasoning.

  3. VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis

    cs.CV 2026-05 unverdicted novelty 7.0

    VGenST-Bench is a new video benchmark for MLLM spatio-temporal reasoning built via generative synthesis, a multi-agent pipeline with human oversight, a 3x2x2 taxonomy, and hierarchical tasks separating perception from...

  4. Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models

    cs.CV 2026-05 conditional novelty 7.0

    SP-CoR is a multimodal LLM framework using dynamics-aware sampling, spectral-physics view fusion, and prompt distillation that outperforms baselines on the new CoopSR benchmark and EgoTeam dataset for multi-robot coop...

  5. SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting

    cs.CV 2025-11 unverdicted novelty 7.0

    SFHand presents the first streaming language-guided autoregressive framework for 3D hand forecasting, achieving up to 35.8% gains over prior methods and 13.4% better downstream embodied task performance.

  6. Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos

    cs.CV 2026-07 conditional novelty 6.5

    EgoMemo uses multi-scale temporal summaries, a knowledge graph, and visual archives to decide whether and when to intervene proactively on continuous egocentric video, setting baselines on the new EgoServe benchmark o...

  7. DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation

    cs.CV 2026-04 unverdicted novelty 6.0

    A scalable training-free pipeline using video segmentation, filtering, and off-the-shelf multimodal models creates DenseStep2M, a dataset of 100K videos and 2M detailed instructional steps that improves dense captioni...

  8. From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation

    cs.CV 2026-04 unverdicted novelty 6.0

    Interpolating exo and ego videos into a single continuous sequence lets diffusion sequence models generate more coherent first-person videos than direct conditioning, even without pose interpolation.

  9. From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation

    cs.CV 2026-04 conditional novelty 6.0

    Interpolating only the video frames between synchronized exo and ego clips already turns discontinuous cross-view generation into continuous sequence modeling and measurably improves diffusion-based Exo2Ego synthesis.

  10. Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed Spaces

    cs.CV 2026-02 conditional novelty 6.0

    A new human-curated benchmark of 1,000 ranking questions on real-world engineering structures shows the best vision-language model reaches 33.6% accuracy where humans reach 91.6%.

  11. The N-Body Problem: Parallel Execution from Single-Person Egocentric Video

    cs.CV 2025-12 conditional novelty 6.0

    Structured prompting for a VLM, fed with ground-truth spatial zone schedules, can generate more feasible multi-agent parallel executions from single-person egocentric videos.

  12. EgoExo-Con: Exploring View-Invariant Video Temporal Understanding

    cs.CV 2025-10 conditional novelty 6.0

    Most Video-LLMs answer temporal questions far less consistently when the same event is shown from ego and exo views, and a GRPO variant with a reasoning-similarity reward partially closes the gap.

  13. LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video

    cs.CV 2026-06 unverdicted novelty 5.0

    Presents LongSpace-Bench benchmark and LongSpace framework that chunks long videos, adds 3D structural cues, and builds layer-aware memory to improve spatial reasoning in multimodal LLMs.

  14. EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs

    cs.CV 2026-05 unverdicted novelty 5.0

    EgoCoT-Bench provides 3,172 verifiable QA pairs across perception, anticipation, and reasoning tasks on egocentric videos, revealing that many MLLMs give answer-correct but evidence-inconsistent explanations.