Pith. sign in

Paper Citation Record · LEDGER

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

As of 18 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 15 inbound Pith citation observations for arXiv:2507.18342.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18342 v1

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:20:58.963152Z

measured 110 of 110 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:32:27.422997Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:49:17.558383Z

Reference resolution

95 of 95 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f2310575-32fd-4f90-bef6-c88311fa63f3 · outbound

This paper cites Visual-policy learning through multi-camera view to single-camera view knowledge distillation for robot manipulation tasks.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Visual-policy learning through multi-camera view to single-camera view knowledge distillation for robot manipulation tasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.617371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.617371Z digest=sha256:9a0cbbb26dcf87f44c12f3c07df4f55e4ceab8894f1b9093caa8ee85f55d0dd3

Observation 32b58db8-8323-46c7-8617-6ad42750aa93 · outbound

This paper cites Claude 3.7 sonnet and claude code, 2025.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Claude 3.7 sonnet and claude code, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.621745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.621745Z digest=sha256:e4dd378d6ea3c622b316ae8b070983797a98d4da67855b94b7bf4d9eec04fb86

Observation 67f8bc59-1544-40c0-82e4-0bc7bdb8fd29 · outbound

This paper cites Observational learning.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Observational learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.625637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.625637Z digest=sha256:ebd17d12d4fea1ca6baa61c469f31249b3cf2da5efdcfd80ef4787de5e58f35c

Observation a3a41bf0-8eaa-4dc2-baec-da47b7e8841a · outbound

This paper cites an unresolved cited work.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.629556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.629556Z digest=sha256:a0801962f6238f7d464f23a2ae67c8a5706bd8f65047ed053048f9116a4d0fb5

Observation 195e7901-9056-485b-bf87-80485589499a · outbound

This paper cites In your place: neuropsychological evidence for altercentric remapping in embodied perspective taking.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs In your place: neuropsychological evidence for altercentric remapping in embodied perspective taking

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.633385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.633385Z digest=sha256:b9130c4d7f3acf7c7aee23e74e4c60ad1e63a2452e88af6e892f641f1a6ddfe6

Observation a56965a1-0cd4-41bb-b21e-754d1fe5c59f · outbound

This paper cites Spatial memory: how egocentric and allocentric combine.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Spatial memory: how egocentric and allocentric combine

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.637270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.637270Z digest=sha256:f000eea3a9aef1354ee55c052d9d0b04d23cf153fe568459edb8e7dd6b66cb45

Observation 500c462b-1311-4611-b691-1b571ffc462d · outbound

This paper cites Hourvideo: 1-hour video-language understanding.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Hourvideo: 1-hour video-language understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.641150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.641150Z digest=sha256:f18fb4ffd9069ee69c507d02cd00db7973668e79292efaa260c84a6422be2ce5

Observation 40bfb0cd-576b-4e7e-be9a-a50316f37e40 · outbound

This paper cites CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.644807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.644807Z digest=sha256:f234229ff9242654a709e6e42916c774f4c7a1074bd08177b5d0a8923880b66c

Observation 7c8cf13b-8852-4646-bd10-07e2c70b9d5b · outbound

This paper cites put myself into your place.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs put myself into your place

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.648836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.648836Z digest=sha256:4d15c9a60e5c3146890a6673f65d7e6192e8d6bc7a588deac3e91410cbdd9f4b

Observation 1ffa20f7-275f-40f7-9773-25e0efae1da7 · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Scaling egocentric vision: The epic-kitchens dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.652640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.652640Z digest=sha256:0b885306b87db0f35938d235931287989f9ac114bafaddffcd7879a7b4a8ec4e

Observation 343622d4-2d36-4e23-bbfa-a25025e05a77 · outbound

This paper cites Epic-kitchens visor benchmark: Video segmentations and object relations.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Epic-kitchens visor benchmark: Video segmentations and object relations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.656210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.656210Z digest=sha256:827c2d93777f5b2620596bd9211b3acb88708409ad415d653168c2fd417f9f93

Observation 1e2b9dac-14ca-4ba8-88f9-950b36361009 · outbound

This paper cites Gemini 2.5 pro, 2025.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Gemini 2.5 pro, 2025

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.659999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.659999Z digest=sha256:4c99660f08576a0e646e8568b0f4c0608dde90535421c6ab911338209dd71bc6

Observation fd5530cb-e27c-4997-a7b7-8aa2d3169a73 · outbound

This paper cites Interacting networks of brain regions underlie human spatial navigation: a review and novel synthesis of the literature.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Interacting networks of brain regions underlie human spatial navigation: a review and novel synthesis of the literature

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.663511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.663511Z digest=sha256:7fc6cfbe55f646c15fa92fa2891bec3dfe22a23fda255e497ca616e8ad5d7b66

Observation 798a96e5-3f76-4954-8086-ca2d90d20441 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.667394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.667394Z digest=sha256:6c6da8c9ea42698155088cdbab039a437f636fdbbabd4b4be2f1280bea021a9b

Observation 69c1d61f-42a0-4578-9dea-3a8343feaf39 · outbound

This paper cites Complexity-Based Prompting for Multi-Step Reasoning.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Complexity-Based Prompting for Multi-Step Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.671208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.671208Z digest=sha256:cd816d283e86df30ce6644b6f031337fff2fbc59d0b80891bf96bf858d3eff4a

Observation 3f37f27f-952d-4e41-a810-a0d090650b14 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Ego4d: Around the world in 3,000 hours of egocentric video

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.674985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.674985Z digest=sha256:1008307dff103b6cdf101529c63f14a05d5bb1535eaad3bafb2379a5dee5cd43

Observation 34a89421-6e29-4280-aacb-d753faa57df3 · outbound

This paper cites Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.678833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.678833Z digest=sha256:87260b0549720178c90b5b734c7ce28635744b69cff96db66e38a8c1ac3c9e0a

Observation 8ea24c61-3eb8-4d72-a07a-cae17bb42941 · outbound

This paper cites Ava: A video dataset of spatio-temporally localized atomic visual actions.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Ava: A video dataset of spatio-temporally localized atomic visual actions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.682780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.682780Z digest=sha256:b263df75b317e01513412097c3470d00a0ee81fdcd918ca4d4d77092375f28cd

Observation ec7a0169-2b96-4690-ab8c-ea79d5195bde · outbound

This paper cites Multiple human association and tracking from egocentric and complementary top views.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Multiple human association and tracking from egocentric and complementary top views

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.686233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.686233Z digest=sha256:e851586c865d1722f26ac2980fe613780987582be0c62b5c6777ca238aca0b4e

Observation 4ce2ff04-61e6-4fb5-b252-e953c23626aa · outbound

This paper cites What is modelled during observational learning? Journal of sports sciences, 25(5):531–545, 2007.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs What is modelled during observational learning? Journal of sports sciences, 25(5):531–545, 2007

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.691028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.691028Z digest=sha256:4eea4b1cd0830e60c2639b12bba2be41b5e67227858d7588d9d4ba6810b007fb

Observation 6c900cf2-3c50-4a7f-a6a1-e3dbd1d9f79d · outbound

This paper cites An ego-vision system for discovering human joint attention.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs An ego-vision system for discovering human joint attention

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:21:00.165437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.694818Z digest=sha256:fdb721f11a7cc9cd7f478f9fe8a827f4dd1aaefda8ee1850424a58c09885be84

Observation 0bf67345-564d-4955-b7fb-7278172bdb99 · outbound

This paper cites Improving action segmentation via graph-based temporal reasoning.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Improving action segmentation via graph-based temporal reasoning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:21:00.151308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.698351Z digest=sha256:b0532a228b408a60111b74df7a39629cb873356d202b05d4dcde244da72e0719

Observation 5a595925-4655-4dc4-9c46-8cf167c170e7 · outbound

This paper cites Egoexolearn: A dataset for bridging asynchronous ego-and exo-centric view of procedural activities in real world.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Egoexolearn: A dataset for bridging asynchronous ego-and exo-centric view of procedural activities in real world

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:21:00.132433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.702040Z digest=sha256:21cafea110eb540dd7693d93ec5d462afb483aeac0f41cb2c0678ea191d9f563

Observation b5ad38a0-fd46-4150-b409-72d8b5758dce · outbound

This paper cites Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.705374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.705374Z digest=sha256:b4e58893a87f19b09eac934f598363ee82dc8763f4653e380c95e9d74555cfa2

Observation 55a6e0b9-07ba-42fc-8449-267a40846e03 · outbound

This paper cites GPT-4o System Card.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs GPT-4o System Card

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.710025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.710025Z digest=sha256:a3bd27bcbe2f01c274e3d5cc15bb963df7797add636f757a84e381060b459c7e

Observation ac2851b2-345a-4d95-b3d8-a5dc82ee5977 · outbound

This paper cites Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:21:00.114741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.714181Z digest=sha256:d1292d4c4a3db61f1e2a99068e1453828731dc8c2e763052a1683ad47f6cefe4

Observation f1e05fce-ca9f-4c0a-a436-620e78c46e79 · outbound

This paper cites Egotaskqa: Understanding human tasks in egocentric videos.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Egotaskqa: Understanding human tasks in egocentric videos

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:21:00.100849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.717632Z digest=sha256:e3d71517d0c2be412f14162bed07e765a7b4bd7bdc2465df208c5fb8d2e5fc2f

Observation fa0dee72-f454-4fa2-9618-c7dced4d2c91 · outbound

This paper cites Large language models are zero-shot reasoners.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Large language models are zero-shot reasoners

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.722270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.722270Z digest=sha256:a176d4e162e938ff8c7ddf678f818ac57fca762d9608aa85a3022e6d2ffded9f

Observation 322f785f-d3a1-4aa0-b614-7d8e03ee7c5b · outbound

This paper cites A neural code for egocentric spatial maps in the human medial temporal lobe.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs A neural code for egocentric spatial maps in the human medial temporal lobe

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:21:00.070602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.725834Z digest=sha256:fbb133ab246dd2455cf55bc8948bc5deb1dd3d5a514e1ddc76c92efe06e6ec0e

Observation a6d62541-ba34-49f7-8966-6f94c836dac0 · outbound

This paper cites H2o: Two hands manipulating objects for first person interaction recognition.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs H2o: Two hands manipulating objects for first person interaction recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:21:00.047449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.729432Z digest=sha256:e1531e5a3be11ac09447abb661c09511502a239d6e224943e85ff00428addd8f

Observation 3e8c75a4-a3f7-4ae7-8709-f41783d74906 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.732844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.732844Z digest=sha256:865102e14fb4d2868ed6cf5327eb9ec2348f404337a34c1d42e59b05bf5ef823

Observation 5cd4631f-db5c-43c6-b8be-e66c8270d8b0 · outbound

This paper cites iGibson 2.0: Object-Centric Simulation for Robot Learning of Everyday Household Tasks.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs iGibson 2.0: Object-Centric Simulation for Robot Learning of Everyday Household Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.736826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.736826Z digest=sha256:4110525aeb8d41c0470bef1edd4112665c65b38693a438a1162b74f59ac5c895

Observation 1cf0da7c-2ee8-49dc-960f-bb2c2f3371d5 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs VideoChat: Chat-Centric Video Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.740653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.740653Z digest=sha256:54130d58125fd744b8e1cdd0d67438b4247073d9f3d324e4a11ddb0c7dbae748

Observation abe12be3-7929-4c38-aff4-45b63c0ef68d · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.744663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.744663Z digest=sha256:f8aec91d00d6a6af43f566377586db178f58d6bca6122916a61ff6fe0edcc894

Observation b84f0f9b-786a-4894-b00b-49e17b95a2e0 · outbound

This paper cites Ego-exo: Transferring visual representa- tions from third-person to first-person videos.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Ego-exo: Transferring visual representa- tions from third-person to first-person videos

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:21:00.002322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.747904Z digest=sha256:a649d9711627733d4a112dd4e7700a6615b9b2f4703f6696a49ee46ddda68cc6

Observation 40fbdaf5-3353-4bcf-a1f1-e526057788ce · outbound

This paper cites DeepSeek-V3 Technical Report.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs DeepSeek-V3 Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.751334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.751334Z digest=sha256:b8ed9473e3dc819233f3c67d4ffa5c4d40fd920bbdbcad9acc4f8f99c3abd2a0

Observation b762a704-7475-4755-a799-80d8f08b9f74 · outbound

This paper cites Skill transfer learning for autonomous robots and human–robot cooperation: A survey.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Skill transfer learning for autonomous robots and human–robot cooperation: A survey

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.989510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.755144Z digest=sha256:ec3a709f4444a5474d2d00cbab443697c6db921fe07e22196046b793fb6e7628

Observation 24ef29ac-de58-4737-831d-e7a86a718373 · outbound

This paper cites Hoi4d: A 4d egocentric dataset for category-level human-object interaction.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Hoi4d: A 4d egocentric dataset for category-level human-object interaction

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.758843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.758843Z digest=sha256:e7fb0a3210d1af3571fb0c844b2fdb726e61dd7124b9ad9d2bd6598d09b93656

Observation b2c29678-f4a0-4ad3-b74e-81883267bf42 · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs NVILA: Efficient Frontier Visual Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.762270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.762270Z digest=sha256:f3069fc60fdad747f90ba2cc39695a1372d47687e1a1c93001e1f8fb258c5495

Observation e7d8ac96-6b1e-4ad1-a8f5-806e8f3421ca · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.970058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.766048Z digest=sha256:6cfd419a8d53d6ac6376afde2b64a5abee87c2e8669dac08017ec7d76567cd91

Observation a5af1db1-350c-4fce-a7fb-e54fb3c7546b · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.769555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.769555Z digest=sha256:4b0cf88c1bdbece9ea030ae1c23508ca6868e192ca878a986d386e69ae611c81

Observation 45887d3f-a001-41de-9ec1-2a5e5fea8579 · outbound

This paper cites Gpt-4o mini: advancing cost-efficient intelligence, 07 2024.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Gpt-4o mini: advancing cost-efficient intelligence, 07 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.952062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.773316Z digest=sha256:43d5304eda484c0ddb35e4c512e920e64df849b019a93ccf6573d9090a569f5f

Observation d4856754-4a7b-4506-bd2e-896f5c28f8d7 · outbound

This paper cites Egovlpv2: Egocentric video-language pre-training with fusion in the backbone.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.776896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.776896Z digest=sha256:9034683807dff32ecb8518db3fe9980d5e16548fd03f6e6e33552537d685974a

Observation 7a573c7e-9f12-4203-97d3-0fdadb4bdb56 · outbound

This paper cites EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.780213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.780213Z digest=sha256:9ff28464bb944c0cb90dcc0a2fbae04721c7ad5601f72d1a9887505cf307c722

Observation 922753bc-70f8-40d7-b25d-958dfa7bb363 · outbound

This paper cites Synchronization is All You Need: Exocentric-to-Egocentric Transfer for Temporal Action Segmentation with Unlabeled Synchronized Video Pairs.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Synchronization is All You Need: Exocentric-to-Egocentric Transfer for Temporal Action Segmentation with Unlabeled Synchronized Video Pairs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.784500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.784500Z digest=sha256:34c975959229b861503de3465dbbbe7d5da9fccb6111c00449c44d93f2785100

Observation 42fb4af6-ce29-4ddd-97e6-b9505cb88f72 · outbound

This paper cites The meccano dataset: Understanding human-object interactions from egocentric videos in an industrial-like domain.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs The meccano dataset: Understanding human-object interactions from egocentric videos in an industrial-like domain

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.933386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.788336Z digest=sha256:70c7789252ebd4e6506618ac321c794ca3f18bf2c1bdf76bbb6f3895fcaa292c

Observation 33f12641-5c0c-4e36-b1c9-49807e2031d9 · outbound

This paper cites Home action genome: Cooperative compositional action understanding.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Home action genome: Cooperative compositional action understanding

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.922384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.791786Z digest=sha256:ba3a4020ce302c959a4adf415a93f4d2c68e4d8929eeaa71bb7cd1a6914480a2

Observation 2d56ccb1-da42-45fb-8f52-bcac6016e0d7 · outbound

This paper cites Watch and learn: the cognitive neuroscience of learning from others’ actions.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Watch and learn: the cognitive neuroscience of learning from others’ actions

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.910231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.795433Z digest=sha256:dbb3e8e5ee7f3e31be01c0aa9fbc15c42ebb200902b6d8bf8d7b05cf7fc5ec6d

Observation f1b7abe3-f17f-4e4d-a059-ab720e673389 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.798955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.798955Z digest=sha256:b1c0c7947d390c57cc748a6535dabbccefb7a0c1b4b5b5dd5b2cce5722ba558c

Observation 1f4cd9b7-d2ee-4c50-bc57-68db6b5e666b · outbound

This paper cites A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.802920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.802920Z digest=sha256:fbd9632ebb07fed8b3d3e58fe4fc8b3ce20713ee3a8fecb364cd2814836316a7

Observation 3a7c4021-baec-4106-a456-4cbce44a88a1 · outbound

This paper cites Sener, D.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Sener, D

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.898336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.806519Z digest=sha256:3e6a092b21d3d60f6ae1347a449f4989c3dc82961608f24d208a2d894984bcbb

Observation acded3b6-d552-48e6-ae4b-96d140bfdd78 · outbound

This paper cites Self-supervised disentangled representation learning for third-person imitation learning.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Self-supervised disentangled representation learning for third-person imitation learning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.886385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.809903Z digest=sha256:f158dbe4b32c989ddd50bd819ea94d78b7227bad8d7993bd6d9c3bce937f757a

Observation 6c33268e-7a7f-4cba-86ea-05c5b4051135 · outbound

This paper cites Third-person visual imitation learning via decoupled hierarchical controller.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Third-person visual imitation learning via decoupled hierarchical controller

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.875278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.813364Z digest=sha256:bc079d565cad0af1086390107a015bf16209f492276bd52a5ca466c3c6c8b47a

Observation defd28b4-0e73-4694-a699-09591d2c2629 · outbound

This paper cites Actor and observer: Joint modeling of first and third-person videos.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Actor and observer: Joint modeling of first and third-person videos

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.863436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.816505Z digest=sha256:bd04eb64cd85b227ec9fa3efc130db628c9eacba8c14f88657d968d162d25bd5

Observation 08cbaa41-583a-40de-8add-546444c82e9a · outbound

This paper cites Ego4d goal-step: Toward hierarchical understanding of procedural activities.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Ego4d goal-step: Toward hierarchical understanding of procedural activities

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.851917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.819680Z digest=sha256:1832ce76b97808194033538f3464822f7de958f41da90f4ea569cde7fb650a94

Observation 3199f91d-1a2c-4add-8a29-6bf69c572b99 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.822961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.822961Z digest=sha256:08aa71e57ca71dc189e3de9e674fb60dd92c138061f238ae7c837e543a7865ee

Observation 4a022816-f474-4dee-a8d4-91f07793b44c · outbound

This paper cites Qwen2.5-vl, January 2025.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Qwen2.5-vl, January 2025

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.826248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.826248Z digest=sha256:5b941d17d17e77e6da4e14a853941cc7fe908a368653aa44f43dbb1d39f87d76

Observation 7df283ff-6895-4807-82ea-4fc136f625cf · outbound

This paper cites Learning from semantic alignment between unpaired multiviews for egocentric video recognition.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Learning from semantic alignment between unpaired multiviews for egocentric video recognition

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.833739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.830251Z digest=sha256:50eace55a29f6f436de374ce35063884b92ce737657e767fb019d87c1164215e

Observation cb910283-684b-4147-a1b4-b580779b192c · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs LVBench: An Extreme Long Video Understanding Benchmark

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.833497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.833497Z digest=sha256:e625cc4b0d3552b871412c28ee1fd7e17c358dcec22469f6143c28e6758271be

Observation df3eec7f-ade8-4dd1-aa0c-724fef2db1ab · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.837266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.837266Z digest=sha256:3cca2a0b5b934f96e58a18f1c7f9e5b7c877bd268a363951b823208879f6b248

Observation d85953e3-5241-408d-bb23-e2caedd1daad · outbound

This paper cites See what i see: Enabling user-centric robotic assistance using first-person demonstrations.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs See what i see: Enabling user-centric robotic assistance using first-person demonstrations

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.821670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.840921Z digest=sha256:94f131b7a0f829917d0f9bcffbb4d5b8724ded49ac513e2a261904a119d509cc

Observation 64bbbf1c-abf3-445b-b7c0-9aa7b7079950 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Chain-of-thought prompting elicits reasoning in large language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.844180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.844180Z digest=sha256:f9f62e2e4d037be14033d56a16c08ebf914a465b367c487d2c1fc1f612bc97a3

Observation 8b429db5-0f38-4f96-b01b-9955e9e23fba · outbound

This paper cites Incomplete multi-view domain adaptation via channel enhancement and knowledge transfer.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Incomplete multi-view domain adaptation via channel enhancement and knowledge transfer

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.802996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.847409Z digest=sha256:59ba0151a729bf27ccdb2b4c2e436823b0ef17e01be4ded9050c740b0853093a

Observation b44d3de0-c7ab-4f5a-a924-ee2341e72d14 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Next-qa: Next phase of question-answering to explaining temporal actions

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.850945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.850945Z digest=sha256:40654b7efac8808043e6bc0cb61bf2f12bdfb579346c8e81d7e5d0efd634ba84

Observation 518387af-c61e-4d7e-9ad7-7eadf9dcc0c3 · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Can i trust your answer? visually grounded video question answering

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.854230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.854230Z digest=sha256:66b1e9bae1d5bbbd27d31e13cff93420434c1a958d9d3063570d1d41b7763068

Observation f7a65284-674c-459a-819d-2171ddf567cc · outbound

This paper cites Pov: Prompt-oriented view-agnostic learning for egocentric hand-object interaction in the multi-view world.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Pov: Prompt-oriented view-agnostic learning for egocentric hand-object interaction in the multi-view world

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.777635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.857657Z digest=sha256:0700522787bcd83f4a4b942c1bd14103fa6a9fa50a07cc16024212156ac770e7

Observation 8e7d8b54-8e8c-480d-8c44-f16eef383f26 · outbound

This paper cites Retrieval-augmented egocentric video captioning.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Retrieval-augmented egocentric video captioning

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.766885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.860995Z digest=sha256:bb12fe204e831ce102673fb1201029c9134aed635d4969d2f6b8d3610a20aac2

Observation 5dcc23a0-8e38-46b0-88af-9278962a70a8 · outbound

This paper cites EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.864319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.864319Z digest=sha256:f361f38e0f97c25a4a5712b8537973d141e9a2372f9fd271a094e0f4de882d8d

Observation 834eaabb-6377-4be5-a98a-e4bd841d5949 · outbound

This paper cites Advancing high-resolution video-language representation with large-scale video transcriptions.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Advancing high-resolution video-language representation with large-scale video transcriptions

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.756120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.867879Z digest=sha256:67d148c0a27f926261d4efd205d23d2e53947d50ab37f0cef39ab0d68821cb92

Observation 596d892d-73ae-473a-b003-c984dd3d0763 · outbound

This paper cites Qwen2.5 Technical Report.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Qwen2.5 Technical Report

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.871155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.871155Z digest=sha256:f9bd39449954542261c0dc9eedb4e375fec5dd67b58894daa763fcc2aff30c6c

Observation c9136d0f-360b-4248-943e-1fa6c86beb79 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.874554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.874554Z digest=sha256:570c298e0544577f9afec41d8b8b8ab63d7696166bd80e2d81cc4c08fc823aeb

Observation dc6b34ac-2dc9-48c3-b89d-d1a83c4a3c1a · outbound

This paper cites Egolife: Towards egocentric life assistant.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Egolife: Towards egocentric life assistant

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.878139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.878139Z digest=sha256:ee52f26591da8a136f3a4fda89779981d02cbf70735cc161330564d7024c9b24

Observation c8b282ef-3762-4dd1-8294-00a2c0a0d065 · outbound

This paper cites Interact before align: Leveraging cross-modal knowledge for domain adaptive action recognition.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Interact before align: Leveraging cross-modal knowledge for domain adaptive action recognition

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.744952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.881438Z digest=sha256:c7224c10f81e01a24be57c3de3c5607e50d4c9d11728b2408ca9dd31c4c75a31

Observation 0580ab10-c00b-4d0d-a4f4-59aefd1d49d3 · outbound

This paper cites Fine-grained affordance annotation for egocentric hand-object interaction videos.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Fine-grained affordance annotation for egocentric hand-object interaction videos

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.734002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.884973Z digest=sha256:af8fcd970bd5ab43eabdb514d631af9921490fa493666bf14d0407bcb7a03e72

Observation 7a9b7ffe-15ef-43b2-8076-e826bd277000 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.888448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.888448Z digest=sha256:cafee2166716f7cfab1574b776227f0477874a701afd92281c7eb13187b94b81

Observation 375f7e57-2fd9-4c80-ac91-7aa5e1ac6ea9 · outbound

This paper cites Fusing personal and environmental cues for identification and segmentation of first-person camera wearers in third-person views.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Fusing personal and environmental cues for identification and segmentation of first-person camera wearers in third-person views

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.722825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.892194Z digest=sha256:4653be2cae87503cfa34fc610a2662e9fc7842a5458c746ae7c31e05c08ea10d

Observation c3c402f3-cf2f-4f73-ae74-400c88736e56 · outbound

This paper cites Progressive-Hint Prompting Improves Reasoning in Large Language Models.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Progressive-Hint Prompting Improves Reasoning in Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.895550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.895550Z digest=sha256:3390bf0c2fd2826f3dc835dc248472fa87bcf303a06624aa43ff611597cb089e

Observation c27682d3-9ae7-456c-a81e-129857e924a0 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs MLVU: Benchmarking Multi-task Long Video Understanding

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.899257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.899257Z digest=sha256:8c4adfc976a802e54ae1301540572930ee8f82a9d11cc37bee6406a466886e4c

Observation 808b2c60-d78f-4e2e-9c40-7b72bb381f05 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T18:20:58.902748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:20:58.902748Z digest=sha256:37c0599078e7d985a7f795265939a47a32faa4062696946002b1aee36f90b28b

Observation f24dfd04-ffee-4dec-9b00-3bb9237752ab · outbound

This paper cites an unresolved cited work.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:20:59.711139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.906977Z digest=sha256:0a0f9bc2a1b1323751b2be14a43d7067bd867ac3db50bab47280a8b03138e2b6

Observation d183422f-1521-4f58-b83f-5aced6b6d0e7 · outbound

This paper cites an unresolved cited work.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:20:59.700559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.910865Z digest=sha256:689f8d6234c8798a40f7e3f4d5ec23a8ccf24c5c5cbbf25c64cd61535e8d80c9

Observation 275dc58b-0209-4a90-9ede-77d618066f60 · outbound

This paper cites an unresolved cited work.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:20:59.689958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.914567Z digest=sha256:bd47124d39b6e1553baf20dce71d794551c6d1368a71cedec913e2eef4f955b5

Observation 991598cc-7fe2-4422-956b-b2630b39cd97 · outbound

This paper cites an unresolved cited work.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:20:59.679413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.917998Z digest=sha256:5125be05d84d8733e59de9b5ef395a1df9c6e0986e0e95e3f9fac777346434b8

Observation 3b0e048e-607e-41fe-8c3c-3f58574c13de · outbound

This paper cites Question.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Question

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.668778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.921699Z digest=sha256:0fa1496bd283c14f523438cc26d6f4943f67708587dc22eaa1b52337306b48dc

Observation 0f43294f-02f9-478d-831f-db6855b19560 · outbound

This paper cites an unresolved cited work.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:20:59.658026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.925362Z digest=sha256:ce76ca7aaacd7cb89fc5f3435e5b936b7b4540a74db311826516f53bc064c977

Observation 282679b6-7fd9-4da3-9f4a-9086d0ab3dfb · outbound

This paper cites Video 1" and the second video as.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Video 1" and the second video as

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.647911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.929948Z digest=sha256:78d8c146418019a94fb12013403914e4aa61aa1ce56271be36a4378223590fbf

Observation a3489ba2-d2e2-4206-affa-6062b7194dbf · outbound

This paper cites an unresolved cited work.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:20:59.636852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.933851Z digest=sha256:3d247c429004b1b95d12c4601a31141149879bf2286a3fb07429312d3b2f3596

Observation 785ac986-5495-4259-bf75-900708b3c800 · outbound

This paper cites an unresolved cited work.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:20:59.625591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.937450Z digest=sha256:97101f1b2c2f51a59cf906b79afbb95dea0fa11bf7882880b44e300619b0fa6e

Observation b4d7efa8-aef5-4ec6-9178-0a70a44bceb6 · outbound

This paper cites an unresolved cited work.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:20:59.615422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.940976Z digest=sha256:86f64a7225ff53e2f6472616c604d5d6521e126c077faca1263aa24029097874

Observation 4e6a24d5-3896-44e0-99fc-5e66bbd0fca1 · outbound

This paper cites Video 2:.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Video 2:

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.603633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.944844Z digest=sha256:653f3207ba3e755b15a741550e889c6ead3f09b9bf22b421af45e1ab0d22041d

Observation d323f144-c266-40cf-9e2e-cdd767555048 · outbound

This paper cites an unresolved cited work.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work

Reference 91

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:20:59.591877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.948326Z digest=sha256:f993639887367342f3d3c25b5bf2dd9cc53b2bdc1c6f7a6d97296d2e86f1e63c

Observation 5ff29efd-15c9-4534-afab-5b658022d0fb · outbound

This paper cites an unresolved cited work.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work

Reference 92

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:20:59.580983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.951798Z digest=sha256:caf3f69076ef784ff544c0c061809122766418913390f712072d7e4bbfad990c

Observation 828da7f9-eb37-450d-bc93-7f7af754d9d5 · outbound

This paper cites Conclusion: After reviewing the sequences, Option B correctly describes the actions.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Conclusion: After reviewing the sequences, Option B correctly describes the actions

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.568475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.955466Z digest=sha256:45c3dd0551a3af6ac0c760527814838c030368c2fd727a640cf603afb52a69b2

Observation d9295f79-8092-4f58-9914-c2b31e807622 · outbound

This paper cites an unresolved cited work.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work

Reference 94

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:20:59.557191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.959906Z digest=sha256:3cdad2629565bfda69524917a7ca0badf729d67959cd1931279b471eec1b42b4

Observation a60e31f7-1397-42b0-b532-4d555af401ab · outbound

This paper cites Video 2: The person within the bounding box is positioned at the bottom of the scene, facing upwards, and is in a different area relative to others.

EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Video 2: The person within the bounding box is positioned at the bottom of the scene, facing upwards, and is in a different area relative to others

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:20:59.545118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:20:58.963152Z digest=sha256:d2cd4d94eef6c67dbc23bc57c0e7b0189f94ced74300f8c51f714f19ddf4261a

Pith citing papers

Observation ded4086a-74cc-4ade-a563-ab2d0cdba6ca · inbound

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding cites this paper.

EgoExo-Con: Exploring View-Invariant Video Temporal Understanding EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T07:23:04.393109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:23:04.393109Z digest=sha256:00bf7dc5e61daaab8909ea719d80cf5f870dfb79afa044429214fc6f063c5bb5

Observation e0c2ee20-3038-46d8-8b9a-1409b44e840b · inbound

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting cites this paper.

SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:54:18.229368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T17:53:19.002603Z digest=sha256:2844439cc287dfd791c8eb3346135e62da96fbcb2c88b5ccec52fd1de55ea2af

Observation 2ecf6caa-271e-43b3-bc76-5a0fff9478e5 · inbound

The N-Body Problem: Parallel Execution from Single-Person Egocentric Video cites this paper.

The N-Body Problem: Parallel Execution from Single-Person Egocentric Video EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T16:55:20.757154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:55:20.757154Z digest=sha256:b1112bba86b1778b9fd2a12d3d9bf46a412c7d7636cae3b72d26bb409c3ae4db

Observation bb7a0976-f2e6-4f81-9200-b8ed1f5c95a0 · inbound

Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed Spaces cites this paper.

Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed Spaces EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T03:29:06.350552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:29:06.350552Z digest=sha256:999a7b2af26cf437c5f1367c520659244396ed139149409d3e25d1cdd378707e

Observation 13a31d8e-665d-4506-b2d4-d4238d16157d · inbound

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation cites this paper.

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:25:26.186466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T13:23:35.738229Z digest=sha256:428a8cfd2f8aa5f69495f1ed738eb5b7196ba6b68dfb0190f78d181c83efff22

Observation 46bb81c8-d476-42cf-a242-af7a1fed4cd8 · inbound

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation cites this paper.

From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T20:34:35.048469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:34:35.048469Z digest=sha256:26ecbdc3aa3f1a0cced2b515dcc0cfcb41dd17aa34232936aa52f16c50af2fab

Observation c59f0476-3ed6-4de5-bfbe-dea78629240a · inbound

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation cites this paper.

DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:26.137872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-07T13:50:23.835653Z digest=sha256:a1a34295639d6ca1290dd41953e2e68172eb2c389b59fd5a1d2c2e6bc5d9e482

Observation 37c82ded-8913-4a0d-9f9a-8e0e4ce53947 · inbound

Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models cites this paper.

Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:38:12.458834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T10:36:20.388169Z digest=sha256:8966c01e6d3cc55a223360e670471c685fc820de88a73c94fbd2e81650da91e6

Observation a2990647-57cd-42d9-985a-36ff731ae53d · inbound

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs cites this paper.

EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:22.284466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T05:53:05.450946Z digest=sha256:854a7e2424d09d56983e58f58774d2c6c2b74d004eb5999af9551daecdc0da3a

Observation dc39be0c-1d2b-44a0-be0c-473551593e09 · inbound

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis cites this paper.

VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:14:42.751993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T07:12:02.612292Z digest=sha256:74c1d06a11056952a62dc25a228e84782b74b133263fcb9306032780dc6afb28

Observation 4ed66bc6-0995-4ad6-b27a-fd3fa3cd3cc3 · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:55:23.228583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T04:54:23.077914Z digest=sha256:f7a41ffc0c16587b5ead83ebeb824b8b0a6d813589e596dafdc5847421eb34cd

Observation 7761571f-da78-4229-90ec-ab3281be825a · inbound

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering cites this paper.

CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:17.561558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-04T00:41:02.284215Z digest=sha256:e69bf7b4f212a11177d0553f0ea79d893897a1f5c61863cc8007b841e184038b

Observation aaba518a-16eb-410b-a63c-70c166379fb3 · inbound

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video cites this paper.

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:57.252484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T02:16:25.730555Z digest=sha256:1cbda7805cdcc40bd13879dec6207fc62c2976502d40f02ea027b49554767836

Observation 1ec72eaf-36e5-4e72-9cb2-369079c40204 · inbound

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos cites this paper.

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T05:01:06.200663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T05:01:06.200663Z digest=sha256:3e6904878ccf6749de6a9d9fc60aa123daed91e4c5b49a0f823c97a64a38b386

Observation 4cc2130b-ef76-41f9-831b-afa575490aff · inbound

Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests cites this paper.

Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:32:27.422997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:32:27.422997Z digest=sha256:25b8f8d33c1766bbb7d9c52a2ac6e2707d8c4c8b4139d0f86adfcf78731939c9