Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:20:58.963152Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 15 inbound Pith citation observations for arXiv:2507.18342.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:20:58.963152Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:32:27.422997Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T00:49:17.558383Z
95 of 95 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f2310575-32fd-4f90-bef6-c88311fa63f3 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Visual-policy learning through multi-camera view to single-camera view knowledge distillation for robot manipulation tasks
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32b58db8-8323-46c7-8617-6ad42750aa93 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Claude 3.7 sonnet and claude code, 2025
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67f8bc59-1544-40c0-82e4-0bc7bdb8fd29 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Observational learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3a41bf0-8eaa-4dc2-baec-da47b7e8841a · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 195e7901-9056-485b-bf87-80485589499a · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs In your place: neuropsychological evidence for altercentric remapping in embodied perspective taking
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a56965a1-0cd4-41bb-b21e-754d1fe5c59f · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Spatial memory: how egocentric and allocentric combine
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 500c462b-1311-4611-b691-1b571ffc462d · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Hourvideo: 1-hour video-language understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40bfb0cd-576b-4e7e-be9a-a50316f37e40 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs CG-Bench: Clue-grounded Question Answering Benchmark for Long Video Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c8cf13b-8852-4646-bd10-07e2c70b9d5b · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs put myself into your place
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ffa20f7-275f-40f7-9773-25e0efae1da7 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Scaling egocentric vision: The epic-kitchens dataset
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 343622d4-2d36-4e23-bbfa-a25025e05a77 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Epic-kitchens visor benchmark: Video segmentations and object relations
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e2b9dac-14ca-4ba8-88f9-950b36361009 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Gemini 2.5 pro, 2025
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd5530cb-e27c-4997-a7b7-8aa2d3169a73 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Interacting networks of brain regions underlie human spatial navigation: a review and novel synthesis of the literature
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 798a96e5-3f76-4954-8086-ca2d90d20441 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69c1d61f-42a0-4578-9dea-3a8343feaf39 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Complexity-Based Prompting for Multi-Step Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f37f27f-952d-4e41-a810-a0d090650b14 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Ego4d: Around the world in 3,000 hours of egocentric video
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34a89421-6e29-4280-aacb-d753faa57df3 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea24c61-3eb8-4d72-a07a-cae17bb42941 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Ava: A video dataset of spatio-temporally localized atomic visual actions
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec7a0169-2b96-4690-ab8c-ea79d5195bde · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Multiple human association and tracking from egocentric and complementary top views
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ce2ff04-61e6-4fb5-b252-e953c23626aa · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs What is modelled during observational learning? Journal of sports sciences, 25(5):531–545, 2007
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c900cf2-3c50-4a7f-a6a1-e3dbd1d9f79d · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs An ego-vision system for discovering human joint attention
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0bf67345-564d-4955-b7fb-7278172bdb99 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Improving action segmentation via graph-based temporal reasoning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5a595925-4655-4dc4-9c46-8cf167c170e7 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Egoexolearn: A dataset for bridging asynchronous ego-and exo-centric view of procedural activities in real world
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b5ad38a0-fd46-4150-b409-72d8b5758dce · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Vinci: A Real-time Embodied Smart Assistant based on Egocentric Vision-Language Model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55a6e0b9-07ba-42fc-8449-267a40846e03 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs GPT-4o System Card
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac2851b2-345a-4d95-b3d8-a5dc82ee5977 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f1e05fce-ca9f-4c0a-a436-620e78c46e79 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Egotaskqa: Understanding human tasks in egocentric videos
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fa0dee72-f454-4fa2-9618-c7dced4d2c91 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Large language models are zero-shot reasoners
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 322f785f-d3a1-4aa0-b614-7d8e03ee7c5b · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs A neural code for egocentric spatial maps in the human medial temporal lobe
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a6d62541-ba34-49f7-8966-6f94c836dac0 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs H2o: Two hands manipulating objects for first person interaction recognition
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3e8c75a4-a3f7-4ae7-8709-f41783d74906 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs LLaVA-OneVision: Easy Visual Task Transfer
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cd4631f-db5c-43c6-b8be-e66c8270d8b0 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs iGibson 2.0: Object-Centric Simulation for Robot Learning of Everyday Household Tasks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cf0da7c-2ee8-49dc-960f-bb2c2f3371d5 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs VideoChat: Chat-Centric Video Understanding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abe12be3-7929-4c38-aff4-45b63c0ef68d · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b84f0f9b-786a-4894-b00b-49e17b95a2e0 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Ego-exo: Transferring visual representa- tions from third-person to first-person videos
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 40fbdaf5-3353-4bcf-a1f1-e526057788ce · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs DeepSeek-V3 Technical Report
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b762a704-7475-4755-a799-80d8f08b9f74 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Skill transfer learning for autonomous robots and human–robot cooperation: A survey
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 24ef29ac-de58-4737-831d-e7a86a718373 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Hoi4d: A 4d egocentric dataset for category-level human-object interaction
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2c29678-f4a0-4ad3-b74e-81883267bf42 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs NVILA: Efficient Frontier Visual Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7d8ac96-6b1e-4ad1-a8f5-806e8f3421ca · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Egoschema: A diagnostic benchmark for very long-form video language understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a5af1db1-350c-4fce-a7fb-e54fb3c7546b · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Howto100m: Learning a text-video embedding by watching hundred million narrated video clips
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45887d3f-a001-41de-9ec1-2a5e5fea8579 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Gpt-4o mini: advancing cost-efficient intelligence, 07 2024
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d4856754-4a7b-4506-bd2e-896f5c28f8d7 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Egovlpv2: Egocentric video-language pre-training with fusion in the backbone
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a573c7e-9f12-4203-97d3-0fdadb4bdb56 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs EgoMe: A New Dataset and Challenge for Following Me via Egocentric View in Real World
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 922753bc-70f8-40d7-b25d-958dfa7bb363 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Synchronization is All You Need: Exocentric-to-Egocentric Transfer for Temporal Action Segmentation with Unlabeled Synchronized Video Pairs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42fb4af6-ce29-4ddd-97e6-b9505cb88f72 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs The meccano dataset: Understanding human-object interactions from egocentric videos in an industrial-like domain
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 33f12641-5c0c-4e36-b1c9-49807e2031d9 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Home action genome: Cooperative compositional action understanding
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2d56ccb1-da42-45fb-8f52-bcac6016e0d7 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Watch and learn: the cognitive neuroscience of learning from others’ actions
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f1b7abe3-f17f-4e4d-a059-ab720e673389 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f4cd9b7-d2ee-4c50-bc57-68db6b5e666b · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a7c4021-baec-4106-a456-4cbce44a88a1 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Sener, D
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation acded3b6-d552-48e6-ae4b-96d140bfdd78 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Self-supervised disentangled representation learning for third-person imitation learning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6c33268e-7a7f-4cba-86ea-05c5b4051135 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Third-person visual imitation learning via decoupled hierarchical controller
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation defd28b4-0e73-4694-a699-09591d2c2629 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Actor and observer: Joint modeling of first and third-person videos
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 08cbaa41-583a-40de-8add-546444c82e9a · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Ego4d goal-step: Toward hierarchical understanding of procedural activities
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3199f91d-1a2c-4add-8a29-6bf69c572b99 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a022816-f474-4dee-a8d4-91f07793b44c · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Qwen2.5-vl, January 2025
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7df283ff-6895-4807-82ea-4fc136f625cf · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Learning from semantic alignment between unpaired multiviews for egocentric video recognition
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation cb910283-684b-4147-a1b4-b580779b192c · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs LVBench: An Extreme Long Video Understanding Benchmark
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df3eec7f-ade8-4dd1-aa0c-724fef2db1ab · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d85953e3-5241-408d-bb23-e2caedd1daad · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs See what i see: Enabling user-centric robotic assistance using first-person demonstrations
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 64bbbf1c-abf3-445b-b7c0-9aa7b7079950 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Chain-of-thought prompting elicits reasoning in large language models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b429db5-0f38-4f96-b01b-9955e9e23fba · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Incomplete multi-view domain adaptation via channel enhancement and knowledge transfer
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b44d3de0-c7ab-4f5a-a924-ee2341e72d14 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Next-qa: Next phase of question-answering to explaining temporal actions
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 518387af-c61e-4d7e-9ad7-7eadf9dcc0c3 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Can i trust your answer? visually grounded video question answering
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7a65284-674c-459a-819d-2171ddf567cc · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Pov: Prompt-oriented view-agnostic learning for egocentric hand-object interaction in the multi-view world
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8e7d8b54-8e8c-480d-8c44-f16eef383f26 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Retrieval-augmented egocentric video captioning
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5dcc23a0-8e38-46b0-88af-9278962a70a8 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 834eaabb-6377-4be5-a98a-e4bd841d5949 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Advancing high-resolution video-language representation with large-scale video transcriptions
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 596d892d-73ae-473a-b003-c984dd3d0763 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Qwen2.5 Technical Report
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9136d0f-360b-4248-943e-1fa6c86beb79 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc6b34ac-2dc9-48c3-b89d-d1a83c4a3c1a · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Egolife: Towards egocentric life assistant
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8b282ef-3762-4dd1-8294-00a2c0a0d065 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Interact before align: Leveraging cross-modal knowledge for domain adaptive action recognition
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0580ab10-c00b-4d0d-a4f4-59aefd1d49d3 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Fine-grained affordance annotation for egocentric hand-object interaction videos
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7a9b7ffe-15ef-43b2-8076-e826bd277000 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 375f7e57-2fd9-4c80-ac91-7aa5e1ac6ea9 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Fusing personal and environmental cues for identification and segmentation of first-person camera wearers in third-person views
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c3c402f3-cf2f-4f73-ae74-400c88736e56 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Progressive-Hint Prompting Improves Reasoning in Large Language Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c27682d3-9ae7-456c-a81e-129857e924a0 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs MLVU: Benchmarking Multi-task Long Video Understanding
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 808b2c60-d78f-4e2e-9c40-7b72bb381f05 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f24dfd04-ffee-4dec-9b00-3bb9237752ab · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d183422f-1521-4f58-b83f-5aced6b6d0e7 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 275dc58b-0209-4a90-9ede-77d618066f60 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 991598cc-7fe2-4422-956b-b2630b39cd97 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3b0e048e-607e-41fe-8c3c-3f58574c13de · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Question
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0f43294f-02f9-478d-831f-db6855b19560 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 282679b6-7fd9-4da3-9f4a-9086d0ab3dfb · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Video 1" and the second video as
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a3489ba2-d2e2-4206-affa-6062b7194dbf · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 785ac986-5495-4259-bf75-900708b3c800 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b4d7efa8-aef5-4ec6-9178-0a70a44bceb6 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4e6a24d5-3896-44e0-99fc-5e66bbd0fca1 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Video 2:
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d323f144-c266-40cf-9e2e-cdd767555048 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5ff29efd-15c9-4534-afab-5b658022d0fb · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 828da7f9-eb37-450d-bc93-7f7af754d9d5 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Conclusion: After reviewing the sequences, Option B correctly describes the actions
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d9295f79-8092-4f58-9914-c2b31e807622 · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Unresolved cited work
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a60e31f7-1397-42b0-b532-4d555af401ab · outbound
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs Video 2: The person within the bounding box is positioned at the bottom of the scene, facing upwards, and is in a different area relative to others
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ded4086a-74cc-4ade-a563-ab2d0cdba6ca · inbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0c2ee20-3038-46d8-8b9a-1409b44e840b · inbound
SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2ecf6caa-271e-43b3-bc76-5a0fff9478e5 · inbound
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb7a0976-f2e6-4f81-9200-b8ed1f5c95a0 · inbound
Thinking in Structures: Evaluating Spatial Intelligence in Constraint-Governed Spaces EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13a31d8e-665d-4506-b2d4-d4238d16157d · inbound
From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 46bb81c8-d476-42cf-a242-af7a1fed4cd8 · inbound
From Synchrony to Sequence: Exo-to-Ego Generation via Interpolation EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c59f0476-3ed6-4de5-bfbe-dea78629240a · inbound
DenseStep2M: A Scalable, Training-Free Pipeline for Dense Instructional Video Annotation EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 37c82ded-8913-4a0d-9f9a-8e0e4ce53947 · inbound
Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a2990647-57cd-42d9-985a-36ff731ae53d · inbound
EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dc39be0c-1d2b-44a0-be0c-473551593e09 · inbound
VGenST-Bench: A Benchmark for Spatio-Temporal Reasoning via Active Video Synthesis EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4ed66bc6-0995-4ad6-b27a-fd3fa3cd3cc3 · inbound
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7761571f-da78-4229-90ec-ab3281be825a · inbound
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation aaba518a-16eb-410b-a63c-70c166379fb3 · inbound
LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1ec72eaf-36e5-4e72-9cb2-369079c40204 · inbound
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cc2130b-ef76-41f9-831b-afa575490aff · inbound
Spatial-IQ: Deconstructing Spatial Intelligence via Hierarchical Capability Tests EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.