Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T07:23:09.304749Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 4 inbound Pith citation observations for arXiv:2510.26113.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T07:23:09.304749Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T20:52:48.921328Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T11:18:13.630096Z
64 of 64 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c67ea8dd-45f5-495f-ab41-260ab92a9656 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6ba7251-ab3b-4e11-bf02-140b50c8fb7b · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 818ad9fd-dacf-47d6-baf8-31ac7beebc4c · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4362ca1e-a179-4a0b-bbd6-3a949bbea7b0 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Egothink: Evaluating first-person perspective thinking capability of vision-language models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 924251dc-ef6d-422f-9cfc-fcf651a4d180 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65aeb1ee-2fe7-48cd-a543-de5a1b93bcda · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bda65d5-9efd-44e8-89ea-24f8062f3b04 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Lost in Time: A New Temporal Benchmark for VideoLLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1cdb779-09db-4959-a7cf-87bfa3f5cdb8 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Grounded question-answering in long egocentric videos
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00828e5b-4d62-4af0-8a61-3127e25064f2 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24f6e97c-0632-4ae9-967e-a0d4b26fa599 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e522de50-7231-4245-8e73-fed5048b2eec · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Tall: Temporal activity localization via language query
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b0f31df-a307-4b22-a7c5-068bec89bbaa · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40da9d03-56e1-44bc-8e62-0621b546d254 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6230194e-7128-41ae-9e16-2bf574635c34 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Ego-exo4d: Understanding skilled human activity from first-and third-person perspectives
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 204b3c85-64a2-4538-ab67-1203535fc192 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ded4086a-74cc-4ade-a563-ab2d0cdba6ca · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f04e3edc-26d8-48cc-9088-56a511600f8b · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Lora: Low-rank adaptation of large language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3bcb6c2-a76f-4a19-99c6-c2d2e8cfc221 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Vtimellm: Empower llm to grasp video moments
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9995c41-a9e8-4e83-9bf5-67a22059a791 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efe5f9e7-559c-47b5-981e-e66e4ce67e81 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Background-aware moment detection for video moment retrieval
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72182d6e-0cb4-4c48-8c12-713a858c1fd5 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding On the consistency of video large language models in temporal comprehension
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b956a3b-ee12-4fd9-9ec4-217354adda91 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7eb983c-3338-409c-9783-3fcefa8a01d5 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Ego-exo: Transferring visual representations from third-person to first-person videos
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9ea4eec-9311-461a-8e4c-a37c6bb2a7d6 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Egoexo-fitness: Towards egocentric and exocentric full-body action understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b220119e-658b-4c3a-b551-f46cb3136cad · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Universal video temporal grounding with generative multi-modal large language models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85ebea19-bc89-4b8f-a73c-c97ba66a00b4 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Improved Visual-Spatial Reasoning via R1-Zero-Like Training
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e29ac5f-6fcf-45c7-a84d-1367cdd3281a · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Is Your Video Language Model a Reliable Judge?
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 576e35da-897f-4682-987b-2841d57b5bef · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Put myself in your shoes: Lifting the egocentric perspective from exocentric videos
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28cdc54f-0505-495b-b2ab-f5e82d964198 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Viewpoint rosetta stone: Unlocking unpaired ego-exo videos for view-invariant representation learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc64ff55-0d78-41f6-ac22-f0bb340aafde · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Egoschema: A diagnostic benchmark for very long-form video language understanding
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afac3c60-b6db-49b0-b1f4-a3103d8b5ea1 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Chrono: A simple blueprint for representing time in mllms
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 398e02e3-7de2-4b68-b19a-4efd16ad01ba · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Introducing gpt-5
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0d565e6-f2a1-4b25-91ef-a253ef8806e0 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a1bc3f7-50ca-4392-8f64-b594b993cd5a · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Momentor: Advancing Video Large Language Model with Fine-Grained Temporal Reasoning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a609b8e1-bbcd-48cf-955f-ee55acf01d50 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7998972-807d-446f-a68c-d266f04d7fc0 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62aa2089-f183-4bf9-adaa-b73f4c9b87ad · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Assembly101: A large-scale multi-view video dataset for understanding procedural activities
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07c9a36e-0566-4394-bdbe-94e24e9b92c1 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 058739ef-4f61-4791-94b1-27af30bbb50d · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebf6d6a1-6ad0-4b0a-aa6d-c88e799f7210 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Actor and observer: Joint modeling of first and third-person videos
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75fc005a-618e-4748-866b-52ccca80e918 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Learning from semantic alignment between unpaired multiviews for egocentric video recognition
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 451f6820-401e-48a2-80a4-a0c50b14720d · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1154590-e902-4fc2-94e3-dfcfe8eeecbe · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 092dee13-de5c-4c06-89c5-20e794ea9026 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b064333-15cc-4f7f-9284-4c40a6317838 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fa2f2a3-cd9b-42d7-ba04-97b5d38e51c1 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding HawkEye: Training Video-Text LLMs for Grounding Text in Videos
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21ea69d5-bc46-41b4-9c8f-a27a96bdae55 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Longvideobench: A benchmark for long-context interleaved video-language understanding
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7089af4-5a67-4c21-b09b-25ae1ad9482c · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Next-qa: Next phase of question-answering to explaining temporal actions
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 538d22a1-53db-4d21-a70a-015b0e21c93e · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Can i trust your answer? visually grounded video question answering
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d49cc283-871f-40cb-8554-8c45fbb62f66 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Egoblind: Towards egocentric visual assistance for the blind people
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b3121aa-73da-422d-9565-dc6e6dfe64e0 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Text2Reward: Reward Shaping with Language Models for Reinforcement Learning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3c79df9-f7ee-4172-89b7-79e38a469b74 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Video question answering via gradually refined attention over appearance and motion
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec8daf29-db2e-48ef-a8c7-c7a745bd8571 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Learning fine-grained view-invariant representations from unpaired ego-exo videos via temporal alignment
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da52ac15-2194-4fb7-adb0-42283e78c112 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Qwen3 Technical Report
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e57d479-bd9e-4d1f-9d1e-c94d17e07899 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Mmego: Towards building egocentric multimodal llms for video qa
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d68df159-68ed-458a-89c6-b4686498c6a7 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5502050-69e0-428c-8a3b-03e29de7de9a · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a4eb131-d608-4e6a-b021-23fcbd145c25 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6803488-262b-4bec-b68a-a90d1ae860ff · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Exo2ego: Exocentric knowledge guided mllm for egocentric video understanding
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 738943dd-7f6f-4f19-83b6-12b008fa9194 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Long Context Transfer from Language to Vision
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f503adc7-1cf5-4832-9b23-ae3118e44777 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eaca711-a83a-4a5a-a29c-ad30b18f7463 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 510b3cf6-f4ea-4ced-8f42-9f164f620dfd · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e63d3dd1-f44a-4a8c-b06f-f90c0545d318 · outbound
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding Mlvu: Benchmarking multi-task long video understanding
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 886e0750-ebdb-4523-9acd-ecf55f158a3e · inbound
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 109af9ed-c9ae-4cbe-83e4-953f092c6891 · inbound
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 765d3134-f777-456a-9bb8-1b59c2d926c8 · inbound
Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 883e2ed7-8b35-4b16-afe6-26bb39d7abb0 · inbound
EgoExoMem: Cross-View Memory Reasoning over Synchronized Egocentric and Exocentric Videos EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.