Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:11:15.122695Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 4 inbound Pith citation observations for arXiv:2506.17629.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:11:15.122695Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:21:29.776719Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-17T04:08:59.824731Z
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c88ea755-e6c0-420f-a647-b26e2366fe41 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Remembr: Building and reasoning over long- horizon spatio-temporal memory for robot navigation
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0ddeae4d-d437-4f04-9da2-a723437a9cf8 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f914b26-6b94-4f76-98d4-fcb94b281d5e · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9be3426a-da38-44c2-a187-b1000d024a2a · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Cafuser: Condition-aware multimodal fusion for robust semantic perception of driving scenes.IEEE Robotics and Automation Letters, 2025
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d7e7d189-72dd-4cd7-a9b0-cc308bb2bd38 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea286289-857c-4ef9-8025-4c5b2cf3e324 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Embodied question answer- ing
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation eceea4e6-d709-4315-abac-b2fb5d308170 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cf23b82-8498-474f-b127-b2271b84c125 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Videoagent: A memory-augmented mul- timodal agent for video understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9b97ae9f-b38a-4e0e-8eae-176329a48208 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d2c6250-70f5-4be3-9c5a-cddea093f0e6 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Objectrelator: Enabling cross-view object rela- tion understanding across ego-centric and exo-centric per- spectives
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 42d19641-6415-495a-93d9-097d6f6852ac · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Ego4d: Around the world in 3,000 hours of egocentric video
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a33312fa-4b74-44fc-a458-c3804a03c990 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning A survey of deep learning techniques for autonomous driving.Journal of field robotics, 37(3):362– 386, 2020
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b04b94a0-179f-4609-bd50-62918aa1362b · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Open vocabulary multi- label video classification
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1e1e9735-5ac2-4121-9370-5d275f21e18e · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Sequential multi-object grasping with one dexterous hand
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2d2b7f61-358d-43a7-881e-f901f3eba15a · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning ChatDB: Augmenting LLMs with Databases as Their Symbolic Memory
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4739f48-111e-47b9-a529-13c37b7e8077 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Vtimellm: Empower llm to grasp video moments
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c97ef85-f5ba-4fc2-926f-7ae99b48003c · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Egotaskqa: Understanding human tasks in egocentric videos
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d6159f40-45fc-4137-b1c8-affd4ba9e33c · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Context-aware planning and environment-aware memory for instruction following em- bodied agents
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ff7b1043-b31e-419f-8804-8e98638da4c8 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning DeepSeek-V3 Technical Report
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83a1f811-dee2-4958-b4fe-fe50e014fbc2 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed89a511-4efe-4444-bcd2-fbbc92b60f6a · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Aligning cyber space with physical world: A comprehensive survey on embodied ai.IEEE/ASME Transactions on Mechatronics, 2025
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e3e2a5a0-9fc5-4b46-a50b-6f6492bce91e · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning A Survey on Vision-Language-Action Models for Embodied AI
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3378b811-4abc-4fe7-afda-5c2cd7f7406b · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Openeqa: Embodied question answering in the era of foun- dation models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 27352482-e06b-40f3-a7bc-1ade7ebc192c · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Egoschema: A diagnostic benchmark for very long- form video language understanding.Advances in Neural In- formation Processing Systems, 36:46212–46244, 2023
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c3d13be7-5da6-4c23-9a6a-337731811b1d · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning MemGPT: Towards LLMs as Operating Systems
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb242e3-a4ea-457c-9d14-d078fbc2945b · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning V2-sam: Mar- rying sam2 with multi-prompt experts for cross-view object correspondence.CVPR, 2026
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7abbdd78-a7b5-4447-a4f6-a9625c299c44 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f1d8620f-2305-4dcb-9292-d78a4b4f7b79 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77d510d2-e44b-40d6-9ba0-57e218681358 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Omnia de egotempo: Benchmarking temporal understanding of multi-modal llms in egocentric videos
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 37a66466-2bb3-4eeb-916d-dcd03f3ee448 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Improving language understanding by gen- erative pre-training
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f8b06ee-aa50-4191-86e9-99aa0e57a90f · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Character-llm: A trainable agent for role-playing
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e58510fb-2bab-47e2-868e-a3cc19b0bdd9 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Moviechat: From dense token to sparse memory for long video understanding
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fe840ef-5720-4867-90a8-8014fa7b8bac · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Part, Ioannis Papaioannou, Arash Eshghi, Ioannis Konstas, and Oliver Lemon
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8b145558-78f5-46ec-a715-a83df2731935 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Mac: a meta-learning approach for fea- ture learning and recombination.Pattern Analysis and Ap- plications, 27(2):63, 2024
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a6def184-c419-49f3-a854-58a7b1b38159 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning LLaMA: Open and Efficient Foundation Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76c53c38-ea5b-4444-a9d7-10b86e9e747f · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Ocra: Object-centric learning with 3d and tactile priors for human-to-robot action transfer.arXiv preprint arXiv:2603.14401, 2026
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de32a317-a7c5-4105-beeb-3824e4cea282 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af51e76d-04ab-457a-8371-9b9b83cba00c · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Videotree: Adaptive tree-based video representation for llm reasoning on long videos
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a3563688-65bb-4c91-80b4-0552db3507bb · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Lighttrack: Finding lightweight neural net- works for object tracking via one-shot architecture search
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 171999ef-dca6-4d0e-9888-6e7024ac3186 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Qwen2.5 Technical Report
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa6c9143-a734-4197-9b9c-8265a34361f5 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning So- cratic models: Composing zero-shot multimodal reasoning with language
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b627dae2-a079-487f-8340-cce3aa692acd · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9de21945-e071-4fa5-b778-410270d01b50 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 490c09d0-828f-4d01-a647-ed0d7877173e · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 368d900c-88a4-42bd-bd2c-2196ed159593 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Considering hardware limitations and ef- ficiency, we preprocess videos by sampling frames at 0.5 FPS on OpenEQA and EgoSchema, while limiting frame count to 32 for EgoTempo
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation cd6243c8-5e00-4e97-b4b1-c09ac06c76ec · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a500b113-a6fc-41ff-89eb-4a0f364785ed · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning The results are presented in Table 6
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 580dda91-a3db-45a6-89c0-5bb6bf94f7e0 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7791ed57-e883-47e7-9ad0-03d8c3784990 · outbound
CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning We highlight relevant infor- mation throughout the reasoning trace, marking correct de- tails in green and erroneous ones in red
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8b986300-5312-4992-9628-5e010a6e1a45 · inbound
V$^{2}$-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7c1461e4-1bbd-49a4-854b-b25ce684d2b6 · inbound
ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d7f51606-6c0e-42af-99f7-38c41c4575c6 · inbound
EgoSound: Benchmarking Sound Understanding in Egocentric Videos CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e5bd53a0-5e08-49a5-8aaa-a277d9c73123 · inbound
The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.