Pith. sign in

Paper Citation Record · LEDGER

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning

As of 16 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 4 inbound Pith citation observations for arXiv:2506.17629.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17629 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:11:15.122695Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:21:29.776719Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T04:08:59.824731Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved22
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c88ea755-e6c0-420f-a647-b26e2366fe41 · outbound

This paper cites Remembr: Building and reasoning over long- horizon spatio-temporal memory for robot navigation.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Remembr: Building and reasoning over long- horizon spatio-temporal memory for robot navigation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.890596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:14.919634Z digest=sha256:86e78232f2245a9dc1fffe0537788dc5e8ba4f9df1c5be3f417345905ffc2e5b

Observation 0ddeae4d-d437-4f04-9da2-a723437a9cf8 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:14.924830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:14.924830Z digest=sha256:8d49a06afbec8f34fd52ae64ca8dc226304e64b18cd6f2425f7710beb73d0f3f

Observation 4f914b26-6b94-4f76-98d4-fcb94b281d5e · outbound

This paper cites Qwen2.5-VL Technical Report.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:14.929732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:14.929732Z digest=sha256:4ecba3423afbdfd48a79874810992c8b7082f73286d2613970e417b17c61af9c

Observation 9be3426a-da38-44c2-a187-b1000d024a2a · outbound

This paper cites Cafuser: Condition-aware multimodal fusion for robust semantic perception of driving scenes.IEEE Robotics and Automation Letters, 2025.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Cafuser: Condition-aware multimodal fusion for robust semantic perception of driving scenes.IEEE Robotics and Automation Letters, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.877183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:14.934080Z digest=sha256:bf1666ff3bc3131df1b597f1115993ad99dfce9d96b01c0b3f8f5c1a9d932789

Observation d7e7d189-72dd-4cd7-a9b0-cc308bb2bd38 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:14.938127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:14.938127Z digest=sha256:a90c7e059c7d14c3ff1de36445ff182e67ac6f3bc7aa3b6f26b6acf4151de7dc

Observation ea286289-857c-4ef9-8025-4c5b2cf3e324 · outbound

This paper cites Embodied question answer- ing.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Embodied question answer- ing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.863511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:14.942516Z digest=sha256:4ccc626ac6aedabcbc3ae7523bd5374d08c0dba3169d3f7c31a4cab3a674853b

Observation eceea4e6-d709-4315-abac-b2fb5d308170 · outbound

This paper cites an unresolved cited work.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:14.947269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:14.947269Z digest=sha256:15e55208d17b2966fd31b753587e6b328f8f22773d40ba841ff16ce011352d8d

Observation 3cf23b82-8498-474f-b127-b2271b84c125 · outbound

This paper cites Videoagent: A memory-augmented mul- timodal agent for video understanding.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Videoagent: A memory-augmented mul- timodal agent for video understanding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.840637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:14.951436Z digest=sha256:2e97d0c61727262588b8e905ccfdbb9ef435f1dd50d637db4263a9f2a2d29a40

Observation 9b97ae9f-b38a-4e0e-8eae-176329a48208 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:14.955692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:14.955692Z digest=sha256:de77c76896d9ad325ce4a16c35bdda46209d0999c247101fcc3be33b721c9a53

Observation 4d2c6250-70f5-4be3-9c5a-cddea093f0e6 · outbound

This paper cites Objectrelator: Enabling cross-view object rela- tion understanding across ego-centric and exo-centric per- spectives.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Objectrelator: Enabling cross-view object rela- tion understanding across ego-centric and exo-centric per- spectives

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.825270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:14.960706Z digest=sha256:d9050176d735bb9476dc56ba23c15eccfaf7989007aab750c3c797332691cc84

Observation 42d19641-6415-495a-93d9-097d6f6852ac · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Ego4d: Around the world in 3,000 hours of egocentric video

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.811481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:14.964752Z digest=sha256:686beb8f61d83a53ba12619230833669bd903a1ae8dfe8f7554623628ebcf778

Observation a33312fa-4b74-44fc-a458-c3804a03c990 · outbound

This paper cites A survey of deep learning techniques for autonomous driving.Journal of field robotics, 37(3):362– 386, 2020.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning A survey of deep learning techniques for autonomous driving.Journal of field robotics, 37(3):362– 386, 2020

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.797307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:14.968640Z digest=sha256:f0cbcdd77fb4cfc37701c042c4a113d474621e7536274219cf4c92146aa42700

Observation b04b94a0-179f-4609-bd50-62918aa1362b · outbound

This paper cites Open vocabulary multi- label video classification.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Open vocabulary multi- label video classification

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.784150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:14.972620Z digest=sha256:5edbf6e52f154bbac04f214ac32479b8db3c6577c8332c106a27f80087c3466c

Observation 1e1e9735-5ac2-4121-9370-5d275f21e18e · outbound

This paper cites Sequential multi-object grasping with one dexterous hand.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Sequential multi-object grasping with one dexterous hand

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.771258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:14.976838Z digest=sha256:20d3302c61475ce066bf265a1b045a27a46786a7ff8233e804ed79fc2f3a6dd7

Observation 2d2b7f61-358d-43a7-881e-f901f3eba15a · outbound

This paper cites ChatDB: Augmenting LLMs with Databases as Their Symbolic Memory.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning ChatDB: Augmenting LLMs with Databases as Their Symbolic Memory

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:14.980788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:14.980788Z digest=sha256:afcf1e17d7ec987c1ddb0736fa854175130866117497dc6ab1d80700928b92a8

Observation c4739f48-111e-47b9-a529-13c37b7e8077 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Vtimellm: Empower llm to grasp video moments

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:14.984953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:14.984953Z digest=sha256:93491fa87da963676ec0a0eea119ff18c3cfec093bac7fb4b9f5fa0410373e02

Observation 7c97ef85-f5ba-4fc2-926f-7ae99b48003c · outbound

This paper cites Egotaskqa: Understanding human tasks in egocentric videos.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Egotaskqa: Understanding human tasks in egocentric videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.749389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:14.988940Z digest=sha256:20c387952216a0e164baf406368d62cabad108a0807df9bef6b26f02d74bd1ec

Observation d6159f40-45fc-4137-b1c8-affd4ba9e33c · outbound

This paper cites Context-aware planning and environment-aware memory for instruction following em- bodied agents.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Context-aware planning and environment-aware memory for instruction following em- bodied agents

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.735632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:14.993172Z digest=sha256:17e321aae023bcb18337f3e56895b0c90d0ce8ddd9d3ce9c800d8628d0a7582c

Observation ff7b1043-b31e-419f-8804-8e98638da4c8 · outbound

This paper cites DeepSeek-V3 Technical Report.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning DeepSeek-V3 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:14.997487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:14.997487Z digest=sha256:b8e70c754270d862730c8e5c88541cc3db231541166c1c0abd2989aeed8fa164

Observation 83a1f811-dee2-4958-b4fe-fe50e014fbc2 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.001732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.001732Z digest=sha256:a0217f874488197a7b7997c5233169d4c2e06287b95605d87f349a72f1b53ba0

Observation ed89a511-4efe-4444-bcd2-fbbc92b60f6a · outbound

This paper cites Aligning cyber space with physical world: A comprehensive survey on embodied ai.IEEE/ASME Transactions on Mechatronics, 2025.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Aligning cyber space with physical world: A comprehensive survey on embodied ai.IEEE/ASME Transactions on Mechatronics, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.710217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.005682Z digest=sha256:77557cfc93f709e26293fec3b7881236a2e327507c77922ad417d661d9cbbbf3

Observation e3e2a5a0-9fc5-4b46-a50b-6f6492bce91e · outbound

This paper cites A Survey on Vision-Language-Action Models for Embodied AI.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning A Survey on Vision-Language-Action Models for Embodied AI

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.009577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.009577Z digest=sha256:62ec9e5e3a75a257ee92a80bac39d6325cb2d6ea6543983059b045ea46cab147

Observation 3378b811-4abc-4fe7-afda-5c2cd7f7406b · outbound

This paper cites Openeqa: Embodied question answering in the era of foun- dation models.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Openeqa: Embodied question answering in the era of foun- dation models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.695086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.013900Z digest=sha256:a1d62fba0ce5e578ef393498dd5677f0b951044ba94d9fa3cf0349357762ce09

Observation 27352482-e06b-40f3-a7bc-1ade7ebc192c · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.Advances in Neural In- formation Processing Systems, 36:46212–46244, 2023.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Egoschema: A diagnostic benchmark for very long- form video language understanding.Advances in Neural In- formation Processing Systems, 36:46212–46244, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.681595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.017972Z digest=sha256:3e295ff786ec26cd2e6c5b39ee0f9790fe969d4aaed9eaf978b474774bb88b58

Observation c3d13be7-5da6-4c23-9a6a-337731811b1d · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning MemGPT: Towards LLMs as Operating Systems

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.021957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.021957Z digest=sha256:ca76fca0db1282e221aa0236137f4a4b06b72c66027621d4cf3dd6d354d23091

Observation bfb242e3-a4ea-457c-9d14-d078fbc2945b · outbound

This paper cites V2-sam: Mar- rying sam2 with multi-prompt experts for cross-view object correspondence.CVPR, 2026.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning V2-sam: Mar- rying sam2 with multi-prompt experts for cross-view object correspondence.CVPR, 2026

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.667593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.026221Z digest=sha256:929d8419068e33d70a29e1e7f2f9168e1957d685699ece05a1c702724ff35c8b

Observation 7abbdd78-a7b5-4447-a4f6-a9625c299c44 · outbound

This paper cites an unresolved cited work.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:11:15.652598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.030352Z digest=sha256:b82562cd274d12f489fcd8a1e8f5ce8398ebcb1394ffc6d6c278af4b237fa201

Observation f1d8620f-2305-4dcb-9292-d78a4b4f7b79 · outbound

This paper cites Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.034372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.034372Z digest=sha256:a5d8b57bccf1dcdfb540ce6a40d0387e64245b0c1f993d49c3c77e29e2e9eb4c

Observation 77d510d2-e44b-40d6-9ba0-57e218681358 · outbound

This paper cites Omnia de egotempo: Benchmarking temporal understanding of multi-modal llms in egocentric videos.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Omnia de egotempo: Benchmarking temporal understanding of multi-modal llms in egocentric videos

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.639266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.038662Z digest=sha256:225f539065ba6492e1690f5a13c91148f6a33d9b8c4b3cdd3e7e2bf8bd60e988

Observation 37a66466-2bb3-4eeb-916d-dcd03f3ee448 · outbound

This paper cites Improving language understanding by gen- erative pre-training.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Improving language understanding by gen- erative pre-training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.042504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.042504Z digest=sha256:131bb6b322fb22a374d6cf5f0c96302e860f12ee7076729a2078b2dc49d4087d

Observation 6f8b06ee-aa50-4191-86e9-99aa0e57a90f · outbound

This paper cites Character-llm: A trainable agent for role-playing.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Character-llm: A trainable agent for role-playing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.616461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.046328Z digest=sha256:911afaa06af190081b571e70f3d0ad29ded6dd47a556c25060f5b7980ad5c754

Observation e58510fb-2bab-47e2-868e-a3cc19b0bdd9 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Moviechat: From dense token to sparse memory for long video understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.050586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.050586Z digest=sha256:979ad83f8290db6e80802b470255ebfb9ef2ec819510a787377b91645e969ce3

Observation 4fe840ef-5720-4867-90a8-8014fa7b8bac · outbound

This paper cites Part, Ioannis Papaioannou, Arash Eshghi, Ioannis Konstas, and Oliver Lemon.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Part, Ioannis Papaioannou, Arash Eshghi, Ioannis Konstas, and Oliver Lemon

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.594670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.054683Z digest=sha256:db0caee80a89377aee9b77ee1ab10dc285eec9dbea6f9ae3d7c7734e422cbe52

Observation 8b145558-78f5-46ec-a715-a83df2731935 · outbound

This paper cites Mac: a meta-learning approach for fea- ture learning and recombination.Pattern Analysis and Ap- plications, 27(2):63, 2024.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Mac: a meta-learning approach for fea- ture learning and recombination.Pattern Analysis and Ap- plications, 27(2):63, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.581786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.058811Z digest=sha256:cfee2a05928de64710ad231c9d4baad1aed5b7903cc456a10b67e85aba1d32d0

Observation a6def184-c419-49f3-a854-58a7b1b38159 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning LLaMA: Open and Efficient Foundation Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.062766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.062766Z digest=sha256:9d9969a848a1571a708d7b51b1515745c991133fc1dfec86c7ec077adb8d784e

Observation 76c53c38-ea5b-4444-a9d7-10b86e9e747f · outbound

This paper cites Ocra: Object-centric learning with 3d and tactile priors for human-to-robot action transfer.arXiv preprint arXiv:2603.14401, 2026.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Ocra: Object-centric learning with 3d and tactile priors for human-to-robot action transfer.arXiv preprint arXiv:2603.14401, 2026

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.067040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.067040Z digest=sha256:586655630ef2726b79fed4a900096fc35b7ded01fd2ba687b464b07bf8eafb17

Observation de32a317-a7c5-4105-beeb-3824e4cea282 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.070895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.070895Z digest=sha256:3f266349a90eafc3248ad9114b3b10d216b5780b45908c2d72b8e4eb3a7aec33

Observation af51e76d-04ab-457a-8371-9b9b83cba00c · outbound

This paper cites Videotree: Adaptive tree-based video representation for llm reasoning on long videos.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Videotree: Adaptive tree-based video representation for llm reasoning on long videos

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.568874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.075030Z digest=sha256:9a9d26725d171ea3819a26fe68f230c7a61d05571f0081a3ebe9bd186f0c6dc5

Observation a3563688-65bb-4c91-80b4-0552db3507bb · outbound

This paper cites Lighttrack: Finding lightweight neural net- works for object tracking via one-shot architecture search.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Lighttrack: Finding lightweight neural net- works for object tracking via one-shot architecture search

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.554970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.079136Z digest=sha256:0f0a6cf23a9b9476f36b6c33a9afe000243ab536b929b351c153b8763ade0731

Observation 171999ef-dca6-4d0e-9888-6e7024ac3186 · outbound

This paper cites Qwen2.5 Technical Report.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Qwen2.5 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.082999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.082999Z digest=sha256:ea288caff3236877f03575b9b17cb4b147566a5b253a86ae98232b531cf7d8df

Observation aa6c9143-a734-4197-9b9c-8265a34361f5 · outbound

This paper cites So- cratic models: Composing zero-shot multimodal reasoning with language.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning So- cratic models: Composing zero-shot multimodal reasoning with language

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.541247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.087046Z digest=sha256:8c8411ced8d2424bc54b2191fc187fc6da6fe28eb66ebc09edfea4e09263407d

Observation b627dae2-a079-487f-8340-cce3aa692acd · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.092018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.092018Z digest=sha256:31f31b0fe1e4cbb7af3f88fd9875e43c6f5817d884876d53df61fd8b1d61d6dc

Observation 9de21945-e071-4fa5-b778-410270d01b50 · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.526735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.096261Z digest=sha256:cba5924e6236837a45fca342131bf2a6f231978a3dcc14c0fba0b0b749efe0df

Observation 490c09d0-828f-4d01-a647-ed0d7877173e · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.100496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.100496Z digest=sha256:a15a39d087a70ab90fe5d6ef230a6eebd5e61b6850bd38159348ffc7ce013144

Observation 368d900c-88a4-42bd-bd2c-2196ed159593 · outbound

This paper cites Considering hardware limitations and ef- ficiency, we preprocess videos by sampling frames at 0.5 FPS on OpenEQA and EgoSchema, while limiting frame count to 32 for EgoTempo.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Considering hardware limitations and ef- ficiency, we preprocess videos by sampling frames at 0.5 FPS on OpenEQA and EgoSchema, while limiting frame count to 32 for EgoTempo

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.512053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.105067Z digest=sha256:57527445b57fb6964c28edf7c0b7509c806bb2f32d2183a3366988933d955d54

Observation cd6243c8-5e00-4e97-b4b1-c09ac06c76ec · outbound

This paper cites an unresolved cited work.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:11:15.498046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.109245Z digest=sha256:324f1db97b909a387fcdd69f4af6301a921a532c7db573f71d3d6e977fbf4661

Observation a500b113-a6fc-41ff-89eb-4a0f364785ed · outbound

This paper cites The results are presented in Table 6.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning The results are presented in Table 6

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.483608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.113320Z digest=sha256:ee74dd3906cc27f1a5f499bf72e9015272ace6dde0cc47372e85b5212f2db017

Observation 580dda91-a3db-45a6-89c0-5bb6bf94f7e0 · outbound

This paper cites an unresolved cited work.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Unresolved cited work

Reference 48

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T19:11:15.469738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.117778Z digest=sha256:d8c98aaed3ae1d8b9dd68a24a17b9e1149a26209a5b6039235f6175968ac7e76

Observation 7791ed57-e883-47e7-9ad0-03d8c3784990 · outbound

This paper cites We highlight relevant infor- mation throughout the reasoning trace, marking correct de- tails in green and erroneous ones in red.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning We highlight relevant infor- mation throughout the reasoning trace, marking correct de- tails in green and erroneous ones in red

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.455018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T19:11:15.122695Z digest=sha256:c636ab6f28c2676666c40332f1db37cd8e5b63ae9819e02d713e371ef8508b7f

Pith citing papers

Observation 8b986300-5312-4992-9628-5e010a6e1a45 · inbound

V$^{2}$-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence cites this paper.

V$^{2}$-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:02.909446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T04:08:32.413555Z digest=sha256:05f1d07645e5ec4b89acb90a647db3b5473e3732b812979b6193533328bb7205

Observation 7c1461e4-1bbd-49a4-854b-b25ce684d2b6 · inbound

ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos cites this paper.

ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:02.909446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T02:50:17.955287Z digest=sha256:77afa8c0419c643fba5d663ab32fcf7b98b4609ae9383aed8aa0e9fe51f08af1

Observation d7f51606-6c0e-42af-99f7-38c41c4575c6 · inbound

EgoSound: Benchmarking Sound Understanding in Egocentric Videos cites this paper.

EgoSound: Benchmarking Sound Understanding in Egocentric Videos CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:02.909446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T21:44:47.636912Z digest=sha256:a51a33905224607f62d2b96d6f7320b8168b7e4e35e77afb8aec50baa5f680ab

Observation e5bd53a0-5e08-49a5-8aaa-a277d9c73123 · inbound

The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering cites this paper.

The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:21:29.776719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:21:29.776719Z digest=sha256:ed597b304943d86e7ac9dfa45fcf1b4e464c37c376817ed0670c00925520d03a