Pith. sign in

Paper Citation Record · LEDGER

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning

As of 15 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 4 inbound Pith citation observations for arXiv:2506.17629.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17629 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:11:15.122695Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:21:29.776719Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T04:08:59.824731Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy26
  • unresolved22
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c88ea755-e6c0-420f-a647-b26e2366fe41 · outbound

This paper cites Remembr: Building and reasoning over long- horizon spatio-temporal memory for robot navigation.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Remembr: Building and reasoning over long- horizon spatio-temporal memory for robot navigation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.890596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:14.919634Z digest=sha256:4d4c643673800fd99bfaa9ca7706b2577f86efc9e213dee77f87a09c9190d3f3

Observation 0ddeae4d-d437-4f04-9da2-a723437a9cf8 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:14.924830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:14.924830Z digest=sha256:8d49a06afbec8f34fd52ae64ca8dc226304e64b18cd6f2425f7710beb73d0f3f

Observation 4f914b26-6b94-4f76-98d4-fcb94b281d5e · outbound

This paper cites Qwen2.5-VL Technical Report.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:14.929732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:14.929732Z digest=sha256:4ecba3423afbdfd48a79874810992c8b7082f73286d2613970e417b17c61af9c

Observation 9be3426a-da38-44c2-a187-b1000d024a2a · outbound

This paper cites Cafuser: Condition-aware multimodal fusion for robust semantic perception of driving scenes.IEEE Robotics and Automation Letters, 2025.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Cafuser: Condition-aware multimodal fusion for robust semantic perception of driving scenes.IEEE Robotics and Automation Letters, 2025

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.877183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:14.934080Z digest=sha256:68c51bef74e794453078f33f14231638cd0ba832edad9e0d240602e7fcfeed15

Observation d7e7d189-72dd-4cd7-a9b0-cc308bb2bd38 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:14.938127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:14.938127Z digest=sha256:a90c7e059c7d14c3ff1de36445ff182e67ac6f3bc7aa3b6f26b6acf4151de7dc

Observation ea286289-857c-4ef9-8025-4c5b2cf3e324 · outbound

This paper cites Embodied question answer- ing.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Embodied question answer- ing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.863511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:14.942516Z digest=sha256:c341b3ac7d5e1909f788749f8b571a90ed84e806c91c9de4d60bfb81a913fb64

Observation eceea4e6-d709-4315-abac-b2fb5d308170 · outbound

This paper cites an unresolved cited work.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:14.947269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:14.947269Z digest=sha256:15e55208d17b2966fd31b753587e6b328f8f22773d40ba841ff16ce011352d8d

Observation 3cf23b82-8498-474f-b127-b2271b84c125 · outbound

This paper cites Videoagent: A memory-augmented mul- timodal agent for video understanding.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Videoagent: A memory-augmented mul- timodal agent for video understanding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.840637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:14.951436Z digest=sha256:e6e5392a812f8d18659c7bbb14094301f8d111020836882bc57bbe0b014b518b

Observation 9b97ae9f-b38a-4e0e-8eae-176329a48208 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:14.955692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:14.955692Z digest=sha256:de77c76896d9ad325ce4a16c35bdda46209d0999c247101fcc3be33b721c9a53

Observation 4d2c6250-70f5-4be3-9c5a-cddea093f0e6 · outbound

This paper cites Objectrelator: Enabling cross-view object rela- tion understanding across ego-centric and exo-centric per- spectives.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Objectrelator: Enabling cross-view object rela- tion understanding across ego-centric and exo-centric per- spectives

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.825270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:14.960706Z digest=sha256:58e688854622e5c121085153795a2fca811202b2b7238c1445a05a0db0737e30

Observation 42d19641-6415-495a-93d9-097d6f6852ac · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Ego4d: Around the world in 3,000 hours of egocentric video

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.811481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:14.964752Z digest=sha256:6a8d0e721d6c1719f1f423362c3f42bc1067d6166c83f7e50257141e4477c7fc

Observation a33312fa-4b74-44fc-a458-c3804a03c990 · outbound

This paper cites A survey of deep learning techniques for autonomous driving.Journal of field robotics, 37(3):362– 386, 2020.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning A survey of deep learning techniques for autonomous driving.Journal of field robotics, 37(3):362– 386, 2020

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.797307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:14.968640Z digest=sha256:d261c2b6a4988d7ed3ec158fc6e9054ec8b63141eb2c6bd3e6d9b0202ff8111b

Observation b04b94a0-179f-4609-bd50-62918aa1362b · outbound

This paper cites Open vocabulary multi- label video classification.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Open vocabulary multi- label video classification

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.784150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:14.972620Z digest=sha256:3549081a4bb58373ed8d142368e58d0d1b06138729ebe7cd7dfb6b583fea3f09

Observation 1e1e9735-5ac2-4121-9370-5d275f21e18e · outbound

This paper cites Sequential multi-object grasping with one dexterous hand.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Sequential multi-object grasping with one dexterous hand

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.771258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:14.976838Z digest=sha256:6a41b3f49bcec806944b79bbbf781beac9cb188f46656da4cfbb682ee4ffa06b

Observation 2d2b7f61-358d-43a7-881e-f901f3eba15a · outbound

This paper cites ChatDB: Augmenting LLMs with Databases as Their Symbolic Memory.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning ChatDB: Augmenting LLMs with Databases as Their Symbolic Memory

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:14.980788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:14.980788Z digest=sha256:afcf1e17d7ec987c1ddb0736fa854175130866117497dc6ab1d80700928b92a8

Observation c4739f48-111e-47b9-a529-13c37b7e8077 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Vtimellm: Empower llm to grasp video moments

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:14.984953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:14.984953Z digest=sha256:93491fa87da963676ec0a0eea119ff18c3cfec093bac7fb4b9f5fa0410373e02

Observation 7c97ef85-f5ba-4fc2-926f-7ae99b48003c · outbound

This paper cites Egotaskqa: Understanding human tasks in egocentric videos.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Egotaskqa: Understanding human tasks in egocentric videos

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.749389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:14.988940Z digest=sha256:ec01a8ce597ed3b812fe70a77d57379fbd545587be1d85b04dfd647f6e3d0da9

Observation d6159f40-45fc-4137-b1c8-affd4ba9e33c · outbound

This paper cites Context-aware planning and environment-aware memory for instruction following em- bodied agents.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Context-aware planning and environment-aware memory for instruction following em- bodied agents

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.735632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:14.993172Z digest=sha256:17cd57eb32d132958f36441eac2fcd8f62521394726843ae3fb6d722e511c3b7

Observation ff7b1043-b31e-419f-8804-8e98638da4c8 · outbound

This paper cites DeepSeek-V3 Technical Report.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning DeepSeek-V3 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:14.997487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:14.997487Z digest=sha256:b8e70c754270d862730c8e5c88541cc3db231541166c1c0abd2989aeed8fa164

Observation 83a1f811-dee2-4958-b4fe-fe50e014fbc2 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.001732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.001732Z digest=sha256:a0217f874488197a7b7997c5233169d4c2e06287b95605d87f349a72f1b53ba0

Observation ed89a511-4efe-4444-bcd2-fbbc92b60f6a · outbound

This paper cites Aligning cyber space with physical world: A comprehensive survey on embodied ai.IEEE/ASME Transactions on Mechatronics, 2025.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Aligning cyber space with physical world: A comprehensive survey on embodied ai.IEEE/ASME Transactions on Mechatronics, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.710217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.005682Z digest=sha256:102d77ed77478fa233989a428fdedfe225481a291116a86d48a9db1c84b5ad65

Observation e3e2a5a0-9fc5-4b46-a50b-6f6492bce91e · outbound

This paper cites A Survey on Vision-Language-Action Models for Embodied AI.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning A Survey on Vision-Language-Action Models for Embodied AI

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.009577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.009577Z digest=sha256:62ec9e5e3a75a257ee92a80bac39d6325cb2d6ea6543983059b045ea46cab147

Observation 3378b811-4abc-4fe7-afda-5c2cd7f7406b · outbound

This paper cites Openeqa: Embodied question answering in the era of foun- dation models.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Openeqa: Embodied question answering in the era of foun- dation models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.695086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.013900Z digest=sha256:fbde6ad2026bd1f68a4057801c45f3d7add175b188d33994fa0d139174a5c4af

Observation 27352482-e06b-40f3-a7bc-1ade7ebc192c · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.Advances in Neural In- formation Processing Systems, 36:46212–46244, 2023.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Egoschema: A diagnostic benchmark for very long- form video language understanding.Advances in Neural In- formation Processing Systems, 36:46212–46244, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.681595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.017972Z digest=sha256:1f28dd43338366723702f6c8dd30afb8f1dc79541da922f459dc1e352867a7e6

Observation c3d13be7-5da6-4c23-9a6a-337731811b1d · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning MemGPT: Towards LLMs as Operating Systems

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.021957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.021957Z digest=sha256:ca76fca0db1282e221aa0236137f4a4b06b72c66027621d4cf3dd6d354d23091

Observation bfb242e3-a4ea-457c-9d14-d078fbc2945b · outbound

This paper cites V2-sam: Mar- rying sam2 with multi-prompt experts for cross-view object correspondence.CVPR, 2026.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning V2-sam: Mar- rying sam2 with multi-prompt experts for cross-view object correspondence.CVPR, 2026

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.667593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.026221Z digest=sha256:c56d5db98c79a151aa4b4c0b7a21d751a250332ff5452adbff707c9fcbff6aeb

Observation 7abbdd78-a7b5-4447-a4f6-a9625c299c44 · outbound

This paper cites an unresolved cited work.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:11:15.652598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.030352Z digest=sha256:3f5245782d66696c94d1fcb7882ce5e4017a8de3e2650711cbe1ec321272883f

Observation f1d8620f-2305-4dcb-9292-d78a4b4f7b79 · outbound

This paper cites Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Position: Episodic Memory is the Missing Piece for Long-Term LLM Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.034372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.034372Z digest=sha256:a5d8b57bccf1dcdfb540ce6a40d0387e64245b0c1f993d49c3c77e29e2e9eb4c

Observation 77d510d2-e44b-40d6-9ba0-57e218681358 · outbound

This paper cites Omnia de egotempo: Benchmarking temporal understanding of multi-modal llms in egocentric videos.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Omnia de egotempo: Benchmarking temporal understanding of multi-modal llms in egocentric videos

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.639266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.038662Z digest=sha256:13d203e370b62099c55da930bb5d02e4a53bb04b71dc33369960f769e248c9ba

Observation 37a66466-2bb3-4eeb-916d-dcd03f3ee448 · outbound

This paper cites Improving language understanding by gen- erative pre-training.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Improving language understanding by gen- erative pre-training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.042504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.042504Z digest=sha256:131bb6b322fb22a374d6cf5f0c96302e860f12ee7076729a2078b2dc49d4087d

Observation 6f8b06ee-aa50-4191-86e9-99aa0e57a90f · outbound

This paper cites Character-llm: A trainable agent for role-playing.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Character-llm: A trainable agent for role-playing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.616461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.046328Z digest=sha256:feefbcbc12b66093e3e65db91c20f2f6527d29a2599962b328d16cdee8b03a55

Observation e58510fb-2bab-47e2-868e-a3cc19b0bdd9 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Moviechat: From dense token to sparse memory for long video understanding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.050586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.050586Z digest=sha256:979ad83f8290db6e80802b470255ebfb9ef2ec819510a787377b91645e969ce3

Observation 4fe840ef-5720-4867-90a8-8014fa7b8bac · outbound

This paper cites Part, Ioannis Papaioannou, Arash Eshghi, Ioannis Konstas, and Oliver Lemon.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Part, Ioannis Papaioannou, Arash Eshghi, Ioannis Konstas, and Oliver Lemon

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.594670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.054683Z digest=sha256:d53f1b552b8e5e65931200e80ff502a4640711bfb9ca7e429129a7c4a44d46d9

Observation 8b145558-78f5-46ec-a715-a83df2731935 · outbound

This paper cites Mac: a meta-learning approach for fea- ture learning and recombination.Pattern Analysis and Ap- plications, 27(2):63, 2024.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Mac: a meta-learning approach for fea- ture learning and recombination.Pattern Analysis and Ap- plications, 27(2):63, 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.581786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.058811Z digest=sha256:fcdddf8b257f60ad5ae26f68edce4ac11ed1ed5d24e02c44f922d5b0442ab74f

Observation a6def184-c419-49f3-a854-58a7b1b38159 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning LLaMA: Open and Efficient Foundation Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.062766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.062766Z digest=sha256:9d9969a848a1571a708d7b51b1515745c991133fc1dfec86c7ec077adb8d784e

Observation 76c53c38-ea5b-4444-a9d7-10b86e9e747f · outbound

This paper cites Ocra: Object-centric learning with 3d and tactile priors for human-to-robot action transfer.arXiv preprint arXiv:2603.14401, 2026.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Ocra: Object-centric learning with 3d and tactile priors for human-to-robot action transfer.arXiv preprint arXiv:2603.14401, 2026

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.067040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.067040Z digest=sha256:586655630ef2726b79fed4a900096fc35b7ded01fd2ba687b464b07bf8eafb17

Observation de32a317-a7c5-4105-beeb-3824e4cea282 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.070895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.070895Z digest=sha256:3f266349a90eafc3248ad9114b3b10d216b5780b45908c2d72b8e4eb3a7aec33

Observation af51e76d-04ab-457a-8371-9b9b83cba00c · outbound

This paper cites Videotree: Adaptive tree-based video representation for llm reasoning on long videos.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Videotree: Adaptive tree-based video representation for llm reasoning on long videos

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.568874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.075030Z digest=sha256:1b87e4ffb85b966f38351f98ea32bd4cdbabcef438aca32d0d5f86f50fa9a063

Observation a3563688-65bb-4c91-80b4-0552db3507bb · outbound

This paper cites Lighttrack: Finding lightweight neural net- works for object tracking via one-shot architecture search.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Lighttrack: Finding lightweight neural net- works for object tracking via one-shot architecture search

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.554970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.079136Z digest=sha256:4944a9a7b94b452c85451a6f6b09b098e44d55319e621a9db9ddaeae5ef117a0

Observation 171999ef-dca6-4d0e-9888-6e7024ac3186 · outbound

This paper cites Qwen2.5 Technical Report.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Qwen2.5 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.082999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.082999Z digest=sha256:ea288caff3236877f03575b9b17cb4b147566a5b253a86ae98232b531cf7d8df

Observation aa6c9143-a734-4197-9b9c-8265a34361f5 · outbound

This paper cites So- cratic models: Composing zero-shot multimodal reasoning with language.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning So- cratic models: Composing zero-shot multimodal reasoning with language

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.541247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.087046Z digest=sha256:b79b77a7571e7247f22261efda45fb71379eb4cfcdf5f070e2b08177f9bfea79

Observation b627dae2-a079-487f-8340-cce3aa692acd · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.092018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.092018Z digest=sha256:31f31b0fe1e4cbb7af3f88fd9875e43c6f5817d884876d53df61fd8b1d61d6dc

Observation 9de21945-e071-4fa5-b778-410270d01b50 · outbound

This paper cites Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Video-LLaMA: An instruction-tuned audio-visual language model for video un- derstanding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.526735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.096261Z digest=sha256:7a4c67760d2ea282ca0d697c2b13bb8abee5961eb9f9d85a0d43ac61d73b8672

Observation 490c09d0-828f-4d01-a647-ed0d7877173e · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T19:11:15.100496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:11:15.100496Z digest=sha256:b0eb549c4d95e650e25e77019cb947d7393cb04a7092cc05fa935ad697b0a690

Observation 368d900c-88a4-42bd-bd2c-2196ed159593 · outbound

This paper cites Considering hardware limitations and ef- ficiency, we preprocess videos by sampling frames at 0.5 FPS on OpenEQA and EgoSchema, while limiting frame count to 32 for EgoTempo.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Considering hardware limitations and ef- ficiency, we preprocess videos by sampling frames at 0.5 FPS on OpenEQA and EgoSchema, while limiting frame count to 32 for EgoTempo

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.512053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.105067Z digest=sha256:de7e68c8ea750aea28db8532454cc150a29fbd33f6e4862596f18701518376a4

Observation cd6243c8-5e00-4e97-b4b1-c09ac06c76ec · outbound

This paper cites an unresolved cited work.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:11:15.498046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.109245Z digest=sha256:20902eed33783903f587083a2e9d1ab2a22af998a6de86be018cd51762297541

Observation a500b113-a6fc-41ff-89eb-4a0f364785ed · outbound

This paper cites The results are presented in Table 6.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning The results are presented in Table 6

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.483608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.113320Z digest=sha256:c55a6d97339bda63de301666e7c21b36612cc1bc5e2a30ba39d15d8e51509f04

Observation 580dda91-a3db-45a6-89c0-5bb6bf94f7e0 · outbound

This paper cites an unresolved cited work.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning Unresolved cited work

Reference 48

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T19:11:15.469738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.117778Z digest=sha256:9d08d5dae6bb08c993eaebc19238399e215f0ce23896c96e983e3caad8caedc9

Observation 7791ed57-e883-47e7-9ad0-03d8c3784990 · outbound

This paper cites We highlight relevant infor- mation throughout the reasoning trace, marking correct de- tails in green and erroneous ones in red.

CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning We highlight relevant infor- mation throughout the reasoning trace, marking correct de- tails in green and erroneous ones in red

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:11:15.455018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T19:11:15.122695Z digest=sha256:094ee13336dab38347322cd986f2e19bf1daf468da09d2b6db5dd3b09c76997a

Pith citing papers

Observation 8b986300-5312-4992-9628-5e010a6e1a45 · inbound

V$^{2}$-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence cites this paper.

V$^{2}$-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:02.909446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T04:08:32.413555Z digest=sha256:723203d57e3eb583f9c988955f8ccd23dad811fc543625a6b9b65c85921a4b57

Observation 7c1461e4-1bbd-49a4-854b-b25ce684d2b6 · inbound

ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos cites this paper.

ToG-Bench: Task-Oriented Spatio-Temporal Grounding in Egocentric Videos CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:02.909446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T02:50:17.955287Z digest=sha256:1cdc287302f80c27dab30c79b3d0fde1b6ae063edb7fd85ed1a110f7ebd1dbf8

Observation d7f51606-6c0e-42af-99f7-38c41c4575c6 · inbound

EgoSound: Benchmarking Sound Understanding in Egocentric Videos cites this paper.

EgoSound: Benchmarking Sound Understanding in Egocentric Videos CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:02.909446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T21:44:47.636912Z digest=sha256:39287c512919541477cc7acf48e4f2a407e307655484cb93b3a0c5154c1a4043

Observation e5bd53a0-5e08-49a5-8aaa-a277d9c73123 · inbound

The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering cites this paper.

The First EgoCross Challenge at EgoVis 2026: Cross-Domain Egocentric Video Question Answering CLiViS: Unleashing Cognitive Map through Linguistic-Visual Synergy for Embodied Visual Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:21:29.776719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:21:29.776719Z digest=sha256:21fcf34148ff0a6857bd57f08427df4eb3859dbb8163566cc61cc671c8430bc7