Pith. sign in

Paper Citation Record · LEDGER

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models

As of 7 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2605.07148.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.07148 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T00:56:06.243890Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:23:00.372251Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-01T09:35:41.079576Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact28
  • verified fuzzy25
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08c3477f-89d2-4975-8896-0434b2e46ef6 · outbound

This paper cites Cognitive maps in rats and men.Psychological review, 55(4):189.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Cognitive maps in rats and men.Psychological review, 55(4):189

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.835669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:4981150914497151c74175f187aa0188968144737aa705fa1e2180764c7a3bab

Observation 0f1b2dc9-aa01-48ea-92ff-d63faf01c576 · outbound

This paper cites Oxford university press.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Oxford university press

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.864222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:be4191aa0ae8839b72cd7f65d647085e49f1ad0dde20e0baf39931443339ecc7

Observation f6200967-202b-452d-8c8f-86f4c8d91f81 · outbound

This paper cites Non-euclidean navigation.Journal of Experimental Biology, 222(Suppl_1):jeb187971.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Non-euclidean navigation.Journal of Experimental Biology, 222(Suppl_1):jeb187971

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.776402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:4ed031168794aad0d8d7f1ae0f7e1f04e3bafc90f5868c6da76eee3b7f973c7a

Observation cd19fa2b-720d-4c83-96cd-69a6048d47bc · outbound

This paper cites Structuring knowledge with cognitive maps and cognitive graphs.Trends in cognitive sciences, 25(1):37–54.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Structuring knowledge with cognitive maps and cognitive graphs.Trends in cognitive sciences, 25(1):37–54

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.750044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:66073cb61aef19bc68ca25e5d2d11ab49d8d67d8444934ac3da57692d710a441

Observation 27281b9e-500c-480a-bf0e-caf4a31cd9d7 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:00:55.062047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:4b611a6fcd20855503c936b555a25aa9ad595fa6a6b0f861e4a96d98b582484c

Observation 5155a600-edc0-47e2-8798-bfc8cfff8a9d · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Video-chatgpt: Towards detailed video understanding via large vision and language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.830776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:20f444d083565b9dae9a57e63271635c60946a8162cdb119d223f8ec43a4bb10

Observation ee7d748a-dd6c-4c90-a4d3-f1fe73c1ad7f · outbound

This paper cites GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models GSR-BENCH: A Benchmark for Grounded Spatial Reasoning Evaluation via Multimodal LLMs

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.057963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:8f54a96f8004914021552db43ccac3332b2858b7ec5f095acfe1eb5d233520a4

Observation ef9c6ca6-02f4-4ae1-ab5f-df39ffd73df5 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.089512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:634801f594aaf99b9dbf8fb6baf273174f5d4bc520b4efbdc6704b91ba898b7a

Observation 8e743d46-bcb1-403b-86ee-2961d9043cd8 · outbound

This paper cites Reasoning paths with reference objects elicit quantitative spatial reasoning in large vision-language models.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Reasoning paths with reference objects elicit quantitative spatial reasoning in large vision-language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.787266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:f77543d9f98de162c4b848e206684636ea85a5fde60fe14d584094a009b6b8d6

Observation 927ebeab-a5bc-4728-99de-eb2cbe62b1dc · outbound

This paper cites Spatial reasoning in multimodal large language models: A survey of tasks, benchmarks and methods.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Spatial reasoning in multimodal large language models: A survey of tasks, benchmarks and methods

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.110438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:5d7ab50b1893aa31415b5645ef01bddf09d7d9bfb8b34a545186cab9fcbba12e

Observation 6b8a87e7-0520-4788-87c1-62fa8f53e181 · outbound

This paper cites Infinibench: Infinite benchmarking for visual spatial reasoning with customizable scene complexity.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Infinibench: Infinite benchmarking for visual spatial reasoning with customizable scene complexity

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.177210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:00e99a97be7e352671adfb4730f207e305f7ea73751b0550d85e1a19f4ed9036

Observation 8e86367b-a397-4a40-bd3a-17dffc49096b · outbound

This paper cites Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Generic attention-model explainability for interpreting bi-modal and encoder-decoder transformers

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.855459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:a9dc619eb874a511a19756a03586e9154299eea89d3404e031f30f5b39cf1865

Observation 85e2888c-989e-427a-b271-75f040e6cd36 · outbound

This paper cites Why is spatial reasoning hard for vlms? an attention mechanism perspective on focus areas.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Why is spatial reasoning hard for vlms? an attention mechanism perspective on focus areas

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.093786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:e2191bc453875e7e586168bbb89597fa5a7953798f2f947079994abeda2d1fdb

Observation 614fd597-f426-4038-920b-02175952746f · outbound

This paper cites Beyond semantics: Rediscovering spatial awareness in vision-language models.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Beyond semantics: Rediscovering spatial awareness in vision-language models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.140796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:ef8484b37c6029be14329b0ad9f9f4fe5d625979b6f4c37e67557d8c2194b060

Observation c2cc8c22-7df3-492c-83e6-ca6d44185f43 · outbound

This paper cites The Geometry of Categorical and Hierarchical Concepts in Large Language Models.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models The Geometry of Categorical and Hierarchical Concepts in Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:54.996003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:6ad151dd5bfa0c4900ccc2b2344d47d967a3d5c02a77099af8f2fb601f0a4838

Observation c6a61006-fe26-4a2a-82bb-a79f507e37c3 · outbound

This paper cites Linear mechanisms for spa- tiotemporal reasoning in vision language models.arXiv preprint arXiv:2601.12626.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Linear mechanisms for spa- tiotemporal reasoning in vision language models.arXiv preprint arXiv:2601.12626

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.078944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:53899d7847f0166fb149b7f0787d87942a34dfa9904fddf0091cac1517bf5254

Observation 5d1eb770-c5a5-4726-b3d8-4ce8a463a431 · outbound

This paper cites Visual symbolic mechanisms: Emergent sym- bol processing in vision language models.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Visual symbolic mechanisms: Emergent sym- bol processing in vision language models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.074503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:b89e2153f3b4126bb5d040b6ad597865b3ad08d29e8b4bec160d7e14a44b3f3c

Observation 49f331ee-9b96-4d8e-a761-040cd6933770 · outbound

This paper cites Analyzing the behavior of visual question answering models.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Analyzing the behavior of visual question answering models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.821555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:de42ad6dcdc2ca0b1634e74c5f92cf47170c54192ba3d6a402211bf27aae9fb3

Observation 9bae1465-c803-42dc-adac-5b739f011eab · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.797458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:9c3d07b8c8068a1add063f43d034aa1e15718f69478793e8b04552abdd83948e

Observation 037a3f59-e3e9-40d4-9a48-0538a26a2384 · outbound

This paper cites Vision Transformers Need More Than Registers.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Vision Transformers Need More Than Registers

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:00:55.183613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:15ff20f4f4e3145fd6f66ec43b73cef25508b11f85c639ad3bfa6d4f301e79eb

Observation 5b3b7ce2-3770-4241-add7-677b6fecd296 · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.817349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:3f707279380bc669af5fb27cf77871e3125e1fc9883f84345e064f9b5f53c5f7

Observation a3c8bfe4-cf8e-4839-a798-c2f68178f098 · outbound

This paper cites Mindcube: Spatial mental modeling from limited views.arXiv e-prints, pages arXiv–2506.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Mindcube: Spatial mental modeling from limited views.arXiv e-prints, pages arXiv–2506

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.840123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:285a8ed7e75c484dd56b11ad940a1d332b44cb7444b7a58179bbfc0f6aeb0c14

Observation 87938ae6-6ca2-4747-b61f-b550c280af50 · outbound

This paper cites arXiv preprint arXiv:2602.07082 (2026),https://arxiv.org/ abs/2602.070824.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models arXiv preprint arXiv:2602.07082 (2026),https://arxiv.org/ abs/2602.070824

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:55.121495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:e9adb9d82d62333ae5e0a0555e5d9a82a680dfdc4aad61dab2036ec03a8f4ce9

Observation 181f7c00-1c9a-4297-9606-1a175e9b9ef2 · outbound

This paper cites Spatialrgpt: Grounded spatial reasoning in vision-language models.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Spatialrgpt: Grounded spatial reasoning in vision-language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.850247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:f35f4077809336c159276ca2b7a28288a5a118128760ede2d971ad239384e2ee

Observation 338621f3-c352-40ec-801e-fa934f5f5090 · outbound

This paper cites Scaling spatial intelligence with multimodal foundation models.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Scaling spatial intelligence with multimodal foundation models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.013102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:f1d4bebba42fb18ae827b5bc598756cedb404b2ff12dc2edbb477094284bc856

Observation ca4585f3-9bbd-455f-99f5-b88ee5681fbf · outbound

This paper cites SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-07T01:16:00.407423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:b8779434caca806a9ecd74725b28473d3538b499860006844b70b490ac78ab1e

Observation 2d33efec-352e-454e-86c7-bb00797f260e · outbound

This paper cites Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:58:10.438799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:2b2c82a93fd0dcd33986bddfd72483e26ae7c100a6ba1488aa218400e2ab3142

Observation 37da16b4-2c4e-4f13-9efc-754e4fa9f5c7 · outbound

This paper cites Causal abstractions of neural networks.Advances in neural information processing systems, 34:9574–9586.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Causal abstractions of neural networks.Advances in neural information processing systems, 34:9574–9586

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.812012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:943c91dccccaebae2a9188fee54d609cd5d7a298a76dfd6b55b9a3d45c145184

Observation b410ba3a-7c41-4aed-b402-320a88e5580e · outbound

This paper cites Investigating gender bias in language models using causal mediation analysis.Advances in neural information processing systems, 33:12388–12401.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Investigating gender bias in language models using causal mediation analysis.Advances in neural information processing systems, 33:12388–12401

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.771191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:38eb5ccca6b079bb26fc016751ef62fdbaaffa5f84f17f28467fd04cdee68f31

Observation b743b19c-20f2-4c89-b492-e9dba5bb284c · outbound

This paper cites Emergent linear representations in world models of self-supervised sequence models.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Emergent linear representations in world models of self-supervised sequence models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.760153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:2b3787e87a0f51c3814dbcd2e8c86a85e5a90c9cb127b7100b6c7d71a8ae66c0

Observation cfd72f98-76e0-4803-b261-18f66a79e08b · outbound

This paper cites Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Decomposing Representation Space into Interpretable Subspaces with Unsupervised Learning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:43:10.747895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:e71587ec3626e9e6c8a14b68e6fed5ed4c3b7a52b212c9edddd24d241c83fb16

Observation 0c5c4716-1a45-4ccb-b2ce-10b46640e1ac · outbound

This paper cites Does object binding naturally emerge in large pretrained vi- sion transformers?.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Does object binding naturally emerge in large pretrained vi- sion transformers?

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.021812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:f69db5bb50123495220f714eec638fd377aac00368954a6f7ee0f16679b065d9

Observation cbe57a86-3ae8-4a3f-a2d5-0c3ca7acd9a5 · outbound

This paper cites SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:00:55.152020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:0faff8f652b3a2251981061ed9603a7cafaab7e8a292005199cf6357b4e15e23

Observation 058cd24e-667c-4fd4-a396-ed5a7c8120bb · outbound

This paper cites Spatialvid: A large-scale video dataset with spatial annotations.arXiv preprint arXiv:2509.09676, 2025a.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Spatialvid: A large-scale video dataset with spatial annotations.arXiv preprint arXiv:2509.09676, 2025a

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.048555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:107aa00ad1329fa1510ef77bb821f026966ef5e3a584234de930809332a75434

Observation 4ef4ad3a-7120-4fda-a182-e177201cabc0 · outbound

This paper cites Qwen2.5-vl, January 2025.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Qwen2.5-vl, January 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.803547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:2d6fd12e49d7bbbd295e22ff7e661b6abbae7e6365afdfcd419fff891d4fcd4a

Observation 0bbd56e6-483b-4a62-8228-81bf73abc4eb · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:00:55.145785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:dd1761d46ba83ca8b0c2f1fd60c235ce334bbce1f5babd375a3e68c1363417c0

Observation 6493fa68-41b8-4a0b-a46c-d4e7f916c00e · outbound

This paper cites Linear Representations of Sentiment in Large Language Models.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Linear Representations of Sentiment in Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:43:06.705954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:10b104a857bf4f3f6b6161e53703c7a724008068b66a91e67b3d9524368ed3ed

Observation 7813ad13-0ad4-425f-bb5b-dd6db8a09fc0 · outbound

This paper cites Linear spaces of meanings: compositional structures in vision-language models.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Linear spaces of meanings: compositional structures in vision-language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.825709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:8225b36d6dcb215bc8dd20fb960df8151897e6e754af838ecbe5c4f0379a333b

Observation c4936031-72cf-45e6-bedf-49071b82a667 · outbound

This paper cites Deciphering personalization: Towards fine-grained explainability in natural language for personalized image generation models.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Deciphering personalization: Towards fine-grained explainability in natural language for personalized image generation models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.156748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:327e513dfec76451d45562643c51c4a246328a8236cc0727f40b488802a586b4

Observation 910488ed-2c93-4e9a-8e62-f61b379587e9 · outbound

This paper cites The Linear Representation Hypothesis and the Geometry of Large Language Models.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models The Linear Representation Hypothesis and the Geometry of Large Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:43:32.339354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:37d9f7f7655a7ca42c955aa6b8c08395a1d63918e4d2315854d72f72035ce013

Observation 9fe6844e-8149-443b-8446-9f62bf4d1938 · outbound

This paper cites On the Origins of Linear Representations in Large Language Models.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models On the Origins of Linear Representations in Large Language Models

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:55.069345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:8d7e02c8456937c0973e8543166018dc26fc5dbc4b20aa0c83d57add97245a8e

Observation cad0295a-cb29-49e2-b2e8-09af52fc21fc · outbound

This paper cites Line of Sight: On Linear Representations in VLLMs.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Line of Sight: On Linear Representations in VLLMs

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.043918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:9a5a0c4f9c97835a67ed6b2304ae6ddcbd0b2bb63d2bf13bfc0d6815d42eac2a

Observation 203ca3fe-6312-480a-8f52-b33d225eff26 · outbound

This paper cites Interpreting clip with sparse linear concept embeddings (splice).Advances in Neural Information Processing Systems, 37:84298–84328.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Interpreting clip with sparse linear concept embeddings (splice).Advances in Neural Information Processing Systems, 37:84298–84328

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.766403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:bab0b1fc580ad63f1c3941eba4cd8606907139ed6750f40f46f5c62622c5fc12

Observation 0d0fdd4e-6d1a-4a93-9455-6b94734e17c1 · outbound

This paper cites Linear Spatial World Models Emerge in Large Language Models.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Linear Spatial World Models Emerge in Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:54.983163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:da4d4de3f059fd7cf2084a5fbfee4c13cb90059769af0af3d06121ba5001a15f

Observation fc1baed4-e6aa-4312-a630-77e07bc6c1e4 · outbound

This paper cites Laplacian eigenmaps for dimensionality reduction and data representation.Neural computation, 15(6):1373–1396.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Laplacian eigenmaps for dimensionality reduction and data representation.Neural computation, 15(6):1373–1396

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.742602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:e12eeb0813d57753d4a9ff019c579ab2331cdaf4e3e0c4d10cbb0b27175a9251

Observation 689aa3bb-4778-4f13-9d7f-a1543783fed9 · outbound

This paper cites ICLR: In-Context Learning of Representations.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models ICLR: In-Context Learning of Representations

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.026979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:0f049a205fb1eb2d1517bcea2ef5960343292f7a75feab0c742d50173331c9de

Observation cabe0242-351e-46ee-8f59-8b5a6bfc806d · outbound

This paper cites Transformers as Unrolled Inference in Probabilistic Laplacian Eigenmaps: An Interpretation and Potential Improvements.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Transformers as Unrolled Inference in Probabilistic Laplacian Eigenmaps: An Interpretation and Potential Improvements

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.098204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:49644e5d7b3de0168516e5d697629b87a5b38705cbca03efc722844743ae1cfc

Observation 5ec14619-68aa-49f3-8e0b-34c367b7f2d9 · outbound

This paper cites Learning Continually by Spectral Regularization.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Learning Continually by Spectral Regularization

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:00:54.975635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:d9f6f9fed7d3d33365f70bd020093a5a6529591a8f94e98b9354d252c6e125c9

Observation 6d13ab07-fc8f-4fbc-8322-3c90687f59e2 · outbound

This paper cites Principal spectral regularization makes momentum surpass adam for llm training.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Principal spectral regularization makes momentum surpass adam for llm training

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.860201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:92ca843bcc1260cfc05f96bf3e754b25edea5d3d54338351b6a38c70bdc69af8

Observation 337f0e07-c135-47fd-a09f-8002e729bb5c · outbound

This paper cites Representation Learning on Graphs: Methods and Applications.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Representation Learning on Graphs: Methods and Applications

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.162225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:42dfb1fb0a2658834c020cc33f5bb3227279dcbee5456d5768d7174f6f10978c

Observation 94272ea8-8158-4e44-879e-c6267b7de011 · outbound

This paper cites Semi-Supervised Classification with Graph Convolutional Networks.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Semi-Supervised Classification with Graph Convolutional Networks

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-11T05:00:55.187460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:2fc2ffcc880e44a5f87ad5d1d2fb710b3c7c6df914383ddf82d1e1791fba606e

Observation 2d70da38-6d48-4dc0-9a45-ffe457edcad4 · outbound

This paper cites cognitive-map.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models cognitive-map

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.735211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:7053dc6a5dce1fa7fbf9816efb4e7fe46a0d1d37eca2c1ab3ddc2d7f0753cd8f

Observation 5c985aba-b092-4bf8-8ed1-a8b7c7065b12 · outbound

This paper cites an unresolved cited work.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-16T00:51:59.868005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:766b9242be5c09f7fc5daa5289a49cc7ac5e669bad337c9492499ac7f6747418

Observation 5b09b197-e4c4-48ec-a358-ca4d05ba1cbe · outbound

This paper cites an unresolved cited work.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-16T00:51:59.807668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:5b75003bf0eb095080b0a589f93a3ee838be7ad51054ff248fc47456d7c93719

Observation 65e336db-7681-45b7-82c0-613d46b187b9 · outbound

This paper cites Each row is rescaled to unit norm.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Each row is rescaled to unit norm

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.781489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:55cf54f707d65ff48efc4ce0e0aa2d97c2182de64547b5f73bd31cc22cecc151

Observation c8697e43-b4f4-4bc0-8f19-1de257aafae3 · outbound

This paper cites Drop columns with diagonal |Rkk| below 10−6 of the maximum (rank-deficient class directions).

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Drop columns with diagonal |Rkk| below 10−6 of the maximum (rank-deficient class directions)

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.791989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:281241d830a63864de3f78ebc57d99e70786fefb1b7aaa2cf464a03c6e5b9420

Observation 512bef4f-93ca-4709-b63c-60a3c64a58b3 · outbound

This paper cites Algorithm summary.The end-to-end per-step computation is: 1.Forward pass, capturingH ℓ ∈R B×T×d at each Dirichlet-target layer via forward hooks.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models Algorithm summary.The end-to-end per-step computation is: 1.Forward pass, capturingH ℓ ∈R B×T×d at each Dirichlet-target layer via forward hooks

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.754920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:7179d74d0f77316d55c550e68d8faf03c88ac9a9971d335a8cf298780eda49d3

Observation 512158b3-f242-451b-a783-c838eca33480 · outbound

This paper cites For the residMulti variant, steps 1–8 run independently at 5 layers and step 9 averages the resulting ratios.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models For the residMulti variant, steps 1–8 run independently at 5 layers and step 9 averages the resulting ratios

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T00:51:59.844562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:3575d562e54466bc38b3459ee3168e04d15ab5defda722b0df0dc889f49fb8dd

Pith citing papers

Observation b1f00412-5aeb-4dde-92e4-cf5cb283b46c · inbound

Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning cites this paper.

Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:35:41.081680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T06:23:00.372251Z digest=sha256:7fa3db9f181cf8e3f7cf73e7f33eac8a4df8a8d7f2bfc12bbc11babbe4f54ed5