Pith. sign in

Paper Citation Record · LEDGER

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models

As of 13 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 0 inbound Pith citation observations for arXiv:2608.10864.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10864 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:33:29.808181Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact2
  • verified fuzzy14
  • unresolved65
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 54a1a48e-541c-4a26-99a9-ac63c5153957 · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.118241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.118241Z digest=sha256:c29be05c75d2e6367943f43408cf9a93f661da14b7411f9cc8046c43666371b5

Observation 1f84e78f-dfcc-4e6e-8447-694e34bbad57 · outbound

This paper cites 3drs: Mllms need 3d-aware representation supervision for scene understanding.Advances in Neural Information Processing Systems, 38: 67961–67988, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models 3drs: Mllms need 3d-aware representation supervision for scene understanding.Advances in Neural Information Processing Systems, 38: 67961–67988, 2026

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.168523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.168523Z digest=sha256:f2afe520724b7de6162940fc74101cf1a4ac29cddb9cd57a0dbbb6fc03e11012

Observation 72786601-5a8c-4972-bbd3-73db14d1f467 · outbound

This paper cites CUP Archive, 1967.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models CUP Archive, 1967

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.208043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.208043Z digest=sha256:ff7625902026da667f0d5bb4fb3f9afc36f4906075b729a1ceb2f3873be096de

Observation bb9cd8f1-1932-4808-9c04-adda660b79d2 · outbound

This paper cites Henry Holt and Co., Inc., New York, NY , USA, 1982.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Henry Holt and Co., Inc., New York, NY , USA, 1982

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:32.032179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:27.258250Z digest=sha256:ffb4fcf67fecaefca078447e5aa704235561e11f21dd7f37becfc316de2e2801

Observation 8b75e6f2-0426-4b6a-847d-ff22eaf69d19 · outbound

This paper cites Number 6.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Number 6

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.296641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.296641Z digest=sha256:96033ead0d23bed00fcd8b4938ba0dd574413d6f88ce42a01ac8bce1c80c73bd

Observation 9c1485a0-d1ed-4f5d-898d-49463389cea6 · outbound

This paper cites Separate visual pathways for perception and action.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Separate visual pathways for perception and action

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.334241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.334241Z digest=sha256:713793f61b98d6c4df0eae90efc7162ec97b0410accda148647e1d2539b3573e

Observation c70e74d4-ee1f-4bb0-a477-e732b8540558 · outbound

This paper cites Mental rotation of three-dimensional objects.Science, 171(3972):701–703, 1971.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Mental rotation of three-dimensional objects.Science, 171(3972):701–703, 1971

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.374416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.374416Z digest=sha256:af3d1023b168433afba7837218e9be3a254974ad318e2ebebd620b18aba2626a

Observation 53fa584c-0db6-46e2-a28b-fd03061b08d9 · outbound

This paper cites Oxford university press, 1978.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Oxford university press, 1978

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.403655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.403655Z digest=sha256:6d9e41d70a995646b39ac46357734c4f3c4080ff7b864c2bd84dca43fa4e3554

Observation 3d47e3f0-36cc-4b90-b178-965802fae45f · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.434416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.434416Z digest=sha256:4deec22f32e70708c2422edd525227fc85e2334fda98267362d51e805488e718

Observation 04ee055b-8d94-4616-9a27-fa3eac91f6a3 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.467684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.467684Z digest=sha256:55d12ca421cf774e229a244bcf7861a3cbe11506bfe68f800df70e71d2a0f3b0

Observation c274fcc8-ac9b-4312-b3d0-4d118e2a2623 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.488894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.488894Z digest=sha256:627b7625d98e1f4cdaf5dd908e09506cb66573eba9bcb7ca1d61c365e13b7a30

Observation f4020c15-82c7-4b85-b0ef-3ff96bf5ab12 · outbound

This paper cites Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.513243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.513243Z digest=sha256:83f5e00023853a54655e862bfc1b376b60fa940fa7677742cbb518c098950b98

Observation ecd75291-ac67-40ae-9e16-39afc05ac79f · outbound

This paper cites Are vlms ready for autonomous driving? an empirical study from the reliability, data and metric perspectives.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Are vlms ready for autonomous driving? an empirical study from the reliability, data and metric perspectives

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.536117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.536117Z digest=sha256:1f0d8cb96d1fd545b11a80fc1ab1d81a4e08b5abb80c75a60576a4d81a93e622

Observation 0f5fb8ed-eb7f-4d78-a561-d48b5f067980 · outbound

This paper cites Qwen2.5-VL Technical Report.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Qwen2.5-VL Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.559154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.559154Z digest=sha256:5c0ba54e998e587bd86d28366bd13e781919bcd1a5be8e14435b39e04012833a

Observation 205a6793-af32-450d-9f26-0bdfd8316df2 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.597794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.597794Z digest=sha256:2aec58032593d389cef95f3e7335c05192e12973582d953fff76051c00437cd7

Observation 61b7547e-1d9b-4a8b-9164-36bb4225850c · outbound

This paper cites LLaV A- video: Video instruction tuning with synthetic data.Transactions on Machine Learning Research, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models LLaV A- video: Video instruction tuning with synthetic data.Transactions on Machine Learning Research, 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.861793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:27.627853Z digest=sha256:cf84a57c06b2d9db9345c82478f3b4b1a78dd063e906b4f4ed624fcb735d358a

Observation 7bb53d38-73a0-4035-a903-7194cffba8ed · outbound

This paper cites Probing the 3d awareness of visual foundation models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Probing the 3d awareness of visual foundation models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.851699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:27.662301Z digest=sha256:5ab9dd5249d2752195388e9e8f2e8d4f25966bf9496dc86ff308eb96aa905208

Observation 5cf0e01c-42fd-495e-a90d-a30d3d2c7c98 · outbound

This paper cites Sti- bench: Are mllms ready for precise spatial-temporal world understanding? InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 5622–5632, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Sti- bench: Are mllms ready for precise spatial-temporal world understanding? InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 5622–5632, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.693368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.693368Z digest=sha256:7eb68fc1913cdeac78136d9e06049b2182d018d17e985b55a3ecbf3b256dd17c

Observation 87d6ce1f-049d-4151-902c-6f4b92e319bd · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.750364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.750364Z digest=sha256:950c4156ea1ed8e7bda94eafe6d4116a1fc2f883bd79cc87b3879b484c3708c6

Observation ccddb12c-aba2-4058-990e-f3875f62cebe · outbound

This paper cites Vlm4d: Towards spatiotem- poral awareness in vision language models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Vlm4d: Towards spatiotem- poral awareness in vision language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.802712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.802712Z digest=sha256:a2ef7ee53104310cabf997625a5900e3bc842f4d67ee35af2fb7be6344e1371d

Observation 4a06d401-c430-4d9c-bbd9-3f0656262306 · outbound

This paper cites Cambrian-s: Towards spatial supersensing in video.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambrian-s: Towards spatial supersensing in video

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.915456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.915456Z digest=sha256:095c56fe9f79d87f7caaafb24136eeb62f51785a33e0ab9a2bc5037b24e62704

Observation beac19f4-b222-40b6-9e7e-69cf97e0fcbe · outbound

This paper cites Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.973683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.973683Z digest=sha256:bce0ab3202ae50dd315b7cc29598af52c7832303fe05341837377e8adcc996c7

Observation 4647a97b-0e1f-4c31-bcbd-a85f98f5f94e · outbound

This paper cites Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.004225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.004225Z digest=sha256:9854530130894b70019914f11c6ad17c850117befcbdacdc7b719ab2155bb29e

Observation f434eff9-3b11-4021-8934-ecd778e162be · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Distilling the Knowledge in a Neural Network

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.057213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.057213Z digest=sha256:dc1149eb033fd6fbb5c541c8abcba5eb0552e88aa9c1b1cb6744e1b9677e01e5

Observation c976a4dc-f4a9-40f3-b939-dd1c9a60943b · outbound

This paper cites Unsupervised natural experience rapidly alters invariant object representation in visual cortex.science, 321(5895):1502–1507, 2008.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unsupervised natural experience rapidly alters invariant object representation in visual cortex.science, 321(5895):1502–1507, 2008

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.803965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:28.107936Z digest=sha256:45e43180dc61a2c91b397c8bde8e203eca5b544bcadcd05e482cdb28b75c42ee

Observation a77fa326-3b6b-490e-bc6b-4cf2003d5c93 · outbound

This paper cites Slow feature analysis: Unsupervised learning of invariances.Neural computation, 14(4):715–770, 2002.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Slow feature analysis: Unsupervised learning of invariances.Neural computation, 14(4):715–770, 2002

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.688307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:28.143957Z digest=sha256:ae38d6751bb62acc7bbf1be2a5d8957777075c233e21f35dc0d1708cae1a296e

Observation f1a1f640-c5ea-4332-b960-e17a12b4f2f1 · outbound

This paper cites The tolman-eichenbaum machine: unifying space and relational memory through generalization in the hippocampal formation.Cell, 183(5):1249– 1263, 2020.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models The tolman-eichenbaum machine: unifying space and relational memory through generalization in the hippocampal formation.Cell, 183(5):1249– 1263, 2020

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.182390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.182390Z digest=sha256:465fdd556b72dcd89acc19ffcab2ad130e5f67b14970b5ea6e36a66f47fda200

Observation 0b40abdd-a5db-4460-8a8c-2ef393025aae · outbound

This paper cites Psychology of spatial cognition.Wiley Interdisciplinary Reviews: Cognitive Science, 3(6):565–580, 2012.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Psychology of spatial cognition.Wiley Interdisciplinary Reviews: Cognitive Science, 3(6):565–580, 2012

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.556519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:28.212915Z digest=sha256:3aa9c7db4a73b279ce5848cd06ec7d76313297f486e8aec4cdc3fd5a4d067031

Observation a6faf50b-53c7-4948-b050-b6ca50e69b8f · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.247419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.247419Z digest=sha256:3a0d5e3c71fda977e47e86af4f5682db35cd07eb3530953f7918631f4ac6c1b5

Observation 67da9146-8180-461b-920c-a09a20989136 · outbound

This paper cites GPT-4o System Card.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models GPT-4o System Card

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.289469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.289469Z digest=sha256:09ae43177370f9c9ed26f3202728a531177c137ad34068d355cd67eda3da82b8

Observation 7c938b5e-59a1-4a40-bc81-2a4b0f23cbd7 · outbound

This paper cites STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.317477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.317477Z digest=sha256:2758798e46aff60b7c45a5cc8feeba6922bc594892b41a8b4bf2a87b2aafaa16

Observation d9ce87d5-a09d-4627-9a8c-8812128504eb · outbound

This paper cites VLM4D: Towards Spatiotemporal Awareness in Vision Language Models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.342219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.342219Z digest=sha256:07c6265e92f1b97a984b559e3d27ab952b4a88544917097b0ffeca892e1e5421

Observation 6ccf82cd-9f78-4f08-9861-33308bddca74 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Learning Transferable Visual Models From Natural Language Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.364382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.364382Z digest=sha256:4a2e885f33504995e372cded03bb8e206018d51e4cd06744fcc6e5bd4f2352b8

Observation 1c083e24-1267-4b2f-94ec-a3fe365ffef1 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Sigmoid Loss for Language Image Pre-Training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.381618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.381618Z digest=sha256:05a162d7e7f9379a3de5bb5902108c64a819f0d3c3b759792fb26ad49bdffca7

Observation 8e3a871a-0a6a-473f-a70b-2be1a1d4fa0f · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.399065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.399065Z digest=sha256:6682551065c565c3b663e52d6752553b4e071e78efe23fe3a71e1d484d795d19

Observation 96ab4e6c-53d2-426a-abc4-f0e554584a1f · outbound

This paper cites O’Sullivan, Fang Cao, Tahoura Nedaee, Kamyar Rajabalifardi, Fei-Fei Li, Ehsan Adeli, and Euan Ashley.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models O’Sullivan, Fang Cao, Tahoura Nedaee, Kamyar Rajabalifardi, Fei-Fei Li, Ehsan Adeli, and Euan Ashley

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.460314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.460314Z digest=sha256:b127247543a58367be22f04fb4a37a785fbe68008aec4c3c6efef5f07dc5b67e

Observation f65e9dfb-4032-429b-9eef-45796e4626c8 · outbound

This paper cites From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2026

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.508799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.508799Z digest=sha256:c983b825cdd89800a986de662e7be397a051e3b71a3e50653bd6ad9fad6f792e

Observation ebcc31ef-8308-4671-a011-7d8922eaf6e2 · outbound

This paper cites Visual spatial tuning.arXiv preprint arXiv:2511.05491, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Visual spatial tuning.arXiv preprint arXiv:2511.05491, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.562622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.562622Z digest=sha256:905d1138823514c7796cddf60358b4a17281ccc9006467f16577b3ea5c120c2f

Observation 06b1887d-f586-4ca0-8901-29e9cdbf6dd2 · outbound

This paper cites Cambrian-S: Towards Spatial Supersensing in Video.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambrian-S: Towards Spatial Supersensing in Video

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.585047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.585047Z digest=sha256:06c43e8d6076c411f7514461d0673923379be3b822255a7c5b12c0da293fe82c

Observation da5f7c1d-cf31-4b45-84aa-571d67dcbbf6 · outbound

This paper cites Vggt: Visual geometry grounded transformer.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Vggt: Visual geometry grounded transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.614240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.614240Z digest=sha256:bf149f29c080efc596e2e02a40d3cc3ce5e25e8e8f413810275c786a68f0e1c1

Observation 10bccea0-33e6-4810-9190-2e9cd6e583d6 · outbound

This paper cites Continuous 3d perception model with persistent state.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Continuous 3d perception model with persistent state

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.541921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:28.641131Z digest=sha256:e4cb5fcdc35b373908038e8922fa697937f93821c16effd6ec33085c9102f29f

Observation 084fcc2b-0416-453a-88f8-c2da350228e7 · outbound

This paper cites SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.651082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.651082Z digest=sha256:287a45c512b24b7e0f845b2400bea6227d0e8bb867e50dd210b4bea368985d77

Observation bb3b1a12-f7d2-46a3-a3ef-2d507a968d3b · outbound

This paper cites G2vlm: Geometry grounded vision language model with unified 3d reconstruction and spatial reasoning.arXiv preprint arXiv:2511.21688, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models G2vlm: Geometry grounded vision language model with unified 3d reconstruction and spatial reasoning.arXiv preprint arXiv:2511.21688, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.682876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.682876Z digest=sha256:84de5464e1e0f671e021ab06fe109754f9b19af6c7203f2efbde933d1e991171

Observation a6476444-98df-496a-8dec-752400d760ba · outbound

This paper cites Think with 3d: Geometric imagination grounded spatial reasoning from limited views.arXiv preprint arXiv:2510.18632, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Think with 3d: Geometric imagination grounded spatial reasoning from limited views.arXiv preprint arXiv:2510.18632, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.731186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.731186Z digest=sha256:bb2f3186250cbba15a357a37ed093f31fabe2985701cc7ff6c371515db5ac733

Observation a7c3b56b-2384-4993-af06-ca9fde4930fb · outbound

This paper cites Vision-aligned Latent Reasoning for Multi-modal Large Language Model.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Vision-aligned Latent Reasoning for Multi-modal Large Language Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.783675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.783675Z digest=sha256:737c865e771328ce171e3e8d649c9d0ee2c7ea70f338d9a6f1b38ddcb512c392

Observation 420ee94e-dab1-4a97-b330-7b95dfc9d6f4 · outbound

This paper cites Cambridge University Press, 2001.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambridge University Press, 2001

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.508886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:28.816106Z digest=sha256:b330535ca865a0c00a6eeb4d4a25a18f74f2580758481a59f1f084f3802d3c64

Observation a1e511ba-3ad4-4ca0-be74-f96a6923cdf0 · outbound

This paper cites Cambridge University Press, 2014.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambridge University Press, 2014

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.837244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.837244Z digest=sha256:6cc7bc4c88ca7440e81abf761674ab96a2ba5ceb7b543618073aa5a7b6ce9be0

Observation 58daaab1-74dc-4075-be49-655b718d8179 · outbound

This paper cites Relational knowledge distillation.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Relational knowledge distillation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.847346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.847346Z digest=sha256:ae52b85f98a8f8f09e68d0797ff25dd61c06e81e87640efc4ae1d6a1395b4e96

Observation c5c2a94d-f3f4-41d6-824f-99fc63b9abef · outbound

This paper cites DINOv3.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models DINOv3

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.881926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.881926Z digest=sha256:0f8744958b74751ce8b0f174f97f8967bdced0e9bfdb33fb2d43ad98dd9fb432

Observation ceadf3e1-0473-4f8c-8652-8cbde385cbba · outbound

This paper cites Perception Encoder: The best visual embeddings are not at the output of the network.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Perception Encoder: The best visual embeddings are not at the output of the network

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.893432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.893432Z digest=sha256:2e4fb09ad3a0ba3b96536ce38caf3477c357d5af14a1328dc1423c37f77c0bb1

Observation b7f2d45a-2f0f-4047-ba08-e8d01fcba513 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.442289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:28.897760Z digest=sha256:6a942e1f1db74a7c71c0f112621a1411a331ba8c25668912a67431580c9bef44

Observation 7d9ab556-828c-43a9-95f8-2f452d8dee24 · outbound

This paper cites Videorepa: Learning physics for video generation through relational alignment with foundation models.Advances in Neural Information Processing Systems, 38:122647– 122676, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Videorepa: Learning physics for video generation through relational alignment with foundation models.Advances in Neural Information Processing Systems, 38:122647– 122676, 2026

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.383300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:28.927333Z digest=sha256:89039e0a85b62b88de623e8734ae7dd25daf49fb58db083e0c8db34d57451570

Observation ae99a4e1-13d8-4c11-97b7-8dc38ff6acaa · outbound

This paper cites Moalign: Motion-centric representation alignment for video diffusion models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Moalign: Motion-centric representation alignment for video diffusion models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.272471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:28.983429Z digest=sha256:7af97220452de07608e06b7b946a794eafc3079695181c39bc16990bdb9168a9

Observation 34256c33-f0dd-4f2e-83d0-80391481943f · outbound

This paper cites Lever- aging vision-language models for improving domain generalization in image classification.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Lever- aging vision-language models for improving domain generalization in image classification

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.153761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:29.047125Z digest=sha256:fdd4af0b964b6ebed3586bc976c5278ca45c8670a610d77e3ff0dc6309e359f1

Observation b33995ff-cbda-4283-84df-d2ae3fc1b135 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.100129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.100129Z digest=sha256:576a3b54fbb1deaa2c390ad3ee51322b2348a3d306c24671675363c24c18ffd5

Observation be2c11c2-5218-40a3-9139-f67c19d73469 · outbound

This paper cites ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.148328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.148328Z digest=sha256:81dd827ad7dd1e4efa87eef6421e948f6e96c72e489190d1ecf88b3420c16873

Observation b9bca34a-8a20-46eb-b56e-c75c6970ab98 · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.171413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.171413Z digest=sha256:58d2c784af5da121f5fbf61b351cfec32fdc85376f76feb35246b363f5ea26b8

Observation 8a819d40-0fd1-4343-a531-38dee812edd0 · outbound

This paper cites Scaling spatial intelligence with multimodal foundation models.arXiv preprint arXiv:2511.13719, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Scaling spatial intelligence with multimodal foundation models.arXiv preprint arXiv:2511.13719, 2026

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.175284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.175284Z digest=sha256:72dec091c76b8a9b3353231a1639636f55fca71c7df5044103648fe5b5511b64

Observation f08d0c3c-1a81-4e3a-8b09-b52449f00a60 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.181722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.181722Z digest=sha256:94c5aa882105985c5221e05ff87b6b1fc9e7b8c8a710a93e85987ea8b5cb2f13

Observation 7d3efa8d-16ba-4abc-a9bd-0e210dd59505 · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Grounded 3D-LLM with Referent Tokens

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.184264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.184264Z digest=sha256:0f392582e16745a3bdd5ddf6eb67e24f54f9b5c5046bef72a14a2bc01cb8bbcc

Observation d48dd1ca-9762-44c5-a86c-d64e53f01c20 · outbound

This paper cites Unifying 3D Vision-Language Understanding via Promptable Queries.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unifying 3D Vision-Language Understanding via Promptable Queries

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.188680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.188680Z digest=sha256:f6963b04093eafa6c92824fee46047a65eb3ab0bf46d9f1e18857ae674468678

Observation 80814b0b-2ab0-46c3-a529-d5a44720240d · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.215429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.215429Z digest=sha256:08725364a5086d4383c9469e34cc7d1095bb72aa9c575a0036004c0bc1ffdb40

Observation 8f0e4626-515e-4a87-82c9-1358dc631503 · outbound

This paper cites Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.277581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.277581Z digest=sha256:e74a5de5ae5af9ff6ba0303d0d4dda3d00fe4cf6987ad3e80cd060a3d71d2cc7

Observation 08779163-fc63-47d5-92c3-8e12b46253ba · outbound

This paper cites 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.320043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.320043Z digest=sha256:57f95ffce2abc6c8233ecab417b76ec8d898f0c10d7709b6bb559b38e90fdc37

Observation 3a337eac-a6ff-4da2-9b54-bef0828e8066 · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.367094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.367094Z digest=sha256:6c1efe7aa583b2cdb6400283de68b49b40173b6495fec59f4b23959b70570c7c

Observation a34fcf9c-3f7f-4b4b-b317-abf2803f0958 · outbound

This paper cites Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.383190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.383190Z digest=sha256:5c088e3f67b85970996b438bdd19b01a0274ff68e5f19c7de725ade8db8d1525

Observation 42c428b8-ffda-4025-99c0-b416bcb1efcc · outbound

This paper cites Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.389249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.389249Z digest=sha256:8a09d8281340668a974ad23cafabbb7e847d17c2cb339c0a14625a5cd2ed66f1

Observation 5a57a09e-1a22-4cf4-aabe-9ef8fe282024 · outbound

This paper cites ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.392263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.392263Z digest=sha256:82d518ae77b2efb50cefeb80cfa1e5d5c74b9d4976138c92537a6a2bb23238f6

Observation 59ee0b39-5d19-4f57-8997-2578b3ce65f0 · outbound

This paper cites Multi3DRefer: Grounding Text Description to Multiple 3D Objects.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Multi3DRefer: Grounding Text Description to Multiple 3D Objects

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.395853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.395853Z digest=sha256:c0b7ee15f18c9e19a3dbc0a15d4941ae0612d71a634a2dc3b6ab04963a1ebe12

Observation 8ebeea71-6527-44ba-a6a1-93d19a00a5f6 · outbound

This paper cites Scan2Cap: Context-aware Dense Captioning in RGB-D Scans.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Scan2Cap: Context-aware Dense Captioning in RGB-D Scans

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:33:29.990241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:29.399144Z digest=sha256:3987711db7792a4e7c75aa3eb14a4103172855ae99e0fa78e0535deaf8008eaa

Observation 8636b1ed-3519-4725-a936-864a867a45f7 · outbound

This paper cites ScanQA: 3D Question Answering for Spatial Scene Understanding.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models ScanQA: 3D Question Answering for Spatial Scene Understanding

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:33:29.939265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:29.425799Z digest=sha256:55005f2100f1becc11067023c8e6fe1eff41a4843ba80eeea0f55f7707f7fe84

Observation 32cd7efc-d531-4675-b0dc-53185f7009b0 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.481703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.481703Z digest=sha256:c117d0cd34dda2348d7400c4f9042eea80d22cb9ce788e1dba1325adb2894387

Observation 42ce05c5-32e2-4ad2-9e09-cdabc4cfc9f1 · outbound

This paper cites Mask3D: Mask Transformer for 3D Semantic Instance Segmentation.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Mask3D: Mask Transformer for 3D Semantic Instance Segmentation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.538485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.538485Z digest=sha256:45d8b228deb7e386cfc66499a03e77a6fb92db55fd844bac613a20193252db49

Observation f6c8680a-2431-4bca-ac4a-f9586d1bbc40 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:31.122134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:29.598617Z digest=sha256:a113c84369ff2acf127ca9b9c7ee266164c89c98630ef9ecf86f48c64619cc71

Observation d4b87380-ad14-4e41-9610-156e596fdc85 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:31.075624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:29.662827Z digest=sha256:2e9385f6c1dab9fdcb25aff55e1ae3453fe1261f3cf7166cf16c3d8cdf0d6042

Observation 282e49fa-6365-421a-adb5-a8656ab047ed · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:30.938959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:29.699550Z digest=sha256:6d9d112ff05e12c71ebab820bbe3af36ee3f02042f5f7516bbaa1331bdfb0e17

Observation 16b01bc7-69f0-4b7c-8490-4910cfaabd85 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:30.824082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:29.734512Z digest=sha256:606bb7dcdb3e3c56659251d185029c1f3cf49972eeb284a937b36e8822086577

Observation 7b4a6676-c5b5-4cf8-a6a7-226466578864 · outbound

This paper cites We retain only patches that contain at least one valid-depth pixel; let V ⊆ {1,.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models We retain only patches that contain at least one valid-depth pixel; let V ⊆ {1,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:30.815753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:29.739770Z digest=sha256:8c62ee174cce1575367035668b905205c9bfb0b0e485fadf62908c9c19f5961c

Observation cc25501e-7ad7-4c4b-a7c8-1fa287ff6a73 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:30.807455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:29.742791Z digest=sha256:673baef71cddb5e99c8d31ad78673289d4829e370dd8f4a0406ec5036adc4d98

Observation 34ded46c-8425-4293-aec7-003f99db8477 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:30.730339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:29.745742Z digest=sha256:3d6e9fb0d4753ce77612859c46e1cfaaa6c5e1f5042cc1b49def1518b4458a5f

Observation 79576ef0-f845-47b8-ae6a-14427c427d17 · outbound

This paper cites Dist., Room Size, Rel.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Dist., Room Size, Rel

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:30.611920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:29.749409Z digest=sha256:2d7a173526d2f4080fee77e6cae505f3b7ec079e01155a0ce1b6b22facc2f772

Observation 6aed1b9d-374a-4b8b-8a21-a7fb85a65085 · outbound

This paper cites Feature. Dist.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Feature. Dist

Reference 86

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T15:33:30.561614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:33:29.808181Z digest=sha256:93c3201cf65e4bbde793ed0b99cd854f3289b71896500bc50cc25a0de9b28f98

Pith citing papers

No inbound Pith citation observations are available.