Pith. sign in

Paper Citation Record · LEDGER

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models

As of 13 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 0 inbound Pith citation observations for arXiv:2608.10864.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.10864 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:33:29.808181Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact2
  • verified fuzzy14
  • unresolved65
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 54a1a48e-541c-4a26-99a9-ac63c5153957 · outbound

This paper cites VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.118241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.118241Z digest=sha256:e8b0d10de08660a1b3df8d168b2d76bc776f3c4491efbd3bdbfc601d71a8f117

Observation 1f84e78f-dfcc-4e6e-8447-694e34bbad57 · outbound

This paper cites 3drs: Mllms need 3d-aware representation supervision for scene understanding.Advances in Neural Information Processing Systems, 38: 67961–67988, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models 3drs: Mllms need 3d-aware representation supervision for scene understanding.Advances in Neural Information Processing Systems, 38: 67961–67988, 2026

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.168523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.168523Z digest=sha256:25934208bc33f38d79b91d36f8e2fe5d278945c2614d01f55e530c42db177749

Observation 72786601-5a8c-4972-bbd3-73db14d1f467 · outbound

This paper cites CUP Archive, 1967.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models CUP Archive, 1967

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.208043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.208043Z digest=sha256:24780e5238e169ae6986a5586539af2ef083666a9410413151dffdf1eb2467b6

Observation bb9cd8f1-1932-4808-9c04-adda660b79d2 · outbound

This paper cites Henry Holt and Co., Inc., New York, NY , USA, 1982.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Henry Holt and Co., Inc., New York, NY , USA, 1982

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:32.032179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:27.258250Z digest=sha256:4f462e93a3d63f5dfc261fa8820e86926c16551b6179b485fc134802780e635b

Observation 8b75e6f2-0426-4b6a-847d-ff22eaf69d19 · outbound

This paper cites Number 6.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Number 6

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.296641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.296641Z digest=sha256:c4a7eb54aa33732c1c3e37dbf7e5245fdb828e8b7e770487743c4313903008a7

Observation 9c1485a0-d1ed-4f5d-898d-49463389cea6 · outbound

This paper cites Separate visual pathways for perception and action.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Separate visual pathways for perception and action

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.334241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.334241Z digest=sha256:37c83f0423bbd567b9a367065dc231df7cf3fcfd0f1eb522b4379783d1853a23

Observation c70e74d4-ee1f-4bb0-a477-e732b8540558 · outbound

This paper cites Mental rotation of three-dimensional objects.Science, 171(3972):701–703, 1971.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Mental rotation of three-dimensional objects.Science, 171(3972):701–703, 1971

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.374416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.374416Z digest=sha256:35f07e27b331306c41511c4da96b2e0ed903ea809d7f91f9c09e29ea02f1db20

Observation 53fa584c-0db6-46e2-a28b-fd03061b08d9 · outbound

This paper cites Oxford university press, 1978.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Oxford university press, 1978

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.403655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.403655Z digest=sha256:2deb4233bb5f3989a04bc0745774316785dd8b6f04457b9a891b8f33b7345c2c

Observation 3d47e3f0-36cc-4b90-b178-965802fae45f · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.434416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.434416Z digest=sha256:5d12b1d7e538902c9e4315e52fd18c578c120ac73c01fb55d983ee378eb702b0

Observation 04ee055b-8d94-4616-9a27-fa3eac91f6a3 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.467684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.467684Z digest=sha256:7d8281dddbeecd81d20761d8aec9c7d95864dd1c59e019dc16de2c6d456fefe4

Observation c274fcc8-ac9b-4312-b3d0-4d118e2a2623 · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.488894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.488894Z digest=sha256:cc2e9f74424bf861f95948c7695734e2018e348ca78a2f7c1145c6839f82ae10

Observation f4020c15-82c7-4b85-b0ef-3ff96bf5ab12 · outbound

This paper cites Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Open x- embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.513243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.513243Z digest=sha256:b0e3224f76817583c3757d91a8043a6791da9d1945b42877e4e5c39fe5864d59

Observation ecd75291-ac67-40ae-9e16-39afc05ac79f · outbound

This paper cites Are vlms ready for autonomous driving? an empirical study from the reliability, data and metric perspectives.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Are vlms ready for autonomous driving? an empirical study from the reliability, data and metric perspectives

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.536117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.536117Z digest=sha256:b59a836008b686e8c6cc46b7bd923c15ea3d5809c99aec990eb9ebd8f12fd148

Observation 0f5fb8ed-eb7f-4d78-a561-d48b5f067980 · outbound

This paper cites Qwen2.5-VL Technical Report.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Qwen2.5-VL Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.559154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.559154Z digest=sha256:fc1bd6b798b154a9c73005349ea1364071cb9c6311460d6d35ea157aa92bbe33

Observation 205a6793-af32-450d-9f26-0bdfd8316df2 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.597794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.597794Z digest=sha256:c4fdd1cadbc0451978b89487b00cbad5d113c9ad00374091b5f33b787ce33eff

Observation 61b7547e-1d9b-4a8b-9164-36bb4225850c · outbound

This paper cites LLaV A- video: Video instruction tuning with synthetic data.Transactions on Machine Learning Research, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models LLaV A- video: Video instruction tuning with synthetic data.Transactions on Machine Learning Research, 2025

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.861793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:27.627853Z digest=sha256:33194044defc104be3c183e7e0a8526b682d75e682be289a11057f55294e659c

Observation 7bb53d38-73a0-4035-a903-7194cffba8ed · outbound

This paper cites Probing the 3d awareness of visual foundation models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Probing the 3d awareness of visual foundation models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.851699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:27.662301Z digest=sha256:9cf0db83b3bf785d76ce268c6eb68e8e3d801daaf14491fb0a41fa12bfcf63ad

Observation 5cf0e01c-42fd-495e-a90d-a30d3d2c7c98 · outbound

This paper cites Sti- bench: Are mllms ready for precise spatial-temporal world understanding? InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 5622–5632, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Sti- bench: Are mllms ready for precise spatial-temporal world understanding? InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 5622–5632, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.693368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.693368Z digest=sha256:2c73f8c6ec1ceb987f44465999fc85872834c16738d23d5a520d5e50289d7ae2

Observation 87d6ce1f-049d-4151-902c-6f4b92e319bd · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.750364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.750364Z digest=sha256:0d1df6044320b526cc2ec5bdf9d37341d4266799e3163de0c290bef6d579aeb4

Observation ccddb12c-aba2-4058-990e-f3875f62cebe · outbound

This paper cites Vlm4d: Towards spatiotem- poral awareness in vision language models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Vlm4d: Towards spatiotem- poral awareness in vision language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.802712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.802712Z digest=sha256:41c16f7e13826714384580017a914eb25818d55e84bab469d118a653110a6c32

Observation 4a06d401-c430-4d9c-bbd9-3f0656262306 · outbound

This paper cites Cambrian-s: Towards spatial supersensing in video.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambrian-s: Towards spatial supersensing in video

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.915456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.915456Z digest=sha256:b04d46cf95ea58583fd4208e4dc20e08411c99794c8d190152c9ce09a517e2a1

Observation beac19f4-b222-40b6-9e7e-69cf97e0fcbe · outbound

This paper cites Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Questioning Representational Optimism in Deep Learning: The Fractured Entangled Representation Hypothesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:27.973683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:27.973683Z digest=sha256:ec108c12e01ad53c336c1ed75df43992b6df14d355f28e3accb24861235f4e05

Observation 4647a97b-0e1f-4c31-bcbd-a85f98f5f94e · outbound

This paper cites Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625, 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.004225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.004225Z digest=sha256:93756bce5d27ee005a27a657ba536cfdad52d0482ef95e794c60346b8aa58dc1

Observation f434eff9-3b11-4021-8934-ecd778e162be · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Distilling the Knowledge in a Neural Network

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.057213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.057213Z digest=sha256:7ab013504e2093e208e32fae847cd4efdfc89362e96da53b8cf658d1c0b8d3b2

Observation c976a4dc-f4a9-40f3-b939-dd1c9a60943b · outbound

This paper cites Unsupervised natural experience rapidly alters invariant object representation in visual cortex.science, 321(5895):1502–1507, 2008.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unsupervised natural experience rapidly alters invariant object representation in visual cortex.science, 321(5895):1502–1507, 2008

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.803965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:28.107936Z digest=sha256:e513d60ccaeb1bcd09ced02a37a7d0dcd435cc07da1f1895cda64d68e9a17bb4

Observation a77fa326-3b6b-490e-bc6b-4cf2003d5c93 · outbound

This paper cites Slow feature analysis: Unsupervised learning of invariances.Neural computation, 14(4):715–770, 2002.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Slow feature analysis: Unsupervised learning of invariances.Neural computation, 14(4):715–770, 2002

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.688307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:28.143957Z digest=sha256:d03bbb21979495d07096ed8e3ba6ae3d622646c9ee5deafaa2a63f94924c07dd

Observation f1a1f640-c5ea-4332-b960-e17a12b4f2f1 · outbound

This paper cites The tolman-eichenbaum machine: unifying space and relational memory through generalization in the hippocampal formation.Cell, 183(5):1249– 1263, 2020.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models The tolman-eichenbaum machine: unifying space and relational memory through generalization in the hippocampal formation.Cell, 183(5):1249– 1263, 2020

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.182390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.182390Z digest=sha256:6d6757e577489a31dff22398d78dc13ba7fda718fd6ad82d8cb0ba904c877f22

Observation 0b40abdd-a5db-4460-8a8c-2ef393025aae · outbound

This paper cites Psychology of spatial cognition.Wiley Interdisciplinary Reviews: Cognitive Science, 3(6):565–580, 2012.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Psychology of spatial cognition.Wiley Interdisciplinary Reviews: Cognitive Science, 3(6):565–580, 2012

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.556519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:28.212915Z digest=sha256:70e1399a24ebc35e8c6d3cee3e7a6b9babb32eb366b1259662146bdcf796e2d4

Observation a6faf50b-53c7-4948-b050-b6ca50e69b8f · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.247419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.247419Z digest=sha256:7117004e4d3c9ef355e3e20f610fbd30912355ae85f0134cdec56f4b366e484a

Observation 67da9146-8180-461b-920c-a09a20989136 · outbound

This paper cites GPT-4o System Card.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models GPT-4o System Card

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.289469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.289469Z digest=sha256:2d586f82e74b97d8e6a455894cfefa4783454d7fa18ab0e7d9297887db9730c7

Observation 7c938b5e-59a1-4a40-bc81-2a4b0f23cbd7 · outbound

This paper cites STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.317477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.317477Z digest=sha256:6a15264b6fac91e4fc6de656b9653033ef4407da0f3bf15048b4e0f942a2eedc

Observation d9ce87d5-a09d-4627-9a8c-8812128504eb · outbound

This paper cites VLM4D: Towards Spatiotemporal Awareness in Vision Language Models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models VLM4D: Towards Spatiotemporal Awareness in Vision Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.342219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.342219Z digest=sha256:30a20bb8c32d301f577a1024987c20dcc58f11b7878aa5d7d6d6b161fbf50843

Observation 6ccf82cd-9f78-4f08-9861-33308bddca74 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Learning Transferable Visual Models From Natural Language Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.364382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.364382Z digest=sha256:96290fef3d50c4819ed7009db860ca91a6b7a857132df08764c74cd059080da6

Observation 1c083e24-1267-4b2f-94ec-a3fe365ffef1 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Sigmoid Loss for Language Image Pre-Training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.381618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.381618Z digest=sha256:5a93686567afdba75803b6909623a53855831fdc46ee7481c88d6d8536e5edbb

Observation 8e3a871a-0a6a-473f-a70b-2be1a1d4fa0f · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.399065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.399065Z digest=sha256:d0518c1c26155cb0089f018e69d06132023cfed5c069b5c36d07cb9a105a8986

Observation 96ab4e6c-53d2-426a-abc4-f0e554584a1f · outbound

This paper cites O’Sullivan, Fang Cao, Tahoura Nedaee, Kamyar Rajabalifardi, Fei-Fei Li, Ehsan Adeli, and Euan Ashley.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models O’Sullivan, Fang Cao, Tahoura Nedaee, Kamyar Rajabalifardi, Fei-Fei Li, Ehsan Adeli, and Euan Ashley

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.460314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.460314Z digest=sha256:85b711f169e009c8379da09b28489307e057a14591da5f4112c7a96122642882

Observation f65e9dfb-4032-429b-9eef-45796e4626c8 · outbound

This paper cites From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models From flatland to space: Teaching vision-language models to perceive and reason in 3d.arXiv preprint arXiv:2503.22976, 2026

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.508799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.508799Z digest=sha256:d942b92433cab17d47f7289675edcf0b5a22f1ce429c91388158ca4686ea1c39

Observation ebcc31ef-8308-4671-a011-7d8922eaf6e2 · outbound

This paper cites Visual spatial tuning.arXiv preprint arXiv:2511.05491, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Visual spatial tuning.arXiv preprint arXiv:2511.05491, 2025

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.562622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.562622Z digest=sha256:6ed8d5e26fa7f1169ba0cc2e32df6d53af3f0a49bff84033876517bdfd468cda

Observation 06b1887d-f586-4ca0-8901-29e9cdbf6dd2 · outbound

This paper cites Cambrian-S: Towards Spatial Supersensing in Video.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambrian-S: Towards Spatial Supersensing in Video

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.585047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.585047Z digest=sha256:5d5088e8bd09cecf500d1e90b11fa8ad0c60b58091f8becfba876a7a65472139

Observation da5f7c1d-cf31-4b45-84aa-571d67dcbbf6 · outbound

This paper cites Vggt: Visual geometry grounded transformer.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Vggt: Visual geometry grounded transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.614240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.614240Z digest=sha256:9aabbf21e8d07fc983e24f2d83b61b3cc4384055a2bcc4877eceb6a3bbba0df1

Observation 10bccea0-33e6-4810-9190-2e9cd6e583d6 · outbound

This paper cites Continuous 3d perception model with persistent state.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Continuous 3d perception model with persistent state

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.541921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:28.641131Z digest=sha256:da7cb1a8b1f759b97bef0eee5ce345cdb0aa481d3f34be108731156648762374

Observation 084fcc2b-0416-453a-88f8-c2da350228e7 · outbound

This paper cites SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.651082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.651082Z digest=sha256:1367495a86eeaa1b9f761f45560377e49d1144b5667cc45f2c46f52c22ab9abf

Observation bb3b1a12-f7d2-46a3-a3ef-2d507a968d3b · outbound

This paper cites G2vlm: Geometry grounded vision language model with unified 3d reconstruction and spatial reasoning.arXiv preprint arXiv:2511.21688, 2025.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models G2vlm: Geometry grounded vision language model with unified 3d reconstruction and spatial reasoning.arXiv preprint arXiv:2511.21688, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.682876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.682876Z digest=sha256:a87aa68abcc4ca4236fb4376caf2912cdb8a0da57f02d2edeb23d607c37a2a9b

Observation a6476444-98df-496a-8dec-752400d760ba · outbound

This paper cites Think with 3d: Geometric imagination grounded spatial reasoning from limited views.arXiv preprint arXiv:2510.18632, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Think with 3d: Geometric imagination grounded spatial reasoning from limited views.arXiv preprint arXiv:2510.18632, 2026

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.731186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.731186Z digest=sha256:de55738da3f8b0398654d8d034bd5a6267edcedc79306ea77567d17a54bfd454

Observation a7c3b56b-2384-4993-af06-ca9fde4930fb · outbound

This paper cites Vision-aligned Latent Reasoning for Multi-modal Large Language Model.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Vision-aligned Latent Reasoning for Multi-modal Large Language Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.783675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.783675Z digest=sha256:89e09a0db484382cdb36ac2d3facd2f50fed38ba6651634df9941356188c1948

Observation 420ee94e-dab1-4a97-b330-7b95dfc9d6f4 · outbound

This paper cites Cambridge University Press, 2001.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambridge University Press, 2001

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.508886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:28.816106Z digest=sha256:56340a77e8862de973f83c60ba8c9d2bfe1fc8b2b109f5c89d3d03be817e2c6e

Observation a1e511ba-3ad4-4ca0-be74-f96a6923cdf0 · outbound

This paper cites Cambridge University Press, 2014.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Cambridge University Press, 2014

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.837244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.837244Z digest=sha256:7e614d283c359d2388a7a7e6c4bfed6040939558d4bbe4f0a8078b5fccfba26e

Observation 58daaab1-74dc-4075-be49-655b718d8179 · outbound

This paper cites Relational knowledge distillation.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Relational knowledge distillation

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.847346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.847346Z digest=sha256:3204b4c152d5e4504d560f3f7200ab2b9cd8ea89152daf7ec6b3477fbcfd3056

Observation c5c2a94d-f3f4-41d6-824f-99fc63b9abef · outbound

This paper cites DINOv3.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models DINOv3

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.881926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.881926Z digest=sha256:248b0b58d98d6807dcd5121118085146c3bda45f1f98bd9b40f6069237fe0676

Observation ceadf3e1-0473-4f8c-8652-8cbde385cbba · outbound

This paper cites Perception Encoder: The best visual embeddings are not at the output of the network.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Perception Encoder: The best visual embeddings are not at the output of the network

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:28.893432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:28.893432Z digest=sha256:0342976aff92bf627798a4eea62a23128c70fe6b66bb2de388ab08a3efda61ee

Observation b7f2d45a-2f0f-4047-ba08-e8d01fcba513 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.442289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:28.897760Z digest=sha256:7da5f4d71ed157c05e5f7647236ef6c9265a4b918aa6bbc97040d70d7ba371fa

Observation 7d9ab556-828c-43a9-95f8-2f452d8dee24 · outbound

This paper cites Videorepa: Learning physics for video generation through relational alignment with foundation models.Advances in Neural Information Processing Systems, 38:122647– 122676, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Videorepa: Learning physics for video generation through relational alignment with foundation models.Advances in Neural Information Processing Systems, 38:122647– 122676, 2026

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.383300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:28.927333Z digest=sha256:44af0750b6433b5bc69274cce625aa411eadeb8e78b6ff1646f7ec4b8c40ab24

Observation ae99a4e1-13d8-4c11-97b7-8dc38ff6acaa · outbound

This paper cites Moalign: Motion-centric representation alignment for video diffusion models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Moalign: Motion-centric representation alignment for video diffusion models

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.272471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:28.983429Z digest=sha256:ac7855ac934aa51817a235de579b308ddb67e87ed10fdd0aec08530884203747

Observation 34256c33-f0dd-4f2e-83d0-80391481943f · outbound

This paper cites Lever- aging vision-language models for improving domain generalization in image classification.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Lever- aging vision-language models for improving domain generalization in image classification

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:31.153761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.047125Z digest=sha256:cb091d296115b0411f691682b32e31df51479839cf60d418e4602aa8839a8613

Observation b33995ff-cbda-4283-84df-d2ae3fc1b135 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.100129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.100129Z digest=sha256:37e6bec57a118a3d5a4f3811bc1f369b2fd008a6256b9a83e8223a13eda7f307

Observation be2c11c2-5218-40a3-9139-f67c19d73469 · outbound

This paper cites ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models ScanNet++: A High-Fidelity Dataset of 3D Indoor Scenes

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.148328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.148328Z digest=sha256:bd606aa2b369dece992202f0f099c0c6ac3198d35798c55472c56bb802f63188

Observation b9bca34a-8a20-46eb-b56e-c75c6970ab98 · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.171413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.171413Z digest=sha256:eba5cc8608b0dc5b058f7af1cade393463683066710accafea40bc584772e51b

Observation 8a819d40-0fd1-4343-a531-38dee812edd0 · outbound

This paper cites Scaling spatial intelligence with multimodal foundation models.arXiv preprint arXiv:2511.13719, 2026.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Scaling spatial intelligence with multimodal foundation models.arXiv preprint arXiv:2511.13719, 2026

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.175284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.175284Z digest=sha256:a517f2a788e19e0035138448c33b7f5756076fa158559c2df69bdb69f301d2d6

Observation f08d0c3c-1a81-4e3a-8b09-b52449f00a60 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.181722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.181722Z digest=sha256:42de506a035c4c9cd72eb2ba9680abd10c4c2546196a4632761c6794ae7f49a1

Observation 7d3efa8d-16ba-4abc-a9bd-0e210dd59505 · outbound

This paper cites Grounded 3D-LLM with Referent Tokens.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Grounded 3D-LLM with Referent Tokens

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.184264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.184264Z digest=sha256:97fcaead68260d0cb5b1ef0baf8d93e3dbfa66a6bb235ee056db9181e77f819a

Observation d48dd1ca-9762-44c5-a86c-d64e53f01c20 · outbound

This paper cites Unifying 3D Vision-Language Understanding via Promptable Queries.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unifying 3D Vision-Language Understanding via Promptable Queries

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.188680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.188680Z digest=sha256:659ce75b9cfbab214b88c792c0a63b783c5ad3f271be4d6daa12fc50d91dfb7f

Observation 80814b0b-2ab0-46c3-a529-d5a44720240d · outbound

This paper cites Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.215429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.215429Z digest=sha256:2307da5d7136f473dd8868a643596d811d500051966f8dd9a62b5e03dff8fcdc

Observation 8f0e4626-515e-4a87-82c9-1358dc631503 · outbound

This paper cites Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.277581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.277581Z digest=sha256:d844bcea23d7e49cd6aae01b13a0f2bcc21a3e71a81ed7c835f0abb30555bd65

Observation 08779163-fc63-47d5-92c3-8e12b46253ba · outbound

This paper cites 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.320043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.320043Z digest=sha256:fc3794fa2e74c41694cfe1b14a0428cc90e4f9ba71ef99712f044991d493185a

Observation 3a337eac-a6ff-4da2-9b54-bef0828e8066 · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.367094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.367094Z digest=sha256:cdcaeca08fbeaa067b2e69c335d11e1bdb4c0b748ea15c927c0d0ab9d02b43d0

Observation a34fcf9c-3f7f-4b4b-b317-abf2803f0958 · outbound

This paper cites Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.383190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.383190Z digest=sha256:b0d10bface458aa3a01b5a783106441b8e01d333b0fb910c17bd0043b1ed1c23

Observation 42c428b8-ffda-4025-99c0-b416bcb1efcc · outbound

This paper cites Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.389249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.389249Z digest=sha256:bd057783b20a1ea86f686dce36e5f26cd4d2007612c7649f11ca3432b9c0b344

Observation 5a57a09e-1a22-4cf4-aabe-9ef8fe282024 · outbound

This paper cites ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.392263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.392263Z digest=sha256:67e1feaed3b07ec24614d2b3dc96d1beef50f80613642cc45b7774650fdc6b38

Observation 59ee0b39-5d19-4f57-8997-2578b3ce65f0 · outbound

This paper cites Multi3DRefer: Grounding Text Description to Multiple 3D Objects.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Multi3DRefer: Grounding Text Description to Multiple 3D Objects

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.395853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.395853Z digest=sha256:30bfad3f7b7e17d77077dc47e6e17a319e56a9592442244421aa8ff6dfbe8c54

Observation 8ebeea71-6527-44ba-a6a1-93d19a00a5f6 · outbound

This paper cites Scan2Cap: Context-aware Dense Captioning in RGB-D Scans.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Scan2Cap: Context-aware Dense Captioning in RGB-D Scans

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:33:29.990241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.399144Z digest=sha256:de509a2e811b0197ea078e17d399fcf03eca9ccb1228927e06b433d2635c5fe4

Observation 8636b1ed-3519-4725-a936-864a867a45f7 · outbound

This paper cites ScanQA: 3D Question Answering for Spatial Scene Understanding.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models ScanQA: 3D Question Answering for Spatial Scene Understanding

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:33:29.939265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.425799Z digest=sha256:a07216db538265e162e643bd1f3d49cee54c3c9558033d0d2427d0a2624c7e66

Observation 32cd7efc-d531-4675-b0dc-53185f7009b0 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.481703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.481703Z digest=sha256:8b6f2e55766a64adcb1913e16028ad0b55fd2785cbba2fe07634dba923e1e246

Observation 42ce05c5-32e2-4ad2-9e09-cdabc4cfc9f1 · outbound

This paper cites Mask3D: Mask Transformer for 3D Semantic Instance Segmentation.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Mask3D: Mask Transformer for 3D Semantic Instance Segmentation

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.538485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.538485Z digest=sha256:d1808822d5b80f7aa1361c17018fa48a2d40b446f5a2d9c6506eae481328d8ad

Observation f6c8680a-2431-4bca-ac4a-f9586d1bbc40 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:31.122134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.598617Z digest=sha256:682b194004922e80929445f2580ce2d7ebb2d64122f64cd480e5f44d8ab136e4

Observation d4b87380-ad14-4e41-9610-156e596fdc85 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:31.075624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.662827Z digest=sha256:fef48443d87fd20d78b2f79f29e096a19ccbb702b9331e195e6320dc5841e3a8

Observation 282e49fa-6365-421a-adb5-a8656ab047ed · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:30.938959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.699550Z digest=sha256:edfa47bf4c8e828acb7f16a6d64bb536abd48cd87ac4ab6920736057493e535a

Observation 16b01bc7-69f0-4b7c-8490-4910cfaabd85 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:30.824082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.734512Z digest=sha256:bf8962321ed336c85f72da7b5d20455ff39b71a3532af371e1d4dd66d5b1b686

Observation 7b4a6676-c5b5-4cf8-a6a7-226466578864 · outbound

This paper cites We retain only patches that contain at least one valid-depth pixel; let V ⊆ {1,.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models We retain only patches that contain at least one valid-depth pixel; let V ⊆ {1,

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:30.815753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.739770Z digest=sha256:346cf5f79960c6ab849589cac645c2b1cb68783f8410061308fe355edb89983d

Observation cc25501e-7ad7-4c4b-a7c8-1fa287ff6a73 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:30.807455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.742791Z digest=sha256:b62ad16c9cdc128ae33ce928eececc5f2e24d1ed1fad82859e92778f2522f80d

Observation 34ded46c-8425-4293-aec7-003f99db8477 · outbound

This paper cites an unresolved cited work.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:33:30.730339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.745742Z digest=sha256:2b403112a3961be8e8edc314fa35faf11cc3a3fa520113dd68acfe46f8d9311b

Observation 79576ef0-f845-47b8-ae6a-14427c427d17 · outbound

This paper cites Dist., Room Size, Rel.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Dist., Room Size, Rel

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:33:30.611920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.749409Z digest=sha256:21e4677c7569ea70c51f935280bf3805ca1989861fb437919c89d0bd19c61698

Observation 6aed1b9d-374a-4b8b-8a21-a7fb85a65085 · outbound

This paper cites Feature. Dist.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models Feature. Dist

Reference 86

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T15:33:30.561614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T15:33:29.808181Z digest=sha256:5e19dae25aeb48e663edfbb317ba6b3a174383a1483313fc5540ef2c5c301079

Pith citing papers

No inbound Pith citation observations are available.