Pith. sign in

Paper Citation Record · LEDGER

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

As of 5 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 71 inbound Pith citation observations for arXiv:2505.20279.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.20279 v5

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T12:54:01.013242Z

measured 166 of 166 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 71 of 71 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:38:55.151754Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

95 of 95 outbound references displayed

  • verified exact29
  • verified fuzzy63
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 288cff19-520e-4cfa-b0d2-7c93e1465652 · outbound

This paper cites an unresolved cited work.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-19T13:02:19.416499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:0760f33d479545181f9576934e0874870f284ffca5c6bca46deb570becb2cc8e

Observation d2fc3e26-dbaa-4493-a5a5-8f78ac1cf3e8 · outbound

This paper cites Specificity of learning: Why infants fall over a veritable cliff.Psychological Science, 11(4):290–295.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Specificity of learning: Why infants fall over a veritable cliff.Psychological Science, 11(4):290–295

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.395095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:edf36296064bae533815de207d532f13e19f63a0a5bb5af63bc2d23ed220d8ed

Observation e7e5e756-a79e-4fb7-b8ed-ef4fbfbbfd6b · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Flamingo: a visual language model for few-shot learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.390471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:e848b3e13e827f9744341f220dd0cb76c93b4e0ea74a25ecb0bc8b49bf1b9910

Observation 639cbea1-0a31-4838-af80-3f66b605261b · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Scanqa: 3d question answering for spatial scene understanding

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.370909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:36f9cca39c376325662ee1b544752067cc378e888be09debfb36b5dce5a7409a

Observation 6ce9edef-0c4f-464a-9457-60fab7defb66 · outbound

This paper cites Qwen Technical Report.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Qwen Technical Report

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.946057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:fd984fec70a46bbcae76b1eea98f81e47349a7a6883d0c7a5dca465aebfda0c7

Observation af1f33da-af6b-41b8-8518-180d7e997bc4 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.937834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:bfbcd7a62325bf7dcf39c3a174aa80bf21535b94aaa8507b9229c5fc6a00dc18

Observation f4be9339-441c-49d9-a6de-86ce00b5b57c · outbound

This paper cites Qwen2.5-VL Technical Report.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Qwen2.5-VL Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.971216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:401bea8bfefee865057009d016dad23e8eb981ae7dd297fd5ff856be06e81c1e

Observation 7df4d349-4fbd-4a94-b5af-b136e2a2bd04 · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.979751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:14913c4fcac984253fef23187f2a6e581d0764d4dd5ea144e48d90ae36d5c423

Observation ed22b408-5285-4d4d-a914-e5c3bcf60c0c · outbound

This paper cites Large Linguistic Models: Investigating LLMs' metalinguistic abilities.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Large Linguistic Models: Investigating LLMs' metalinguistic abilities

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.919160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:ccd7f82e3a35280707dda52f5922b2f31b995a1b32da83658a29e2ab3c65cb76

Observation 982e7e28-d815-48d4-9b1e-1abb639db716 · outbound

This paper cites Language models are few-shot learners.NeurIPS.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Language models are few-shot learners.NeurIPS

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.283045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:eef059040c2b1782e20bcd709a57adf659674d917304e60453ac4bfbb9ec3d37

Observation db9b0fdf-6b0e-419c-98d0-760af2ad7471 · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.280271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:d708c0dfec538271c3e91f679a4467657c71e45720935a74fd1f988dcbfcd2c9

Observation 76bdf84f-f44b-44b6-bead-d5983a2b7160 · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.277972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:05fac12b6b8908f51c4edaf5e5e122207b09e669bc0bccdbf850307a1aa11c98

Observation 4a38ee90-8e7a-4a1b-be44-bb4fedba1558 · outbound

This paper cites Longvila: Scaling long-context vi- sual language models for long videos.arXiv.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Longvila: Scaling long-context vi- sual language models for long videos.arXiv

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.275426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:4f989c073c8eda2c59dae94e7820853286997b6eb5e32461e9a1939763ff91d6

Observation 0ec60a93-f700-414b-b19c-533701813241 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.450434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:7b2d1bf0ccc0c8a2ae1ea3a40077bdee0f831dce2725eac27abba38a0b0d1d77

Observation fc7d53ed-46b1-451f-b050-68c2c4a909f5 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.447169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:74cc994358a195ec6eea40b1f46b3bacffa458a1389e0909998edd836db7f48d

Observation 6180e137-9b41-42aa-b94d-628554fc5bda · outbound

This paper cites Spatial- rgpt: Grounded spatial reasoning in vision language models.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Spatial- rgpt: Grounded spatial reasoning in vision language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.444453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:67128f9cf4fcfee56bf946c8b7f70dd62a97cf9500e7a9ca5a98c210c85853e0

Observation 2c7c9a71-c1c5-4a81-b5bd-50c603d98855 · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.441733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:3b1b358782762f14e38779f06c4cb8dbdf7c3f136033a95dfa7db62a392919f4

Observation e2097ed8-b3b4-4820-8212-72a6f32c502a · outbound

This paper cites 3d-llava: Towards generalist 3d lmms with omni superpoint transformer.arXiv.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction 3d-llava: Towards generalist 3d lmms with omni superpoint transformer.arXiv

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.439059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:1dbc3249ff4266bd06b2ee0eba657078292a085f4908b4242e25ec911f6c1349

Observation 26ba6bdf-b15c-4173-86e0-16b7abba719e · outbound

This paper cites Palm-e: An embodied multimodal language model.arxiv.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Palm-e: An embodied multimodal language model.arxiv

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.435994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:f8ffd81d980af1c1437fb7328cb3067e7c729187b453e282b6e870eef95f19a4

Observation 91cc8df9-50c0-4aff-adee-bdc9dd7bf212 · outbound

This paper cites Large spatial model: End-to-end unposed images to semantic 3d.Advances in neural information processing systems, 37:40212–40229.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Large spatial model: End-to-end unposed images to semantic 3d.Advances in neural information processing systems, 37:40212–40229

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.433172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:ccc24a41a01c50c2852ee60f0c8843d555c4897d5f02e268800a318be064a523

Observation 7c163f6c-01ff-40f3-8fbc-9329d33b7fa5 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.965601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:dcdba80a1f7ecad076f3a8aebf2e613c5d23e8b8cb7153d4f199a0b8912ba38d

Observation f35a9cfc-7b31-49b3-85b6-8d2168020559 · outbound

This paper cites Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.430621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:3b229029b20155a66df8679ef304761ae31bcb5f554cb83a729dc48b556c9dce

Observation ddd3158c-4c61-43b6-bba2-fb0f62201ce4 · outbound

This paper cites an unresolved cited work.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-19T13:02:19.428030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:f0aabd281aa31eff69b79b55709a1ea41103d9351393561e7bfc226299c0a8be

Observation b22ee976-8d20-40a3-be7e-7adab94fe5e0 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.425555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:6bcd8d71a97893dba894d926aaee83b130b73c40cdcef369c02dd5623d6a6b23

Observation eef2d304-8096-4309-a9d5-7da150b0f885 · outbound

This paper cites Cascade cost volume for high-resolution 9 multi-view stereo and stereo matching.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Cascade cost volume for high-resolution 9 multi-view stereo and stereo matching

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.422187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:fc4458a874eecb25d69eb5355258a014ed199760ac4ab5f96b9b649e5849a50a

Observation b56cf3a5-89bd-42f3-b48f-6e0a5385f393 · outbound

This paper cites Cambridge university press.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Cambridge university press

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.419449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:a69827188ba569de9be974b5ed0ef052da554063ab0d62379eb0d7c3f8ccaaeb

Observation e7b214bc-6fd8-4003-9ebc-9130ae929dcd · outbound

This paper cites Individual differences in spatial abilities.The Cambridge handbook of visuospatial thinking, pages 121–169.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Individual differences in spatial abilities.The Cambridge handbook of visuospatial thinking, pages 121–169

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.414137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:cc2f0436f409c22c040da19543afe9d76fef6fb96229442b7d7ccd518f53a5df

Observation 85ac9295-25a7-4836-aacc-09ec91d410e5 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Measuring Massive Multitask Language Understanding

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.976699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:c0d58b13eebf1a495e87bf46427bfb399b3186cebf08380cbac4698f10ccc077

Observation b959b21d-cfcc-4970-a531-70beba8b3ee1 · outbound

This paper cites 3d-llm: Inject- ing the 3d world into large language models.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction 3d-llm: Inject- ing the 3d world into large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.411041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:1c15c2a4cde5f2d72473dc8dd46aeff86520945a51d5ccfdeb7371df2d473a53

Observation 8956cff6-dc7b-4017-8ecf-7489da190266 · outbound

This paper cites Multiply: A multisensory object- centric embodied large language model in 3d world.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Multiply: A multisensory object- centric embodied large language model in 3d world

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.407256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:b4cb96bff6f0f31da489f1557121b52477cd0cef9dd605fae911098977e9051d

Observation 0b072315-f3a0-41d2-80b2-c8005b6f7366 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1(2):3.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Lora: Low-rank adaptation of large language models.ICLR, 1(2):3

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.404467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:4f95b08eade13c660e4c8a1db38e35252eb59d681283efeb434bfb9ddd0262f3

Observation a36efbde-83f7-4cc1-945c-3f65fe7e747c · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Chat-scene: Bridging 3d scene and large language models with object identifiers

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.401520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:256f9cb0333c863b25d5ad10de5932b47ca56378e01e316030ca58c8ed80e3dc

Observation 7c3fa2df-ace9-4161-b619-cff13a410cf1 · outbound

This paper cites Think- ing in dynamics: How multimodal large language models perceive, track, and reason dynamics in physical 4d world.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Think- ing in dynamics: How multimodal large language models perceive, track, and reason dynamics in physical 4d world

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.398597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:d6c2d1814978f28fe610df48c6ddf750e9cbb77270f04022abb0adaaffbc08c5

Observation 13b7a12b-88d9-40d8-9b2c-d6b87400b72e · outbound

This paper cites Gpt-4o system card.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Gpt-4o system card

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.392838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:4eef6884d4d9312b4e872a9309dcfe8df1641b133124a85b3fc6520080cfdfd4

Observation 176b21c4-a32a-48e2-82cd-5929fbabfabf · outbound

This paper cites GPT-4o System Card.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction GPT-4o System Card

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.934955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:3ac736aadd708159cfeade1ffec23de7ec58f59e2704977f51f9804fe4b1518d

Observation 2c14db9d-9c5f-49d7-afa7-dfbbafa5de74 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Scaling up visual and vision-language representation learning with noisy text supervision

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.387705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:20ce20f901a8d6c8325cca8b8919e6a04c67b1aa62a045d25cd85369c431c21b

Observation 2397de9d-ad86-46d4-bfd0-59f3ea0cdc34 · outbound

This paper cites Language models with rationality.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Language models with rationality

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.384751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:f4fd88da8f6e8f4622d2a49dc9c19d6a7f5d5ea345c7d89a9758547fb7864791

Observation 92002432-7ab6-4ae7-9617-0f8d134f234f · outbound

This paper cites Ground- ing image matching in 3d with mast3r.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Ground- ing image matching in 3d with mast3r

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.382302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:f8bb46a46436b84a4ddcf8a0217243ab49e4527a7c8fbd7501ee6f900a2d4b83

Observation e1fb0817-cae5-436a-b16c-0832312d0733 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction LLaVA-OneVision: Easy Visual Task Transfer

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.951456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:762d505fb5167300033b4923284093089d657dbe34e2ca7810c28ae0274c5e20

Observation a0edd5e1-3abc-409c-9775-a45302aaebb2 · outbound

This paper cites Llava-onevision: Easy visual task transfer.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Llava-onevision: Easy visual task transfer

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.379855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:84d751a924c66c087f0ffabd9b85db06bb5283588f18b3b36b4003d12d63465e

Observation a92127dc-e1d9-4520-a9f7-5fd2aafb18a4 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.377361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:016ea218766a415e2434d383196f7b458d3608d5ac761d676d693cc87c65360f

Observation 4c4ae958-d758-4598-99ea-ad5c82545015 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.374251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:0ef7ce4c74032f16bcf7ea197181d195d6ec763d46eb4ec30e1b6d9979e4abe1

Observation 3b714487-2121-435b-a022-cf9295b61b17 · outbound

This paper cites Vila: On pre-training for visual language models.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Vila: On pre-training for visual language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.368407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:2b79ea9ff44bb402fb8ca6a417ffe87fe137ebfcb72ac852fee2ed4a012b4081

Observation 4f9fc90f-956d-4066-bb00-6a7ff5f040e0 · outbound

This paper cites Ost-bench: Evaluating the capabilities of mllms in online spatio-temporal scene understanding.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Ost-bench: Evaluating the capabilities of mllms in online spatio-temporal scene understanding

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.928522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:f4a827958deed200ac536fd70bdfa32c7523040debe7d9ee349437bace0b3736

Observation 4946b5cd-c496-47bd-9de3-29f191b7d7eb · outbound

This paper cites Visual instruction tuning.NeurIPS.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Visual instruction tuning.NeurIPS

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.366087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:8108a1a3458b0c5f224bd6139cf686fe42418173915d91025096465bd132e62b

Observation c77b11ea-bf67-4cfe-b28a-8fe3ed3e4a48 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.995688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:296f54e87d815abd5df525b795c61cbf487282423f6deb05ad906b7a66c9f811

Observation c2108898-ca17-4368-9202-129e1c1152ea · outbound

This paper cites Unified-io: A unified model for vision, language, and multi-modal tasks.arXiv.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Unified-io: A unified model for vision, language, and multi-modal tasks.arXiv

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.363730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:7211df9da3bf856cc87b9a41d93b66b17184e775adde4ef73c38b104a02a9860

Observation c627b4dc-78c6-4e8b-a85c-543e476b10c9 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction SQA3D: Situated Question Answering in 3D Scenes

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.925336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:f8b00f4a633f41c8dfe0e99c811a7ce0a95e23d87afc8dac19692ae3b74af5a9

Observation f927bbe9-841f-456f-bae3-1d998d683bdc · outbound

This paper cites Openeqa: Embodied question answering in the era of foun- dation models.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Openeqa: Embodied question answering in the era of foun- dation models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.360925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:392df1dce955f20c44d1b50ec1761b36c2f77358f010a81c2510c4eb46ef0e20

Observation 9d43534f-83a9-4cf0-bd8e-23df09169d75 · outbound

This paper cites Spa- tiallm: Training large language models for structured in- door modeling.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Spa- tiallm: Training large language models for structured in- door modeling

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.932334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:68133af6009e5839a063f9e60c9dedc15de25500946a382737dbc5afb9243bc2

Observation b5736758-3f31-4fac-a9f7-e3333cedabde · outbound

This paper cites How are the locations of objects in the environment represented in memory? InInternational conference on spatial cognition, pages 174–191.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction How are the locations of objects in the environment represented in memory? InInternational conference on spatial cognition, pages 174–191

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.358474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:09463f37fc3dae40f333a8c2df7ae1f310ed183f6a395f2a3c04269c4c6c0a5a

Observation 19b0647b-b903-4de6-877a-00fbe927077d · outbound

This paper cites Individual differences in navigation: an introductory overview.Prime archives in psychology.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Individual differences in navigation: an introductory overview.Prime archives in psychology

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.355982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:e808a4bd05d9707eef3fadb96133eb180a230ac8282d05fdc23194e36b139d2d

Observation 283d2cf4-5cc5-4656-8047-51b32808da02 · outbound

This paper cites Mast3r-slam: Real-time dense slam with 3d reconstruction 10 priors.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Mast3r-slam: Real-time dense slam with 3d reconstruction 10 priors

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.353271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:245f8cfb6090aae74ff1310bb6c4bdd721edc074e8042f05e3dedefe32ba13b2

Observation c43ae4d9-456b-4dae-a9f5-b21879f2b106 · outbound

This paper cites A Comprehensive Overview of Large Language Models.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction A Comprehensive Overview of Large Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:28:39.632097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:6c12730f64626bb2d923a633eacfe4412a3a04a004347dcff1613260e2d9c542

Observation f66ca052-d8d4-409a-8fbb-064208529145 · outbound

This paper cites Kosmos-2: Grounding multimodal large language models to the world.arXiv.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Kosmos-2: Grounding multimodal large language models to the world.arXiv

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.350951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:3bb0aa03ac32b4866fe6741777a9252ea9ea9158c8204d510ed1d002689a5d58

Observation f580f4c9-297d-4e79-9d3e-64ccb0abde5f · outbound

This paper cites Improving language understanding by gener- ative pre-training.OpenAI Blog.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Improving language understanding by gener- ative pre-training.OpenAI Blog

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.348497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:27a25829cea95c7a461f2d1fdaea6ce7dbd1e918bac2cca9b12de4c52951f530

Observation 06d58fa1-eea7-43f7-92a0-5d7737e15429 · outbound

This paper cites Language models are unsu- pervised multitask learners.OpenAI blog, 1(8):9.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Language models are unsu- pervised multitask learners.OpenAI blog, 1(8):9

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.345965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:9a6a70a5693323c6819475c62a452d4bc68febdc526089a25db67ab79d7fdaab

Observation 1be725fc-8a8f-42d8-b96d-a4ba772a9216 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Learn- ing transferable visual models from natural language super- vision

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.343196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:5483246726bf3eac21b900d49584906ff371963f2a82bcd3769f56fef0fb11c3

Observation f9af211a-6592-49ff-9c64-2a7ff77b0f6e · outbound

This paper cites Habitat: A platform for embodied ai research.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Habitat: A platform for embodied ai research

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.340296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:f94c89ee54ee4d34c7fdd642f22aa4f796276b25eb57e5aafdcfb8acb504cdba

Observation ab55bd1e-d896-4e6b-be9c-7d65a3ea14a1 · outbound

This paper cites Structure- from-motion revisited.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Structure- from-motion revisited

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.337960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:ce53a12f85e48624464ce2263920122415637af5afdd768f575638e383d6a657

Observation d76d7748-9158-4771-b9a3-69c86b2fb3be · outbound

This paper cites MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.954380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:1038baa23b60580fd17ab5e783d2f4ba5657f147ac96fc97021eed98c22c1142

Observation dcb81cea-be14-4721-9888-3d547935b109 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Gemini: A Family of Highly Capable Multimodal Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.948707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:a699c293846825a7b656bcc4cbe0cbb2b46d23dedec2ae4b27b2c23ec590551d

Observation a1ec1067-4ff1-4553-917d-b4d1fd4a3d83 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of con- text.arXiv.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Gemini 1.5: Unlocking multimodal understanding across millions of tokens of con- text.arXiv

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.335354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:f2aba87b20e97d94d186313d1bbe78d3bcacd109d39078fe119e53657ec78da4

Observation 2d0dd26b-8296-47f4-9ccb-fc0c0eba4d91 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction LLaMA: Open and Efficient Foundation Language Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.959632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:761c7663599e8eb1ecf2c7574e20a1f2bb7dfdb2efcf678e9c2e2a98b0fabf95

Observation a48a28e7-5e0c-45b7-b853-3485cc1dc09a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.974006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:7de6542b18c1a683d79cfa3174f4522b833fb4e28c8580de0e18092b57d5310a

Observation a08a204b-1749-46f7-adf6-aac496c39be1 · outbound

This paper cites an unresolved cited work.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-05-19T13:02:19.333007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:2f8b9c5e98f86c625c333740b3bd2b88e95c2fe275502be047fdda7792dadee4

Observation 03683f5f-842f-4921-9417-097cfe935b31 · outbound

This paper cites 3D Reconstruction with Spatial Memory.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction 3D Reconstruction with Spatial Memory

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.998536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:6ab1aa697eb82cd79aa68afeb70e344c3e500191f2ac3e017d1e605065f5e506

Observation 8bfcbaab-93bb-47aa-a2c7-679f23797e2e · outbound

This paper cites Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.982684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:eb79a0a5fd825e0e4016080be8637b0a0efd648c8fdf6b1bd1e82532853144fa

Observation 2f1d1b15-0de5-45ac-8c0a-daec510929df · outbound

This paper cites VGGT: Visual Geometry Grounded Transformer.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction VGGT: Visual Geometry Grounded Transformer

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.992322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:50ec30bb0d2545628b93328bc2be9da03eedea67b1a53e88d4ef9fa542145b05

Observation 5d5fcb24-c093-4280-a9a8-5134bfc89e16 · outbound

This paper cites A survey on large language model based au- tonomous agents.Frontiers of Computer Science, 18(6): 186345.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction A survey on large language model based au- tonomous agents.Frontiers of Computer Science, 18(6): 186345

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.330759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:0c4f04f20f83fa011ec83bdefee18b8edcf3dc84abaf346cea69ef6f8074e785

Observation 6867501f-769a-4719-a96c-8430b0beb986 · outbound

This paper cites Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.328176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:9310e20c6534b5db300ee117e8d87afc78da204d275d4d00887e37a11849bd8f

Observation 7f4afd21-f9f3-41f2-9cd6-eedb231e61ca · outbound

This paper cites Continuous 3D Perception Model with Persistent State.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Continuous 3D Perception Model with Persistent State

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.957147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:1de51bfc062a84e3607059122901270776263c7577a7326f721abe2725951e73

Observation 62118cfe-8cf9-414d-a5b3-bf7c7510b21f · outbound

This paper cites Dust3r: Geometric 3d vi- sion made easy.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Dust3r: Geometric 3d vi- sion made easy

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.325908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:fd222d63ba3e3d9f20b87199d25a8e27534842fc703dfa18d028a8a5bdef7e23

Observation b87f94a7-a7f7-455e-a5ca-bb963b40f335 · outbound

This paper cites Emergent abilities of large language models.TMLR.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Emergent abilities of large language models.TMLR

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.323032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:0fd36009ef668f1bb4242787c564340b925954c42a543bc25db0389826764b39

Observation 464c7c19-7436-4453-817f-9704be2f3139 · outbound

This paper cites Dynamicverse: A physically- aware multimodal framework for 4d world modeling.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Dynamicverse: A physically- aware multimodal framework for 4d world modeling

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.943520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:16e1f1744a1474f0bf87bcd8ad9a6b9a4855286e7d27db987dd12c2d29fccb9f

Observation 075cb424-c74a-4abc-b2a4-31ba790d8c5d · outbound

This paper cites Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.988712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:65c63de5b8570e2fdf2012668010f059b146963177b3f5b93efcdca4e846b05b

Observation 381d89a9-f92c-435a-80ea-61ed73fb442d · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.916212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:50880801489a17f5c2ea721f179a72c63a31c4c4da9355f248db2cf0eeacb1a3

Observation f2f7711f-1566-4c1b-ae81-9b1467a86c9f · outbound

This paper cites Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.962691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:e0ba50cccf12ce07a6738f54f337e5ea873ea28bb9032fa6a5d828cea195dc3e

Observation 7c0e173d-c1f5-406a-a593-ae4c0d7b5607 · outbound

This paper cites Mvsnet: Depth inference for unstructured multi-view stereo.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Mvsnet: Depth inference for unstructured multi-view stereo

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.320163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:5de0f88429533a678e2446ccf54ad3e2c788949da6fbb2af41075db194daf60d

Observation 0ec9e627-5adf-4b66-b74c-86fbda9ac8b7 · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d in- door scenes.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Scannet++: A high-fidelity dataset of 3d in- door scenes

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.317225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:cb6cbbbc31759ff66667f3c67ef05eaba32722498e6e4946265fe46a90151c9a

Observation 04289116-dcf4-4568-9715-6b70ce94f633 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.313956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:0f293ee2aa16ac14fdd329e7078c48d7281501263b6a525b9b80d2cd389afa9f

Observation 2bdd7e1f-26f0-429c-95fe-e645d9aab789 · outbound

This paper cites MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.921894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:247d6bdea6470e030d98e6364ded1600933168eb9a93bde4740c35f3e6069211

Observation af60ba3e-af7f-46ec-887c-445a66aa29be · outbound

This paper cites Spatialstack: Layered geometry-language fusion for 3d vlm spatial reasoning.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Spatialstack: Layered geometry-language fusion for 3d vlm spatial reasoning

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.311221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:d324f3da03456b5af40f524cb5811d2162454a09f289429ee8eb981272d76665

Observation 0ca71d35-d57a-46c1-b31b-74d6af952ada · outbound

This paper cites Long context transfer from language to vision.arXiv.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Long context transfer from language to vision.arXiv

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.308776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:a064e4891200299b98bb0f14aecafbf04eb05ee9dfdb0783fa66f1f1806f0193

Observation cc288160-2790-4e00-b581-d2b03553eec8 · outbound

This paper cites Llava- next: A strong zero-shot video understanding model.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Llava- next: A strong zero-shot video understanding model

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.306070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:7e8e2c501964344f83023b6932239527b0c46a30aa47d1639c2a6789cdb175e1

Observation 9f3976c4-30c8-426f-a89b-b2cd0d54fcb2 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:57:17.940393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:68ab73034acf40609aa0dacf48bdc862d87558a7d1acad6d78171500c39e3931

Observation 7fc8ac22-0502-4f8b-832c-97fea333e413 · outbound

This paper cites Unveiling linguistic regions in large language mod- els.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Unveiling linguistic regions in large language mod- els

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.303448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:a1b7db056f4517547baebf3ef690acc3dbe41c4bbded6a9f404969d051a953eb

Observation 0e73a075-a863-4f68-ac5a-18f3ac68fd2f · outbound

This paper cites Video-3d llm: Learning position-aware video representation for 3d scene understanding.arXiv.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Video-3d llm: Learning position-aware video representation for 3d scene understanding.arXiv

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.300884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:c48a0875c609341abeecbebabe23ec79a7dad1e9170b4b1f67ca7635cc37db37

Observation 1afbfa50-2e6b-4cf1-8682-475436557316 · outbound

This paper cites Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.985781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:9cb644940223d46760ef832655825cf2521356e6d187ec272cde522508d816ab

Observation d4efdad9-bc2c-491b-8bea-7585c10c65a9 · outbound

This paper cites Video-3d llm: Learning position-aware video representation for 3d scene understanding.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Video-3d llm: Learning position-aware video representation for 3d scene understanding

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.298301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:d7223a9d0df45e1d493c160a71ca27609382e5f71b6123f99f3e6ed86ecff931

Observation 30577fc6-a5f6-4ed3-9ead-8dac2f83a044 · outbound

This paper cites Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.295076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:5c8bd9a846b0302b7f214bffa9e590bb461e96a819be2ae1bc623643ce2edfa5

Observation 7e33c036-c44a-4149-b231-72d3b52bb796 · outbound

This paper cites Feature4x: Bridging any monocular video to 4d agentic ai with versatile gaussian feature fields.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Feature4x: Bridging any monocular video to 4d agentic ai with versatile gaussian feature fields

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.292195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:e4f5d7b40e7fbaa5bd896e739ed240a1327051e067b7841ba3bdc30a82c4debf

Observation 7367b6aa-65a0-4729-a323-d1b27d74bcf5 · outbound

This paper cites Vlm4d: To- wards spatiotemporal awareness in vision language models.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Vlm4d: To- wards spatiotemporal awareness in vision language models

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.289863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:70907b9753610b01ca63f19f987bf4e30226455765599ae0da11fbf8c57427d0

Observation 1cfd3034-3377-453c-950e-c1aff81858b6 · outbound

This paper cites Llava-3d: A simple yet effective pathway to empowering lmms with 3d-awareness.arXiv.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Llava-3d: A simple yet effective pathway to empowering lmms with 3d-awareness.arXiv

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.287629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:2a037438d3161cd0550b2761bbd0959dcc6bcef8a32dcc973c25b6157aede020

Observation c92deb70-9b8e-4a07-971f-6e7b8f9ad8d1 · outbound

This paper cites cut3r"). •Spatial tower feature selection:all (-spatial_tower_select_feature.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction cut3r"). •Spatial tower feature selection:all (-spatial_tower_select_feature

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:02:19.285410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:860cdb5b0a2eacc6a44e3fad40f1292bcdb7331e934e7e904937222e92ac8da7

Pith citing papers

Observation 1748d8d8-dcb6-4c53-96ee-9ac2ae4ff15a · inbound

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling cites this paper.

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:17:06.515129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-19T05:13:28.767788Z digest=sha256:bdba406a0444d1c75de2a7d813a717eb12c2ebad48da3aaf0dd52b1a97f22e73

Observation 18e16956-dcb9-415c-98e7-ffb48e60eabc · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:42:05.749042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:2e51fd0e0c7d847d15aa14865653b2a3e6fd066118baf1703937faf109a63aea

Observation 8beef7b3-9c22-424a-9346-c780f52d2a60 · inbound

POMA-3D: The Point Map Way to 3D Scene Understanding cites this paper.

POMA-3D: The Point Map Way to 3D Scene Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:30:11.587250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T20:27:27.347592Z digest=sha256:944c1447a502d1163a3aa3d431ef308753000adc42e79b4273e07728a75c00be

Observation 4b37eb10-8f7d-4afa-a7bd-7b914bb2bdd3 · inbound

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention cites this paper.

AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T06:29:09.920022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T06:28:22.652509Z digest=sha256:e205f8a14b1ca9201d039cab85218041a5094bc11bef5e75ee3c63a6644ebaab

Observation 72fc96a9-1852-4a86-a0bd-ed1dcd958dde · inbound

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images cites this paper.

MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T20:38:55.151754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:38:55.151754Z digest=sha256:e7e2ca733da5e288a0d50291b003f3ce225b84a3c49c3629b385ec2ef37e515a

Observation 21301638-439e-4870-8602-cdf4736ed86b · inbound

Vision-Language Memory for Spatial Reasoning cites this paper.

Vision-Language Memory for Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:31.896130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:31.896130Z digest=sha256:a90b82f53c50ae8d2b5bf5ed5c164275ed5ffd5b9f54162eec70535782e81ff6

Observation 8ce4afd6-147d-4bd5-84f1-9cdda824591c · inbound

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics cites this paper.

Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T16:27:29.726654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:27:29.726654Z digest=sha256:2f5b7fcb4d1cca98c2d77e23ed66aa22bef81ecb4d1b849f80c603a56220a689

Observation a1831e13-f921-45a6-bf70-d21f12765136 · inbound

4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation cites this paper.

4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-16T21:21:16.837689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T21:20:38.952416Z digest=sha256:3fb8bbb88e00e15a9c3fadcccc474c1f283ae6f3e128979d7bc3d5d560ecd2b5

Observation f3930752-0673-4b19-bf00-18b286a6e413 · inbound

SpatialMosaic: A Multiview VLM Dataset for Partial Visibility cites this paper.

SpatialMosaic: A Multiview VLM Dataset for Partial Visibility VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:43:20.635922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T19:42:25.626448Z digest=sha256:9441c464eabe2bbbc198780dee2ea333f143bd2f4be7e59f2a25469886438dcb

Observation c41cab4d-fb61-4081-a3a1-95893414ea94 · inbound

Thinking with Geometry: Active Geometry Integration for Spatial Reasoning cites this paper.

Thinking with Geometry: Active Geometry Integration for Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:40:42.249141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T06:39:29.010937Z digest=sha256:2d050852805dc27cd17dc4eebbf23b69ae573153cbd9596b5e0eb8202206710f

Observation da88302d-5f16-4054-a7fa-7860c3ea0ee1 · inbound

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models cites this paper.

From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-03T03:16:49.980712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:16:49.980712Z digest=sha256:ed34955af07301661d015ef56bc8a754e450c53049b28626e791b53607849fde

Observation 2ecc57ab-2221-4b00-8e81-684fbf08f0ed · inbound

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models cites this paper.

GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T18:06:09.863681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:06:09.863681Z digest=sha256:b3895a5f9154dea4e3ee02552c58979834f07ad1fdcb346853968294185f7e89

Observation 52691746-4362-4d09-b91e-28edaf28870f · inbound

Lifting Unlabeled Internet-level Data for 3D Scene Understanding cites this paper.

Lifting Unlabeled Internet-level Data for 3D Scene Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:18:21.027476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T22:16:57.890955Z digest=sha256:28c9caba9f6bc7052b0481d3b7fa409887ebd3dd9739b7b4e43ca96c3fb2e8b9

Observation a8a92a83-2a37-4d29-9491-ddb7fae99dd0 · inbound

Token Warping Helps MLLMs Look from Nearby Viewpoints cites this paper.

Token Warping Helps MLLMs Look from Nearby Viewpoints VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:08:17.319333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T21:07:55.062113Z digest=sha256:7939885263848c4d2c98a782d476ed0f1144b426eb54dafec89f7643ba8df68b

Observation 9540fc1d-5bca-4c5f-baa7-747a4ab62bfb · inbound

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs cites this paper.

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:43:22.849324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T22:41:09.840792Z digest=sha256:09db6fe7767ea1726f1c313aceedc31e47d59c179baeefb884b59cceb6dbefa0

Observation 6304c5ad-c398-44bc-a868-2b2cec354c16 · inbound

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs cites this paper.

EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T14:39:55.177552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T14:39:55.177552Z digest=sha256:dcf2b93f01a7b8a039a05a4726272883668da4a286d20f4973964484d9d32aa9

Observation 59ef32d2-5eae-40c0-ab84-b4d452b401e4 · inbound

Let Geometry GUIDE: Layer-wise Unrolling of Geometric Priors in Multimodal LLMs cites this paper.

Let Geometry GUIDE: Layer-wise Unrolling of Geometric Priors in Multimodal LLMs VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:50:48.937546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T19:32:59.381755Z digest=sha256:5c09b1c0c91f56c3eed8dccf72fc7c3173f26ca0bc86f4894dc21722197c1980

Observation 4e775d07-3f13-4c16-8a62-3d50c5bdac36 · inbound

OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence cites this paper.

OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T06:20:56.452510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:41:31.544716Z digest=sha256:cca60c77be6db1e025f836678daa5d265ebc06f0cf490547f7e1d8d0f61ac94d

Observation 988ed2c0-e3b9-4738-b948-e4023014ec2a · inbound

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding cites this paper.

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:51:22.687350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:25:31.097385Z digest=sha256:d126885d9c6d7203df7dbf0244a19978aad233c539ba49521b557468863a6fbd

Observation 208a0fa6-3e3a-41ec-92d0-ee4dfb2f89f6 · inbound

EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks cites this paper.

EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:26:02.915690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:10:36.209152Z digest=sha256:1b6a4eafd6965ee0af546200de19824655916e59ea12c8bed759797f1162dafa

Observation b7a3b696-8d0f-4da9-99d8-6762a4fe11bb · inbound

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents cites this paper.

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T07:51:01.840538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:55:48.506049Z digest=sha256:b9c4eaf3004ffcbcd78ff04526418883a28d3b231145a0aba5568cf79e874262

Observation 669f4dd9-308c-4e4f-b456-43437c4d6cb5 · inbound

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents cites this paper.

Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T23:07:40.699400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:07:40.699400Z digest=sha256:29f5cd39549cabf52239305e75fe39247c3037dd458ecc2f07798af58c3f7662

Observation 8cd39307-67d1-4280-a214-54ed885adc0a · inbound

FF3R: Feedforward Feature 3D Reconstruction from Unconstrained views cites this paper.

FF3R: Feedforward Feature 3D Reconstruction from Unconstrained views VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:05:59.356995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:47:52.901901Z digest=sha256:bf38cc02add90b8664c33c11078dbf12763cd52b4cb63e6ac0c11d199402fb23

Observation 5148c8c7-2edf-4964-bdd9-757a2185c3e0 · inbound

Any 3D Scene is Worth 1K Tokens: 3D-Grounded Representation for Scene Generation at Scale cites this paper.

Any 3D Scene is Worth 1K Tokens: 3D-Grounded Representation for Scene Generation at Scale VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T08:45:59.615137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:30:33.113292Z digest=sha256:6e547658e4254c3ce064a26d11b913c4c84e7ebd1199bcf67f73d22a93fa2bc3

Observation c6c57998-7704-4748-ba5c-7844ba140ca5 · inbound

SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning cites this paper.

SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-10T06:41:37.173534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T06:31:30.778309Z digest=sha256:ee97637fdf5b870675062c06be2a7b67075eec4cc0ab49c9a55225c7d6b1ff1f

Observation abf82160-1912-4bee-bb4f-9d665413e1bc · inbound

$M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills cites this paper.

$M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-11T22:11:15.955057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T03:13:36.437080Z digest=sha256:858711a5fbf77c7be4ef2778d493acc2a488e3d50c15a69fb4751ff493ab1656

Observation 7d95985a-5bfc-45e6-9901-549d7b439c0e · inbound

$M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills cites this paper.

$M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T18:17:43.564400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T18:17:43.564400Z digest=sha256:8bedb19d650302a14da5cff1113746bd91eacd74bd3e30d9522f8243b0d26453

Observation d56513e2-ee0b-438f-95ff-6b0c9f902ef0 · inbound

From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs cites this paper.

From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-11T16:26:07.551530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T17:01:38.481471Z digest=sha256:54d9ca6c2e686aa08d8c6741a28b1f60efde11dc5c91112bac06bde1f5765369

Observation 6ba8aabf-f7a2-4852-ae40-767ea41d21eb · inbound

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding cites this paper.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:11.622389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T14:20:08.404090Z digest=sha256:feec4bf3eb6fe63d427cf1d64ebfba21a7782992f78d6601b6c07a63a6cead83

Observation 7b1db54e-df36-495e-8276-412ce2a58024 · inbound

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding cites this paper.

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-25T06:16:39.987125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-25T06:15:33.062980Z digest=sha256:b15f8b85b67eedcfde15c7b3ae4d8412bfe33e55897eac9403c2e04a24f37965

Observation 9daa15a9-0708-4dc3-a37e-eee7b909be5b · inbound

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models cites this paper.

ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:46:36.484052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:00:23.681682Z digest=sha256:6f09a03e1a95100507b1b0459b8d49f1d615bab8890a075073b3966e318d4ab4

Observation 4b6c8d52-2ab3-43dc-a562-34b6ec611bd7 · inbound

Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence cites this paper.

Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:26:19.434881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T03:24:41.877312Z digest=sha256:a2ec859383448721c05585981667cfb0c45274cb2120085acfe8849f699ba794

Observation cb433cfb-8682-4833-960d-027191258f21 · inbound

Unlocking Dense Metric Depth Estimation in VLMs cites this paper.

Unlocking Dense Metric Depth Estimation in VLMs VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T19:23:40.983094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T19:20:04.468206Z digest=sha256:26eb5a6870884f075e3185a579c99948579805282e3480dccd68c31f803cf18f

Observation a871081d-438e-475e-83b5-81e377616e16 · inbound

Unlocking Dense Metric Depth Estimation in VLMs cites this paper.

Unlocking Dense Metric Depth Estimation in VLMs VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-21T07:59:51.135163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T07:54:52.926995Z digest=sha256:38a6aedaca465c8919f5c33bf29db67e76dcf0b99170d2725aa3765f10ca1103

Observation 2af36fcf-4064-4925-a018-20e654f26702 · inbound

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop cites this paper.

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T10:53:13.231273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T10:52:22.778489Z digest=sha256:b958c1c2e939c83469da7113b88dd8f48c53439400c3636824911e8ae7d1f024

Observation 78c983b0-1385-49ef-8794-399535244996 · inbound

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop cites this paper.

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T15:05:47.166596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T18:25:17.831116Z digest=sha256:ccb1e4e21593db84e1fe3f0fc2af9eed96cd24bbc1e0503a849de56ab94c4ea9

Observation 9a01ac1f-6c16-4883-8bf5-7a95a003cd18 · inbound

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs cites this paper.

Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T06:33:05.687991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T06:31:04.432486Z digest=sha256:7971f980f4ef65f5a001acea6cda30edf4e2b8b9ec685aa93641b4f321007661

Observation de6b68d9-e1e1-43f7-97a1-4e02afaee201 · inbound

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models cites this paper.

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T05:28:04.482966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T05:27:30.938311Z digest=sha256:341921b4e99daf68a8e3bc112069cf668ff078d428292c8a0fe2b75dd0aaca96

Observation cd044ad7-e6bb-41fe-869f-895663ee4ec4 · inbound

GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning cites this paper.

GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-22T07:24:43.013284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T07:23:53.811794Z digest=sha256:a7c3cb6dea35a35882a99172e40048e69c07695be67d0b87d184e6e2f2f1f2bf

Observation 0d1b0f23-44eb-4190-aae2-67b781e2cf15 · inbound

FOUND-IT: Foundation-model-first Task-driven 3D Scene Graphs with Granularity on Demand cites this paper.

FOUND-IT: Foundation-model-first Task-driven 3D Scene Graphs with Granularity on Demand VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:14:00.409776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:04:21.666471Z digest=sha256:34276f42852c2b9d2e0ac67a140e8079705522397a861881cbbffd0cf0166f19

Observation 65344bf0-5fc9-41cb-a2fb-ba53442bd7f1 · inbound

ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs cites this paper.

ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:23:59.822765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:23:38.178195Z digest=sha256:6606860bde0ab200f8cba8f5b34efd1b23cd07a0df4019b5241334e625b5298b

Observation fc22489d-b3ad-44fa-b356-3890aac17da1 · inbound

Rethinking VLM Representation for VLA Initialization cites this paper.

Rethinking VLM Representation for VLA Initialization VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-29T22:24:00.117539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T22:21:29.733181Z digest=sha256:53893ba79f88bed84c4fdca190bceafca9047ca1b22c29281dba3527509d9333

Observation 02ae9c95-43bc-457a-9ec5-5106e037d117 · inbound

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning cites this paper.

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:51.003446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T18:15:28.261185Z digest=sha256:55d277d5f0b5139b42d0ded77701ff6adaef7cfc5a58e734c050034a4913e9c5

Observation 410835b6-7d69-46ff-b0b7-0d3755c92003 · inbound

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning cites this paper.

Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T15:54:12.909974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T15:54:12.909974Z digest=sha256:fd2a77c91e5bbac8aab6cfda13f281d71318e527aceb7c537668057b4ecc3773

Observation 5ad13a23-9cba-4100-a90f-dd043cf2f2db · inbound

GEM: Generative Supervision Helps Embodied Intelligence cites this paper.

GEM: Generative Supervision Helps Embodied Intelligence VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:43:28.785956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:38:27.263726Z digest=sha256:77e2d4ebe1222522931eed3b10f03c98d613d6fe1421a523e0b31e90a58ed5d5

Observation 836f2d3b-50e8-4b3f-9ee7-e0b3200285d5 · inbound

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning cites this paper.

Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:53:13.591518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:47:52.739735Z digest=sha256:403753913645408c84de4be077069d706c8840fe7136afcc00a88e6d944b51c0

Observation a5efadf7-31ca-43a6-8bbf-9a9b7633b7db · inbound

VLM3: Vision Language Models Are Native 3D Learners cites this paper.

VLM3: Vision Language Models Are Native 3D Learners VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-29T07:53:14.035073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:45:31.978215Z digest=sha256:696b66417c6f0e4ddad4135350758ea72fc432797d1e0746813c82484892f969

Observation 9d4695e3-ec45-4763-b24e-d4daa812bde7 · inbound

Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning cites this paper.

Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-01T20:46:13.705434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T17:46:21.821764Z digest=sha256:ea4fe29a12ee2a003d374bf055f6f6816c6443fd4f5fbd5aaeb3fe051e202d4f

Observation 172430c8-6156-4f93-b18f-16d0396fe560 · inbound

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video cites this paper.

LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T12:16:57.284944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T02:16:25.730555Z digest=sha256:c4fbc0039f610f7fc13aaab87e54ac51a915fb79570887e73db344d8b45f8001

Observation 97f70fd8-d53f-49d8-95d7-7d4ff37a028d · inbound

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models cites this paper.

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:16:57.501253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T02:13:53.219095Z digest=sha256:71eccf8d506a70c4b7af2eea793b45c925e86bba851600e3d2548dcabf43d44f

Observation 2ad5078d-463c-4a61-8820-64f322518b5e · inbound

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators cites this paper.

Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.185660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T02:25:30.998989Z digest=sha256:1faf9d58a456b686846c515a6a91f99bd008a03ea16abcf2ff437d9b1aff92ac

Observation c8b7796c-cea6-4bb4-a106-b26f1bb6a626 · inbound

Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors cites this paper.

Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T16:17:09.596982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T22:49:02.846428Z digest=sha256:a23629b4d3cd816361ce405b9a1e00da15433e4cc0697c2b98ada09d2af999e0

Observation 49f4cc9a-3a38-4f48-a4cc-9fbf10a12dba · inbound

MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model cites this paper.

MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-02T21:57:25.759560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T19:24:59.678266Z digest=sha256:0aa019634acc0d13ebf3f508b36545b01720f179c2f812e9c77aaab63978da66

Observation 005223eb-9046-492a-bdcc-9e2bad9af23d · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-03T06:07:41.198931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T12:55:47.754632Z digest=sha256:8de2a12cf9a1eaddaa51ddc9e9f1b2f7c1dad19254ba7fcbbb2ee873b3cde0b0

Observation 96c73d22-5ff5-409a-b42d-5c9ea1ff33c1 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-06-27T13:10:56.965482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T12:55:47.754632Z digest=sha256:2c2620d28c12485fcbd6688817ffcda9415929a4fc50e5fe34f581c5201d6bfe

Observation 769b4eb0-af1d-4d83-9d63-b66266a758b5 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T18:07:09.018997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T18:07:09.018997Z digest=sha256:1c71765cc4f13ccb84c635928637092fd0aee7342d3aa7cb52354b3866e8960b

Observation ebd187e6-062b-4691-82ef-abcb8fe59e2c · inbound

4DP-QA: Scalable QA for 4D Perception in Vision Language Models cites this paper.

4DP-QA: Scalable QA for 4D Perception in Vision Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T08:17:45.089423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T10:56:34.577237Z digest=sha256:bc7c2fe957ff294dff7d8fcc491bb926723b7f500391f43fada4caca54fc0661

Observation 176e599f-f951-4c2f-b405-889d38af0ada · inbound

Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning cites this paper.

Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-07-03T09:17:48.664257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T10:26:51.429436Z digest=sha256:83c55475d611e497c67d9af2cbdb2ef0fecde86edc94222d3f9fe51bdf12e521

Observation 4736f1a5-253d-44b8-aeaa-f8ed487e9985 · inbound

Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning cites this paper.

Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T11:51:33.410398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:51:33.410398Z digest=sha256:2979d6224fd717428d1e704c7b02e5e2735c6a3599d9c8b75b6c3bf157d852c8

Observation ecf77e8a-a6be-4908-aff0-fa8ba5799aed · inbound

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views cites this paper.

Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:39:45.344628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T08:38:46.044079Z digest=sha256:0acbf163ca699ab477ff4beeb2d78f9250d98a5739c4e0b445c6bb08e2604f66

Observation 1520916c-287b-4dee-8ae8-afd507451ec0 · inbound

HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory cites this paper.

HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-04T10:49:46.507869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T08:27:15.889885Z digest=sha256:d707f2afda79ef022b700554e7e5324bb0dbc9b03e71b286c3919595d0e65a1c

Observation 019ffe8c-56e7-4c01-8659-fb28a604c152 · inbound

Natural Language Camera Movement Understanding cites this paper.

Natural Language Camera Movement Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T05:16:19.750976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:16:19.750976Z digest=sha256:32d0d631295996489364a3729952a0b47ab5d50b12a978ee0ef5980f76d74cb4

Observation 29d6508d-02ce-49ee-b531-beff7c611609 · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 157

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:979ba445a50dad299aaa8313d43c6d85cb64733a6e16a685eca96cc2eb4d1d7d

Observation b5130446-0bca-45bf-bbb3-b20c569965ec · inbound

GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding cites this paper.

GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T05:13:00.188880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:13:00.188880Z digest=sha256:d5527474f25e0278b36a94a2a0a60fe70bf73054a132e9a78b2c6209500edc4e

Observation 3a027346-4e82-4a16-966d-da6358e1aa4e · inbound

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding cites this paper.

Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T00:23:01.277839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:23:01.277839Z digest=sha256:8ca384bb75de1bc10346072944274658cd01ed57c047167695625e514ad83c8a

Observation 16e7b789-74c0-470a-b4fb-88a0ae418c67 · inbound

LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning cites this paper.

LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T18:40:13.176413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:40:13.176413Z digest=sha256:ae9d13150dba6114ed6980fd9a8588fcd5bd362f39c433f0239c5cce1baacd3c

Observation 3fa0a7a7-c2fc-46e4-9139-d2db48040c5a · inbound

ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning cites this paper.

ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T17:36:37.411186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:36:37.411186Z digest=sha256:2c28634e2c0fd4e9668f26fef6185cd945c022c325954cea0d75bc9caa5688d0

Observation 88ca5a98-8e67-494d-91a3-64c1160421e3 · inbound

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model cites this paper.

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T16:32:47.404785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:32:47.404785Z digest=sha256:af201f96caf72d725bbf31c0d4af9233a434876953cd1919b8c29c88d4a8447b

Observation c9f46a6b-c319-4bc8-8a57-883213c90388 · inbound

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model cites this paper.

RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T01:57:28.855191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:57:28.855191Z digest=sha256:5aeeae8c9884ba3ff551c0a8188252622fb9eaf0158819f87d302d3830c06d69

Observation b57eb1c8-c706-43c6-9117-8cb5704b1a7c · inbound

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? cites this paper.

ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T09:09:24.856306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:09:24.856306Z digest=sha256:597e6f18d0ca5ad23f7dac326cbf845c2b37540164038a68d35ce0ae947a22d0

Observation 21c8bbc7-3954-478e-bafd-bb833e4922bd · inbound

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text cites this paper.

Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T08:39:37.227319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:39:37.227319Z digest=sha256:3486d42295f9c28919914558ad60c817c92a6fdb0389e6fcb3857e986f3c2352