Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T12:54:01.013242Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 71 inbound Pith citation observations for arXiv:2505.20279.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T12:54:01.013242Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T20:38:55.151754Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
95 of 95 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 288cff19-520e-4cfa-b0d2-7c93e1465652 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d2fc3e26-dbaa-4493-a5a5-8f78ac1cf3e8 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Specificity of learning: Why infants fall over a veritable cliff.Psychological Science, 11(4):290–295
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e7e5e756-a79e-4fb7-b8ed-ef4fbfbbfd6b · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Flamingo: a visual language model for few-shot learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 639cbea1-0a31-4838-af80-3f66b605261b · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Scanqa: 3d question answering for spatial scene understanding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6ce9edef-0c4f-464a-9457-60fab7defb66 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Qwen Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation af1f33da-af6b-41b8-8518-180d7e997bc4 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f4be9339-441c-49d9-a6de-86ce00b5b57c · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Qwen2.5-VL Technical Report
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7df4d349-4fbd-4a94-b5af-b136e2a2bd04 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ed22b408-5285-4d4d-a914-e5c3bcf60c0c · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Large Linguistic Models: Investigating LLMs' metalinguistic abilities
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 982e7e28-d815-48d4-9b1e-1abb639db716 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Language models are few-shot learners.NeurIPS
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation db9b0fdf-6b0e-419c-98d0-760af2ad7471 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 76bdf84f-f44b-44b6-bead-d5983a2b7160 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4a38ee90-8e7a-4a1b-be44-bb4fedba1558 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Longvila: Scaling long-context vi- sual language models for long videos.arXiv
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0ec60a93-f700-414b-b19c-533701813241 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fc7d53ed-46b1-451f-b050-68c2c4a909f5 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6180e137-9b41-42aa-b94d-628554fc5bda · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Spatial- rgpt: Grounded spatial reasoning in vision language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2c7c9a71-c1c5-4a81-b5bd-50c603d98855 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Scannet: Richly-annotated 3d reconstructions of indoor scenes
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e2097ed8-b3b4-4820-8212-72a6f32c502a · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction 3d-llava: Towards generalist 3d lmms with omni superpoint transformer.arXiv
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 26ba6bdf-b15c-4173-86e0-16b7abba719e · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Palm-e: An embodied multimodal language model.arxiv
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 91cc8df9-50c0-4aff-adee-bdc9dd7bf212 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Large spatial model: End-to-end unposed images to semantic 3d.Advances in neural information processing systems, 37:40212–40229
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7c163f6c-01ff-40f3-8fbc-9329d33b7fa5 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f35a9cfc-7b31-49b3-85b6-8d2168020559 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ddd3158c-4c61-43b6-bba2-fb0f62201ce4 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b22ee976-8d20-40a3-be7e-7adab94fe5e0 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation eef2d304-8096-4309-a9d5-7da150b0f885 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Cascade cost volume for high-resolution 9 multi-view stereo and stereo matching
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b56cf3a5-89bd-42f3-b48f-6e0a5385f393 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Cambridge university press
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e7b214bc-6fd8-4003-9ebc-9130ae929dcd · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Individual differences in spatial abilities.The Cambridge handbook of visuospatial thinking, pages 121–169
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 85ac9295-25a7-4836-aacc-09ec91d410e5 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Measuring Massive Multitask Language Understanding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b959b21d-cfcc-4970-a531-70beba8b3ee1 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction 3d-llm: Inject- ing the 3d world into large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8956cff6-dc7b-4017-8ecf-7489da190266 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Multiply: A multisensory object- centric embodied large language model in 3d world
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0b072315-f3a0-41d2-80b2-c8005b6f7366 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Lora: Low-rank adaptation of large language models.ICLR, 1(2):3
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a36efbde-83f7-4cc1-945c-3f65fe7e747c · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Chat-scene: Bridging 3d scene and large language models with object identifiers
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7c3fa2df-ace9-4161-b619-cff13a410cf1 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Think- ing in dynamics: How multimodal large language models perceive, track, and reason dynamics in physical 4d world
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 13b7a12b-88d9-40d8-9b2c-d6b87400b72e · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Gpt-4o system card
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 176b21c4-a32a-48e2-82cd-5929fbabfabf · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction GPT-4o System Card
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2c14db9d-9c5f-49d7-afa7-dfbbafa5de74 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Scaling up visual and vision-language representation learning with noisy text supervision
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2397de9d-ad86-46d4-bfd0-59f3ea0cdc34 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Language models with rationality
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 92002432-7ab6-4ae7-9617-0f8d134f234f · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Ground- ing image matching in 3d with mast3r
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e1fb0817-cae5-436a-b16c-0832312d0733 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction LLaVA-OneVision: Easy Visual Task Transfer
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a0edd5e1-3abc-409c-9775-a45302aaebb2 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Llava-onevision: Easy visual task transfer
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a92127dc-e1d9-4520-a9f7-5fd2aafb18a4 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c4ae958-d758-4598-99ea-ad5c82545015 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 3b714487-2121-435b-a022-cf9295b61b17 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Vila: On pre-training for visual language models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4f9fc90f-956d-4066-bb00-6a7ff5f040e0 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Ost-bench: Evaluating the capabilities of mllms in online spatio-temporal scene understanding
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4946b5cd-c496-47bd-9de3-29f191b7d7eb · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Visual instruction tuning.NeurIPS
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c77b11ea-bf67-4cfe-b28a-8fe3ed3e4a48 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c2108898-ca17-4368-9202-129e1c1152ea · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Unified-io: A unified model for vision, language, and multi-modal tasks.arXiv
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c627b4dc-78c6-4e8b-a85c-543e476b10c9 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction SQA3D: Situated Question Answering in 3D Scenes
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f927bbe9-841f-456f-bae3-1d998d683bdc · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Openeqa: Embodied question answering in the era of foun- dation models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9d43534f-83a9-4cf0-bd8e-23df09169d75 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Spa- tiallm: Training large language models for structured in- door modeling
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b5736758-3f31-4fac-a9f7-e3333cedabde · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction How are the locations of objects in the environment represented in memory? InInternational conference on spatial cognition, pages 174–191
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 19b0647b-b903-4de6-877a-00fbe927077d · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Individual differences in navigation: an introductory overview.Prime archives in psychology
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 283d2cf4-5cc5-4656-8047-51b32808da02 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Mast3r-slam: Real-time dense slam with 3d reconstruction 10 priors
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c43ae4d9-456b-4dae-a9f5-b21879f2b106 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction A Comprehensive Overview of Large Language Models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f66ca052-d8d4-409a-8fbb-064208529145 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Kosmos-2: Grounding multimodal large language models to the world.arXiv
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f580f4c9-297d-4e79-9d3e-64ccb0abde5f · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Improving language understanding by gener- ative pre-training.OpenAI Blog
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 06d58fa1-eea7-43f7-92a0-5d7737e15429 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Language models are unsu- pervised multitask learners.OpenAI blog, 1(8):9
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1be725fc-8a8f-42d8-b96d-a4ba772a9216 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Learn- ing transferable visual models from natural language super- vision
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f9af211a-6592-49ff-9c64-2a7ff77b0f6e · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Habitat: A platform for embodied ai research
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ab55bd1e-d896-4e6b-be9c-7d65a3ea14a1 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Structure- from-motion revisited
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d76d7748-9158-4771-b9a3-69c86b2fb3be · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dcb81cea-be14-4721-9888-3d547935b109 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Gemini: A Family of Highly Capable Multimodal Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a1ec1067-4ff1-4553-917d-b4d1fd4a3d83 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Gemini 1.5: Unlocking multimodal understanding across millions of tokens of con- text.arXiv
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2d0dd26b-8296-47f4-9ccb-fc0c0eba4d91 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction LLaMA: Open and Efficient Foundation Language Models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a48a28e7-5e0c-45b7-b853-3485cc1dc09a · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a08a204b-1749-46f7-adf6-aac496c39be1 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 03683f5f-842f-4921-9417-097cfe935b31 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction 3D Reconstruction with Spatial Memory
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8bfcbaab-93bb-47aa-a2c7-679f23797e2e · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Ross3D: Reconstructive Visual Instruction Tuning with 3D-Awareness
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2f1d1b15-0de5-45ac-8c0a-daec510929df · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction VGGT: Visual Geometry Grounded Transformer
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5d5fcb24-c093-4280-a9a8-5134bfc89e16 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction A survey on large language model based au- tonomous agents.Frontiers of Computer Science, 18(6): 186345
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6867501f-769a-4719-a96c-8430b0beb986 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Ofa: Unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning framework
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7f4afd21-f9f3-41f2-9cd6-eedb231e61ca · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Continuous 3D Perception Model with Persistent State
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 62118cfe-8cf9-414d-a5b3-bf7c7510b21f · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Dust3r: Geometric 3d vi- sion made easy
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b87f94a7-a7f7-455e-a5ca-bb963b40f335 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Emergent abilities of large language models.TMLR
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 464c7c19-7436-4453-817f-9704be2f3139 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Dynamicverse: A physically- aware multimodal framework for 4d world modeling
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 075cb424-c74a-4abc-b2a4-31ba790d8c5d · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 381d89a9-f92c-435a-80ea-61ed73fb442d · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f2f7711f-1566-4c1b-ae81-9b1467a86c9f · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7c0e173d-c1f5-406a-a593-ae4c0d7b5607 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Mvsnet: Depth inference for unstructured multi-view stereo
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0ec9e627-5adf-4b66-b74c-86fbda9ac8b7 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Scannet++: A high-fidelity dataset of 3d in- door scenes
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 04289116-dcf4-4568-9715-6b70ce94f633 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for ex- pert agi
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2bdd7e1f-26f0-429c-95fe-e645d9aab789 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation af60ba3e-af7f-46ec-887c-445a66aa29be · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Spatialstack: Layered geometry-language fusion for 3d vlm spatial reasoning
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0ca71d35-d57a-46c1-b31b-74d6af952ada · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Long context transfer from language to vision.arXiv
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cc288160-2790-4e00-b581-d2b03553eec8 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Llava- next: A strong zero-shot video understanding model
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9f3976c4-30c8-426f-a89b-b2cd0d54fcb2 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7fc8ac22-0502-4f8b-832c-97fea333e413 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Unveiling linguistic regions in large language mod- els
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0e73a075-a863-4f68-ac5a-18f3ac68fd2f · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Video-3d llm: Learning position-aware video representation for 3d scene understanding.arXiv
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1afbfa50-2e6b-4cf1-8682-475436557316 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d4efdad9-bc2c-491b-8bea-7585c10c65a9 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Video-3d llm: Learning position-aware video representation for 3d scene understanding
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 30577fc6-a5f6-4ed3-9ead-8dac2f83a044 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Feature 3dgs: Supercharging 3d gaussian splatting to enable distilled feature fields
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7e33c036-c44a-4149-b231-72d3b52bb796 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Feature4x: Bridging any monocular video to 4d agentic ai with versatile gaussian feature fields
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7367b6aa-65a0-4729-a323-d1b27d74bcf5 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Vlm4d: To- wards spatiotemporal awareness in vision language models
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1cfd3034-3377-453c-950e-c1aff81858b6 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction Llava-3d: A simple yet effective pathway to empowering lmms with 3d-awareness.arXiv
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c92deb70-9b8e-4a07-971f-6e7b8f9ad8d1 · outbound
VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction cut3r"). •Spatial tower feature selection:all (-spatial_tower_select_feature
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1748d8d8-dcb6-4c53-96ee-9ac2ae4ff15a · inbound
Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 18e16956-dcb9-415c-98e7-ffb48e60eabc · inbound
MiMo-Embodied: X-Embodied Foundation Model Technical Report VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 8beef7b3-9c22-424a-9346-c780f52d2a60 · inbound
POMA-3D: The Point Map Way to 3D Scene Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4b37eb10-8f7d-4afa-a7bd-7b914bb2bdd3 · inbound
AVA-VLA: Improving Vision-Language-Action models with Active Visual Attention VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 72fc96a9-1852-4a86-a0bd-ed1dcd958dde · inbound
MonoSR: Open-Vocabulary Spatial Reasoning from Monocular Images VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21301638-439e-4870-8602-cdf4736ed86b · inbound
Vision-Language Memory for Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ce4afd6-147d-4bd5-84f1-9cdda824591c · inbound
Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1831e13-f921-45a6-bf70-d21f12765136 · inbound
4D-RGPT: Toward Region-level 4D Understanding via Perceptual Distillation VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation f3930752-0673-4b19-bf00-18b286a6e413 · inbound
SpatialMosaic: A Multiview VLM Dataset for Partial Visibility VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c41cab4d-fb61-4081-a3a1-95893414ea94 · inbound
Thinking with Geometry: Active Geometry Integration for Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation da88302d-5f16-4054-a7fa-7860c3ea0ee1 · inbound
From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 1997
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ecc57ab-2221-4b00-8e81-684fbf08f0ed · inbound
GAP-MLLM: Geometry-Aligned Pre-training for Activating 3D Spatial Perception in Multimodal Large Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52691746-4362-4d09-b91e-28edaf28870f · inbound
Lifting Unlabeled Internet-level Data for 3D Scene Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a8a92a83-2a37-4d29-9491-ddb7fae99dd0 · inbound
Token Warping Helps MLLMs Look from Nearby Viewpoints VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9540fc1d-5bca-4c5f-baa7-747a4ab62bfb · inbound
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6304c5ad-c398-44bc-a868-2b2cec354c16 · inbound
EgoMind: Activating Spatial Cognition through Linguistic Reasoning in MLLMs VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59ef32d2-5eae-40c0-ab84-b4d452b401e4 · inbound
Let Geometry GUIDE: Layer-wise Unrolling of Geometric Priors in Multimodal LLMs VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4e775d07-3f13-4c16-8a62-3d50c5bdac36 · inbound
OpenSpatial: A Principled Data Engine for Empowering Spatial Intelligence VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 988ed2c0-e3b9-4738-b948-e4023014ec2a · inbound
MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 208a0fa6-3e3a-41ec-92d0-ee4dfb2f89f6 · inbound
EgoTL: Egocentric Think-Aloud Chains for Long-Horizon Tasks VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation b7a3b696-8d0f-4da9-99d8-6762a4fe11bb · inbound
Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 669f4dd9-308c-4e4f-b456-43437c4d6cb5 · inbound
Text-Guided 6D Object Pose Rearrangement via Closed-Loop VLM Agents VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cd39307-67d1-4280-a214-54ed885adc0a · inbound
FF3R: Feedforward Feature 3D Reconstruction from Unconstrained views VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5148c8c7-2edf-4964-bdd9-757a2185c3e0 · inbound
Any 3D Scene is Worth 1K Tokens: 3D-Grounded Representation for Scene Generation at Scale VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c6c57998-7704-4748-ba5c-7844ba140ca5 · inbound
SpatialImaginer: Towards Adaptive Visual Imagination for Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation abf82160-1912-4bee-bb4f-9d665413e1bc · inbound
$M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7d95985a-5bfc-45e6-9901-549d7b439c0e · inbound
$M^2$-VLA: Boosting Vision-Language Models for Generalizable Manipulation via Layer Mixture and Meta-Skills VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d56513e2-ee0b-438f-95ff-6b0c9f902ef0 · inbound
From Where Things Are to What They Are For: Benchmarking Spatial-Functional Intelligence in Multimodal LLMs VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 6ba8aabf-f7a2-4852-ae40-767ea41d21eb · inbound
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7b1db54e-df36-495e-8276-412ce2a58024 · inbound
4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9daa15a9-0708-4dc3-a37e-eee7b909be5b · inbound
ViSRA: A Video-based Spatial Reasoning Agent for Multi-modal Large Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4b6c8d52-2ab3-43dc-a562-34b6ec611bd7 · inbound
Thinking with Novel Views: A Systematic Analysis of Generative-Augmented Spatial Intelligence VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cb433cfb-8682-4833-960d-027191258f21 · inbound
Unlocking Dense Metric Depth Estimation in VLMs VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a871081d-438e-475e-83b5-81e377616e16 · inbound
Unlocking Dense Metric Depth Estimation in VLMs VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2af36fcf-4064-4925-a018-20e654f26702 · inbound
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 78c983b0-1385-49ef-8794-399535244996 · inbound
ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9a01ac1f-6c16-4883-8bf5-7a95a003cd18 · inbound
Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation de6b68d9-e1e1-43f7-97a1-4e02afaee201 · inbound
CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cd044ad7-e6bb-41fe-869f-895663ee4ec4 · inbound
GeoWeaver: Grounding Visual Tokens with Geometric Evidence before Scene Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0d1b0f23-44eb-4190-aae2-67b781e2cf15 · inbound
FOUND-IT: Foundation-model-first Task-driven 3D Scene Graphs with Granularity on Demand VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 65344bf0-5fc9-41cb-a2fb-ba53442bd7f1 · inbound
ProSR: Process-Shaped Spatial Reasoning for Reliable Chain-of-Thought in VLMs VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fc22489d-b3ad-44fa-b356-3890aac17da1 · inbound
Rethinking VLM Representation for VLA Initialization VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 02ae9c95-43bc-457a-9ec5-5106e037d117 · inbound
Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 410835b6-7d69-46ff-b0b7-0d3755c92003 · inbound
Q-GeoMem: Question-Guided Geometric Memory for Video Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ad13a23-9cba-4100-a90f-dd043cf2f2db · inbound
GEM: Generative Supervision Helps Embodied Intelligence VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 836f2d3b-50e8-4b3f-9ee7-e0b3200285d5 · inbound
Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation a5efadf7-31ca-43a6-8bbf-9a9b7633b7db · inbound
VLM3: Vision Language Models Are Native 3D Learners VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9d4695e3-ec45-4763-b24e-d4daa812bde7 · inbound
Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 172430c8-6156-4f93-b18f-16d0396fe560 · inbound
LongSpace: Exploring Long-Horizon Spatial Memory from Perception to Recall in Video VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 97f70fd8-d53f-49d8-95d7-7d4ff37a028d · inbound
Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 2ad5078d-463c-4a61-8820-64f322518b5e · inbound
Thinking with Imagination: Agentic Visual Spatial Reasoning with World Simulators VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation c8b7796c-cea6-4bb4-a106-b26f1bb6a626 · inbound
Stream3D-VLM: Online 3D Spatial Understanding with Incremental Geometry Priors VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 49f4cc9a-3a38-4f48-a4cc-9fbf10a12dba · inbound
MotionVLA: Injecting Geometric Motion into Vision-Language-Action Model VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 005223eb-9046-492a-bdcc-9e2bad9af23d · inbound
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 96c73d22-5ff5-409a-b42d-5c9ea1ff33c1 · inbound
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 769b4eb0-af1d-4d83-9d63-b66266a758b5 · inbound
Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebd187e6-062b-4691-82ef-abcb8fe59e2c · inbound
4DP-QA: Scalable QA for 4D Perception in Vision Language Models VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 176e599f-f951-4c2f-b405-889d38af0ada · inbound
Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4736f1a5-253d-44b8-aeaa-f8ed487e9985 · inbound
Ouroboros-Spatial: Closing the Data-Model Loop for Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecf77e8a-a6be-4908-aff0-fa8ba5799aed · inbound
Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1520916c-287b-4dee-8ae8-afd507451ec0 · inbound
HoloAgent-0: A Unified Embodied Agent Framework with 3D Spatial Memory VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 019ffe8c-56e7-4c01-8659-fb28a604c152 · inbound
Natural Language Camera Movement Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29d6508d-02ce-49ee-b531-beff7c611609 · inbound
ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 157
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5130446-0bca-45bf-bbb3-b20c569965ec · inbound
GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a027346-4e82-4a16-966d-da6358e1aa4e · inbound
Beyond Single Expert: Harmonizing Diverse Visual Priors in MLLMs for Spatial Understanding VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16e7b789-74c0-470a-b4fb-88a0ae418c67 · inbound
LenGuard-GPC: Length Guarding with Guided-Prompt Consistency for Spatial Reasoning Reinforce Learning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fa0a7a7-c2fc-46e4-9139-d2db48040c5a · inbound
ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88ca5a98-8e67-494d-91a3-64c1160421e3 · inbound
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9f46a6b-c319-4bc8-8a57-883213c90388 · inbound
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b57eb1c8-c706-43c6-9117-8cb5704b1a7c · inbound
ViSTR-Bench: Can MLLMs Reason from Continuous Visual Cues in Dynamic Scenes? VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21c8bbc7-3954-478e-bafd-bb833e4922bd · inbound
Show, Don't Tell: Evaluating Spatial Cognition in Generative Pixels Rather Than LLM Text VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.