Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T04:27:29.029373Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2608.02980.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T04:27:29.029373Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation aeb6dbff-f449-4e49-b91a-9442bef094cc · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding ReferIt3D: Neural Listen- ers for Fine-Grained 3D Object Identification in Real-World Scenes
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1d6a1d8e-9086-4208-94a4-ffcdb6f39801 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Llama 3: The llama-3 herd of models.https: //ai.meta.com/llama/, 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2705b699-df6a-4def-ba9e-90ca31a5b5a5 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Locate 3d: Real-world ob- ject localization via self-supervised learning in 3d, 2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bf16414f-d807-45ae-885d-41f2b0be43cc · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scanqa: 3d question answering for spatial scene understanding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a53fa677-6872-40d5-97c4-0fd172862f17 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 07d077cc-e702-40ff-b6d3-9e1d824d8e99 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Token merging: Your ViT but faster
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0c2c2a72-6b94-49ed-823a-92146925c91d · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding From thousands to billions: 3d visual language grounding via render-supervised distillation from 2d vlms, 2025
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 558c4219-0243-4c7c-a0e5-5490ec30a8e7 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding End- to-End Object Detection with Transformers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ec62c190-df38-4cbf-8493-7fc304c59ed1 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Matterport3D: Learning from RGB-D Data in Indoor Environments
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bb4ba43-8281-40f9-b5dc-512086bd13d1 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding ScanRefer: 3D Object Localization in RGB-D Scans using Natural Language
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ca779a4b-4832-461b-8f06-d871b5697ef7 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Ll3da: Visual interactive instruction tuning for omni-3d understanding reasoning and planning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 619f4d38-021e-4b34-a0e4-bcd3f7a9c974 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Grounded 3D-LLM with Referent Tokens
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7196c203-57d3-4550-bb01-9805a5c02bfb · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Schwing, Alexan- der Kirillov, and Rohit Girdhar
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ff2b045f-ce34-4731-8b73-0618518d85e3 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scannet: Richly-annotated 3d reconstructions of indoor scenes
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e479c723-70a6-4685-a861-2e2752513adf · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface re-integration.ACM Transactions on Graphics 2017 (TOG),
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 335a98b9-e5cf-4641-8d30-f44a2c5d3c9b · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d43c921-b730-4e1a-a0e7-c6ff566238d8 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3d-llm: In- jecting the 3d world into large language models.Advances in Neural Information Processing Systems, 36:20482–20494,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bcbb5c7-fc2b-41e5-a0ee-43312f09a4b7 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0639deb9-e224-41af-832e-f8b83e7fa876 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Chat-scene: Bridging 3d scene and large language models with object identifiers
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 826b1c76-0114-4b09-a29b-1337a1ccb579 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding An Embodied Generalist Agent in 3D World
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da6b087d-9517-436b-b3b5-7d670b149057 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Revisiting multimodal positional encoding in vision-language models, 2026
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dc93f922-558f-492f-95c5-56de9aac65c1 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Reason3d: Searching and reasoning 3d segmentation via large language model.3DV, 2025
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6d18d69e-bbb6-403e-a26e-21a3aeeb3ec8 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Bottom up top down detection transform- ers for language grounding in images and point clouds
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 00114ae5-a656-48af-af38-86faf6e5ad96 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Odin: A single model for 2d and 3d segmentation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1f15f975-3b82-4750-8445-9ef82ff9df72 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unifying 2d and 3d vision-language un- derstanding, 2025
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6594ca3a-8222-4c0b-9389-eafbf8bf2361 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding MDETR - Modulated Detection for End-to-End Multi-Modal Under- standing
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 54add548-3d69-4b75-b0ed-14a7a4c44850 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding ReferItGame: Referring to objects in pho- tographs of natural scenes
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4df9d2ee-444b-4b06-8fcf-80a493fc90c8 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Restr: Convolution-free referring image segmentation using transformers, 2022
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7a97f72d-02bf-433a-9bf6-2406accc3ba8 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Mask-attention-free transformer for 3d in- stance segmentation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fbc27d2-3930-4d67-b19e-c3dcb341710d · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Lisa: Reasoning segmenta- tion via large language model, 2024
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2bc1063e-20fe-48e6-a533-28a0ba3899a5 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 990060e3-85a9-4452-b100-14f7a6892a78 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Grounded language-image pre-training
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d6cdb1fe-4044-4bc4-84f5-41148dfb81d4 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3eed: Ground everything everywhere in 3d
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fff96ced-d235-4551-8a5b-736904880d32 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Microsoft coco: Common objects in context
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 679bd081-9eab-4124-b646-e150e493fdae · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 86f5f433-2743-4c66-9336-3161c0ddaef9 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding View-on- graph: Zero-shot 3d visual grounding via vision-language reasoning on scene graphs
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e7e3f48c-5888-47b3-99cf-02e104416982 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3d-sps: Single-stage 3d visual grounding via referred point progressive selection
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation eaa2cf8c-ab49-4f40-9e22-9dd632d5699f · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding SQA3D: Situated Question Answering in 3D Scenes
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f5db08d-fb08-4d91-97e2-a62a87f4d3cf · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Goucher, Adam Perelman, Aditya Ramesh, and Aidan Clark et al
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 89d94f3b-0fd8-4117-84e4-7eb4187182f3 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Languagerefer: Spatial-language model for 3d visual grounding
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 804b1176-f87a-4488-82aa-f34154d5fdb3 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Language- grounded indoor 3d semantic segmentation in the wild
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e8181c31-468e-47a5-9aaa-621d58963501 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Mask3d: Mask trans- former for 3d semantic instance segmentation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e85d7520-83ca-4c62-a919-c53c0a1894e0 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Evaluating zero-shot gpt-4v performance on 3d vi- sual question answering benchmarks, 2024
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fb31858b-2724-4037-a0f0-14fc4a0e0baa · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Hashimoto
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation e7588881-b2a1-403d-981b-8ab81995c8c4 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Gemini: A family of highly capable multimodal models, 2025
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 00d5c209-b687-46ef-930b-6564ef3fe0c7 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 40bfa89e-34f3-4e62-89db-af3f72fd4069 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9dfa92bb-95b8-4e8f-a1ee-e45f1e9792f5 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Realworldqa
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6889bd48-7bd7-473d-ad12-834ff21fc8c2 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Qwen2.5 Technical Report
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73561748-6326-40c7-8411-bbce61381ead · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Sat: 2d semantics assisted training for 3d visual grounding
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 08c79efa-76bc-4065-9f52-83fed13b6322 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8e754b19-6361-4677-b4d2-9b35de80d58c · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Scannet++: A high-fidelity dataset of 3d in- door scenes
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72dadb02-5187-48be-8e89-950c586ccd5d · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual refer- ring
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a4be2bb5-2751-4d74-a4b4-c5f4ac0c264a · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Multi3drefer: Grounding text description to multiple 3d ob- jects
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7ba9a399-a68b-4a16-b07c-e8d9b19d776e · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Towards learning a generalist model for embod- ied navigation
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 473c6c45-a62f-48ee-946f-aeb34c1cb832 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Learning from videos for 3d world: Enhancing mllms with 3d vision geometry priors.arXiv preprint arXiv:2505.24625,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d49da31-7603-4865-8b38-9e7ca9fa0cb9 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Video-3d llm: Learning position-aware video representation for 3d scene understanding
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bc1e7350-31e4-4b27-9f50-6a1eb06190e7 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Llava-3d: A simple yet effective pathway to empowering lmms with 3d-awareness, 2024
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation aede57e3-b847-485b-8cae-c13d3ab279a2 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding 3d-vista: Pre-trained transformer for 3d vision and text alignment
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 87b2bc23-7840-435c-aa55-b96ad2a15573 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unifying 3D Vision-Language Understanding via Promptable Queries
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5d76e898-d0c7-45e5-a8e8-862ac983013b · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Generalized decoding for pixel, image, and lan- guage
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 78a29522-3b1a-456e-9cc1-91dc6ac60d28 · outbound
Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding Unresolved cited work
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
No inbound Pith citation observations are available.