Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:34:54.732094Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 6 inbound Pith citation observations for arXiv:2506.17545.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:34:54.732094Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T10:49:47.113836Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T21:16:14.446824Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ed8e7168-1445-4b8f-b31a-47ef08ff9f0a · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9226bca7-4524-49b8-8091-469a56ba4367 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scanqa: 3d question answering for spatial scene understanding
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ad1c548-4f4c-4ed3-a81e-26e7e230146d · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3061aad6-af57-490d-8a9f-02616bb75303 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a866976b-438a-45c9-a23b-79a47a285988 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Do as i can, not as i say: Grounding language in robotic affordances
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8886deb5-0959-48ae-b47a-f5854c681229 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scanrefer: 3d object localization in rgb-d scans using natural language
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 962879af-d356-4703-98e2-e3a1c2052ec4 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations R1-v: Reinforcing super generalization ability in vision-language models with less than $3
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 279ff78f-65b3-41ff-b81e-d483adeb9d9a · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Language conditioned spatial relation reasoning for 3d object grounding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d0ab54f3-d442-4292-b066-44c289b2dfec · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations End-to-end 3d dense captioning with vote2cap-detr
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 285d4ae2-ca5a-45aa-8e85-414c5362e4b1 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 28a4d20a-9564-41cb-88a2-23ecabea99b9 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scan2cap: Context-aware dense captioning in rgb-d scans
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 28484630-a5d9-4b57-a6a9-8d67a508401c · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Functionality understanding and segmentation in 3d scenes
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7e5f8c23-6c1c-48b2-b91e-11f77dea73ad · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scannet: Richly-annotated 3d reconstructions of indoor scenes
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b253733e-2e3f-400e-8a32-56f2b7b5b772 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scenefun3d: fine-grained functionality and affordance understanding in 3d scenes
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7acd6dc5-915f-4e15-b313-13cd14a70da8 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d14b9ef8-5e7f-4d64-b888-4815955d8d59 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f54d23d-d774-4890-8f84-618e83889bcc · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations The ecological approach to visual perception: classic edition
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0353bd0-68cd-4d0e-b8aa-ee09a817f793 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 005ced98-3c95-48b2-9856-325d9f985566 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Viewrefer: Grasp the multi-view knowledge for 3d visual grounding
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b9d52e38-0094-444f-aa1f-65c4f97ce632 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14668d63-bd0c-49fa-be77-36c52ef52330 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Chat-scene: Bridging 3d scene and large language models with object identifiers
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8c1c1e4c-3e24-4bd9-b9eb-e8941c7b684e · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations An Embodied Generalist Agent in 3D World
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 353c30ed-8014-4f78-8788-fbacb505715b · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Text-guided graph neural networks for referring 3d instance segmentation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fbc57c2c-1005-4aa4-8900-2cdb4a8f0385 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Multi-view transformer for 3d visual grounding
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1e59bc04-336c-47d1-8f1d-34864794b21f · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c7ebe5b-73ec-46c8-b30a-254cba943b27 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Bottom up top down detection transformers for language grounding in images and point clouds
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8b4f9729-5c84-4207-bc3a-ddc83b247789 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Lerf: Language embedded radiance fields
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ff9199b-effe-478f-8564-e0160dc4a7d3 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations VideoChat: Chat-Centric Video Understanding
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e234c58c-8045-4310-aedb-429afb0f45c3 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5228ce6a-ba5c-47d6-9e27-6a5406917599 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Sqa3d: Situated question answering in 3d scenes
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ecf0d1f-e033-4411-999d-0edd71d31830 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Introducing openai o1
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1ce32298-8785-4299-8788-d676de5a47d2 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Openscene: 3d scene understanding with open vocabularies
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e405d8fc-4ca3-4d86-bc3f-00c087ee3ed9 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Learning transferable visual models from natural language supervision
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 69fe2126-b29c-4af0-95a6-dd0306ff9b4e · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations SAM 2: Segment Anything in Images and Videos
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3fa0c00-fe7f-4026-abb8-a8b4ab348688 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations High-resolution image synthesis with latent diffusion models, 2021
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82a2a810-cd69-4659-bd7e-d28eb2c4c698 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Proximal Policy Optimization Algorithms
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bd2fa6a-d9fb-41b5-a809-f01b28e6e8d0 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Mask3d: Mask transformer for 3d semantic instance segmentation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 45df17a8-1db0-49b3-bf58-476513600938 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2563195-10da-4789-9fb6-4383212f9d6a · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Chatgpt for robotics: Design principles and model abilities
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 67ca3aea-21d9-42f0-82e8-0d874e7f886f · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78530adf-24b4-4652-b6b5-b232e77c5410 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bf266ad-34ce-4b37-a064-63a11da1e28c · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Distilling coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f3770be4-4d4c-46c4-ab61-fa2d0697575d · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Pointllm: Empower- ing large language models to understand point clouds
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 05da5af7-46a8-4a2d-af28-13db6388021c · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7a2cf311-ede6-42a4-a5a4-ce6053d4b62d · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Visual programming for zero-shot open-vocabulary 3d visual grounding
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a39f9cf-49bf-44fc-bc8c-6d39f64c550d · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Empowering large language models with 3d situation awareness
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 494884c1-1a1c-49cf-a2f6-49713e22119b · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations 3dvg-transformer: Relation modeling for visual grounding on point clouds
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3df51210-40c0-4bee-ae88-a9c4dfca88e0 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca3eb5f5-d067-4075-b7e1-c7514176b17b · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Uni3D: Exploring Unified 3D Representation at Scale
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64d34c28-84ba-4862-84ad-17a14bb5ecd8 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations 3d-vista: Pre-trained transformer for 3d vision and text alignment
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad927a9d-0c9f-4e92-80ff-8f16ffaec360 · outbound
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations [EVENT]" in the video, determine the precise time period of the occurrence of the object. Provide the start and end times (in seconds, precise to one decimal place) in the format
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a03740dc-991f-46e0-b652-48ad2b8c165e · inbound
3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc1d0c90-7740-4d73-9f6f-43a60c62986c · inbound
What if? Emulative Simulation with World Models for Situated Reasoning Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
Reference 115
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6c9fc21-e333-4f3e-a41d-595aa2598a86 · inbound
Token Warping Helps MLLMs Look from Nearby Viewpoints Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f454b41e-fca4-4ae5-8299-3dba8be8881e · inbound
CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8b7f3447-d5a3-469b-a3dd-700edaba72ba · inbound
Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5e2d5bd9-7531-4e87-b024-96a063370336 · inbound
Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.