Pith. sign in

Paper Citation Record · LEDGER

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

As of 8 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 6 inbound Pith citation observations for arXiv:2506.17545.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17545 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:34:54.732094Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:49:47.113836Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T21:16:14.446824Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ed8e7168-1445-4b8f-b31a-47ef08ff9f0a · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:01.552537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:50.085085Z digest=sha256:fe4ae4d59b3dcb31faf4b92af09fcc8b32beec83816b31a131fd373823946582

Observation 9226bca7-4524-49b8-8091-469a56ba4367 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scanqa: 3d question answering for spatial scene understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.187260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.187260Z digest=sha256:3db280330cb03491a6c4bd3cca34d11722b2f5334769a632f06b5fb1c9f7b384

Observation 4ad1c548-4f4c-4ed3-a81e-26e7e230146d · outbound

This paper cites Qwen2.5-VL Technical Report.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.293925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.293925Z digest=sha256:4ce0774a1124a7bfc3f32df01e6a0bcf080dd5699025ae9852ed7a7017744acb

Observation 3061aad6-af57-490d-8a9f-02616bb75303 · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.424528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.424528Z digest=sha256:36ddf0ffd42f5a70e5bfa1ae2a04074669f39bd38c51066ea67d4be1df8ce95e

Observation a866976b-438a-45c9-a23b-79a47a285988 · outbound

This paper cites Do as i can, not as i say: Grounding language in robotic affordances.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Do as i can, not as i say: Grounding language in robotic affordances

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.502146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.502146Z digest=sha256:0f66c42c23ab4c69d7a34dec1cc3ceb51f14239cabdb5c257f5de21dc84e6a90

Observation 8886deb5-0959-48ae-b47a-f5854c681229 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.590561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.590561Z digest=sha256:2ef7638eafd7601ba760a849b7203bb9d46a9401234782e59305d4f516a4c317

Observation 962879af-d356-4703-98e2-e3a1c2052ec4 · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than $3.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations R1-v: Reinforcing super generalization ability in vision-language models with less than $3

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:50.674777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:50.674777Z digest=sha256:cc67ab6e669e334b505f96beb2c4f61e73c2ce5b6798f83027772c61229ce896

Observation 279ff78f-65b3-41ff-b81e-d483adeb9d9a · outbound

This paper cites Language conditioned spatial relation reasoning for 3d object grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Language conditioned spatial relation reasoning for 3d object grounding

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:01.194834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:50.786284Z digest=sha256:6797a015dccf7d29a62c8794d7a0ce4d87abeb27a7db88ac65112abea341f331

Observation d0ab54f3-d442-4292-b066-44c289b2dfec · outbound

This paper cites End-to-end 3d dense captioning with vote2cap-detr.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations End-to-end 3d dense captioning with vote2cap-detr

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:00.957226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:50.864776Z digest=sha256:943607d8c4cd9c75e3bfc9244102c4dffe912fc1d53556941cb96fd37ca80050

Observation 285d4ae2-ca5a-45aa-8e85-414c5362e4b1 · outbound

This paper cites V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:00.706023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.009696Z digest=sha256:82b911162150d329d1c21c02dd433ee5ddcc6939ab5234ac69e83f85fec74ae4

Observation 28a4d20a-9564-41cb-88a2-23ecabea99b9 · outbound

This paper cites Scan2cap: Context-aware dense captioning in rgb-d scans.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scan2cap: Context-aware dense captioning in rgb-d scans

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:00.402504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.112310Z digest=sha256:a46914943bd95b6c1d9e34feb7d138f5772e47d505b0971131259687e430dc33

Observation 28484630-a5d9-4b57-a6a9-8d67a508401c · outbound

This paper cites Functionality understanding and segmentation in 3d scenes.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Functionality understanding and segmentation in 3d scenes

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:35:00.222262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.198878Z digest=sha256:38ad6fb690baf609daf6e778edebc6918a08d4bdde48c7d014fb022a7be1f3e4

Observation 7e5f8c23-6c1c-48b2-b91e-11f77dea73ad · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:59.955079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.302373Z digest=sha256:ea9dc27ebb5906e88770c32b502b25205f6d73f86a995e285c73af55638f4342

Observation b253733e-2e3f-400e-8a32-56f2b7b5b772 · outbound

This paper cites Scenefun3d: fine-grained functionality and affordance understanding in 3d scenes.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scenefun3d: fine-grained functionality and affordance understanding in 3d scenes

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:59.666719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.373750Z digest=sha256:fff29770399c9de0aa99018c12edc81d3eae37f22c838ab7d6e7aace4d5a566a

Observation 7acd6dc5-915f-4e15-b313-13cd14a70da8 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.432004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.432004Z digest=sha256:46e17a5d6c5a9127593364d60a9727a89a906feaab226eb66bea611221c169a9

Observation d14b9ef8-5e7f-4d64-b888-4815955d8d59 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.467823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.467823Z digest=sha256:b25847d717e2d5f1e10817ec4899281f3d1737773279270b27f2c5909fd53880

Observation 7f54d23d-d774-4890-8f84-618e83889bcc · outbound

This paper cites The ecological approach to visual perception: classic edition.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations The ecological approach to visual perception: classic edition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.538593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.538593Z digest=sha256:b42feddb2d88c0e1768e701eba4f26931ecc4e6afac0f1cf9f0eede62c47b46d

Observation f0353bd0-68cd-4d0e-b8aa-ee09a817f793 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.580536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.580536Z digest=sha256:1360c186272854a489e3a33a6d433a9ff0442c11b6de027f01014cde535336fc

Observation 005ced98-3c95-48b2-9856-325d9f985566 · outbound

This paper cites Viewrefer: Grasp the multi-view knowledge for 3d visual grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Viewrefer: Grasp the multi-view knowledge for 3d visual grounding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:59.393195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.630120Z digest=sha256:fa0423b4bd23e50da11ff3d952bda935a5a66096a3dd761d99ad74870078cb88

Observation b9d52e38-0094-444f-aa1f-65c4f97ce632 · outbound

This paper cites Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Point-Bind & Point-LLM: Aligning Point Cloud with Multi-modality for 3D Understanding, Generation, and Instruction Following

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.681819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.681819Z digest=sha256:307585d16be46aec4c2b04536c337708e81e52baff37e9d928ef2bfbc7d4447f

Observation 14668d63-bd0c-49fa-be77-36c52ef52330 · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Chat-scene: Bridging 3d scene and large language models with object identifiers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:59.186224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.737382Z digest=sha256:695e58eab4e237cd6edca7fff817ca98f2e868bb63beb6b8b5315d0348bd1cca

Observation 8c1c1e4c-3e24-4bd9-b9eb-e8941c7b684e · outbound

This paper cites An Embodied Generalist Agent in 3D World.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations An Embodied Generalist Agent in 3D World

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:51.827382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:51.827382Z digest=sha256:73aaa444138df8ae44aa1084521d0bc67e186e917474cbb267a6b92ba87966cb

Observation 353c30ed-8014-4f78-8788-fbacb505715b · outbound

This paper cites Text-guided graph neural networks for referring 3d instance segmentation.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Text-guided graph neural networks for referring 3d instance segmentation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:58.944257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.886696Z digest=sha256:463aefb8422363f77da9c6fc7c2fecfb023d120c28117291188e145997fc0dfb

Observation fbc57c2c-1005-4aa4-8900-2cdb4a8f0385 · outbound

This paper cites Multi-view transformer for 3d visual grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Multi-view transformer for 3d visual grounding

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:58.705701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:51.957621Z digest=sha256:075be6e6f1c3ff112b3f4879fc61f68765d0ebda5cacad6128f79eba2c685b14

Observation 1e59bc04-336c-47d1-8f1d-34864794b21f · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:52.047794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:52.047794Z digest=sha256:0da411944f00ca0d8340162c9593eedabc851ab237633c075871fbafd498713d

Observation 4c7ebe5b-73ec-46c8-b30a-254cba943b27 · outbound

This paper cites Bottom up top down detection transformers for language grounding in images and point clouds.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Bottom up top down detection transformers for language grounding in images and point clouds

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:58.471374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:52.121287Z digest=sha256:bf000bb30996afd1fb40b6a8e8ea29689234a0321b8ec8004da87f208823e33a

Observation 8b4f9729-5c84-4207-bc3a-ddc83b247789 · outbound

This paper cites Lerf: Language embedded radiance fields.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Lerf: Language embedded radiance fields

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:52.196735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:52.196735Z digest=sha256:b59e5dd507cf5ca1a02a3389f65ece086685e3b64f97fe1c48d7f4de36aa634f

Observation 5ff9199b-effe-478f-8564-e0160dc4a7d3 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations VideoChat: Chat-Centric Video Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:52.248369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:52.248369Z digest=sha256:3b28fff209b387af15244578ec2eaddf1680dd785b6b0ab5d9b092c73156a621

Observation e234c58c-8045-4310-aedb-429afb0f45c3 · outbound

This paper cites SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:52.340466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:52.340466Z digest=sha256:9ce58e2384d0baccb85d4c702c899a8df39fa6e1eb45cc6a41f915e301a320eb

Observation 5228ce6a-ba5c-47d6-9e27-6a5406917599 · outbound

This paper cites Sqa3d: Situated question answering in 3d scenes.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Sqa3d: Situated question answering in 3d scenes

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:52.404835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:52.404835Z digest=sha256:e905a9f79518a15101133e23548c1f00d37a6a9c015d1b86ae554235a8372bd7

Observation 7ecf0d1f-e033-4411-999d-0edd71d31830 · outbound

This paper cites Introducing openai o1.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Introducing openai o1

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:58.064915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:52.525748Z digest=sha256:475869d437f1d3cb58432e4161b4066b1b2e54c73a9ce42c6b2e3357e04ec544

Observation 1ce32298-8785-4299-8788-d676de5a47d2 · outbound

This paper cites Openscene: 3d scene understanding with open vocabularies.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Openscene: 3d scene understanding with open vocabularies

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:57.587518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:52.613220Z digest=sha256:7272a133d1c68fa0ca598e7280043b5bb3c87c23777c6ae23632264751100b7c

Observation e405d8fc-4ca3-4d86-bc3f-00c087ee3ed9 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Learning transferable visual models from natural language supervision

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:57.324756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:52.768733Z digest=sha256:d4738297665258f3b2673098f717ec0ad024df2bc5e5d586aa98533832c6b04d

Observation 69fe2126-b29c-4af0-95a6-dd0306ff9b4e · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations SAM 2: Segment Anything in Images and Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.004776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.004776Z digest=sha256:559e76d2900972a181d63269029cd0f847205f4b9650ea054ce34d5e79171bc1

Observation c3fa0c00-fe7f-4026-abb8-a8b4ab348688 · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2021.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations High-resolution image synthesis with latent diffusion models, 2021

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.075943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.075943Z digest=sha256:ef2f79364103147cb3ceb327f0a66ba2d6364752953799d7393a0eebc5b27bae

Observation 82a2a810-cd69-4659-bd7e-d28eb2c4c698 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.147630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.147630Z digest=sha256:f8f003c61ef8a73b3b0d3eabf24b13a18f69a8b53243e12f9dc456b7c76aa575

Observation 4bd2fa6a-d9fb-41b5-a809-f01b28e6e8d0 · outbound

This paper cites Mask3d: Mask transformer for 3d semantic instance segmentation.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Mask3d: Mask transformer for 3d semantic instance segmentation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:57.078831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:53.210278Z digest=sha256:1413a87a2fcc26f7372d3feec1501d8ed5835409e617b3599ef12adfb74f7375

Observation 45df17a8-1db0-49b3-bf58-476513600938 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.304750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.304750Z digest=sha256:ef3b3934db81755352240a568b808e0f1773027ee68430a6cfcdf251701fb353

Observation b2563195-10da-4789-9fb6-4383212f9d6a · outbound

This paper cites Chatgpt for robotics: Design principles and model abilities.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Chatgpt for robotics: Design principles and model abilities

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:56.868146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:53.418030Z digest=sha256:da40d614fb1e9a6af812f0e85665c4744d2de2bfd43b062034ded324a0e8b6f3

Observation 67ca3aea-21d9-42f0-82e8-0d874e7f886f · outbound

This paper cites Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Grounded-VideoLLM: Sharpening Fine-grained Temporal Grounding in Video Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.529692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.529692Z digest=sha256:3e52564d3be5a04c3a7dea75d203474b1ff6c747c120dc3c9d110dca62745c83

Observation 78530adf-24b4-4652-b6b5-b232e77c5410 · outbound

This paper cites Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:53.624919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:53.624919Z digest=sha256:7ade7fef5b4cff15b0357702083cf79b880652e2c6ee4af8bdb0bdd8453cdc25

Observation 6bf266ad-34ce-4b37-a064-63a11da1e28c · outbound

This paper cites Distilling coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Distilling coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:56.650264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:53.702757Z digest=sha256:f5cf0419847b22c101c74cfeb6963226bbf08ee611fcd81013eafa9a070be892

Observation f3770be4-4d4c-46c4-ab61-fa2d0697575d · outbound

This paper cites Pointllm: Empower- ing large language models to understand point clouds.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Pointllm: Empower- ing large language models to understand point clouds

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:56.416286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:53.845949Z digest=sha256:c8fbd300bb2ffd1b31f00f76ede2accc3c232965d0e11ef30997a8fc2ad8c368

Observation 05da5af7-46a8-4a2d-af28-13db6388021c · outbound

This paper cites Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Instancerefer: Cooperative holistic understanding for visual grounding on point clouds through instance multi-level contextual referring

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:56.111190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:53.974587Z digest=sha256:229ba015dbf9ee0b5231fbb989f3c4f666298a3af3448e4f98346771c095a7ec

Observation 7a2cf311-ede6-42a4-a5a4-ce6053d4b62d · outbound

This paper cites Visual programming for zero-shot open-vocabulary 3d visual grounding.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Visual programming for zero-shot open-vocabulary 3d visual grounding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:54.128620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:54.128620Z digest=sha256:26640b62caefa1882b4229c3f9007111df01f14a38c4d3ffd8be5336bf94da3a

Observation 1a39f9cf-49bf-44fc-bc8c-6d39f64c550d · outbound

This paper cites Empowering large language models with 3d situation awareness.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Empowering large language models with 3d situation awareness

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:55.935976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:54.272841Z digest=sha256:61360e06b8a35c3a2ca9321f5e9158b1e13fc87b988bd15fd734aebd58e92f39

Observation 494884c1-1a1c-49cf-a2f6-49713e22119b · outbound

This paper cites 3dvg-transformer: Relation modeling for visual grounding on point clouds.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations 3dvg-transformer: Relation modeling for visual grounding on point clouds

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:55.760801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:54.384753Z digest=sha256:f06fc4699ff394322fa33c4134b1f1320a6857abb8e63a5a10129df5ff27bd6c

Observation 3df51210-40c0-4bee-ae88-a9c4dfca88e0 · outbound

This paper cites BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:54.466311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:54.466311Z digest=sha256:728f8345d83b2674e5c95af013f9fff0128c52b57427d071975f9e11fbd89118

Observation ca3eb5f5-d067-4075-b7e1-c7514176b17b · outbound

This paper cites Uni3D: Exploring Unified 3D Representation at Scale.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations Uni3D: Exploring Unified 3D Representation at Scale

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:54.555746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:34:54.555746Z digest=sha256:3a8d3b877b9ca34331121644602ca7001185f3fcbf9dbfa23dbc16071f81702f

Observation 64d34c28-84ba-4862-84ad-17a14bb5ecd8 · outbound

This paper cites 3d-vista: Pre-trained transformer for 3d vision and text alignment.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations 3d-vista: Pre-trained transformer for 3d vision and text alignment

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:55.621691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:54.649230Z digest=sha256:2f6f8564a617a2f75f95df579f908475d927342f4af13af53eec042861fc6c0b

Observation ad927a9d-0c9f-4e92-80ff-8f16ffaec360 · outbound

This paper cites [EVENT]" in the video, determine the precise time period of the occurrence of the object. Provide the start and end times (in seconds, precise to one decimal place) in the format.

Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations [EVENT]" in the video, determine the precise time period of the occurrence of the object. Provide the start and end times (in seconds, precise to one decimal place) in the format

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:34:55.426439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T23:34:54.732094Z digest=sha256:b7181da268fd587014e6fcabf949cb96a19123da8b9a84c21996ef693313026f

Pith citing papers

Observation a03740dc-991f-46e0-b652-48ad2b8c165e · inbound

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding cites this paper.

3D-R1: Enhancing Reasoning in 3D VLMs for Unified Scene Understanding Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T10:49:47.113836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:49:47.113836Z digest=sha256:58b6923936d30c9217a218287bd6dd7be81331f1aacca0133698a71f1adafa9b

Observation dc1d0c90-7740-4d73-9f6f-43a60c62986c · inbound

What if? Emulative Simulation with World Models for Situated Reasoning cites this paper.

What if? Emulative Simulation with World Models for Situated Reasoning Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 115

Resolution
unresolved
no resolver link, observed 2026-07-15T13:51:30.008232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:51:30.008232Z digest=sha256:c34db8e1ff689af5674db58613b50b4f3e93ef2ea33c08c0dd770317e0827207

Observation b6c9fc21-e333-4f3e-a41d-595aa2598a86 · inbound

Token Warping Helps MLLMs Look from Nearby Viewpoints cites this paper.

Token Warping Helps MLLMs Look from Nearby Viewpoints Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:08:17.348876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T21:07:55.062113Z digest=sha256:80b39fbaec72e5a9f37b08efee81535015c29a04cc68ba95b474f35183892927

Observation f454b41e-fca4-4ae5-8299-3dba8be8881e · inbound

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models cites this paper.

CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:28:04.506945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T05:27:30.938311Z digest=sha256:1d07e48cb5d8460f3a4648155b60ba791a4f3506e8c5001ee75c9ff9a2173c15

Observation 8b7f3447-d5a3-469b-a3dd-700edaba72ba · inbound

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs cites this paper.

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:14.448373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:14:48.328093Z digest=sha256:dc2cfdd0f4f56fed23e8d5d5526418a26a3ba4ed82eb7d884990e065acf619e6

Observation 5e2d5bd9-7531-4e87-b024-96a063370336 · inbound

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs cites this paper.

Distilling Neuro-Symbolic Programs into 3D Multi-modal LLMs Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:34:38.162717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T11:26:47.680406Z digest=sha256:8f844d76c66b50b4d1a6355f622d32ac11a68b2c4beacd471485c42df22390d0