Pith. sign in

Paper Citation Record · LEDGER

SQA3D: Situated Question Answering in 3D Scenes

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 65 inbound Pith citation observations for arXiv:2210.07474.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2210.07474 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 65 of 65 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:58:28.023924Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:37:43.684631Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 819c36a9-47cb-438d-93f7-8b8516abab46 · inbound

Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents cites this paper.

Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents SQA3D: Situated Question Answering in 3D Scenes

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:27:40.683807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T03:27:40.524895Z digest=sha256:f573d3efb76212beb104fe60cc7abec6b281303733d094196192a44b800ab2b6

Observation 436a7c1b-a225-465c-8ada-000d3831bd31 · inbound

PerLA: Perceptive 3D Language Assistant cites this paper.

PerLA: Perceptive 3D Language Assistant SQA3D: Situated Question Answering in 3D Scenes

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T05:52:39.078828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:52:39.078828Z digest=sha256:7f3902c89b17bd828ac6f01254711d0e3cb36a039c2cf2b2aa66e8a65d1c8279

Observation b423ee5f-888c-4ac2-9c5e-6767df90a78b · inbound

Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation cites this paper.

Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation SQA3D: Situated Question Answering in 3D Scenes

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:25:28.485261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-23T07:23:51.435139Z digest=sha256:f5485f108a59e43f154c81698c0c2c6f571539e6e0cbbc4a46b5b41758d16a4a

Observation e21d37cc-8cd7-4298-8ba7-7674f125bd40 · inbound

NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries cites this paper.

NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries SQA3D: Situated Question Answering in 3D Scenes

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:51.702124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:51.702124Z digest=sha256:4ee05697d9ee69b0b0d94d7074f45b640afaba10c3aa4f600fb4047ae31d21ae

Observation e8c46256-0ae8-43b0-aca2-222628685921 · inbound

3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding cites this paper.

3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:45:36.728239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:45:36.728239Z digest=sha256:c29bc59dfff2fa6f7081a6359f40066189f612d155f9f8660578d8b1e59231b4

Observation 6a38dbac-d56f-413e-bffd-407f73c31438 · inbound

Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models cites this paper.

Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models SQA3D: Situated Question Answering in 3D Scenes

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T04:39:55.834755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:39:55.834755Z digest=sha256:05220117821ceb3696a1b573fe1083af48bd77404502a0f3a59251cd9bf8eacb

Observation 88c484b1-9a0e-4456-9a8e-bce52bb091a6 · inbound

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer cites this paper.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer SQA3D: Situated Question Answering in 3D Scenes

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.541249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.541249Z digest=sha256:7cf3e627cb4f03ad652ca6d486698eba48e294e218903927a3598d02bc6576c7

Observation f7ae0f13-03a1-47e9-829e-8ac4747cb558 · inbound

CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds cites this paper.

CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds SQA3D: Situated Question Answering in 3D Scenes

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T21:48:27.366951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:48:27.366951Z digest=sha256:c24c5d3782011de56d2b91738b8a84a24aa9967f42fbef62b94cf6d337e2a052

Observation a9548686-ba0b-4b81-b9d0-bb565b686606 · inbound

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding cites this paper.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.511670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.511670Z digest=sha256:5cefce2d95468b511ec85bd14b0c7ef0b40777939be14c7cccbd1332450c1092

Observation 8c2750c8-b066-42ba-af1e-1a379520b0bc · inbound

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering cites this paper.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering SQA3D: Situated Question Answering in 3D Scenes

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.800718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.800718Z digest=sha256:84ccfed8e8826f606dda6e98d5469f231f1d850e1a4cc032e891c270d29f47ff

Observation 92abf978-0943-444b-8104-eeb60dacb6da · inbound

Hypo3D: Exploring Hypothetical Reasoning in 3D cites this paper.

Hypo3D: Exploring Hypothetical Reasoning in 3D SQA3D: Situated Question Answering in 3D Scenes

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T17:12:16.477250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:12:16.477250Z digest=sha256:4c77649b6e4c3c0908f852386eb0c654257dd9eec2b0fbd8c7170ed0d279509f

Observation a8e05f33-b0bc-4aec-9385-989f13418cc5 · inbound

Revisiting 3D LLM Benchmarks: Are We Really Testing 3D Capabilities? cites this paper.

Revisiting 3D LLM Benchmarks: Are We Really Testing 3D Capabilities? SQA3D: Situated Question Answering in 3D Scenes

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:02.077308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T04:54:02.077308Z digest=sha256:2a6679c2008be7ab3a8014afab5b95493ac99787c68798d9fbfd6b17134ef399

Observation 91930c99-fbf9-4179-abcd-b596c95a3ee6 · inbound

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding cites this paper.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.023924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.023924Z digest=sha256:448ec36111aa35136228900739251b87a6b130abc4101b19a76f1f5ba741afbc

Observation 6210f3c7-3150-4d1c-8a14-250ccb5f87f9 · inbound

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models cites this paper.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.251882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.251882Z digest=sha256:09bd9010e20adf2355197a48f393ac0f8dbbbecd3b3e78b6c4b118021a0d4df4

Observation 68b11d85-261f-4f2c-9ef5-cbf1b05bcec1 · inbound

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding cites this paper.

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:46.850936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:46.850936Z digest=sha256:08c9aaf9e3409b89269fb357751e4b6f0c8814a36ac856b01860af3161f920f5

Observation 1f6ad87a-adb8-4a2f-bf28-f37fa25e739c · inbound

AdaToken-3D: Dynamic Spatial Gating for Efficient 3D Large Multimodal-Models Reasoning cites this paper.

AdaToken-3D: Dynamic Spatial Gating for Efficient 3D Large Multimodal-Models Reasoning SQA3D: Situated Question Answering in 3D Scenes

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:05.196777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:05.196777Z digest=sha256:03db25101572e4693a96934c6ea47c838978cf2801bb27c4679ffc5b23deb125

Observation 167fd060-4958-473b-b1af-9bbd11db0d31 · inbound

DC-Scene: Data-Centric Learning for 3D Scene Understanding cites this paper.

DC-Scene: Data-Centric Learning for 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:51.975970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:51.975970Z digest=sha256:8d068b7f3695257bf0b99a2e9b0a84169039f38b88bac4e28199fb3169c3d999

Observation 3037b647-cbac-4793-bddc-87e095cc4d53 · inbound

CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation cites this paper.

CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation SQA3D: Situated Question Answering in 3D Scenes

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:55.344695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:55.344695Z digest=sha256:9cbd4ae1c4b9c7506647244009ad588be7302ec696ba8b188034244a4a96e96c

Observation c627b4dc-78c6-4e8b-a85c-543e476b10c9 · inbound

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction cites this paper.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction SQA3D: Situated Question Answering in 3D Scenes

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.925336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:3554646b93cacaaaadc6d9ff2f6cee7337c55a055e84e64e344af0fe42a71635

Observation 0c12676b-a1a9-41a5-b575-7c99e3bc2948 · inbound

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts cites this paper.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts SQA3D: Situated Question Answering in 3D Scenes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.614106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.614106Z digest=sha256:8ceedb83548640c229a5ca9bd299c2a84202bc63bdc87bfd846c87b57c0d61f9

Observation 35281398-dec0-4f5a-9f3e-b9ae96e21f3e · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence SQA3D: Situated Question Answering in 3D Scenes

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.901724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:07d6b8ff6a401bb3be4fa02574d5cf08e507ab66200a34c34c55bbac8d2c578c

Observation b1371454-5560-4685-be67-40b01adbe23b · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence SQA3D: Situated Question Answering in 3D Scenes

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.433617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:c2b3280bfaa19baade13c60e44656a2e78a3226e17d5f75009720a6bb84c08d3

Observation 5e872cd2-3e60-48b5-b16e-e9c2b767fd75 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces SQA3D: Situated Question Answering in 3D Scenes

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.281354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.281354Z digest=sha256:d9f83aa6c115c45d909eaf796c93e5a1251ec11496147e64f71bdde1f5beffcd

Observation 7ce951ce-6584-460d-a315-39895b64c3f3 · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? SQA3D: Situated Question Answering in 3D Scenes

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.844021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.844021Z digest=sha256:1e39c274b3a004390800107216c4aa5cb4185aa5179a34dd551fa95f039a9b96

Observation fd4dec3b-1a3b-4b5f-b142-d4d15b470c63 · inbound

Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs cites this paper.

Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs SQA3D: Situated Question Answering in 3D Scenes

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:45.371810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:45.371810Z digest=sha256:03eeb7625687423536073947f9747f2eca5da2460f0a29936c2955416ee8b198

Observation 27d46f3b-d777-4076-b839-5e64e9a65f93 · inbound

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation cites this paper.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation SQA3D: Situated Question Answering in 3D Scenes

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.971816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.971816Z digest=sha256:4fa0854543c28e7f3bd8210235b12283f9cf6179420407a26c89afe5521953de

Observation c05b8ad3-7462-4d43-a348-e4e2897b91a0 · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:02.275383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:02.275383Z digest=sha256:b79be554fc977ed235f003a63e3e51370701b843710ec7d6624bc65bbbb47dca

Observation 68b78ee9-4bc6-4ab9-9f87-2c5e4c9eaa73 · inbound

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models cites this paper.

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:16.929007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:16.929007Z digest=sha256:78c255df1e3c2e2a4adf5e5567f98930e0fdb9da182423bd92782b7c58e982fd

Observation 16a7dcca-1648-4e11-a8d9-68a24647f1d1 · inbound

Reinforced Embodied Active Defense: Exploiting Adaptive Interaction for Robust Visual Perception in Adversarial 3D Environments cites this paper.

Reinforced Embodied Active Defense: Exploiting Adaptive Interaction for Robust Visual Perception in Adversarial 3D Environments SQA3D: Situated Question Answering in 3D Scenes

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T18:19:07.226106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:19:07.226106Z digest=sha256:c71056109130aa9ec0fe91976f5fd6bc7439664eb0020bfa6ac4359c5b8b2855

Observation 6fe358d8-c9ab-4143-9eb9-a6092229c0b9 · inbound

NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding cites this paper.

NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding SQA3D: Situated Question Answering in 3D Scenes

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T13:48:03.529117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:48:03.529117Z digest=sha256:66a015b9f4453fe700b502be8af2a3e3c7fab1992db7fd77e47c4ada66c8034f

Observation db95da7b-c382-4b00-a051-48f98fe660d1 · inbound

Agentic Services Computing cites this paper.

Agentic Services Computing SQA3D: Situated Question Answering in 3D Scenes

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-04T14:41:50.768259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:41:50.768259Z digest=sha256:6e5754f87347f9ba99096fbc38ffdf23fb03696a4004a4c3d25c542bb49d8625

Observation 959ea317-4073-4a7c-9901-74d0171b7331 · inbound

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert cites this paper.

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert SQA3D: Situated Question Answering in 3D Scenes

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:13.152895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:38:13.152895Z digest=sha256:a892eb96a661d3759e329b262c8cac1f0027ebbb36cd9d5145342abd53e2dd2e

Observation 4110447d-fc49-4567-9278-45ac00f965b6 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report SQA3D: Situated Question Answering in 3D Scenes

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:42:05.711355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:2334340f22c636ba9f5eedb7325cc39c61c1de90ace7c97eb96d78f7c1af986e

Observation b197cc77-cd8e-4bac-8457-6fb4a9deafa2 · inbound

POMA-3D: The Point Map Way to 3D Scene Understanding cites this paper.

POMA-3D: The Point Map Way to 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.613552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T20:27:27.347592Z digest=sha256:6fbf40cfbcd6d22c8d4a57fce3bc076b038c3fd94baf7d5d7a86d28f69607132

Observation 0dc1fc8b-b746-4df7-a58d-114e42ef3a65 · inbound

Vision-Language Memory for Spatial Reasoning cites this paper.

Vision-Language Memory for Spatial Reasoning SQA3D: Situated Question Answering in 3D Scenes

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:34.304977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:34.304977Z digest=sha256:70e62800e675cba3277ed3ff9e1ed3166313b978f02fc42bcdde6329b245dac0

Observation a5cd0976-4699-43a0-9fbc-cd9b660cc7f7 · inbound

SpatialMosaic: A Multiview VLM Dataset for Partial Visibility cites this paper.

SpatialMosaic: A Multiview VLM Dataset for Partial Visibility SQA3D: Situated Question Answering in 3D Scenes

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:43:20.617257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-16T19:42:25.626448Z digest=sha256:61c08b6aec54fc673684cf652f15127cb8a116523c8597dd5dfb7ed1dc8b6cbd

Observation 28c0cfce-8169-4eb5-8926-21892d136bf1 · inbound

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding cites this paper.

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:30:22.455334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T09:30:03.668178Z digest=sha256:791de3bfca272956d6da8d80ce26c9573ec3d5f08505505ba1bab38ca48bed72

Observation 6692311b-c692-414b-9e12-bbd1b51aa19b · inbound

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM cites this paper.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM SQA3D: Situated Question Answering in 3D Scenes

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.431171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:fc9532f694829adf833cdedac5332a59e097b949912b3e1f3ae526027a81ceeb

Observation ae767bea-78e7-4d99-a4fd-658e6227ffd2 · inbound

RGB-Pointmap Pretraining for Unified 3D Scene Understanding cites this paper.

RGB-Pointmap Pretraining for Unified 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:23:17.922593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T21:19:49.421653Z digest=sha256:da7e452a7198d2c009191bb494c75056f64a5822d8456ce48f53ff842215f9d4

Observation 2a22fb0e-cdbe-4f8f-b2e3-4f2abead8eca · inbound

Geometry-Guided 3D Visual Token Pruning for Video-Language Models cites this paper.

Geometry-Guided 3D Visual Token Pruning for Video-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:51:09.960260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T05:49:38.346274Z digest=sha256:5d057929df6b0f21df4173dbd663430a43d8099a17c71d41823b1afcbd397f36

Observation ef9c6ca6-02f4-4ae1-ab5f-df39ffd73df5 · inbound

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models cites this paper.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.089512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:6a021c84a742793e67742ec2eb44a8ac13f2955268ee615cb65d2c912f7c63ac

Observation ef774b3a-3006-4e17-b46c-3109d04b72bf · inbound

Unlocking Dense Metric Depth Estimation in VLMs cites this paper.

Unlocking Dense Metric Depth Estimation in VLMs SQA3D: Situated Question Answering in 3D Scenes

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:23:40.988324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T19:20:04.468206Z digest=sha256:3769315b36c96d36ddf3ae61878b6266e38c9362c65060ab293bb802be238915

Observation 11710254-5529-4b0e-8a7f-ec29e86173f0 · inbound

Unlocking Dense Metric Depth Estimation in VLMs cites this paper.

Unlocking Dense Metric Depth Estimation in VLMs SQA3D: Situated Question Answering in 3D Scenes

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:51.100352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T07:54:52.926995Z digest=sha256:8a266ca2c621792dd182d7076523f591a97b42ed3337ae3b82427f0f92e522ce

Observation 357819a6-6195-4f6b-8724-839622430391 · inbound

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models cites this paper.

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:59:36.535565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-21T04:55:40.370796Z digest=sha256:d0f060214d394afab2bd1dc9c0beddd8a14487f8df253ba1301967051fa5d29b

Observation 9f8efa9a-6140-4949-8425-2f04eabe57c1 · inbound

Extending Embodied Question Answering from Perception to Decision cites this paper.

Extending Embodied Question Answering from Perception to Decision SQA3D: Situated Question Answering in 3D Scenes

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:58.998391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T21:30:40.182958Z digest=sha256:334c615dcb6b652cc75fd4f40ab104165818ba9071e6d1439a74ef240d71db36

Observation 65436e5f-a2d6-426a-9bb6-6f72fb166a94 · inbound

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)? cites this paper.

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)? SQA3D: Situated Question Answering in 3D Scenes

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.561481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T07:48:19.295578Z digest=sha256:819641760456ee9850225ecd7745db47a44de3c338cc2eb94d3de14e0e72c2a7

Observation c4294406-1598-4665-9516-753a97861a12 · inbound

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration? cites this paper.

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration? SQA3D: Situated Question Answering in 3D Scenes

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:12:24.824379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T17:08:59.111306Z digest=sha256:c3371b4cfc71785eaf46728eacbd3c7885063c72ca44eeae4c9d420ed4a125fb

Observation 26e71fa9-63e9-477d-9549-df1090713b02 · inbound

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation cites this paper.

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation SQA3D: Situated Question Answering in 3D Scenes

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:27.034459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T10:54:02.188634Z digest=sha256:baaaf69cb382af673d052304db2788ee37a2001c0729de4206c2d2cc7ef1ffc0

Observation 3290b45a-ea4f-4d17-ac26-f0b2dceadf3c · inbound

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models cites this paper.

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:57.525718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T02:13:53.219095Z digest=sha256:a1794b034f1a18a4993012f07b5e5402ea89aafbdf0dd1cf0bf77863927f70f1

Observation 9a827449-de59-4e92-a137-7a6db53cd253 · inbound

PAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene Understanding cites this paper.

PAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:57.091368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T02:17:30.190068Z digest=sha256:eff9c9868d4a2824810a753a787f9a1747535720ec8978cb2b3f7b98239710c7

Observation e3bbb6fe-fb9b-44d0-8d3b-509856a63db1 · inbound

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks cites this paper.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks SQA3D: Situated Question Answering in 3D Scenes

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.647837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:819979dd17fcea4fd8201da2742da38987db6ad08e95946a9f1f9db1ef325cd2

Observation 9db02e74-6048-4617-85ec-1b3aea2b3f17 · inbound

Intelligent Automation for Embodied Benchmark Construction: Pipelines, Embodiments, Simulators, and Trends cites this paper.

Intelligent Automation for Embodied Benchmark Construction: Pipelines, Embodiments, Simulators, and Trends SQA3D: Situated Question Answering in 3D Scenes

Reference 113

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T11:58:05.934971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T09:17:50.473747Z digest=sha256:7541777a0a5a109c85d2dc9b23f31d8391dfc42ea49ae3d01b28957b1745c939

Observation c57f703c-ac56-45fd-a4df-08ab3dffc33d · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI SQA3D: Situated Question Answering in 3D Scenes

Reference 131

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:fcdaa04989558bd26519c7100814aa8b1741f200b2aa90c8fea82bf83fb196ca

Observation 5a206202-79d4-4d82-9ccd-deef6721c0ae · inbound

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models cites this paper.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models SQA3D: Situated Question Answering in 3D Scenes

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.689493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-08T02:44:00.608590Z digest=sha256:86fb617196f66887e9c21ed319aaef8d1ec81fdd0cf24b5d7baff27cb46f3058

Observation 6d7833d2-f947-402c-8f2f-9c41c8434e18 · inbound

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models cites this paper.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models SQA3D: Situated Question Answering in 3D Scenes

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:026427af54ebe81d711ac908fb1d65f97051347622a2cc205588ae29ca8c3e19

Observation 6c66308c-2870-44c7-824a-b686389d25ba · inbound

SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts cites this paper.

SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts SQA3D: Situated Question Answering in 3D Scenes

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-11T01:37:43.718773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-11T01:29:46.709388Z digest=sha256:72a40f53725b6e3af55cdceee4619694c20c49946774ffc4e430da32c75bf2cf

Observation e3603877-e820-49b5-abad-1c6e8531855f · inbound

IMBench: A Benchmark for Intuitive Robotic Manipulation cites this paper.

IMBench: A Benchmark for Intuitive Robotic Manipulation SQA3D: Situated Question Answering in 3D Scenes

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T22:45:08.351210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:45:08.351210Z digest=sha256:0eccbf3f1b1b8a2bdeb4f9b9023e1092e87e690ab8840676f05c77b96154cbc9

Observation c77eab79-ef14-4e5a-92a2-79d6ed19107c · inbound

Data Pyramid for Embodied Manipulation: A Survey cites this paper.

Data Pyramid for Embodied Manipulation: A Survey SQA3D: Situated Question Answering in 3D Scenes

Reference 268

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.880583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.880583Z digest=sha256:53f5664b17362aa241b71c43680ad870c6bd72f06635e4291bb999b4da39f534

Observation ba8df6b1-1262-4c46-9504-2301ce664a74 · inbound

Distill What RGB Can Recover: Privileged 3D Evidence for RGB-Only Vision-Language Models cites this paper.

Distill What RGB Can Recover: Privileged 3D Evidence for RGB-Only Vision-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T00:41:37.886561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:41:37.886561Z digest=sha256:4b74e23367411cb57ee240b72d267d787affdf9eecd7041ee40fa662a7b1bc97

Observation 1c97c294-4d7f-4147-b2f8-df8ad2e07620 · inbound

SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models cites this paper.

SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T15:08:10.235780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:08:10.235780Z digest=sha256:d8150f7cae6f479e9ff42bdcb47f0e71eae5ffd7a204b12dd962908ab071be84

Observation eaa2cf8c-ab49-4f40-9e22-9dd632d5699f · inbound

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding cites this paper.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.908340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.908340Z digest=sha256:440d2931804509a6a388011047f924ec140aa078f66d21ee615802ef60605d3f

Observation 9883fbc9-cda3-4c97-9331-2b087a2e0870 · inbound

HiSC: Hierarchical Spatial Clustering Token Compression for Efficient 3D Scene Understanding cites this paper.

HiSC: Hierarchical Spatial Clustering Token Compression for Efficient 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:51:52.908285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:51:52.908285Z digest=sha256:05ded93a0f3ca33129a76c73b47c62e99a5278b1529ca48e992bf021fc7c1a21

Observation 32cd7efc-d531-4675-b0dc-53185f7009b0 · inbound

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models cites this paper.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.481703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.481703Z digest=sha256:9abd5029027e0721b040ce26034379ad40a13444d0c6d55baa47896a6197532b

Observation 9109863f-f2b3-4b28-9562-c3096244ddf3 · inbound

CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting cites this paper.

CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting SQA3D: Situated Question Answering in 3D Scenes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T05:15:36.107424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:15:36.107424Z digest=sha256:f2bd7f8cc3c0bb2323dbdd3b278371187cc24d5b8c206fd4fbb39e50f866c169

Observation db4ce132-ef9d-4b62-91fb-afec3b931859 · inbound

CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting cites this paper.

CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting SQA3D: Situated Question Answering in 3D Scenes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T14:18:39.145114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:18:39.145114Z digest=sha256:8c15d7eb869a9e68cb323c457beeba1d83db7b1655921f7d51c2709c5f12fce7