Pith. sign in

Paper Citation Record · LEDGER

SQA3D: Situated Question Answering in 3D Scenes

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 65 inbound Pith citation observations for arXiv:2210.07474.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2210.07474 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 65 of 65 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:58:28.023924Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:37:43.684631Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 819c36a9-47cb-438d-93f7-8b8516abab46 · inbound

Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents cites this paper.

Describe, Explain, Plan and Select: Interactive Planning with Large Language Models Enables Open-World Multi-Task Agents SQA3D: Situated Question Answering in 3D Scenes

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:27:40.683807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T03:27:40.524895Z digest=sha256:19fb37b4f9d9d547d5972c9ca2f49bacfc72be5e9e8d06db097ce827d8f5c5dd

Observation 436a7c1b-a225-465c-8ada-000d3831bd31 · inbound

PerLA: Perceptive 3D Language Assistant cites this paper.

PerLA: Perceptive 3D Language Assistant SQA3D: Situated Question Answering in 3D Scenes

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T05:52:39.078828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:52:39.078828Z digest=sha256:7f3902c89b17bd828ac6f01254711d0e3cb36a039c2cf2b2aa66e8a65d1c8279

Observation b423ee5f-888c-4ac2-9c5e-6767df90a78b · inbound

Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation cites this paper.

Motion-aware Contrastive Learning for Temporal Panoptic Scene Graph Generation SQA3D: Situated Question Answering in 3D Scenes

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:25:28.485261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-23T07:23:51.435139Z digest=sha256:4ba4dae8016c0441d47a14270f6ac5e9125b89fa18ad63cfaeec33abd722dfba

Observation e21d37cc-8cd7-4298-8ba7-7674f125bd40 · inbound

NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries cites this paper.

NoisyEQA: Benchmarking Embodied Question Answering Against Noisy Queries SQA3D: Situated Question Answering in 3D Scenes

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:43:51.702124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:43:51.702124Z digest=sha256:4ee05697d9ee69b0b0d94d7074f45b640afaba10c3aa4f600fb4047ae31d21ae

Observation e8c46256-0ae8-43b0-aca2-222628685921 · inbound

3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding cites this paper.

3DGraphLLM: Combining Semantic Graphs and Large Language Models for 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:45:36.728239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:45:36.728239Z digest=sha256:c29bc59dfff2fa6f7081a6359f40066189f612d155f9f8660578d8b1e59231b4

Observation 6a38dbac-d56f-413e-bffd-407f73c31438 · inbound

Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models cites this paper.

Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models SQA3D: Situated Question Answering in 3D Scenes

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T04:39:55.834755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:39:55.834755Z digest=sha256:05220117821ceb3696a1b573fe1083af48bd77404502a0f3a59251cd9bf8eacb

Observation 88c484b1-9a0e-4456-9a8e-bce52bb091a6 · inbound

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer cites this paper.

3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer SQA3D: Situated Question Answering in 3D Scenes

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T22:38:21.541249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:38:21.541249Z digest=sha256:7cf3e627cb4f03ad652ca6d486698eba48e294e218903927a3598d02bc6576c7

Observation f7ae0f13-03a1-47e9-829e-8ac4747cb558 · inbound

CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds cites this paper.

CL3DOR: Contrastive Learning for 3D Large Multimodal Models via Odds Ratio on High-Resolution Point Clouds SQA3D: Situated Question Answering in 3D Scenes

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T21:48:27.366951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:48:27.366951Z digest=sha256:c24c5d3782011de56d2b91738b8a84a24aa9967f42fbef62b94cf6d337e2a052

Observation a9548686-ba0b-4b81-b9d0-bb565b686606 · inbound

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding cites this paper.

3UR-LLM: An End-to-End Multimodal Large Language Model for 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:12.511670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:12.511670Z digest=sha256:5cefce2d95468b511ec85bd14b0c7ef0b40777939be14c7cccbd1332450c1092

Observation 8c2750c8-b066-42ba-af1e-1a379520b0bc · inbound

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering cites this paper.

Embodied Intelligence for 3D Understanding: A Survey on 3D Scene Question Answering SQA3D: Situated Question Answering in 3D Scenes

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T19:24:17.800718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:24:17.800718Z digest=sha256:84ccfed8e8826f606dda6e98d5469f231f1d850e1a4cc032e891c270d29f47ff

Observation 92abf978-0943-444b-8104-eeb60dacb6da · inbound

Hypo3D: Exploring Hypothetical Reasoning in 3D cites this paper.

Hypo3D: Exploring Hypothetical Reasoning in 3D SQA3D: Situated Question Answering in 3D Scenes

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T17:12:16.477250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:12:16.477250Z digest=sha256:4c77649b6e4c3c0908f852386eb0c654257dd9eec2b0fbd8c7170ed0d279509f

Observation a8e05f33-b0bc-4aec-9385-989f13418cc5 · inbound

Revisiting 3D LLM Benchmarks: Are We Really Testing 3D Capabilities? cites this paper.

Revisiting 3D LLM Benchmarks: Are We Really Testing 3D Capabilities? SQA3D: Situated Question Answering in 3D Scenes

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T04:54:02.077308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T04:54:02.077308Z digest=sha256:2a6679c2008be7ab3a8014afab5b95493ac99787c68798d9fbfd6b17134ef399

Observation 91930c99-fbf9-4179-abcd-b596c95a3ee6 · inbound

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding cites this paper.

Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:28.023924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:58:28.023924Z digest=sha256:448ec36111aa35136228900739251b87a6b130abc4101b19a76f1f5ba741afbc

Observation 6210f3c7-3150-4d1c-8a14-250ccb5f87f9 · inbound

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models cites this paper.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.251882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.251882Z digest=sha256:09bd9010e20adf2355197a48f393ac0f8dbbbecd3b3e78b6c4b118021a0d4df4

Observation 68b11d85-261f-4f2c-9ef5-cbf1b05bcec1 · inbound

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding cites this paper.

LLaVA-4D: Embedding SpatioTemporal Prompt into LMMs for 4D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T20:42:46.850936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:42:46.850936Z digest=sha256:08c9aaf9e3409b89269fb357751e4b6f0c8814a36ac856b01860af3161f920f5

Observation 1f6ad87a-adb8-4a2f-bf28-f37fa25e739c · inbound

AdaToken-3D: Dynamic Spatial Gating for Efficient 3D Large Multimodal-Models Reasoning cites this paper.

AdaToken-3D: Dynamic Spatial Gating for Efficient 3D Large Multimodal-Models Reasoning SQA3D: Situated Question Answering in 3D Scenes

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:31:05.196777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:31:05.196777Z digest=sha256:4dba5f115f516c8432292b83748facb8bf95b69faf342a40082cd263fa76aa03

Observation 167fd060-4958-473b-b1af-9bbd11db0d31 · inbound

DC-Scene: Data-Centric Learning for 3D Scene Understanding cites this paper.

DC-Scene: Data-Centric Learning for 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:26:51.975970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:26:51.975970Z digest=sha256:8d068b7f3695257bf0b99a2e9b0a84169039f38b88bac4e28199fb3169c3d999

Observation 3037b647-cbac-4793-bddc-87e095cc4d53 · inbound

CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation cites this paper.

CoNav: Collaborative Cross-Modal Reasoning for Embodied Navigation SQA3D: Situated Question Answering in 3D Scenes

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:55.344695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:02:55.344695Z digest=sha256:9cbd4ae1c4b9c7506647244009ad588be7302ec696ba8b188034244a4a96e96c

Observation c627b4dc-78c6-4e8b-a85c-543e476b10c9 · inbound

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction cites this paper.

VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction SQA3D: Situated Question Answering in 3D Scenes

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T12:57:17.925336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T12:54:01.013242Z digest=sha256:e1ca0d759016577dcad72f711f29fae3da89cfffa48bf63fe7e3c9be29d9a51b

Observation 0c12676b-a1a9-41a5-b575-7c99e3bc2948 · inbound

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts cites this paper.

Uni3D-MoE: Scalable Multimodal 3D Scene Understanding via Mixture of Experts SQA3D: Situated Question Answering in 3D Scenes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:46:13.614106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:46:13.614106Z digest=sha256:8ceedb83548640c229a5ca9bd299c2a84202bc63bdc87bfd846c87b57c0d61f9

Observation 35281398-dec0-4f5a-9f3e-b9ae96e21f3e · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence SQA3D: Situated Question Answering in 3D Scenes

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:34:36.901724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T08:34:36.824053Z digest=sha256:a5971accf409fcab99354e54ae5f70dea14e31db63b0fdd1ff4668e2307ab3cd

Observation b1371454-5560-4685-be67-40b01adbe23b · inbound

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence cites this paper.

Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial Intelligence SQA3D: Situated Question Answering in 3D Scenes

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:00:51.433617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-22T00:59:13.826054Z digest=sha256:ea5d92e76d92ba0bdec7059131e5d77cc7a1554574c896f735eb6896d205a53d

Observation 5e872cd2-3e60-48b5-b16e-e9c2b767fd75 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces SQA3D: Situated Question Answering in 3D Scenes

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:13.281354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:13.281354Z digest=sha256:d9f83aa6c115c45d909eaf796c93e5a1251ec11496147e64f71bdde1f5beffcd

Observation 7ce951ce-6584-460d-a315-39895b64c3f3 · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? SQA3D: Situated Question Answering in 3D Scenes

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:07.844021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:07.844021Z digest=sha256:1e39c274b3a004390800107216c4aa5cb4185aa5179a34dd551fa95f039a9b96

Observation fd4dec3b-1a3b-4b5f-b142-d4d15b470c63 · inbound

Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs cites this paper.

Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs SQA3D: Situated Question Answering in 3D Scenes

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:25:45.371810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:25:45.371810Z digest=sha256:f64855b461655386e4985ef0e76a4fa2b8aa5e334ae770675ff0cb91d3dc1c5d

Observation 27d46f3b-d777-4076-b839-5e64e9a65f93 · inbound

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation cites this paper.

Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation SQA3D: Situated Question Answering in 3D Scenes

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:53:00.971816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:53:00.971816Z digest=sha256:4fa0854543c28e7f3bd8210235b12283f9cf6179420407a26c89afe5521953de

Observation c05b8ad3-7462-4d43-a348-e4e2897b91a0 · inbound

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding cites this paper.

Fast3D: Accelerating 3D Multi-modal Large Language Models for Efficient 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:02.275383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:02.275383Z digest=sha256:b79be554fc977ed235f003a63e3e51370701b843710ec7d6624bc65bbbb47dca

Observation 68b78ee9-4bc6-4ab9-9f87-2c5e4c9eaa73 · inbound

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models cites this paper.

Spatial 3D-LLM: Exploring Spatial Awareness in 3D Vision-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:16.929007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:16.929007Z digest=sha256:78c255df1e3c2e2a4adf5e5567f98930e0fdb9da182423bd92782b7c58e982fd

Observation 16a7dcca-1648-4e11-a8d9-68a24647f1d1 · inbound

Reinforced Embodied Active Defense: Exploiting Adaptive Interaction for Robust Visual Perception in Adversarial 3D Environments cites this paper.

Reinforced Embodied Active Defense: Exploiting Adaptive Interaction for Robust Visual Perception in Adversarial 3D Environments SQA3D: Situated Question Answering in 3D Scenes

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T18:19:07.226106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:19:07.226106Z digest=sha256:c71056109130aa9ec0fe91976f5fd6bc7439664eb0020bfa6ac4359c5b8b2855

Observation 6fe358d8-c9ab-4143-9eb9-a6092229c0b9 · inbound

NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding cites this paper.

NeuroVoxel-LM: Language-Aligned 3D Perception via Dynamic Voxelization and Meta-Embedding SQA3D: Situated Question Answering in 3D Scenes

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T13:48:03.529117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:48:03.529117Z digest=sha256:66a015b9f4453fe700b502be8af2a3e3c7fab1992db7fd77e47c4ada66c8034f

Observation db95da7b-c382-4b00-a051-48f98fe660d1 · inbound

Agentic Services Computing cites this paper.

Agentic Services Computing SQA3D: Situated Question Answering in 3D Scenes

Reference 117

Resolution
unresolved
no resolver link, observed 2026-08-04T14:41:50.768259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:41:50.768259Z digest=sha256:6e5754f87347f9ba99096fbc38ffdf23fb03696a4004a4c3d25c542bb49d8625

Observation 959ea317-4073-4a7c-9901-74d0171b7331 · inbound

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert cites this paper.

GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert SQA3D: Situated Question Answering in 3D Scenes

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T11:38:13.152895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:38:13.152895Z digest=sha256:a892eb96a661d3759e329b262c8cac1f0027ebbb36cd9d5145342abd53e2dd2e

Observation 4110447d-fc49-4567-9278-45ac00f965b6 · inbound

MiMo-Embodied: X-Embodied Foundation Model Technical Report cites this paper.

MiMo-Embodied: X-Embodied Foundation Model Technical Report SQA3D: Situated Question Answering in 3D Scenes

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:42:05.711355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T20:40:54.096289Z digest=sha256:2dc7855ec5eb64a5d5e8169b2a173133e3bd6d8c744f1c5a86d6390eb7b49095

Observation b197cc77-cd8e-4bac-8457-6fb4a9deafa2 · inbound

POMA-3D: The Point Map Way to 3D Scene Understanding cites this paper.

POMA-3D: The Point Map Way to 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:30:11.613552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T20:27:27.347592Z digest=sha256:df86e4fde9ff58a5678515ff46baac27ae59603ec90248670e37c28975f52a00

Observation 0dc1fc8b-b746-4df7-a58d-114e42ef3a65 · inbound

Vision-Language Memory for Spatial Reasoning cites this paper.

Vision-Language Memory for Spatial Reasoning SQA3D: Situated Question Answering in 3D Scenes

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:34.304977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:34.304977Z digest=sha256:70e62800e675cba3277ed3ff9e1ed3166313b978f02fc42bcdde6329b245dac0

Observation a5cd0976-4699-43a0-9fbc-cd9b660cc7f7 · inbound

SpatialMosaic: A Multiview VLM Dataset for Partial Visibility cites this paper.

SpatialMosaic: A Multiview VLM Dataset for Partial Visibility SQA3D: Situated Question Answering in 3D Scenes

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:43:20.617257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T19:42:25.626448Z digest=sha256:78e25b3ba173910e54e9a117f5020c0d21643e594e69cfb48f886c9f3ae6ec4a

Observation 28c0cfce-8169-4eb5-8926-21892d136bf1 · inbound

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding cites this paper.

Feeling the Space: Egomotion-Aware Video Representation for Efficient and Accurate 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:30:22.455334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T09:30:03.668178Z digest=sha256:edc98a59f25036de61aea61fee1341f6df5cfc6677fa2f10208377e7d6595323

Observation 6692311b-c692-414b-9e12-bbd1b51aa19b · inbound

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM cites this paper.

Chat-Scene++: Exploiting Context-Rich Object Identification for 3D LLM SQA3D: Situated Question Answering in 3D Scenes

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:32:59.431171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T21:31:00.764078Z digest=sha256:49302f26a46b6c92de2dbeb5d130c6197fee9367974b487f50d39a0289285ad8

Observation ae767bea-78e7-4d99-a4fd-658e6227ffd2 · inbound

RGB-Pointmap Pretraining for Unified 3D Scene Understanding cites this paper.

RGB-Pointmap Pretraining for Unified 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:23:17.922593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T21:19:49.421653Z digest=sha256:083461dc0a76d1c8127d493d2fa0d39539d8919caba83cdeb02c407efbc5d037

Observation 2a22fb0e-cdbe-4f8f-b2e3-4f2abead8eca · inbound

Geometry-Guided 3D Visual Token Pruning for Video-Language Models cites this paper.

Geometry-Guided 3D Visual Token Pruning for Video-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:51:09.960260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T05:49:38.346274Z digest=sha256:b1dbb5f705292a5ae3e658a6e39cd57fbc1a0eb35bc10554c6b1244523c1265e

Observation ef9c6ca6-02f4-4ae1-ab5f-df39ffd73df5 · inbound

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models cites this paper.

Uncovering and Shaping the Latent Representation of 3D Scene Topology in Vision-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.089512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T00:56:06.243890Z digest=sha256:ff7d7cc96ee5ca86edfc58b16132d90414d8890969db574f9bedff0297d7399c

Observation ef774b3a-3006-4e17-b46c-3109d04b72bf · inbound

Unlocking Dense Metric Depth Estimation in VLMs cites this paper.

Unlocking Dense Metric Depth Estimation in VLMs SQA3D: Situated Question Answering in 3D Scenes

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-20T19:23:40.988324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-20T19:20:04.468206Z digest=sha256:1a5d3ee85b937c70e3869b46b90bf61e117ca88ab40e165600d1ae61e49e8411

Observation 11710254-5529-4b0e-8a7f-ec29e86173f0 · inbound

Unlocking Dense Metric Depth Estimation in VLMs cites this paper.

Unlocking Dense Metric Depth Estimation in VLMs SQA3D: Situated Question Answering in 3D Scenes

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:51.100352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T07:54:52.926995Z digest=sha256:4fc947e8093ec7cab5a1afeeb6fe455ac82a570526bb1fd886e817c20a1949f7

Observation 357819a6-6195-4f6b-8724-839622430391 · inbound

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models cites this paper.

ArchSIBench: Benchmarking the Architectural Spatial Intelligence of Vision-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:59:36.535565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T04:55:40.370796Z digest=sha256:e5b0c8bcc25cfdeae299b092e363925c5eb4e2f86740f8f75ba462e8f64ffa7d

Observation 9f8efa9a-6140-4949-8425-2f04eabe57c1 · inbound

Extending Embodied Question Answering from Perception to Decision cites this paper.

Extending Embodied Question Answering from Perception to Decision SQA3D: Situated Question Answering in 3D Scenes

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:33:58.998391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T21:30:40.182958Z digest=sha256:bfb55e63d20ba028ab46529421577add03ca010e1d8e3a87a8e2f7a525b39f3d

Observation 65436e5f-a2d6-426a-9bb6-6f72fb166a94 · inbound

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)? cites this paper.

Seeing Isn't Knowing: Do VLMs Know When Not to Answer Spatial Questions (and Why)? SQA3D: Situated Question Answering in 3D Scenes

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.561481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T07:48:19.295578Z digest=sha256:7dfa886adaaa20045b54037980529aa910464aad343728919a815a0ade8b1a76

Observation c4294406-1598-4665-9516-753a97861a12 · inbound

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration? cites this paper.

Where to Look: Can Foundation Models Reach a Target Viewpoint Through Active Exploration? SQA3D: Situated Question Answering in 3D Scenes

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:12:24.824379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T17:08:59.111306Z digest=sha256:25fafec15e1abf1e78a4fb2ce3e3a98e40f29ad794b2b458537751d9a95e4ecd

Observation 26e71fa9-63e9-477d-9549-df1090713b02 · inbound

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation cites this paper.

Zero-Shot 3D Question Answering via Hierarchical View-to-Token Transportation SQA3D: Situated Question Answering in 3D Scenes

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:27.034459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-28T10:54:02.188634Z digest=sha256:45ecb5d890b7377709c2523684b1909594f69445c4364ee052f2318e824ef9b7

Observation 3290b45a-ea4f-4d17-ac26-f0b2dceadf3c · inbound

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models cites this paper.

Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:57.525718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T02:13:53.219095Z digest=sha256:0fd9d0f693ef93e21cd5156def86a0ea9890604690fa4338d4951a20680a3439

Observation 9a827449-de59-4e92-a137-7a6db53cd253 · inbound

PAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene Understanding cites this paper.

PAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:57.091368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T02:17:30.190068Z digest=sha256:9ed93a4aff21984ded2425ba459b387fc66e48435d52b0b03525428ce25566b4

Observation e3bbb6fe-fb9b-44d0-8d3b-509856a63db1 · inbound

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks cites this paper.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks SQA3D: Situated Question Answering in 3D Scenes

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.647837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:4e7494027c61c18f8a1640a1eabadc01aa48b32859e1d3072d35b684494f05b2

Observation 9db02e74-6048-4617-85ec-1b3aea2b3f17 · inbound

Intelligent Automation for Embodied Benchmark Construction: Pipelines, Embodiments, Simulators, and Trends cites this paper.

Intelligent Automation for Embodied Benchmark Construction: Pipelines, Embodiments, Simulators, and Trends SQA3D: Situated Question Answering in 3D Scenes

Reference 113

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T11:58:05.934971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T09:17:50.473747Z digest=sha256:624dc09fff2f7e00ffd73a027403ca4f6f931a411ba6ec02d59d1d002994fa45

Observation c57f703c-ac56-45fd-a4df-08ab3dffc33d · inbound

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI cites this paper.

ACE-Brain-0.5: A Unified Embodied Foundational Model for Physical Agentic AI SQA3D: Situated Question Answering in 3D Scenes

Reference 131

Resolution
unresolved
no resolver link, observed 2026-07-11T19:16:57.396710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:16:57.396710Z digest=sha256:476564c2282a0d4a00bb723d114ea384fa3fdd778ee8f21feae647b89f8fcb09

Observation 5a206202-79d4-4d82-9ccd-deef6721c0ae · inbound

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models cites this paper.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models SQA3D: Situated Question Answering in 3D Scenes

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.689493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-08T02:44:00.608590Z digest=sha256:cba5a17d2739f872a9ed9d3d83ec9069972797a806aa4c25375d726f65a447d2

Observation 6d7833d2-f947-402c-8f2f-9c41c8434e18 · inbound

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models cites this paper.

CAIRN: Cross-Room 3D Scene Understanding with Topology-Aware Large Multimodal Models SQA3D: Situated Question Answering in 3D Scenes

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T16:00:13.133298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:00:13.133298Z digest=sha256:910a9ac67c1d5d5a56b73813371388fdd2a94b85ce9f0c058280cd89e32503c4

Observation 6c66308c-2870-44c7-824a-b686389d25ba · inbound

SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts cites this paper.

SpaR3D-MoE: Adaptive 3D Spatial Reasoning from Sparse Views Meets Geometry-Inductive Mixture-of-Experts SQA3D: Situated Question Answering in 3D Scenes

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-11T01:37:43.718773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-07-11T01:29:46.709388Z digest=sha256:c757a89edac86cf71d69c780dd7d521b06ebba67448feb00419b9c28973c3919

Observation e3603877-e820-49b5-abad-1c6e8531855f · inbound

IMBench: A Benchmark for Intuitive Robotic Manipulation cites this paper.

IMBench: A Benchmark for Intuitive Robotic Manipulation SQA3D: Situated Question Answering in 3D Scenes

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T22:45:08.351210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:45:08.351210Z digest=sha256:0eccbf3f1b1b8a2bdeb4f9b9023e1092e87e690ab8840676f05c77b96154cbc9

Observation c77eab79-ef14-4e5a-92a2-79d6ed19107c · inbound

Data Pyramid for Embodied Manipulation: A Survey cites this paper.

Data Pyramid for Embodied Manipulation: A Survey SQA3D: Situated Question Answering in 3D Scenes

Reference 268

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.880583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.880583Z digest=sha256:a34cd961af00a344c12353a8e3da440253ae751a0a0efbb8a3ee52d1e8f6f282

Observation ba8df6b1-1262-4c46-9504-2301ce664a74 · inbound

Distill What RGB Can Recover: Privileged 3D Evidence for RGB-Only Vision-Language Models cites this paper.

Distill What RGB Can Recover: Privileged 3D Evidence for RGB-Only Vision-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T00:41:37.886561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:41:37.886561Z digest=sha256:4b74e23367411cb57ee240b72d267d787affdf9eecd7041ee40fa662a7b1bc97

Observation 1c97c294-4d7f-4147-b2f8-df8ad2e07620 · inbound

SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models cites this paper.

SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-15T15:08:10.235780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:08:10.235780Z digest=sha256:d8150f7cae6f479e9ff42bdcb47f0e71eae5ffd7a204b12dd962908ab071be84

Observation eaa2cf8c-ab49-4f40-9e22-9dd632d5699f · inbound

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding cites this paper.

Qwen-3D: A Generalist 3D Vision-Language Model for Spatial Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T04:27:28.908340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:27:28.908340Z digest=sha256:440d2931804509a6a388011047f924ec140aa078f66d21ee615802ef60605d3f

Observation 9883fbc9-cda3-4c97-9331-2b087a2e0870 · inbound

HiSC: Hierarchical Spatial Clustering Token Compression for Efficient 3D Scene Understanding cites this paper.

HiSC: Hierarchical Spatial Clustering Token Compression for Efficient 3D Scene Understanding SQA3D: Situated Question Answering in 3D Scenes

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:51:52.908285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:51:52.908285Z digest=sha256:05ded93a0f3ca33129a76c73b47c62e99a5278b1529ca48e992bf021fc7c1a21

Observation 32cd7efc-d531-4675-b0dc-53185f7009b0 · inbound

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models cites this paper.

Multi-View Relational Distillation for Spatial Reasoning with Vision-Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-12T15:33:29.481703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:33:29.481703Z digest=sha256:9abd5029027e0721b040ce26034379ad40a13444d0c6d55baa47896a6197532b

Observation 9109863f-f2b3-4b28-9562-c3096244ddf3 · inbound

CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting cites this paper.

CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting SQA3D: Situated Question Answering in 3D Scenes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T05:15:36.107424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:15:36.107424Z digest=sha256:f2bd7f8cc3c0bb2323dbdd3b278371187cc24d5b8c206fd4fbb39e50f866c169

Observation db4ce132-ef9d-4b62-91fb-afec3b931859 · inbound

CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting cites this paper.

CausalSplat: Towards Comprehensive Hierarchical Reasoning in 3D Gaussian Splatting SQA3D: Situated Question Answering in 3D Scenes

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T14:18:39.145114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T14:18:39.145114Z digest=sha256:8c15d7eb869a9e68cb323c457beeba1d83db7b1655921f7d51c2709c5f12fce7