Pith. sign in

Paper Citation Record · LEDGER

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes

As of 4 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2605.31148.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.31148 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T22:47:46.542267Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact13
  • verified fuzzy0
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9defce82-bb36-428b-bbcc-9d34b1558c03 · outbound

This paper cites Placeit3d: Language-guided object placement in real 3d scenes.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Placeit3d: Language-guided object placement in real 3d scenes

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:ffff4c10b259896f3731c710f9f1ce5f3ab8aea07fe5cd37d8dd8fc3f56e9c4f

Observation c88463bc-9241-4132-9c8e-455242d8ecf2 · outbound

This paper cites Gemini-3.1 pro.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Gemini-3.1 pro

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:559337be7a7ede3613f9e6940205b13f57170f3311d8f4ce5a068c2e8d238a4b

Observation 395fc00e-7781-4061-a7c8-bff1666a8a2b · outbound

This paper cites Scanedit: Hierarchically-guided functional 3d scan editing.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Scanedit: Hierarchically-guided functional 3d scan editing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:fa0fe8c6fa28c311932877add8f0957e550ca65204aecc38054ec5c90380e99b

Observation 183ade7e-ec5a-4b58-a3d4-b0274a29fa85 · outbound

This paper cites Repurposing 3D Generative Model for Autoregressive Layout Generation.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Repurposing 3D Generative Model for Autoregressive Layout Generation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:52:45.649456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:d305a75262fa81438e534f7c016396012402be5ee71d3f3cd8ee8db1080fed1d

Observation cb976fd4-dd60-4d99-b28e-56a9e309d5db · outbound

This paper cites Layoutgpt: Compositional visual planning and generation with large language models.Advances in Neural Information Processing Systems, 36:18225–18250, 2023.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Layoutgpt: Compositional visual planning and generation with large language models.Advances in Neural Information Processing Systems, 36:18225–18250, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:80b09f38a706c934aedc8d0690e806e7ed3d43dce226a9f04f0b177437eee256

Observation 58e58408-69c2-4b73-bdf8-af4a224bb64d · outbound

This paper cites GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:52:45.651730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:5eca104b6c4b71ab34d1ce6801ed2cfe8c5c2ac00b7fd676b25c7d9344aeb95d

Observation 527711ea-d5d9-43bc-a49a-6af4ab464910 · outbound

This paper cites Fireplace: Geometric refinements of llm common sense reasoning for 3d object placement.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Fireplace: Geometric refinements of llm common sense reasoning for 3d object placement

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:df401b8821b63c6b324a2f11770bb510a30c22222fae5428435b4b07a182f8c9

Observation 889aa6f0-bf32-4c0d-a6bb-86e02a2f4be3 · outbound

This paper cites Spatial-dise: A unified benchmark for evaluating spatial reasoning in vision-language models.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Spatial-dise: A unified benchmark for evaluating spatial reasoning in vision-language models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:45.654255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:8be54ff23ff15e66e0b467f2d93a038653a1c529b9eebb0ffce4de7db792732d

Observation 4e520516-14ae-456a-a8f0-c68dd39449be · outbound

This paper cites Do you see me: A multidimensional benchmark for evaluating visual perception in multimodal llms.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Do you see me: A multidimensional benchmark for evaluating visual perception in multimodal llms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:e8b417b4dbdb6f8be37a81cca38375aaa761f1b5eeca725524491bedb6ae22d9

Observation 41cfa93a-29cd-418a-ab55-bd0bb52586af · outbound

This paper cites Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models.ArXiv, abs/2505.21500.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models.ArXiv, abs/2505.21500

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:45.656882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:20ea582eda9e597341d35a3194b9996d211e4fb89fbe10d9018db15bdd9a0be4

Observation 9c1c9787-8eb6-4c9b-8ecf-4579c5edccf8 · outbound

This paper cites Embodied agent interface: Benchmarking llms for embodied decision making.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Embodied agent interface: Benchmarking llms for embodied decision making

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:0545a65da4cb5efa19d72af5f8c7ddab5cdf2505ea5baa75c16372ab7ff8610d

Observation 34651ba5-d003-4d40-9e8f-04ed8d779eda · outbound

This paper cites Core Knowledge Deficits in Multi-Modal Language Models.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Core Knowledge Deficits in Multi-Modal Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:01.160739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:268ff5fc0472101dd29d52b8c71a7b21455055fe980ef1fa29b9cc433009375f

Observation 87be5ba3-0d53-4c46-a9cc-05ccccf50ec2 · outbound

This paper cites Spatial reasoning in multimodal large language models: A survey of tasks, benchmarks and methods.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Spatial reasoning in multimodal large language models: A survey of tasks, benchmarks and methods

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:01.155277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:9b6adc5117c79f2061e1862f2373e267c14797f50b4158e98106452d36536b46

Observation 402da64e-ed34-4f2a-83f1-6e795d61cbe1 · outbound

This paper cites Openeqa: Embodied question answering in the era of foundation models.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Openeqa: Embodied question answering in the era of foundation models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:eeee2b7108bc317f1f5596c9ed4b2811c23799c07fa1e9d1565450c736defa68

Observation 28fc1866-c84b-4107-b232-3b95bced9820 · outbound

This paper cites Gpt-5.4.https://openai.com/index/introducing-gpt-5-4/, 2026.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Gpt-5.4.https://openai.com/index/introducing-gpt-5-4/, 2026

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:13f70d0fdf4ceeee56945a5b0f00379ab204303272afed6d583b1ac64bbbb794

Observation c33b1710-2012-4dc1-8b5f-fb3176e6aa64 · outbound

This paper cites Qwen3.6-27B: Flagship-level coding in a 27B dense model, April 2026.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Qwen3.6-27B: Flagship-level coding in a 27B dense model, April 2026

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:298d3aeab02c2287c8221fd2c09a9d283d1aedcc9e72c0a21f4cd411fe6e09f3

Observation 8adb51d0-8853-4ffe-9bde-69537a55485b · outbound

This paper cites Qwen3.6-35B-A3B: Agentic coding power, now open to all, April 2026.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Qwen3.6-35B-A3B: Agentic coding power, now open to all, April 2026

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:027a8613bda9b77f9ac36320f26e489320979dce77bea09ad8d6c74c76d4e445

Observation d137ed69-6c52-4aad-b193-8cee47b04f23 · outbound

This paper cites Vision language models are blind: Failing to translate detailed visual features into words.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Vision language models are blind: Failing to translate detailed visual features into words

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:45.666857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:03e375f4f97f2665c02ad53095b95da6e8673716ee83a6a431c4df56d31d8470

Observation d5fbae3c-6a59-457f-bf9d-62b66a9d59b1 · outbound

This paper cites Does Spatial Cognition Emerge in Frontier Models?.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Does Spatial Cognition Emerge in Frontier Models?

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:16:01.154262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:467464526ccc62d4ab1914cfa71da5d83116fff67399fe8cb0e2e4789c8faf2a

Observation 7d6b5f4d-bd6e-4fd2-ae96-bc525ddac2ef · outbound

This paper cites Layoutvlm: Differentiable optimization of 3d layout via vision- language models.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Layoutvlm: Differentiable optimization of 3d layout via vision- language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:5184f5d3b2f371a21cdfa4e5c19bc73325e7f14e6cd5c051088fa2cd5bdc7725

Observation 3abb5795-8fb8-4a04-83e3-f8def5e15c0f · outbound

This paper cites SpaceVista: All-Scale Visual Spatial Reasoning from mm to km.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes SpaceVista: All-Scale Visual Spatial Reasoning from mm to km

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:16:01.156980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:6804121b79961ee987e771af9234d5773b63620158733fafac4f75b51f604a6b

Observation 39770aa2-0783-4818-b3e7-fbb41d9fc165 · outbound

This paper cites Kimi k2.5: Visual agentic intelligence, 2026.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Kimi k2.5: Visual agentic intelligence, 2026

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:d8539017d542706339ab5f0cf8df726d10fbc4a297555e1a61338a8017b80de0

Observation 062363fa-4c4f-4810-9295-82c8a5ebe7c1 · outbound

This paper cites Is a picture worth a thousand words? delving into spatial reasoning for vision language models.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Is a picture worth a thousand words? delving into spatial reasoning for vision language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:51f34b6fedab37d20699169aca12c16061e6ca6a3ac00db8f39275e160065afa

Observation eec17679-0fa4-4514-8190-c4f898e91af4 · outbound

This paper cites Raisecity: A multimodal agent framework for reality-aligned 3d world generation at city-scale.arXiv preprint arXiv:2511.18005, 2025.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Raisecity: A multimodal agent framework for reality-aligned 3d world generation at city-scale.arXiv preprint arXiv:2511.18005, 2025

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:45.667323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:647663807bdae107fefc32371ac7d0d6881289df7fe8a592bb3e5a8c5529504c

Observation d6a48103-f77f-449d-a06a-aa163513ea57 · outbound

This paper cites Embodied scene understanding for vision language models via metavqa.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Embodied scene understanding for vision language models via metavqa

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:4897edbf94b245ff8a9e194e2e4adb25e3a89308838e380bb5c4f39dca2be0cb

Observation 7a9bc7f3-b327-4305-a735-fd06fa2f5f9a · outbound

This paper cites Site: towards spatial intelligence thorough evaluation.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Site: towards spatial intelligence thorough evaluation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:ec324911b7c7f093292f1761b9449cc1d2447ea33ff1d83a4e9213eaf7b2244c

Observation c24158f4-458a-4e7e-9f64-0924bf6b9e9f · outbound

This paper cites Spatial457: A diagnostic benchmark for 6d spatial reasoning of large mutimodal models.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Spatial457: A diagnostic benchmark for 6d spatial reasoning of large mutimodal models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:961316b67e6f8b37ad032647e9802ed6e2671f8c348bc8fcbfd77e73ce498745

Observation b2325f0d-1813-431c-8b06-bdc3f95f59af · outbound

This paper cites Visual room rearrange- ment.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Visual room rearrange- ment

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:c5da519bdcf8d6f474a368ec5464de489b159ee6f48e79a4ad141baba8a3cc2e

Observation 35661b05-bba2-4e93-90c8-8ca555bb050c · outbound

This paper cites Spatialtree : How spatial abilities branch out in MLLMs.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Spatialtree : How spatial abilities branch out in MLLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:f3232e8248033012e2bcdcc31bcb71033b13a7387f76789b6f2a5a5dd1df2bff

Observation f1a7da44-6e1a-4a81-acc0-c9c1886e1067 · outbound

This paper cites CityCube: Benchmarking cross-view spatial reasoning on vision-language models in urban environments.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes CityCube: Benchmarking cross-view spatial reasoning on vision-language models in urban environments

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:45.661726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:760830318f5dda246c6fc1884093234ed51f51c84c1cdbaac457ea9869901904

Observation 4afe071f-2a15-43f4-8ea3-71ba977465bb · outbound

This paper cites Defining and evaluating visual language models’ basic spatial abilities: A perspective from psychometrics.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Defining and evaluating visual language models’ basic spatial abilities: A perspective from psychometrics

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:5ccce781bbcc5ff085ff72b635927ff4410f99b0ac4d4daeb5b2750c25de35a9

Observation dc9d0108-100e-4a2e-9837-0e8876ac5551 · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:4991a384350b93159101421d414e23663c6ad0c97304a91a8dc1d4e2686bea0f

Observation 75bb97db-8f79-4e4b-9735-31f12aaee948 · outbound

This paper cites Holodeck: Language guided generation of 3d embodied ai environments.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Holodeck: Language guided generation of 3d embodied ai environments

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:e6d601aa5e7127b5ec0b68965d24df5885a450bf6f490b360b1b67fc2fc2644b

Observation e3e8ec7c-3825-49d0-b65b-4c3078aea8ec · outbound

This paper cites Spatial mental modeling from limited views.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Spatial mental modeling from limited views

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:cd39f92fa4528d42d8b94014a0ad29bc90f44ff804169dd18418f603fee0b5c5

Observation e4e13388-034e-4a2c-9609-348c0e85f46d · outbound

This paper cites arXiv:2509.18905 (2025) 6, 9, 17.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes arXiv:2509.18905 (2025) 6, 9, 17

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T22:52:45.662040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:c72dbda3c368950ac088b004c4059878b53f88c6e78f79eaa772988b16defed5

Observation 15ad3398-d24d-48d0-8787-6dd67acdcac4 · outbound

This paper cites Et-plan- bench: Embodied task-level planning benchmark towards spatial-temporal cognition with foundation models.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Et-plan- bench: Embodied task-level planning benchmark towards spatial-temporal cognition with foundation models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:34eee61f20b3f4f086d93385a3b421b097b76a2465c95daa710673ba47936ada

Observation ecb79f32-8aa3-4e05-9154-9da03777d856 · outbound

This paper cites Theory of space: Can foundation models construct spatial beliefs through active exploration? InThe Fourteenth International Conference on Learning Representations, 2026.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Theory of space: Can foundation models construct spatial beliefs through active exploration? InThe Fourteenth International Conference on Learning Representations, 2026

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:80e4a742e98100781b26309a768481c33204cd0f5b81b0bcdcbbd3dce5cd1d94

Observation 6cdeccad-6bfc-4953-b7e2-69469410851f · outbound

This paper cites SPHERE: Unveiling spatial blind spots in vision-language models through hierarchical evaluation.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes SPHERE: Unveiling spatial blind spots in vision-language models through hierarchical evaluation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:a4fd9cee56ebe03ea5cb1c69fa1b8cf767430bca7d4a9db030c040211ea0c6eb

Observation 331979fc-5656-4070-abbf-1e0d018efdb4 · outbound

This paper cites an unresolved cited work.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:bcd063e19b292a7e3cd7208263bfeb8ebcfe8fccb0e6bf042082fdf16f41b841

Observation d1670f69-e4a6-4a2d-96d0-9f1da2a7c030 · outbound

This paper cites Cityeqa: A hierarchical llm agent on embodied question answering benchmark in city space.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes Cityeqa: A hierarchical llm agent on embodied question answering benchmark in city space

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T22:47:46.542267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:efcaecdb283b41cde5b5fa4cd7431a2ff6ebde34837786789c0d6fee66d221e9

Observation e187df64-a2be-4430-b8dd-084a787d2d91 · outbound

This paper cites 3d-layout-r1: Structured reasoning for language-instructed spatial editing.arXiv preprint arXiv:2603.22279, 2026.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes 3d-layout-r1: Structured reasoning for language-instructed spatial editing.arXiv preprint arXiv:2603.22279, 2026

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-28T22:52:45.664638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:8fc6ca9ace69d71a58d7aa4900a5b1fb907f499358147a49379fc406124a3599

Observation 2c07ae8b-5559-4680-848b-53e748b90322 · outbound

This paper cites InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts.

SpatialAct: Probing Spatial Reasoning-to-Action Capabilities of VLM Agents in 3D Scenes InternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layouts

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:52:45.669731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:47:46.542267Z digest=sha256:7d1c30d8c3be443c859f604c34fbaf808df34956869b5960ac98f43724a13eea

Pith citing papers

No inbound Pith citation observations are available.