Pith. sign in

Paper Citation Record · LEDGER

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

As of 6 August 2026, this Paper Citation Record lists 94 of 94 outbound references and 1 inbound Pith citation observation for arXiv:2606.09669.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.09669 v1

Coverage vector

measured 94 of 94 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T16:35:14.099586Z

measured 95 of 95 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-31T05:01:26.556632Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

94 of 94 outbound references displayed

  • verified exact43
  • verified fuzzy0
  • unresolved48
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e19ec2f7-38ca-407d-87ce-cbeaff72d213 · outbound

This paper cites Introducing claude opus 4.5.https://www.anthropic.com/news/claude-opus-4-5, 2025.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Introducing claude opus 4.5.https://www.anthropic.com/news/claude-opus-4-5, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:031a90cf59be05fb00c22b6959163837ffe25488a60f38854fe3e0782c8998b2

Observation 9c66e2c6-d9d5-4130-b27c-cbd96bf48858 · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Scanqa: 3d question answering for spatial scene understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:55db9be8f6a954853e5ebf8a9d90ee51d5b7b908745843868658cc6cb644e68c

Observation b7d38452-163f-4b8e-893e-4b9bf688b039 · outbound

This paper cites Qwen2.5-vl technical report,.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Qwen2.5-vl technical report,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:d3745020a14a824a9a53b159b08b5ec47c15990d81e6c3913d1d7c17c8c33ffa

Observation f66a7bef-b29a-450e-888c-b40a0bdbe18e · outbound

This paper cites Qwen2.5-VL Technical Report.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Qwen2.5-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.698652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:98ba1929e797e1709c2979974ecf8823aaef6aeaf824d6e364815434a8e8a38d

Observation f421a06f-5405-4497-aec1-c48a90d8a332 · outbound

This paper cites AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks AgiBot World Colosseo: A Large-scale Manipulation Platform for Scalable and Intelligent Embodied Systems

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.701092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:11a8d8ec74096587e71a256c023840472f60f59ba228d12a91f0f31f218d6244

Observation 355f85ad-613a-4f62-b0ee-dd78a9934bf3 · outbound

This paper cites Seed2.0, 2026.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Seed2.0, 2026

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:b3fa516a7b880b116d9da3bb7deeaaa07e4c694f65e0e5b30e786026bcb337a9

Observation d520075e-3622-4749-b350-114a78798bb0 · outbound

This paper cites Scaling spatial intelligence with multimodal foundation models.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Scaling spatial intelligence with multimodal foundation models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.693069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:95a8fe7f0172e9a6388def41db7c8ec4ef4edcedee5a0b732191b3741c9ea7bf

Observation f886e9f8-1cbc-460d-b9f4-95836b54f964 · outbound

This paper cites Has gpt-5 achieved spatial intelligence? an empirical study.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Has gpt-5 achieved spatial intelligence? an empirical study

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.696168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:d180ae6c991412fc0378e3f49854dfdf1112a9396f2e0ecc18ca5b97347d038c

Observation b60b1f06-9489-4891-be2e-2bbb9c5e9fb1 · outbound

This paper cites Spider2-v: How far are multi- modal agents from automating data science and engineering workflows?Advances in Neural Information Processing Systems, 37:107703–107744, 2024.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Spider2-v: How far are multi- modal agents from automating data science and engineering workflows?Advances in Neural Information Processing Systems, 37:107703–107744, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:3a539da861f59f6d5019683371ec2c1ddec93881d2602b99b5f50244b215bcef

Observation 282f1d41-ddf1-4534-9ee6-c80143644165 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:9ce247af44401db1f019c9353cc19a337a2c8949c58bcaa5456ed76a474cf172

Observation 70485e45-96d9-4d73-af6f-bcdc83d9b780 · outbound

This paper cites Robogpt: an llm-based long-term decision-making embodied agent for instruction following tasks.IEEE Transactions on Cognitive and Developmental Systems, 2025.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Robogpt: an llm-based long-term decision-making embodied agent for instruction following tasks.IEEE Transactions on Cognitive and Developmental Systems, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:14743054ec541cc5ba3ecef3b8ce9d31ff1acd435d333ff59723d35c2f4c2bf7

Observation 4cc676b1-2f72-46a1-9140-b6a3be5cac7e · outbound

This paper cites EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.581129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:6ef0a3363f1067c8a397292b8db707753dc45cc3da24b5f4162aacc3bb89b188

Observation 95f6c51d-73b5-43a5-b907-d6a099730053 · outbound

This paper cites Gemini 3 pro best for complex tasks and bringing creative concepts to life.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Gemini 3 pro best for complex tasks and bringing creative concepts to life

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:dc2762d0f2e3c7704c41c3c033838382b6d03d1c77bd98f367e6a9e37678fdcf

Observation 1c5b0d4f-ea57-4159-9c19-ba9e32c7aee4 · outbound

This paper cites Proc- thor: Large-scale embodied AI using procedural generation.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Proc- thor: Large-scale embodied AI using procedural generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:538ab566b583f6dbef1ce1a3f63388195c83a98b5845eb737d20d6fd9bec51d1

Observation 490fab34-15f3-4c17-bb02-7553be026abf · outbound

This paper cites Carla: An open urban driving simulator.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Carla: An open urban driving simulator

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:844d47b3b025939beedd302105215d0bc065c339fd4e71b97026fe8176891dde

Observation 437053b7-29fa-4d35-af72-a8c6b716f2b6 · outbound

This paper cites Palm-e: an embodied multimodal language model.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Palm-e: an embodied multimodal language model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:f04ef496e775e373e7424ce82714bbe9342b5eac31f2296f7dd0dd0c39d936a3

Observation ab5f3e42-3b70-4dca-82ef-2a13b1f33127 · outbound

This paper cites Embspatial-bench: Benchmarking spatial understanding for embodied tasks with large vision-language models.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Embspatial-bench: Benchmarking spatial understanding for embodied tasks with large vision-language models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:ff812960e574c5d1431bea09e4e4b261eaa42930bb23fe24d8faaa48afa153ba

Observation 22d25092-d524-4f01-ab54-53ee1e75f7e7 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:0645f55a4aa8af4ee7259dfbabab1d9709710b796f4bd277413a27921a43c5e6

Observation 8b3900d9-8a1d-4a28-afd3-3a29b74e65a3 · outbound

This paper cites Vlm-gronav: Robot naviga- tion using physically grounded vision-language models in outdoor environments.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Vlm-gronav: Robot naviga- tion using physically grounded vision-language models in outdoor environments

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:e0cc7220627da09b49887a08c40aa3762dd8a982b631c467e72f4cb32141c262

Observation bad98602-e31d-453c-a353-99be9a3793c6 · outbound

This paper cites Minedojo: Building open-ended embodied agents with internet-scale knowledge.Advances in Neural Information Processing Systems, 35:18343–18362, 2022.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Minedojo: Building open-ended embodied agents with internet-scale knowledge.Advances in Neural Information Processing Systems, 35:18343–18362, 2022

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:38b104559360d8b7a4abcebf7531f357e410316c7453af937399e4aa4dc61745

Observation 4a70c945-6d58-4446-8d2b-ec4e86e6c48d · outbound

This paper cites Videoagent: A memory-augmented multimodal agent for video understanding.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Videoagent: A memory-augmented multimodal agent for video understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:8ed475a5994ed4cad3e0134cbc81da90789c5290ce06ad5fc481ff427a5c9a01

Observation 99aeff15-ca31-49d5-9d8f-ec285f719cc9 · outbound

This paper cites EmbodiedCity: A Benchmark Platform for Embodied Agent in Real-world City Environment.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks EmbodiedCity: A Benchmark Platform for Embodied Agent in Real-world City Environment

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.594785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:fc6728fc713ba1e1f775eaad2a524c7fbada3229f61e865be5ce3d89969b9433

Observation 904e2b40-3fee-4344-9257-618725ee1d69 · outbound

This paper cites Spatial reasoning with vision-language models in ego-centric multi-view scenes.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Spatial reasoning with vision-language models in ego-centric multi-view scenes

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.619892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:132f65292f224a3f201f406744dfbdbb8e3963c77b2e593111cf6488ab7145ba

Observation a66b4439-5621-4308-bdb7-0f430c52ce57 · outbound

This paper cites Seed1.5-VL Technical Report.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Seed1.5-VL Technical Report

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.632774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:0c1756537fa4ce8bfa1cfae1defb576ce0b406418dd639b6dcdacd5eb6cde715

Observation e558c513-fe5f-47b4-9d07-c872f0cba296 · outbound

This paper cites WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.597343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:fa94e4c1bdb6bdd573671522ed1ec716237a13c0da4ea4e7d9077db2c43ba9bc

Observation 3ec8fe6e-85ed-4eea-a2c5-4ad0290399f3 · outbound

This paper cites Cogagent: A visual language model for gui agents.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Cogagent: A visual language model for gui agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:2472c155dbaa0a8650d22ceff6b6efe153b02c348ede37c4c39739e948203b17

Observation f58e2aa4-ffbb-43da-b36e-23e0596ffde8 · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.638601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:432671d684789aa31bc08a2cc3bc1501ec02418ed83a23237c2eff507e897c14

Observation 201398b8-5fd3-43ff-9d82-33be8351ed0d · outbound

This paper cites 3d concept learning and reasoning from multi-view images.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks 3d concept learning and reasoning from multi-view images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:9766f0de028f5324b528ea3dd314a2f4ca41a12115f0b33663db8eb23ec9f928

Observation 23405645-c976-4165-bf18-35fe376620f6 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:a9a550194f838799d9daf4b67eed67a9280d72d60e890d42736a755ebd89657d

Observation c0a17fe4-11fb-4d73-9ca2-553efb806878 · outbound

This paper cites OmniSpatial: Towards comprehensive spatial reasoning benchmark for vision language models.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks OmniSpatial: Towards comprehensive spatial reasoning benchmark for vision language models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:c45f7e6e763089f44c5ac7bbca65ac4d5f5713e7598ccfbfaed7269269b28db5

Observation 0ba344a5-5b0b-45f8-adc4-bea208ebcc72 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Clevr: A diagnostic dataset for compositional language and elementary visual reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:ab92110de6b826c65eb240493f3f5d224b18268a4d2213f6ae33ec9b9cdc5615

Observation 0e36c439-f820-4daf-88cc-a87165ded012 · outbound

This paper cites Visualwebarena: Evaluating multimodal agents on realistic visual web tasks.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Visualwebarena: Evaluating multimodal agents on realistic visual web tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:aa5080dc0af42cc81eb2f2a71cbd2107a8b6e0a9d9f5d22b089fcc1cccbbffef

Observation a0271b8b-b6f8-42c8-89e6-5a7b5f579fab · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.665294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:3503cd0f30086dbeb19a71be55a5c40c2eef2b75e935fdabd41591f34283f095

Observation c5769039-0a8d-443a-be16-af765d62c9b4 · outbound

This paper cites AutoBio: A Simulation and Benchmark for Robotic Automation in Digital Biology Laboratory.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks AutoBio: A Simulation and Benchmark for Robotic Automation in Digital Biology Laboratory

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.574810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:f24e7b7986b6eff2e929bff93560f92c87f6fc45f9546e96f026f4cf2fa42138

Observation d883ac61-6992-4006-8511-d1557f1e8af1 · outbound

This paper cites iGibson 2.0: Object-Centric Simulation for Robot Learning of Everyday Household Tasks.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks iGibson 2.0: Object-Centric Simulation for Robot Learning of Everyday Household Tasks

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.653610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:94975fd3784142da77bf023c57c2275230e29a6b82c328a6eb864907947a81c2

Observation 5b0fd86d-b567-49ba-a5e0-9faa1c78e798 · outbound

This paper cites Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Embodied Agent Interface: Benchmarking LLMs for Embodied Decision Making

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:27:30.604976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:e857ffba1da7e0d9176c44b2586550f734130b0ad6e7df15aa2b335a9efc7c85

Observation 33d6404e-0cbb-4f70-9f7c-e847df3e248b · outbound

This paper cites M3DBench: Let's Instruct Large Models with Multi-modal 3D Prompts.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks M3DBench: Let's Instruct Large Models with Multi-modal 3D Prompts

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:27:30.616917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:a11977172b423e0b0577faeec29d7db84dc264d0e7b07f1d5d8442ed25fdfae6

Observation 4493eb09-8c01-489f-86f2-0d0d2b90c1a2 · outbound

This paper cites Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:75bfa750071e50af05719a803dcb57b7af51b20651a63c06a99d395b82c654fa

Observation 84dc1d62-ee12-427a-a3da-6415615ffdf3 · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.592032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:5750d7f3cfef9c0e8656f0b948c6e69398ad1c23914d858f60ba55d8e3a488d9

Observation fadd6697-439e-4a60-9b06-9cd7213e8244 · outbound

This paper cites Vila: On pre-training for visual language models.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Vila: On pre-training for visual language models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:37f4dddf8870f7a45cba8d8a32b47df3d7e0f6ae5f376e1298699e94a244c8fc

Observation 9cca9819-9e55-453f-830a-1775486ac4df · outbound

This paper cites Mmsi-video-bench: A holistic benchmark for video-based spatial intelligence.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Mmsi-video-bench: A holistic benchmark for video-based spatial intelligence

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.676637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:0322d6bad0c91044cc92a1bdc669641b5d434be4827c4312eef073e73c256996

Observation 7faaa7df-a02b-4549-9973-be32c17c1caa · outbound

This paper cites Ost-bench: Evaluating the capabilities of mllms in online spatio-temporal scene understanding.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Ost-bench: Evaluating the capabilities of mllms in online spatio-temporal scene understanding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.589508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:0bb882cb3a0f47cf2ea553706b354d64e9ff553feebd4f0f82ecc2cbe180a38f

Observation 390a2ffa-94be-478f-b236-e0d49b6011c1 · outbound

This paper cites Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 11:635–651, 2023.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Visual spatial reasoning.Transactions of the Association for Computational Linguistics, 11:635–651, 2023

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:948f4aa323bcad3ba4f63d8ce656823dd733d8d29b4f039e1e50055991cad566

Observation 9409299b-06ca-4713-86bb-867596d94f2e · outbound

This paper cites Llava-plus: Learning to use tools for creating multimodal agents.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Llava-plus: Learning to use tools for creating multimodal agents

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:1ebdc3c0bcd28a62f288a5a8e1a145bfe9e76d36e73979d39ac327d0d10b2998

Observation 4463b37a-7455-440d-b0ce-ba3746587120 · outbound

This paper cites Spatial reasoning in multimodal large language models: A survey of tasks, benchmarks and methods.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Spatial reasoning in multimodal large language models: A survey of tasks, benchmarks and methods

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.687723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:1e3f063e3b79f2c5b90d2228286a990067f4b7a8fdd8edf7f2558a14f1281b32

Observation 8cc528be-b179-4200-b688-23ea88a620e4 · outbound

This paper cites SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks SpatialCoT: Advancing Spatial Reasoning through Coordinate Alignment and Chain-of-Thought for Embodied Task Planning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.607933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:3383e52ae37abe06c4ffb68399fdfa585f97edf4b69204699f5a850b8a7d5416

Observation 2edc5e87-a03e-4724-8dd3-17f7e341d694 · outbound

This paper cites 3DSRBench: A comprehensive 3D spatial reasoning benchmark.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks 3DSRBench: A comprehensive 3D spatial reasoning benchmark

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:aa9581d0ac37571b5aeb7d5ff5dc3ff7198feee19dbd8aae0ebd727a6d738903

Observation e3bbb6fe-fb9b-44d0-8d3b-509856a63db1 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks SQA3D: Situated Question Answering in 3D Scenes

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.647837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:710ced11322fc1bd56829c39ced2ba4edc38422b3a475840265ffb9ab9c1d57a

Observation 52e7fa26-025b-4aa9-b779-206c23cbe5ae · outbound

This paper cites Introducing gpt-5.2.https://openai.com/index/introducing-gpt-5-2/, 2025.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Introducing gpt-5.2.https://openai.com/index/introducing-gpt-5-2/, 2025

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:8fa5d29580b81a6003eaea5a4b5449ef9d4c3d9edb82510d92fdaa806b71c973

Observation 7fc78741-d1e9-457b-8856-4f4cdde1eb20 · outbound

This paper cites Gpt -5.4 thinking system card, 2026.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Gpt -5.4 thinking system card, 2026

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:c75b61b246ca17fa55c6bf9b264d5dcaf3cfe6adf5550d73e9bf3d46d3c1bb54

Observation e57cfd88-4ad6-49ce-b2c1-0c25249a8f5d · outbound

This paper cites VirtualHome: SimulatingHouseholdActivitiesViaPrograms.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks VirtualHome: SimulatingHouseholdActivitiesViaPrograms

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-27T16:41:03.041882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:f15630db865e04577fee04877db85335993185edcb20c2293ccccca1fc9e96ad

Observation 716a930f-e5ec-463b-87f8-6fa3ecfa264c · outbound

This paper cites Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.670960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:c18cc5366df2c4a21adb9d0a26261a82ac15a8cb3185179022c2b4df4d06b5c2

Observation 8be52b7c-ecef-4d84-97a5-41d585b2d87f · outbound

This paper cites AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.659784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:9db58d799c77b6c065974bb0b360b41dab88284526bf0bf6702eb9d614924fc6

Observation f5e1b9ed-7880-456f-8f2f-c6e270ddbd85 · outbound

This paper cites Habitat: A platform for embodied ai research.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Habitat: A platform for embodied ai research

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:04caeb773e73461d93025a3cf74d092c84cef1a22c4d2acf91c2645d0f59cadf

Observation 7d6136a1-e4e2-4bf0-b42c-b023e1276659 · outbound

This paper cites ALFRED: A benchmark for interpreting grounded instructions for everyday tasks.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks ALFRED: A benchmark for interpreting grounded instructions for everyday tasks

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:446a834e949f176947e1ab33f75d2a31233ddbc1350979d29702cd10fc3a97ac

Observation 730ebfde-bd4a-4a2d-912c-fb186366aa0d · outbound

This paper cites OpenAI GPT-5 System Card.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks OpenAI GPT-5 System Card

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.684772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:423ee6e06d8dbfb846ca068a06e646c6ed7b349212606aef325c5533d4f3aa28

Observation 37545a29-3088-4d74-88eb-b82e4ada0810 · outbound

This paper cites Corso, and Eric Sax.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Corso, and Eric Sax

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.635879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:0e5319d8c073ad4eb8912f19f33cb575e3990f4ff59a2b4747d57099c0f7fb32

Observation 261c5356-9983-4d93-b8bb-a3696dff785a · outbound

This paper cites Gemini 3 pro: the frontier of vision ai, 2025b.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Gemini 3 pro: the frontier of vision ai, 2025b

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:2909262b3e1da2642cea023d8c30afaef4826ebbe11d382dbf05cb4d22e0728c

Observation 7e6567fb-826a-42d5-bf77-390d819a8e1a · outbound

This paper cites Gemini 3 flash, 2025b.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Gemini 3 flash, 2025b

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:8f2c8b66630b0652811105faaeac19f80c06a3b509d7ea71e147b5f26af9f19b

Observation 682748b4-3303-4ab4-969c-23004fdc38be · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Gemini: A Family of Highly Capable Multimodal Models

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.650444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:b8b3bbe4f8a80cf4049d79f98cbba80b9f3ff4e46df48d7307372712b21981da

Observation 0f504f61-0745-49ee-a3b9-7063321a9ba5 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.690088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:cc657f4d5b5ded075db299c80ead1769ed2788e493681ff32f7a1c6931b30dc9

Observation 3e1e8a3a-1162-4726-89c9-471d29186c50 · outbound

This paper cites Glm-4.6v: Open source multimodal models with native tool use, 2025a.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Glm-4.6v: Open source multimodal models with native tool use, 2025a

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:cbc4f21eb2b84177d65e7e206c2c6b7331974f16da744c8c4c73bf034315e18d

Observation c6e77ae3-82e5-426a-bb9f-e5f6c8b2ce78 · outbound

This paper cites Kimi-VL Technical Report.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Kimi-VL Technical Report

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.679664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:4cdb41dc7929244b17c2bbd25d795f2b8480b10e914ac0338564a734e9f4cfc2

Observation 608bc7a3-0356-43cf-9214-46a9e2e05c9d · outbound

This paper cites Kimi K2.5: Visual Agentic Intelligence.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Kimi K2.5: Visual Agentic Intelligence

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.673556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:537265492d3eac049a7e2c4678553897d8724d1c4111cb74e99a1a95247988de

Observation 76141400-ff9a-43f4-b460-a978f86dd3ba · outbound

This paper cites Qwen3.5: Accelerating productivity with native multimodal agents, February.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Qwen3.5: Accelerating productivity with native multimodal agents, February

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:065ea422afe530a2fdaa68d6bff6dbe717a6df3ece8db4a3535909236cca4aa2

Observation 03f265a7-e397-4ef6-9d3d-6aed2eec7870 · outbound

This paper cites an unresolved cited work.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:1e337881300485d05b8b2fadb4187ac061bd0c3ae90e92f4ed94a4bb6a907394

Observation 3b536508-5ced-4349-901c-1d1cfa4170e4 · outbound

This paper cites UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.662721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:1ac7d02530506ccbe90c8cf1dadfd079e63b2b1a02bea6711436c0dedded7f60

Observation fc543fad-a748-4ef5-882d-5a3642b2d122 · outbound

This paper cites Is a picture worth a thousand words? delving into spatial reasoning for vision language models.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Is a picture worth a thousand words? delving into spatial reasoning for vision language models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:7417fd9a7fb71ca7b88d7e64735c9d49883b1c04c86add94365a437b6ef2ea8c

Observation 435b9a5a-a0f0-486c-b535-5ecca914db45 · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.628006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:3fc3716b0478f9c6963f76f4389d400b3435460377ebabb3ee07c786d556073e

Observation 21213f33-8e68-4f38-aa59-41b378b34bf4 · outbound

This paper cites Drivemlm: Aligning multi-modal large language models with behavioral planning states for au- tonomous driving.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Drivemlm: Aligning multi-modal large language models with behavioral planning states for au- tonomous driving

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.668210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:f4af4bcf04283447fe838902c965dba6d49aa7cc765569af9b4d416039c3caa0

Observation 4966711e-80b7-4215-ae12-4ddab166b700 · outbound

This paper cites SITE: Towards spatial intelligence thorough evaluation.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks SITE: Towards spatial intelligence thorough evaluation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:de4755a55cd1584ccd23b7dee41ecd1eaadeb3db73ec11a857f1a53162063088

Observation b95a30a8-a914-455c-a66f-bed685e87999 · outbound

This paper cites Mobile-agent-v2: Mobile device operation assistant with effective navigation via multi-agent collaboration.Advances in Neural Information Processing Systems, 37:2686–2710, 2024a.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Mobile-agent-v2: Mobile device operation assistant with effective navigation via multi-agent collaboration.Advances in Neural Information Processing Systems, 37:2686–2710, 2024a

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.602431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:3b05f04a88c7138df8069f7d233543c173d562fd1859c9c6ef82a068107c616f

Observation 3c35a0f5-bc28-4e26-a091-49f7a7036520 · outbound

This paper cites SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks SpatialScore: Towards Comprehensive Evaluation for Spatial Intelligence

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.583555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:9512cec28de1e7af864d328722728c2b42a9d7112dce5fd4a723379342ba677e

Observation 30b9fc91-27f8-4b12-8a65-e9f40ce592be · outbound

This paper cites Gibson env: Real-world perception for embodied agents.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Gibson env: Real-world perception for embodied agents

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:6ae2804158190bb18bb0e09dc7943711745c871d51a950f373c4087f4a66a0ce

Observation 9384b297-522b-47a9-b872-689bf72615bb · outbound

This paper cites Sapien: A simulated part-based interactive environ- ment.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Sapien: A simulated part-based interactive environ- ment

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:e44edb70ed0bf96980e2b6e0a10a97ecb6dcaa42cf3b5f40c99714ba346ab123

Observation a0fa33ef-b1e2-44be-a90e-0a4a59c37316 · outbound

This paper cites Large Multimodal Agents: A Survey.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Large Multimodal Agents: A Survey

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.599891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:66acb2c6e48b4fb53ad360b87e91dd962fb2a4f8838fbede964b607f3825e936

Observation 23390021-1232-4d1f-8378-e34d27642b40 · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.Advances in Neural Information Processing Systems, 37:52040–52094, 2024.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.Advances in Neural Information Processing Systems, 37:52040–52094, 2024

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:25e7bedc0eb48d5148b251d2214e93d5867a19f8b6cf103890f65437e8c74243

Observation 8f60c5b8-8177-45be-a72d-bd80e57f6758 · outbound

This paper cites SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks SpatialBench: Benchmarking Multimodal Large Language Models for Spatial Cognition

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.644667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:37c6494acf7598cbb77008c1da10a8988b297e2a1993b1ad5fa1d9946dce2f1d

Observation f8eeb98d-69b3-4123-b257-cc669f1713e9 · outbound

This paper cites Pointllm: Empowering large language models to understand point clouds.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Pointllm: Empowering large language models to understand point clouds

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:1a94ce2cabf470377db285d188f401a783a791d21a0abe90db00a3c5299c9109

Observation 5e37bb14-aaaf-485c-837b-d1260cde98ac · outbound

This paper cites Qwen3 Technical Report.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Qwen3 Technical Report

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.656994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:dbd6211de77b3507900d3ff8ad06d227aafe19d3052d59096f3899684374c781

Observation 16fb1da4-b107-451d-a9e8-9f658e839d15 · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 81

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:a0eb9993a2ec8e5145bd639b1e09121d3aba083103cabf2a19de2a6468eeebf9

Observation e4bdfe65-5a21-4455-9f69-007194d073a2 · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.613186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:2879c884a868bdf5358a5ca1a1760b76c429002ef222b85c8617accb178926b1

Observation 9c0d1755-532a-4557-9ae5-bda7c82698d9 · outbound

This paper cites MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.641550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:7ef9b44ba6eb97010fe16d4bc9911dbe5b8c5cb048928b106fb7a61d7fdf5316

Observation 1cca129d-e0bc-4fd0-8e6c-0255cea06803 · outbound

This paper cites Drivearena: A closed-loop generative simulation platform for autonomous driving.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Drivearena: A closed-loop generative simulation platform for autonomous driving

Reference 84

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:0dea37d3a47f99d9e761d2f0f36d0d4ff1175b36283cc4abf8b5785bb6fdb614

Observation cffc97cb-c8d0-412c-af02-d858758f39c9 · outbound

This paper cites Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Meta-world: A benchmark and evaluation for multi-task and meta reinforcement learning

Reference 85

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:425913de66ad32f0b2306bc895b1224bd8ca577e6dfff72effeff0243ea2c4cb

Observation 3ddff7b5-55cb-44d7-8fe0-8bc68c1087ed · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.610598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:78f78983398f608d6d210047e200b3be7196096f9e5307adac90081823a32cd7

Observation 26ed2b1d-1f73-49cf-9e10-c90681737d50 · outbound

This paper cites VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks VLABench: A Large-Scale Benchmark for Language-Conditioned Robotics Manipulation with Long-Horizon Reasoning Tasks

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:27:30.622739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:7b61e959c155040340659a96845010c5364f12ee5723036ba1ed7046e367a576

Observation 2dd119b8-baaf-4135-a0e9-9f9daf40e8aa · outbound

This paper cites Thyme: Think Beyond Images.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Thyme: Think Beyond Images

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.578474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:cbccbcdb7b04a3591c2a21126295b888312b9999184c487bd10ffcae5d5dd2e9

Observation 675f39fe-051c-432c-ad1a-7f5a508eb2c5 · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 89

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.625350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:067b0752ea008dac75d105e678edddea65e0af526d98b9701f195aa44b680a99

Observation 976c17db-01b1-48af-825c-55d53e007d73 · outbound

This paper cites Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Towards Spatial Trace with Reasoning in Vision-Language Models for Robotics

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-07-03T01:27:30.630444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:0767dd514cbfeb17b0af1180397904a4830bfab6a8354f06a68ae964e1c76501

Observation 75ee56c0-637b-4033-af2f-4716bda53048 · outbound

This paper cites LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness

Reference 91

Resolution
malformed identifier
arxiv_id, observed 2026-07-03T01:27:30.682419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:e7cfaafc392745d95b9cda236d1639493ec55c832279c34ce7af037805a23354

Observation b870de9b-894b-492f-9077-6e0e1b349423 · outbound

This paper cites an unresolved cited work.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Unresolved cited work

Reference 92

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:bb88feac1806abdfd602456fa686d0cd2f81931cca82f5cf033a3808ce94be41

Observation 8cf7a7c7-f8ea-4dde-97ac-e4a64d326be7 · outbound

This paper cites As introduced in Section 2.3, this interface abstracts raw backend commands into high-level text primitives to form a unified MLLM-native action space.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks As introduced in Section 2.3, this interface abstracts raw backend commands into high-level text primitives to form a unified MLLM-native action space

Reference 93

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:37e13de7a7bfb0db90e63d8a97cae1e40f3576046850c3dead4782db5e0bb46b

Observation 1a4b8ee3-f5a1-47ee-8a26-aaa2958fef65 · outbound

This paper cites Step sizes vary by environment: AI2-THOR / Proc- THOR / VirtualHome use Small = 0.25 m, Medium = 0.5 m, Large = 1 m.

SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks Step sizes vary by environment: AI2-THOR / Proc- THOR / VirtualHome use Small = 0.25 m, Medium = 0.5 m, Large = 1 m

Reference 94

Resolution
unresolved
no resolver link, observed 2026-06-27T16:35:14.099586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T16:35:14.099586Z digest=sha256:01b068c39daf7231fb223c32bbd9ddbc79244cd06943ac67899ff841019f5e70

Pith citing papers

Observation 38384984-b82d-4c47-a2c8-d907d08dc60d · inbound

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models cites this paper.

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models SpatialWorld: Benchmarking Interactive Spatial Reasoning of Multimodal Agents in Real-World Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T05:01:26.556632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T05:01:26.556632Z digest=sha256:a6d5318debc39d0eae6de8eeee9c7851c633f56fddf4dd6e0fbc94ec6aa0f323