Pith. sign in

Paper Citation Record · LEDGER

Warehouse Spatial Question Answering with LLM Agent

As of 7 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2507.10778.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10778 v2

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:29:40.087802Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e2ed1f8c-2a8d-4720-84e0-5316589660b5 · outbound

This paper cites SpatialBot: Precise Spatial Understanding with Vision Language Models.

Warehouse Spatial Question Answering with LLM Agent SpatialBot: Precise Spatial Understanding with Vision Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T17:29:38.536084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:29:38.536084Z digest=sha256:7166a598742aa42b832ebafd5f1534c8d45a0a6322d74168faa05b104200aaa6

Observation 8546b2a5-113b-4722-93f7-8bacdf1ba08e · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

Warehouse Spatial Question Answering with LLM Agent Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T17:29:38.601314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:29:38.601314Z digest=sha256:afc754a2b991506dc4177401464b2a7443036d5b0d67aca607f0fe10e77ddefa

Observation 454134af-8b06-4101-84ee-8f8796bd87a7 · outbound

This paper cites Spatial- rgpt: Grounded spatial reasoning in vision-language mod- els.

Warehouse Spatial Question Answering with LLM Agent Spatial- rgpt: Grounded spatial reasoning in vision-language mod- els

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:41.903708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:29:38.675477Z digest=sha256:47291a56547ec715f2909c754e23987160c0cb8dab34975dba46715a2a7bc0b0

Observation 82dec2ed-f7d3-4dc7-b385-6488c149875e · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Warehouse Spatial Question Answering with LLM Agent Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:29:38.751572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:29:38.751572Z digest=sha256:e33f573fe1d602a179cba02a8456622fa88c0b67dff17ee83d4f6410cf3f3c9a

Observation b2839409-843e-4a79-8ea7-8a2741ae333b · outbound

This paper cites Deep residual learning for image recognition.

Warehouse Spatial Question Answering with LLM Agent Deep residual learning for image recognition

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:29:38.866567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:29:38.866567Z digest=sha256:1735db1ee06460cedb0e7c9458a5e217893923a60404f52ed1d928845aafc95c

Observation 91eac91c-5b43-4613-918e-ed4b68643fa5 · outbound

This paper cites ToSA: Token Merging with Spatial Awareness.

Warehouse Spatial Question Answering with LLM Agent ToSA: Token Merging with Spatial Awareness

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:29:40.213090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:29:38.951111Z digest=sha256:d0e7f34d09233b8d2fd3c09f47dcfe6f1ba9052db2303d92bf4e3e91ea9117fc

Observation 9ab03bd5-fe61-43cb-9a85-e8ea79c7a311 · outbound

This paper cites Zero-shot 3d question answering via voxel-based dynamic token compres- sion.

Warehouse Spatial Question Answering with LLM Agent Zero-shot 3d question answering via voxel-based dynamic token compres- sion

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:41.636975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:29:39.046690Z digest=sha256:d486e3f53f3d87608df993e4354e2e2a36b7e99b0490e2d9fb6fe26d7c11d8fe

Observation bea3ddf5-48cb-4953-a0a4-9a820e2facfb · outbound

This paper cites Embodied agent inter- face: Benchmarking llms for embodied decision making.

Warehouse Spatial Question Answering with LLM Agent Embodied agent inter- face: Benchmarking llms for embodied decision making

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:41.484786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:29:39.147806Z digest=sha256:e46036049c0b370a3baba71596e1b5489f5b819aff7e7cea4d31724d5ae136a6

Observation c0645438-3102-44c6-be13-edce72644785 · outbound

This paper cites Seeground: See and ground for zero-shot open- vocabulary 3d visual grounding.

Warehouse Spatial Question Answering with LLM Agent Seeground: See and ground for zero-shot open- vocabulary 3d visual grounding

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:41.340612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:29:39.240988Z digest=sha256:83dbf39fd28887683a60123aaf234a54e34c7b32105f6150857e507beddee939

Observation 95ded4ad-249c-406d-b44a-c81945463e3c · outbound

This paper cites Focal loss for dense object detection.

Warehouse Spatial Question Answering with LLM Agent Focal loss for dense object detection

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T17:29:39.339118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:29:39.339118Z digest=sha256:b3ad55bc80c3738007cdcc5cdff9afddf8cc50a6b529114be7ef0b151d16f498

Observation b91428a8-af37-48e5-8393-4014bd586534 · outbound

This paper cites an unresolved cited work.

Warehouse Spatial Question Answering with LLM Agent Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:29:41.171583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:29:39.441137Z digest=sha256:c3ae36dcc4f9d3b608c94c63376cf3a5eb3c9b86a5567f139e67d46438387e36

Observation 5fab6fd2-c228-419e-a7fa-02f60e1b36c3 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

Warehouse Spatial Question Answering with LLM Agent Videoagent: Long-form video understanding with large language model as agent

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:41.004358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:29:39.531046Z digest=sha256:c6885a47d2f6f2b8b519293638d26bd7118074301151bf9eef2cd551eada1051

Observation 0eb380ba-f903-4438-8451-a846145b6362 · outbound

This paper cites Vlm-grounder: A vlm agent for zero-shot 3d visual grounding.

Warehouse Spatial Question Answering with LLM Agent Vlm-grounder: A vlm agent for zero-shot 3d visual grounding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:40.857991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:29:39.602826Z digest=sha256:1fd2db0918d1fcb0c66dddc0ebf4742ef4ec6304f1cef7f319b075a5277cb7a9

Observation fdd43a2f-1baa-469b-91e6-839abcf414f5 · outbound

This paper cites Fouhey, and Joyce Chai.

Warehouse Spatial Question Answering with LLM Agent Fouhey, and Joyce Chai

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:40.724684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:29:39.706694Z digest=sha256:9f83629b806670f4ca5b18465bd93d3bf326347a484cb2796c872dd6b0750aac

Observation 8860a391-0e75-410b-af83-022e58746082 · outbound

This paper cites Visual programming for zero-shot open-vocabulary 3d visual grounding.

Warehouse Spatial Question Answering with LLM Agent Visual programming for zero-shot open-vocabulary 3d visual grounding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T17:29:39.837375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:29:39.837375Z digest=sha256:f34af8a1b82d275bd83e7e436d00fc1fde610b7e44ac7bc6a292ebeb14482acc

Observation 1d12a078-4ada-4793-ac93-cdb943b2dae1 · outbound

This paper cites Agent3d-zero: An agent for zero-shot 3d understanding.

Warehouse Spatial Question Answering with LLM Agent Agent3d-zero: An agent for zero-shot 3d understanding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:40.590524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:29:39.930124Z digest=sha256:b5d042f6b8f28b3c6ed7fe009b4a8e33dc3c0183f6c4db5cb942c9e80e22b636

Observation a4424dfe-9a38-4859-b3f6-9f00496230c5 · outbound

This paper cites See and think: Embodied agent in virtual environment.

Warehouse Spatial Question Answering with LLM Agent See and think: Embodied agent in virtual environment

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:40.473966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:29:40.024063Z digest=sha256:64810ef326f20e361f986245defe1bbf75bf26c9d6a413a54d1906a5b00c4e06

Observation bf329b4c-22fd-4676-8123-d4de0cb898f8 · outbound

This paper cites Video-3d llm: Learning position-aware video representation for 3d scene understanding.

Warehouse Spatial Question Answering with LLM Agent Video-3d llm: Learning position-aware video representation for 3d scene understanding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:29:40.349978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T17:29:40.087802Z digest=sha256:5136684ba8680ae14150019355486915cf2ae478a6d936ad95c724919977cae1

Pith citing papers

No inbound Pith citation observations are available.