Pith. sign in

Paper Citation Record · LEDGER

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models

As of 18 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 6 inbound Pith citation observations for arXiv:2505.04911.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.04911 v1

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:22:03.380823Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-02T14:16:39.649823Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:57.806668Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy32
  • unresolved24
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e86eb903-a87a-4e2f-9bf5-fc7b6fbca0c8 · outbound

This paper cites Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Referit3d: Neural listeners for fine-grained 3d object identification in real-world scenes

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.307473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.135735Z digest=sha256:cc27342ec84266d3e1b28fe9ac55443a3dfdc78f6692e39d506418175b3ce587

Observation 1d12989b-4334-4396-9791-e2829238a99f · outbound

This paper cites Scanqa: 3d question answering for spatial scene understanding.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Scanqa: 3d question answering for spatial scene understanding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.293839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.140661Z digest=sha256:52e94a82d799b7a7440da4fe8e9a836068d27470f44d17b8d3b286ede3823926

Observation 9605b7ab-546f-4a40-87e2-f74e0f8f9b1d · outbound

This paper cites Language models are few-shot learners.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Language models are few-shot learners

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.145122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.145122Z digest=sha256:b2825ad85734c4fa52aa442cef5085b72b406d8965ff0d5e652a0b4556cde392

Observation 12bd605f-b8da-4fb9-8a85-6ae7a8c52104 · outbound

This paper cites Scanrefer: 3d object localization in rgb-d scans using natural language.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Scanrefer: 3d object localization in rgb-d scans using natural language

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.270591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.149341Z digest=sha256:9589c7784b44cfab329cd7a693d9c3062506642f54e61abff49ab3661aa6e381

Observation c77c582b-ceee-4e72-acfd-2c9c91dd2bb9 · outbound

This paper cites Language conditioned spatial relation reasoning for 3d object grounding.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Language conditioned spatial relation reasoning for 3d object grounding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.154205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.154205Z digest=sha256:e632e80f8b5d1e90736bea00baed0b78a82aaabaf6786af30c48b6e5d711b5ab

Observation a3a7b681-8ca4-4306-b2c5-817da5cf11a2 · outbound

This paper cites End-to-end 3d dense captioning with vote2cap-detr.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models End-to-end 3d dense captioning with vote2cap-detr

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.248135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.158388Z digest=sha256:d6402f63e3a9758373667ab14f890d11de6841071d5b5f49289c92c3fc625717

Observation 62a696d9-2606-4d80-8f36-054def1b838e · outbound

This paper cites Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Ll3da: Visual interactive instruction tuning for omni-3d understand- ing reasoning and planning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.234010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.162684Z digest=sha256:af7a3546b2f399b0419dade8efcb7cf7d61967e5d28d60b08b869ec3009898a2

Observation 19477877-9569-4c18-9160-a0cba3708f80 · outbound

This paper cites V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models V ote2cap-detr++: Decoupling localization and describing for end-to-end 3d dense captioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.166781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.166781Z digest=sha256:d060677740f0a20633e89bac13f2a5776fea285f573bf2fe4a2eae1094b4ee5c

Observation ddab9ca1-522b-407d-be4f-3b801bf96576 · outbound

This paper cites Scan2cap: Context-aware dense captioning in rgb-d scans.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Scan2cap: Context-aware dense captioning in rgb-d scans

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.210852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.171354Z digest=sha256:3c8d5c08e66900aa304ea99ff071b16137f2eb9090855b2b7d73d3bb6bfd86b6

Observation bb8a2e0d-e198-4779-91a4-10508fb01ecb · outbound

This paper cites Zero-Shot Video Question Answering with Procedural Programs.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Zero-Shot Video Question Answering with Procedural Programs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.175282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.175282Z digest=sha256:afe8f775b090bbceb7570714cdccc3c90b0c819795d3bdd56ab5070ba93377d8

Observation e62622ba-bc4b-46ce-8b0e-a6f2555e6e67 · outbound

This paper cites Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Chang, Manolis Savva, Maciej Hal- ber, Thomas Funkhouser, and Matthias Nießner

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.196609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.179532Z digest=sha256:c5f0c33320e6a9d586923521830d7ac9fe7c7ae6e32ad2b8a0ad7d5ee6cfc396

Observation 97757f96-34c9-402a-b4fa-55bc56daca0e · outbound

This paper cites Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface reintegration.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Bundlefusion: Real-time globally consistent 3d reconstruction using on-the-fly surface reintegration

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.182664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.183532Z digest=sha256:afbcd886f44eb49f5ef71bd2aaadb4bc3d5a7dcfe0134c1d93d5bd00116e3a95

Observation 73d615a7-3f14-4ce8-9fae-f1fe4c60a5d1 · outbound

This paper cites Lan- guage to map: Topological map generation from natural lan- guage path instructions.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Lan- guage to map: Topological map generation from natural lan- guage path instructions

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.169169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.187554Z digest=sha256:722b9e176c8441c96175a61ef100b6a3a046e90bec526c5a0ba5089a301dda41

Observation 6ab4f79a-5cf5-4133-b3a8-8ed308676a1d · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.191900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.191900Z digest=sha256:1c01f434250f23c3b4037fc3b9d893f2942c83dc6294331d747f68448e79cda6

Observation 2c85059b-8684-45bd-bf00-4b842fe50835 · outbound

This paper cites Gemini 2.0.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Gemini 2.0

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.155379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.196198Z digest=sha256:bb27ac7c07794807db5fa9bd037a8aa18735b714c13ed911aac972d69723dfef

Observation ea930b38-9c25-4aa5-b206-aa8d476712b7 · outbound

This paper cites 3d-llm: Inject- ing the 3d world into large language models.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models 3d-llm: Inject- ing the 3d world into large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.127711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.204752Z digest=sha256:411eb9d1be25e5c4e031495340e6a397bb45fc470dfb13d2f1c201df9c8cbbd5

Observation 6c524b60-e4dc-4b7c-9069-91dc333d6502 · outbound

This paper cites Vtimellm: Empower llm to grasp video moments.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Vtimellm: Empower llm to grasp video moments

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.209045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.209045Z digest=sha256:d6747f908671e3c73adb9a0eb38af689d0131f70183b079f1f0639b394191651

Observation bc4bb9a6-c3aa-46c9-8ffe-59e94d7d12ea · outbound

This paper cites Chat-scene: Bridging 3d scene and large language models with object identifiers.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Chat-scene: Bridging 3d scene and large language models with object identifiers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.103954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.213504Z digest=sha256:4a7ee853ad531fe9ac0005f670ecbe0f01054b8512877910032b785e3c45c991

Observation 42dea5ca-3c0f-4f88-bcd5-4a9a421e304e · outbound

This paper cites An embodied generalist agent in 3d world.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models An embodied generalist agent in 3d world

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.089392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.218020Z digest=sha256:3e34269798d51f7a4f71c3cf8c8989972b5ed5119a86632823c2455e1af86f0e

Observation 08f25081-872c-45dc-8d72-7b56624a3114 · outbound

This paper cites Multi- view transformer for 3d visual grounding.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Multi- view transformer for 3d visual grounding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.074082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.222848Z digest=sha256:db97aa55c059cbe865c9f1f9c4d8dae6698ebcc5e85a4433994cdee777f9c04b

Observation feba9d65-0368-403c-9c7d-88c1af251926 · outbound

This paper cites More: Multi-order relation mining for dense captioning in 3d scenes.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models More: Multi-order relation mining for dense captioning in 3d scenes

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.059085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.227014Z digest=sha256:54a043579ba6552c737877c7ac609e94c67f0a9a3c4bb1bff47ebe1bace57ddf

Observation 5fc57cdc-e54a-4e75-9c31-5352eecc6956 · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Chat-univi: Unified visual representation em- powers large language models with image and video under- standing

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.045298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.231008Z digest=sha256:1d25e99597a29dbe3887f36b2b8ff0a46a861c2da36d84cbe1f31c65b650df22

Observation 409bcd8b-76b4-459e-a068-4c1804712b8d · outbound

This paper cites Context-aware alignment and mutual masking for 3d- language pre-training.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Context-aware alignment and mutual masking for 3d- language pre-training

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.031214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.234828Z digest=sha256:5415e48da8810b3352d47e71e13e858c852a968dab4214082d441d57414f56df

Observation 5873b649-ecec-4ba3-b20b-2261d623c629 · outbound

This paper cites Large language models are zero-shot reasoners.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Large language models are zero-shot reasoners

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.016161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.238906Z digest=sha256:0d3fa192b520dce6e3cdc748d4957a239de34b650c8ba7ee047d7ec9bd93b472

Observation 22f3ec89-007b-4aff-a790-82c9e75bbebb · outbound

This paper cites VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.243306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.243306Z digest=sha256:3f3880de8072e3333573d1c58b0d3631bda1d4b36cecb17ca42539068f1b4a4b

Observation 2e386186-e9eb-44c0-be87-a92f0d04684e · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.247560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.247560Z digest=sha256:e0cb175ca9c1fee4970bc44bf57431e151f60bb5f8d5657d8a0f8a66c53b9e8e

Observation 6210f3c7-3150-4d1c-8a14-250ccb5f87f9 · outbound

This paper cites SQA3D: Situated Question Answering in 3D Scenes.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models SQA3D: Situated Question Answering in 3D Scenes

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.251882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.251882Z digest=sha256:09bd9010e20adf2355197a48f393ac0f8dbbbecd3b3e78b6c4b118021a0d4df4

Observation ecff8102-9088-4f9f-9f71-53aed34d39fa · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.256603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.256603Z digest=sha256:445864fa11b37c08c79501410679b1d867b1bd1a2effaac67fcaf2326c55b7dc

Observation d79439f8-c82d-4088-9bde-eab93ec71203 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:04.001310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.261081Z digest=sha256:967097d3b267d7ee64130094ca89ac67233b517fee6de1db67431e09c8bcf393

Observation a9d9fd3d-cb6b-4465-8f5a-79d7cdf080b5 · outbound

This paper cites Large Language Models Know Your Contextual Search Intent: A Prompting Framework for Conversational Search.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Large Language Models Know Your Contextual Search Intent: A Prompting Framework for Conversational Search

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.264822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.264822Z digest=sha256:d3186753efc4a243d724a79de07ee1afa65a74b185b6570b8674d5c6c23b5797

Observation bbb705d2-7184-44fc-ad5d-4e7927e915e2 · outbound

This paper cites Morevqa: Exploring modular reason- ing models for video question answering.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Morevqa: Exploring modular reason- ing models for video question answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.269541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.269541Z digest=sha256:e32de38d04544251bbd02a9244ae38dd002c581f3c542948dec91939c09d3a10

Observation 977261d6-c621-4116-9ad3-86628a6ccff0 · outbound

This paper cites an unresolved cited work.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:22:03.979809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.273740Z digest=sha256:d16f6328424d8c5c11f8ba26b3642e1a19b036654e9484b082801564aaa67ef3

Observation 838691c0-94e7-48c6-b3d8-a0537adfce99 · outbound

This paper cites Clip-guided vision-language pre-training for question answering in 3d scenes.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Clip-guided vision-language pre-training for question answering in 3d scenes

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:03.966685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.277716Z digest=sha256:b5b9601bd39aa9556f60bd4826c342cfa94da6e5fdfaf5de3664cbc0f85c2c31

Observation 83064e19-086e-483f-937f-3acc166af40d · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Learn- ing transferable visual models from natural language super- vision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.282100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.282100Z digest=sha256:3b260c38c304c65eaefbacbb215416f1638ec0a12efb04b3bbec1e36f9120b76

Observation 117e1819-8a16-454e-b764-f39d5d31578e · outbound

This paper cites Traveler: A multi-lmm agent frame- work for video question-answering.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Traveler: A multi-lmm agent frame- work for video question-answering

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:03.943681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.286034Z digest=sha256:791ab41e93e10b3c92946b8f1d34cdd63456f59a4a98afb1b03d04e4497c839c

Observation 617918b5-632b-4f33-b8ac-2610ed8fa4ec · outbound

This paper cites Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Droid-slam: Deep visual slam for monocular, stereo, and rgb-d cameras

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:03.930471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.290745Z digest=sha256:cd53174671e02d9a3e424eebd74a957b0c1accdfba1ed7a30475088f0ac42148

Observation 058ab9a4-c8b9-4231-a66f-c261116658f5 · outbound

This paper cites Four ways to improve verbo-visual fusion for dense 3d visual grounding.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Four ways to improve verbo-visual fusion for dense 3d visual grounding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:03.916949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.295087Z digest=sha256:bf7daa01f11367863f419ba36a2cd277bde3bd88015088116bd6ddf2756a839c

Observation 11f386b9-62d8-44ef-b37e-3e39a81c415f · outbound

This paper cites Vamos: Versatile action models for video understanding.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Vamos: Versatile action models for video understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:03.903477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.299492Z digest=sha256:3dd3f47d75815bbd784644e6d2007feba7f135ef799fdc3218919562d91866e3

Observation 3dec7b14-3bf8-4c59-a6e4-ea69e32e4772 · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Videoagent: Long-form video understanding with large language model as agent

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:03.889114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.304088Z digest=sha256:9f294f42cc4fbfa00b65d1a514a9e0ad3304579b9831d300f5562116dee93bd1

Observation 73ba2d20-d31d-4ac9-b98f-ffd17c52e0f2 · outbound

This paper cites Language models with im- age descriptors are strong few-shot video-language learners.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Language models with im- age descriptors are strong few-shot video-language learners

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:03.875883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.308641Z digest=sha256:2eecff57bb70de07002a4d8ad377d34d45358784f3acdb5e206288a64308e20c

Observation 1e267989-021d-45ee-acc4-cd5c5fb3e726 · outbound

This paper cites 3DRP-Net: 3D Relative Position-aware Network for 3D Visual Grounding.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models 3DRP-Net: 3D Relative Position-aware Network for 3D Visual Grounding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.313320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.313320Z digest=sha256:96fd2e7541ba26a8e0d2ffa8f6d643ad8535f77492d896cb4b08287d3ab1ef0b

Observation e35fece7-918a-49a2-a437-e9dea1860898 · outbound

This paper cites Distill- ing coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Distill- ing coarse-to-fine semantic matching knowledge for weakly supervised 3d visual grounding

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:03.861754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.317838Z digest=sha256:f1815b53730a6160a79a7419688b124c09c6014ee95f32e8d95c377e0660e28f

Observation 9be380f3-06e8-44da-92d1-7c819c976757 · outbound

This paper cites Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Chat-3D: Data-efficiently Tuning Large Language Model for Universal Dialogue of 3D Scenes

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.322405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.322405Z digest=sha256:6850ab460bbac18504aff3d2a3e3063ad8e798e678556171a040854a4b55b877

Observation 57f36b66-8254-4401-8b8b-d408b660de53 · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.327442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.327442Z digest=sha256:fb87187cca145c5e6e86344a06d3145cc6e22852dd071981b839c6b19c1dc2b5

Observation bdc57e75-91e2-4629-951f-215073b6f9c0 · outbound

This paper cites Emergent Abilities of Large Language Models.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Emergent Abilities of Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.332050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.332050Z digest=sha256:36cf4f4ffd51073c6bc8e8362251bb0ef16294a1181a27f4da3e890327ea6d5f

Observation f7c67614-691b-4381-b798-a71f74fdb12e · outbound

This paper cites Retrieval-based video language model for efficient long video question answering.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Retrieval-based video language model for efficient long video question answering

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.337025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.337025Z digest=sha256:3d2118dcd54ed6167286d93be6aa9aa46f9581778fd87b6f31addf45e59e3cd3

Observation a0477564-9838-4399-9a69-280a98f48bce · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Depth anything: Unleashing the power of large-scale unlabeled data

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:03.846832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.341891Z digest=sha256:1445314053e12a9497ce4e984eda8ef5ff590887b01fd01a97fb76c19dbb8991

Observation cd0b139d-0cd1-4d05-b8a0-31aff332231e · outbound

This paper cites Self-chained image-language model for video localization and question answering.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Self-chained image-language model for video localization and question answering

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:03.832157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.346055Z digest=sha256:fd82c72318b2267eca870dd6b02f17262960ab95051733313f5d90490e74a245

Observation a80e4e70-9be1-4353-ae75-2bd9952b0c13 · outbound

This paper cites X-trans2cap: Cross-modal knowledge transfer using transformer for 3d dense caption- ing.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models X-trans2cap: Cross-modal knowledge transfer using transformer for 3d dense caption- ing

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:03.816395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.350124Z digest=sha256:6e2a5ca797ea029cf1e80817758ca7ec70a208e9221515bcf5b75963fc46f7f5

Observation 5026e5e9-11f5-4cbb-aea8-2b920f8e913a · outbound

This paper cites 3dgraphllm: Combin- ing semantic graphs and large language models for 3d scene understanding, 2024.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models 3dgraphllm: Combin- ing semantic graphs and large language models for 3d scene understanding, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:03.801083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.354228Z digest=sha256:b6aac60a0608b030ba0127f8e6d732c44a1941cd3a9781a6ed0eae92ba5dfe44

Observation b4df0afd-5be2-4674-9677-a4a40005770a · outbound

This paper cites A Simple LLM Framework for Long-Range Video Question-Answering.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models A Simple LLM Framework for Long-Range Video Question-Answering

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.358265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.358265Z digest=sha256:b5fa3e06b7a304e662c8141bc307858433ed97e5f08f19b68353a647be8a4121

Observation 411d878c-4c64-4dfa-9de8-1318d557804d · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.362685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.362685Z digest=sha256:181d565b7a67be3cbc82e32676b86b064ebadbf2bc6c371df75a856d01dc8ab5

Observation 60b24c58-b124-4ebe-914c-b2bf96d06de6 · outbound

This paper cites Multi3drefer: Grounding text description to multiple 3d ob- jects.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Multi3drefer: Grounding text description to multiple 3d ob- jects

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:22:03.786738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.367022Z digest=sha256:fb3963b83b28530bafa8459e340b806d38f272fde507a00eee66d8481ca861d5

Observation fc122c2b-4a4e-443a-98a1-d9ac1e2b8279 · outbound

This paper cites 3dvg- transformer: Relation modeling for visual grounding on point clouds.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models 3dvg- transformer: Relation modeling for visual grounding on point clouds

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.371148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.371148Z digest=sha256:208a18cf55be35bccc1a6d03b619247ef616635b2435685b9ca9dc2f88cd4094

Observation 7e8b58bb-e232-48c0-9cc4-1cc51e9b19ae · outbound

This paper cites LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models LSceneLLM: Enhancing Large 3D Scene Understanding Using Adaptive Visual Preferences

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T23:22:03.376214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:22:03.376214Z digest=sha256:633262ac5d7f0f2dfe882c8f8f3507514a43b27f576dd8e39fb2ebb5cf9a0d52

Observation b11c3423-8cde-4b5f-a60b-8b72dcd0781d · outbound

This paper cites Is,” “Can,.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Is,” “Can,

Reference 56

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T23:22:03.760691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.380823Z digest=sha256:8bcb9bf9a5163e0021b4c89231945e1108bcddb1b269f1f85b31aa70c85fe126

Observation b5082d10-da5c-4036-998f-6a28e8f99f7a · outbound

This paper cites an unresolved cited work.

SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:22:04.141936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:22:03.200342Z digest=sha256:7b442b5599c08f03177aa4fb6cbcb7e1ef034ff92516cb929ca51f133705716f

Pith citing papers

Observation 8e256f1d-9d5c-47c7-a611-d24fdaf9da67 · inbound

Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models cites this paper.

Flame3D: Zero-shot Compositional Reasoning of 3D Scenes with Agentic Language Models SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:11:15.844978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T02:09:09.684786Z digest=sha256:85c783702e95df39a82f3fe520aa3aaa7394dc7001ff00d194faedf9449340ff

Observation cbafe699-c661-4036-a0bf-7b6d2b5e497a · inbound

SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning cites this paper.

SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:08:13.503564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T11:06:29.022994Z digest=sha256:39951be5287c1fa704b51b93bab9f919a81ac8d25d5ba2882ea15a015fc42580

Observation cffbfe24-0b3a-4e70-ac24-9e6f7f0ccbae · inbound

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning cites this paper.

Reason, Then Re-reason: Cross-view Revisiting Improves Spatial Reasoning SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:57:47.457896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T10:41:41.697216Z digest=sha256:91e0248a146c7081c605563e228a75968871cea5c055284d2f01daca3aa739a4

Observation e4675a4d-f28b-497f-9265-d25caba2b4a3 · inbound

Agentic Collaborative Cognition for Zero-Shot 3D Understanding cites this paper.

Agentic Collaborative Cognition for Zero-Shot 3D Understanding SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:39:57.808440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T00:22:06.082183Z digest=sha256:95009224c88e65cb9be03c147048b3f29ce20240d91c0c73d64049b1091c9f15

Observation b2aa3805-169c-484a-9519-0475390ff47f · inbound

Agentic Collaborative Cognition for Zero-Shot 3D Understanding cites this paper.

Agentic Collaborative Cognition for Zero-Shot 3D Understanding SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:59:52.650682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T05:37:41.407624Z digest=sha256:208e660bbd9cd356316786b28f585a16b5e1b3ce107e0f287a677d9e6169ce52

Observation ffce045f-5314-4807-ace6-774b901ddb76 · inbound

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping cites this paper.

OmniView-Space: Reinforcing Spatial Reasoning via Multi-Perspective Spatial Mapping SpatialPrompting: Keyframe-driven Zero-Shot Spatial Reasoning with Off-the-Shelf Multimodal Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:17:02.457657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-02T14:16:39.649823Z digest=sha256:2dc1f229803e30223cfce081b31281383a972184bb9474a89541781af78c7071