Pith. sign in

Paper Citation Record · LEDGER

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning

As of 3 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2606.25360.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.25360 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-25T21:23:44.253416Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T15:45:43.483298Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact6
  • verified fuzzy0
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa56e012-6ce6-40a4-b48b-d5e80e006be1 · outbound

This paper cites π 0: A vision-language-action flow model for general robot control,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning π 0: A vision-language-action flow model for general robot control,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:8b858519d5f94460e9b30239a59f69832a33cd2e1e8920b3cbffceafb6fbc779

Observation 034c7b5b-f9de-4e31-82d0-fc7538c38318 · outbound

This paper cites OpenVLA: An open-source vision-language-action model,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning OpenVLA: An open-source vision-language-action model,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:150e336f0f81999e1c849cdaadcde2449e441688f2d7c7a251db7a98d19b7bd2

Observation c6e1ab2f-2136-4386-9e4c-5fea8e64ece3 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning SAM 3: Segment Anything with Concepts

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:20:07.157528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:f80eeeb135a205d582bd2e634dfa0c7715bc985965e9338df0694438a99497f7

Observation 74b78b5e-a78e-4084-aa7a-c55801a491c4 · outbound

This paper cites ALVINN: An autonomous land vehicle in a neural network,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning ALVINN: An autonomous land vehicle in a neural network,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:690a6275a27407b51becef55c7beec49629fed96bb638901f1650ed3e06481bd

Observation 3239d6f8-2b3f-4200-a385-b3431935b6c8 · outbound

This paper cites Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Robotic Manipulation via Imitation Learning: Taxonomy, Evolution, Benchmark, and Challenges

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:07.152003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:c34cc84541a8b65e9ca2614e070a3bb04a8380721a0ed9723871bd8a215fa8b8

Observation bb255f11-f354-47d4-84f7-7d2e69fc3472 · outbound

This paper cites Learning fine-grained bimanual manipulation with low-cost hardware,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Learning fine-grained bimanual manipulation with low-cost hardware,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:9a24a74833a142dcb867dfa258059d0f38e0be576625b457443aab69b2f67de3

Observation 6a2444a1-2ec6-403e-b4a2-ba672a804eca · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Diffusion policy: Visuomotor policy learning via action diffusion,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:fee543a7d0109ce9a821150e4bbfa2cb8ad6ea88d7a535de1e6a23eef8be21a3

Observation 0d921627-8eff-4cbe-a854-7f02fd358531 · outbound

This paper cites FiLM: Visual reasoning with a general conditioning layer,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning FiLM: Visual reasoning with a general conditioning layer,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:1ec9014bc3cf793e802bc214561900f58fe7ba2a911ea9ccab59ae4a7a78d6dc

Observation 988e34b6-fd3d-4425-9b4a-a98309529043 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Learning transferable visual models from natural language supervision,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:1784742358aa6ac21c8ad385e519716b09ab41341e04558aab5b6c4872b3acd0

Observation fb07ed91-7e93-4a20-a6f1-c0e0ac4f51c6 · outbound

This paper cites RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:20:07.149183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:06d70282b3362ead065206c772dcee2215ba0f63e241c2f4bc6b2cacbf485f9c

Observation 2c10252d-d4a1-421f-a607-2e71311693ec · outbound

This paper cites 3D diffusion policy: Generalizable visuomotor policy learning via simple 3D representations,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning 3D diffusion policy: Generalizable visuomotor policy learning via simple 3D representations,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:3a44c83beabfa8794185e6e6c38da093c4636ec6a7553eea0b7bd99fd7955955

Observation 3a72c63c-cce9-42f8-88fb-9f599b8d6e4c · outbound

This paper cites RISE: 3D perception makes real-world robot imitation simple and effective,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning RISE: 3D perception makes real-world robot imitation simple and effective,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:b6fa34eb25b933cd8bcb8b19931011a5d39c8ec802aad0e2a66511216d6e65cd

Observation ed1ddf3b-764b-42c5-9e64-83c472a43a57 · outbound

This paper cites RT-2: Vision-language-action models transfer web knowledge to robotic control,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning RT-2: Vision-language-action models transfer web knowledge to robotic control,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:29daef8b5bf5505bc0ecc15d8c2129e5de1737f99dd7348c3ac8bc2a434d6c50

Observation e85e71e3-8618-49c5-a24a-8e2cc70719ec · outbound

This paper cites Dense object nets: Learn- ing dense visual object descriptors by and for robotic manipulation,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Dense object nets: Learn- ing dense visual object descriptors by and for robotic manipulation,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:09c216e70fa01d84ff9210d9f4e73bcef1eba6160af59ef414a3b63f2f401cfc

Observation f67ef300-7f55-4893-b806-b931d961480b · outbound

This paper cites VIMA: General robot manipu- lation with multimodal prompts,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning VIMA: General robot manipu- lation with multimodal prompts,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:338e41555a0dcbb8f8a374bf6279a3fdc954a6f1d110a8102918dc6653908364

Observation 68320909-8301-4fd9-8cb8-0ffbc1d4656d · outbound

This paper cites MOKA: Open-vocabulary robotic manipulation through mark-based visual prompting,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning MOKA: Open-vocabulary robotic manipulation through mark-based visual prompting,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:2c9ab84de4f911d6b58aaab5885511585bf31f7354a0d2f0d54faec7c005f3be

Observation 5ad94462-be3c-472f-a3b8-350a8bda9438 · outbound

This paper cites ReKep: Spatio- temporal relational keypoint constraints for generalizable robotic ma- nipulation,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning ReKep: Spatio- temporal relational keypoint constraints for generalizable robotic ma- nipulation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:4a55e0b09fdd864b9235bfcf2d355dbafd3987ddf785d5bfe3addde11ccb9c2e

Observation 596185ce-c585-4a7d-952d-58411ee7f0e5 · outbound

This paper cites ProtCLIP: Function-Informed Protein Multi-Modal Learning.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning ProtCLIP: Function-Informed Protein Multi-Modal Learning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:07.154799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:4d0107c587fbbb990dfa539f99f51f01b9ebd680d98458bba3d71c388537f711

Observation 1001b5fe-7477-4d1d-80f9-c7768c154381 · outbound

This paper cites Spatial forcing: Implicit spatial representation alignment for vision- language-action model.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Spatial forcing: Implicit spatial representation alignment for vision- language-action model

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:07.143749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:1a656f1bd8466873a28b348c5e61a76923634d312fe1544ce36afe7e9ab68a13

Observation 26a75245-f23b-4850-902a-b68bd98c7382 · outbound

This paper cites More than a point: Capturing uncertainty with adaptive affordance heatmaps for spatial grounding in robotic tasks,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning More than a point: Capturing uncertainty with adaptive affordance heatmaps for spatial grounding in robotic tasks,

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:07.146577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:f255ea5c689eeca1c7a40801d1eebf4f535264668e8403fbd139119e5d2b5990

Observation 5d95502a-351d-48bb-abf4-29b917a635b1 · outbound

This paper cites SAM2Act: Integrating visual foundation model with a memory architecture for robotic manipulation,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning SAM2Act: Integrating visual foundation model with a memory architecture for robotic manipulation,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:f4d1585da8226966a027adbc6c7f841ea4eb5f12025f51f55008a393b866ed91

Observation 31ecf627-bde9-4b78-adf1-12745b35a77c · outbound

This paper cites Deep residual learning for image recognition,.

Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning Deep residual learning for image recognition,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-25T21:23:44.253416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-25T21:23:44.253416Z digest=sha256:a80c8b7270d21013465f675d9c7d2e0376479959b1377911529dcd5b7577f738

Pith citing papers

Observation eb82373d-f568-47a7-b9b6-aecb5a80be73 · inbound

KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation cites this paper.

KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation Decoupling Semantics and Geometric Grounding: Spatial Visual Prompts for Language-Conditioned Imitation Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-11T15:45:43.483298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T15:45:43.483298Z digest=sha256:8596cce678650ff004770608e3b7b1cd9a7ea2cf2e874757600f04705e89d41e