Pith. sign in

Paper Citation Record · LEDGER

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization

As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2506.06196.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06196 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:04:16.844847Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-17T06:02:31.638120Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T06:04:09.181694Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1b2a33bc-b820-4e31-aadc-427dcb8b041f · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:11.033917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:11.033917Z digest=sha256:8f3ac07fd9872b48fa97d720e4e000189ad3c3b3347931f40e12561d06276460

Observation 48b2af8f-acb5-4e04-8eb6-180a621a7250 · outbound

This paper cites RT-H: Action Hierarchies Using Language.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization RT-H: Action Hierarchies Using Language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:11.137684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:11.137684Z digest=sha256:0c760fc7b2355bdf04ec81586d71063296674eb8093751ea7c8ebf64c76ae2e5

Observation 04e9c524-ce43-4847-808b-a700a037d53b · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:11.265627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:11.265627Z digest=sha256:ea734e7707f74383e6bc14b7ca475008b61c123630917c346b2bfb840535389f

Observation ba69c4b4-c5c9-4d6e-8086-3653d68b3a13 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:11.406735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:11.406735Z digest=sha256:9b2e52f9da85e1b855a9298f12df6f03faf164d479d84cefc88b62d603306bc3

Observation e94d5aa8-d95d-48e7-9cf2-7ff61fd6eec7 · outbound

This paper cites Spatialvlm: Endowing vision- language models with spatial reasoning capabilities,.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Spatialvlm: Endowing vision- language models with spatial reasoning capabilities,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:20.371729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:04:11.538502Z digest=sha256:d88a26e90f84bb0ca697bde0c104d2ea4a117fec9b20b35a50daf317b1db9d45

Observation dd2bad11-c156-4878-a78f-7290705ce8b8 · outbound

This paper cites Towards Human-Level Bimanual Dexterous Manipulation with Reinforcement Learning.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Towards Human-Level Bimanual Dexterous Manipulation with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:11.859563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:11.859563Z digest=sha256:880d49d8106cf95cb997bd0a5fc972eafd32d7287e7e7a5f9cbb1d0b4b2b6406

Observation a32a3ecd-42e8-4a43-9a48-07bcfe48b8b9 · outbound

This paper cites Open x- embodiment: Robotic learning datasets and rt-x models,.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Open x- embodiment: Robotic learning datasets and rt-x models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:20.078039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:04:12.067288Z digest=sha256:9d196bcf62f44615db54a9d0843bbbb59c516688ed51382a0c47e2fe2e759df5

Observation ef7fdd6a-cfee-4057-a51e-109cb29c73be · outbound

This paper cites RoboNet: Large-Scale Multi-Robot Learning.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization RoboNet: Large-Scale Multi-Robot Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:12.307676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:12.307676Z digest=sha256:6aac256a5d2878e71c29606a07b09c6a6e4c5163a32676f54ed848872f064de1

Observation 96c4315e-3e64-40b0-83bb-a3d7c5e414e2 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:12.206866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:12.206866Z digest=sha256:0249d339cf28f32b4becb9a7187c4e4e84021d7f4be6b38ca76e4d5569248236

Observation e04d1035-3dcc-4afd-a3fb-70d09f167aff · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization PaLM-E: An Embodied Multimodal Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:12.847476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:12.847476Z digest=sha256:c01594b6137e9b8da613ec8967f3f79b0668deeb1721099c074d20c6af353cca

Observation 71e35357-4ddc-4a40-9075-d55046e143b0 · outbound

This paper cites Goal-conditioned imitation learning,.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Goal-conditioned imitation learning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:19.782268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:04:12.506939Z digest=sha256:baa5d770037614ba13368286174093492121140a2e381205386f62b183085424

Observation f8e0dc2a-dc31-4c6c-baa0-fe14db1d8574 · outbound

This paper cites Bridge data: Boosting generalization of robotic skills with cross- domain datasets, 2021.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Bridge data: Boosting generalization of robotic skills with cross- domain datasets, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:19.483094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:04:13.114027Z digest=sha256:6eb022a3c03349b63be91fecb92daabafb63e87e402dc819b395e1e3c3b0a3db

Observation 7bc2c181-e75b-4b2c-be2c-be5eedf95333 · outbound

This paper cites Zero-shot Task Adaptation using Natural Language.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Zero-shot Task Adaptation using Natural Language

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:13.256781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:13.256781Z digest=sha256:9f0a9a5faa7b07df1f84d13794a437da25efdb9aef3673deaa43d883e4790c09

Observation 98af9672-62c5-4c32-8e14-9b839cba2c5a · outbound

This paper cites Tenenbaum, Dale Schuurmans, and Pieter Abbeel.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Tenenbaum, Dale Schuurmans, and Pieter Abbeel

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:13.008703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:13.008703Z digest=sha256:36ed894bd0b3821e220c4d20660b03c687d2942b5aa609968bc5e567eab8c7eb

Observation bbab7591-1d4b-449b-800b-4d99923cb8a1 · outbound

This paper cites Scaling up and distilling down: Language-guided robot skill acquisition,.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Scaling up and distilling down: Language-guided robot skill acquisition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:19.083800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:04:13.641938Z digest=sha256:977d96a1f039c72620c9601d793e273f7489080707064e4a8951552bef579c49

Observation 77d7b19a-06fe-44d6-b545-81a41046f6ba · outbound

This paper cites Hierarchical Few-Shot Imitation with Skill Transition Models.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Hierarchical Few-Shot Imitation with Skill Transition Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:04:17.299566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:04:13.916780Z digest=sha256:ae14495e511ebdb20fd7893b8c998948ea86f40fddf4788289d1a65212a13d03

Observation af306728-a8d2-493e-bef3-befb557e26f9 · outbound

This paper cites Rt-trajectory: Robotic task generalization via hindsight trajectory sketches,.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Rt-trajectory: Robotic task generalization via hindsight trajectory sketches,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:13.396260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:13.396260Z digest=sha256:f730911e989f9a59ac3e478cda62442481e4ab65876f04580d23c9c58330dd09

Observation 3ab4ef90-0d42-499e-890e-2265f8db6e41 · outbound

This paper cites Inner Monologue: Embodied Reasoning through Planning with Language Models.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Inner Monologue: Embodied Reasoning through Planning with Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:14.300460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:14.300460Z digest=sha256:8ff924c4c52182e2d134d626b640b6dc1c04a6d21713961d5cc1c61e233adeda

Observation d1b67de7-7075-4b12-a775-bf24e5ce6635 · outbound

This paper cites Davi- son.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Davi- son

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:18.676762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:04:14.423649Z digest=sha256:3b2a85557ccab6614e201d2ebf0c8cd19a4dde0a19d381ba797c21407694da1c

Observation 301e6b42-0a85-4c6a-b92b-082e95a0d7a5 · outbound

This paper cites Scaling Up and Distilling Down: Language-Guided Robot Skill Acquisition.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Scaling Up and Distilling Down: Language-Guided Robot Skill Acquisition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:13.792473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:13.792473Z digest=sha256:60eee37fc80f3fa36a1d37dfa37ebab7063de2e863bfe108d0a00e2420568a9d

Observation 4499946b-792c-45f3-a8ee-e712764fee94 · outbound

This paper cites Egomimic: Scaling imitation learning via egocentric video, 2024.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Egomimic: Scaling imitation learning via egocentric video, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:14.632736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:14.632736Z digest=sha256:efb7aa7a4eaa4328b7f6873c6b6191401a54a135b78937711cdf1c4d05541a64

Observation eb8e90ef-5c1e-42c4-843a-b877d33a5a80 · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents,.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Language models as zero-shot planners: Extracting actionable knowledge for embodied agents,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:14.064861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:14.064861Z digest=sha256:0b9382759f0d43805165abd6ef492bef75cd48eeb89079dc151b31d3597b5028

Observation 64a33463-34f3-4274-a335-c93734941f1e · outbound

This paper cites Openvla: An open-source vision-language-action model,.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Openvla: An open-source vision-language-action model,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:18.399852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:04:14.874563Z digest=sha256:718e736b6471c8569e3513c615e7fe8654f389a90257d768cc4576376cf93b1f

Observation ddc84c13-9c55-4c89-8003-930e5b9430ff · outbound

This paper cites Code as Policies: Language Model Programs for Embodied Control.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Code as Policies: Language Model Programs for Embodied Control

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:15.152345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:15.152345Z digest=sha256:4296e058ef52fdf266864b2b870cf284ce2ca67e33e48c89f258133c13854460

Observation 094e848f-e007-4bf2-a841-b7ed10906f90 · outbound

This paper cites MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:15.276212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:15.276212Z digest=sha256:1af7eb433d05ebf648a26a3960636ab56405757e72b6d305eb94e13f3fd81d34

Observation d40ce355-46d7-4239-bb10-ed404cb7b352 · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning, 2022.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Bc-z: Zero-shot task generalization with robotic imitation learning, 2022

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:14.540458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:14.540458Z digest=sha256:bc6a94b66a2acdea8875f626fc246bad7acabb2fe9e8b501c2f145c031dcf6e6

Observation a2f8b3cc-c20c-44cb-ac76-a04d0f763938 · outbound

This paper cites Learning latent plans from play.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Learning latent plans from play

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:18.177456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:04:15.530183Z digest=sha256:148b8e7bfaa1d5bb98f198787fe34b85626f57fe9ac78cd06afeef4e45c3813f

Observation 8f362bea-e46d-4cea-a396-ee2e0a5f58dc · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:14.738359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:14.738359Z digest=sha256:4d7a595fdffbd6432d659e8ae7f025fec845d45f38fa7de0155a658e9cdc14d4

Observation 4aa998ca-47b7-4171-9228-3d09c27d40f0 · outbound

This paper cites Visual Reinforcement Learning with Imagined Goals.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Visual Reinforcement Learning with Imagined Goals

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:15.752484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:15.752484Z digest=sha256:d21adcfed3e2e187ec2bc19ce5dbd0e6b508b3e9e2162963f3ee48425e5487cb

Observation c258ec24-6ed2-4454-aacd-133ee025ff65 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization OpenVLA: An Open-Source Vision-Language-Action Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:14.996024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:14.996024Z digest=sha256:5806a1995f1fff10f2a7d09a9224c6808e490b959f7bf6f61150ceb1dac86a6c

Observation 9e66d2b2-3332-410e-b1d7-9f79c874c03c · outbound

This paper cites LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:16.013005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:16.013005Z digest=sha256:380c32d386964b80cb5608a745bade7175fea4da1e676c1d7230f62e6c04c058

Observation 4e0423d0-c961-43ae-944f-95f520db5b34 · outbound

This paper cites an unresolved cited work.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:04:18.027711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:04:16.122763Z digest=sha256:c2ab25081a01cf89672a4499303a30c0fc472ca5564ab0406aa112da5c89b9ce

Observation 587726b3-0826-4a81-afb2-4cc58c84d4e1 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:15.419455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:15.419455Z digest=sha256:6138b78a1f778977f8de58c6f94c76dd2a9cacf6a90f20b4781b0a7311dba3ed

Observation c30664b7-ea4d-4650-9262-cd75fe29750e · outbound

This paper cites Skill Induction and Planning with Latent Language.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Skill Induction and Planning with Latent Language

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:16.313044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:16.313044Z digest=sha256:716e7e568bd7aa82069fdd0637e6b3afbd12fdb9ee3f6d4b5681a071f96d16af

Observation 335c160c-5635-429c-8999-124272021dc1 · outbound

This paper cites Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:15.631738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:15.631738Z digest=sha256:91bd0a16fca82ae291c3329c14193e1e73d8cd0af4a9dace18a076b5e74e4165

Observation c3d5ce3a-d78e-4923-82b3-774bae28362f · outbound

This paper cites KITE: Keypoint-Conditioned Policies for Semantic Manipulation.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization KITE: Keypoint-Conditioned Policies for Semantic Manipulation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:16.485673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:16.485673Z digest=sha256:e6ad27c4fb70ab8bd4e49e7e901d8d5660477a6fdfc1d1e23ba831f120798cd4

Observation 0257a728-d1a0-4ecd-aa02-63c2951c0e17 · outbound

This paper cites RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:15.921883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:15.921883Z digest=sha256:ae8d7ec824a6f1f2617b7890cc709c6ade8aa1c6badad587baa45770b5b8bb71

Observation 769b965c-2dee-46ec-8b43-21acd8735ce3 · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:16.624067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:16.624067Z digest=sha256:df3f69716e48647af5b762c6d36b71a7a102a4e5fab9be45991e094e524d8207

Observation dd3f4caf-acb7-4b27-b876-81fd35ced15e · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:16.711246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:16.711246Z digest=sha256:ee4b2b532766817c706a2a430eb2d1f1e09308fe1c58809e8dab2671c08e2657

Observation 5bd28e61-5b4f-4f08-a081-3bfc54cceac9 · outbound

This paper cites Discovering motor programs by re- composing demonstrations.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Discovering motor programs by re- composing demonstrations

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:17.857100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:04:16.248457Z digest=sha256:74fd02ea8a8dbe61a923b3a1fdf38b60ad4ea720ec5fb8922fe296f064ededbc

Observation d2eee499-57e4-4e5f-8b69-f99ad7ae9380 · outbound

This paper cites Bridging Language and Action: A Survey of Language-Conditioned Robot Manipulation.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Bridging Language and Action: A Survey of Language-Conditioned Robot Manipulation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:16.844847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:16.844847Z digest=sha256:b0304835d0be7fc5b500e62a0f98a8b9e487877544e1707b2c7215a6d8b934e3

Observation 04ea60fa-6bd8-421d-b5fd-28f30bf65f3a · outbound

This paper cites RoboCLIP: One Demonstration is Enough to Learn Robot Policies.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization RoboCLIP: One Demonstration is Enough to Learn Robot Policies

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:16.384793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:16.384793Z digest=sha256:d91d4ebe3461cff1227d0d67fbacfa2cda8c22f2aaed9008e0fe1130a73a9533

Observation 5f61b262-e051-453d-ab56-4b5be6ccd34f · outbound

This paper cites Robust imitation of diverse behaviors, 2017.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Robust imitation of diverse behaviors, 2017

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:17.688878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T06:04:16.546319Z digest=sha256:554231678bdd0846b84e7d079391a44e0cd147bf9e02c0233440e1cebfe66030

Observation 7af48a98-fc57-4197-9039-0db98dc9e37a · outbound

This paper cites Zhao, Jonathan Tompson, Danny Driess, Pete Florence, Kamyar Ghasemipour, Chelsea Finn, and Ayzaan Wahid.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Zhao, Jonathan Tompson, Danny Driess, Pete Florence, Kamyar Ghasemipour, Chelsea Finn, and Ayzaan Wahid

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:16.776984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:16.776984Z digest=sha256:d35cad4d816993b403e92c7fc97e0e75bab2968790e9df2c30e5c3078b616d30

Observation 3337a975-e389-4e13-8b0d-db66e74bdc7e · outbound

This paper cites Goal-conditioned Imitation Learning.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Goal-conditioned Imitation Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:12.655115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:12.655115Z digest=sha256:36d417486c72d4c438bf52fe32613c3f2d4dcdf2904805e81017add6542efb98

Observation bc320820-7c37-4171-98ce-06cae42dc222 · outbound

This paper cites Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:14.191632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:14.191632Z digest=sha256:3739bdb9a20f890752fefe32960e714a84713274250742ddb56333774e693f46

Observation 7ccb436b-01f9-486c-a6c5-f9be3b136d8c · outbound

This paper cites RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:13.524282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:13.524282Z digest=sha256:eda58f2f81d372bdffbbc12661bbe25535b5bd0e109e58b65d3ebb5e1718499c

Observation a82e9a24-ed16-4b9a-a427-25cd479fe798 · outbound

This paper cites SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:11.687427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:11.687427Z digest=sha256:1613fb0a6435616632c437afa7bfa7d9133fc18987149df86a7ae55896efb528

Pith citing papers

Observation 2b3285ce-41ec-4537-81f0-c8891e2ac14d · inbound

Continually Evolving Skill Knowledge in Vision Language Action Model cites this paper.

Continually Evolving Skill Knowledge in Vision Language Action Model Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:04:09.185337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T06:02:31.638120Z digest=sha256:577195234e3b6731b8399b3e30b17b157f0e71f21fa9227f8cc7ed8723b3be2f