Pith. sign in

Paper Citation Record · LEDGER

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization

As of 19 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 1 inbound Pith citation observation for arXiv:2506.06196.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.06196 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:04:16.844847Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-17T06:02:31.638120Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T06:04:09.181694Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1b2a33bc-b820-4e31-aadc-427dcb8b041f · outbound

This paper cites Do As I Can, Not As I Say: Grounding Language in Robotic Affordances.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Do As I Can, Not As I Say: Grounding Language in Robotic Affordances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:11.033917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:11.033917Z digest=sha256:95e80b49299208d5a1462be1f48dee042b377de75d446dbc3b026a2616377e49

Observation 48b2af8f-acb5-4e04-8eb6-180a621a7250 · outbound

This paper cites RT-H: Action Hierarchies Using Language.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization RT-H: Action Hierarchies Using Language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:11.137684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:11.137684Z digest=sha256:ee8cd06d7c444224f6f1967dddf0b483d9e3b3abdcfe5e13b3095228adb41a93

Observation 04e9c524-ce43-4847-808b-a700a037d53b · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:11.265627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:11.265627Z digest=sha256:9fb11958a0e9ebbb675b5f40b6eec7efce27b4ff59e25f12245e9d42b47de24e

Observation ba69c4b4-c5c9-4d6e-8086-3653d68b3a13 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:11.406735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:11.406735Z digest=sha256:fb5da0678fc4cc508a5241cfbde2d0dbac429727d0abe26848f488d8a4f98057

Observation e94d5aa8-d95d-48e7-9cf2-7ff61fd6eec7 · outbound

This paper cites Spatialvlm: Endowing vision- language models with spatial reasoning capabilities,.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Spatialvlm: Endowing vision- language models with spatial reasoning capabilities,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:20.371729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T06:04:11.538502Z digest=sha256:6e1cc79837d77d9d4ecd17e1ccb054c6dc59d897384160f208c3fece2e8216d8

Observation dd2bad11-c156-4878-a78f-7290705ce8b8 · outbound

This paper cites Towards Human-Level Bimanual Dexterous Manipulation with Reinforcement Learning.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Towards Human-Level Bimanual Dexterous Manipulation with Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:11.859563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:11.859563Z digest=sha256:ca9f27c59541adf4b80a401357f75a9d4c9f3805026d9a0f1148f297abc1a6d8

Observation a32a3ecd-42e8-4a43-9a48-07bcfe48b8b9 · outbound

This paper cites Open x- embodiment: Robotic learning datasets and rt-x models,.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Open x- embodiment: Robotic learning datasets and rt-x models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:20.078039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T06:04:12.067288Z digest=sha256:22ef22fb735e4718be369837b87b4542da6167c4f9c797be8554ab65578363d4

Observation ef7fdd6a-cfee-4057-a51e-109cb29c73be · outbound

This paper cites RoboNet: Large-Scale Multi-Robot Learning.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization RoboNet: Large-Scale Multi-Robot Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:12.307676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:12.307676Z digest=sha256:21de5c072e3c05393c475e03f5663db8083a2f5ffa45bdc422ebcaa46e0b82ff

Observation 96c4315e-3e64-40b0-83bb-a3d7c5e414e2 · outbound

This paper cites Open X-Embodiment: Robotic Learning Datasets and RT-X Models.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Open X-Embodiment: Robotic Learning Datasets and RT-X Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:12.206866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:12.206866Z digest=sha256:7eb500bf7e162f62b2c64770c61e42441f48914cadd8bd3df11c2c4144af23c5

Observation e04d1035-3dcc-4afd-a3fb-70d09f167aff · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization PaLM-E: An Embodied Multimodal Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:12.847476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:12.847476Z digest=sha256:bd14305913e7889a4ccd3033f441ea1f3b9f05b4856dbd47dd5f9f1daebd14f8

Observation 71e35357-4ddc-4a40-9075-d55046e143b0 · outbound

This paper cites Goal-conditioned imitation learning,.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Goal-conditioned imitation learning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:19.782268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T06:04:12.506939Z digest=sha256:c9c16a29cfa83ec9c4030a8922b6abaa0314b565f571b7fce138fa5225e07653

Observation f8e0dc2a-dc31-4c6c-baa0-fe14db1d8574 · outbound

This paper cites Bridge data: Boosting generalization of robotic skills with cross- domain datasets, 2021.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Bridge data: Boosting generalization of robotic skills with cross- domain datasets, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:19.483094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T06:04:13.114027Z digest=sha256:c9f5e9c6ee96dd773ce940deb4e0925761e40a501bff6582a198d1a23a7f7d12

Observation 7bc2c181-e75b-4b2c-be2c-be5eedf95333 · outbound

This paper cites Zero-shot Task Adaptation using Natural Language.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Zero-shot Task Adaptation using Natural Language

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:13.256781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:13.256781Z digest=sha256:be4a12bf71d51b7971cf38693188923e5d3fcc460e2dfb5e5d7ec7dc9b4e6b1d

Observation 98af9672-62c5-4c32-8e14-9b839cba2c5a · outbound

This paper cites Tenenbaum, Dale Schuurmans, and Pieter Abbeel.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Tenenbaum, Dale Schuurmans, and Pieter Abbeel

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:13.008703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:13.008703Z digest=sha256:9884c6ff28d8e5673148ce155e58ead40bb93b4d88437f8613e75af25cfdbc35

Observation bbab7591-1d4b-449b-800b-4d99923cb8a1 · outbound

This paper cites Scaling up and distilling down: Language-guided robot skill acquisition,.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Scaling up and distilling down: Language-guided robot skill acquisition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:19.083800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T06:04:13.641938Z digest=sha256:4dbadfdecc0e25e41838d802efa25247f514d3e06be5b2e966e53b34833d7e1f

Observation 77d7b19a-06fe-44d6-b545-81a41046f6ba · outbound

This paper cites Hierarchical Few-Shot Imitation with Skill Transition Models.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Hierarchical Few-Shot Imitation with Skill Transition Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:04:17.299566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T06:04:13.916780Z digest=sha256:5e1d3f2e070e84b050fc464b1ba750eac02a69e8447862e10e4d2b1071fc25d3

Observation af306728-a8d2-493e-bef3-befb557e26f9 · outbound

This paper cites Rt-trajectory: Robotic task generalization via hindsight trajectory sketches,.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Rt-trajectory: Robotic task generalization via hindsight trajectory sketches,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:13.396260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:13.396260Z digest=sha256:5feff542252eb2f094869a8a404261596d18f32d3f95e6ef1982b9715eda65d8

Observation 3ab4ef90-0d42-499e-890e-2265f8db6e41 · outbound

This paper cites Inner Monologue: Embodied Reasoning through Planning with Language Models.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Inner Monologue: Embodied Reasoning through Planning with Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:14.300460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:14.300460Z digest=sha256:bc7ecc1c400e7f5460ea2c0bb22dbec79962328cac8c2abd8dd9fe7b86120c42

Observation d1b67de7-7075-4b12-a775-bf24e5ce6635 · outbound

This paper cites Davi- son.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Davi- son

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:18.676762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T06:04:14.423649Z digest=sha256:45c76fdce7fd059ed4a774b0d6e0323c5eb3b34c178a5c18bb68e352d130a59d

Observation 301e6b42-0a85-4c6a-b92b-082e95a0d7a5 · outbound

This paper cites Scaling Up and Distilling Down: Language-Guided Robot Skill Acquisition.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Scaling Up and Distilling Down: Language-Guided Robot Skill Acquisition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:13.792473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:13.792473Z digest=sha256:98598a2d386f8ea450fffb90708abaaee04ba35c3ce887aaccfafb2692e10496

Observation 4499946b-792c-45f3-a8ee-e712764fee94 · outbound

This paper cites Egomimic: Scaling imitation learning via egocentric video, 2024.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Egomimic: Scaling imitation learning via egocentric video, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:14.632736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:14.632736Z digest=sha256:7cb3434d03851d309dfa4168030a278e82634a119aec492a696950f0912f9b3e

Observation eb8e90ef-5c1e-42c4-843a-b877d33a5a80 · outbound

This paper cites Language models as zero-shot planners: Extracting actionable knowledge for embodied agents,.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Language models as zero-shot planners: Extracting actionable knowledge for embodied agents,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:14.064861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:14.064861Z digest=sha256:4e18260c35187cb3922841acefee550eda18283f500f827396c2121e69de770f

Observation 64a33463-34f3-4274-a335-c93734941f1e · outbound

This paper cites Openvla: An open-source vision-language-action model,.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Openvla: An open-source vision-language-action model,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:18.399852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T06:04:14.874563Z digest=sha256:e6c12335d11ae0fc2415fb1909aed62b69a32210102482eef3b9e8c756a8276f

Observation ddc84c13-9c55-4c89-8003-930e5b9430ff · outbound

This paper cites Code as Policies: Language Model Programs for Embodied Control.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Code as Policies: Language Model Programs for Embodied Control

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:15.152345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:15.152345Z digest=sha256:846d866f939a4cefef39acee3eea11ff3277cd06f85492d90fe7f96416fcf185

Observation 094e848f-e007-4bf2-a841-b7ed10906f90 · outbound

This paper cites MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization MOKA: Open-World Robotic Manipulation through Mark-Based Visual Prompting

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:15.276212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:15.276212Z digest=sha256:47115b65174265227b5c8169ef9ac44f56d6e6d1139c96e479e671a70d183727

Observation d40ce355-46d7-4239-bb10-ed404cb7b352 · outbound

This paper cites Bc-z: Zero-shot task generalization with robotic imitation learning, 2022.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Bc-z: Zero-shot task generalization with robotic imitation learning, 2022

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:14.540458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:14.540458Z digest=sha256:5975b9d0bca0047587c0050eb6d0c82ad73a1c3b5e5eb34a0f2c08289d721d2e

Observation a2f8b3cc-c20c-44cb-ac76-a04d0f763938 · outbound

This paper cites Learning latent plans from play.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Learning latent plans from play

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:18.177456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T06:04:15.530183Z digest=sha256:d16e60887ad7aed31222b1a0e224dab53113c5af201a9c68dd6b0c2657b8a422

Observation 8f362bea-e46d-4cea-a396-ee2e0a5f58dc · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:14.738359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:14.738359Z digest=sha256:617293ab96b05f005f88cafb9313da1119c01f3aac9edfbcf3d46baec3d98b77

Observation 4aa998ca-47b7-4171-9228-3d09c27d40f0 · outbound

This paper cites Visual Reinforcement Learning with Imagined Goals.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Visual Reinforcement Learning with Imagined Goals

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:15.752484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:15.752484Z digest=sha256:107ffa5577c2a9e4c2597e513e82656316e0a71c1837e96fe304afa281dc0080

Observation c258ec24-6ed2-4454-aacd-133ee025ff65 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization OpenVLA: An Open-Source Vision-Language-Action Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:14.996024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:14.996024Z digest=sha256:caa734d496fea196adb49b2559dba47fde93908ce977b362b31f6da7553255ad

Observation 9e66d2b2-3332-410e-b1d7-9f79c874c03c · outbound

This paper cites LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:16.013005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:16.013005Z digest=sha256:b8fec5fc60c7d981e6adfaaefc0bfdff810f3856ee831b3470d560128aabc2be

Observation 4e0423d0-c961-43ae-944f-95f520db5b34 · outbound

This paper cites an unresolved cited work.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T06:04:18.027711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T06:04:16.122763Z digest=sha256:c77176ba772f2735f767934730bf6352b327a039ec3be6fe56d00180bfec863b

Observation 587726b3-0826-4a81-afb2-4cc58c84d4e1 · outbound

This paper cites RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization RDT-1B: a Diffusion Foundation Model for Bimanual Manipulation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:15.419455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:15.419455Z digest=sha256:ca03f174f237bbc16da478c94496e91242bc2ba137ca941502f4fd2eba165c9d

Observation c30664b7-ea4d-4650-9262-cd75fe29750e · outbound

This paper cites Skill Induction and Planning with Latent Language.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Skill Induction and Planning with Latent Language

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:16.313044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:16.313044Z digest=sha256:b06249bcdab0a908e46ab169b5bf1151d2c7ed2e881f5d6b8cd4bf0d5b26b12a

Observation 335c160c-5635-429c-8999-124272021dc1 · outbound

This paper cites Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Dex-Net 2.0: Deep Learning to Plan Robust Grasps with Synthetic Point Clouds and Analytic Grasp Metrics

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:15.631738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:15.631738Z digest=sha256:52fbb397425effc38851bd51135c6982c01649f2ef6bc27077dcb626bff5177b

Observation c3d5ce3a-d78e-4923-82b3-774bae28362f · outbound

This paper cites KITE: Keypoint-Conditioned Policies for Semantic Manipulation.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization KITE: Keypoint-Conditioned Policies for Semantic Manipulation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:16.485673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:16.485673Z digest=sha256:1e8a373a94180e963480bb4e764822307d958d9647c3601ba0ba0d31554c33de

Observation 0257a728-d1a0-4ecd-aa02-63c2951c0e17 · outbound

This paper cites RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:15.921883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:15.921883Z digest=sha256:369e73e9b0e85353f2c3688e4684f778d8b6f2d877be90b9d0e8406301f02f97

Observation 769b965c-2dee-46ec-8b43-21acd8735ce3 · outbound

This paper cites Robotic Control via Embodied Chain-of-Thought Reasoning.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Robotic Control via Embodied Chain-of-Thought Reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:16.624067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:16.624067Z digest=sha256:1a25d447858cbada52c9b75373c0c78bb7ef9d8af1627eaeb89faa272248cd9a

Observation dd3f4caf-acb7-4b27-b876-81fd35ced15e · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:16.711246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:16.711246Z digest=sha256:7b897656a8028b2132ee6e2ba0f22c715bcd9e8b3cb9d42407b42c7a449b2643

Observation 5bd28e61-5b4f-4f08-a081-3bfc54cceac9 · outbound

This paper cites Discovering motor programs by re- composing demonstrations.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Discovering motor programs by re- composing demonstrations

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:17.857100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T06:04:16.248457Z digest=sha256:d656e6073d68ccba2caedb53034fbe7c88586b0aea24b489b22c06172f9727b6

Observation d2eee499-57e4-4e5f-8b69-f99ad7ae9380 · outbound

This paper cites Bridging Language and Action: A Survey of Language-Conditioned Robot Manipulation.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Bridging Language and Action: A Survey of Language-Conditioned Robot Manipulation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:16.844847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:16.844847Z digest=sha256:5fdd47541a6b767d83a5f2930e4489e21bc67ac0317f0b0938c5b891378f0d4c

Observation 04ea60fa-6bd8-421d-b5fd-28f30bf65f3a · outbound

This paper cites RoboCLIP: One Demonstration is Enough to Learn Robot Policies.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization RoboCLIP: One Demonstration is Enough to Learn Robot Policies

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:16.384793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:16.384793Z digest=sha256:c676f883e2113a78ec38c54fff9abbcc501dc51effaa63eb862e8a63d9e2535e

Observation 5f61b262-e051-453d-ab56-4b5be6ccd34f · outbound

This paper cites Robust imitation of diverse behaviors, 2017.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Robust imitation of diverse behaviors, 2017

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T06:04:17.688878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T06:04:16.546319Z digest=sha256:64bb2b6d90e6e42a4ac3a0ea73f3dac45e6b8786593a3c685a7202d56cb02a7e

Observation 7af48a98-fc57-4197-9039-0db98dc9e37a · outbound

This paper cites Zhao, Jonathan Tompson, Danny Driess, Pete Florence, Kamyar Ghasemipour, Chelsea Finn, and Ayzaan Wahid.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Zhao, Jonathan Tompson, Danny Driess, Pete Florence, Kamyar Ghasemipour, Chelsea Finn, and Ayzaan Wahid

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:16.776984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:16.776984Z digest=sha256:83d0103a3a9c981755d84b7357acf7cc0830b22c5c62ecaf8f8840ca18909dac

Observation 3337a975-e389-4e13-8b0d-db66e74bdc7e · outbound

This paper cites Goal-conditioned Imitation Learning.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Goal-conditioned Imitation Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:12.655115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:12.655115Z digest=sha256:50045369c873445c29d98a63e4e12e1d4c9b153f2abd5c35bd0d7ff890fd2657

Observation bc320820-7c37-4171-98ce-06cae42dc222 · outbound

This paper cites Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied Agents

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:14.191632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:14.191632Z digest=sha256:dacda137c9170fcd2af018280f8455157a9aef700fd34a588e461b30d85e6bbf

Observation 7ccb436b-01f9-486c-a6c5-f9be3b136d8c · outbound

This paper cites RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:13.524282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:13.524282Z digest=sha256:048e1e71b4209806bd7859099a8227b7ec6cef65e55a648979041f35360d5b2b

Observation a82e9a24-ed16-4b9a-a427-25cd479fe798 · outbound

This paper cites SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities.

Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T06:04:11.687427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:04:11.687427Z digest=sha256:98f4a28997613a6850719ebdb5245a34b670d2ad83a9da0143f81009bbcad8c5

Pith citing papers

Observation 2b3285ce-41ec-4537-81f0-c8891e2ac14d · inbound

Continually Evolving Skill Knowledge in Vision Language Action Model cites this paper.

Continually Evolving Skill Knowledge in Vision Language Action Model Bridging Perception and Action: Spatially-Grounded Mid-Level Representations for Robot Generalization

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:04:09.185337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T06:02:31.638120Z digest=sha256:39e5f02472931ae89440e7b7131ef6a594c5962c8dbe8635c3e1d3e6bd6a4eaf