Pith. sign in

Paper Citation Record · LEDGER

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors

As of 22 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 1 inbound Pith citation observation for arXiv:2504.21530.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.21530 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:07:05.277615Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:57:26.612635Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T00:57:31.251745Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f074e8e6-36cd-4847-9400-191526bc441e · outbound

This paper cites GPT-4 Technical Report.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.014417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.014417Z digest=sha256:60556e7a18627b2ddec7910c9288d717bc5286f42c4270cdde3c2d992769b70b

Observation 0d461be6-308f-4049-b4c4-d4bf31ce7dd9 · outbound

This paper cites RT-H: Action Hierarchies Using Language.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors RT-H: Action Hierarchies Using Language

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.020787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.020787Z digest=sha256:1d30fcf5c68d2b362c2e41efeb10eed7e819eb3f5534a966d25e8996104705b5

Observation 0ef83400-73fb-493f-aebc-6736262fa895 · outbound

This paper cites Track2act: Predicting point tracks from internet videos enables diverse zero-shot robot manip- ulation, 2024.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Track2act: Predicting point tracks from internet videos enables diverse zero-shot robot manip- ulation, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:06.109718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.026019Z digest=sha256:5241767ec32c0e99c034be73b940a8fc53b86c28d1eb28429a7574e5887ad01e

Observation 4836559d-3557-4069-804f-920f4084e771 · outbound

This paper cites Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Zero-Shot Robotic Manipulation with Pretrained Image-Editing Diffusion Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.030891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.030891Z digest=sha256:b86dacf3671ad9bf6f778b343f1906e2b32e9a52fea20733a9c3a323a9d2c1cb

Observation 8ce6648c-e311-409c-9a1d-44ef3561d056 · outbound

This paper cites RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.036033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.036033Z digest=sha256:8813c0395e44385da41879eba9da5804cada7276ae7f4c819cd0dd706f6b8aba

Observation 2824a6d8-0541-4ffe-89e4-b040ff4cfe62 · outbound

This paper cites Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Closed-Loop Visuomotor Control with Generative Expectation for Robotic Manipulation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.041332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.041332Z digest=sha256:f7c9fa7e34b441092d100f3c1d740eb76f2371f4894bb1ec9a6181f2990d8f30

Observation 46e141c7-077d-4b04-9206-4332712dbd01 · outbound

This paper cites Shikra: Unleashing multimodal llm’s referential dialogue magic, 2023.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Shikra: Unleashing multimodal llm’s referential dialogue magic, 2023

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:06.093942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.046596Z digest=sha256:c2fcd87a0e4c24f976c681e3938c41677ffc7caee691f8c9eaabe5963e454b3f

Observation 69db985e-c967-4495-bffa-78c239e509c5 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action dif- fusion.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Diffusion policy: Visuomotor policy learning via action dif- fusion

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:06.078264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.051125Z digest=sha256:1496617c4aa5c194ae397a1cc8736e0de5320a74b3a81ca9670423fac70fe9c3

Observation 61ba1ce3-9d6c-4b66-b236-ad5704926dfb · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.056513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.056513Z digest=sha256:ad42529a85a771016277f1670d63ff09c1763a6f7e569f22f233596026a3660e

Observation cb84a066-c793-414a-9616-46a80439ccf0 · outbound

This paper cites Imitating Task and Motion Planning with Visuomotor Transformers.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Imitating Task and Motion Planning with Visuomotor Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.061342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.061342Z digest=sha256:ecd8b24b04e5509c7e78d8f70e5433a52a989de70ba8197e83a3b5a0fe8aa3c0

Observation 5fb40c13-a3c4-425a-9999-85db878aed17 · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Objaverse: A universe of annotated 3d objects

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.066352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.066352Z digest=sha256:d3ea4b82456deffc30a275b82e1948e61ae3c4c54e63afe4860d4aa6859bbf93

Observation ee4c90f3-db52-4217-9d5d-df1c49c4340b · outbound

This paper cites FlowBot3D: Learning 3D Articulation Flow to Manipulate Articulated Objects.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors FlowBot3D: Learning 3D Articulation Flow to Manipulate Articulated Objects

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.071158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.071158Z digest=sha256:f74c10686237a97269d8fa834a1fb34ce6d48d901245065aff35382a8d4030cf

Observation db5e5724-441e-4e18-a801-d1e2dfe57d50 · outbound

This paper cites Anygrasp: Robust and efficient grasp perception in spa- tial and temporal domains.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Anygrasp: Robust and efficient grasp perception in spa- tial and temporal domains

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.076200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.076200Z digest=sha256:30851870e6610843e3f9aa76dbfed34773511073ff6ef0cd1040f08b78187e7b

Observation 67075254-6d1b-4dfb-b3a4-b71e2eac5973 · outbound

This paper cites Moka: Open-world robotic manipulation through mark- based visual prompting.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Moka: Open-world robotic manipulation through mark- based visual prompting

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:06.033158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.081370Z digest=sha256:d6e7481d37fc9fb0ec94964895972f425d2ddebded345d4b642890b94a29cc88

Observation dce6ae1b-891a-4f92-acf5-519d726b0f15 · outbound

This paper cites RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors RT-Trajectory: Robotic Task Generalization via Hindsight Trajectory Sketches

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.085786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.085786Z digest=sha256:51d057245844053f8ccd5ada99a7fbd45c78abfd76835bfdeb110a3e35532666

Observation 1c440137-15a8-49d7-8e4e-808e75f09171 · outbound

This paper cites Masked autoencoders are scalable vision learners.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Masked autoencoders are scalable vision learners

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.090760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.090760Z digest=sha256:8233068f0aae27c93662b716c9076bb3c339a5b30ea6d1a84b67b42381d2fe9e

Observation 94e26d11-60ca-461b-9501-8ae73a4335c9 · outbound

This paper cites Perceiver: General perception with iterative attention.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Perceiver: General perception with iterative attention

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.095555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.095555Z digest=sha256:9c57ecdbf1a5a89a6e6d359f2694b584648d427992a20dbcab059e28659a5db1

Observation 2064a1ad-e8e5-4fd1-90f3-81836ef34b23 · outbound

This paper cites Rlbench: The robot learning benchmark & learning environment.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Rlbench: The robot learning benchmark & learning environment

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.998329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.100013Z digest=sha256:09709b8e4a0423f0e7f310c61398987ce203815a2564f1f0fd0618f351b3004c

Observation 6dcf33ae-c757-4c84-9011-f240efd95410 · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors OpenVLA: An Open-Source Vision-Language-Action Model

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.104695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.104695Z digest=sha256:d2b8d1c6254027c9401f7ee87f4edb68c1a6aa287759c773fe9ebdb46d446fd6

Observation 2a032ea2-8a3c-4f81-9829-da414ccef191 · outbound

This paper cites Segment Anything.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Segment Anything

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.109396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.109396Z digest=sha256:95b400f4f906d4d2330e68b4c69a4f8443689e1fbcf42cd99c12a593abfd0557

Observation cb65e6f4-5702-4e09-908b-9d2c06834f18 · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Lisa: Reasoning segmentation via large language model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.983337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.114382Z digest=sha256:9135f7571f03e8ea5cdcba50d0dcf0d01ddf5a6ef08111903dd1dc4c723301c4

Observation 8b745074-d37b-41b5-bb85-14f9fa64895e · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.118904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.118904Z digest=sha256:ea780f4066da5a9ce8130b2379d23af5646823e0e7a33563aa55219a9a3162a5

Observation e331e974-d1ec-41bc-b1a0-a22077e7d56e · outbound

This paper cites Visual instruction tuning, 2023.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Visual instruction tuning, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.123999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.123999Z digest=sha256:10dbb6aac01a683f2051fe1ba42d8fc5f00344f785735284c7643b09df2d0e05

Observation e4e6f5ae-d174-4aec-aec3-fd7c21ea9676 · outbound

This paper cites Groma: Localized visual tokenization for grounding multimodal large language models.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Groma: Localized visual tokenization for grounding multimodal large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.948299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.128581Z digest=sha256:9e2e338f1dc93605bd544a54eb6985cd51cf11ed2c02ba579dde23e9967ff10f

Observation 03b09d70-72d5-49f7-8dcf-899593fbd439 · outbound

This paper cites What Matters in Learning from Offline Human Demonstrations for Robot Manipulation.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors What Matters in Learning from Offline Human Demonstrations for Robot Manipulation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.133010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.133010Z digest=sha256:12f643ae329e7cc4a74f0a2a32e4b7aa9246644e458ede0390c311e9e20f2176

Observation d7ad62f4-ff13-4623-9aaf-bdb8c8051f59 · outbound

This paper cites MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.137762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.137762Z digest=sha256:8418cbd44169d4b4e286602a12aa3104d95d2ca8efdd01d27d8c8fcaad677dc6

Observation f5354ceb-51cd-440b-9a55-78753364ebba · outbound

This paper cites Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manip- ulation tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.143620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.143620Z digest=sha256:0dc3bd26d35391756786be0d3fbac65272f98d4a8b06b0e0f733d7e20e69df6a

Observation 837e1179-462f-4f19-a917-b055c634894f · outbound

This paper cites RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.148191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.148191Z digest=sha256:aa363512144e12ace3caf433a839d3681fd104b3f3e142126005255db7b47f0e

Observation c1715922-d1dd-4aab-9682-659c4505358a · outbound

This paper cites Kosmos-2: Ground- ing multimodal large language models to the world.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Kosmos-2: Ground- ing multimodal large language models to the world

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.922762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.153083Z digest=sha256:85c954be4e4f8eb3ffb281994708afd054c88773871f325109712a8e1b9457c8

Observation ddd45b25-6d60-4b09-a78c-13707d575483 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Film: Visual reasoning with a general conditioning layer

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.907274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.157617Z digest=sha256:b0e84c2dfabf1b9ac9a93a74afc4ee971a06bf9283b1d579bea9b9b9b764247e

Observation b2365aef-51a9-4f86-b86b-8a121f1917be · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Learning transferable visual models from natural language supervi- sion

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.162217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.162217Z digest=sha256:799c2db4dd5ca0dc591d7e4cdcc688fe58f089a13ad7938a2e56aa7d1df2dda4

Observation 0eefa914-6846-4363-95fe-df26598a7180 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Glamm: Pixel grounding large multimodal model

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.881722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.166814Z digest=sha256:e324ae884253e2b08285baa2dfe4ab3a040b930184dda17fea1b88cec22a1d54

Observation 4775d94b-b1f0-4892-a54a-ec08ba8c47ff · outbound

This paper cites Toolflownet: Robotic manipulation with tools via predicting tool flow from point clouds.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Toolflownet: Robotic manipulation with tools via predicting tool flow from point clouds

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.171351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.171351Z digest=sha256:4ca2b6df0e92f3fe9901aea578122563fee427c280b604112320aa253644fc27

Observation 6ea02a72-57ef-4be3-876c-634121302f77 · outbound

This paper cites Cliport: What and where pathways for robotic manipulation.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Cliport: What and where pathways for robotic manipulation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.856222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.175851Z digest=sha256:df5e1fbb4a7ec6f120beaaa196083b76c86d80cf758bd27eb509b6a4db4fae78

Observation c4303925-de3e-4ff7-b20f-e2d8fda771b3 · outbound

This paper cites Open-World Object Manipulation using Pre-trained Vision-Language Models.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Open-World Object Manipulation using Pre-trained Vision-Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.180657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.180657Z digest=sha256:d08b9b1f9f1514257326b2bd1caa77ca58a99ad173e55fea744d0b07ada36f3c

Observation 06adf63a-fb02-466e-917e-62a3edf1ff75 · outbound

This paper cites KITE: Keypoint-Conditioned Policies for Semantic Manipulation.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors KITE: Keypoint-Conditioned Policies for Semantic Manipulation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.185666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.185666Z digest=sha256:7a8ac6e0a1f74863c48113225c696ef90dc25fdd69af8cced2fb6c2fcf2e5f06

Observation 99ce5b5c-6f34-4606-8ad6-b484b9192006 · outbound

This paper cites Rt-sketch: Goal-conditioned imitation learning from hand-drawn sketches.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Rt-sketch: Goal-conditioned imitation learning from hand-drawn sketches

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.840640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.190494Z digest=sha256:6e286e9dd62505cb3cd1f4e015a9f103bc0e3d9dbbc39ff985a5cd8d112da943

Observation 3c951082-1103-41b6-a1fd-93714512b2d4 · outbound

This paper cites Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.195110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.195110Z digest=sha256:2dd7f8119a2d3fd46de7d7d0366047c41e78dc98024c568c3c5bb550d5fdf9ef

Observation aa50c0e3-b952-419d-8fb3-7bb6e17d6376 · outbound

This paper cites Robotap: Tracking arbitrary points for few-shot visual imitation.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Robotap: Tracking arbitrary points for few-shot visual imitation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.825296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.200214Z digest=sha256:8a709c8d551ada1da2a9b63ac85ab0b0aff97d4775128b6383e9fa7cde0c39bf

Observation 1f055d2a-0859-4feb-8b4d-d4d2e90921cb · outbound

This paper cites GenSim: Generating Robotic Simulation Tasks via Large Language Models.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors GenSim: Generating Robotic Simulation Tasks via Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.205184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.205184Z digest=sha256:ca539157b37137a3b4ed2986188fc2f518f32b14475d4f9c8d4141bace61a84f

Observation c4c29b91-de0c-400b-b238-62856187c9ae · outbound

This paper cites The All-Seeing Project V2: Towards General Relation Comprehension of the Open World.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors The All-Seeing Project V2: Towards General Relation Comprehension of the Open World

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.209945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.209945Z digest=sha256:5362327a03ad24c24de9f39aef9c3588f1c4aec4f5ffdaa7cb2a1af50e7234c7

Observation 9c87fe91-406e-482f-9f33-dd3d4e553154 · outbound

This paper cites Any-point Trajectory Modeling for Policy Learning.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Any-point Trajectory Modeling for Policy Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.214926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.214926Z digest=sha256:473960eb4fff63c3ef657d5ededd997cc14acd64b9e02617715dd49afc111040

Observation 38f1035b-a4ea-4dd2-91ed-9445a9c65e77 · outbound

This paper cites Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.219984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.219984Z digest=sha256:7c5a9a6517607f42108d591cc8c3761ded8ffc8a55d15b31e8f9890f079f77d1

Observation 7f07915e-3277-4d46-a68e-a438cba7f124 · outbound

This paper cites Florence-2: Advancing a unified representation for a variety of vision tasks.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Florence-2: Advancing a unified representation for a variety of vision tasks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.809925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.224660Z digest=sha256:87feeaf61d25d915e829b237d4c86fefee1c0e5b1b11b851f2bc91d4929a7092

Observation 029ea9f6-7685-4cf3-9bce-2e865967315b · outbound

This paper cites Flow as the cross-domain manipulation interface.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Flow as the cross-domain manipulation interface

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.794214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.229333Z digest=sha256:97925f35e9e8bca6c0c25c6ee5416329c6bda5ca394b6249bd9901ba909ec1b3

Observation ecbbc05a-5278-44fc-8150-7e04daaf905a · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.233862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.233862Z digest=sha256:226a6a7e8d6f2ce886e1943a32ab4a89e67dc9deea0e5244d6fde9564ac363f6

Observation e7e61a4b-e04a-4e2f-8d5c-e7aee350d623 · outbound

This paper cites General Flow as Foundation Affordance for Scalable Robot Learning.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors General Flow as Foundation Affordance for Scalable Robot Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.238702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.238702Z digest=sha256:5cc4796705a63ad1078998cff442506fad0cd8053af0027485547f490b9ded88

Observation 169caaa6-595d-4751-b87d-a9bd1b76cd8c · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.244349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.244349Z digest=sha256:0717aca7bce6ac5b22f5db5666c490041ccc597e6eb360f2cf75559514cbc115

Observation 7c18a2de-149d-4851-8d80-6174614df944 · outbound

This paper cites Sprint: Scalable policy pre-training via language instruc- 10 tion relabeling.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Sprint: Scalable policy pre-training via language instruc- 10 tion relabeling

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.776430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.249417Z digest=sha256:83767727c32f856636b0b6c94871c08b3126eb9f4f68876d1397719d2990b39f

Observation 6872fe4e-93ad-48de-b552-228b094bfe3d · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.254066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.254066Z digest=sha256:8d8b8d3c35d8a7217d067f16e7e27ca9b6ff7814616759a23c851bdff7150145

Observation 9ae6efb6-9784-4b30-86a4-c8cf4450b3b9 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T05:07:05.258931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:07:05.258931Z digest=sha256:23577a27ba5f9cd75a44a432df0b479b48e895017183fea285f7f6b88b4a10a9

Observation 4c314359-1979-466b-9693-1c318cd8b848 · outbound

This paper cites an unresolved cited work.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:07:05.760836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.263792Z digest=sha256:505ba21f299fcedd8bebc2fe108306d6fbac5156da731fb9281bc10f067219f7

Observation c5833f51-b78b-47a2-9536-ba3c34ab632e · outbound

This paper cites an unresolved cited work.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:07:05.745175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.268368Z digest=sha256:41dff3ee4c043c9d57075d4460593fc09d4c00dc5e39d38a015ac2ee6d31bdd7

Observation 71b8116e-2fb6-44ad-bc13-402a9bbaa1d2 · outbound

This paper cites an unresolved cited work.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-16T05:07:05.729786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.273044Z digest=sha256:8a99662be799f7249f262a63647561762b563de7c7e5973a16665e2e7f418e38

Observation a133c291-10b0-4dfa-b7c8-e906210a3e78 · outbound

This paper cites yellow surface with brown spots.

RoboGround: Robotic Manipulation with Grounded Vision-Language Priors yellow surface with brown spots

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T05:07:05.714651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-16T05:07:05.277615Z digest=sha256:194c2b2a292e7ac61e96bd40266aada455d7d51761261f3f31306b1658a021a9

Pith citing papers

Observation 7c29c10a-ade1-4c72-91e6-45a7d04c8a88 · inbound

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making cites this paper.

AntiGrounding: Lifting Robotic Actions into VLM Representation Space for Decision Making RoboGround: Robotic Manipulation with Grounded Vision-Language Priors

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:57:31.256067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T00:57:26.612635Z digest=sha256:d4f58e3d700ed5d7c1ac5696661458da0578967b380e0163a75b040ebac1bed4