Pith. sign in

Paper Citation Record · LEDGER

Unify Robot Actions in Camera Frame

As of 4 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 1 inbound Pith citation observation for arXiv:2511.17001.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.17001 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T21:16:51.909363Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T18:31:03.548002Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T22:57:26.662723Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact18
  • verified fuzzy26
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f81cab21-7dac-4095-a40d-f2f7f79f23f3 · outbound

This paper cites $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control.

Unify Robot Actions in Camera Frame $\pi_0$: A Vision-Language-Action Flow Model for General Robot Control

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:20:17.151567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:91c4a444cbacf879b2a3b03fda17459bcd0b4477f2c2fc9f9a3e0ebc7c9b7c47

Observation 38ef90de-ef96-441f-a665-65d3fe7dd03f · outbound

This paper cites Easyhec: Accurate and automatic hand-eye calibration via differentiable rendering and space exploration.RA-L.

Unify Robot Actions in Camera Frame Easyhec: Accurate and automatic hand-eye calibration via differentiable rendering and space exploration.RA-L

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.208930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:6039cc142255c3faea53e83720bc64fdde6d11bb272b6079421b1a35a48b804a

Observation 780e7b2f-fc98-4ce2-b994-419c07228316 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action dif- fusion.IJRR.

Unify Robot Actions in Camera Frame Diffusion policy: Visuomotor policy learning via action dif- fusion.IJRR

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.127132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:17c1959028ee332d91dc44f2f2ad58374ed4b1810a1ebd932e315dbfc0f957f2

Observation 88619db3-1388-43cc-ab3f-ec35c8d52919 · outbound

This paper cites Hand-eye calibration using dual quaternions.IJRR.

Unify Robot Actions in Camera Frame Hand-eye calibration using dual quaternions.IJRR

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.211698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:b1933ecb4727d7bc9fd56e8d503b58ab63b1bb31f433fe8843edc04e282845a2

Observation 8720b644-0c86-4d29-821c-3682f636cf8d · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Unify Robot Actions in Camera Frame An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:20:17.123725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:dcc2d3f508f53ed435ebd96c8ee3a80a0c658f5ff8bf4d19ac0154fa7c50b08d

Observation 6b282e29-46cd-4f02-b159-3e2005516051 · outbound

This paper cites AirExo-2: Scaling up Generalizable Robotic Imitation Learning with Low-Cost Exoskeletons.

Unify Robot Actions in Camera Frame AirExo-2: Scaling up Generalizable Robotic Imitation Learning with Low-Cost Exoskeletons

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:20:17.146749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:9502164071c4f41ddf2a76c79c2ea240a5d1f56d9c7a9ba266965263e3de6ab6

Observation d7b43fd3-fb90-4764-b56a-4be4f8848d0a · outbound

This paper cites Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation.

Unify Robot Actions in Camera Frame Mobile ALOHA: Learning Bimanual Mobile Manipulation with Low-Cost Whole-Body Teleoperation

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:20:17.129381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:82bc332400244a9f9a16e9f76fea48c947af527c5114ba3f8190b6366b7fe719

Observation 7647d8b8-a983-48de-939a-5cd6d28bc174 · outbound

This paper cites Easy- hec++: Fully automatic hand-eye calibration with pretrained image models.

Unify Robot Actions in Camera Frame Easy- hec++: Fully automatic hand-eye calibration with pretrained image models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.185674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:79f03ad1b2571959f1aed2bd5b5aee0e345295344005728118c7f16605866cca

Observation ef677e73-2099-4909-a657-2a023b16610c · outbound

This paper cites Robust robot-camera calibra- tion.

Unify Robot Actions in Camera Frame Robust robot-camera calibra- tion

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.136996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:30873c207925d482568bbf95190c33ef2672a589f4b3df06c267cb6f34fa9efc

Observation b1fd7e4e-a5f9-4452-8d6e-e485c52fed88 · outbound

This paper cites $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization.

Unify Robot Actions in Camera Frame $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:20:17.191917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:61a0d23eb6634bdff8d9ebc1c0c615d622571b064590461551a01fedfdbdf474

Observation 31c82a33-22a7-4095-86f9-3d7fd1823a77 · outbound

This paper cites Rlbench: The robot learning benchmark & learning environment.RA-L.

Unify Robot Actions in Camera Frame Rlbench: The robot learning benchmark & learning environment.RA-L

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.178754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:cb93305916bb6cf373ad51f8579aad7d67e10909e8566a7639003f10af30aa00

Observation 909777c9-1159-426c-8a6b-6158a56d539c · outbound

This paper cites Do you know where your camera is? view-invariant pol- icy learning with camera conditioning.

Unify Robot Actions in Camera Frame Do you know where your camera is? view-invariant pol- icy learning with camera conditioning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:20:17.224020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:973440c1393e82bafa3b19cc9b186e9dcb572ce13f8fda31a60d4d2460f0df79

Observation 89153cf0-89d3-4c83-a7c1-03f8dd87abaf · outbound

This paper cites Co- tracker: It is better to track together.

Unify Robot Actions in Camera Frame Co- tracker: It is better to track together

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.175591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:09b3dc81165e2db4201849240570130f632b8cd3eb88f433753e8d08c04b5abe

Observation 85290c57-0c12-464d-b487-6b0d0b534827 · outbound

This paper cites Co- tracker3: Simpler and better point tracking by pseudo- labelling real videos.

Unify Robot Actions in Camera Frame Co- tracker3: Simpler and better point tracking by pseudo- labelling real videos

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.169060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:5078ec3533ddf2435463a7dfdb83787f6623b824f9a85c58d6eedd351acc3f13

Observation 1a0d4e30-6cb4-4b1c-91dc-8e0b214fa3c4 · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Unify Robot Actions in Camera Frame DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:20:17.171732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:49a1a86f91b76d2ea37794929273499480c34b8e9b232f08db2240ff2d872a9c

Observation fd491480-8f23-492c-96b4-866fd7c5e58f · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

Unify Robot Actions in Camera Frame OpenVLA: An Open-Source Vision-Language-Action Model

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:20:17.197592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:f45b0343acfa5d3fc71827fe4ba198936aa25ebf231074973750b6b9b05ee7e5

Observation 00484e7f-0c49-4a7f-9d22-147e7e5aa4a0 · outbound

This paper cites Single-view robot pose and joint angle estimation via render & compare.

Unify Robot Actions in Camera Frame Single-view robot pose and joint angle estimation via render & compare

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.165401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:50544b774cec60c0c1c6d864ce86c192d06f083a4479dd28340aef51d2ea8ce6

Observation d3ae63f6-68cd-46dd-aa60-cfefc9f6ee9e · outbound

This paper cites Modular primitives for high-performance differentiable rendering.ToG.

Unify Robot Actions in Camera Frame Modular primitives for high-performance differentiable rendering.ToG

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.188981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:c0bf052b6dad5aa313ce63ccfd781159edaa5746eca9014f7a1441d95e21f715

Observation 1ed9ed54-ab30-4dbb-a016-7b723836a8e1 · outbound

This paper cites Camera-to-robot pose estimation from a single image.

Unify Robot Actions in Camera Frame Camera-to-robot pose estimation from a single image

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.161508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:8077cabe145f9bd070116263a70e651b1207543ebd6d692fa71e49b0f5503efa

Observation 839e34d6-b2e6-4f67-80d8-6e6eab7e1d64 · outbound

This paper cites Phantom: Training Robots Without Robots Using Only Human Videos.

Unify Robot Actions in Camera Frame Phantom: Training Robots Without Robots Using Only Human Videos

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:52.652914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:07baf08ad931245fec7220e9c4aa8c61ecdd56fdf073a469d633af6418afa388

Observation 016ceb0a-e572-421a-a3aa-79ccee7d8e79 · outbound

This paper cites Ep n p: An accurate o (n) solution to the p n p problem.IJCV.

Unify Robot Actions in Camera Frame Ep n p: An accurate o (n) solution to the p n p problem.IJCV

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.182376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:12d147785cf07e9de80e94223aeb2d5ff91b1877c9cd3d49f7ff95181c36dda6

Observation f8d4e1b1-2085-4573-9f20-5119585b0bf9 · outbound

This paper cites Prompting depth anything for 4k resolution accurate metric depth estimation.

Unify Robot Actions in Camera Frame Prompting depth anything for 4k resolution accurate metric depth estimation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.201871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:6cb0504e1cf0dac804ed63ef62feb1749515ca499f6cf8a79666fa1daaecfe0b

Observation e14ff6a7-a002-4121-8172-cff574436896 · outbound

This paper cites Libero: Benchmarking knowl- edge transfer for lifelong robot learning.NeurIPS.

Unify Robot Actions in Camera Frame Libero: Benchmarking knowl- edge transfer for lifelong robot learning.NeurIPS

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.198892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:d289f22e6674e67dfaa68a2d0f94314f2db243845bc62779f970218cc6337272

Observation 37007c2a-30a1-4e07-97bb-7902ab881c7a · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Unify Robot Actions in Camera Frame Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.205553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:17965cb9aa5dfd9da132cc8c7b14ca421b84c9377557546409b582caf2d0bc2c

Observation a8716ab8-2613-4ae6-854e-de9ede3a3d7d · outbound

This paper cites Markerless camera-to-robot pose estimation via self-supervised sim-to- real transfer.

Unify Robot Actions in Camera Frame Markerless camera-to-robot pose estimation via self-supervised sim-to- real transfer

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.195706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:22c19b5c48a50ace2f92281c1ec9870905995d4c4a2e1a7cdebe7a2ec911a3f1

Observation ae5df544-dd5f-40f2-9117-76b753a3c153 · outbound

This paper cites Ctrnet-x: Camera-to- robot pose estimation in real-world conditions using a single camera.

Unify Robot Actions in Camera Frame Ctrnet-x: Camera-to- robot pose estimation in real-world conditions using a single camera

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.154481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:2744c3d2bb0d5cb82e006871de11c0dd5a863d22fcd7fb4af8cce62bd3f6317c

Observation 34f6a455-e7b5-4b54-bb26-253997eb95a6 · outbound

This paper cites Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning.

Unify Robot Actions in Camera Frame Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:20:17.161329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:2f5fbcb11a4d630c25f8b83ef696ee302c91b73142c232cd5c2a9bfe37615393

Observation ccab7834-03c7-40eb-8797-d220e13d6fa5 · outbound

This paper cites Segic: Unleashing the emergent correspondence for in-context segmentation.

Unify Robot Actions in Camera Frame Segic: Unleashing the emergent correspondence for in-context segmentation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.172371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:ebf97c446277bf15a5d4854264ee1a0f17396710a7139434ec45b8529078a00b

Observation e11f9d58-8445-4869-a31f-17b4dd4dc9e9 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Unify Robot Actions in Camera Frame DINOv2: Learning Robust Visual Features without Supervision

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:20:17.185441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:36f3daf0f81d902890a24c113e6d0c3a10785f29f9886f5c266776cbee77ebd3

Observation 058f4355-3ed5-4a0e-acb3-c45fefb00675 · outbound

This paper cites Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0.

Unify Robot Actions in Camera Frame Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration 0

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.144371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:aa068514cde9e680eca07455e4de4a7f5c3ddb4cf2a373b42e0193cf8c38f75f

Observation 5a1b8cef-fa45-4139-b815-1d86e0f54c18 · outbound

This paper cites Robot sensor calibration: solving ax= xb on the euclidean group.T-RO.

Unify Robot Actions in Camera Frame Robot sensor calibration: solving ax= xb on the euclidean group.T-RO

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.157984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:ea66efb72e0e7f0217901cbf83ead1ae08b01ea253934f712f7c4689852048e6

Observation 881dde1f-d917-4901-83f7-da6e5209dc4f · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Unify Robot Actions in Camera Frame Learn- ing transferable visual models from natural language super- vision

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.192438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:0dc23c6f67c6f38c4da31fba45d3e67af24d5c9a951615aa06d23ba390b8e0b9

Observation 95741c77-8fa2-4f39-ab28-ff3158b1269d · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Unify Robot Actions in Camera Frame SAM 2: Segment Anything in Images and Videos

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:20:17.209821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:7bc6e6828b8df0224326833f5cb590a669ee39a6fbf882ce0d38a07efbc3d76b

Observation 1c5b9076-7145-495c-a71b-30b77f83cc11 · outbound

This paper cites Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection.

Unify Robot Actions in Camera Frame Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:20:17.141250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:49a9f8094bacd20123ebe3b744e5e9529cc991bfb15c880d1ab1253db14ccab8

Observation e5080c59-02ac-42f9-be76-c9243527bf6b · outbound

This paper cites Emergent correspondence from image diffusion.NeurIPS.

Unify Robot Actions in Camera Frame Emergent correspondence from image diffusion.NeurIPS

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.130455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:38ddb104e367059d81139ac7058d3965e15d6fa25eec5dfaa0d2028fd0829150

Observation fca4bb70-35cc-4783-a231-de22c7753f3e · outbound

This paper cites A new technique for fully autonomous and efficient 3 d robotics hand/eye calibra- tion.T-RO.

Unify Robot Actions in Camera Frame A new technique for fully autonomous and efficient 3 d robotics hand/eye calibra- tion.T-RO

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.133627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:d110aa2f9375ff71e41ac5196556a0831314b40c07368e0cd0bfe28bec4f7bd1

Observation dc195d3c-e434-4879-ba25-b139876d5c2c · outbound

This paper cites MimicPlay: Long-Horizon Imitation Learning by Watching Human Play.

Unify Robot Actions in Camera Frame MimicPlay: Long-Horizon Imitation Learning by Watching Human Play

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:20:17.176691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:a723821e3da18bf93842578b0d5128e5c2b69563cb2912b891dfa6ada102269d

Observation 7c5c7f4a-d5f5-4190-a1c3-fd3505ed2c76 · outbound

This paper cites Meta- world: A benchmark and evaluation for multi-task and meta reinforcement learning.

Unify Robot Actions in Camera Frame Meta- world: A benchmark and evaluation for multi-task and meta reinforcement learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.140875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:886acb6a8958595582d108ae105607ca9a4a610ea108f7a5f09eede6f6e72441

Observation 61b2f714-d90f-426f-88b0-858e4e4c65a7 · outbound

This paper cites DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection.

Unify Robot Actions in Camera Frame DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:20:17.135058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:d8af70015a06ccb72696494ffbb9876476793995c97201b887169112557dbe52

Observation 89b7172b-7578-4d7b-8eab-2af44d50c043 · outbound

This paper cites Vlabench: A large-scale benchmark for language-conditioned robotics manipulation with long- horizon reasoning tasks.

Unify Robot Actions in Camera Frame Vlabench: A large-scale benchmark for language-conditioned robotics manipulation with long- horizon reasoning tasks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.147803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:8d7578e0633302b73fc5be92c43349c62bfa3ce34fed961c584c2b482bdff2b2

Observation cb4fe4bc-0492-4874-81b6-d0d012b26025 · outbound

This paper cites Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy.

Unify Robot Actions in Camera Frame Grounding Actions in Camera Space: Observation-Centric Vision-Language-Action Policy

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:20:17.204124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:c2e3690db4de874d508d3ac0f16ff2be8fb596b5150c8da4bb75c8647b51d2ea

Observation 91a16a7c-8394-4e46-bd3a-1e3a07129cff · outbound

This paper cites Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware.

Unify Robot Actions in Camera Frame Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:20:17.156363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:898bc57eec2bbe0c09e6acb4c0e0fa073e2a9718fed651f14212e1e76ed3a0fc

Observation 0ffd5098-eef0-4947-8c97-86959e49bf50 · outbound

This paper cites robosuite: A Modular Simulation Framework and Benchmark for Robot Learning.

Unify Robot Actions in Camera Frame robosuite: A Modular Simulation Framework and Benchmark for Robot Learning

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:20:17.214948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:076a0df648e3d87b973d167592dadc4ac873fe74989e1b9d9d4758a686b734a7

Observation 937a12df-0b77-448c-95bb-2f756752c669 · outbound

This paper cites Rt-2: Vision-language-action models transfer web knowledge to robotic control.

Unify Robot Actions in Camera Frame Rt-2: Vision-language-action models transfer web knowledge to robotic control

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T21:20:18.151326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T21:16:51.909363Z digest=sha256:5e53f0458b162084dee9ef564dcb9c78e6abd3a18faf7e4256c16303c68bfacb

Pith citing papers

Observation 9d7360df-869f-4e25-88a0-e7c80dd2894c · inbound

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data cites this paper.

Two Bridges, One Pathway: From VLMs to Generalizable VLAs with Embodied Trajectory-Coupled Data Unify Robot Actions in Camera Frame

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:57:26.663913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T18:31:03.548002Z digest=sha256:012bdd85b8087f29ff331fe153bdbcc778dbedf6b3de25c230accdbaacc71fe7