Pith. sign in

Paper Citation Record · LEDGER

Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2403.12943.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.12943 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:51:36.054289Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T22:16:36.381522Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1f54d54d-777c-4a64-807e-a2aad2cf201b · inbound

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation cites this paper.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.425532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:aedfc568da73cc626ecd162a6c3753a98a6b4ad8c6ef7219cb495afd66eaa05f

Observation 10acbeeb-755b-47a5-817c-206621baabf8 · inbound

MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos cites this paper.

MimicDroid: In-Context Learning for Humanoid Robot Manipulation from Human Play Videos Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T18:51:36.054289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:51:36.054289Z digest=sha256:f789818ca28d20a71a2b150eec3f0c9c310fc1fe40bb9540f6c664ded89e6176

Observation 6d471f07-c92a-459a-99ea-cfb68383ce45 · inbound

X-Diffusion: Training Diffusion Policies on Cross-Embodiment Human Demonstrations cites this paper.

X-Diffusion: Training Diffusion Policies on Cross-Embodiment Human Demonstrations Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:40:33.758626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T00:37:12.170711Z digest=sha256:0d38ae390140d9da212b3b2d4f92173e22812809aa9a86f87f550b8335d7bfbf

Observation 9822926b-549b-4efd-8df1-35e56a64c3a1 · inbound

MonoDuo: Using One Robot Arm to Learn Bimanual Policies cites this paper.

MonoDuo: Using One Robot Arm to Learn Bimanual Policies Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:23:12.504103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:22:17.679021Z digest=sha256:4a09dc047dc259c843b80fc01b52c2feb98c848c504f949eee816254b12eaaf9

Observation 22f8a5fe-0331-4c21-bd76-2aad2605fb70 · inbound

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling cites this paper.

General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 180

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:33:27.849572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T13:33:03.368006Z digest=sha256:9bfb2a4aada7fbe29dfd0364a0467b66cde86388923a768d9a7d36e5af3e5c40

Observation a5ae3bd1-d1c1-412e-9acc-fa82ad88fd8c · inbound

SynthICL: Scalable In-context Imitation Learning with Synthetic Data cites this paper.

SynthICL: Scalable In-context Imitation Learning with Synthetic Data Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:57:25.875163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T19:23:45.402022Z digest=sha256:a759540df8c47ab675643f78ad7bd934897d084e218f9540249c5e36a486af79

Observation bb398fcd-6270-4c20-b429-1c078dc459b9 · inbound

EgoInfinity: A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot Learning cites this paper.

EgoInfinity: A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot Learning Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:08:56.423229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T01:35:25.740110Z digest=sha256:92041566049dd048294bb7812acfcca7c9d7c42b959d90353d37a31894f1364f

Observation 3dbc32e8-6aaa-4d65-bf1b-f936ee1834f3 · inbound

In-Context World Modeling for Robotic Control cites this paper.

In-Context World Modeling for Robotic Control Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:00:09.584646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T19:12:22.513577Z digest=sha256:79ed470174141677cae6c35a3251441c508ed3615762da2605fd97684f293a92

Observation 984a43b8-6ba8-40d4-b44f-acfa9417ff3f · inbound

In-Context World Modeling for Robotic Control cites this paper.

In-Context World Modeling for Robotic Control Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:29:51.616019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T05:11:07.089829Z digest=sha256:36ec350e420bd234b933840a1cd0f72c2ec403e6e5ff31fe535b1d6d3a3beacf

Observation a3d179df-d985-4252-8814-069ac488d1a8 · inbound

In-Context World Modeling for Robotic Control cites this paper.

In-Context World Modeling for Robotic Control Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T12:05:57.682386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:05:57.682386Z digest=sha256:43c3ce5833593e54596e5f51c3e6c539a4447fd28ce8efe38b7bbd0245deafbd

Observation 0d801c73-4984-416d-8b0f-34d423ede1ed · inbound

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time cites this paper.

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-09T22:16:36.382745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-09T22:16:31.529359Z digest=sha256:9b8d3448a462790d26cf796e4919e20ea2dbe5a2553e51060258053bbcf42183

Observation fcbc64e7-e5f8-4c8e-b8bd-09bcc44480ca · inbound

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time cites this paper.

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T06:48:14.554799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T06:48:14.554799Z digest=sha256:4d3e750b8de31bca1c7d6a84cffb8d5e83a6b8434c96cdd648bf7435b7ec5a6f