Pith. sign in

Paper Citation Record · LEDGER

VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2503.07135.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.07135 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:01:16.616474Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T11:01:02.081938Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3122f98a-e85a-494a-94b6-ed1f2baa0346 · inbound

OpenTie: Open-vocabulary Sequential Rebar Tying System cites this paper.

OpenTie: Open-vocabulary Sequential Rebar Tying System VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T16:19:17.903642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:19:17.903642Z digest=sha256:cf98eeb79ce4b38ff0089d6b0c1080b1f2b7929dac2515e2a610f784aecab57d

Observation 2be61ed7-57b6-4d4a-bee5-4819472138c9 · inbound

Calib3R: A 3D Foundation Model for Multi-Camera to Robot Calibration and 3D Metric-Scaled Scene Reconstruction cites this paper.

Calib3R: A 3D Foundation Model for Multi-Camera to Robot Calibration and 3D Metric-Scaled Scene Reconstruction VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:10:47.067346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:10:47.067346Z digest=sha256:b447dc7060aab806872d36c8868d8615ec1534c05698ce87e2cf09e153207db4

Observation f72c2d64-4e90-42a6-9ae8-083e281d9cf0 · inbound

3PoinTr: 3D Point Tracks for Learning Manipulation from Unconstrained Human Videos cites this paper.

3PoinTr: 3D Point Tracks for Learning Manipulation from Unconstrained Human Videos VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-15T12:36:50.019492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:36:50.019492Z digest=sha256:ab4323f22d5284be540ad0cddcdc7d522e62ab8ead874bb93925aa84757d034d

Observation 6fe2d1e5-deb1-4a79-a5dc-1aa4945546b0 · inbound

WARPED: Wrist-Aligned Rendering for Robot Policy Learning from Egocentric Human Demonstrations cites this paper.

WARPED: Wrist-Aligned Rendering for Robot Policy Learning from Egocentric Human Demonstrations VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:01:02.088600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:14:24.932972Z digest=sha256:ed75f14cff190569430d5f41dfc30f842ec6dc5219ad473d678020f79918b305

Observation 1a14e57d-602e-4b17-9612-b82b3bf0bf1c · inbound

VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances cites this paper.

VLAff: Vision-Language-Affordance Model for Unified Actionable Affordances VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T18:01:16.616474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:01:16.616474Z digest=sha256:f992a0335bf8c6180525bdabd4112621398f70b1314aebd525f792a77cee4151