Pith. sign in

Paper Citation Record · LEDGER

Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2405.01527.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.01527 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:25:28.265492Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T00:59:19.447450Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a301df75-c8b9-42a5-bdcd-b266e91c134f · inbound

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation cites this paper.

Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T12:17:01.361481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T12:17:01.294466Z digest=sha256:b6aef1eba7686b326f00c1f8f2874810276b7c9afb9dbeb6ab37fc29f3491078

Observation 5dab8f85-de31-4452-99a7-9f3f3b0d10ef · inbound

Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets cites this paper.

Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:25:00.455812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T16:25:00.365534Z digest=sha256:f72b6c8474201de4d592e79e40881905af37a50b81174bda25304d5ea8a5a2ca

Observation 00d3f223-6257-4740-a84f-450d192cd370 · inbound

DreamGen: Unlocking Generalization in Robot Learning through Video World Models cites this paper.

DreamGen: Unlocking Generalization in Robot Learning through Video World Models Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:50:45.584784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T23:50:45.332466Z digest=sha256:dea738e91dccc4d77e129304d7bd52d404699e53736f1f3130c9a5eafb8c4b8e

Observation 3fa80b8f-6101-4a69-ad93-300ccb2033eb · inbound

AMPLIFY: Actionless Motion Priors for Robot Learning from Videos cites this paper.

AMPLIFY: Actionless Motion Priors for Robot Learning from Videos Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:25:28.265492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:25:28.265492Z digest=sha256:99eca0b341b418e405fff737fb16ac67a3562cdfb96b35edbe010f7d449bde7c

Observation 5e5329b4-b67c-421f-8cc4-6118f9da2f45 · inbound

GR-3 Technical Report cites this paper.

GR-3 Technical Report Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T08:04:12.537478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T08:04:12.433863Z digest=sha256:379da0c158d4d676035ef2778ec35522f5b1a263b68979dc8161d85741327781

Observation b5b9a121-0666-4845-8bf0-3fff5f18d14d · inbound

Deep Sensorimotor Control by Imitating Predictive Models of Human Motion cites this paper.

Deep Sensorimotor Control by Imitating Predictive Models of Human Motion Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T16:22:54.459581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:22:54.459581Z digest=sha256:7125c6407f140f49d3bbc91d71370e77653ffe4aee16b312d87b6a10692bbf42

Observation 60f3705c-ae27-4808-a058-02e31495b9c4 · inbound

AFFORD2ACT: Affordance-Guided Automatic Keypoint Selection for Generalizable and Lightweight Robotic Manipulation cites this paper.

AFFORD2ACT: Affordance-Guided Automatic Keypoint Selection for Generalizable and Lightweight Robotic Manipulation Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:21:15.239747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T10:17:53.226866Z digest=sha256:e74886cb3aba90c2d4049c8b896c94f97fcde1ebac294a721e183ab18582a93c

Observation 099d3ba7-accf-445f-925f-2e6523593aa6 · inbound

X-Diffusion: Training Diffusion Policies on Cross-Embodiment Human Demonstrations cites this paper.

X-Diffusion: Training Diffusion Policies on Cross-Embodiment Human Demonstrations Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:40:33.732996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T00:37:12.170711Z digest=sha256:bdb1efce22d3429eabc4d31ebbbeb9b408ec20cf2b4ea3bb05c2540909d14383

Observation 9a8c9833-1ac9-4c71-97d7-58e59ce29a84 · inbound

3PoinTr: 3D Point Tracks for Learning Manipulation from Unconstrained Human Videos cites this paper.

3PoinTr: 3D Point Tracks for Learning Manipulation from Unconstrained Human Videos Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-15T12:36:50.019492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:36:50.019492Z digest=sha256:2d8137519847eda2c9d496c9ec44eb8198df7a3288007431f5a922b3004c4740

Observation b71cd742-32d9-4c7d-9528-783fd72a2936 · inbound

From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data cites this paper.

From Video to Control: A Survey of Learning Manipulation Interfaces from Temporal Visual Data Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:03:01.152804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T17:02:18.358675Z digest=sha256:e3450e9982fa7eee13c968c8350d931d5714464474f22aba4b2a1a9898c79b86

Observation 0acdc831-d45a-4b54-9d8a-f90c63bbef19 · inbound

WARPED: Wrist-Aligned Rendering for Robot Policy Learning from Egocentric Human Demonstrations cites this paper.

WARPED: Wrist-Aligned Rendering for Robot Policy Learning from Egocentric Human Demonstrations Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:01:02.286265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:14:24.932972Z digest=sha256:f11c59a3189e40b1159246089b16ff60447cd0e1e612882c7a7431780561c096

Observation 92964978-90f2-4e73-a6fa-0d41ec92645e · inbound

Dexterous Point Policy: Learning Point-based Dexterous Hand Policies from Human Demonstrations cites this paper.

Dexterous Point Policy: Learning Point-based Dexterous Hand Policies from Human Demonstrations Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:57:38.562581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T13:31:04.351106Z digest=sha256:f9c8de0442f1753169c1d377b45663b5ce15b8d3ce10b36c4e04a1e564a9e644

Observation 992a482a-38aa-44c1-9021-a17f6032637b · inbound

Do as I Do: Dexterous Manipulation Data from Everyday Human Videos cites this paper.

Do as I Do: Dexterous Manipulation Data from Everyday Human Videos Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:59:19.450427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T20:51:21.209882Z digest=sha256:0b67d885b330c483066c9880114ea3a8c8ddaed4b81a98cb25392acc673de603

Observation 4aa3c5c8-c4ad-49c4-8595-64319aa71636 · inbound

KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation cites this paper.

KAM-WM: Kinematic Affordance Maps from Latent World Models for Robot Manipulation Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-11T15:45:43.483298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T15:45:43.483298Z digest=sha256:51046f49a10101d4b68201f528e46be8c6384f571a411f29efe904864e3ce21c

Observation 60085264-cc85-4c81-a0bc-5fa5b733b795 · inbound

Data Pyramid for Embodied Manipulation cites this paper.

Data Pyramid for Embodied Manipulation Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T06:18:55.238293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:18:55.238293Z digest=sha256:60a22b2f2678bce1732324df73fdff39e2cd7bf05c24073a380042f38d50dad9

Observation ded7e344-99c3-4e20-ba39-6e761221cd4b · inbound

Track4Action: Distilling World-Centric 3D Tracker into Vision-Language-Action Policies cites this paper.

Track4Action: Distilling World-Centric 3D Tracker into Vision-Language-Action Policies Track2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot Manipulation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T14:05:04.733720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:05:04.733720Z digest=sha256:93889a4fd2d3d25c3b17ac69636b8ca64c0560b99b51310351d96d32f3a81bdd