Pith. sign in

REVIEW 2 major objections 5 minor 1 cited by

AMPLIFY: Actionless Motion Priors for Robot Learning from Videos

T0 review · 2 major / 5 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Conditioning a policy on latent motion tokens predicted from video-only training yields 1.2-2.2x low-data gains and the first zero-action-data LIBERO generalization, this paper argues.

desk verdict A well-built latent motion token framework with a genuinely interesting zero-shot LIBERO result, but the headline human-video transfer claim needs a missing control before it can be taken at face value. read the letter →

arxiv 2506.14198 v1 pith:IBGA2462 submitted 2025-06-17 cs.RO cs.CVcs.LG

classification cs.ROcs.CVcs.LG
keywords behaviorcloningaction-freevideolearningkeypointtrajectorypredictionlatentmotiontokensforwardandinversedynamicsrobotmanipulationcross-embodimentworldmodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

AMPLIFY sets out to show that the information needed to control a robot can be carried by action-free video, provided the video is distilled into compact discrete motion tokens. The paper splits policy learning into a forward dynamics model that any videos can train, predicting latent keypoint-motion tokens, and an inverse dynamics model that a small set of action-labeled interactions can train, decoding those tokens into action chunks. On this interface the paper reports 1.2-2.2x improvements in low-data policy learning, a 1.4x average real-world gain from adding human videos over the same action data, and an average 60.5% success rate on LIBERO tasks for which no target-task action data was ever seen. A sympathetic reading is that action generalization is limited less by the volume of action labels and more by the quality of the action-free motion prior.

What carries the argument

The load-bearing object is the latent motion token: single-step velocities extracted from CoTracker tracks of a re-initialized 20 by 20 grid of points, encoded by a causally-masked transformer, quantized with Finite Scalar Quantization (FSQ) into a fixed 2048-code space, and decoded through a local-window classification head that scores each point's next motion inside a 15 by 15 pixel window around its previous location. An autoregressive transformer predicts these tokens from the current image and a language task description, and a separate cross-attention transformer decoder, the inverse dynamics model, maps image, proprioception, and predicted tokens into a distribution over a 16-step action chunk, with temporal ensembling at inference. The discrete codebook and the local-window classification turn motion prediction into a tractable classification problem; the frozen token interface is what allows the forward and inverse stages to train on entirely different datasets.

What would settle it

Construct a pair of tasks whose keypoint motion is nearly identical over the 16-frame prediction horizon but whose correct actions diverge (for example, two LIBERO-Goal tasks that differ only in which of two objects is the target, filmed so the first 0.8 seconds of motion look alike). If AMPLIFY still solves both at a rate near its reported 60% average, the motion-token interface is more informative than the 2D-ambiguity caveat suggests; if success collapses toward the near-zero inverse-only baseline, the information bottleneck sits in the 2D track representation rather than in the action head.

Watch

Extended reading notes

Core claim

The central claim is that latent motion tokens are a sufficient interface between visual dynamics and control: single-step velocities of a re-initialized 20 by 20 keypoint grid are compressed into a discrete 2048-code space, a forward model trained only on videos and task descriptions predicts the next token sequence, and an inverse model that never sees the goal decodes those tokens into an action chunk. The paper argues this decomposition cleanly separates what motion defines a task from how a robot can perform it, letting video data and interaction data scale independently. Empirically the learned dynamics are more accurate than prior keypoint and full-video-prediction baselines, and the gains concentrate exactly where action labels are scarce: few-shot learning, cross-embodiment transfer from human video, and zero-shot generalization on LIBERO suites for which no action data was ever provided.

Load-bearing premise

The pipeline stands or falls on the assumption that a re-initialized 20 by 20 grid of 2D keypoint tracks, compressed into 2048 discrete codes, preserves the information needed to choose the right action, and that the point tracker's tracks can be treated as ground truth — even though the inverse model never sees the goal and the paper's Limitations section concedes that 2D tracks leave ambiguity between actions.

Editorial extensions

If this is right

  • Few-shot learning scales with video rather than action count: with only 2 demonstrations per task, AMPLIFY reaches roughly 1.94x the success rate of the keypoint baseline ATM on LIBERO.
  • Action-free human video transfers to robot control: adding human demonstrations to the forward model improves real-world success by 1.32x to 1.5x over the same policy fed only robot data.
  • Zero-action-data generalization is possible: trained on LIBERO-90 actions only, AMPLIFY averages 60.5% success on four unseen LIBERO suites while behavior-cloning baselines score near zero.
  • Latent motion tokens function as a world-model artifact: conditioning the AVDC video predictor on them improves PSNR, LPIPS, and SSIM on BridgeData v2.
  • More video, same actions: with action data held at 2 trajectories, task success rises from 0.12 with 2 training videos to 0.55 with 50, evidence that the prior keeps improving as video scales.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A consequence the paper motivates but does not test: if the motion token is the whole task channel, the inverse model should train on undirected play or exploration data as well as on expert demos, since it never sees goals.
  • The 2D ambiguity the paper concedes predicts where the method should fail — tasks whose distinguishing information sits in the goal rather than in pixel motion; the LIBERO-Goal set being the hardest zero-shot target (41%) is consistent with that, and swapping in 3D point tracking would be the direct stress test.
  • The interface framing suggests motion tokens could serve as a reference channel inside other policy families, since the ablations show Gaussian, diffusion, and flow-matching action heads perform nearly identically.
  • The non-monotonic video-scaling curve (0.34 at 5 videos vs 0.23 at 10) hints that the forward model's benefit depends on track diversity, not raw clip count; a scaling study over internet-scale, scene-diverse video would show whether the cross-embodiment gain saturates.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper introduces AMPLIFY, a three-stage framework for robot learning from action-free videos. It first tokenizes CoTracker keypoint trajectories into discrete FSQ codes via an autoencoder, then trains an autoregressive forward dynamics model to predict future motion tokens from an image and task description, and finally trains an inverse dynamics model to map predicted tokens to actions. Experiments evaluate track prediction accuracy on LIBERO, BridgeData v2, and Something-Something v2; downstream policy learning in in-distribution, few-shot, cross-embodiment, and zero-shot action-label settings; and conditional video prediction. The reported results include a 3.7x reduction in track prediction MSE over ATM, 1.2-2.2x improvements in few-shot policy learning, a 1.4x average real-world improvement when human videos are added, and 60.5% average success on LIBERO target tasks with zero target-task action data.

Significance. If the results hold, AMPLIFY makes a strong contribution by demonstrating that a compact latent motion representation learned from arbitrary videos can provide a useful prior for downstream control, enabling substantial gains in low-data regimes and a novel zero-shot action-label generalization result. The paper is well-structured, includes extensive ablations, and compares against strong baselines. The few-shot and zero-shot experiments include control variants (inverse-only and w/o-tracks) that isolate the motion-token contribution, and the video-scaling experiment in Appendix F.1 is a valuable diagnostic. However, the headline cross-embodiment claim is not supported by the presented comparison, and the track-prediction evaluation split is missing.

major comments (2)
  1. [§3.2, Table 4; Appendix F.1] The cross-embodiment transfer claim is not isolated: the 1.4x average improvement compares an AMPLIFY policy whose forward dynamics is trained on human and robot videos against a Diffusion Policy trained only on robot demonstrations. This differs in two ways at once (motion-token conditioning and additional video data), so the gain cannot be attributed to human videos. The paper's own video-scaling experiment (Table 14) shows that adding robot videos alone raises success from 0.12 to 0.55 in a low-data LIBERO setting, demonstrating that extra video data alone can produce large gains. To support the claim that action-free human videos provide a cross-embodiment benefit, the authors should add an AMPLIFY variant whose forward dynamics is trained only on robot videos and show that the addition of human videos yields a further improvement over that control.
  2. [§3.1, Table 2; Appendix D] The track-prediction evaluation split is not stated anywhere in the paper. It is unclear whether the forward dynamics model is evaluated on videos that were also used for training; if it is, the reported 3.7x MSE improvement and 2.5x pixel-accuracy improvement over ATM are not evidence of generalization. The paper must specify how the videos in each dataset (LIBERO, BridgeData v2, Something-Something v2) were partitioned into training and evaluation sets for the motion tokenizer and forward dynamics model, and confirm that the evaluation rollouts are disjoint from the training rollouts.
minor comments (5)
  1. [Section 1, Contribution 1] The claim of 'the first latent keypoint dynamics model' should be qualified with respect to prior work such as Moto [34] and Latent Action Pretraining [33], which also learn latent motion or action representations from videos; the novelty should be positioned more precisely.
  2. [Section 3.2, Table 4] The sentence 'The average improvements of 1.32×, 1.4×, and 1.5×' is confusing because the table reports a single average of 0.42 vs 0.58; the per-task or per-demonstration-count ratios should be reported explicitly.
  3. [Section 3.2, Cross-Embodiment Transfer] The text says 'we evaluate AMPLIFY in both the few-shot setting and the full demonstration setting,' but the main text does not define the exact demonstration counts for 'All' for each task; these are only listed in Appendix Table 9, so a cross-reference would help.
  4. [Section 3.2, real-world results] Success rates are reported without confidence intervals despite being averaged over 10 rollouts; standard errors or per-seed results would make the comparison more rigorous.
  5. [Appendix D.4] The preprocessing paragraph introduces the window subscript without defining it; e.g., 'length-τ video' and 'τ length-T windows' should be defined precisely.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: AMPLIFY's modular forward/inverse dynamics chain is independently benchmarked; the Table 4 human-video confound is an attribution gap, not a constructed equivalence.

full rationale

AMPLIFY's derivation chain is modular and not circular by construction. Motion tokens are produced by an FSQ autoencoder from CoTracker keypoint velocities (Eq. 1); the forward dynamics model predicts those tokens from image and language (Eq. 2); the inverse dynamics model maps tokens to action chunks using an external NLL action loss (Eq. 3). Each stage is trained on its own target, and none of the paper's equations define a predicted quantity as the fitted input. Downstream policy claims are checked against external baselines (Diffusion Policy, ATM, Track2Act, UniPi, BAKU, QueST) and include control variants (inverse-only, w/o tracks) that isolate the motion-token contribution. The only self-citation is QueST, which appears solely as a baseline in Table 3 and is not load-bearing for any method choice or uniqueness argument. The Limitations section openly acknowledges the 2D-track ambiguity, which is a scope caveat rather than a circular step. The cross-embodiment comparison in Table 4 lacks an AMPLIFY variant trained without human videos, so the 1.4x gain over Diffusion Policy is confounded by added robot-video data; Table 14 shows robot-video scaling alone can improve success. This is an experimental attribution gap and a correctness risk, not a construction-level circularity: the claim is not equivalent to its inputs by definition or equation. The paper is therefore self-contained against external benchmarks, with no significant circularity.

Assumptions & free parameters 8 free parameters · 6 assumptions · 1 invented entities

The core architecture is trained end-to-end from data; no physical constants are invoked. The method does rely on several domain assumptions and a set of hyperparameters selected by ablation on LIBERO-Long, plus an internal latent representation that has no external falsifiable handle. The most consequential unstated risk is the assumption that 2D keypoint tracks from CoTracker contain enough information for action inference.

free parameters (8)
  • Prediction horizon T = 16
    Chosen as a tradeoff: inverse dynamics improves with longer horizons, forward prediction accuracy improves with shorter horizons; 16 maximizes downstream policy success in Appendix E.
  • Local window size W = 15
    Decoder classifies velocities over a 15x15 pixel window; motions larger than the window cannot be represented.
  • FSQ codebook size = 2048
    Selected by ablation; 2048 marginally outperforms 512 and 1024 on LIBERO-Long delta AUC.
  • Latent code sequence length = 16
    Selected by ablation; 16 outperforms 2, 4, 8, and 32 on delta AUC.
  • Hidden dimension = 768
    Selected by ablation; 768 marginally beats 512 and 1024.
  • Transformer depths = 2 tokenizer, 8 forward, 4 inverse
    From the hyperparameter table; forward depth 8 versus 4 is an ablation choice.
  • Keypoint grid size N = 400 (20x20)
    Uniform grid reinitialized each frame; no ablation is shown for grid density.
  • Action loss temporal discount gamma = 0.99
    Borrowed from ACT; decays the influence of errors later in action chunks.
assumptions (6)
  • domain assumption CoTracker tracks are treated as ground truth for all downstream training and evaluation.
    Sec 2.1 and D.3 define target tracks from CoTracker without quantifying tracker error; any systematic tracker failure propagates into tokens and metrics.
  • domain assumption A reinitialized uniform 20x20 grid captures task-relevant motion across diverse tasks and embodiments.
    Sec 2.1 motivates the grid by simplicity; no coverage analysis or task-relevant point selection is provided.
  • domain assumption 2D pixel-space tracks suffice to disambiguate robot actions.
    Sec 5 concedes that multiple actions can map to the same 2D tracks; the inverse model also never sees goals.
  • domain assumption Environment dynamics are deterministic given observation and goal.
    Sec 5 states that stochastic settings require separating agent actions from exogenous noise; the current forward model ignores that separation.
  • domain assumption Human video motion priors transfer to robot action inference.
    Sec 3.2 relies on human videos in forward training, but no ablation isolates their contribution.
  • standard math FSQ and transformer training assumptions from prior work hold as described.
    FSQ implicit codebook with straight-through gradients and autoregressive causal transformers are used as black boxes in Sec 2.2.
invented entities (1)
  • Latent motion tokens (FSQ-quantized keypoint velocity codes)
    purpose: Compact intermediate representation that decouples forward dynamics from inverse action inference.
    The codebook is learned from CoTracker tracks and has no external falsifiable handle; its usefulness is shown only through internal reconstruction, policy success, and video prediction metrics in this paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of AMPLIFY: Actionless Motion Priors for Robot Learning from Videos." pith.science (2026). https://pith.science/paper/IBGA2462

@misc{pith2026250614198,
  author       = {Pith},
  title        = {Pith review of: AMPLIFY: Actionless Motion Priors for Robot Learning from Videos},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IBGA2462}},
  note         = {Machine review of arXiv:2506.14198}
}
read the original abstract

Action-labeled data for robotics is scarce and expensive, limiting the generalization of learned policies. In contrast, vast amounts of action-free video data are readily available, but translating these observations into effective policies remains a challenge. We introduce AMPLIFY, a novel framework that leverages large-scale video data by encoding visual dynamics into compact, discrete motion tokens derived from keypoint trajectories. Our modular approach separates visual motion prediction from action inference, decoupling the challenges of learning what motion defines a task from how robots can perform it. We train a forward dynamics model on abundant action-free videos and an inverse dynamics model on a limited set of action-labeled examples, allowing for independent scaling. Extensive evaluations demonstrate that the learned dynamics are both accurate, achieving up to 3.7x better MSE and over 2.5x better pixel prediction accuracy compared to prior approaches, and broadly useful. In downstream policy learning, our dynamics predictions enable a 1.2-2.2x improvement in low-data regimes, a 1.4x average improvement by learning from action-free human videos, and the first generalization to LIBERO tasks from zero in-distribution action data. Beyond robotic control, we find the dynamics learned by AMPLIFY to be a versatile latent world model, enhancing video prediction quality. Our results present a novel paradigm leveraging heterogeneous data sources to build efficient, generalizable world models. More information can be found at https://amplify-robotics.github.io/.

Figures

Figures reproduced from arXiv: 2506.14198 by the authors.

Figure 1
Figure 1. Overview. AMPLIFY decomposes policy learning into forward and inverse dynamics, using latent keypoint motion as an intermediate representation. The forward model can be trained on any video data, while the inverse model can be trained any interaction data. In contrast with behavior cloning (BC), AMPLIFY requires fewer demos, can generalize to tasks for which we have zero action data, and learn from human videos. Abs… view at source ↗
Figure 2
Figure 2. Architecture. AMPLIFY consists of a three-stage decomposition: (a) keypoint tracks are compressed into a discrete latent space using FSQ. For each timestep and each point, the decoder outputs a distribution in a local window centered around each point to reconstruct the instantaneous velocities, (b) a forward dynamics model is trained to predict the latent codes for the next T timesteps given an input image and task… view at source ↗
Figure 3
Figure 3. Decoded keypoint trajectory predictions from AMPLIFY. Zero-movement points are not shown. transformer using the flattened feature map from a pre-trained ResNet-18 [60] to generate 7 × 7 = 49 vision tokens per image. The summary token from a T5 [61] text embedding of the task description is used to tokenize language inputs. These conditioning tokens are then concatenated with a start of sequence (SOS) token and the l… view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: LIBERO few-shot. Comparison of AMPLIFY against ATM [54] and a no-video-pre-training baseline. Our forward model is trained on all videos, and the inverse model is only trained on a limited number of demos. evaluate AMPLIFY along four dimensions measuring (1) in-distrib…
Figure 5
Figure 5. Figure 5: A sample of the 130 diverse tasks and environment configurations in LIBERO. 2 RGB Corner Cameras 1 RGB Front Camera [PITH_FULL_IMAGE:figures/full_fig_p020_5.png]
Figure 6
Figure 6. Figure 6: We use three static RGB cameras as input observations for both human and robot (UR5) data. 20 [PITH_FULL_IMAGE:figures/full_fig_p020_6.png]
Figure 7
Figure 7. Figure 7: Real-World Tasks and Cameraviews. Each row shows front and corner camera views for three different tasks: Task 1: “Put the Rubik’s Cube on the Box”, Task 2: “Stack the Green and Blue Cups in the Orange Cup”, and Task 3: “Open the Box and Move the Eggplant into the Bowl…
Figure 8
Figure 8. Figure 8: Track Predictions from AMPLIFY on Real-World Robot Data. 29 [PITH_FULL_IMAGE:figures/full_fig_p029_8.png]
Figure 9
Figure 9. Figure 9: Track Predictions from AMPLIFY on Real-World Human Data. 30 [PITH_FULL_IMAGE:figures/full_fig_p030_9.png]
Figure 10
Figure 10. Figure 10: Video Predictions from AVDC [44] conditioned on AMPLIFY, trained on Bridge Data. 31 [PITH_FULL_IMAGE:figures/full_fig_p031_10.png]
Figure 11
Figure 11. Figure 11: Track Predictions from AMPLIFY on Bridge Data. 32 [PITH_FULL_IMAGE:figures/full_fig_p032_11.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. 3PoinTr: 3D Point Tracks for Learning Manipulation from Unconstrained Human Videos

    cs.RO 2026-03 conditional novelty 6.0 of 10

    Dense 3D point-track prediction from unconstrained human videos plus a track-conditioned closed-loop policy yields large sample-efficiency gains over BC and video-pretraining baselines.

Reference graph

Works this paper leans on

107 extracted references · 18 canonical work pages · cited by 1 Pith paper

  1. [1]

    Radford, J

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever. Language models are unsupervised multitask learners. 2019

  2. [2]

    Brown, B

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Nee- lakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-V oss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Rad- ford, I. Sutskever, and D. Am...

  3. [3]

    Achiam, S

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. Gpt-4 technical report.arXiv preprint arXiv:2303.08774, 2023

  4. [4]

    Radford, J

    A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever. Learning transferable visual models from natural language supervision, 2021. URLhttps://arxiv.org/abs/2103.00020

  5. [5]

    Ramesh, P

    A. Ramesh, P. Dhariwal, A. Nichol, C. Chu, and M. Chen. Hierarchical text-conditional image generation with clip latents, 2022. URLhttps://arxiv.org/abs/2204.06125

  6. [6]

    Rombach, A

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-resolution image synthesis with latent diffusion models, 2022. URL https://arxiv.org/abs/2112. 10752

  7. [7]

    Brohan, N

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, J. Dabis, C. Finn, K. Gopalakrishnan, K. Haus- man, A. Herzog, J. Hsu, J. Ibarz, B. Ichter, A. Irpan, T. Jackson, S. Jesmonth, N. J. Joshi, R. Julian, D. Kalashnikov, Y . Kuang, I. Leal, K.-H. Lee, S. Levine, Y . Lu, U. Malla, D. Manju- nath, I. Mordatch, O. Nachum, C. Parada, J. Peralta, E. Perez, K. Pertsc...

  8. [8]

    Brohan, N

    A. Brohan, N. Brown, J. Carbajal, Y . Chebotar, X. Chen, K. Choromanski, T. Ding, D. Driess, A. Dubey, C. Finn, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control.arXiv preprint arXiv:2307.15818, 2023

Show all 107 references
  1. [9]

    Padalkar, A

    A. Padalkar, A. Pooley, A. Jain, A. Bewley, A. Herzog, A. Irpan, A. Khazatsky, A. Rai, A. Singh, A. Brohan, et al. Open x-embodiment: Robotic learning datasets and rt-x models. arXiv preprint arXiv:2310.08864, 2023

  2. [11]

    Pertsch, K

    K. Pertsch, K. Stachowicz, B. Ichter, D. Driess, S. Nair, Q. Vuong, O. Mees, C. Finn, and S. Levine. Fast: Efficient action tokenization for vision-language-action models.arXiv preprint arXiv:2501.09747, 2025

  3. [12]

    L. X. Shi, B. Ichter, M. Equi, L. Ke, K. Pertsch, Q. Vuong, J. Tanner, A. Walling, H. Wang, N. Fusai, et al. Hi robot: Open-ended instruction following with hierarchical vision-language- action models.arXiv preprint arXiv:2502.19417, 2025

  4. [13]

    Khazatsky, K

    A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, P. D. Fagan, J. Hejna, M. Itkina, M. Lepert, Y . J. Ma, P. T. Miller, J. Wu, S. Belkhale, S. Dass, H. Ha, A. Jain, A. Lee, Y . Lee, M. Memmel, S. Pa...

  5. [14]

    M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn. OpenVLA: An Open-Source Vision-Language-Action Model, June

  6. [15]

    Black, N

    K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, J. Tanner, Q. Vuong, A. Walling, H. Wang, and U. Zhilinsky. $pi_0$: A Visi...

  7. [16]

    Y . Du, M. Yang, B. Dai, H. Dai, O. Nachum, J. B. Tenenbaum, D. Schuurmans, and P. Abbeel. Learning Universal Policies via Text-Guided Video Generation, Nov. 2023. URL http: //arxiv.org/abs/2302.00111. arXiv:2302.00111 [cs]

  8. [17]

    McCarthy, D

    R. McCarthy, D. C. H. Tan, D. Schmidt, F. Acero, N. Herr, Y . Du, T. G. Thuruthel, and Z. Li. Towards Generalist Robot Learning from Internet Video: A Survey, June 2024. URL http://arxiv.org/abs/2404.19664. arXiv:2404.19664 [cs]

  9. [18]

    Rybkin, K

    O. Rybkin, K. Pertsch, K. G. Derpanis, K. Daniilidis, and A. Jaegle. Learning what you can do before doing anything, Feb. 2019. URL http://arxiv.org/abs/1806.09655. arXiv:1806.09655 [cs, stat]. 10

  10. [19]

    Agarwal, A

    NVIDIA, N. Agarwal, A. Ali, M. Bala, Y . Balaji, E. Barker, T. Cai, P. Chattopadhyay, Y . Chen, Y . Cui, Y . Ding, D. Dworakowski, J. Fan, M. Fenzi, F. Ferroni, S. Fidler, D. Fox, S. Ge, Y . Ge, J. Gu, S. Gururani, E. He, J. Huang, J. Huffman, P. Jannaty, J. Jin, S. W. Kim, G....

  11. [20]

    J. Ho, W. Chan, C. Saharia, J. Whang, R. Gao, A. Gritsenko, D. P. Kingma, B. Poole, M. Norouzi, D. J. Fleet, and T. Salimans. Imagen video: High definition video generation with diffusion models, 2022. URLhttps://arxiv.org/abs/2210.02303

  12. [21]

    W. Kong, Q. Tian, Z. Zhang, R. Min, Z. Dai, J. Zhou, J. Xiong, X. Li, B. Wu, J. Zhang, et al. Hunyuanvideo: A systematic framework for large video generative models.arXiv preprint arXiv:2412.03603, 2024

  13. [22]

    Gupta, A

    Veo-Team, :, A. Gupta, A. Razavi, A. Toor, A. Gupta, D. Erhan, E. Shaw, E. Lau, F. Belletti, G. Barth-Maron, G. Shaw, H. Erdogan, H. Sidahmed, H. Nandwani, H. Moraldo, H. Kim, I. Blok, J. Donahue, J. Lezama, K. Mathewson, K. David, M. K. Lorrain, M. van Zee, M. Narasimhan, M. ...

  14. [23]

    Bar-Tal, H

    O. Bar-Tal, H. Chefer, O. Tov, C. Herrmann, R. Paiss, S. Zada, A. Ephrat, J. Hur, G. Liu, A. Raj, et al. Lumiere: A space-time diffusion model for video generation. InSIGGRAPH Asia 2024 Conference Papers, pages 1–11, 2024

  15. [24]

    Y . J. Ma, S. Sodhani, D. Jayaraman, O. Bastani, V . Kumar, and A. Zhang. Vip: Towards universal visual reward and representation via value-implicit pre-training.arXiv preprint arXiv:2210.00030, 2022

  16. [25]

    Ghosh, C

    D. Ghosh, C. Bhateja, and S. Levine. Reinforcement Learning from Passive Data via Latent Intentions, Apr. 2023. URLhttp://arxiv.org/abs/2304.04782. arXiv:2304.04782 [cs, stat]

  17. [26]

    Bhateja, D

    C. Bhateja, D. Guo, D. Ghosh, A. Singh, M. Tomar, Q. Vuong, Y . Chebotar, S. Levine, and A. Kumar. Robotic Offline RL from Internet Videos via Value-Function Pre-Training, Sept

  18. [27]

    Dashora, D

    N. Dashora, D. Ghosh, and S. Levine. Viva: Video-trained value functions for guiding online rl from diverse data.arXiv preprint arXiv:2503.18210, 2025

  19. [28]

    Sermanet, C

    P. Sermanet, C. Lynch, Y . Chebotar, J. Hsu, E. Jang, S. Schaal, S. Levine, and G. Brain. Time-contrastive networks: Self-supervised learning from video. In2018 IEEE international conference on robotics and automation (ICRA), pages 1134–1141. IEEE, 2018

  20. [29]

    S. Nair, A. Rajeswaran, V . Kumar, C. Finn, and A. Gupta. R3m: A universal visual representa- tion for robot manipulation.arXiv preprint arXiv:2203.12601, 2022

  21. [30]

    T. M. Moerland, J. Broekens, A. Plaat, C. M. Jonker, et al. Model-based reinforcement learning: A survey.Foundations and Trends® in Machine Learning, 16(1):1–118, 2023. 11

  22. [31]

    Y . Hu, Y . Guo, P. Wang, X. Chen, Y .-J. Wang, J. Zhang, K. Sreenath, C. Lu, and J. Chen. Video prediction policy: A generalist robot policy with predictive visual representations.arXiv preprint arXiv:2412.14803, 2024

  23. [32]

    Bruce, M

    J. Bruce, M. D. Dennis, A. Edwards, J. Parker-Holder, Y . Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, et al. Genie: Generative interactive environments. In Forty-first International Conference on Machine Learning, 2024

  24. [33]

    S. Ye, J. Jang, B. Jeon, S. Joo, J. Yang, B. Peng, A. Mandlekar, R. Tan, Y .-W. Chao, B. Y . Lin, et al. Latent action pretraining from videos.arXiv preprint arXiv:2410.11758, 2024

  25. [34]

    Y . Chen, Y . Ge, Y . Li, Y . Ge, M. Ding, Y . Shan, and X. Liu. Moto: Latent Motion Token as the Bridging Language for Robot Manipulation, Dec. 2024. URL http://arxiv.org/abs/ 2412.04445. arXiv:2412.04445 [cs]

  26. [35]

    Z. Cui, H. Pan, A. Iyer, S. Haldar, and L. Pinto. Dynamo: In-domain dynamics pretraining for visuo-motor control.Advances in Neural Information Processing Systems, 37:33933–33961, 2024

  27. [36]

    Karaev, I

    N. Karaev, I. Rocco, B. Graham, N. Neverova, A. Vedaldi, and C. Rupprecht. CoTracker: It is better to track together. 2023

  28. [37]

    Wang, Y .-Y

    Q. Wang, Y .-Y . Chang, R. Cai, Z. Li, B. Hariharan, A. Holynski, and N. Snavely. Tracking everything everywhere all at once. InInternational Conference on Computer Vision, 2023

  29. [38]

    Doersch, Y

    C. Doersch, Y . Yang, M. Vecerik, D. Gokay, A. Gupta, Y . Aytar, J. Carreira, and A. Zisserman. Tapir: Tracking any point with per-frame initialization and temporal refinement. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 10061–10072, 2023

  30. [39]

    Y . Xiao, Q. Wang, S. Zhang, N. Xue, S. Peng, Y . Shen, and X. Zhou. Spatialtracker: Tracking any 2d pixels in 3d space. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2024

  31. [40]

    Vecerik, C

    M. Vecerik, C. Doersch, Y . Yang, T. Davchev, Y . Aytar, G. Zhou, R. Hadsell, L. Agapito, and J. Scholz. RoboTAP: Tracking Arbitrary Points for Few-Shot Visual Imitation, Aug. 2023. URLhttp://arxiv.org/abs/2308.15975. arXiv:2308.15975 [cs]

  32. [41]

    Z. Qin, K. Fang, Y . Zhu, L. Fei-Fei, and S. Savarese. Keto: Learning keypoint representations for tool manipulation. In2020 IEEE International Conference on Robotics and Automation (ICRA), pages 7278–7285. IEEE, 2020

  33. [42]

    C. Yuan, C. Wen, T. Zhang, and Y . Gao. General flow as foundation affordance for scalable robot learning.arXiv preprint arXiv:2401.11439, 2024

  34. [43]

    L.-H. Lin, Y . Cui, A. Xie, T. Hua, and D. Sadigh. FlowRetrieval: Flow-Guided Data Retrieval for Few-Shot Imitation Learning, Oct. 2024. URL http://arxiv.org/abs/2408. 16944. arXiv:2408.16944

  35. [44]

    P.-C. Ko, J. Mao, Y . Du, S.-H. Sun, and J. B. Tenenbaum. Learning to Act from Actionless Videos through Dense Correspondences.arXiv:2310.08576, 2023

  36. [45]

    Kareer, D

    S. Kareer, D. Patel, R. Punamiya, P. Mathur, S. Cheng, C. Wang, J. Hoffman, and D. Xu. Egomimic: Scaling imitation learning via egocentric video, 2024. URL https://arxiv. org/abs/2410.24221

  37. [46]

    J. Gao, Z. Tao, N. Jaquier, and T. Asfour. K-VIL: Keypoints-Based Visual Imitation Learn- ing.IEEE Transactions on Robotics, 39(5):3888–3908, Oct. 2023. ISSN 1941-0468. doi: 10.1109/TRO.2023.3286074. URL https://ieeexplore.ieee.org/abstract/ document/10189175. Conference Name:...

  38. [47]

    Fang, B.-R

    X. Fang, B.-R. Huang, J. Mao, J. Shone, J. B. Tenenbaum, T. Lozano-Pérez, and L. P. Kaelbling. Keypoint Abstraction using Large Models for Object-Relative Imitation Learning, Oct. 2024. URLhttp://arxiv.org/abs/2410.23254. arXiv:2410.23254

  39. [48]

    C. Gao, H. Zhang, Z. Xu, C. Zhehao, and L. Shao. Flip: Flow-centric generative planning as general-purpose manipulation world model. InThe Thirteenth International Conference on Learning Representations, 2025

  40. [49]

    Manuelli, W

    L. Manuelli, W. Gao, P. Florence, and R. Tedrake. kpam: Keypoint affordances for category- level robotic manipulation. InThe International Symposium of Robotics Research, pages 132–157. Springer, 2019

  41. [50]

    Guzey, Y

    I. Guzey, Y . Dai, G. Savva, R. Bhirangi, and L. Pinto. Bridging the Human to Robot Dexterity Gap through Object-Oriented Rewards, Oct. 2024. URL http://arxiv.org/abs/ 2410.23289. arXiv:2410.23289

  42. [51]

    M. Xu, Z. Xu, Y . Xu, C. Chi, G. Wetzstein, M. Veloso, and S. Song. Flow as the Cross-Domain Manipulation Interface, July 2024. URL http://arxiv.org/abs/2407.15208. arXiv:2407.15208 [cs]

  43. [52]

    Bharadhwaj, D

    H. Bharadhwaj, D. Dwibedi, A. Gupta, S. Tulsiani, C. Doersch, T. Xiao, D. Shah, F. Xia, D. Sadigh, and S. Kirmani. Gen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot Manipulation, Sept. 2024. URL https://arxiv.org/abs/2409. 16283v1

  44. [53]

    Bharadhwaj, R

    H. Bharadhwaj, R. Mottaghi, A. Gupta, and S. Tulsiani. Track2act: Predicting point tracks from internet videos enables diverse zero-shot robot manipulation.arXiv preprint arXiv:2405.01527, 2024

  45. [54]

    C. Wen, X. Lin, J. So, K. Chen, Q. Dou, Y . Gao, and P. Abbeel. Any-point Trajectory Mod- eling for Policy Learning, Feb. 2024. URL http://arxiv.org/abs/2401.00025. arXiv:2401.00025 [cs]

  46. [55]

    Hansen, X

    N. Hansen, X. Wang, and H. Su. Temporal difference learning for model predictive control. In ICML, 2022

  47. [56]

    Hansen, H

    N. Hansen, H. Su, and X. Wang. TD-MPC2: Scalable, Robust World Models for Continuous Control, Mar. 2024. URL http://arxiv.org/abs/2310.16828. arXiv:2310.16828 [cs]

  48. [57]

    Scannell, M

    A. Scannell, M. Nakhaei, K. Kujanpää, Y . Zhao, K. S. Luck, A. Solin, and J. Pajarinen. Discrete codebook world models for continuous control.arXiv preprint arXiv:2503.00653, 2025

  49. [58]

    Mentzer, D

    F. Mentzer, D. Minnen, E. Agustsson, and M. Tschannen. Finite Scalar Quantization: VQ-V AE Made Simple, Oct. 2023. URL http://arxiv.org/abs/2309.15505. arXiv:2309.15505 [cs]

  50. [59]

    A. v. d. Oord, O. Vinyals, and K. Kavukcuoglu. Neural Discrete Representation Learning, May 2018. URLhttp://arxiv.org/abs/1711.00937. arXiv:1711.00937 [cs]

  51. [60]

    K. He, X. Zhang, S. Ren, and J. Sun. Deep Residual Learning for Image Recognition. In2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, June

  52. [61]

    Raffel, N

    C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y . Zhou, W. Li, and P. J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer.Journal of machine learning research, 21(140):1–67, 2020. 13

  53. [62]

    Haldar, Z

    S. Haldar, Z. Peng, and L. Pinto. Baku: An efficient transformer for multi-task policy learning. arXiv preprint arXiv:2406.07539, 2024

  54. [63]

    T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware

  55. [64]

    H. R. Walke, K. Black, T. Z. Zhao, Q. Vuong, C. Zheng, P. Hansen-Estruch, A. W. He, V . Myers, M. J. Kim, M. Du, et al. Bridgedata v2: A dataset for robot learning at scale. In Conference on Robot Learning, pages 1723–1736. PMLR, 2023

  56. [65]

    something something

    R. Goyal, S. E. Kahou, V . Michalski, J. Materzy ´nska, S. Westphal, H. Kim, V . Haenel, I. Fruend, P. Yianilos, M. Mueller-Freitag, F. Hoppe, C. Thurau, I. Bax, and R. Memisevic. The "something something" video database for learning and evaluating visual common sense,

  57. [66]

    B. Liu, Y . Zhu, C. Gao, Y . Feng, Q. Liu, Y . Zhu, and P. Stone. Libero: Benchmarking knowledge transfer for lifelong robot learning.arXiv preprint arXiv:2306.03310, 2023

  58. [67]

    X. Gu, C. Wen, W. Ye, J. Song, and Y . Gao. Seer: Language instructed video prediction with latent diffusion models.arXiv preprint arXiv:2303.14897, 2023

  59. [69]

    A. Mete, H. Xue, A. Wilcox, Y . Chen, and A. Garg. QueST: Self-Supervised Skill Abstractions for Learning Continuous Control, Sept. 2024. URL http://arxiv.org/abs/2407. 15840. arXiv:2407.15840 [cs]

  60. [70]

    T. D. Ngo, P. Zhuang, C. Gan, E. Kalogerakis, S. Tulyakov, H.-Y . Lee, and C. Wang. Delta: Dense efficient long-range 3d tracking for any video.arXiv preprint arXiv:2410.24211, 2024

  61. [71]

    Misra, A

    D. Misra, A. Saran, T. Xie, A. Lamb, and J. Langford. Towards principled representation learning from videos for reinforcement learning.arXiv preprint arXiv:2403.13765, 2024

  62. [72]

    M. Yang, D. Schuurmans, P. Abbeel, and O. Nachum. Dichotomy of control: Separating what you can control from what you cannot.arXiv preprint arXiv:2210.13435, 2022

  63. [73]

    S. Park, D. Ghosh, B. Eysenbach, and S. Levine. Hiql: Offline goal-conditioned rl with latent states as actions.Advances in Neural Information Processing Systems, 36:34866–34891, 2023

  64. [74]

    C. Wang, L. Fan, J. Sun, R. Zhang, L. Fei-Fei, D. Xu, Y . Zhu, and A. Anandkumar. Mimicplay: Long-horizon imitation learning by watching human play.arXiv preprint arXiv:2302.12422, 2023

  65. [75]

    L. Chen, S. Bahl, and D. Pathak. PlayFusion: Skill Acquisition via Diffusion from Language- Annotated Play. InProceedings of The 7th Conference on Robot Learning, pages 2012–2029. PMLR, Dec. 2023. URL https://proceedings.mlr.press/v229/chen23c. html. ISSN: 2640-3498

  66. [76]

    Lynch, M

    C. Lynch, M. Khansari, T. Xiao, V . Kumar, J. Tompson, S. Levine, and P. Sermanet. Learning Latent Plans from Play. InProceedings of the Conference on Robot Learning, pages 1113–1132. PMLR, May 2020. URL https://proceedings.mlr.press/v100/lynch20a. html. ISSN: 2640-3498

  67. [77]

    Bharadhwaj, A

    H. Bharadhwaj, A. Gupta, S. Tulsiani, and V . Kumar. Zero-shot robot manipulation from passive human videos.arXiv preprint arXiv:2302.02011, 2023

  68. [78]

    Xiong, Q

    H. Xiong, Q. Li, Y .-C. Chen, H. Bharadhwaj, S. Sinha, and A. Garg. Learning by watching: Physical imitation of manipulation skills from human videos. In2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 7827–7834. IEEE, 2021. 14

  69. [79]

    K. Shaw, S. Bahl, and D. Pathak. Videodex: Learning dexterity from internet videos. In Conference on Robot Learning, pages 654–665. PMLR, 2023

  70. [80]

    Mandikal and K

    P. Mandikal and K. Grauman. Dexvip: Learning dexterous grasping with human hand pose priors from video. InConference on Robot Learning, pages 651–661. PMLR, 2022

  71. [81]

    S. Bahl, A. Gupta, and D. Pathak. Human-to-robot imitation in the wild.arXiv preprint arXiv:2207.09450, 2022

  72. [82]

    S. Bahl, R. Mendonca, L. Chen, U. Jain, and D. Pathak. Affordances from human videos as a versatile representation for robotics. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13778–13790, 2023

  73. [83]

    Mendonca, S

    R. Mendonca, S. Bahl, and D. Pathak. Structured world models from human videos.arXiv preprint arXiv:2308.10901, 2023

  74. [84]

    S. Nair, E. Mitchell, K. Chen, S. Savarese, C. Finn, et al. Learning language-conditioned robot behavior from offline data and crowd-sourced annotation. InConference on Robot Learning, pages 1303–1315. PMLR, 2022

  75. [85]

    Schmeckpeper, O

    K. Schmeckpeper, O. Rybkin, K. Daniilidis, S. Levine, and C. Finn. Reinforcement learning with videos: Combining offline observations with interaction.arXiv preprint arXiv:2011.06507, 2020

  76. [86]

    Bhateja, D

    C. Bhateja, D. Guo, D. Ghosh, A. Singh, M. Tomar, Q. Vuong, Y . Chebotar, S. Levine, and A. Kumar. Robotic offline rl from internet videos via value-function pre-training.arXiv preprint arXiv:2309.13041, 2023

  77. [87]

    Ghosh, C

    D. Ghosh, C. A. Bhateja, and S. Levine. Reinforcement learning from passive data via latent intentions. InInternational Conference on Machine Learning, pages 11321–11339. PMLR, 2023

  78. [88]

    Bruce, M

    J. Bruce, M. Dennis, A. Edwards, J. Parker-Holder, Y . Shi, E. Hughes, M. Lai, A. Mavalankar, R. Steigerwald, C. Apps, Y . Aytar, S. Bechtle, F. Behbahani, S. Chan, N. Heess, L. Gonzalez, S. Osindero, S. Ozair, S. Reed, J. Zhang, K. Zolna, J. Clune, N. de Freitas, S. Singh, an...

  79. [89]

    Escontrela, A

    A. Escontrela, A. Adeniji, W. Yan, A. Jain, X. B. Peng, K. Goldberg, Y . Lee, D. Hafner, and P. Abbeel. Video prediction models as rewards for reinforcement learning.arXiv preprint arXiv:2305.14343, 2023

  80. [90]

    Huang, G

    T. Huang, G. Jiang, Y . Ze, and H. Xu. Diffusion reward: Learning rewards via conditional video diffusion.arXiv preprint arXiv:2312.14134, 2023

  81. [91]

    Schmidt and M

    D. Schmidt and M. Jiang. Learning to act without actions. InThe Twelfth International Conference on Learning Representations (ICLR), 2024

  82. [92]

    Y . Wen, J. Lin, Y . Zhu, J. Han, H. Xu, S. Zhao, and X. Liang. Vidman: Exploiting implicit dynamics from video diffusion model for effective robot manipulation, 2024. URL https: //arxiv.org/abs/2411.09153

  83. [93]

    Liang, R

    J. Liang, R. Liu, E. Ozguroglu, S. Sudhakar, A. Dave, P. Tokmakov, S. Song, and C. V ondrick. Dreamitate: Real-world visuomotor policy learning via video generation, 2024. URL https: //arxiv.org/abs/2406.16862

  84. [94]

    Jaegle, F

    A. Jaegle, F. Gimeno, A. Brock, A. Zisserman, O. Vinyals, and J. Carreira. Perceiver: General Perception with Iterative Attention, June 2021. URL http://arxiv.org/abs/2103. 03206. arXiv:2103.03206 [cs, eess]. 15

  85. [95]

    Oquab, T

    M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby, et al. Dinov2: Learning robust visual features without supervision.arXiv preprint arXiv:2304.07193, 2023

  86. [96]

    Dosovitskiy, L

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. De- hghani, M. Minderer, G. Heigold, S. Gelly, et al. An image is worth 16x16 words: Transformers for image recognition at scale.arXiv preprint arXiv:2010.11929, 2020

  87. [97]

    A. Veit, M. Wilber, and S. Belongie. Residual Networks Behave Like Ensembles of Rel- atively Shallow Networks, Oct. 2016. URL http://arxiv.org/abs/1605.06431. arXiv:1605.06431 [cs]

  88. [98]

    Lipman, R

    Y . Lipman, R. T. Q. Chen, H. Ben-Hamu, M. Nickel, and M. Le. Flow matching for generative modeling, 2023. URLhttps://arxiv.org/abs/2210.02747

  89. [99]

    Bjorck, F

    NVIDIA, :, J. Bjorck, F. Castañeda, N. Cherniadev, X. Da, R. Ding, L. J. Fan, Y . Fang, D. Fox, F. Hu, S. Huang, J. Jang, Z. Jiang, J. Kautz, K. Kundalia, L. Lao, Z. Li, Z. Lin, K. Lin, G. Liu, E. Llontop, L. Magne, A. Mandlekar, A. Narayan, S. Nasiriany, S. Reed, Y . L. Tan, ...

  90. [100]

    C. Chi, S. Feng, Y . Du, Z. Xu, E. Cousineau, B. Burchfiel, and S. Song. Diffusion Policy: Visuomotor Policy Learning via Action Diffusion, Mar. 2023. URL http://arxiv.org/ abs/2303.04137. arXiv:2303.04137 [cs]. 16 Appendix In this document, we provide detailed supplementary m...

  91. [106]

    Require:DatasetsV,R 1:Preprocess keypoint tracksκ t inV 2:Learn latent motion encoding to compressκ t into discrete tokensz t using Eq

    Use Action-Free Videos (and Expert Demonstrations) to learn how observations evolve with respect to a goal 17 Algorithm 1AMPLIFYTraining. Require:DatasetsV,R 1:Preprocess keypoint tracksκ t inV 2:Learn latent motion encoding to compressκ t into discrete tokensz t using Eq. 1 3...

  92. [107]

    Put the Rubik’s Cube on the Box

    Use Interaction Data (Undirected and Demonstrations) to learn how a sequence of observations maps to a sequence of actions This decomposition effectively decouplestask understanding(the sequence of observations that correspond to a goal) andtask execution(translating a referen...

  93. [108]

    MSE: We take normalized track predictions in the range [−1,1] and compute MSE between predicted(x, y)values for each point and the corresponding ground-truth point from CoTracker

  94. [109]

    3.∆ AUC: This metric was originally introduced by works presenting point tracking [ 38, 36] and later used for track prediction [53]

    Pixel Accuracy: We measure the (normalized) percentage of predictions that are pixel-perfect compared to ground-truth. 3.∆ AUC: This metric was originally introduced by works presenting point tracking [ 38, 36] and later used for track prediction [53]. The metric is computed a...

  95. [110]

    ground truth

    video generation model by conditioning it on the motion tokens produced by the forward dynamics model. These motion tokens serve as a conditioning signal that guides video prediction based on the expected dynamics. To condition the video generation model, motion tokens generat...

  96. [2016]

    ISSN: 1063-6919

    doi:10.1109/CVPR.2016.90. ISSN: 1063-6919

  97. [2017]

    URLhttps://arxiv.org/abs/1706.04261

  98. [2020]

    URL https://proceedings.neurips.cc/paper_files/paper/2020/ file/1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf

  99. [2024]

    arXiv:2406.09246 [cs]

    URLhttp://arxiv.org/abs/2406.09246. arXiv:2406.09246 [cs]

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.