Pith. sign in

Video pretraining (vpt): Learning to act by watching unlabeled online videos

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

citation-role summary

background 1

citation-polarity summary

fields

cs.CV 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Playing with Transformer at 30+ FPS via Next-Frame Diffusion

cs.CV · 2025-06-02 · conditional · novelty 6.0

Next-Frame Diffusion combines block-wise causal attention, consistency distillation, and action-based speculative sampling to generate action-conditioned Minecraft video at over 30 FPS on an A100 with a 310M parameter model.

citing papers explorer

Showing 1 of 1 citing paper.

  • Playing with Transformer at 30+ FPS via Next-Frame Diffusion cs.CV · 2025-06-02 · conditional · none · ref 4

    Next-Frame Diffusion combines block-wise causal attention, consistency distillation, and action-based speculative sampling to generate action-conditioned Minecraft video at over 30 FPS on an A100 with a 310M parameter model.