Pith. sign in
Pith Number

pith:G5ILXTEH

pith:2022:G5ILXTEHRFZAB7INVLY4NHPO3C
not attested not anchored not stored refs resolved

Video Diffusion Models

Alexey Gritsenko, David J. Fleet, Jonathan Ho, Mohammad Norouzi, Tim Salimans, William Chan

A diffusion model extended from images generates high-fidelity coherent videos using joint training and conditional sampling.

arxiv:2204.03458 v2 · 2022-04-07 · cs.CV · cs.AI · cs.LG

Add to your LaTeX paper
\usepackage{pith}
\pithnumber{G5ILXTEHRFZAB7INVLY4NHPO3C}

Prints a linked badge after your title and injects PDF metadata. Compiles on arXiv. Learn more · Embed verified badge

Record completeness

1 Bitcoin timestamp
2 Internet Archive
3 Author claim open · sign in to claim
4 Citations open
5 Replications open
Portable graph bundle live · download bundle · merged state
The bundle contains the canonical record plus signed events. A mirror can host it anywhere and recompute the same current state with the deterministic merge algorithm.

Claims

C1strongest claim

We present the first results on a large text-conditioned video generation task, as well as state-of-the-art results on established benchmarks for video prediction and unconditional video generation.

C2weakest assumption

That treating video as an extension of image diffusion (with joint training and the new conditional sampling) is sufficient to produce temporally coherent high-fidelity output without major additional architectural changes for motion modeling.

C3one line summary

A diffusion model for video generation extends image architectures with joint image-video training and improved conditional sampling, delivering first large-scale text-to-video results and state-of-the-art performance on video prediction and unconditional generation benchmarks.

References

65 extracted · 65 resolved · 5 Pith anchors

[1] https://www.tensorflow.org/ datasets 2022
[2] ViViT: A video vision transformer 2021
[3] Stochastic Variational Video Prediction 2017 · arXiv:1710.11252
[4] Fitvid: Overfitting in pixel-level video prediction.arXiv preprint arXiv:2106.13195 2021
[5] Is space-time attention all you need for video understanding? 2021

Formal links

2 machine-checked theorem links

Cited by

65 papers in Pith

Receipt and verification
First computed 2026-07-05T04:34:16.381834Z
Builder pith-number-builder-2026-05-17-v1
Signature Pith Ed25519 (pith-v1-2026-05) · public key
Schema pith-number/v1.0

Canonical hash

3750bbcc87897200fd0daaf1c69deed89e9b6124a5cd7a28a2a4b6ddd5dedcff

Aliases

arxiv: 2204.03458 · arxiv_version: 2204.03458v2 · doi: 10.48550/arxiv.2204.03458 · pith_short_12: G5ILXTEHRFZA · pith_short_16: G5ILXTEHRFZAB7IN · pith_short_8: G5ILXTEH
Agent API
Verify this Pith Number yourself
curl -sH 'Accept: application/ld+json' https://pith.science/pith/G5ILXTEHRFZAB7INVLY4NHPO3C \
  | jq -c '.canonical_record' \
  | python3 -c "import sys,json,hashlib; b=json.dumps(json.loads(sys.stdin.read()), sort_keys=True, separators=(',',':'), ensure_ascii=False).encode(); print(hashlib.sha256(b).hexdigest())"
# expect: 3750bbcc87897200fd0daaf1c69deed89e9b6124a5cd7a28a2a4b6ddd5dedcff
Canonical record JSON
{
  "metadata": {
    "abstract_canon_sha256": "5087b293a4582c56e830cacecfb1fbff56e5337a2d10670ddffcdcae3ccc649a",
    "cross_cats_sorted": [
      "cs.AI",
      "cs.LG"
    ],
    "license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
    "primary_cat": "cs.CV",
    "submitted_at": "2022-04-07T14:08:02Z",
    "title_canon_sha256": "298b61f8bcb493c054d01d107797de5bc9d7b8fc05e162952a1b302faddb38f4"
  },
  "schema_version": "1.0",
  "source": {
    "id": "2204.03458",
    "kind": "arxiv",
    "version": 2
  }
}