Pith. sign in
Pith Number

pith:RMT4HSRN

pith:2022:RMT4HSRNJNUJ2HPIDCIAY7G3LH
not attested not anchored not stored refs resolved

Make-A-Video: Text-to-Video Generation without Text-Video Data

Adam Polyak, Devi Parikh, Harry Yang, Jie An, Oran Gafni, Oron Ashual, Qiyuan Hu, Sonal Gupta, Songyang Zhang, Thomas Hayes, Uriel Singer, Xi Yin, Yaniv Taigman

A method turns text into videos by extending image generators with motion learned separately from unlabeled footage.

arxiv:2209.14792 v1 · 2022-09-29 · cs.CV · cs.AI · cs.LG

Add to your LaTeX paper
\usepackage{pith}
\pithnumber{RMT4HSRNJNUJ2HPIDCIAY7G3LH}

Prints a linked badge after your title and injects PDF metadata. Compiles on arXiv. Learn more · Embed verified badge

Record completeness

1 Bitcoin timestamp
2 Internet Archive
3 Author claim open · sign in to claim
4 Citations open
5 Replications open
Portable graph bundle live · download bundle · merged state
The bundle contains the canonical record plus signed events. A mirror can host it anywhere and recompute the same current state with the deterministic merge algorithm.

Claims

C1strongest claim

Make-A-Video sets the new state-of-the-art in text-to-video generation, as determined by both qualitative and quantitative measures.

C2weakest assumption

That the proposed spatial-temporal decomposition and pipeline can transfer motion dynamics learned from unsupervised video to text-conditioned generation without introducing visible artifacts or losing text faithfulness.

C3one line summary

Make-A-Video achieves state-of-the-art text-to-video generation by decomposing temporal U-Net and attention structures to add space-time modeling to text-to-image models, trained without any paired text-video data.

References

15 extracted · 15 resolved · 12 Pith anchors

[2] Language Models are Few-Shot Learners 2005 · arXiv:2005.14165
[3] arXiv preprint arXiv:2204.14217 , eprint =
[4] Make-a-Scene
[5] Score-cam: Score-weighted visual explanations for convolutional neural net- works 2020 · doi:10.1109/cvprw50498.2020.00193
[6] Denoising Diffusion Probabilistic Models 2006 · arXiv:2006.11239

Formal links

2 machine-checked theorem links

Cited by

132 papers in Pith

Receipt and verification
First computed 2026-07-05T05:02:05.789195Z
Builder pith-number-builder-2026-05-17-v1
Signature Pith Ed25519 (pith-v1-2026-05) · public key
Schema pith-number/v1.0

Canonical hash

8b27c3ca2d4b689d1de818900c7cdb59cd80b5b83a6ad45b0ed02c8147182790

Aliases

arxiv: 2209.14792 · arxiv_version: 2209.14792v1 · doi: 10.48550/arxiv.2209.14792 · pith_short_12: RMT4HSRNJNUJ · pith_short_16: RMT4HSRNJNUJ2HPI · pith_short_8: RMT4HSRN
Agent API
Verify this Pith Number yourself
curl -sH 'Accept: application/ld+json' https://pith.science/pith/RMT4HSRNJNUJ2HPIDCIAY7G3LH \
  | jq -c '.canonical_record' \
  | python3 -c "import sys,json,hashlib; b=json.dumps(json.loads(sys.stdin.read()), sort_keys=True, separators=(',',':'), ensure_ascii=False).encode(); print(hashlib.sha256(b).hexdigest())"
# expect: 8b27c3ca2d4b689d1de818900c7cdb59cd80b5b83a6ad45b0ed02c8147182790
Canonical record JSON
{
  "metadata": {
    "abstract_canon_sha256": "aba523ef24a47557c4cc45d2005e48691e4fb7259af54dcf216ca2581522599f",
    "cross_cats_sorted": [
      "cs.AI",
      "cs.LG"
    ],
    "license": "http://arxiv.org/licenses/nonexclusive-distrib/1.0/",
    "primary_cat": "cs.CV",
    "submitted_at": "2022-09-29T13:59:46Z",
    "title_canon_sha256": "08eba9c4a0fc0b51bf8cecc8724418772c8143eb298c2f0dbc5fb27811b88099"
  },
  "schema_version": "1.0",
  "source": {
    "id": "2209.14792",
    "kind": "arxiv",
    "version": 1
  }
}