Pith. sign in

Paper Citation Record · LEDGER

Movie Gen: A Cast of Media Foundation Models

As of 10 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 100 inbound Pith citation observations for arXiv:2410.13720.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.13720 v2

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-11T14:16:18.521699Z

measured 188 of 188 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 100 of 257 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:46:08.177469Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

88 of 88 outbound references displayed

  • verified exact66
  • verified fuzzy4
  • unresolved0
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch16

External citation measurements

8
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation b4f40554-53e5-46c6-a8c9-86d765e9c2dc · outbound

This paper cites Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation.

Movie Gen: A Cast of Media Foundation Models Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:26.008358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:8f45204c9d473a01fa0c6cb765b8382bf9dad250207b75b0c41c68f490c0e0a1

Observation 758571fd-a831-4955-8a9d-c18f47ea0350 · outbound

This paper cites eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers.

Movie Gen: A Cast of Media Foundation Models eDiff-I: Text-to-Image Diffusion Models with an Ensemble of Expert Denoisers

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-15T01:44:22.991611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:8a7b65b1e8fc93be0c99ef64f2267b34e5d956edaf1fe5f1df0e9944bfda5289

Observation 56de7286-24de-4857-8c1d-0491cb7bb1b3 · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

Movie Gen: A Cast of Media Foundation Models Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:19.190665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:117ce6e7106ab63405f7b10ce9482ddde49900d6db0ad36299eb668d65527e71

Observation f67b0561-a37c-4020-8b0b-cee280a14226 · outbound

This paper cites A Note on the Inception Score.

Movie Gen: A Cast of Media Foundation Models A Note on the Inception Score

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:19.488521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:dfbfc1075765e7a6c995c4a646272a70f0b5d2c7c4f7d75145080c6edb678e88

Observation 10aaf17d-a409-483c-b356-a9847c8ba46b · outbound

This paper cites Meta Open Compute Project, Grand Teton AI platform.https://engineering.

Movie Gen: A Cast of Media Foundation Models Meta Open Compute Project, Grand Teton AI platform.https://engineering

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T14:16:26.315128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:45f07893f7d2b9067f075f59a33fe278f4e549de5a41897aa133d9b8e0c39c2c

Observation 10a9d913-aaf0-4b72-8189-aac96663ddfa · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Movie Gen: A Cast of Media Foundation Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:16:20.818910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:94b6e065777da0416acd491febd922b1f7565014ceb7cc142b877623c7d47ab0

Observation 10de9f58-95d8-495c-aeb6-78fac5139772 · outbound

This paper cites Language Models are Few-Shot Learners.

Movie Gen: A Cast of Media Foundation Models Language Models are Few-Shot Learners

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:16:21.075360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:5dcefad0d1f2509696e830b7ca2a649408fdd9d7ab95c796d9bde3fb28dc5c64

Observation 2d70a758-8848-4e50-ac5c-6a570fd5265a · outbound

This paper cites Still-Moving: Customized Video Generation without Customized Video Data.

Movie Gen: A Cast of Media Foundation Models Still-Moving: Customized Video Generation without Customized Video Data

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:21.204598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:7fc49428c64ece267f6ba3f4c4e7590cdd038dac24c96d16da293c988895b321

Observation c0b07ee2-4b25-4c54-b1b0-414e9ea2a2ff · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

Movie Gen: A Cast of Media Foundation Models VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:40:44.284627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:bdf1e980cb1e15d9141df0785ab6043e9c8ebde695b24a44c697fec13b041009

Observation acd00a39-f0ad-4f1d-b9e1-b19623ca0bf6 · outbound

This paper cites PhotoVerse: Tuning-Free Image Customization with Text-to-Image Diffusion Models.

Movie Gen: A Cast of Media Foundation Models PhotoVerse: Tuning-Free Image Customization with Text-to-Image Diffusion Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:21.576682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:0f6d360f214070f9948278d0a79f429cbaff388f8f41a86ed8d4a2021433fea7

Observation 4dcbda89-585f-4398-99ce-c15c70908c72 · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

Movie Gen: A Cast of Media Foundation Models Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:21.686371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:460fffe612834f674a571ae25695ff65eaeb2436446b3db14dc361da3724be3d

Observation e305e4cb-e195-4a41-8417-0657d1a9265e · outbound

This paper cites Vision Transformers Need Registers.

Movie Gen: A Cast of Media Foundation Models Vision Transformers Need Registers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T09:41:38.514155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:cec00a3602e99041fa23462e871d001880087eef5dc1b598e4ac280533cb6e08

Observation 125e7e4e-6b60-4b8f-bc3f-abe3ff5efc77 · outbound

This paper cites Diffusion Models Beat GANs on Image Synthesis.

Movie Gen: A Cast of Media Foundation Models Diffusion Models Beat GANs on Image Synthesis

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-13T11:16:28.664120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:3d1a8851cf2276bdd036bb2ffdd572103c7dfc1a3031d7a29466c73c21e41eb1

Observation 0a21340f-8060-4650-b419-127689400f4b · outbound

This paper cites The Llama 3 Herd of Models.

Movie Gen: A Cast of Media Foundation Models The Llama 3 Herd of Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:16:22.245197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:8bf64d03936270323d22897a92672d56953fe6bd00b2ba746279a47f99b19706

Observation 96de64f4-a267-465e-a626-22a6b65b4c74 · outbound

This paper cites High Fidelity Neural Audio Compression.

Movie Gen: A Cast of Media Foundation Models High Fidelity Neural Audio Compression

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:49:52.306785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:7baa02841cb477db7a66d0719701777616839581c8d8fe2034def282c6eae0da

Observation 46c02703-c910-4bf1-859a-ba8113296125 · outbound

This paper cites Understanding Back-Translation at Scale.

Movie Gen: A Cast of Media Foundation Models Understanding Back-Translation at Scale

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:22.608280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:9eb1d0b0b5e6e5e0fca0938aeea9751f22ca268a331497c07b97ad620784e2d6

Observation e1adce8f-d3f3-440e-88d8-a21f1d72211b · outbound

This paper cites Structure and Content-Guided Video Synthesis with Diffusion Models.

Movie Gen: A Cast of Media Foundation Models Structure and Content-Guided Video Synthesis with Diffusion Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:22.728880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:a18f10ad45b16abaf8f6bfc004ce4f0072c99fea01f418c7be25f9f3cb410c88

Observation 43201648-175b-40c1-840e-753a491ee7e6 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Movie Gen: A Cast of Media Foundation Models An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:08:55.677650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:6df5408aafa77c9847b2f27066c8a77320e3c023acdd30ee1d142289513104e9

Observation 611fa5cd-30ef-4f5c-8471-f57c7199ca61 · outbound

This paper cites TokenFlow: Consistent Diffusion Features for Consistent Video Editing.

Movie Gen: A Cast of Media Foundation Models TokenFlow: Consistent Diffusion Features for Consistent Video Editing

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:17:47.115702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:6d77eebbe8da344eb5460fd33d1a3b77e0b497d8f9c56f5c365e4335d4d0db21

Observation 707a0d2c-634c-4acf-a713-b365b0ec29d6 · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

Movie Gen: A Cast of Media Foundation Models Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:23.275961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:ca4e81513ec86d8aea6fd65007e3ef84a2e5c78bcc84e48070e395643197a5c1

Observation 481d2025-b1f3-4238-806e-8ffa6a9ef252 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Movie Gen: A Cast of Media Foundation Models AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:16:23.380792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:dde22620f6bf5fbcb00748be67654d934223128c4028f758788bf3428ca0a1e5

Observation fcfa539f-3a60-4c06-91f7-08487f1a35db · outbound

This paper cites Photorealistic Video Generation with Diffusion Models.

Movie Gen: A Cast of Media Foundation Models Photorealistic Video Generation with Diffusion Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:23.388206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:0cb4268034025c0482372c6e42df8b597da215d2312ae5aa07401002bb207355

Observation 5a949d7d-8578-42dc-8a45-39e2139e0c3b · outbound

This paper cites ID-Animator: Zero-Shot Identity-Preserving Human Video Generation.

Movie Gen: A Cast of Media Foundation Models ID-Animator: Zero-Shot Identity-Preserving Human Video Generation

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:23.486698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:25f8e45c214f618e8f38f19ee422a55c090d0c41acbe81b638b4ead51ce7ddb2

Observation 91629891-0dbe-46b6-b66b-b2a5d7f34509 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Movie Gen: A Cast of Media Foundation Models Imagen Video: High Definition Video Generation with Diffusion Models

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:16:23.497613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:5601b1c0d07369506670264f2ddcd044b5ab14189f2abb113b3752e50c7cb5f8

Observation 9a1e8cfd-c000-4d99-84b6-247470282895 · outbound

This paper cites DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation.

Movie Gen: A Cast of Media Foundation Models DirecT2V: Large Language Models are Frame-Level Directors for Zero-Shot Text-to-Video Generation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:23.576244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:c4c68474b7fa46d056c7b19340cfcba9e67041e6476489cd7ef41eae3403dddb

Observation 07805306-e726-4c76-b86f-a99a6bf6d7dd · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

Movie Gen: A Cast of Media Foundation Models CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:16:23.725103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:876c5c2f387385a04901e02f8f3325a0b2dc43f0ee256a4c0e41c67cf99d2255

Observation 6a113585-72d9-4523-8412-f1324f1b0c02 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Movie Gen: A Cast of Media Foundation Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:16:23.840348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:c701a37d3781706ec4f6d3dd3b22923b93c6319f43aaf12a2fb6d85b51f04b23

Observation eca944cf-56c8-4bf3-8dff-1ab0605dd869 · outbound

This paper cites Noise2Music: Text-conditioned Music Generation with Diffusion Models.

Movie Gen: A Cast of Media Foundation Models Noise2Music: Text-conditioned Music Generation with Diffusion Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:23.975684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:706cbf5f274df64ad070d1369d304a96936d079004b2dfd69e07da9755d8911f

Observation e940fd72-b93e-406a-867d-451dbb5e3c57 · outbound

This paper cites Video2Music: Suitable Music Generation from Videos using an Affective Multimodal Transformer model.

Movie Gen: A Cast of Media Foundation Models Video2Music: Suitable Music Generation from Videos using an Affective Multimodal Transformer model

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:24.096063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:927517954d2717b40a35919f54986ca4f1aa74c54597fcf00385bbfc99eb5080

Observation 86ab49e8-d686-4f02-83cb-ced1d7aae8f0 · outbound

This paper cites RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models.

Movie Gen: A Cast of Media Foundation Models RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:24.296364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:bb969452ed3377367f20781c94f41a6521665aefa07613adc6bfc120cd7cf8e1

Observation fc7a146b-07fd-453d-bbd4-02e2af7e4ee1 · outbound

This paper cites Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators.

Movie Gen: A Cast of Media Foundation Models Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video Generators

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:24.575795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:c24bd7846e28abb50a10cef3bf665dced7202899226ca9eb5712103994e56425

Observation 357b476f-a538-4eba-a9ca-98310347690b · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Movie Gen: A Cast of Media Foundation Models Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:24.646511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:be1deab421ed3d6ca78dcae9e1e62626711d0ce451e9d4e0523ea883b39918ee

Observation 2aa304d7-ad32-4dff-81c2-567c3d6a7c5b · outbound

This paper cites FIFO-Diffusion: Generating Infinite Videos from Text without Training.

Movie Gen: A Cast of Media Foundation Models FIFO-Diffusion: Generating Infinite Videos from Text without Training

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:24.703349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:7ab0cb0eeb973c898764a2c9e76ff4733a0954c96edf09eb8c16d309e5f94546

Observation 4881b9ae-c92b-430f-acc1-7c9177e5e105 · outbound

This paper cites Auto-Encoding Variational Bayes.

Movie Gen: A Cast of Media Foundation Models Auto-Encoding Variational Bayes

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:16:24.856379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:36465fba54d2b6fe65fc0d10391ee6ee9708c345cc8136aa8199039f00486ef8

Observation 8f8008db-29b7-4c2f-8ea4-107feed8b09a · outbound

This paper cites Imagine Flash: Accelerating Emu Diffusion Models with Backward Distillation.

Movie Gen: A Cast of Media Foundation Models Imagine Flash: Accelerating Emu Diffusion Models with Backward Distillation

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:24.984341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:154ecc420ada802a08983e1789bdab5e7abd4047c9a0650e491aaf3310be9d37

Observation 92e00695-367b-4952-9570-f4dc5192e3c9 · outbound

This paper cites VideoPoet: A Large Language Model for Zero-Shot Video Generation.

Movie Gen: A Cast of Media Foundation Models VideoPoet: A Large Language Model for Zero-Shot Video Generation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:51:05.830483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:c518050534abf68e60ecb2d5e01293d9a064154d7908c92667259c5e94614dc8

Observation 32b56a7a-2b0d-42f3-9330-44fc7b8f3be1 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

Movie Gen: A Cast of Media Foundation Models AudioGen: Textually Guided Audio Generation

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:24.998932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:67d570f1b18aade3550a1239776c28f02b801d92b838b03d93f7b1d1570b091c

Observation ba1ee809-b7f2-494e-8cff-6c3cbe16a9b2 · outbound

This paper cites High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching.

Movie Gen: A Cast of Media Foundation Models High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:25.009553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:73ab8e3e6cc793beee1dc20551d2a00c0badb1380cd220fe797e6843ac140b34

Observation f456c7f4-541e-47cd-856c-1b8ad3962f2c · outbound

This paper cites Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale.

Movie Gen: A Cast of Media Foundation Models Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:25.046954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:7e93605ebab77d85928a8a23535c83e53057adc55ecad1a2ce3c1f3d846af90a

Observation d1f8a6af-4077-4cfd-922b-e4af6a01f3a2 · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

Movie Gen: A Cast of Media Foundation Models BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:25.090823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:6494fa971467d0c299736edfc01b9f0af21062360a31612afa3092e7341b71e7

Observation 9f24a98e-73f5-4c9e-8e51-773911d0a315 · outbound

This paper cites VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation.

Movie Gen: A Cast of Media Foundation Models VideoGen: A Reference-Guided Latent Diffusion Approach for High Definition Text-to-Video Generation

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:25.098681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:785b342562cb6d7b1ba5c797609035d48383999330c09eea4b35729ce6d4d822

Observation 831561ab-e0b9-430b-ae2e-7424cf18adfc · outbound

This paper cites PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding.

Movie Gen: A Cast of Media Foundation Models PhotoMaker: Customizing Realistic Human Photos via Stacked ID Embedding

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:25.109885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:aa67e609629d0108a832c612229823c619552e5a08c0ca6ab6da35c0dae5bf47

Observation b2972eb6-45e4-43df-8fdd-a5e8694e5a67 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Movie Gen: A Cast of Media Foundation Models Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:28:28.418803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:23dfdd91c291a1d52a3be67611c916cde262ff5f2dfb12294600de69f7db868d

Observation 9bb7485d-25d2-4dd5-b681-ee0fd4fb6716 · outbound

This paper cites Dream Machine, 2024.https://lumalabs.ai/dream-machine.

Movie Gen: A Cast of Media Foundation Models Dream Machine, 2024.https://lumalabs.ai/dream-machine

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T14:16:26.755064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:e787b436805b9b8fedb10731d5a98a46d2fe79d6a1077e3c1df85d5b13978e55

Observation 6d5d256a-5199-48db-903f-0e6f59148992 · outbound

This paper cites VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation.

Movie Gen: A Cast of Media Foundation Models VideoFusion: Decomposed Diffusion Models for High-Quality Video Generation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:25.123993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:d3e143fabcdf2c0eca825d5f18c73ea4a59e6e3d22cf7c8a01d16e9aeefb2163

Observation 3ea8ea41-e9b2-4782-a974-32caf4efad47 · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

Movie Gen: A Cast of Media Foundation Models Latte: Latent Diffusion Transformer for Video Generation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:45:35.835056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:d15f3329709ba44a51abe4aa861017ff0e81ff44adfef24a5de4e2e06fd37c90

Observation 660d8a96-811c-40bf-88fb-5779fd49f8e0 · outbound

This paper cites Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization.

Movie Gen: A Cast of Media Foundation Models Tango 2: Aligning Diffusion-based Text-to-Audio Generations through Direct Preference Optimization

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:25.186220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:449393aea512fdf85f9f475bc8300de9141f5b9473d9252dc236081873e5596b

Observation ae437815-b30d-4f71-ab37-1ba08bbe4f93 · outbound

This paper cites FoleyGen: Visually-Guided Audio Generation.

Movie Gen: A Cast of Media Foundation Models FoleyGen: Visually-Guided Audio Generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:25.365557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:35dd7babd4044c9af55d86cf9abfb58fa9bc71c347402124e7f56f5e75b60d0a

Observation 0ac9f63f-b997-4a5f-86d4-481b2492225b · outbound

This paper cites Midjourney, 2024.https://www.midjourney.com/.

Movie Gen: A Cast of Media Foundation Models Midjourney, 2024.https://www.midjourney.com/

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T14:16:26.776652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:24439c57dfdc17ca21beae41c0e7c385f9833fc45c9982a95f0d96c82341120f

Observation b9a4cb6f-2ddc-4641-bdf0-8484336b2e6a · outbound

This paper cites Dall-E 3, 2024.https://openai.com/index/dall-e-3/.

Movie Gen: A Cast of Media Foundation Models Dall-E 3, 2024.https://openai.com/index/dall-e-3/

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-11T14:16:26.342590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:5a27314e258d3c2e78fb9bf217d8db5ab570630cab234f94f126dc260c323656

Observation 5aff580c-4b2c-4911-ba14-a07eb89cf218 · outbound

This paper cites MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation.

Movie Gen: A Cast of Media Foundation Models MoA: Mixture-of-Attention for Subject-Context Disentanglement in Personalized Image Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:25.567978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:896f364ec7791d62f1e6e1b097d544f0d4f0d24f79f9ebd03b88035e48eefe8f

Observation 8e086225-2a11-4d16-b618-041cdd2f1083 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Movie Gen: A Cast of Media Foundation Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:16:25.652695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:9548803dbeca5f9c9748e665f3403ba24a0979ec9f577a6d9607bffea79c4f33

Observation 887b98e6-a3de-4a9f-b974-282d74b91c5b · outbound

This paper cites InstructVid2Vid: Controllable Video Editing with Natural Language Instructions.

Movie Gen: A Cast of Media Foundation Models InstructVid2Vid: Controllable Video Editing with Natural Language Instructions

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:25.765970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:4f8a769d5ab61aaa4f1bf6bc0b8be1ec94f232486088b5bea8d45e4df07e8180

Observation 7d6d1bbb-bfe4-4228-8b6d-3fa8cf4ca357 · outbound

This paper cites FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling.

Movie Gen: A Cast of Media Foundation Models FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:25.800863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:77445085887d65e5e4b667af3742e7f48613570bee68c3c5c80b975bfdd045b5

Observation 53186047-57d1-4956-868e-db57d2d2864a · outbound

This paper cites ZeRO: Memory Optimizations Toward Training Trillion Parameter Models.

Movie Gen: A Cast of Media Foundation Models ZeRO: Memory Optimizations Toward Training Trillion Parameter Models

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:24:35.924862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:143ce49fd47ecdab577fd61b6106fe3697049d2ada9e5e15674858f2466879fd

Observation 1a7118b6-1eb4-4633-84a3-1b4629c6c3b5 · outbound

This paper cites SelfEval: Leveraging the discriminative nature of generative models for evaluation.

Movie Gen: A Cast of Media Foundation Models SelfEval: Leveraging the discriminative nature of generative models for evaluation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:25.866446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:57c102a70a240edb536cabc5b0ba625754cb18560ffc6211f13f7d9648762341

Observation bb25cb3d-836f-45aa-84b9-1edcdb2a1352 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Movie Gen: A Cast of Media Foundation Models Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:16:25.872212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:00aef0132d70432c8100f56cf111c0f4f0d2883296b4236b93bc20f67e7b01a1

Observation 41a104e2-9a64-4a5a-acb8-61a944f11321 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Movie Gen: A Cast of Media Foundation Models SAM 2: Segment Anything in Images and Videos

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:16:25.876967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:2a977f11b4a907c4587fe99379e7826ca47f0e2fbc1dfc4089719531ff6abc44

Observation 6f09a7bd-06a9-4200-b46f-e61178dc17e7 · outbound

This paper cites ZeRO-Offload: Democratizing Billion-Scale Model Training.

Movie Gen: A Cast of Media Foundation Models ZeRO-Offload: Democratizing Billion-Scale Model Training

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:25.884336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:3a9d5ca6fd603aaa60276fc651fbf4e89f610e79ad957d92bdfaa355aa219687

Observation 993aad31-2db0-44d6-a98e-2b6f38318d86 · outbound

This paper cites HyperDreamBooth: HyperNetworks for Fast Personalization of Text-to-Image Models.

Movie Gen: A Cast of Media Foundation Models HyperDreamBooth: HyperNetworks for Fast Personalization of Text-to-Image Models

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:25.895979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:326288b777156050548da17ca3de7cb1d87f0d1cc49be5b97cfb8c79138e7473

Observation e166fa77-6b7b-4e61-8217-80499a309916 · outbound

This paper cites GLU Variants Improve Transformer.

Movie Gen: A Cast of Media Foundation Models GLU Variants Improve Transformer

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:16:25.901409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:7d81d0803b140cb39f7e44e6e3332c680b72e3c48dadc550629230f9c075fd7f

Observation 0c322811-7749-48e4-bc4b-4936f82b6af3 · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

Movie Gen: A Cast of Media Foundation Models NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:25.914341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:8df4f16441b9ac86f7188780ac1df4f84930cfa58126423d81c312e63617cee1

Observation e2a7c3e8-d3c0-4ab9-8695-9d710bf77841 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Movie Gen: A Cast of Media Foundation Models Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:16:25.965837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:23c28a7f448fc2b0c269eb0d8b402c449921bb668fa75ab36439a59d427fe453

Observation a071d6f9-bf51-46d3-b3b8-51387df54c6c · outbound

This paper cites Denoising Diffusion Implicit Models.

Movie Gen: A Cast of Media Foundation Models Denoising Diffusion Implicit Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:16:25.970806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:bcf0fc149c68b6873649d8a59464f7a9d57e3742641c569f6a19f510fe380850

Observation f1062967-4a22-4856-b582-4dab59e1dc67 · outbound

This paper cites UL2: Unifying Language Learning Paradigms.

Movie Gen: A Cast of Media Foundation Models UL2: Unifying Language Learning Paradigms

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:25.979094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:9f11ff0d683a585c5fccee4a6edc9e0f8eac163f19a99730a1eae289251d4b6f

Observation f645beed-5c44-42e3-9c5a-88de7fa1b484 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Movie Gen: A Cast of Media Foundation Models Gemini: A Family of Highly Capable Multimodal Models

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:16:18.885225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:c2fd2a3e099ca5b09a4d636717bea03cf3af77736c31a0f66839419a3fb7611c

Observation 9f54eb01-8fd8-4a42-9a8b-610b770f4a19 · outbound

This paper cites VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling.

Movie Gen: A Cast of Media Foundation Models VidMuse: A Simple Video-to-Music Generation Framework with Long-Short-Term Modeling

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:26.050224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:a940a92603a0f70965b32e3106cf2947c8ba1981cf838f74a4ba77f0f462fba1

Observation 79fdecf5-0fdb-4d88-8b3c-335bfad65e33 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Movie Gen: A Cast of Media Foundation Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:16:26.056021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:de6ad6cf8d85cb7a014006a74275d571929ae0900c808872b613f98f7e94bc25

Observation c1ec0f3c-064d-4904-b008-082d18969987 · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

Movie Gen: A Cast of Media Foundation Models Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:26.065148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:380643a6250b14e00a9198d44f2508efc7667eeced2917bf8ddb9cbbf7bfa903

Observation 35c5ac29-625f-4074-ba40-ac750fb6c7c7 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Movie Gen: A Cast of Media Foundation Models ModelScope Text-to-Video Technical Report

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:47:29.701736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:81ace0754817aed149b5558fe47f195c2918f7379ebb2e429dd94dd2023bb360

Observation bb53d8dc-8d8b-4238-93d8-991e0a4f701a · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

Movie Gen: A Cast of Media Foundation Models LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:26.088346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:54929943f3747fa5ee267cc77b0f5b8614d62b6d740368ac1c6937f471f30305

Observation 2bb6ef7b-b685-4a05-b161-1c8d3ac8d042 · outbound

This paper cites Fairy: Fast Parallelized Instruction-Guided Video-to-Video Synthesis.

Movie Gen: A Cast of Media Foundation Models Fairy: Fast Parallelized Instruction-Guided Video-to-Video Synthesis

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:26.095848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:632d656ecc2d648ac0c2a499860338d7527b755a4fc5eb1744e89a49286a7e75

Observation 03869817-10c4-4591-a028-89fdaf83f69b · outbound

This paper cites CVPR 2023 Text Guided Video Editing Competition.

Movie Gen: A Cast of Media Foundation Models CVPR 2023 Text Guided Video Editing Competition

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:26.103933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:7963aa5dfed17817031243740d2af85ddc375b3016ec4b1df038a37f7745367c

Observation cbed8755-5703-4767-8392-ac2162c8b1b2 · outbound

This paper cites DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors.

Movie Gen: A Cast of Media Foundation Models DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:26.115356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:b8963d866fb1cf84812d8defb56dd1ed4781c6743fa2513058a29c29cd748312

Observation 6a07cdce-6986-4bcd-8f44-a161a423fb90 · outbound

This paper cites Demystifying CLIP Data.

Movie Gen: A Cast of Media Foundation Models Demystifying CLIP Data

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:20:20.637207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:d3ae4fecded4a8625a8687b928029a76bf6beb53c344b93de51424ec67023104

Observation 07b7ac75-c128-441c-a497-5f555665cb84 · outbound

This paper cites Video-to-Audio Generation with Hidden Alignment.

Movie Gen: A Cast of Media Foundation Models Video-to-Audio Generation with Hidden Alignment

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:26.132308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:8c1e54ebec0dbe498efeb08439ab8ace00a2d70e9e973108e3e90aa8896286d7

Observation 8cbc9df8-350b-4c85-8651-22f58ff425ba · outbound

This paper cites VideoGPT: Video Generation using VQ-VAE and Transformers.

Movie Gen: A Cast of Media Foundation Models VideoGPT: Video Generation using VQ-VAE and Transformers

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:24:33.857072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:a96301196f4a013a43c3fdcfc577e24fabc39d64bb5a2499c5a6931d4d77c1c8

Observation 145f84c5-5b18-4668-9a5b-338d7de8ca5d · outbound

This paper cites Motion-Conditioned Image Animation for Video Editing.

Movie Gen: A Cast of Media Foundation Models Motion-Conditioned Image Animation for Video Editing

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:26.175885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:aeb8f04c783008e00985e50da2d6db03502dc9c6f7fc5006f60f9f76c7da76a6

Observation 59064447-c7da-474e-b8cd-b9705acf7a61 · outbound

This paper cites Space-Time Diffusion Features for Zero-Shot Text-Driven Motion Transfer.

Movie Gen: A Cast of Media Foundation Models Space-Time Diffusion Features for Zero-Shot Text-Driven Motion Transfer

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:26.240376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:9480d2efa61b0c4f8376e34f4152bcf19c5895198a7d3e30d74acc6059f53f1d

Observation 51de4d56-1e0e-4dfd-8222-31d2b3940a51 · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Movie Gen: A Cast of Media Foundation Models IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:16:26.257948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:1cfd99339f09c32752a62cfc89f55c5189ab926e9a42c71432bbf9dfbb5a6c19

Observation 7e63f620-5c23-4b22-8d79-8cb28d10a96d · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

Movie Gen: A Cast of Media Foundation Models Vector-quantized Image Modeling with Improved VQGAN

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:40:37.571629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:5d00151045fed79c333434575c3507e77eaedac9530b35d3086c3ec0936b5432

Observation 86105db8-b705-4c88-b1ba-c5b8893b03e3 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Movie Gen: A Cast of Media Foundation Models Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 82

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:06:45.090605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:ccb455552a3d68065d4b85d1074b17928a6e361ba04961d4746142d5c218caeb

Observation b7633aa8-744c-4b5d-a698-2e0585f569d6 · outbound

This paper cites I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models.

Movie Gen: A Cast of Media Foundation Models I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-17T23:51:24.156433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:e9ca70b42c5117c880521f9b3a2c61e6e345cedfee7ea839cb7373bbc5240c73

Observation 068449c1-770e-4e67-ab01-2487027d9d06 · outbound

This paper cites Real-Time Video Generation with Pyramid Attention Broadcast.

Movie Gen: A Cast of Media Foundation Models Real-Time Video Generation with Pyramid Attention Broadcast

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:20.266812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:74a7a60bbf009e3d989e92ca6aaa81c9e634015dd04c569f3ccce969b6d82410

Observation 8e6105ec-4847-4d40-814d-98941fcfc30b · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

Movie Gen: A Cast of Media Foundation Models PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:15:20.153254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:f5683101da2785f4eb77a3f9dfd5fe35be25493e464e653be1780d8192aa7e0d

Observation 25ffd022-69b0-4bda-a8d1-7ba83cf4be25 · outbound

This paper cites Quantized GAN for Complex Music Generation from Dance Videos.

Movie Gen: A Cast of Media Foundation Models Quantized GAN for Complex Music Generation from Dance Videos

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:20.657018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:8ec55b073cdceb445edc00d06f275597dc2d0a0300b54b8d84a37f386dbe1b5b

Observation 8aeff77e-6c0e-4244-93af-762991a51bc6 · outbound

This paper cites Hence, we include the dates on which we collected the videos from website.

Movie Gen: A Cast of Media Foundation Models Hence, we include the dates on which we collected the videos from website

Reference 87

Resolution
malformed identifier
raw_fallback, observed 2026-05-11T14:16:26.365134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:40a351512fb80c78c76cd200ddf04142d121e165f3ed7430ca7ad29cb5eac1b6

Observation ce07aaf1-b868-4697-8855-0a28d4439aa8 · outbound

This paper cites Overall” audio quality for the Audio Quality pairwise task and the “Correctness.

Movie Gen: A Cast of Media Foundation Models Overall” audio quality for the Audio Quality pairwise task and the “Correctness

Reference 88

Resolution
malformed identifier
raw_fallback, observed 2026-05-11T14:16:26.524980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:289f1f488d1358a90d82d163458d243a96fb820f2184aa4befbabb397c9d9190

Pith citing papers

Observation c9813c43-2e1e-4bac-b39d-cbe66d9ae298 · inbound

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control cites this paper.

$\pi_0$: A Vision-Language-Action Flow Model for General Robot Control Movie Gen: A Cast of Media Foundation Models

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:26.778393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T12:38:24.425784Z digest=sha256:cf4ea797ea5fca088a3ca5b8352941c4325ca0ea9bae2520dbf5570fb0d2229e

Observation 709b48a6-1d7b-44ba-a935-5fc38c489046 · inbound

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models cites this paper.

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models Movie Gen: A Cast of Media Foundation Models

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T02:48:45.117826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T02:48:44.900467Z digest=sha256:12e94cedb4c5706b5b56fc159ffd88c59049c51780a7bee02b8fc040ae094133

Observation aa6ea669-b771-475c-823c-6bf85f8b8403 · inbound

KFC-W: Generating 3D-Consistent Videos from Unposed Internet Photos cites this paper.

KFC-W: Generating 3D-Consistent Videos from Unposed Internet Photos Movie Gen: A Cast of Media Foundation Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-23T16:48:12.425124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T16:47:09.717923Z digest=sha256:57fbc1ab4f4f239f50ee83094f946e9de2a10043585020bd0911114eb2583239

Observation 87101d13-8322-4bb1-9546-2c05cb45c359 · inbound

Open-Sora Plan: Open-Source Large Video Generation Model cites this paper.

Open-Sora Plan: Open-Source Large Video Generation Model Movie Gen: A Cast of Media Foundation Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-23T08:42:45.177050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T08:38:27.946746Z digest=sha256:614f8c50cc84a9d6471bdc962a0dab32a00efeb4fad63edd82b2542f9a23443a

Observation 92fe95a1-d7dd-4183-b56e-83c7f1a21dff · inbound

HunyuanVideo: A Systematic Framework For Large Video Generative Models cites this paper.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Movie Gen: A Cast of Media Foundation Models

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.402396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:91a80164e6de306aa0f0219a8e46f16f00b512e2618a2bd6fd941aad2fd2d794

Observation e07d4a21-f378-42de-ae5a-1257db2e842c · inbound

Flow Matching Guide and Code cites this paper.

Flow Matching Guide and Code Movie Gen: A Cast of Media Foundation Models

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T10:28:14.134797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T10:28:14.014706Z digest=sha256:a297c4e6f6c6b9359456acdb01114791aa7833367e27b3615b5ee7cf1bfd4590

Observation e1fa6baa-56ca-4eb8-9cf2-b71fa160e747 · inbound

DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization cites this paper.

DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization Movie Gen: A Cast of Media Foundation Models

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:02:41.967947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T06:57:50.897865Z digest=sha256:68ec29a974890f3d516eebd6820b91211303f5b8768604853741fd13e9685b0e

Observation 826e3003-77b5-4b7d-8f3c-146c73461f66 · inbound

LTX-Video: Realtime Video Latent Diffusion cites this paper.

LTX-Video: Realtime Video Latent Diffusion Movie Gen: A Cast of Media Foundation Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:26.778393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T10:36:12.058970Z digest=sha256:adda972adee9c4ce5dace7b68e88dcd702d796724d2b3bd1f98c6062c9f7e6d8

Observation 27991783-7e99-4e63-845a-5a77c901b66c · inbound

Cosmos World Foundation Model Platform for Physical AI cites this paper.

Cosmos World Foundation Model Platform for Physical AI Movie Gen: A Cast of Media Foundation Models

Reference 156

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:26.778393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T23:38:44.933410Z digest=sha256:ac9dd3529c7d772f4af8b1c0be9e28d177b1172e3bf45ac46ae17c3f39b9d042

Observation 4f8e4f5b-41ad-478f-98ff-91296eddb570 · inbound

Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps cites this paper.

Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps Movie Gen: A Cast of Media Foundation Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:45:17.654401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T11:45:17.473970Z digest=sha256:07aaa6a70d36f21c72cb2becb4abc358a7601180d77e9ddde870252fdbde523e

Observation 095024fa-01b2-423f-8d21-7c87e87a35f7 · inbound

Improving Video Generation with Human Feedback cites this paper.

Improving Video Generation with Human Feedback Movie Gen: A Cast of Media Foundation Models

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-13T15:30:02.772291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T15:30:02.578430Z digest=sha256:88dc3ea9028b8a889bbb29c54116f207f326ff2feaa84ae9767ca5fb0cb8b157

Observation c4876983-387a-4576-a046-bea434a1ba8d · inbound

LLMs can see and hear without any training cites this paper.

LLMs can see and hear without any training Movie Gen: A Cast of Media Foundation Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T00:46:08.177469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:46:08.177469Z digest=sha256:4e4ec8442274fd075005449002c6cf80e943109bda31e952b447797f16621c4c

Observation 6de2d389-5659-4dca-a761-eb794a9eaf20 · inbound

Diffusion Autoencoders are Scalable Image Tokenizers cites this paper.

Diffusion Autoencoders are Scalable Image Tokenizers Movie Gen: A Cast of Media Foundation Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T22:56:36.153918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:56:36.153918Z digest=sha256:de947abbcdcf08daea0313c75c10dad70d4294adb8ee8a9406b5dd6aebd1e90f

Observation 0950bc12-d5bc-4a4f-978d-90274c9b11bc · inbound

Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation cites this paper.

Video Latent Flow Matching: Optimal Polynomial Projections for Video Interpolation and Extrapolation Movie Gen: A Cast of Media Foundation Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T18:51:12.602618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:51:12.602618Z digest=sha256:c302381cf9f36f026f07908347807b2950fc5446468ccc5b175e531173742f12

Observation 86449761-abe7-4943-b48c-a2a395cd0e7c · inbound

OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models cites this paper.

OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models Movie Gen: A Cast of Media Foundation Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T16:47:13.008763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T16:47:13.008763Z digest=sha256:1bb53fde419b33cb3a2c865aa06cbb693d3e4fa2281d66d6d6103efce398f883

Observation c00aae12-9b48-4c8a-aae7-4be53d98b9a6 · inbound

MotionCanvas: Cinematic Shot Design with Controllable Image-to-Video Generation cites this paper.

MotionCanvas: Cinematic Shot Design with Controllable Image-to-Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T22:56:55.208472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:56:55.208472Z digest=sha256:e401f18de958f439af780ad94aaf753d1bbf5cf03230888b55f06b756eaeb1c1

Observation dcee6775-7798-4992-915b-6afec6cc2381 · inbound

Fast Video Generation with Sliding Tile Attention cites this paper.

Fast Video Generation with Sliding Tile Attention Movie Gen: A Cast of Media Foundation Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T22:37:42.108044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:37:42.108044Z digest=sha256:888b61470cdb8d0a286e45148cf484423e0e2727c2dfabed3c1c1c909369eb61

Observation 3f5005d0-01bc-49b2-b5a9-bbff2e9cf3bf · inbound

Goku: Flow Based Video Generative Foundation Models cites this paper.

Goku: Flow Based Video Generative Foundation Models Movie Gen: A Cast of Media Foundation Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T21:07:32.379704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T21:07:32.379704Z digest=sha256:c1a553dba1dbdca7cb452c1f07819d0313b01cb8dee1611e3f2fa3939bc776cf

Observation e7ecb01a-2f82-4394-9737-e2f9f20c090c · inbound

Latent Swap Joint Diffusion for 2D Long-Form Latent Generation cites this paper.

Latent Swap Joint Diffusion for 2D Long-Form Latent Generation Movie Gen: A Cast of Media Foundation Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T20:12:29.536356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T20:12:29.536356Z digest=sha256:09810f685fca16ee8d286c93593fb292f4e23e0c8eae4c58a299d34cdf4dc3a5

Observation 61a28cc3-124b-4600-be0c-295385ca37f0 · inbound

VFX Creator: Animated Visual Effect Generation with Controllable Diffusion Transformer cites this paper.

VFX Creator: Animated Visual Effect Generation with Controllable Diffusion Transformer Movie Gen: A Cast of Media Foundation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T17:13:02.804779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:13:02.804779Z digest=sha256:ba38d77997b42daebf1f7f0339e7b9bf265ae0dc50e066b884ea03c95453f309

Observation 40317859-3c3c-4ec6-b491-d7a849bf6dbc · inbound

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT cites this paper.

Lumina-Video: Efficient and Flexible Video Generation with Multi-scale Next-DiT Movie Gen: A Cast of Media Foundation Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T14:26:46.802121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:26:46.802121Z digest=sha256:0befac7070adddd39c573d92ed339fb82c97b255ea08d0c73fde0945c265f2f7

Observation d7afcf0f-186b-41af-af0a-f56541d67858 · inbound

Flow Matching for Collaborative Filtering cites this paper.

Flow Matching for Collaborative Filtering Movie Gen: A Cast of Media Foundation Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T13:15:15.235839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:15:15.235839Z digest=sha256:1d564755a99216d01d61f3f835a60b1cd0d8f63c8ac9800f0b581188bc09528c

Observation b30cd212-45ff-4832-aeff-ec2c75b2361f · inbound

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference cites this paper.

PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference Movie Gen: A Cast of Media Foundation Models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-08T12:19:43.212984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:19:43.212984Z digest=sha256:9fc4e9dafd4f3eb2add0707b3da780a2f345dfd8466d498fe9e3018233453208

Observation d0c61e21-094b-4a5f-a7ea-462699649d33 · inbound

Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts cites this paper.

Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts Movie Gen: A Cast of Media Foundation Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T11:21:20.559714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:21:20.559714Z digest=sha256:de43a94d78f84c8df9b2a5292d1fcd311556a4547f163ab9629723b73b477ecd

Observation 9856ffde-2b26-4825-8889-93ed73b16072 · inbound

Pre-Trained Video Generative Models as World Simulators cites this paper.

Pre-Trained Video Generative Models as World Simulators Movie Gen: A Cast of Media Foundation Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T15:14:22.347722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:14:22.347722Z digest=sha256:7fed853e7eefb223cfcd3de8008446233dfc8041ee72c86334b3d304b442977a

Observation 2bae2d54-4924-47d2-b9dc-e509da7c40d9 · inbound

FloVD: Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video Synthesis cites this paper.

FloVD: Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video Synthesis Movie Gen: A Cast of Media Foundation Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T05:55:09.529053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:55:09.529053Z digest=sha256:03ec619604ab49ef65cb9e868faf88968f6e7d7f89c94f86a0bcce0fd89bf625

Observation 61cba769-5fc5-465b-898d-433d921fd919 · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Movie Gen: A Cast of Media Foundation Models

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T08:02:23.924392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:3a9dd68908f49a500b118d8e976486ca43d12e82efdf0309b42859b095f93c2b

Observation 15d0c5c0-dad7-4a5f-b62d-15b00b28d72e · inbound

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions cites this paper.

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions Movie Gen: A Cast of Media Foundation Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-23T00:32:18.245053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T00:30:55.729900Z digest=sha256:b23bc4454bb8553d63360caafb438d5a2ee89ca6092cfc1c03ad81d4ef55cbd3

Observation 62412fd0-b0cf-4025-8f59-2e028fd65dbd · inbound

Wan: Open and Advanced Large-Scale Video Generative Models cites this paper.

Wan: Open and Advanced Large-Scale Video Generative Models Movie Gen: A Cast of Media Foundation Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:07:14.423703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T23:05:32.595632Z digest=sha256:65cdc617be9e33e1d841eca28a47b58dce5f4786cd06ca554b97b431c960edb3

Observation ccc8dbbc-c5fa-4857-9c86-a53272c1d1af · inbound

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness cites this paper.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Movie Gen: A Cast of Media Foundation Models

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.055974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:2b5797710660e4bdc3b324ddff0a2b0b29c15513ab32a42f3a80c5e8cf491c9d

Observation db551e23-efa0-4c74-8d01-ae633181e9f3 · inbound

Transfer between Modalities with MetaQueries cites this paper.

Transfer between Modalities with MetaQueries Movie Gen: A Cast of Media Foundation Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:49:23.248001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T22:49:23.074271Z digest=sha256:8f64ac373d882d670b392b2ff6ae7eba770f1a4f20ac17a185f7983ecc91eae5

Observation 73375d52-4a4f-450e-aef1-81bf8a4ef7b0 · inbound

MAGI-1: Autoregressive Video Generation at Scale cites this paper.

MAGI-1: Autoregressive Video Generation at Scale Movie Gen: A Cast of Media Foundation Models

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T20:31:15.841638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T20:31:15.700943Z digest=sha256:00615d510ef4d369c1e894fdfcdd3c232d6a65d32c1ccc1385221622aab5f3ea

Observation 402c1531-b64c-435c-ad8c-4374b7d67849 · inbound

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos cites this paper.

TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos Movie Gen: A Cast of Media Foundation Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:03.532534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:03:03.532534Z digest=sha256:77a2c5b2da9838e62297dcfc3a3d586625b9716094f2315566bf88ef5a645c06

Observation 52a3b87a-82de-4dc9-84ff-ed35d2ef92f3 · inbound

Long-Context State-Space Video World Models cites this paper.

Long-Context State-Space Video World Models Movie Gen: A Cast of Media Foundation Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:18.814503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:18.814503Z digest=sha256:dc7cb9a1c59c334ca0c891f848ee7a307ea0e0f40f86b9027b01bf3f92b67231

Observation b56701b8-6fae-490a-bb47-531247ec1f90 · inbound

Test-Time Training Done Right cites this paper.

Test-Time Training Done Right Movie Gen: A Cast of Media Foundation Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:25:45.242488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T11:25:45.153563Z digest=sha256:0f21973c1597140b8b9d864e66555da0c1e402241b1cc4cdfdc9b0d68b8fb157

Observation 73135b7e-14d0-4374-a5a9-d0eb02e715f7 · inbound

Accelerated Sampling from Masked Diffusion Models via Entropy Bounded Unmasking cites this paper.

Accelerated Sampling from Masked Diffusion Models via Entropy Bounded Unmasking Movie Gen: A Cast of Media Foundation Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:41.561669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:41.561669Z digest=sha256:b39ebbd5c8781e7e3aa0bbed5bc0c20473605ea124040d50625703f28c62f4f7

Observation 16317c3c-7c29-483a-ab92-48d8f5d94aa1 · inbound

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers cites this paper.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Movie Gen: A Cast of Media Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:07.866515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:07.866515Z digest=sha256:cda382790541dc2c00ee3414f0d155482a96cb4b593a5c0bbab8664b196636fc

Observation 57d1a527-e4bd-44ff-82ca-b4125f600380 · inbound

Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution cites this paper.

Self-supervised ControlNet with Spatio-Temporal Mamba for Real-world Video Super-resolution Movie Gen: A Cast of Media Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:19.077238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:19.077238Z digest=sha256:0a8a72aff42ac3631906f956c6603a9e2cf506bdcb301eaef814e6ad605625a8

Observation 01627c78-aa77-4ca4-8f29-4de28ad38766 · inbound

Humanoid World Models: Open World Foundation Models for Humanoid Robotics cites this paper.

Humanoid World Models: Open World Foundation Models for Humanoid Robotics Movie Gen: A Cast of Media Foundation Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:56.774332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:56.774332Z digest=sha256:10ebf75130a706250f223668fc8f516d0a1734d7143dda0c0b6f2db5b4d3610a

Observation 36275562-bb18-48f9-bcf3-fcbefeaf1c6f · inbound

Playing with Transformer at 30+ FPS via Next-Frame Diffusion cites this paper.

Playing with Transformer at 30+ FPS via Next-Frame Diffusion Movie Gen: A Cast of Media Foundation Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:44.157194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:44.157194Z digest=sha256:bf50eedbc44994bafa4bdd5068bd7027dc184ed4ff7933de6b13bb6a17d6b0ac

Observation b5b36acf-998c-4b4d-a376-8a27a8d6b087 · inbound

RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions cites this paper.

RefEdit: A Benchmark and Method for Improving Instruction-based Image Editing Model on Referring Expressions Movie Gen: A Cast of Media Foundation Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:07:56.716959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:07:56.716959Z digest=sha256:2e35ce77ade5636e94566235e9db117e67d341a8e4aa043c4c636a4089795cfc

Observation e4eac15e-91b5-450b-abcc-7c7a20256a95 · inbound

Follow-Your-Creation: Empowering 4D Creation through Video Inpainting cites this paper.

Follow-Your-Creation: Empowering 4D Creation through Video Inpainting Movie Gen: A Cast of Media Foundation Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:23.757368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:23.757368Z digest=sha256:e61dd1498bd1a8668d6449cb7629d65fc243d54c0f4c4d997059a17a6dcf6832

Observation ab8e86b3-9e41-49df-ad74-80fdcc9da14f · inbound

ContentV: Efficient Training of Video Generation Models with Limited Compute cites this paper.

ContentV: Efficient Training of Video Generation Models with Limited Compute Movie Gen: A Cast of Media Foundation Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:33.500496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:33.500496Z digest=sha256:548023bb71022175a5660abf5bcb61ae4ca067daf9de66faa676e17167bbfc05

Observation a3fe0e59-0e33-44af-aae3-f965ae67eb2d · inbound

From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models cites this paper.

From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models Movie Gen: A Cast of Media Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:30.972332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:30.972332Z digest=sha256:bcf25d130fcb0cbe60c85030d78b1379e6eedc1623c66ee0c84bd219ce6ce796

Observation fe2140d3-0410-4633-b890-35f32edab8e9 · inbound

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement cites this paper.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Movie Gen: A Cast of Media Foundation Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.997934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.997934Z digest=sha256:f41a3efdda60c94d2c868206b2872f2b840cd129b9dfe5edeac5be5fdee9198f

Observation b0563564-1aa4-491d-a10b-6305dffedad5 · inbound

EgoM2P: Egocentric Multimodal Multitask Pretraining cites this paper.

EgoM2P: Egocentric Multimodal Multitask Pretraining Movie Gen: A Cast of Media Foundation Models

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:42.552979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:42.552979Z digest=sha256:6e9a0a24f888b24fc9bf5f8f549a58f43f20e5171d8414f907c4af74b78e530a

Observation 9ef452e8-00bb-4656-8cfa-3fb41a957153 · inbound

Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion cites this paper.

Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion Movie Gen: A Cast of Media Foundation Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:26.778393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:36:53.029590Z digest=sha256:a32ae2113884adb1a6a85cdf5ecdf7abc617ed34ff72891e0454636189e6cb93

Observation 15c9c35d-d7e8-45c9-918e-6c3ee2b72c86 · inbound

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset cites this paper.

Phantom-Data : Towards a General Subject-Consistent Video Generation Dataset Movie Gen: A Cast of Media Foundation Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:10.412598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:10.412598Z digest=sha256:a76ecd52355a08d4ac6aa6b46a6a2e29821247ba7668a416a6b2ba900221e5fb

Observation 186907ed-1502-45d5-82a0-79d9c0153b15 · inbound

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance cites this paper.

Step-by-Step Video-to-Audio Synthesis via Negative Audio Guidance Movie Gen: A Cast of Media Foundation Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:45:07.206934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:45:07.206934Z digest=sha256:f8617c9d49f5338a24420aadd1622996f1fdc2f5b417339cd6bdbbc0404fc373

Observation 38041d61-9b3a-4bfc-848d-e573dd9856d0 · inbound

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation cites this paper.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Movie Gen: A Cast of Media Foundation Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:08.717848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:08.717848Z digest=sha256:14bd29f07c57740f347bf24595b1387fbcb5461c1d3f62bbd361e2bacaf60d2c

Observation bf9b0f74-61d1-49b7-9944-2a664f657eec · inbound

Transition Matching: Scalable and Flexible Generative Modeling cites this paper.

Transition Matching: Scalable and Flexible Generative Modeling Movie Gen: A Cast of Media Foundation Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:46:28.390443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:46:28.390443Z digest=sha256:068165735d030c12a6dd3c6aac3e86aece043aa23a2ff984f97168d456334866

Observation d342d1ca-3627-48bb-950b-6b4d559ff64e · inbound

FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion cites this paper.

FreeLong++: Training-Free Long Video Generation via Multi-band SpectralFusion Movie Gen: A Cast of Media Foundation Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:28:16.552459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:28:16.552459Z digest=sha256:d2dd6e5c013fb8e1d6a2f381eb584d9a3d2eeb9bffd804c301350d11915b411f

Observation 93484021-023b-4bf8-b5db-e9eec2f6ad75 · inbound

Populate-A-Scene: Affordance-Aware Human Video Generation cites this paper.

Populate-A-Scene: Affordance-Aware Human Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.731808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.731808Z digest=sha256:a75d3aff15fbe1b74eb8e56fab182390448046b7bbcc84f4b26bbe314052198c

Observation 71ae8b05-80e9-422f-b504-85dbe3c8382d · inbound

LACONIC: A 3D Layout Adapter for Controllable Image Creation cites this paper.

LACONIC: A 3D Layout Adapter for Controllable Image Creation Movie Gen: A Cast of Media Foundation Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:21:38.696715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:21:38.696715Z digest=sha256:99812780837b753e18b4bc23ba11844b09d4398c720d30ec89ca187cd6c5f5b9

Observation d290f7f3-77c5-4d3d-adbe-9d1b35a8480e · inbound

GraphBrep: Learning B-Rep in Graph Structure for Efficient CAD Generation cites this paper.

GraphBrep: Learning B-Rep in Graph Structure for Efficient CAD Generation Movie Gen: A Cast of Media Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:19.944723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:19.944723Z digest=sha256:e44ec82a3af588c9b99a0f8415bc7b9f30d06e6fabf831f1f0c07597c628418c

Observation 03b093e9-5210-4af3-b679-40633855ac20 · inbound

Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation cites this paper.

Tora2: Motion and Appearance Customized Diffusion Transformer for Multi-Entity Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:19:32.766455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:19:32.766455Z digest=sha256:c757e614f1a4dde5b5d65de7086410b16d09a41601120a505f2fe159336f32fa

Observation 4c1f6efc-ac82-4ad9-b8ea-a7da7c53f5e8 · inbound

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality cites this paper.

A Survey on Long-Video Storytelling Generation: Architectures, Consistency, and Cinematic Quality Movie Gen: A Cast of Media Foundation Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:51:27.992227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:51:27.992227Z digest=sha256:e5a36f73e259b93367d619a05f01c8de46c9fcc901024900b307c940cf5ef9d0

Observation 21790bed-c2d4-4338-b583-b1c99e62288b · inbound

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling cites this paper.

Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling Movie Gen: A Cast of Media Foundation Models

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:17:06.701163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-19T05:13:28.767788Z digest=sha256:5548b7d0897090e7cb63d2c6aaa8ec3cc4d7f858ca2731725785c4564427f79b

Observation a526ccb9-1282-4617-a76c-fb6a4f82f176 · inbound

LoViC: Efficient Long Video Generation with Context Compression cites this paper.

LoViC: Efficient Long Video Generation with Context Compression Movie Gen: A Cast of Media Foundation Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:27.790649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:39:27.790649Z digest=sha256:8fbb90836692fbf627ec555d4c746abb0936fef8663a9dda4cb1892ace7c9d87

Observation efb3fd45-47bb-42d5-8413-fc22bdb2d81d · inbound

Frozen Forecasting: A Unified Evaluation cites this paper.

Frozen Forecasting: A Unified Evaluation Movie Gen: A Cast of Media Foundation Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:42:57.318365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T03:42:54.620069Z digest=sha256:769f36894b442c432d3d7189d1000d5e73362d50fe742a79c4cdbb6d8ac48101

Observation 28d675ee-7aec-49f7-9bc2-044a2206edf3 · inbound

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers cites this paper.

Look, Focus, Act: Efficient and Robust Robot Learning via Human Gaze and Foveated Vision Transformers Movie Gen: A Cast of Media Foundation Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:26:24.834469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:26:24.834469Z digest=sha256:4aa6486d379195bbc511bdb5fe6a90095c546e64922501a4c623e6794a32e9df

Observation bb315865-5db7-4819-bc0d-baafc365e7c2 · inbound

Back to the Features: DINO as a Foundation for Video World Models cites this paper.

Back to the Features: DINO as a Foundation for Video World Models Movie Gen: A Cast of Media Foundation Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:24:20.142280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:24:20.142280Z digest=sha256:c2dc1e091047b3e45aeefefb1446d57afdeb4d00689993e2340c054b0c1ff052

Observation 73329125-5b0e-4757-a629-9c0b07f1bb92 · inbound

Flow Matching Policy Gradients cites this paper.

Flow Matching Policy Gradients Movie Gen: A Cast of Media Foundation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:07:09.488679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:07:09.488679Z digest=sha256:8e97aad669906be375aab9bc14e596b99ec9e703913af17d7d065d8a6df1cfb9

Observation cb12e912-727d-402a-bdf0-3da2ef9a2a5e · inbound

Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm cites this paper.

Scaling Up Audio-Synchronized Visual Animation: An Efficient Training Paradigm Movie Gen: A Cast of Media Foundation Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T01:05:32.217703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T01:05:32.217703Z digest=sha256:3ff4da354e4f22f9fbcfc00316bcb0102759b4de6f5c93d1200feb9a9f1e4984

Observation 086c64f7-92df-4dcc-8745-cc9d75bdf4b7 · inbound

Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers cites this paper.

Consistent and Controllable Image Animation with Motion Linear Diffusion Transformers Movie Gen: A Cast of Media Foundation Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T22:21:05.448579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:21:05.448579Z digest=sha256:84b9050b31c17272ae0d3b08cf2cd964d2217a0d1ba0d99eec687f3e8e7a8e84

Observation ca3414a6-ae03-4966-9e0a-b5e1e3385ccd · inbound

X-UniMotion: Animating Human Images with Expressive, Unified and Identity-Agnostic Motion Latents cites this paper.

X-UniMotion: Animating Human Images with Expressive, Unified and Identity-Agnostic Motion Latents Movie Gen: A Cast of Media Foundation Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T21:09:29.035522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:09:29.035522Z digest=sha256:f9cce9e6751cfb3e0a30ee9b713c6c65bf208e53056d0f3b354d5e9940a5717b

Observation 48bb1691-4b6f-40c7-be7f-ab8867cea4cc · inbound

Waver: Wave Your Way to Lifelike Video Generation cites this paper.

Waver: Wave Your Way to Lifelike Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:17.695941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:17.695941Z digest=sha256:40654dd31d057bc884c386d09f1b75be09e9c5d010ab5682bef12fb4520cf2b4

Observation 4b766f44-b95b-4d07-87df-f4552eaff303 · inbound

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper cites this paper.

Efficient Video-to-Audio Generation via Multiple Foundation Models Mapper Movie Gen: A Cast of Media Foundation Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T05:50:17.658718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:50:17.658718Z digest=sha256:5d7bfadcafbafab0af64f35ce80ffe34e29f7c8116435e4b1b86ba7bdc1d3e52

Observation 6f395c99-31ce-4824-a34a-07a418acced7 · inbound

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts cites this paper.

UniVerse-1: Unified Audio-Video Generation via Stitching of Experts Movie Gen: A Cast of Media Foundation Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T00:05:47.639479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:05:47.639479Z digest=sha256:e4476a8db0f0bed7e16d76d3f4eb8c540d8d778d86c2e647bd1d6f462e36ceec

Observation 931b262e-7efa-478f-8884-de43bff657ef · inbound

UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward cites this paper.

UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward Movie Gen: A Cast of Media Foundation Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T23:06:09.088428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:06:09.088428Z digest=sha256:bd97881e3ae6490245bba95f2d1e0b3d554cb0f1670861984f3471d61babe5fb

Observation dcca47f5-00a3-4da8-9eea-bfe8a32261de · inbound

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning cites this paper.

HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning Movie Gen: A Cast of Media Foundation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T20:35:30.324506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:35:30.324506Z digest=sha256:7f3029f3281927ef7e666a10ba5b32550895be264622a0b0d9749c5cec55ca67

Observation ed61d7da-adb1-4e8f-bfea-3a36ff56f536 · inbound

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders cites this paper.

Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders Movie Gen: A Cast of Media Foundation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T18:58:10.972403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:58:10.972403Z digest=sha256:d87c93324a8e08a23456fc9aba5823a85b74ed559210c8e02d67e6b23a41ad63

Observation fbe8a28d-ca9e-40c0-87b8-409e3c34bea1 · inbound

EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning cites this paper.

EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning Movie Gen: A Cast of Media Foundation Models

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T13:46:25.811975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T13:46:05.547669Z digest=sha256:62f6192804ebf1244c7e60ca91d2c56f2df62ba3661a5bdfd457a5faeaaf69c2

Observation 9657f906-d921-4762-a3f6-0f673c839d23 · inbound

Rolling Forcing: Autoregressive Long Video Diffusion in Real Time cites this paper.

Rolling Forcing: Autoregressive Long Video Diffusion in Real Time Movie Gen: A Cast of Media Foundation Models

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:15:29.186788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T11:15:29.102090Z digest=sha256:78612a06faf4bc4bebbc0b98954213adb056770f6e3ec04f8054ad465d575ecd

Observation 00e3431f-70f8-48dd-9f92-8e16ca0104a7 · inbound

Data-to-Energy Stochastic Dynamics cites this paper.

Data-to-Energy Stochastic Dynamics Movie Gen: A Cast of Media Foundation Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T13:40:54.549271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:40:54.549271Z digest=sha256:892f0e357c3119dc9bcdb031628753507510120dd05fb52dd173a1f883266cf3

Observation cb1b4a5c-1ab0-4c6e-a790-dbde2f3369f5 · inbound

Self-Forcing++: Towards Minute-Scale High-Quality Video Generation cites this paper.

Self-Forcing++: Towards Minute-Scale High-Quality Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:39:54.200182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T22:39:53.995700Z digest=sha256:368892078dd9f572189d6f18963dd8413a20284b471b5d457a75d6a062866de2

Observation 6c47020f-13a8-4617-885d-61bb9e295630 · inbound

Scaling Sequence-to-Sequence Generative Neural Rendering cites this paper.

Scaling Sequence-to-Sequence Generative Neural Rendering Movie Gen: A Cast of Media Foundation Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:11:13.931272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T10:09:27.285710Z digest=sha256:79fa47ac2a56036cb1b93fda460cc78a4d678299e998f103d143856a727848f1

Observation 6b362679-350b-45dc-a69e-b4c80fd4f4b8 · inbound

SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization cites this paper.

SSDD: Single-Step Diffusion Decoder for Efficient Image Tokenization Movie Gen: A Cast of Media Foundation Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T11:25:34.677649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:25:34.677649Z digest=sha256:9b33465de39c9c9548c1366ef1a5a9c5184f9c4428939654559c44fd774d58a7

Observation 6f57bb97-5184-4f30-b402-53c3a92da22a · inbound

UniVideo: Unified Understanding, Generation, and Editing for Videos cites this paper.

UniVideo: Unified Understanding, Generation, and Editing for Videos Movie Gen: A Cast of Media Foundation Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.898580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.898580Z digest=sha256:b7a944de3e36f19936511d51d619aace53a2eb24bd8b8c45d1b520121a2c92d7

Observation eeb0646c-5326-4610-a13f-de4d977799cf · inbound

Exploring Cross-Modal Flows for Few-Shot Learning cites this paper.

Exploring Cross-Modal Flows for Few-Shot Learning Movie Gen: A Cast of Media Foundation Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-18T06:20:58.902679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T06:16:12.608513Z digest=sha256:4b344e1e3cffd8e69895afe99cbdef7518f29149bf8ea61c2dd9d6bd4417b920

Observation ac627e75-7345-4a2f-bc18-0c57f3975541 · inbound

RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling cites this paper.

RAPO++: Cross-Stage Prompt Optimization for Text-to-Video Generation via Data Alignment and Test-Time Scaling Movie Gen: A Cast of Media Foundation Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:15:54.306835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T05:13:42.934115Z digest=sha256:4c51063e122faa4fb714e9b615625adad122ab70938bfb4d139642d69388fb6d

Observation 7f3725a6-18a3-4364-9a46-c4ef55f2f7be · inbound

Epipolar Geometry Improves Video Generation Models cites this paper.

Epipolar Geometry Improves Video Generation Models Movie Gen: A Cast of Media Foundation Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T08:20:09.440233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:20:09.440233Z digest=sha256:b8ae4c060bf69b93050573f258e8619ccddfc1fe0d1b859b705b5bef62957bb3

Observation 74b9026e-75e5-420d-80d1-c74b354abed5 · inbound

World Simulation with Video Foundation Models for Physical AI cites this paper.

World Simulation with Video Foundation Models for Physical AI Movie Gen: A Cast of Media Foundation Models

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-12T23:01:13.725884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T23:01:13.546110Z digest=sha256:b2ad5b3f72b35007603aafe0c86220151307869125938e9a0d10162df5d6d274

Observation 5048904e-4a33-4ec9-a37f-2118fcd68ef3 · inbound

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation cites this paper.

Kandinsky 5.0: A Family of Foundation Models for Image and Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T21:34:10.546872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:34:10.546872Z digest=sha256:55d4d6192d90ed62321488f45fb2f96274cdda79a103ed783c930b7556084366

Observation 0b910c79-e744-42dc-8d28-7e6a43603114 · inbound

Less is More: Data-Efficient Adaptation for Controllable Text-to-Video Generation cites this paper.

Less is More: Data-Efficient Adaptation for Controllable Text-to-Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-17T19:55:09.952289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T19:54:35.190865Z digest=sha256:02105f6cd6de2db2739c8e2254c0e5f8a5682e6ef52f5b2c4da49574a412d282

Observation 809a7fc8-2bec-47d9-a65a-1324ce235a84 · inbound

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer cites this paper.

One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer Movie Gen: A Cast of Media Foundation Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-21T18:25:28.333573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T18:25:22.486891Z digest=sha256:4d74d1463fe78c8b49966918e6dacbc6173002cdf628f6663b38b976b9668618

Observation 3d666953-9a2e-456b-bd4e-4adcb8db919d · inbound

Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation cites this paper.

Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation Movie Gen: A Cast of Media Foundation Models

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-16T18:17:55.002890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T18:17:54.943863Z digest=sha256:d08e5b07a20fc7c21b9e78e3e2be72ee9566713977122ff05537f6f1b73b1393

Observation 3ba89b9a-fdca-4522-8621-8962fb5d3926 · inbound

Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability cites this paper.

Delving into Latent Spectral Biasing of Video VAEs for Superior Diffusability Movie Gen: A Cast of Media Foundation Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T18:25:55.997414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:25:55.997414Z digest=sha256:e26385eb62ed1325ce2e2b17d5745594fd3eda9137976d9c52de81a68853d6cd

Observation 223ff78f-d065-4329-b28c-80834dad8de0 · inbound

VABench: A Comprehensive Benchmark for Audio-Video Generation cites this paper.

VABench: A Comprehensive Benchmark for Audio-Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:08:43.884622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T00:03:45.576961Z digest=sha256:e048b800efe360e4869a0e636b3f16c35bfebdf09eaf76c82a68c65695a76596

Observation 25ac8a62-8686-42e7-943f-88d069e31bff · inbound

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits cites this paper.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Movie Gen: A Cast of Media Foundation Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:37.377778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:37.377778Z digest=sha256:028910ce9f3c27f2f9b0f22471eb43895f4b6d8b6d95aae2a2a4d50b3ade8c9c

Observation 2073f5f2-f292-4d1f-9e60-76c8e47f5ae2 · inbound

The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection cites this paper.

The devil is in the details: Enhancing Video Virtual Try-On via Keyframe-Driven Details Injection Movie Gen: A Cast of Media Foundation Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:08:22.938728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T20:06:30.109866Z digest=sha256:a9816dc8111f81a2e9b455f74f5f5ea3d12df7659d7d02e62d37ef8a4b57bfac

Observation 08d3b635-47c6-4998-8303-4f6116ae4a36 · inbound

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World cites this paper.

DriveLaW:Unifying Planning and Video Generation in a Latent Driving World Movie Gen: A Cast of Media Foundation Models

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:38:21.138670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T19:34:39.518649Z digest=sha256:4b98701c476c2ac39b5301eda05d232456097c07e784ab90187f0c3e74fff46f

Observation e49c83c9-f2fb-4ec1-af9c-020468693ceb · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:48:21.893652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T19:43:37.604351Z digest=sha256:9baa8093bb0d128548bc90dd77eda71360b532a94b3cc17ffecd06d226e424ef

Observation 3be5f3b9-5644-4ea5-8e7e-b634521d05fb · inbound

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation cites this paper.

PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-21T16:14:15.245980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T16:10:31.015783Z digest=sha256:3560329103af7f304ca95871ff1c8ad98c988a74ce0bf6269b249bbf00163fd7

Observation 88698636-bd5f-4316-96ed-c4f6daf41ef8 · inbound

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation cites this paper.

PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T13:24:47.170535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:24:47.170535Z digest=sha256:ab47c5efe9faf80dc9287fd5f7c51ee318712f1f0124420340f26bfdfa291561

Observation 202b0d33-757f-4ab3-8960-6922fcf1e493 · inbound

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation cites this paper.

EchoFoley: Event-Centric Hierarchical Control for Video Grounded Creative Sound Generation Movie Gen: A Cast of Media Foundation Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T13:21:10.238117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:21:10.238117Z digest=sha256:b864db0cb2de4ced95168e850471b82f4cf6953c3ed923877668ec13e5786004

Observation e961e207-426f-4de2-a923-b31f657c04a6 · inbound

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation cites this paper.

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-21T17:32:01.790891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T17:32:01.642256Z digest=sha256:ce946fff4ec19090b7386cb4c093075cc8317567f80afc72b79aa90f417b2137

Observation 1c2cd5da-b115-4633-8bad-a115ae8147f3 · inbound

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation cites this paper.

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-22T11:51:30.012175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T11:48:00.633421Z digest=sha256:d3b75f4c4a738d0f6db4156e1dbf69d88c13fe6c6bb23e609ce8f871318f5a8d

Observation 726170ab-8127-4293-8bea-4e3eaf54f040 · inbound

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation cites this paper.

Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T05:30:27.140435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:30:27.140435Z digest=sha256:beb7d6374c2fa5b186e2c9acf1ec25dd505c1882c554adbc4357e79c27aefa42

Observation fc9bbcd3-c546-4a0c-a0db-7a83059e42dc · inbound

Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization cites this paper.

Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization Movie Gen: A Cast of Media Foundation Models

Reference 11

Resolution
malformed identifier
local_arxiv, observed 2026-05-16T08:00:44.666626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T07:58:24.456859Z digest=sha256:339c84e45f0349d2a79a56e56fce427c76b3270d0010a1649ebc5808505cbfe6