Pith. sign in

Paper Citation Record · LEDGER

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

As of 7 August 2026, this Paper Citation Record lists 100 of 116 outbound references and 100 inbound Pith citation observations for arXiv:2311.15127.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.15127 v1

Coverage vector

measured 100 of 116 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T22:58:51.792047Z

measured 200 of 200 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 100 of 766 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T19:40:21.149980Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

100 of 116 outbound references displayed

  • verified exact32
  • verified fuzzy60
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

67
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 5da6655c-ed78-40ef-bafd-3a4f839b1168 · outbound

This paper cites Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:52.156087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:19995f1a13e48efa618124322985d598dd9f7abf5e537ab1fcd9fc3d9c011543

Observation 144bcf58-117e-4840-8e26-ea1f0d350bab · outbound

This paper cites Renderdiffusion: Image diffusion for 3d reconstruction, in- painting and generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Renderdiffusion: Image diffusion for 3d reconstruction, in- painting and generation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.351279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:543d32ec5a69cf49094cc10ab7b29a8da2640f776a9bcd702e07e8fa34bc1be9

Observation 1369ee33-e07b-4808-a922-19ef6454c28c · outbound

This paper cites A general language assistant as a laboratory for alignment.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets A general language assistant as a laboratory for alignment

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.354002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:f23b2e818e472c7c113727e6e97286adeffed48534586c5f1be34b87dc5db28d

Observation 93cc8df8-95fc-4c9f-be29-a1da5d5b8662 · outbound

This paper cites Campbell, and Sergey Levine.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Campbell, and Sergey Levine

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.356405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:6d738d3efb79b97525f2c1e708149a6f9c989596956f7dd5d04d1f69f80f65ed

Observation 933ef10c-4406-4898-814a-9df65483ee6b · outbound

This paper cites Character region awareness for text de- tection.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Character region awareness for text de- tection

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.360851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:d7c38af38633529e634b00c1fd6a4801c4af8f9a35d5357d171dee6e381a771a

Observation db7d9742-6bab-457f-821b-97f97f659745 · outbound

This paper cites Training a helpful and harm- less assistant with reinforcement learning from human feed- back.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Training a helpful and harm- less assistant with reinforcement learning from human feed- back

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.367993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:b5605384284e857aebd5b2ad4cf2d5bcb9bd10c38dff1e1dbe1e567dc26756d3

Observation 22aa8867-91ab-44a3-a949-85e04be5d2bb · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.372008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:f34b608fc69d429d4ed54ffebc058009e6258e49bcc4cd37834022ad183def2b

Observation 50404684-73cb-43f0-aa23-7cf3fbb80143 · outbound

This paper cites ipoke: Poking a still image for con- trolled stochastic video synthesis.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets ipoke: Poking a still image for con- trolled stochastic video synthesis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.374897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:3758610f98d068e284e89abf860307d22de2521d6e6ab4344fc0ed7af1953dd7

Observation 9911c2c8-7ca3-40c2-a592-bbe9e7d0d173 · outbound

This paper cites Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:52.140275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:883dcca1b39321df940aadc02b0ef1a6a764c3b3c778f4a71c6822ecac9cabf2

Observation 560ec6c5-ca90-4822-bdfa-d8858b739e0e · outbound

This paper cites Generating long videos of dynamic scenes.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Generating long videos of dynamic scenes

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.380162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:a935eb0a38049ff510ec64367f55eec1102db31b2fd43902cd1e7a69c1f837b0

Observation 84289ebc-a5e0-4436-98b1-657f7e6aac5a · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Quo vadis, action recognition? a new model and the kinetics dataset

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.383092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:264fb4f3a8613cdf261db140d48e690ad97e344ff8dd7ff2055d8d4476b36bb8

Observation 3df0282d-bca6-472d-93e9-a5c7ff611852 · outbound

This paper cites Im- proved conditional vrnns for video prediction.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Im- proved conditional vrnns for video prediction

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.389104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:f7a800e2e1cc81772a69896ce0a42185d129d90bcb534bef77bc0a5863ce041b

Observation 3a9c73d5-ff6c-42ac-99ee-0e9da43b4573 · outbound

This paper cites Emu: Enhancing image generation models using photogenic nee- dles in a haystack.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Emu: Enhancing image generation models using photogenic nee- dles in a haystack

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.395124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:bd068c9f0e36a70ed5915dd4c2f8cba77a628ad5f4b3e9ba984c9ceeebcfe80f

Observation d0f5b2f2-63ce-40e6-943e-a29cb2343dad · outbound

This paper cites Objaverse-XL: A Universe of 10M+ 3D Objects.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Objaverse-XL: A Universe of 10M+ 3D Objects

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:02:11.861391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:4d43538f81db47398fe983b1a83976a1b3f569313a724901f9a56de38af6b356

Observation bd7479d9-f02e-4934-9fdd-5e7b8e090b5d · outbound

This paper cites Objaverse: A universe of annotated 3d objects.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Objaverse: A universe of annotated 3d objects

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.398112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:aac00132073deb628708c448b4dbcef218543b4b368608af928888fc0333e729

Observation 5a104dc8-f20a-4b87-be90-f5352dab01a2 · outbound

This paper cites Nerdi: Single-view nerf synthesis with language-guided diffusion as general image priors.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Nerdi: Single-view nerf synthesis with language-guided diffusion as general image priors

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.401071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:477c3a436d9dee5dba928b27620d43baac0a1156c4964c8c5a1b0b426e303dd1

Observation 320fb33a-6144-4d89-9f0b-560514fdc576 · outbound

This paper cites Stochastic video genera- tion with a learned prior.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Stochastic video genera- tion with a learned prior

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.404300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:71e8ec5a8bc5c9c8e96be4d5d151c58755d67b5a9b945e1c941165db4bd7ccb8

Observation 6b2076e6-a43b-4069-99c3-f9aa92e5dbba · outbound

This paper cites Diffusion Models Beat GANs on Image Synthesis.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Diffusion Models Beat GANs on Image Synthesis

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-13T11:16:28.664120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:b41a65b89285c939c04225e6d231c60f7d192894079fc98bb57877a3643f714d

Observation 020a0ee9-fbcb-44c6-af09-1542b02f5622 · outbound

This paper cites Derpanis, and Bj¨orn Om- mer.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Derpanis, and Bj¨orn Om- mer

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.406918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:f71f7b2b5ff5521357f05fc8a33beb23f36b71bbfda6f27fbe04e09abae271e2

Observation 3f825f0d-913a-4c73-8220-2720c572daeb · outbound

This paper cites Google scanned objects: A high-quality dataset of 3d scanned household items.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Google scanned objects: A high-quality dataset of 3d scanned household items

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.410446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:9e6bcfa7fca7de917f95e1c13ea24accd22b649c41625881e8b4c61d38b824dd

Observation 8f93d1e6-83f0-4f7e-943b-9d7a023e0fdd · outbound

This paper cites an unresolved cited work.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-10T22:58:52.413748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:f22f641e2a764d1871dbfc30a53f6571447c8a7c0fe931a26431ded52fb2d0e3

Observation d992ae0f-c056-43e8-8e4b-a13887ea7fe6 · outbound

This paper cites Taming Transformers for High-Resolution Image Synthesis.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Taming Transformers for High-Resolution Image Synthesis

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:58:52.084829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:82c14eb711db00de100dbb16681d3ad13e105abcfa886c266ca373c4d492e531

Observation 9f6ebd11-36e7-4089-9f9e-849e4be37272 · outbound

This paper cites Structure and content-guided video synthesis with diffusion models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Structure and content-guided video synthesis with diffusion models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.417956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:8ce29072f9b8b6efbc3e4d3b03f6a499ddea9ec5f6150c79d9c4a4e6babd6a5a

Observation 47afb22e-77d8-4081-a419-162fe74e5bef · outbound

This paper cites Two-frame motion estimation based on polynomial expansion.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Two-frame motion estimation based on polynomial expansion

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.423853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:77f6e4301ec6f05ba93da4b08c29eda77584886bc860123c9ea0c032f8d14f09

Observation 5c183e3c-648f-488b-8aff-7287d2f187b1 · outbound

This paper cites Stylevideogan: A temporal generative model using a pretrained stylegan.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Stylevideogan: A temporal generative model using a pretrained stylegan

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.430349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:2aeec1d1179cd7510b209463b37ce5c5769c2b485bda9636a3e258cb615db94c

Observation d13f1608-b9be-4757-a8c6-c7bb497ae565 · outbound

This paper cites Stochastic latent residual video prediction.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Stochastic latent residual video prediction

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.433370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:06ea66525a73a631cee9b68e7ad0c9e67fec133ac98f21a8e1ecbf8a7852f6e7

Observation fcf8153a-93af-4fa2-927b-05d260211088 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:52.107162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:c1fa5f6262fb0f9de9e23a99d201c997b68abf89c69eb2f72be1fbaaa60e5f13

Observation 690678e9-dadd-49a1-b826-a3233091cbef · outbound

This paper cites Long video generation with time-agnostic vqgan and time- sensitive transformer.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Long video generation with time-agnostic vqgan and time- sensitive transformer

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.440342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:51ae0ddbf7fb7f31ab35eda3f846a413d6ec2411a99016c851c45c5503df34f4

Observation 9a39a536-4f21-4df2-ab0a-6b7b301a3307 · outbound

This paper cites Preserve your own cor- relation: A noise prior for video diffusion models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Preserve your own cor- relation: A noise prior for video diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.443748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:da08d43ea96d52c195889ea1929717787e06e993f1c50342be2a1ca55013d443

Observation 243d2d62-1813-4d4d-a44c-b8a5abda79b6 · outbound

This paper cites Generative adversarial nets.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Generative adversarial nets

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.446463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:75971204944da768a50940f834fe49bcaa0c9ce20fa5adb9445580c97dc2e058

Observation a5b19aa2-1efc-46ca-bfcf-feb54b55f7e5 · outbound

This paper cites Reuse and Diffuse: Iterative Denoising for Text-to-Video Generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Reuse and Diffuse: Iterative Denoising for Text-to-Video Generation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:52.031672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:cd7aa28dc3f440a60ec381e90f9ab6bd1fc5ac190304de0632a9e10ec6030144

Observation e1db8544-3368-4414-b1c3-2e9bbd7852bb · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:52.081067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:4c30bf28107be0f3300b8e2163cf847551b8106aba0bcdcc6168b98cb4ac8618

Observation 2d998d1c-bee5-4b8c-af61-db9795cfdeb9 · outbound

This paper cites Rv-gan: Recurrent gan for unconditional video generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Rv-gan: Recurrent gan for unconditional video generation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.449183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:f6156259ba931c0c6e3b5a9e6d268ba02960968eb561d20d075400d80b68f412

Observation 91f40faa-f813-4d60-a401-8a1b820a41b4 · outbound

This paper cites Diffusion with offset noise.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Diffusion with offset noise

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.453606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:ed53d56bb7de5f4ea80f2880d4cfdfa15b8063abb14a40f9a986438360d17e26

Observation 7a87548d-6f96-4e28-a5cd-86b861b1dd53 · outbound

This paper cites Latent video diffusion models for high- fidelity long video generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Latent video diffusion models for high- fidelity long video generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.456584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:87c5484b15bc2ef8476496c2fb1ebdd159145b8ef6c5d1f6dd1b3ac1d6ac129c

Observation 47f51d8f-25a7-488a-ad22-bbb4152a5a59 · outbound

This paper cites Classifier-free diffusion guidance.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Classifier-free diffusion guidance

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.459221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:d1430e8ede045811d543cde730cec711062a6b5e19256e26caabcdb6adc86573

Observation 2552e170-3aef-480c-b3d9-88cc076c19d3 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Classifier-Free Diffusion Guidance

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:51.924296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:4c0fbca1a3484473598190b68eb476a32b91f3de781ce9e3d05a891b68e530a4

Observation 93c36309-9772-43af-a905-4895897a2b09 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Denoising dif- fusion probabilistic models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.461828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:a91d30e67da4a89f37f4e9bc361368677c4a73be320db45913c7fa05b3fc9118

Observation 063774ac-f08c-4560-bad1-c08ba661c7d8 · outbound

This paper cites Cascaded Diffusion Models for High Fidelity Image Generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Cascaded Diffusion Models for High Fidelity Image Generation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:51.972482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:9c91a06501486acfdaf7dbd7b6a899d31a720a01fbdd7bd77072c5f2d1d2ef29

Observation bac83294-60f9-4168-9ed8-008d896ab7b0 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Imagen Video: High Definition Video Generation with Diffusion Models

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:31:08.340948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:958b0e0a8dc47ee8891d4f21807e7c2c383de97823105b8609fd705577d8c6d5

Observation e481fdc1-9476-4793-8abc-787900d4fca1 · outbound

This paper cites Video Diffusion Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Video Diffusion Models

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T14:38:28.190773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:7ea908fec37278753974f3257451d989798e7e2487ae44bf067a0383a88a56ac

Observation bfa7556b-0994-4e58-9826-50e9a7be80cd · outbound

This paper cites Cogvideo: Large-scale pretraining for text-to- video generation via transformers.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Cogvideo: Large-scale pretraining for text-to- video generation via transformers

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.465451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:a9277f62435a5f794a059d68846b318b0bb4bf5d1e3128c1766f6bf517a1da32

Observation a2d7b100-b1a6-4787-b0c2-88e567d362a9 · outbound

This paper cites Simple diffusion: End-to-end diffusion for high resolution images.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Simple diffusion: End-to-end diffusion for high resolution images

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:52.051355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:7ef60c67c73630d5198f6b142351168b5290d2b5e49964960fa9b60b7aad2425

Observation 80ea7579-f424-4acd-99ee-ffcb0f3bfe4a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets LoRA: Low-Rank Adaptation of Large Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:52.070238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:3af1cf756c1a34afc3a8e580cd9619e752c4ab7861184c47c05bd61aaeab7e26

Observation 3086b5be-cc65-4653-892e-e2273ce0e237 · outbound

This paper cites Estimation of Non- Normalized Statistical Models by Score Matching.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Estimation of Non- Normalized Statistical Models by Score Matching

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.467988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:12c05a9ba60f8524daf257c9036d972bcf63b2e26cef2b72940ebe6e19c4b9c7

Observation ba979575-0310-45df-ba27-f1508672c505 · outbound

This paper cites Open- clip.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Open- clip

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.470262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:a7b9712a497774b56d0df63e2afaad62f6cab180e7ad4f7a973723d03aff8a1f

Observation f113f498-f494-41e8-8c7a-0e2fb4feba54 · outbound

This paper cites Open source computer vision library.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Open source computer vision library

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.472725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:8e5d473d1c546d8fb14e6ebe98446bb80d33095e642fa7a47241eec87acfb7ed

Observation 850225e8-c922-427e-b7ca-a5de58c892b4 · outbound

This paper cites Shap-e: Generating condi- tional 3d implicit functions.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Shap-e: Generating condi- tional 3d implicit functions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.475372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:25f3876911e3de1bcfdcf63c928617f3d15af9913da0b3c739ec49d4c82ebbfc

Observation a244474e-b377-4cb7-8994-d8f16a92ef16 · outbound

This paper cites Lower dimensional kernels for video discriminators.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Lower dimensional kernels for video discriminators

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.478776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:c95e15c1840a8314364f5f5968071e91d09d7bb867157521cb9fde709248a5a7

Observation 7a1154aa-5e2d-4727-8b2f-4426575ecb7d · outbound

This paper cites Elucidating the Design Space of Diffusion-Based Generative Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Elucidating the Design Space of Diffusion-Based Generative Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:46:53.453910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:970f72e9c15b3712e7f63751b958005b403dc0f46df7777f5572f75fc8ba325a

Observation f98227f2-8af8-4cce-a16a-8e71eb6f1bc0 · outbound

This paper cites Text2video-zero: Text- to-image diffusion models are zero-shot video generators.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Text2video-zero: Text- to-image diffusion models are zero-shot video generators

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.481343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:d4d9d50c7350d7a39db5d1a8e527a7580702db5aa64ddcb19db7c0425e0bcb97

Observation 0b4bce9d-a8b5-4b0a-90ad-b1ee75f5da75 · outbound

This paper cites Variational diffusion models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Variational diffusion models

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.483561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:5d2b256c163976943195c9118ed28c9d40793f54cc131a63557e698396ce88b2

Observation afffa12c-1938-460a-9ce6-adc9320f50a0 · outbound

This paper cites Pika labs, https://www.pika.art/.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Pika labs, https://www.pika.art/

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.485890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:7311e7b5737edbaf3e23e9a82ff67e9f190b28596e0375c03948ae76ffc958f1

Observation 3fca2898-f60b-487f-af8d-341ceb1f05fc · outbound

This paper cites Stochastic Adversarial Video Prediction.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Stochastic Adversarial Video Prediction

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:58:51.976482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:babbf75d190074802eb5a16454511902d927b21275f2f47dde6086b552f67bde

Observation 3551b37e-644c-4587-ad56-bb649e380c8e · outbound

This paper cites Common Diffusion Noise Schedules and Sample Steps are Flawed.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Common Diffusion Noise Schedules and Sample Steps are Flawed

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:58:51.982261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:7c19de31b6c80fe170b76b0914ac4bc7af7eb37aea7afcf940b446fc54efaa79

Observation 3dc4f274-ad4c-421d-98d0-83360ec06777 · outbound

This paper cites Zero-1-to-3: Zero-shot one image to 3d object.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Zero-1-to-3: Zero-shot one image to 3d object

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.488826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:345eebbedee13a802459e5c79135eaa01089748d973d69519e11314d97deb622

Observation 4f7f2d8c-d0c4-4c54-8d56-1dfd30365503 · outbound

This paper cites SyncDreamer: Generating Multiview-consistent Images from a Single-view Image.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets SyncDreamer: Generating Multiview-consistent Images from a Single-view Image

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:35:43.296571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:2eced2f480d5a122c2f55d83c22839a22fe39da3a17ea027e6121e30af999c0c

Observation 0345f43f-3b3e-4fe8-8b28-1b23a2831069 · outbound

This paper cites Decoupled Weight Decay Regularization.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Decoupled Weight Decay Regularization

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:52.010731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:32eb23e310ecb653234ad5baae0271c33512bfecc9fa3c5a85e079dabef9d288

Observation 70223a73-0f29-4875-90ba-cb84849a1dbc · outbound

This paper cites Transformation-based adversarial video predic- tion on large-scale data.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Transformation-based adversarial video predic- tion on large-scale data

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.159200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:3c57b124b5a25672ed0b1368baed441909ae04d3c8c7ee4bccb76aa3fd64927d

Observation 10ba17ca-7d1e-4955-a69d-afc3eb79366b · outbound

This paper cites Kingma, Stefano Ermon, Jonathan Ho, and Tim Salimans.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Kingma, Stefano Ermon, Jonathan Ho, and Tim Salimans

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.162066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:c1b0387ddcd6a52e71841129356068f61c5e4c909ca8deebcd8581b9bf2b4dd0

Observation 786d8902-6c1f-4998-8440-9e7b4ba5be80 · outbound

This paper cites Point-e: A system for generating 3d point clouds from complex prompts.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Point-e: A system for generating 3d point clouds from complex prompts

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.165074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:b6da3fe2b479e1e9487ca2412f9af2350fb9bd35b9b6bf779fc0068acf4cb1ac

Observation 11226907-7541-4eef-8724-e5170230b98f · outbound

This paper cites The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:43:45.999902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:6cd9acb635fdeb4f9052d510789f364c83d44154be4655c14aeb4b56471fedf7

Observation f9985714-33e4-4fea-a5d7-d10d465cff3e · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:52.063212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:0406dd86cf6b77e8eb68d6489248d7cba6b3b124d658d76398cefb3e6b48e4da

Observation 7370bdaa-45ce-411a-8e06-5168591e9001 · outbound

This paper cites Training contrastive captioners.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Training contrastive captioners

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.167803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:275efb37be773ca8f8e32429edc5f4e9e6153f2f4c31fa0a912da1c07d2c37d5

Observation 7d7aaa5e-a0d0-480e-9ad9-22ff8f14c19e · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Learning Transferable Visual Models From Natural Language Supervision

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:52.074775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:33fb75d1cb5cbcc12ffafe1b5922dd0a8e738a219be5e1170c202f931f9d4834

Observation fb34dd4b-1bda-4d08-8db5-3bbd2325aea6 · outbound

This paper cites an unresolved cited work.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-05-10T22:58:52.170395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:53ad5797bb19840c79a27fa67a369543e1e8b749aa2dc2006052e91328d973af

Observation a901f9ff-a6a0-41f0-bdbc-c1991e011740 · outbound

This paper cites How dall·e 2 works.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets How dall·e 2 works

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.175437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:cb69a7288a6c6558cbd7414a8d6a19c7cf4a1158d5328f3438311f14a9078592

Observation e428c13a-b90e-41af-9ee7-d6c201de1a31 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:52.092030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:d9a1a69c4471df5529ada63b5bc51aba7eb866ca28a410b9f38c9e6503499eaa

Observation 2f8e63da-c358-482b-9917-e97dcb5efc17 · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets High-Resolution Image Synthesis with Latent Diffusion Models

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:52.096430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:bc13013bbe7457a048b80795ae93599a5ffb38985c94d089cb57ab9dbea8bf10

Observation 2a66a5ed-5b6f-4efa-9d32-435f2b17b5f2 · outbound

This paper cites U-Net: Convolutional Networks for Biomedical Image Segmentation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets U-Net: Convolutional Networks for Biomedical Image Segmentation

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:52.103805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:e4fd3194e9d05b76e3b3de77bb109f8d9359059af3174b8ae2677d1ae094940e

Observation b7c509d8-9f86-4d8a-adcd-54bcf546cf10 · outbound

This paper cites Gen-2 by runway, https://research.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Gen-2 by runway, https://research

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.178198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:6455799417a383e1c6c6231fe50f507817a0d7c7c3884a550a695daf08039608

Observation 5ed6e231-1a11-4c9f-9a22-312c2450f838 · outbound

This paper cites Image Super-Resolution via Iterative Refinement.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Image Super-Resolution via Iterative Refinement

Reference 75

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:58:52.111625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:b00bc19c0af755442631375b88b8dc838e21002570f3992a9a7b1bed9317da0b

Observation a642a08a-88d5-43fb-b08f-f323c26ee55f · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:38:54.138183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:7578dcbce5e3d2f4b5cad695fc8b80816b53c8f6db436498893c82fc884aa93e

Observation 8c4286f9-9084-4c24-8a4e-97f72bd68b88 · outbound

This paper cites Tempo- ral generative adversarial nets with singular value clipping.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Tempo- ral generative adversarial nets with singular value clipping

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.187263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:74e690f0c4ae0bdacbc311a3aea1e582fbda3279f6366b699752dd2447bf7404

Observation 41416f5f-b999-466b-91a0-e07ec90dc987 · outbound

This paper cites Train sparsely, generate densely: Memory- efficient unsupervised training of high-resolution temporal gan.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Train sparsely, generate densely: Memory- efficient unsupervised training of high-resolution temporal gan

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.190766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:3c887085d296537266c90fbffeaead0861f66c7e9972022b5224e9340b1ab43c

Observation 4ab4e3b3-883d-4fb5-a4a7-c28d21fae26d · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Progressive Distillation for Fast Sampling of Diffusion Models

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:37:44.815756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:865393ff645283b7be79c5ee61453bb431e203ad2fec4318b05abea5cfd918cf

Observation c0620f95-20b6-4d95-8a0d-5ebb612dd05b · outbound

This paper cites Laion-5b: An open large-scale dataset for train- ing next generation image-text models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Laion-5b: An open large-scale dataset for train- ing next generation image-text models

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.195095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:fda5b1862878393de2999bf6c3eb74464b117e2788dab9cce199bdaac9ed45a3

Observation 34feb37c-b32b-4396-a56c-7092ec300e26 · outbound

This paper cites MVDream: Multi-view Diffusion for 3D Generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets MVDream: Multi-view Diffusion for 3D Generation

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:36:20.266102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:26186c26f4da4777925e904029b2550d8a556807e3fc948b7b3a4fb52c3ddd9b

Observation 9889afaa-6ea3-4723-a270-acb42ef8f891 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:13:03.403852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:8b2ecabcbf6f753bedf9a034defc8846cecc7b6b815c5af6f3d3d52f67bc5994

Observation a2c8a657-23a3-4415-bf59-d6f14e87f27b · outbound

This paper cites Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Stylegan-v: A continuous video generator with the price, image quality and perks of stylegan2

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.198584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:189cd77983710a075773a112b845a7664eb198b63f7a22890125e60a45be1244

Observation 49a95cb2-cbf4-435d-9106-4b640c78c46d · outbound

This paper cites Deep Unsupervised Learning using Nonequilibrium Thermodynamics.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Deep Unsupervised Learning using Nonequilibrium Thermodynamics

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:12:28.687532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:3a868fcde357449bd1217a62945122f24b6e95c963f9d41e5941a1bfdfc2ea32

Observation e14a40b7-92cb-4361-9d15-4735feaadbe3 · outbound

This paper cites Understanding and mitigating copying in diffusion models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Understanding and mitigating copying in diffusion models

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.201576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:ef28d2cc9727a49d751a6c6dbbaf6fd75f046d929b7be09601c36976866132d9

Observation 94f7f97f-5ae4-4be9-a194-770bd2b6d255 · outbound

This paper cites Improved Techniques for Training Score-Based Generative Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Improved Techniques for Training Score-Based Generative Models

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:51.954301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:f4535aa5a0f710ada79d610afa0364b7abc3f264f887b501729b19e6f8d7555e

Observation 7e794e55-785f-4775-a471-dc44b8ae111e · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Score-Based Generative Modeling through Stochastic Differential Equations

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:58:51.958806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:78bd425a12c28cb789cd921114f7a7b337316e3daf380c7f40f0b7984a072e45

Observation 12eb4b0f-eb97-4a00-b0c3-e577c4f0923f · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:25:00.378223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:de2f233c1e09f6b6ee165490b64c54c8e0430184526f3a810d84405120506537

Observation 2188a425-4618-4d2e-a17a-c9f8587e3775 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Raft: Recurrent all-pairs field transforms for optical flow

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.209334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:c5c50c91226c8f24ff20a7b7b7cc7bdf19ebcd3ec376af967276282cf507f51e

Observation 10c673ff-2f67-46c1-b8dd-1f2ce189ba37 · outbound

This paper cites Metaxas, and Sergey Tulyakov.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Metaxas, and Sergey Tulyakov

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.212551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:36971701b9224371f42c7a61e352d70e797097aa400aea28b2adfc097c617ea6

Observation 1f66e2b7-e935-4535-9e1f-82f4c0c19672 · outbound

This paper cites Converting video formats with ffmpeg.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Converting video formats with ffmpeg

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.220100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:4540a1f0955c01a8901a6a6bb7308a81c1d4ac48e9a546e19ef063204d3a6cfb

Observation f5a396a0-1a76-49d3-bc6f-7c805e91a29b · outbound

This paper cites Score-based generative modeling in latent space.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Score-based generative modeling in latent space

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.230113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:c3fd26fe1ad502ebd2098306ac5bd906072645e8029f916dd78ef6c0b8ef1322

Observation 1bf1c45a-fbe8-4a90-bbe9-9f7fd079c56a · outbound

This paper cites Decomposing motion and content for natural video sequence prediction.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Decomposing motion and content for natural video sequence prediction

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.234644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:735ae8b6795fc95e3a43f0ec1860eac88c0019d423eff3393983a66f06364031

Observation cd012fbc-e7d0-4450-85b0-cd7b2ad9d52e · outbound

This paper cites Phenaki: Variable Length Video Generation From Open Domain Textual Description.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Phenaki: Variable Length Video Generation From Open Domain Textual Description

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:43:34.268786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:38fe8bac2d48f4c6d5414612167e7416bbea32bbd6220695d54dd83a68f8f4fc

Observation 301c62a7-1e1c-4f9d-b3af-a1b3cce9a07f · outbound

This paper cites Mcvd: Masked conditional video diffusion for prediction, generation, and interpolation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Mcvd: Masked conditional video diffusion for prediction, generation, and interpolation

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.239876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:4bbeced853e67f6aa22208f418247c30108538dfcc46f0906817eeb3ad0fa0bf

Observation 0ede45f9-e289-4435-a4e4-d06e131f0308 · outbound

This paper cites Generating videos with scene dynamics.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Generating videos with scene dynamics

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.248191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:f2a6dfe9866f23b63122f915d7fcde9b0f4757daf3cfef81a8e3b0cdca7aa0c7

Observation 867eb3c0-1e04-4c90-b25a-f9635e1a6e03 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets ModelScope Text-to-Video Technical Report

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-12T19:47:29.701736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:bff89acc75981418994400a870bf9aed004a8b3ea235275ee5d9728b1375b5c2

Observation ddb335e0-703c-4c00-9b86-35e1f131e0e3 · outbound

This paper cites G3an: Disentangling appearance and mo- tion for video generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets G3an: Disentangling appearance and mo- tion for video generation

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.252462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:e07411cec77536eb2efa4c2cde84dcc7d688463878afa68102a631d4a830209d

Observation 16e5198c-4583-429d-bb8b-1eb3433d4008 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:52.038771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:9e3e19159847c43fa138daaff2cb133fec9b6b3a554b28e149a96a84668de657

Observation 53774ddc-6ab3-4442-aa6b-2eb8bccb4cae · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Internvid: A large-scale video-text dataset for multimodal understanding and generation

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.258107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:74f436355fa2fcd3284540fa33b5b7ed6d7a35cf836e78e8c578de90c9b0a02b

Observation 3deb0188-2f50-4613-a69e-4b1ac2472a23 · outbound

This paper cites Novel view synthesis with diffusion models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Novel view synthesis with diffusion models

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.262032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:9231a4ae4cf948a42d4bf754d26b4dcbd48f9db5cd5222f0b7bfd704032c4dd0

Observation 052d33ce-9e4a-4aff-8ad6-a4f078eaff1d · outbound

This paper cites Scaling autoregressive video models.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets Scaling autoregressive video models

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-05-10T22:58:52.270162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:d6a255df12467b07321a785ce6bd8d9519b8499124fe0dc596d94ec2782e2bb9

Observation 498e28fe-59c7-4351-9027-343899b4de02 · outbound

This paper cites GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions.

Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets GODIVA: Generating Open-DomaIn Videos from nAtural Descriptions

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:52.067091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T22:58:51.792047Z digest=sha256:7efcca557863b1d521784cbd9712852e559a019a38ddac17ba4446e3692edb5f

Pith citing papers

Observation 677f517e-ed14-4944-b95a-9fa4b713fb03 · inbound

VideoPoet: A Large Language Model for Zero-Shot Video Generation cites this paper.

VideoPoet: A Large Language Model for Zero-Shot Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:51:05.543020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T17:51:05.465548Z digest=sha256:caac2c153990628a84bd206058b66c9128f4fdfe4e67268e1b723468228fc1af

Observation 75e31660-5491-4556-88cb-ccdecfdb3d77 · inbound

Latte: Latent Diffusion Transformer for Video Generation cites this paper.

Latte: Latent Diffusion Transformer for Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:45:35.812408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T21:45:35.754742Z digest=sha256:e8027dfc1dc8a33dd081e506805704647e0787ab05a51635b9645ef2516cf0d2

Observation 1f3eaacc-c372-456e-9c1d-a941e73eca30 · inbound

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation cites this paper.

NaVid: Video-based VLM Plans the Next Step for Vision-and-Language Navigation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:55:20.419878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T04:55:20.362512Z digest=sha256:e231f4a2eccc7cd24dcabb8571feda7798ef682da8897624bbd47ff47d42cf6f

Observation 321e56ae-80eb-47fa-998c-2420a872f1e1 · inbound

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models cites this paper.

Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-13T13:43:11.088767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T13:43:11.024069Z digest=sha256:309c9ca512899a620777096503213abc0338d92665724a3364d809392bc97d59

Observation 076c8b60-ee63-4f63-8fe3-f9764c5aca6e · inbound

Scaling Rectified Flow Transformers for High-Resolution Image Synthesis cites this paper.

Scaling Rectified Flow Transformers for High-Resolution Image Synthesis Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 114

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:27:53.615646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T08:27:53.446686Z digest=sha256:a92cfec4cddf5e22d1e5d3a093f2e6c6ac20520544503355d03c0fcb71d553a7

Observation 6425987e-38d7-48ff-841b-678099452bc6 · inbound

CameraCtrl: Enabling Camera Control for Text-to-Video Generation cites this paper.

CameraCtrl: Enabling Camera Control for Text-to-Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:06:23.575572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T02:06:23.410241Z digest=sha256:ffed571f0d37fd510a9bbc9530e244b01b6f3108fecbe23b4007f87049f4735e

Observation c828e4c4-3656-491e-84ee-4f08741443cf · inbound

InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models cites this paper.

InstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T21:15:33.457880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T21:15:33.430513Z digest=sha256:02daeace804ed6ff4e971a42bf0345b5ac44dcf2f1ed5922f4148fceccde1fc7

Observation a3382bfb-fddc-4c12-a12d-8968b9ab22e9 · inbound

CAT3D: Create Anything in 3D with Multi-View Diffusion Models cites this paper.

CAT3D: Create Anything in 3D with Multi-View Diffusion Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:29:51.094198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:29:51.062396Z digest=sha256:c4edd79ba1b414f034724bf826c9349f6c2e774ff2d7018fe976cfdad18960a8

Observation 0bf5dc40-2682-4de3-8de1-cb57f57c26ad · inbound

VideoPhy: Evaluating Physical Commonsense for Video Generation cites this paper.

VideoPhy: Evaluating Physical Commonsense for Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:34:37.690990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T11:34:37.599691Z digest=sha256:2effaac38c54be75485901779f2def870c4974092b6703de9498998368878df3

Observation bb8bf142-82cf-43d4-874c-aa0f3ee16a3a · inbound

OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation cites this paper.

OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:34:53.169615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T20:34:53.141898Z digest=sha256:05685ae8821e7b1c747fc66182395968900db9b9a6638aedfd17e40241c417ca

Observation a3cdd265-58fa-4aa8-9359-fb467f8c3595 · inbound

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer cites this paper.

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:58:52.491420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T18:26:22.224924Z digest=sha256:b2d7aa65d07e87ebc02be3d3fe0b4907b1496a034fcebad4012dc9a24cada87f

Observation 94ea79f7-864c-4058-a759-95fb8bf2e203 · inbound

Diffusion Models Are Real-Time Game Engines cites this paper.

Diffusion Models Are Real-Time Game Engines Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:04:44.211533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T12:04:44.162343Z digest=sha256:f9964c87e33b922c5660f162d878e842a0f95ff91204d9d4fecbdb198a1c4077

Observation 20e19241-9f9d-4c4d-8186-82bbbe15b687 · inbound

ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis cites this paper.

ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:59:03.767446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T22:59:03.642189Z digest=sha256:6a512fbeddfbe18165f012fbce2b935b95b643543c90d17a6a6218a67c2c6c41

Observation 83453ea2-74ea-49e4-aa2b-fcc77435730a · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:56:06.835877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:3f603fa473a8515b3b81ff7741affc3736f0c4f2a063444c75e55c2c46ae0ec7

Observation c0cdda6d-db31-45f2-a834-168fdac0b0dc · inbound

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think cites this paper.

Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T15:09:37.066514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T15:09:36.982610Z digest=sha256:7c7104eee2034ded740d1554c9e304b50c55d4fa7f5868a20e68c5d4192e3114

Observation 10a9d913-aaf0-4b72-8189-aac96663ddfa · inbound

Movie Gen: A Cast of Media Foundation Models cites this paper.

Movie Gen: A Cast of Media Foundation Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T14:16:20.818910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T14:16:18.521699Z digest=sha256:009705c824ab82801e00401da58f934593353faaf7b0b1628451bc0464371121

Observation 7ed2eb50-bcf6-42f5-9b9a-992f78905002 · inbound

KFC-W: Generating 3D-Consistent Videos from Unposed Internet Photos cites this paper.

KFC-W: Generating 3D-Consistent Videos from Unposed Internet Photos Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T16:48:12.512254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T16:47:09.717923Z digest=sha256:746e72c67c2862dbd3b420e102984140777f270fe29e5a65023a3b474707548b

Observation a68c16b4-5018-4766-aa85-77a91a44ceab · inbound

Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement cites this paper.

Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-23T08:25:29.354394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T08:25:01.468957Z digest=sha256:23d76d638878fe342adff5c13c7c5328c9228c206de16d570e2452f688892955

Observation a49ec532-a065-4d3f-af8b-b838a26ca687 · inbound

Open-Sora Plan: Open-Source Large Video Generation Model cites this paper.

Open-Sora Plan: Open-Source Large Video Generation Model Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-23T08:42:45.258752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T08:38:27.946746Z digest=sha256:ea854ee286c8d06ad9c619361696386dfee8243d692b7d1c73e7ce3d7da2cc09

Observation d8b4d63a-1499-4c55-a154-535f05bbab0f · inbound

Motion-2-To-3: Leveraging 2D Motion Data for 3D Motion Generations cites this paper.

Motion-2-To-3: Leveraging 2D Motion Data for 3D Motion Generations Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:37:38.934014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T06:36:47.877594Z digest=sha256:e3e2cae6bdf1d702b891205f846aaabdaf3b48bcc436ce54c045523d4b706a7b

Observation ba50ecb8-fb95-471d-8cc3-a381649a280b · inbound

Autoregressive Video Generation without Vector Quantization cites this paper.

Autoregressive Video Generation without Vector Quantization Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T15:07:39.760358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T15:07:39.718555Z digest=sha256:a1de4430fc37a321aea108675a64f672c391dc69a1cc84766de0a7c18cdd66cb

Observation 237ee8d4-f4e8-485d-9b22-6ebb7c73f84f · inbound

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations cites this paper.

Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:38:11.182858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T18:38:11.110166Z digest=sha256:9d982d8a2940add4636c1e642eabc15e36275b67b93e1a6b7e480539c9ef8b43

Observation d2fd1aa2-a656-40b7-806a-a2b09ef43f00 · inbound

DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization cites this paper.

DOLLAR: Few-Step Video Generation via Distillation and Latent Reward Optimization Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:02:41.930633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T06:57:50.897865Z digest=sha256:302879f57f17dea178a773f6491e6194310bdec717095e6e88baa18b30c86bb3

Observation 85b5b442-404c-4433-a092-63a3a6cb2357 · inbound

Open-Sora: Democratizing Efficient Video Production for All cites this paper.

Open-Sora: Democratizing Efficient Video Production for All Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T12:01:51.645355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T12:01:51.366667Z digest=sha256:e5caaf78348497f9e1af3be55770c3ce91bb13a3851223cc2c602b7f5a984dc2

Observation 3d83f2f1-fe1f-47b1-a0a0-9656997fa64f · inbound

LTX-Video: Realtime Video Latent Diffusion cites this paper.

LTX-Video: Realtime Video Latent Diffusion Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:36:12.420709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T10:36:12.058970Z digest=sha256:6ffda969c471abcce2a953d384da3e66454eae0a27841ada48148e10fecc0bdb

Observation 82540ea0-d2ca-4dd0-b9b6-1a4ebb59f932 · inbound

Cosmos World Foundation Model Platform for Physical AI cites this paper.

Cosmos World Foundation Model Platform for Physical AI Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-10T23:38:45.307012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T23:38:44.933410Z digest=sha256:395d734c317192b4dcd9dc6e62e2802490c271e901df39afacec62a104a7ab0f

Observation d1ef86b4-a3aa-426c-a2c5-d3587826002b · inbound

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought cites this paper.

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:09:34.836871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T23:09:34.805552Z digest=sha256:d8c02c3a5c77f743b0090476ee12f6f614d3956011ea1dbb2f971f3e514c016b

Observation c684a96a-4bfc-4307-bf13-66e160873369 · inbound

Do generative video models understand physical principles? cites this paper.

Do generative video models understand physical principles? Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:47:05.901531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T12:47:05.825659Z digest=sha256:aa30ad12202270f2e791fe296739d43b6b78b8ff2da8420bf78d073a00dc96ee

Observation 19b6c2c8-fbcd-475a-a5e6-fc04a0cf07cb · inbound

Improving Video Generation with Human Feedback cites this paper.

Improving Video Generation with Human Feedback Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T15:30:02.639174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T15:30:02.578430Z digest=sha256:2c8fa7f3cf848064aa64da80f795f43f72fa7d270fd4ba40aa8e262b59c05bd4

Observation a07b5a0c-4e76-44ab-8e4d-ebed4f18c25f · inbound

History-Guided Video Diffusion cites this paper.

History-Guided Video Diffusion Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:00:14.707044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T12:00:14.672729Z digest=sha256:ce48685429b6ac2f42176086e793b2fed1ebffd158d16bd4f1c5a55c5756e0f8

Observation c64eb1dd-b605-4c30-9720-0d44fad4a11e · inbound

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control cites this paper.

RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T19:40:21.149980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T19:40:21.149980Z digest=sha256:8785f5bee88761c1e828c99b8a7a48fbff7c113b8de5620f9407469640c8bd35

Observation c9ed1b39-6f05-40ae-a201-0187c86cc3f0 · inbound

Simulus: Combining Improvements in Sample-Efficient World Model Agents cites this paper.

Simulus: Combining Improvements in Sample-Efficient World Model Agents Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:32:26.237116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T02:28:23.113408Z digest=sha256:a696e95954168f4800a5666c21a29d182e68d0903fbfd34453db375c07efaa92

Observation 97c18906-19cc-4605-b984-8d033d62f0eb · inbound

Unified Video Action Model cites this paper.

Unified Video Action Model Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T17:50:29.713280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T17:50:29.675358Z digest=sha256:1b7711514e2716fb251e9a4cf01c6e06de9e014b4a528abc157abd662eb12732

Observation 841fa45f-0dd9-4ff8-a955-580ab137a3e4 · inbound

Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling cites this paper.

Scene-Action Prompt Fusion for Coherent Text-to-Video Storytelling Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-23T00:12:17.842739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T00:07:39.286486Z digest=sha256:d9eca5ad7dc0308c55a25c4f10b563d64ef9aa500bf5376578e0f2ee8573a3fb

Observation 9ee9a3d2-797a-4cec-bc35-b0fa26860ba3 · inbound

MusicInfuser: Making Video Diffusion Listen and Dance cites this paper.

MusicInfuser: Making Video Diffusion Listen and Dance Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:32:15.650590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T23:28:51.767845Z digest=sha256:4ff7fb8f12b7184467f5165cdb9a8fb220a2be4a0b2081c74195de026b2d4228

Observation 54dc07b6-e6ed-4c1f-a7d4-e68f66cbc19e · inbound

Wan: Open and Advanced Large-Scale Video Generative Models cites this paper.

Wan: Open and Advanced Large-Scale Video Generative Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:07:14.286337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T23:05:32.595632Z digest=sha256:bcaf8f25bcb576ec069d6f3eb7c142d74e84e2b365ac001d324d6cd32d1339ef

Observation bd8ea729-be5e-477f-8d2f-4fe0a3743754 · inbound

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness cites this paper.

VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-14T18:42:03.230991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T18:42:02.940250Z digest=sha256:fb50ecb9b05ec45fb31f80ad3e89a37e8f6d535f11a475e7c304b060796b47da

Observation fad49ec8-c893-4753-8d51-e77611dd93de · inbound

Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets cites this paper.

Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-13T16:25:00.439355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T16:25:00.365534Z digest=sha256:20dba3c5776c2506eb6e9b4ce14a4796d12259a61bb311e2f363167729dfd788

Observation 29356544-a18f-4f15-8c3e-a8e86e9056a5 · inbound

Beyond the Frame: Generating 360 Panoramic Videos from Perspective Videos cites this paper.

Beyond the Frame: Generating 360 Panoramic Videos from Perspective Videos Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-22T19:55:05.065284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T19:54:28.310854Z digest=sha256:b83c20de2ff0e4062bffd0eaa6267fc074a7b59af9e7d16252340e31ddfd728b

Observation 59ba75fe-4e2a-4333-9902-6e3955774afd · inbound

SkyReels-V2: Infinite-length Film Generative Model cites this paper.

SkyReels-V2: Infinite-length Film Generative Model Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:23:04.123266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T20:23:04.022599Z digest=sha256:85810b7c17f18b46d4e5d61259b91b0eae2b14a3c68ae2406a44e4d89df810f9

Observation 72bf4b04-d20b-46e0-9f5f-3c17aaf69760 · inbound

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute cites this paper.

Learning Zero-Shot Subject-Driven Video Generation Using 1% Compute Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-22T17:51:54.382055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T17:51:41.947939Z digest=sha256:5554e8e2e14b09d87c9d49e7faca3bc6b73a6bf455c0dfe02a223e1b66b3395d

Observation 50cb4175-3ebd-441b-91b3-c4bc0f19128f · inbound

DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment cites this paper.

DriVerse: Navigation World Model for Driving Simulation via Multimodal Trajectory Prompting and Motion Alignment Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-22T17:51:54.776728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T17:50:59.797593Z digest=sha256:4ea4e59d95187da3a1bdc2d1f0e177cc6df441e1ae5117fa9f1b6593822e009a

Observation 108c7b27-eccc-4983-8141-a90ce92558f3 · inbound

MAGI-1: Autoregressive Video Generation at Scale cites this paper.

MAGI-1: Autoregressive Video Generation at Scale Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T20:31:15.753806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T20:31:15.700943Z digest=sha256:541a90aa160926885cf5d9270eaabf530aa57091dc0aaa0c016d04fc4697ebe5

Observation 79c19ce1-eb41-4073-bc28-c88e517ff778 · inbound

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers cites this paper.

Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:33:08.118350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:33:08.118350Z digest=sha256:35e25b5d71b112effbe20b5934cb2c76d9ebc9f1df344ebc03efa5f653c6612a

Observation 3733353e-b3c8-406b-b5fd-c08f5c811b4d · inbound

Programmatic Video Prediction Using Large Language Models cites this paper.

Programmatic Video Prediction Using Large Language Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T15:30:10.371648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:30:10.371648Z digest=sha256:d750674b3ecc6e5b2fe399cb1526f7397979743ad14d74cb72545fc8c8f33c92

Observation ae5ca992-6ec5-4dca-b6ea-df7be148469d · inbound

Only Large Weights (And Not Skip Connections) Can Prevent the Perils of Rank Collapse cites this paper.

Only Large Weights (And Not Skip Connections) Can Prevent the Perils of Rank Collapse Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:57.287233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:57.287233Z digest=sha256:27c04df5c214dc1d03e8cd7d9234d87e8c26ae6dac15596bcb3c42416582afb3

Observation 80a80146-b1a2-410c-ae16-7f506d0205e6 · inbound

Character-Centered Dialogue Generation from Scene-Level Prompts cites this paper.

Character-Centered Dialogue Generation from Scene-Level Prompts Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-22T13:34:53.693751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T13:31:43.083678Z digest=sha256:4db707c89e1a736830f91e4e5e615b14dfce53636801dedaade758be970accc9

Observation 08de7fa9-7549-45b9-ab5b-35260468a037 · inbound

Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models cites this paper.

Bigger Isn't Always Memorizing: Early Stopping Overparameterized Diffusion Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:46.068628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:57:46.068628Z digest=sha256:170d4f7bdc25a1cac8330e48ecc90cd1d77f9b42ff77ed81db90c63ce2d71da9

Observation cc4d5ea1-9353-4a81-8f99-887df7b99e21 · inbound

FLEX: A Backbone for Diffusion-Based Modeling of Spatio-temporal Physical Systems cites this paper.

FLEX: A Backbone for Diffusion-Based Modeling of Spatio-temporal Physical Systems Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:01.759600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:01.759600Z digest=sha256:e3b8f072f82453eab57fe2036f310ab85ba6686b05a125f8316db03fce0b58c5

Observation 895db193-cda1-4e99-9da2-2ebb221932bf · inbound

SafeMVDrive: Multi-view Safety-Critical Driving Video Synthesis in the Real World Domain cites this paper.

SafeMVDrive: Multi-view Safety-Critical Driving Video Synthesis in the Real World Domain Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:49.802396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:49.802396Z digest=sha256:5bae0834374646e4993097290563ffea36edcff7565b3088a907bdfff9d8f6c2

Observation df3c0f70-42af-4cd7-8b4c-70e5deb96fb6 · inbound

DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation cites this paper.

DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:39:28.316171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:39:28.316171Z digest=sha256:5f148744b7b03e247b2dbfee829879b17bb83ef58bed50a04ee13e13707ee846

Observation e074d6b8-00f9-48bc-9769-ae97454657f6 · inbound

OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data cites this paper.

OmniConsistency: Learning Style-Agnostic Consistency from Paired Stylization Data Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:06.596891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:06.596891Z digest=sha256:c53c0861f1f9fe2495509836f556fbb041667bb04b3abc8736bab6a0ff959380

Observation 595c10fa-b923-4a7b-a05e-9ab37d1921e0 · inbound

ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos cites this paper.

ProphetDWM: A Driving World Model for Rolling Out Future Actions and Videos Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:27.731461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:27.731461Z digest=sha256:e0fa239a3283b95015ff2bfb392a5350a364b50bdc7b6ea7284a7cf406a77bc5

Observation 6286e85a-4e05-42b9-8c58-064be661215b · inbound

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters cites this paper.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:26.769430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:26.769430Z digest=sha256:8368ff9930b5f272edce63df7628d7931926a3b7120e7fb347101f83ac570421

Observation 8277f684-58eb-43e6-b379-5c07c7dac116 · inbound

Long-Context State-Space Video World Models cites this paper.

Long-Context State-Space Video World Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:13.707554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:13.707554Z digest=sha256:e93fd85d229cfb93538d80071cf411e3758932c429d4e40868571d9ff9780fd9

Observation bf386271-6711-498d-a1da-8699ffd25bc2 · inbound

AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models cites this paper.

AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:07.663909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:07.663909Z digest=sha256:8d4f724d3c5fb1a52f0590149e2660eaf65463ca3f920946cfa14b89657af7e3

Observation cb4ddbcb-35d7-4e66-a484-d31ae36417c3 · inbound

MotionPro: A Precise Motion Controller for Image-to-Video Generation cites this paper.

MotionPro: A Precise Motion Controller for Image-to-Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:02.807538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:02.807538Z digest=sha256:ca62058a9c014d55cdfe0e93dbd6195ce0e129f007d2599b40f3544c8c34b463

Observation d58735dd-48a3-4813-89ea-9807fa66f72b · inbound

Frame-Level Captions for Long Video Generation with Complex Multi Scenes cites this paper.

Frame-Level Captions for Long Video Generation with Complex Multi Scenes Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:52:20.736392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:52:20.736392Z digest=sha256:af91f560b5867da97903d51d35012d0a2bdb77cc1b7633f1ff39f2fa78ffac6a

Observation 1ab7e092-0de3-4f98-a2b9-1c61bc9da826 · inbound

RainFusion: Adaptive Video Generation Acceleration via Multi-Dimensional Visual Redundancy cites this paper.

RainFusion: Adaptive Video Generation Acceleration via Multi-Dimensional Visual Redundancy Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:07.580972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:07.580972Z digest=sha256:f04bb2649b891dab7116f2db995efbf4b2c1a5432d0b25c97436235a02cf463f

Observation dae2afd7-09f8-4dc9-90a7-63ee88430f72 · inbound

Advancing high-fidelity 3D and Texture Generation with 2.5D latents cites this paper.

Advancing high-fidelity 3D and Texture Generation with 2.5D latents Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:50.951297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:50.951297Z digest=sha256:a08f566c492546a0be0ca88b4ec8c6d3233c55c1ff0f8ccf3f9ad1aa447b561f

Observation 91cf9163-01d5-42e6-bffd-8366d7b65d4d · inbound

EF-VI: Enhancing End-Frame Injection for Video Inbetweening cites this paper.

EF-VI: Enhancing End-Frame Injection for Video Inbetweening Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:13.282358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:13.282358Z digest=sha256:06ab74a111e9cb36959e52f63955f70c8f2eea523ec9c5aca5471df48a3dc1cf

Observation 722a32e1-129b-4b66-ab7f-3785ae1952e1 · inbound

EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance cites this paper.

EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video Guidance Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:09.174892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:09.174892Z digest=sha256:0de9b78197ff9e4dc5925d2b5255a3db51be259a43012ca9ae9d2b2f2c8d55de

Observation b0850a3b-bf11-459d-92af-ab65d7745cbb · inbound

VRAG: Learning World Models for Interactive Video Generation cites this paper.

VRAG: Learning World Models for Interactive Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T12:37:17.516687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T12:35:19.379364Z digest=sha256:be1e14f1e13842e689b3052f342f76687f7a4f4f30aa5afd5a40544e9c048a19

Observation 852c06ea-7cd2-4d73-aa4e-ee4af2824020 · inbound

VRAG: Learning World Models for Interactive Video Generation cites this paper.

VRAG: Learning World Models for Interactive Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:20.968150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:20.968150Z digest=sha256:79c453db5fb065c46fc341c8f4e297073690d946294359a1c768f9f30d53cd93

Observation ce86fbe5-cf7e-4cc9-b7c6-4b9d09473c20 · inbound

PanoWan: Lifting Diffusion Video Generation Models to 360{\deg} with Latitude/Longitude-aware Mechanisms cites this paper.

PanoWan: Lifting Diffusion Video Generation Models to 360{\deg} with Latitude/Longitude-aware Mechanisms Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:22:47.219879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:22:47.219879Z digest=sha256:648e5ceb88173dfef520c514c344650085f8d83b6d5bf725cbeea3b34a4c384d

Observation 35f54321-9100-4419-8bde-24ea86a4925d · inbound

GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control cites this paper.

GeoDrive: 3D Geometry-Informed Driving World Model with Precise Action Control Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:13:41.554664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:13:41.554664Z digest=sha256:f5e528fc117d53188abc13c27f8e5f86b7672fbcd224c822f45dde634c4d37d6

Observation 45142d58-ff24-4dff-b3ce-ff0fafedebe3 · inbound

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation cites this paper.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.493506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.493506Z digest=sha256:e258a46b02a97a3bcfd0a5df74be57fa5d806ae64842afcca7e04f97e45bc528

Observation 842663ce-81b1-4b8c-8aa1-fa07edf8a6ab · inbound

MOVi: Training-free Text-conditioned Multi-Object Video Generation cites this paper.

MOVi: Training-free Text-conditioned Multi-Object Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:00:56.872776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:00:56.872776Z digest=sha256:df8203fc0f697305341f2c4e72bba01b0ad7e5ec619e109221dc2a564cab16ee

Observation 5eeb489c-94f3-4068-90fd-ef64ce9dc5bb · inbound

GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video Diffusion cites this paper.

GeoMan: Temporally Consistent Human Geometry Estimation using Image-to-Video Diffusion Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:59:21.225920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:59:21.225920Z digest=sha256:cb55ee256e755b638e46a6b4f0b8721796f56fc9645b555b54c8b0c67263d4f3

Observation ea03f960-50ba-41dd-9578-e88f573b1475 · inbound

Zero-to-Hero: Zero-Shot Initialization Empowering Reference-Based Video Appearance Editing cites this paper.

Zero-to-Hero: Zero-Shot Initialization Empowering Reference-Based Video Appearance Editing Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:57:00.561585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:57:00.561585Z digest=sha256:5d84c1136f569667e127bd5b411657cbe53739f1006f3c8fc3da2b1c31b11c48

Observation cfcd3251-fdca-4688-8d8c-323dae142f30 · inbound

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models cites this paper.

VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:05.474837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:05.474837Z digest=sha256:52155fd91c89069475d98ff539932827fd5dc3e64a0f9000b9a7d01a989c470f

Observation eaaec39b-1b71-4bfd-abc2-1d3a0b864fec · inbound

MUSE: Model-Agnostic Tabular Watermarking via Multi-Sample Selection cites this paper.

MUSE: Model-Agnostic Tabular Watermarking via Multi-Sample Selection Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:37:04.776426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:37:04.776426Z digest=sha256:c0ed5a9d6c3a0533b43cd780cb400a8107d0aea21b66494e887a7e3fc7a41ed1

Observation 1abbe0f1-7b19-4c3b-bd4e-b28a6580c3dc · inbound

UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation cites this paper.

UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:26:57.174478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:26:57.174478Z digest=sha256:1fcea789b08d774936cee29d8189c9b91d652dc3b6595772a4dbf13189cb1dad

Observation 86705e77-ec3e-4622-a47d-b9c43116f61b · inbound

DreamDance: Animating Character Art via Inpainting Stable Gaussian Worlds cites this paper.

DreamDance: Animating Character Art via Inpainting Stable Gaussian Worlds Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:22:18.681635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:22:18.681635Z digest=sha256:b4c143d325d5b80e5c0398fa1987ffa51caab035b0efe1ed5a0fcef4623e7e46

Observation ea566ae7-4703-4d94-9f1b-107b32f69767 · inbound

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers cites this paper.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:04.358842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:04.358842Z digest=sha256:b4bbc8d767c49b14f795f13dbcfa8872bbbe0f2fa65f9a32cfded7ee0e96e7de

Observation c7b7212c-e8e4-438b-847a-288dcd97ad3b · inbound

DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion cites this paper.

DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:46:58.622785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:46:58.622785Z digest=sha256:2aa80addccade27ea6a68f198ec9eef5ec2e047d8f6f3ad2197746f9080dcc81

Observation 9b61f0f1-5a4f-4c2f-a8f8-fc1c130bdd69 · inbound

LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model cites this paper.

LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:46:59.862351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:46:59.862351Z digest=sha256:47566f2cef6c512484ce2c26043327413d1242f0cfc55e559e3ab604f0ff3140

Observation 273f6702-3d5c-4f94-9a9b-37922d42db71 · inbound

OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation cites this paper.

OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:46.291047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:46.291047Z digest=sha256:f79e90e6bb2e18b9c70cb070c74cfbc7f189200253bfb5d99d14a39636fe5b19

Observation dbd32306-deb8-46d5-bbf4-c0a563a7253e · inbound

Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning cites this paper.

Medical World Model: Generative Simulation of Tumor Evolution for Treatment Planning Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:31:48.259736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:31:48.259736Z digest=sha256:0658cb93f16b025c9282977225697f4e0deb8f571b81b47c25b5c855ccd1e1f0

Observation f7cb1d75-dffb-4694-8c49-a23ef5ff2ab0 · inbound

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios cites this paper.

SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:59.510279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:59.510279Z digest=sha256:eacd4f193838058ea19bd6a8b441e5452396678754959bcf60f5d388c3ec1757

Observation 24fb8d44-64c0-4720-b435-99a064bf509c · inbound

LumosFlow: Motion-Guided Long Video Generation cites this paper.

LumosFlow: Motion-Guided Long Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:12.327322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:27:12.327322Z digest=sha256:347aa4f22f34c4a263d4679d89d3d276edc3a0f7441d00b17f65dc24a375cd11

Observation d5e3b7bd-97c6-4f1f-b3cf-10468669cf26 · inbound

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences cites this paper.

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:55.883086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:55.883086Z digest=sha256:232bdadc25d064009a629475063f99baa4f768e10f10a613c087e3e971890e05

Observation 1de799a0-4604-4e43-bb5f-e6791f4cddbe · inbound

NTIRE 2025 XGC Quality Assessment Challenge: Methods and Results cites this paper.

NTIRE 2025 XGC Quality Assessment Challenge: Methods and Results Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:44.689256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:44.689256Z digest=sha256:c6d317c9ed30a74f898a68272f164cf2341d02dfe19b861527ab7fee6f40c74c

Observation 11ee65f5-0dba-4961-9e6f-9d909588bd2a · inbound

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers cites this paper.

Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion Transformers Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:15:14.494535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:15:14.494535Z digest=sha256:00126551f7262e4b6173ffc0097352a190bdb5cd5e9ec23f2985c654a435b67f

Observation a4795bd8-7b9c-4035-b482-3636ec5fffa5 · inbound

IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation cites this paper.

IllumiCraft: Unified Geometry and Illumination Diffusion for Controllable Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:42.913552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:42.913552Z digest=sha256:dd5925bbd9bb60402f0afca7ed85d7f8e3f18073862378774bca394dcd293cc1

Observation e3b028e7-f192-4c5c-91a8-a383465afc90 · inbound

Follow-Your-Creation: Empowering 4D Creation through Video Inpainting cites this paper.

Follow-Your-Creation: Empowering 4D Creation through Video Inpainting Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:43:18.898818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:43:18.898818Z digest=sha256:4b2782158754ae8355b7f06d9d2da248e69a0bd515c8475220344d7490132c08

Observation 0a0adfaf-cba0-4e71-8451-ff8bdfed2e20 · inbound

FEAT: Full-Dimensional Efficient Attention Transformer for Medical Video Generation cites this paper.

FEAT: Full-Dimensional Efficient Attention Transformer for Medical Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:36:27.020201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:36:27.020201Z digest=sha256:489f9d464a20a83cb068dbbb6ae1f9a7eeff324917b5ed34b5687a3c6d855bf7

Observation 4862582a-f8fc-41a9-a43b-5be1a04433a4 · inbound

Video World Models with Long-term Spatial Memory cites this paper.

Video World Models with Long-term Spatial Memory Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:36.546649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:36.546649Z digest=sha256:fcecb8ed1120014b7e4dfa0d8ec72414a34a842754903ae26b2d507a6c4efca2

Observation 3dd28e69-e6fd-4311-9756-8f0bc0ac0d65 · inbound

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models cites this paper.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.849703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.849703Z digest=sha256:4aec4304bcdbac0874aaddeb2052d53ac04ea63857a59152d17e01596ca3615d

Observation 832a6ef4-b1f8-436b-9b7c-8c0540a3a82d · inbound

FADE: Frequency-Aware Diffusion Model Factorization for Video Editing cites this paper.

FADE: Frequency-Aware Diffusion Model Factorization for Video Editing Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:03.643392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:03.643392Z digest=sha256:52c4fe80b364df2c91d56a9a54a0d3f2740301ec62ccaa928371174747fef02e

Observation e26b4304-6a79-4051-b170-79441a4a6267 · inbound

Restereo: Diffusion stereo video generation and restoration cites this paper.

Restereo: Diffusion stereo video generation and restoration Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:38.018637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:38.018637Z digest=sha256:84969e100641966d408d9b2930077eeabfc3068c9a29736533179abf9d7e15b9

Observation f5b5f75f-9f0d-438e-97a6-4702d6cbff61 · inbound

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision cites this paper.

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:58.231968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:58.231968Z digest=sha256:cd9749e32c858e973392500d6e7bd71ab7e0589ce832c82f0afe42e0d59d8d17

Observation 6d8877ef-506a-48fd-8791-fda7ea1628c4 · inbound

Vid2Sim: Generalizable, Video-based Reconstruction of Appearance, Geometry and Physics for Mesh-free Simulation cites this paper.

Vid2Sim: Generalizable, Video-based Reconstruction of Appearance, Geometry and Physics for Mesh-free Simulation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:01.159569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:01.159569Z digest=sha256:294f73e94ce9f0c52c18f4a91de376d6e151e1292b90723326a5a580e53c9d47

Observation 995ab86c-90b7-454a-b2e8-1b79ade6f371 · inbound

Noise Consistency Regularization for Improved Subject-Driven Image Synthesis cites this paper.

Noise Consistency Regularization for Improved Subject-Driven Image Synthesis Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:47.428843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:47.428843Z digest=sha256:6199a6ee1fe7c9c3a98f7dea33307644d35f2506419f903d243aeab8ccdbe3e2

Observation 6bdf8b05-6eb3-42cc-80ca-c1bcd9ab4322 · inbound

Identity Deepfake Threats to Biometric Authentication Systems: Public and Expert Perspectives cites this paper.

Identity Deepfake Threats to Biometric Authentication Systems: Public and Expert Perspectives Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:53:38.931603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:53:38.931603Z digest=sha256:245230e413431077694f9dcd261bf5057f0e0aaefc185c02fa340c3125f36229

Observation c4450487-e28a-4b81-b0ca-1c9f7e5ec4b2 · inbound

Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion cites this paper.

Hi-VAE: Efficient Video Autoencoding with Global and Detailed Motion Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:14.543380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:14.543380Z digest=sha256:de408d8e89906c194a7f6c1e3fd2d1c4a2a1cdd1d5aec0aa19f6ac6ed5add887

Observation a7ccd42a-13c9-4c17-b516-d26b57255dcc · inbound

From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models cites this paper.

From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:30.953970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:30.953970Z digest=sha256:8c007758e15333182ef1acc50cea6e146c612facfca2b9066d5bd86fa865eb27

Observation 318fa9be-2fec-4dc7-a37f-9a4da7b95fc2 · inbound

ProSplat: Improved Feed-Forward 3D Gaussian Splatting for Wide-Baseline Sparse Views cites this paper.

ProSplat: Improved Feed-Forward 3D Gaussian Splatting for Wide-Baseline Sparse Views Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:35:51.092440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:35:51.092440Z digest=sha256:af9f3fe5214089a3621dae8ce99866b31d7ae6274a2b1a28906d5eb7c58e16aa

Observation fb951ecb-febf-4e02-a5f7-5404be91d409 · inbound

NOVA3D: Normal Aligned Video Diffusion Model for Single Image to 3D Generation cites this paper.

NOVA3D: Normal Aligned Video Diffusion Model for Single Image to 3D Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:31:47.863427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:31:47.863427Z digest=sha256:82895e181a782f8d109ad5a83d615490329fc77e54f0a446b2f22d5c8e5fe9a6

Observation 65232c13-4a34-4297-b1a9-9fa062e9ebd3 · inbound

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement cites this paper.

PolyVivid: Vivid Multi-Subject Video Generation with Cross-Modal Interaction and Enhancement Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:30:32.880099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:30:32.880099Z digest=sha256:67dcba4259de11307f96171f548c2a2ecd05defc7bce7fc5157a47776a4c1d09