Pith. sign in

Paper Citation Record · LEDGER

Factorized Video Autoencoders for Efficient Generative Modelling

As of 22 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2412.04452.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04452 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:30:03.942929Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa86f80b-4a93-4938-b98c-d8b03c60c694 · outbound

This paper cites Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation.

Factorized Video Autoencoders for Efficient Generative Modelling Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.595252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.595252Z digest=sha256:47f6d2407909bdccb0ed74512d7aab72c74f5389b3c0ef4eee0c2d2d497076ef

Observation 6381c8a7-b767-4739-9705-e4a0a24b720e · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

Factorized Video Autoencoders for Efficient Generative Modelling Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.602198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.602198Z digest=sha256:711e7f5d9bda5af97dafe79f1ebafee55b0beabb6751634ba99f17c84150e9af

Observation c7633e65-5254-4364-85a8-a7e2253132b1 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Factorized Video Autoencoders for Efficient Generative Modelling Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.607910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.607910Z digest=sha256:db13189fa636d3758993da9b0928285161736ea75da72c81013aa644af999f29

Observation a7104834-8ac7-419b-b46b-34f49acab452 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

Factorized Video Autoencoders for Efficient Generative Modelling Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.613538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.613538Z digest=sha256:b4e0974387ac97c76270afc2d10058980382b8ea0b42ce1927a29abebe59ea6c

Observation 252db8f1-7a76-4f7d-9459-7364a68e970f · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

Factorized Video Autoencoders for Efficient Generative Modelling Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:05.013267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.618977Z digest=sha256:c3eda19a8f2eb09df61b87f46b5a48112011cdce3bab9b778d58d07e812fd102

Observation 292a12cb-914f-4042-8834-1a702a8f5ff3 · outbound

This paper cites A short note about kinetics-.

Factorized Video Autoencoders for Efficient Generative Modelling A short note about kinetics-

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.625396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.625396Z digest=sha256:54b5ad098a8c58430b1ce6ba8c33b32d789d4c69d4e5385d6ca4930e7c93837b

Observation 0da9b74b-b9b5-4420-bfa9-41350f215fa6 · outbound

This paper cites Tensorf: Tensorial radiance fields.

Factorized Video Autoencoders for Efficient Generative Modelling Tensorf: Tensorial radiance fields

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.636348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.636348Z digest=sha256:2976c84f62db49dd362148437e4471d6d271925f811ae7891d87dd00fed16df0

Observation 4d548e13-3cf5-4307-aa7d-a1ed5689ff5b · outbound

This paper cites Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning.

Factorized Video Autoencoders for Efficient Generative Modelling Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.640968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.640968Z digest=sha256:dba57869f3824f4b0154a1ea55b793ece919a0b5f90a11899691e4d19fe979d6

Observation 919f90c8-9b29-4470-bcb4-e15df0c00b7a · outbound

This paper cites 3d u-net: learn- ing dense volumetric segmentation from sparse annota- tion.

Factorized Video Autoencoders for Efficient Generative Modelling 3d u-net: learn- ing dense volumetric segmentation from sparse annota- tion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.646167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.646167Z digest=sha256:717643ef4d0d9da2624ce0d0613e143c4790d1fd17cec110c63d0e53d3dbed15

Observation 056dffed-8d84-4955-a67b-df1451abfcd8 · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

Factorized Video Autoencoders for Efficient Generative Modelling Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.651548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.651548Z digest=sha256:9c331d2bf2b33d8a8f095abccf5d9fad537b91c08c1e640e8e89d0d3a71def64

Observation 8ef58888-fdd5-49b2-81e1-ad6428c58a03 · outbound

This paper cites Ldmvfi: Video frame interpolation with latent diffusion models.

Factorized Video Autoencoders for Efficient Generative Modelling Ldmvfi: Video frame interpolation with latent diffusion models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.962108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.657483Z digest=sha256:00148f95e5903fe6aff4db0a6cdc65560bebbdebdfa6c02e8835b6e193e1dc07

Observation 8946269b-0a5f-4905-97bb-429daa074e00 · outbound

This paper cites Diffusion models beat gans on image synthesis.

Factorized Video Autoencoders for Efficient Generative Modelling Diffusion models beat gans on image synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.662268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.662268Z digest=sha256:925e41f2d98881249b755ae5dceb80f5d9f6dcb9b3f992abc6ef19b72efa7086

Observation 915a0a3d-2d8e-47db-ba08-c64866abea16 · outbound

This paper cites Video frame interpolation: A comprehensive survey.

Factorized Video Autoencoders for Efficient Generative Modelling Video frame interpolation: A comprehensive survey

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.932701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.667877Z digest=sha256:b9f068b38a64f52ef1dd0aafd65470bbc7ca6e83a57bf5f452a709b7d9dd69dd

Observation 4c1312d8-7e09-4380-b4ed-1010ae77c4a9 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Factorized Video Autoencoders for Efficient Generative Modelling Taming transformers for high-resolution image synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.672537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.672537Z digest=sha256:d2c5546d1577d75bc3203ea1e87bf69fc549a34f796f391a3d572026c7fd7bd6

Observation d6d89e8b-9947-4e00-ae07-5eaec08b4dce · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Factorized Video Autoencoders for Efficient Generative Modelling Cosmos World Foundation Model Platform for Physical AI

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.677424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.677424Z digest=sha256:92e5015a62656b5a9edbce0d3a6842a9739c1ee57f5040c84d6fc3fe1a947e8a

Observation 680b2c94-eb65-4f5b-aa50-da2e203ccc21 · outbound

This paper cites K-planes: Explicit radiance fields in space, time, and appearance.

Factorized Video Autoencoders for Efficient Generative Modelling K-planes: Explicit radiance fields in space, time, and appearance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.682506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.682506Z digest=sha256:e707b7b24572479cc60f9f4a87cae659401bdccee5fb60a842200ae22e29b0f0

Observation ec7def8e-cb2a-45f5-80aa-6db2cea31c5e · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

Factorized Video Autoencoders for Efficient Generative Modelling Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.687129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.687129Z digest=sha256:5f7d8ec4d3fed546f8d510cc396269110fe4ddcf5754ed32cdb62be7b3626d04

Observation ada504c2-6454-43ac-9c76-7cd4c02c325c · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Factorized Video Autoencoders for Efficient Generative Modelling AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.692294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.692294Z digest=sha256:2c28913f4680b15cb787ad2d9b99f84b05ef121e76fc8280c9f0e40b6f321e5a

Observation ddd0a472-6b8d-4915-97ba-73361a3e8439 · outbound

This paper cites Photorealistic video generation with diffusion models, 2023.

Factorized Video Autoencoders for Efficient Generative Modelling Photorealistic video generation with diffusion models, 2023

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.892252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.697190Z digest=sha256:0e2a4fb488f11a5cfaa9111ddceedd386883e90331273d74ed2b21719387c137

Observation 3e6476b8-a059-42c3-99b7-b13be2bf68a2 · outbound

This paper cites Latent video diffusion models for high-fidelity long video generation.

Factorized Video Autoencoders for Efficient Generative Modelling Latent video diffusion models for high-fidelity long video generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.702493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.702493Z digest=sha256:ee25bd97f68d4f155a6a036889f13437ade2082acb50f613b6e6a46c1aab91be

Observation 3c6490a1-daed-47ad-8345-694d17d96753 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Factorized Video Autoencoders for Efficient Generative Modelling Denoising dif- fusion probabilistic models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.863702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.707231Z digest=sha256:e84857f0ad1810fb4c81c1f4070be6a601d1101aa22a16b3a81b8c1741946045

Observation 801cfd4f-9a6a-458c-a650-270d12715f80 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Factorized Video Autoencoders for Efficient Generative Modelling Imagen Video: High Definition Video Generation with Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.712385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.712385Z digest=sha256:981750d22a5f40dc918e03268248d9c3d00145355974e4271db158c6a9cf47fa

Observation 78898a73-5e60-4ba1-9608-794b6c61d323 · outbound

This paper cites Video dif- fusion models.

Factorized Video Autoencoders for Efficient Generative Modelling Video dif- fusion models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.848158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.717471Z digest=sha256:57319285f343dec1c6f790455a510d4731c882c2fa8faff9b4ac901d85ac8fdd

Observation 1b856818-4376-43aa-bc60-d19b3f88d29a · outbound

This paper cites sim- ple diffusion: End-to-end diffusion for high resolution im- ages.

Factorized Video Autoencoders for Efficient Generative Modelling sim- ple diffusion: End-to-end diffusion for high resolution im- ages

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.831083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.723528Z digest=sha256:d2dd2ee6498985fa3f9a047a9b3ff33259a79ecfc3599a343b23d035b3faf72b

Observation 9362dd53-fa6a-4d1f-9d57-cec8d5f3bce2 · outbound

This paper cites Real-time intermediate flow estimation for video frame interpolation.

Factorized Video Autoencoders for Efficient Generative Modelling Real-time intermediate flow estimation for video frame interpolation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.812209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.728419Z digest=sha256:6214859dba5afc9e80d9e11e6fbb3d75d60b6c77bca1c549d3923d1f1e5092a8

Observation 976a8bc2-a48b-481e-b797-9d52053868cd · outbound

This paper cites Scalable Adaptive Computation for Iterative Generation.

Factorized Video Autoencoders for Efficient Generative Modelling Scalable Adaptive Computation for Iterative Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.733845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.733845Z digest=sha256:d3039553e5d7e5998df0c20a4c48b9c0a45390a7edf185c4016d946e817e675e

Observation c8ce54be-f3a3-4d1f-8a32-23131d49d0ae · outbound

This paper cites Video inter- polation with diffusion models.

Factorized Video Autoencoders for Efficient Generative Modelling Video inter- polation with diffusion models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.795613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.739217Z digest=sha256:66b6fca95bd3b20018e471f0e7a802363e78524023e38232c408c83cba37165e

Observation 19eada01-7cd1-459b-bd51-37fe7aea1571 · outbound

This paper cites Video interpolation with diffu- sion models.

Factorized Video Autoencoders for Efficient Generative Modelling Video interpolation with diffu- sion models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.776767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.743897Z digest=sha256:ee24be8d25210588d927e1b4bf814b5f946cd501aab24d22153f0b8902c8d97e

Observation ed0cd23e-6710-42dc-96e3-cac0cbe6b52c · outbound

This paper cites Benchmarking Video Frame Interpolation.

Factorized Video Autoencoders for Efficient Generative Modelling Benchmarking Video Frame Interpolation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.749078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.749078Z digest=sha256:0888b3c450c859808dd6db6e6c24ff51141032f1cef4d57a69e03d2aa95f3594

Observation 16e0ded2-3661-43db-b8a3-d3bd1ae3457d · outbound

This paper cites Hybrid video diffusion models with 2d triplane and 3d wavelet rep- resentation.

Factorized Video Autoencoders for Efficient Generative Modelling Hybrid video diffusion models with 2d triplane and 3d wavelet rep- resentation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.759875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.754098Z digest=sha256:a7ed1207e1d75e00a014166cdfbbd4c80db0f49e237bca727f61bf69380701b4

Observation d67c5c52-ba39-4c39-a17b-f81b38f57dc1 · outbound

This paper cites Auto-Encoding Variational Bayes.

Factorized Video Autoencoders for Efficient Generative Modelling Auto-Encoding Variational Bayes

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.759220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.759220Z digest=sha256:9b3a77860befbb8c114196d9ae7c68988480778c737bdcde02f93cf9ad978c20

Observation bcf71a2e-c152-455f-b5f9-b1154d11911f · outbound

This paper cites Semcity: Semantic scene gener- ation with triplane diffusion.

Factorized Video Autoencoders for Efficient Generative Modelling Semcity: Semantic scene gener- ation with triplane diffusion

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.742631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.766452Z digest=sha256:92ebf2c6d98c01f800409fce49fdfe69a8b4d08d14c3bc43e5a512dac7a9c0e1

Observation 4d88f02c-cbb5-42d7-93c8-b30f7f7bf6c8 · outbound

This paper cites Amt: All-pairs multi-field transforms for efficient frame interpolation.

Factorized Video Autoencoders for Efficient Generative Modelling Amt: All-pairs multi-field transforms for efficient frame interpolation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.771423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.771423Z digest=sha256:684761a4cae5260808c879647ae3ae7a4a1de4d68cd2721389b25188f3023b5a

Observation e20f972a-40a2-4087-8ac8-b9f900513c40 · outbound

This paper cites WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model.

Factorized Video Autoencoders for Efficient Generative Modelling WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.776579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.776579Z digest=sha256:e494fe9bd305ea6dfc0d57dcf9445f095a13b9a56319ea632d11326c520b36ba

Observation 80853de7-8dda-4104-adcf-8d6a7741b96a · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Factorized Video Autoencoders for Efficient Generative Modelling Open-Sora Plan: Open-Source Large Video Generation Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.782354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.782354Z digest=sha256:bc40d1a3123d4e963e45c9bc64407e4e9f32d98e37bd640f264cc8527626af5c

Observation 58eeaaf1-0466-4cbb-8d24-8f5a961ff568 · outbound

This paper cites Common diffusion noise schedules and sample steps are flawed.

Factorized Video Autoencoders for Efficient Generative Modelling Common diffusion noise schedules and sample steps are flawed

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.715688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.788060Z digest=sha256:49338a60f732caa3ca8abd80ceb089ccd69f3be79ba8db75a2226c310e07f471

Observation 89dd405f-52e9-498f-a10b-eb176716b5ad · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Factorized Video Autoencoders for Efficient Generative Modelling SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.793892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.793892Z digest=sha256:cfcd5d6038df2167ff1ca7137eaa7ed9d90d42d43448da5d717b99e441352f4c

Observation 3b70a688-6960-43da-928f-3946a17155e5 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Factorized Video Autoencoders for Efficient Generative Modelling Movie Gen: A Cast of Media Foundation Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.799772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.799772Z digest=sha256:772ce66e8c3fe23efcc3dcd108980613d88280ddb8adbdb1a7f0727027c788e1

Observation c779d7a1-778d-4067-961f-066afe942f17 · outbound

This paper cites Gener- ating diverse high-fidelity images with vq-vae-2.

Factorized Video Autoencoders for Efficient Generative Modelling Gener- ating diverse high-fidelity images with vq-vae-2

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.806000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.806000Z digest=sha256:c1a2de4eaf57b4937f22eb6ecfbca887a1d0a083dbea7d8fc79580cc142e4f33

Observation 0ee18d90-2f2a-4e4d-8eab-ce93d9dc9c78 · outbound

This paper cites Film: Frame inter- polation for large motion.

Factorized Video Autoencoders for Efficient Generative Modelling Film: Frame inter- polation for large motion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.688919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.811064Z digest=sha256:d54e1750b7aab594a4c9e322e67affee9be14024dffed578f4e178d698562e6c

Observation 0652cc0b-39ec-47ce-9408-5143ba09a015 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models, 2021, 2021.

Factorized Video Autoencoders for Efficient Generative Modelling High-resolution image syn- thesis with latent diffusion models, 2021, 2021

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.673254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.815511Z digest=sha256:aa4cfa3daef97fe3b9d35b1cfc0d6998b231659fb99f5ec8df79c5f70f0df614

Observation 39daf284-1542-433c-bb7e-1efee45d9b3d · outbound

This paper cites U- net: Convolutional networks for biomedical image segmen- tation.

Factorized Video Autoencoders for Efficient Generative Modelling U- net: Convolutional networks for biomedical image segmen- tation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.821864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.821864Z digest=sha256:4fedcf0b537eb17ded223c8ebf5b8f2950225ea669ceeb438f76398492300b7e

Observation dd149793-971a-4dd7-bb34-f8035c81d2b3 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Factorized Video Autoencoders for Efficient Generative Modelling Photorealistic text-to-image diffusion models with deep language understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.827317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.827317Z digest=sha256:1cc7faea736bed7e29e57b049101cb2f78b144e8ceaa16d83729af8d1a2d2af5

Observation c24de059-c635-4dea-9e8e-938683fe7d93 · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Factorized Video Autoencoders for Efficient Generative Modelling Progressive Distillation for Fast Sampling of Diffusion Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.833018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.833018Z digest=sha256:edda7b30c30608575d4775d777b15cc89f9fca2fb6ef7f299aca3c96986047ee

Observation fc48e18a-a25a-4b4c-8049-058f110e5773 · outbound

This paper cites Ryan Shue, Eric Ryan Chan, Ryan Po, Zachary Ankner, Jiajun Wu, and Gordon Wetzstein.

Factorized Video Autoencoders for Efficient Generative Modelling Ryan Shue, Eric Ryan Chan, Ryan Po, Zachary Ankner, Jiajun Wu, and Gordon Wetzstein

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.635684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.837917Z digest=sha256:c87a9ab8798208fd011c07c07fef16fb2fa747eebcdca15ec7a850b07e5b1141

Observation 01088514-0af1-4035-bf64-2e370a04a4b8 · outbound

This paper cites Denoising Diffusion Implicit Models.

Factorized Video Autoencoders for Efficient Generative Modelling Denoising Diffusion Implicit Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.843325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.843325Z digest=sha256:bcc5e3272c2756dcbd5e513b5e6f70bde7183df5f7901bd4880575b8c919d0d3

Observation 172af5ce-4c50-4693-a9c7-856b48ed7dfd · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Factorized Video Autoencoders for Efficient Generative Modelling UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.848635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.848635Z digest=sha256:2dda6e1c7a40a0253dcec29665ea0b7d94159d26842c0b43bb8fbbcb2980a4a7

Observation d46bc2f2-ea24-488c-9762-519615c0e090 · outbound

This paper cites Fvd: A new metric for video generation.

Factorized Video Autoencoders for Efficient Generative Modelling Fvd: A new metric for video generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.854150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.854150Z digest=sha256:e09cc93a27760f64bcd01ef472f0b858f7ad2e185d599db9b4bca2fa2615d870

Observation d4ebc61a-f597-4bf0-aa69-05b2246b6621 · outbound

This paper cites Neural discrete representation learning.

Factorized Video Autoencoders for Efficient Generative Modelling Neural discrete representation learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.859862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.859862Z digest=sha256:5bebfe22bf7af8539c573b7452ed716031e68b446877e1f5788ed7c58716fc53

Observation 52817f95-b4a4-445d-a685-505eb40dbac1 · outbound

This paper cites Attention is all you need.

Factorized Video Autoencoders for Efficient Generative Modelling Attention is all you need

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.865128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.865128Z digest=sha256:83fddca1e292ab88e4bc9021415c261e908122b23445bae0bfd8da7ce915643f

Observation 9b1af98f-a6f0-4a51-867e-7489a0ae56f9 · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

Factorized Video Autoencoders for Efficient Generative Modelling Phenaki: Variable length video generation from open domain textual descriptions

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.586719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.870267Z digest=sha256:29946134c9d026ea543620511597a51217e0e60fd3836f3177f8fa48f7de7093

Observation f0dbc5c7-6fb3-43ce-a36e-a3d2cadf78d2 · outbound

This paper cites Omnitokenizer: A joint image-video tokenizer for visual generation.

Factorized Video Autoencoders for Efficient Generative Modelling Omnitokenizer: A joint image-video tokenizer for visual generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.568768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.874943Z digest=sha256:1677e41f352576d8ef1502036eed309f1afbfd0d5789dd2b05e7068148973a5f

Observation 7f20194b-eaa8-449b-a935-6ab099d44eea · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

Factorized Video Autoencoders for Efficient Generative Modelling Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.880228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.880228Z digest=sha256:2c6cd8bfedac32fc684485ed9cc2a6b373dab96c436ee958cbd508437a585af0

Observation acfc32a9-a8e4-4c21-9880-38064b6f6149 · outbound

This paper cites Sin3DM: Learning a Diffusion Model from a Single 3D Textured Shape.

Factorized Video Autoencoders for Efficient Generative Modelling Sin3DM: Learning a Diffusion Model from a Single 3D Textured Shape

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.885358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.885358Z digest=sha256:c37d5c3f6323f7f1952513d89821b205aa15a633489de4aa15df54ce281578df

Observation f0648488-85b3-4d10-8001-0b30b4e4dd01 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Factorized Video Autoencoders for Efficient Generative Modelling CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.892204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.892204Z digest=sha256:fc4e23c44fd4f2e868a46e1ccf4685a9fadf79b955cc8d065cfdce872e2460ce

Observation b7a1ad7b-ef8a-4421-beca-78e70988da9a · outbound

This paper cites Magvit: Masked generative video transformer.

Factorized Video Autoencoders for Efficient Generative Modelling Magvit: Masked generative video transformer

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.539734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.897772Z digest=sha256:38fd9c0a35bad3c467690ecc03505d20f2b6e62c1ee62a30a1365db790b2b3a1

Observation 3e09ef54-2a39-4041-b5c2-ee92ba68c1e5 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Factorized Video Autoencoders for Efficient Generative Modelling Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.903667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.903667Z digest=sha256:b63be37b9270eb4cfeeb60cf6e2af7d81899cd68b3df9c0995f6b9f9d2f0c3a0

Observation ef971c59-9385-43bd-8e89-e9c6f5e0118c · outbound

This paper cites Language model beats diffusion - tokenizer is key to visual generation.

Factorized Video Autoencoders for Efficient Generative Modelling Language model beats diffusion - tokenizer is key to visual generation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.522805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.910211Z digest=sha256:33b2ef78b49dec1fbc8c1904cc4e6a2a4bbf642dc9d8e7cc846a39d597ea1738

Observation a636f1e9-ef7f-48a5-8d75-9d10fbf6c495 · outbound

This paper cites Video probabilistic diffusion models in projected latent space.

Factorized Video Autoencoders for Efficient Generative Modelling Video probabilistic diffusion models in projected latent space

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.915062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.915062Z digest=sha256:9a1ba0c7e66d88cb5beb5497b9254c846765373a7590eab9650747a5a92f9464

Observation 221e7fb3-0296-4240-b001-28019e484271 · outbound

This paper cites Video probabilistic diffusion models in projected latent space.

Factorized Video Autoencoders for Efficient Generative Modelling Video probabilistic diffusion models in projected latent space

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.493940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.919956Z digest=sha256:3240fe74688818d53fbbaa24b87f5f287be8a9a3d64f4f4ec6dcf0dca65d8c57

Observation d4d32f16-ae4c-44cc-9bc6-61362294856b · outbound

This paper cites Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition.

Factorized Video Autoencoders for Efficient Generative Modelling Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.925490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.925490Z digest=sha256:f4ab010ff64cf10cef0419f367c0c287b4ab5016b540c569a4126b76fff6151b

Observation b23fa786-5c1f-4aaa-8aa9-facd31e6f960 · outbound

This paper cites Cv- vae: A compatible video vae for latent generative video mod- els.

Factorized Video Autoencoders for Efficient Generative Modelling Cv- vae: A compatible video vae for latent generative video mod- els

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.477573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.930974Z digest=sha256:b5a3c3c78fb2c657e4c20051433a390d7e3838183846dc1298a79685812f8b6f

Observation 9ef45027-92b3-45df-a3e2-1f450dd6999e · outbound

This paper cites an unresolved cited work.

Factorized Video Autoencoders for Efficient Generative Modelling Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:30:04.442723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.942929Z digest=sha256:a588163ea6e895cab677a1d16cdabb1efdb3d6b35279686c64302c3aea0b03bf

Observation e6b022e0-87ef-404b-b8cc-f93dc781bac3 · outbound

This paper cites For the video interpolation task, the autoencoder is 2 trained for 450, 000 iterations with the same batch size of.

Factorized Video Autoencoders for Efficient Generative Modelling For the video interpolation task, the autoencoder is 2 trained for 450, 000 iterations with the same batch size of

Reference 256

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.460419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-11T21:30:03.937509Z digest=sha256:a0cf27f42dc6440bf8919affb01436b91a0ce2a50ff4a494028ccfcf0a591da7

Observation 25b2559b-ba51-4350-9f14-9132c51ec185 · outbound

This paper cites A Short Note about Kinetics-600.

Factorized Video Autoencoders for Efficient Generative Modelling A Short Note about Kinetics-600

Reference 600

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.630717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.630717Z digest=sha256:e9ccb8f8c37664373d6ea6cefd76972ec6dd5c4e54f8010604e2ef7225dd1c87

Pith citing papers

No inbound Pith citation observations are available.