Pith. sign in

Paper Citation Record · LEDGER

Factorized Video Autoencoders for Efficient Generative Modelling

As of 15 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2412.04452.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04452 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:30:03.942929Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa86f80b-4a93-4938-b98c-d8b03c60c694 · outbound

This paper cites Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation.

Factorized Video Autoencoders for Efficient Generative Modelling Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.595252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.595252Z digest=sha256:a974fb448b44b80538258e4393e45350afea4e7f3ec33a2c64f5781400c977f2

Observation 6381c8a7-b767-4739-9705-e4a0a24b720e · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

Factorized Video Autoencoders for Efficient Generative Modelling Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.602198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.602198Z digest=sha256:d01f538acacd4c3bf08051a5038aedf85543cc48aeca2de2608146deea249a5e

Observation c7633e65-5254-4364-85a8-a7e2253132b1 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Factorized Video Autoencoders for Efficient Generative Modelling Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.607910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.607910Z digest=sha256:d06035b34622b4cf325deabfbc23d2698532f9c5e5b22b91d957d6fb157d1cd1

Observation a7104834-8ac7-419b-b46b-34f49acab452 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

Factorized Video Autoencoders for Efficient Generative Modelling Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.613538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.613538Z digest=sha256:7f285222f547931d53340e8232464e573097bd9425db0622dcf9bc13a33a5629

Observation 252db8f1-7a76-4f7d-9459-7364a68e970f · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

Factorized Video Autoencoders for Efficient Generative Modelling Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:05.013267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.618977Z digest=sha256:4264b74b590907e5f7bab7aca082f344f1498f0a424bf6a0b876fa51bb444e9d

Observation 292a12cb-914f-4042-8834-1a702a8f5ff3 · outbound

This paper cites A short note about kinetics-.

Factorized Video Autoencoders for Efficient Generative Modelling A short note about kinetics-

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.625396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.625396Z digest=sha256:bf576ac3ea3e110cd5fbe90f0379a34c7789128433e48fda9a5cedbe1666937a

Observation 0da9b74b-b9b5-4420-bfa9-41350f215fa6 · outbound

This paper cites Tensorf: Tensorial radiance fields.

Factorized Video Autoencoders for Efficient Generative Modelling Tensorf: Tensorial radiance fields

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.636348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.636348Z digest=sha256:cd37e0cd503578102e22ff153f66fa5e92871e4bfeb15897c8c58918d3b63672

Observation 4d548e13-3cf5-4307-aa7d-a1ed5689ff5b · outbound

This paper cites Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning.

Factorized Video Autoencoders for Efficient Generative Modelling Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.640968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.640968Z digest=sha256:70fe1a4297868dc7b0bd4c25cf5c8052908b3a5fb90983162ecde7d7a03c1e4d

Observation 919f90c8-9b29-4470-bcb4-e15df0c00b7a · outbound

This paper cites 3d u-net: learn- ing dense volumetric segmentation from sparse annota- tion.

Factorized Video Autoencoders for Efficient Generative Modelling 3d u-net: learn- ing dense volumetric segmentation from sparse annota- tion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.646167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.646167Z digest=sha256:676de525885931f1d18a3c0963d346b95bf83193ce877d611fa71e1f8428de8b

Observation 056dffed-8d84-4955-a67b-df1451abfcd8 · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

Factorized Video Autoencoders for Efficient Generative Modelling Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.651548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.651548Z digest=sha256:d7641f7314d43c267f1c5c08a42b8938f1083d8b54bd39a45718407766ab6b97

Observation 8ef58888-fdd5-49b2-81e1-ad6428c58a03 · outbound

This paper cites Ldmvfi: Video frame interpolation with latent diffusion models.

Factorized Video Autoencoders for Efficient Generative Modelling Ldmvfi: Video frame interpolation with latent diffusion models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.962108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.657483Z digest=sha256:095392ea929c1418cf1e3f45dbeee8a59f7f3b43dcb5748c3d81f628041e42e5

Observation 8946269b-0a5f-4905-97bb-429daa074e00 · outbound

This paper cites Diffusion models beat gans on image synthesis.

Factorized Video Autoencoders for Efficient Generative Modelling Diffusion models beat gans on image synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.662268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.662268Z digest=sha256:dd43a5feb012df53c65f233cac5a26112f87ae9fc953054c1dd24aef789e4658

Observation 915a0a3d-2d8e-47db-ba08-c64866abea16 · outbound

This paper cites Video frame interpolation: A comprehensive survey.

Factorized Video Autoencoders for Efficient Generative Modelling Video frame interpolation: A comprehensive survey

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.932701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.667877Z digest=sha256:7e100d6b435ba70a7590e7c63c154f75da152508c37752513521c3b13f38ea05

Observation 4c1312d8-7e09-4380-b4ed-1010ae77c4a9 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Factorized Video Autoencoders for Efficient Generative Modelling Taming transformers for high-resolution image synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.672537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.672537Z digest=sha256:d24f247007158a27120b61b32144320d2c10905215c9a8ac2c3f96f3fff06363

Observation d6d89e8b-9947-4e00-ae07-5eaec08b4dce · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Factorized Video Autoencoders for Efficient Generative Modelling Cosmos World Foundation Model Platform for Physical AI

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.677424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.677424Z digest=sha256:36543a9811d9e846cfa0bcc9dff6fefe78d61f9a3a8e83cc0f8b8693c431f193

Observation 680b2c94-eb65-4f5b-aa50-da2e203ccc21 · outbound

This paper cites K-planes: Explicit radiance fields in space, time, and appearance.

Factorized Video Autoencoders for Efficient Generative Modelling K-planes: Explicit radiance fields in space, time, and appearance

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.682506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.682506Z digest=sha256:ee50faab3e8de9fa7185601eef89d53ccac6dc3d984a05b5ac2aeb2782fb0c08

Observation ec7def8e-cb2a-45f5-80aa-6db2cea31c5e · outbound

This paper cites Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning.

Factorized Video Autoencoders for Efficient Generative Modelling Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.687129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.687129Z digest=sha256:63fabdd0417f4c52d3e2b676defaf0c0c090e3d776de2c1e34022008b5d4f2df

Observation ada504c2-6454-43ac-9c76-7cd4c02c325c · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Factorized Video Autoencoders for Efficient Generative Modelling AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.692294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.692294Z digest=sha256:e5275fbd940c8859989a796f37a51c4ebe263d34673b174852d8f6286e94421a

Observation ddd0a472-6b8d-4915-97ba-73361a3e8439 · outbound

This paper cites Photorealistic video generation with diffusion models, 2023.

Factorized Video Autoencoders for Efficient Generative Modelling Photorealistic video generation with diffusion models, 2023

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.892252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.697190Z digest=sha256:f6b6d807460866045739951fe87ab7b64545104ac546c953e72b460c45bfecdd

Observation 3e6476b8-a059-42c3-99b7-b13be2bf68a2 · outbound

This paper cites Latent video diffusion models for high-fidelity long video generation.

Factorized Video Autoencoders for Efficient Generative Modelling Latent video diffusion models for high-fidelity long video generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.702493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.702493Z digest=sha256:c2f689437dbefd2bf4907f828e87d0ee16feddbd62d67d0bca8b0ceaff7db1c6

Observation 3c6490a1-daed-47ad-8345-694d17d96753 · outbound

This paper cites Denoising dif- fusion probabilistic models.

Factorized Video Autoencoders for Efficient Generative Modelling Denoising dif- fusion probabilistic models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.863702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.707231Z digest=sha256:f730a950d51f9f9c41403309e2d3cd63756cba304ce493dbdfbd9f0c59ef45bf

Observation 801cfd4f-9a6a-458c-a650-270d12715f80 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Factorized Video Autoencoders for Efficient Generative Modelling Imagen Video: High Definition Video Generation with Diffusion Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.712385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.712385Z digest=sha256:77124a95dead4b9050e2502b66ead1c679caa95b174e60f4290f7ebc7863c957

Observation 78898a73-5e60-4ba1-9608-794b6c61d323 · outbound

This paper cites Video dif- fusion models.

Factorized Video Autoencoders for Efficient Generative Modelling Video dif- fusion models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.848158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.717471Z digest=sha256:77d882c8bba620d27022d980b54a8d5e8a626b21598d556d6e4c3509ae7e3057

Observation 1b856818-4376-43aa-bc60-d19b3f88d29a · outbound

This paper cites sim- ple diffusion: End-to-end diffusion for high resolution im- ages.

Factorized Video Autoencoders for Efficient Generative Modelling sim- ple diffusion: End-to-end diffusion for high resolution im- ages

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.831083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.723528Z digest=sha256:26349122c290119609391278ec009a03b843d6b92b14f1a6fe99af52f637a467

Observation 9362dd53-fa6a-4d1f-9d57-cec8d5f3bce2 · outbound

This paper cites Real-time intermediate flow estimation for video frame interpolation.

Factorized Video Autoencoders for Efficient Generative Modelling Real-time intermediate flow estimation for video frame interpolation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.812209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.728419Z digest=sha256:304924f4a0f9958cd6c6f807eb76eadd3cc422b8692835aa4468fb181de09b77

Observation 976a8bc2-a48b-481e-b797-9d52053868cd · outbound

This paper cites Scalable Adaptive Computation for Iterative Generation.

Factorized Video Autoencoders for Efficient Generative Modelling Scalable Adaptive Computation for Iterative Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.733845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.733845Z digest=sha256:d396a8194161f1f5bdbc96419761c00bae70de3113e4daf827775960902839da

Observation c8ce54be-f3a3-4d1f-8a32-23131d49d0ae · outbound

This paper cites Video inter- polation with diffusion models.

Factorized Video Autoencoders for Efficient Generative Modelling Video inter- polation with diffusion models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.795613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.739217Z digest=sha256:a5383b9a117b0fec4902e12128ee5866514970891de708237675a4b44607ad2c

Observation 19eada01-7cd1-459b-bd51-37fe7aea1571 · outbound

This paper cites Video interpolation with diffu- sion models.

Factorized Video Autoencoders for Efficient Generative Modelling Video interpolation with diffu- sion models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.776767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.743897Z digest=sha256:b181592985d5a61d34b5f287af6d04bb6de488715eb7df242fae083a976b6938

Observation ed0cd23e-6710-42dc-96e3-cac0cbe6b52c · outbound

This paper cites Benchmarking Video Frame Interpolation.

Factorized Video Autoencoders for Efficient Generative Modelling Benchmarking Video Frame Interpolation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.749078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.749078Z digest=sha256:ea8e0074f0a03a460f0386b9c400d68a9a7e00a8ffc9a4bbfd20b7d9a9d0d89d

Observation 16e0ded2-3661-43db-b8a3-d3bd1ae3457d · outbound

This paper cites Hybrid video diffusion models with 2d triplane and 3d wavelet rep- resentation.

Factorized Video Autoencoders for Efficient Generative Modelling Hybrid video diffusion models with 2d triplane and 3d wavelet rep- resentation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.759875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.754098Z digest=sha256:8362d3f97dc25eef8a4228bc63c8ef6d3ba277a57f972e483d019b5cf11c7250

Observation d67c5c52-ba39-4c39-a17b-f81b38f57dc1 · outbound

This paper cites Auto-Encoding Variational Bayes.

Factorized Video Autoencoders for Efficient Generative Modelling Auto-Encoding Variational Bayes

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.759220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.759220Z digest=sha256:855ba2eab09451c1cadd26c91210d9142212327689379b5f16ea917b4a6fa766

Observation bcf71a2e-c152-455f-b5f9-b1154d11911f · outbound

This paper cites Semcity: Semantic scene gener- ation with triplane diffusion.

Factorized Video Autoencoders for Efficient Generative Modelling Semcity: Semantic scene gener- ation with triplane diffusion

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.742631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.766452Z digest=sha256:df95dbf4cde3c1e1b612fb900970c82be1c938c4577cc49f803671fff473d53a

Observation 4d88f02c-cbb5-42d7-93c8-b30f7f7bf6c8 · outbound

This paper cites Amt: All-pairs multi-field transforms for efficient frame interpolation.

Factorized Video Autoencoders for Efficient Generative Modelling Amt: All-pairs multi-field transforms for efficient frame interpolation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.771423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.771423Z digest=sha256:199498a4b196322a0f7820ed519e634b9e897c7ff5f605c1439492ef41d1beda

Observation e20f972a-40a2-4087-8ac8-b9f900513c40 · outbound

This paper cites WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model.

Factorized Video Autoencoders for Efficient Generative Modelling WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.776579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.776579Z digest=sha256:1a5a6e08add31d7ca760bbdf691b06c93127d795f7dce4cb0c91b8fada36db41

Observation 80853de7-8dda-4104-adcf-8d6a7741b96a · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Factorized Video Autoencoders for Efficient Generative Modelling Open-Sora Plan: Open-Source Large Video Generation Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.782354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.782354Z digest=sha256:c8cee6ac31d134455eb188630aedaaf7824ae47b442a7f7a91e6cdbd680958b0

Observation 58eeaaf1-0466-4cbb-8d24-8f5a961ff568 · outbound

This paper cites Common diffusion noise schedules and sample steps are flawed.

Factorized Video Autoencoders for Efficient Generative Modelling Common diffusion noise schedules and sample steps are flawed

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.715688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.788060Z digest=sha256:3e7b755dca0b3310873ccb0f0fe301ed2df3b56657b20e3db04dc49655775de6

Observation 89dd405f-52e9-498f-a10b-eb176716b5ad · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Factorized Video Autoencoders for Efficient Generative Modelling SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.793892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.793892Z digest=sha256:f6d6cb543f9ddf410a66716380662a2adf4ffef2513af151e6babb0715032727

Observation 3b70a688-6960-43da-928f-3946a17155e5 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Factorized Video Autoencoders for Efficient Generative Modelling Movie Gen: A Cast of Media Foundation Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.799772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.799772Z digest=sha256:c2428e60248d605ed7342c274a574f58f444264a37f54355f0ea37bd7d98174e

Observation c779d7a1-778d-4067-961f-066afe942f17 · outbound

This paper cites Gener- ating diverse high-fidelity images with vq-vae-2.

Factorized Video Autoencoders for Efficient Generative Modelling Gener- ating diverse high-fidelity images with vq-vae-2

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.806000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.806000Z digest=sha256:c34aecadc3d864bbe4c154c083eaa4bc073767ed4ab9bd111466742402577e83

Observation 0ee18d90-2f2a-4e4d-8eab-ce93d9dc9c78 · outbound

This paper cites Film: Frame inter- polation for large motion.

Factorized Video Autoencoders for Efficient Generative Modelling Film: Frame inter- polation for large motion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.688919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.811064Z digest=sha256:772b0dd46e10626c0db34f96b8a4eca294d90cf300e1a1d1c10f6b617cc95f45

Observation 0652cc0b-39ec-47ce-9408-5143ba09a015 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models, 2021, 2021.

Factorized Video Autoencoders for Efficient Generative Modelling High-resolution image syn- thesis with latent diffusion models, 2021, 2021

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.673254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.815511Z digest=sha256:5dc5fc1e561b71d136f69b87ac8dbb4fae9bf0d96b13d01c2d5eb0ab77affa5d

Observation 39daf284-1542-433c-bb7e-1efee45d9b3d · outbound

This paper cites U- net: Convolutional networks for biomedical image segmen- tation.

Factorized Video Autoencoders for Efficient Generative Modelling U- net: Convolutional networks for biomedical image segmen- tation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.821864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.821864Z digest=sha256:bf306a968043ed57e9591c669aa60695db704cdefaae2b4ead6eb8d857ee7014

Observation dd149793-971a-4dd7-bb34-f8035c81d2b3 · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Factorized Video Autoencoders for Efficient Generative Modelling Photorealistic text-to-image diffusion models with deep language understanding

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.827317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.827317Z digest=sha256:fb71c27a340ba76beac1429b2efcfd09109ae65220f71b5a01793a4a5caa74e2

Observation c24de059-c635-4dea-9e8e-938683fe7d93 · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Factorized Video Autoencoders for Efficient Generative Modelling Progressive Distillation for Fast Sampling of Diffusion Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.833018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.833018Z digest=sha256:a77b81e7a5c81022457ee9a74e6daa1a3d46b2de00e1e4a67e48258925349e95

Observation fc48e18a-a25a-4b4c-8049-058f110e5773 · outbound

This paper cites Ryan Shue, Eric Ryan Chan, Ryan Po, Zachary Ankner, Jiajun Wu, and Gordon Wetzstein.

Factorized Video Autoencoders for Efficient Generative Modelling Ryan Shue, Eric Ryan Chan, Ryan Po, Zachary Ankner, Jiajun Wu, and Gordon Wetzstein

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.635684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.837917Z digest=sha256:47e9abb303a7a820e044f0375d0cc96abc7ea425610f0f3c7886f0bd762d9fbb

Observation 01088514-0af1-4035-bf64-2e370a04a4b8 · outbound

This paper cites Denoising Diffusion Implicit Models.

Factorized Video Autoencoders for Efficient Generative Modelling Denoising Diffusion Implicit Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.843325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.843325Z digest=sha256:0c69da809111e84742ff5956017180857f3e5f69becd8df6e697f319beb0ef16

Observation 172af5ce-4c50-4693-a9c7-856b48ed7dfd · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Factorized Video Autoencoders for Efficient Generative Modelling UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.848635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.848635Z digest=sha256:b37a285fac82173733bcf0db11058deff1a6a0ae66e51592f2bbe51747aea823

Observation d46bc2f2-ea24-488c-9762-519615c0e090 · outbound

This paper cites Fvd: A new metric for video generation.

Factorized Video Autoencoders for Efficient Generative Modelling Fvd: A new metric for video generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.854150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.854150Z digest=sha256:aac27f2dffbb8898934205004b6e8b642d1ea77917107ea36128d04471f1894b

Observation d4ebc61a-f597-4bf0-aa69-05b2246b6621 · outbound

This paper cites Neural discrete representation learning.

Factorized Video Autoencoders for Efficient Generative Modelling Neural discrete representation learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.859862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.859862Z digest=sha256:2189df9aa49a6a06ace3cc40633e5934f862a49b44cdda9015854cffe22d3a95

Observation 52817f95-b4a4-445d-a685-505eb40dbac1 · outbound

This paper cites Attention is all you need.

Factorized Video Autoencoders for Efficient Generative Modelling Attention is all you need

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.865128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.865128Z digest=sha256:4196baf159c435ed8d4f9ebc21c83e0048c2bb08e8e32f5110a4195e132b4d32

Observation 9b1af98f-a6f0-4a51-867e-7489a0ae56f9 · outbound

This paper cites Phenaki: Variable length video generation from open domain textual descriptions.

Factorized Video Autoencoders for Efficient Generative Modelling Phenaki: Variable length video generation from open domain textual descriptions

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.586719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.870267Z digest=sha256:065da3b541953aba82b9b834ea2c0cb0728097108bf65122069cb271bff14151

Observation f0dbc5c7-6fb3-43ce-a36e-a3d2cadf78d2 · outbound

This paper cites Omnitokenizer: A joint image-video tokenizer for visual generation.

Factorized Video Autoencoders for Efficient Generative Modelling Omnitokenizer: A joint image-video tokenizer for visual generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.568768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.874943Z digest=sha256:28af3672334cb154300a7470d2931d5813dedb53c9b913e6762fa53221834c67

Observation 7f20194b-eaa8-449b-a935-6ab099d44eea · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

Factorized Video Autoencoders for Efficient Generative Modelling Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.880228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.880228Z digest=sha256:65c17ac1f007f03e22f597f6b74d0f242c19c4a3a6830299baf82ba555cc5572

Observation acfc32a9-a8e4-4c21-9880-38064b6f6149 · outbound

This paper cites Sin3DM: Learning a Diffusion Model from a Single 3D Textured Shape.

Factorized Video Autoencoders for Efficient Generative Modelling Sin3DM: Learning a Diffusion Model from a Single 3D Textured Shape

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.885358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.885358Z digest=sha256:4c3a269d656684be575454a23d0107972469b2e692380dcaa3ca5f24c8e930ca

Observation f0648488-85b3-4d10-8001-0b30b4e4dd01 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Factorized Video Autoencoders for Efficient Generative Modelling CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.892204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.892204Z digest=sha256:379bae04d9d3a6c07476a4409c1533a3196a19635f475037362b5ed17547c640

Observation b7a1ad7b-ef8a-4421-beca-78e70988da9a · outbound

This paper cites Magvit: Masked generative video transformer.

Factorized Video Autoencoders for Efficient Generative Modelling Magvit: Masked generative video transformer

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.539734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.897772Z digest=sha256:5bce61dd0a6256b5648392a04e1cd37da6469c1d4273f973b1dc616a33aaeafc

Observation 3e09ef54-2a39-4041-b5c2-ee92ba68c1e5 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Factorized Video Autoencoders for Efficient Generative Modelling Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.903667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.903667Z digest=sha256:585bb15172cb6b5842c8e5e9e247f8078fa77ae56b061baaf6e041e8110b4f74

Observation ef971c59-9385-43bd-8e89-e9c6f5e0118c · outbound

This paper cites Language model beats diffusion - tokenizer is key to visual generation.

Factorized Video Autoencoders for Efficient Generative Modelling Language model beats diffusion - tokenizer is key to visual generation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.522805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.910211Z digest=sha256:42872d710f08a91401138dc100577f863e96a0f0e989da35e0628ab109481a5b

Observation a636f1e9-ef7f-48a5-8d75-9d10fbf6c495 · outbound

This paper cites Video probabilistic diffusion models in projected latent space.

Factorized Video Autoencoders for Efficient Generative Modelling Video probabilistic diffusion models in projected latent space

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.915062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.915062Z digest=sha256:5dc4f3280fb2eb9be910d010e2b6c88a26935174aa1ea4d360f31eefd5f9b933

Observation 221e7fb3-0296-4240-b001-28019e484271 · outbound

This paper cites Video probabilistic diffusion models in projected latent space.

Factorized Video Autoencoders for Efficient Generative Modelling Video probabilistic diffusion models in projected latent space

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.493940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.919956Z digest=sha256:cbd3f465a257150890afbe3374a50c90850c3a87235f6adcec01201c473974b3

Observation d4d32f16-ae4c-44cc-9bc6-61362294856b · outbound

This paper cites Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition.

Factorized Video Autoencoders for Efficient Generative Modelling Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.925490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.925490Z digest=sha256:3c950898c033d1e75e818f593771ae878a26305de0b0654615db31379db97ae4

Observation b23fa786-5c1f-4aaa-8aa9-facd31e6f960 · outbound

This paper cites Cv- vae: A compatible video vae for latent generative video mod- els.

Factorized Video Autoencoders for Efficient Generative Modelling Cv- vae: A compatible video vae for latent generative video mod- els

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.477573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.930974Z digest=sha256:054dfe2963a041a716203f70c32a8d9cb47c833a416ede195d40aabd4c5ec0a6

Observation 9ef45027-92b3-45df-a3e2-1f450dd6999e · outbound

This paper cites an unresolved cited work.

Factorized Video Autoencoders for Efficient Generative Modelling Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:30:04.442723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.942929Z digest=sha256:2f6afc6c326a5421a05b1cf5f1a48c7dc6fab2316cea431301319c986ae43e8c

Observation e6b022e0-87ef-404b-b8cc-f93dc781bac3 · outbound

This paper cites For the video interpolation task, the autoencoder is 2 trained for 450, 000 iterations with the same batch size of.

Factorized Video Autoencoders for Efficient Generative Modelling For the video interpolation task, the autoencoder is 2 trained for 450, 000 iterations with the same batch size of

Reference 256

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:30:04.460419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:30:03.937509Z digest=sha256:e3dd30015474dc6c20b0205a0ac9b70ba8f5856d949af5e17b7ea75727eb8a83

Observation 25b2559b-ba51-4350-9f14-9132c51ec185 · outbound

This paper cites A Short Note about Kinetics-600.

Factorized Video Autoencoders for Efficient Generative Modelling A Short Note about Kinetics-600

Reference 600

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:03.630717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:03.630717Z digest=sha256:0cad2f9e55e38c98c099b2912cfe38bdd6886062f9b57fc6234cfcadb958e611

Pith citing papers

No inbound Pith citation observations are available.