{"paper":{"title":"High-Resolution Image Synthesis with Latent Diffusion Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Diffusion models trained in the latent space of pretrained autoencoders generate high-resolution images with substantially lower computational cost than pixel-space versions.","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Andreas Blattmann, Bj\\\"orn Ommer, Dominik Lorenz, Patrick Esser, Robin Rombach","submitted_at":"2021-12-20T18:55:25Z","abstract_excerpt":"By decomposing the image formation process into a sequential application of denoising autoencoders, diffusion models (DMs) achieve state-of-the-art synthesis results on image data and beyond. Additionally, their formulation allows for a guiding mechanism to control the image generation process without retraining. However, since these models typically operate directly in pixel space, optimization of powerful DMs often consumes hundreds of GPU days and inference is expensive due to sequential evaluations. To enable DM training on limited computational resources while retaining their quality and "},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Our latent diffusion models (LDMs) achieve a new state of the art for image inpainting and highly competitive performance on various tasks, including unconditional image generation, semantic scene synthesis, and super-resolution, while significantly reducing computational requirements compared to pixel-based DMs.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the latent representation produced by the pretrained autoencoder preserves enough perceptual detail and structure for the diffusion process to recover high-fidelity images without introducing artifacts that cannot be corrected by the model.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Latent diffusion models achieve state-of-the-art inpainting and competitive results on unconditional generation, scene synthesis, and super-resolution by performing the diffusion process in the latent space of pretrained autoencoders with cross-attention conditioning, while cutting computational and","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Diffusion models trained in the latent space of pretrained autoencoders generate high-resolution images with substantially lower computational cost than pixel-space versions.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"99b7551f59f8aa842561a20584499a4b0bcf014149e9f44fde7991607438d940"},"source":{"id":"2112.10752","kind":"arxiv","version":2},"verdict":{"id":"7f89f845-d231-4c56-a135-287e9275a3d8","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-11T21:56:11.949352Z","strongest_claim":"Our latent diffusion models (LDMs) achieve a new state of the art for image inpainting and highly competitive performance on various tasks, including unconditional image generation, semantic scene synthesis, and super-resolution, while significantly reducing computational requirements compared to pixel-based DMs.","one_line_summary":"Latent diffusion models achieve state-of-the-art inpainting and competitive results on unconditional generation, scene synthesis, and super-resolution by performing the diffusion process in the latent space of pretrained autoencoders with cross-attention conditioning, while cutting computational and","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the latent representation produced by the pretrained autoencoder preserves enough perceptual detail and structure for the diffusion process to recover high-fidelity images without introducing artifacts that cannot be corrected by the model.","pith_extraction_headline":"Diffusion models trained in the latent space of pretrained autoencoders generate high-resolution images with substantially lower computational cost than pixel-space versions."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2112.10752/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":109,"sample":[{"doi":"","year":2017,"title":"NTIRE 2017 chal- lenge on single image super-resolution: Dataset and study","work_id":"1988748c-8932-4452-9c00-5b88d646b869","ref_index":1,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2017,"title":"Wasserstein gan","work_id":"7206cfc3-64f8-4a1a-9f7a-1048227063a9","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2019,"title":"Large scale GAN training for high ﬁdelity natural image synthe- sis","work_id":"cd5e6558-54bb-4d43-890b-6d9378c7a056","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2018,"title":"Holger Caesar, Jasper R. R. Uijlings, and Vittorio Ferrari. Coco-stuff: Thing and stuff classes in context. In 2018 IEEE Conference on Computer Vision and Pattern Recog- nition, CVPR 2018, Salt Lake C","work_id":"11b04614-feab-4271-92db-c1c096753367","ref_index":4,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2021,"title":"Extracting training data from large language models","work_id":"53aae647-3f34-4f37-82ac-863e218ba0ff","ref_index":5,"cited_arxiv_id":"","is_internal_anchor":false}],"resolved_work":109,"snapshot_sha256":"9efc94f5acb339e36f169b3946bff0497a9d5f359c733a0f72b4a6ce40e9b7ca","internal_anchors":16},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}