{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2021:XFBNG3P6ULQKUALPGAL6PJXPBZ","short_pith_number":"pith:XFBNG3P6","schema_version":"1.0","canonical_sha256":"b942d36dfea2e0aa016f3017e7a6ef0e72610a251e77e9c777257a50a2a7e993","source":{"kind":"arxiv","id":"2112.10752","version":2},"attestation_state":"computed","paper":{"title":"High-Resolution Image Synthesis with Latent Diffusion Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Diffusion models trained in the latent space of pretrained autoencoders generate high-resolution images with substantially lower computational cost than pixel-space versions.","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Andreas Blattmann, Bj\\\"orn Ommer, Dominik Lorenz, Patrick Esser, Robin Rombach","submitted_at":"2021-12-20T18:55:25Z","abstract_excerpt":"By decomposing the image formation process into a sequential application of denoising autoencoders, diffusion models (DMs) achieve state-of-the-art synthesis results on image data and beyond. Additionally, their formulation allows for a guiding mechanism to control the image generation process without retraining. However, since these models typically operate directly in pixel space, optimization of powerful DMs often consumes hundreds of GPU days and inference is expensive due to sequential evaluations. To enable DM training on limited computational resources while retaining their quality and "},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":false},"canonical_record":{"source":{"id":"2112.10752","kind":"arxiv","version":2},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CV","submitted_at":"2021-12-20T18:55:25Z","cross_cats_sorted":[],"title_canon_sha256":"deec5d5eb94f2ddfe005fb1d91fc54db06e956ceabe8801103d39447982b0826","abstract_canon_sha256":"bf3049144b7bde983b5d125b703794b3d857826a6753906bc996f3000aa600d6"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T04:14:27.084698Z","signature_b64":"K8EzV43KbpPYdgqETtgoeN4KsJ4Mj7X3bmt4FSyFVYPQr1jDsMsKP7M5L9RW6gcGU4NFC5aR0DpUenZJbBOxAQ==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"b942d36dfea2e0aa016f3017e7a6ef0e72610a251e77e9c777257a50a2a7e993","last_reissued_at":"2026-07-05T04:14:27.084189Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T04:14:27.084189Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"High-Resolution Image Synthesis with Latent Diffusion Models","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Diffusion models trained in the latent space of pretrained autoencoders generate high-resolution images with substantially lower computational cost than pixel-space versions.","cross_cats":[],"primary_cat":"cs.CV","authors_text":"Andreas Blattmann, Bj\\\"orn Ommer, Dominik Lorenz, Patrick Esser, Robin Rombach","submitted_at":"2021-12-20T18:55:25Z","abstract_excerpt":"By decomposing the image formation process into a sequential application of denoising autoencoders, diffusion models (DMs) achieve state-of-the-art synthesis results on image data and beyond. Additionally, their formulation allows for a guiding mechanism to control the image generation process without retraining. However, since these models typically operate directly in pixel space, optimization of powerful DMs often consumes hundreds of GPU days and inference is expensive due to sequential evaluations. To enable DM training on limited computational resources while retaining their quality and "},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Our latent diffusion models (LDMs) achieve a new state of the art for image inpainting and highly competitive performance on various tasks, including unconditional image generation, semantic scene synthesis, and super-resolution, while significantly reducing computational requirements compared to pixel-based DMs.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the latent representation produced by the pretrained autoencoder preserves enough perceptual detail and structure for the diffusion process to recover high-fidelity images without introducing artifacts that cannot be corrected by the model.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Latent diffusion models achieve state-of-the-art inpainting and competitive results on unconditional generation, scene synthesis, and super-resolution by performing the diffusion process in the latent space of pretrained autoencoders with cross-attention conditioning, while cutting computational and","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Diffusion models trained in the latent space of pretrained autoencoders generate high-resolution images with substantially lower computational cost than pixel-space versions.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"99b7551f59f8aa842561a20584499a4b0bcf014149e9f44fde7991607438d940"},"source":{"id":"2112.10752","kind":"arxiv","version":2},"verdict":{"id":"7f89f845-d231-4c56-a135-287e9275a3d8","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-11T21:56:11.949352Z","strongest_claim":"Our latent diffusion models (LDMs) achieve a new state of the art for image inpainting and highly competitive performance on various tasks, including unconditional image generation, semantic scene synthesis, and super-resolution, while significantly reducing computational requirements compared to pixel-based DMs.","one_line_summary":"Latent diffusion models achieve state-of-the-art inpainting and competitive results on unconditional generation, scene synthesis, and super-resolution by performing the diffusion process in the latent space of pretrained autoencoders with cross-attention conditioning, while cutting computational and","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the latent representation produced by the pretrained autoencoder preserves enough perceptual detail and structure for the diffusion process to recover high-fidelity images without introducing artifacts that cannot be corrected by the model.","pith_extraction_headline":"Diffusion models trained in the latent space of pretrained autoencoders generate high-resolution images with substantially lower computational cost than pixel-space versions."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2112.10752/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":109,"sample":[{"doi":"","year":2017,"title":"NTIRE 2017 chal- lenge on single image super-resolution: Dataset and study","work_id":"1988748c-8932-4452-9c00-5b88d646b869","ref_index":1,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2017,"title":"Wasserstein gan","work_id":"7206cfc3-64f8-4a1a-9f7a-1048227063a9","ref_index":2,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2019,"title":"Large scale GAN training for high ﬁdelity natural image synthe- sis","work_id":"cd5e6558-54bb-4d43-890b-6d9378c7a056","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2018,"title":"Holger Caesar, Jasper R. R. Uijlings, and Vittorio Ferrari. Coco-stuff: Thing and stuff classes in context. In 2018 IEEE Conference on Computer Vision and Pattern Recog- nition, CVPR 2018, Salt Lake C","work_id":"11b04614-feab-4271-92db-c1c096753367","ref_index":4,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2021,"title":"Extracting training data from large language models","work_id":"53aae647-3f34-4f37-82ac-863e218ba0ff","ref_index":5,"cited_arxiv_id":"","is_internal_anchor":false}],"resolved_work":109,"snapshot_sha256":"9efc94f5acb339e36f169b3946bff0497a9d5f359c733a0f72b4a6ce40e9b7ca","internal_anchors":16},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2112.10752","created_at":"2026-07-05T04:14:27.084258+00:00"},{"alias_kind":"arxiv_version","alias_value":"2112.10752v2","created_at":"2026-07-05T04:14:27.084258+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2112.10752","created_at":"2026-07-05T04:14:27.084258+00:00"},{"alias_kind":"pith_short_12","alias_value":"XFBNG3P6ULQK","created_at":"2026-07-05T04:14:27.084258+00:00"},{"alias_kind":"pith_short_16","alias_value":"XFBNG3P6ULQKUALP","created_at":"2026-07-05T04:14:27.084258+00:00"},{"alias_kind":"pith_short_8","alias_value":"XFBNG3P6","created_at":"2026-07-05T04:14:27.084258+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":138,"internal_anchor_count":113,"sample":[{"citing_arxiv_id":"2607.07173","citing_title":"Stage-Aware Adaptation and Distribution Calibration for Subject-Driven Personalized Text-to-Image Generation","ref_index":2,"is_internal_anchor":true},{"citing_arxiv_id":"2607.06445","citing_title":"Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders","ref_index":23,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26312","citing_title":"Tailor Made Embeddings for Quantum Machine Learning","ref_index":39,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26872","citing_title":"SpatialFlow-GRPO: Where Spatial Credit Drives Image Editing","ref_index":36,"is_internal_anchor":true},{"citing_arxiv_id":"2606.20971","citing_title":"UNITY: Attention Flow Networks for Adaptive Conditioning in Diffusion","ref_index":25,"is_internal_anchor":true},{"citing_arxiv_id":"2606.19151","citing_title":"The Market in the Model: Latent Diffusion as Neural Economy","ref_index":24,"is_internal_anchor":true},{"citing_arxiv_id":"2607.01248","citing_title":"A Practice Auditing Framework for Large Language Model Use: Collective Empiricism, Pseudo-Rational Cognition, and Governance of AI-Generated Content","ref_index":10,"is_internal_anchor":true},{"citing_arxiv_id":"2606.20722","citing_title":"Multimodal Image Colorization: Quantifying the Impact of Text-Conditioned Guidance on Grayscale-to-Color Translation","ref_index":12,"is_internal_anchor":true},{"citing_arxiv_id":"2606.17432","citing_title":"Edit3DGS: Unified Framework for Dynamic Head Editing via 2D Instruction-Guided Diffusion and 3D Gaussian Splatting","ref_index":19,"is_internal_anchor":true},{"citing_arxiv_id":"2606.13676","citing_title":"Modality Forcing for Scalable Spatial Generation","ref_index":30,"is_internal_anchor":true},{"citing_arxiv_id":"2606.12365","citing_title":"Ambient Diffusion Policy: Imitation Learning from Suboptimal Data in Robotics","ref_index":66,"is_internal_anchor":true},{"citing_arxiv_id":"2606.11096","citing_title":"IDEAL: In-DEpth ALignment Makes A Discrete Representation AutoEncoder","ref_index":39,"is_internal_anchor":true},{"citing_arxiv_id":"2606.11363","citing_title":"NSVQ: Mitigating Codebook Collapse by Stabilizing Encoder Drift in Vector Quantization","ref_index":12,"is_internal_anchor":true},{"citing_arxiv_id":"2606.10046","citing_title":"Inside the Latent Flow: Causal Deciphering of Attention Dynamics in Audio Separation Foundation Models","ref_index":24,"is_internal_anchor":true},{"citing_arxiv_id":"2606.09670","citing_title":"Visual Prompting Meets Feature Reconstruction-Based Anomaly Detection with Dual-Teacher Supervision","ref_index":12,"is_internal_anchor":true},{"citing_arxiv_id":"2606.08032","citing_title":"Variational Proximal Policy Optimization","ref_index":53,"is_internal_anchor":true},{"citing_arxiv_id":"2606.07117","citing_title":"Native3D: End-to-End 3D Scene Generation via Unified Mesh-Texture Modeling and Semantic Alignment","ref_index":32,"is_internal_anchor":true},{"citing_arxiv_id":"2606.06918","citing_title":"DRIFT: From Robustness Gaps to Invariance Manifolds for AI-Generated Image Detection","ref_index":38,"is_internal_anchor":true},{"citing_arxiv_id":"2607.00402","citing_title":"The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models","ref_index":24,"is_internal_anchor":true},{"citing_arxiv_id":"2607.00527","citing_title":"AI Native Games: A Survey and Roadmap","ref_index":51,"is_internal_anchor":true},{"citing_arxiv_id":"2606.05328","citing_title":"The Invisible Hand of Physics: When Video Diffusion Models Know More Than They Show","ref_index":22,"is_internal_anchor":true},{"citing_arxiv_id":"2606.03617","citing_title":"SA-DTS: Semantic-Aware Digital Twin Synchronization over 6G Networks","ref_index":8,"is_internal_anchor":true},{"citing_arxiv_id":"2606.03512","citing_title":"SPADE: Sketch-guided Path Planning Augmented with Diffusion Experts","ref_index":19,"is_internal_anchor":true},{"citing_arxiv_id":"2606.06518","citing_title":"DiBS: Diffusion-Informed Branch Selection","ref_index":12,"is_internal_anchor":true},{"citing_arxiv_id":"2606.02532","citing_title":"Improving Combined Detection and Classification of TEM Defects via Mask-Conditioned Latent Diffusion Augmentation","ref_index":37,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/XFBNG3P6ULQKUALPGAL6PJXPBZ","json":"https://pith.science/pith/XFBNG3P6ULQKUALPGAL6PJXPBZ.json","graph_json":"https://pith.science/api/pith-number/XFBNG3P6ULQKUALPGAL6PJXPBZ/graph.json","events_json":"https://pith.science/api/pith-number/XFBNG3P6ULQKUALPGAL6PJXPBZ/events.json","paper":"https://pith.science/paper/XFBNG3P6"},"agent_actions":{"view_html":"https://pith.science/pith/XFBNG3P6ULQKUALPGAL6PJXPBZ","download_json":"https://pith.science/pith/XFBNG3P6ULQKUALPGAL6PJXPBZ.json","view_paper":"https://pith.science/paper/XFBNG3P6","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2112.10752&json=true","fetch_graph":"https://pith.science/api/pith-number/XFBNG3P6ULQKUALPGAL6PJXPBZ/graph.json","fetch_events":"https://pith.science/api/pith-number/XFBNG3P6ULQKUALPGAL6PJXPBZ/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/XFBNG3P6ULQKUALPGAL6PJXPBZ/action/timestamp_anchor","attest_storage":"https://pith.science/pith/XFBNG3P6ULQKUALPGAL6PJXPBZ/action/storage_attestation","attest_author":"https://pith.science/pith/XFBNG3P6ULQKUALPGAL6PJXPBZ/action/author_attestation","sign_citation":"https://pith.science/pith/XFBNG3P6ULQKUALPGAL6PJXPBZ/action/citation_signature","submit_replication":"https://pith.science/pith/XFBNG3P6ULQKUALPGAL6PJXPBZ/action/replication_record"}},"created_at":"2026-07-05T04:14:27.084258+00:00","updated_at":"2026-07-05T04:14:27.084258+00:00"}