Pith. sign in

REVIEW 16 cited by

Transparent Image Layer Diffusion using Latent Transparency

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.17113 v4 pith:6QGYFPAJ submitted 2024-02-27 cs.CV cs.GR

classification cs.CVcs.GR
keywords latenttransparentdiffusionlayermodeltransparencyimagegeneration
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present LayerDiffuse, an approach enabling large-scale pretrained latent diffusion models to generate transparent images. The method allows generation of single transparent images or of multiple transparent layers. The method learns a "latent transparency" that encodes alpha channel transparency into the latent manifold of a pretrained latent diffusion model. It preserves the production-ready quality of the large diffusion model by regulating the added transparency as a latent offset with minimal changes to the original latent distribution of the pretrained model. In this way, any latent diffusion model can be converted into a transparent image generator by finetuning it with the adjusted latent space. We train the model with 1M transparent image layer pairs collected using a human-in-the-loop collection scheme. We show that latent transparency can be applied to different open source image generators, or be adapted to various conditional control systems to achieve applications like foreground/background-conditioned layer generation, joint layer generation, structural control of layer contents, etc. A user study finds that in most cases (97%) users prefer our natively generated transparent content over previous ad-hoc solutions such as generating and then matting. Users also report the quality of our generated transparent images is comparable to real commercial transparent assets like Adobe Stock.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LiWi: Layering in the Wild

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    Introduces LiWi-100k dataset via agent-orchestrated synthesis and a decomposition model with shadow-guided learning and boundary correction that claims state-of-the-art RGB L1 and Alpha IoU on natural images.

  2. LiWi: Layering in the Wild

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    LiWi uses an agent-driven data synthesis pipeline to build the LiWi-100k dataset and a model with shadow-guided and degradation-restoration objectives that achieves SoTA performance on RGB L1 and Alpha IoU for natural...

  3. RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    RevealLayer decomposes natural images into multiple RGBA layers using diffusion models with region-aware attention, occlusion-guided adaptation, and a composite loss, outperforming prior methods on a new benchmark dataset.

  4. The Garden of Forking Paths: Narrative Arc-Conditioned Gameplay Planning

    cs.HC 2026-05 unverdicted novelty 7.0 of 10

    Forking Garden generates branching dungeon graphs from user-provided storylines by creating a pool of independent nodes and assembling them via arc-guided constraint algorithms that enforce multimodal alignment of gam...

  5. A Unified and Controllable Framework for Layered Image Generation with Visual Effects

    cs.CV 2026-01 unverdicted novelty 7.0 of 10

    LASAGNA produces layered images with integrated visual effects in a single pass, enabling drift-free edits via alpha compositing while releasing a 48K dataset and a 242-sample benchmark.

  6. LaRender: Training-Free Occlusion Control in Image Generation via Latent Rendering

    cs.CV 2025-08 conditional novelty 7.0 of 10

    LaRender replaces cross-attention layers in a pretrained diffusion model with a latent alpha-compositing operation that renders object features in occlusion order, giving training-free occlusion control.

  7. UniWorld-Design: From Pixel Generation to Layer-Native Design

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A two-model framework generates images as transparent layers and decomposes finished designs into ordered, complete semantic layers, outperforming prior decomposition models on per-layer fidelity and editability.

  8. Parallax Portrait Matting

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Parallax from a casually captured second view, fused asymmetrically (background-aligned pixels plus foreground-aligned cross-attention), yields finer portrait mattes and cleaner foreground colors than strong single-im...

  9. Lighting-Consistent Object Transfer Across Radiance Fields

    cs.GR 2026-06 unverdicted novelty 6.0 of 10

    Diffusion-based per-view harmonization for lighting-consistent object transfer between 3DGS scenes, using heterogeneous training data and final 3D consolidation.

  10. PAI-Studio: Cinematic Video Background Replacement with Camera-Aware Motion

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    PAI-Studio reformulates cinematic background replacement as in-context conditional generation inside a Diffusion Transformer with bidirectional attention, trained on a new 30K film-sourced dataset, and reports better ...

  11. LiWi: Layering in the Wild

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Presents LiWi-100k dataset generated via agent-driven decomposition and a model achieving SoTA on RGB L1 and Alpha IoU for natural image layering.

  12. LimeCross: Context-Conditioned Layered Image Editing with Structural Consistency

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    LimeCross enables text-guided editing of individual layers in composite images by conditioning on cross-layer context via bi-stream attention while preserving layer integrity and introducing the LayerEditBench benchmark.

  13. CatalogStitch: Dimension-Aware and Occlusion-Preserving Object Compositing for Catalog Image Generation

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    CatalogStitch provides dimension-aware mask computation and occlusion-aware hybrid restoration to automate corrections in generative object compositing for catalog images.

  14. BLK-Assist: A Methodological Framework for Artist-Led Co-Creation with Generative AI Models

    cs.CY 2026-03 unverdicted novelty 6.0 of 10

    BLK-Assist is a three-part framework (Conceptor for sketches, Stencil for transparent assets, Upscale for high-res outputs) that fine-tunes public diffusion models on one artist's proprietary corpus for style-faithful...

  15. Text-Conditioned Background Generation for Editable Multi-Layer Documents

    cs.CV 2025-12 conditional novelty 5.0 of 10

    A training-free system combines soft latent masking, WCAG-contrast-optimized semi-transparent text backings, and recursive LLM summaries to generate readable, style-consistent backgrounds for multi-page documents.

  16. All Stories Are One Story: Emotional Arc Guided Procedural Game Level Generation

    cs.AI 2025-08 conditional novelty 5.0 of 10

    A procedural game generation system that uses Rise/Fall emotional arcs to shape LLM-written branching stories and entity difficulty showed higher player enjoyment in a small ARPG user study.

Pith tools