Pith. sign in

REVIEW 11 cited by

Text2Layer: Layered Image Generation using Latent Diffusion Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.09781 v1 pith:UGEQSNY2 submitted 2023-07-19 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords imagelayeredlayercompositingdiffusiongenerationablebenefit
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Layer compositing is one of the most popular image editing workflows among both amateurs and professionals. Motivated by the success of diffusion models, we explore layer compositing from a layered image generation perspective. Instead of generating an image, we propose to generate background, foreground, layer mask, and the composed image simultaneously. To achieve layered image generation, we train an autoencoder that is able to reconstruct layered images and train diffusion models on the latent representation. One benefit of the proposed problem is to enable better compositing workflows in addition to the high-quality image output. Another benefit is producing higher-quality layer masks compared to masks produced by a separate step of image segmentation. Experimental results show that the proposed method is able to generate high-quality layered images and initiates a benchmark for future work.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LiWi: Layering in the Wild

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    Introduces LiWi-100k dataset via agent-orchestrated synthesis and a decomposition model with shadow-guided learning and boundary correction that claims state-of-the-art RGB L1 and Alpha IoU on natural images.

  2. LiWi: Layering in the Wild

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    LiWi uses an agent-driven data synthesis pipeline to build the LiWi-100k dataset and a model with shadow-guided and degradation-restoration objectives that achieves SoTA performance on RGB L1 and Alpha IoU for natural...

  3. RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition

    cs.CV 2026-05 unverdicted novelty 7.0 of 10

    RevealLayer decomposes natural images into multiple RGBA layers using diffusion models with region-aware attention, occlusion-guided adaptation, and a composite loss, outperforming prior methods on a new benchmark dataset.

  4. A Unified and Controllable Framework for Layered Image Generation with Visual Effects

    cs.CV 2026-01 unverdicted novelty 7.0 of 10

    LASAGNA produces layered images with integrated visual effects in a single pass, enabling drift-free edits via alpha compositing while releasing a 48K dataset and a 242-sample benchmark.

  5. ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition

    cs.CV 2026-07 conditional novelty 6.0 of 10

    ReDesign turns screenshots into editable layer hierarchies by having a vision-language agent choose tools step by step and verify each split, beating layered-decomposition baselines on a new Figma edit-replay benchmark.

  6. MRT: Masked Region Transformer for Layered Image Generation and Editing at Scale

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Presents MRT, a 20B-parameter masked region diffusion model unifying text-to-layers, image-to-layers, and layers-to-layers tasks with an overflow-aware canvas layer for complete editable outputs.

  7. LiWi: Layering in the Wild

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Presents LiWi-100k dataset generated via agent-driven decomposition and a model achieving SoTA on RGB L1 and Alpha IoU for natural image layering.

  8. LimeCross: Context-Conditioned Layered Image Editing with Structural Consistency

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    LimeCross enables text-guided editing of individual layers in composite images by conditioning on cross-layer context via bi-stream attention while preserving layer integrity and introducing the LayerEditBench benchmark.

  9. RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    RaDL adds two attention modules to a Stable Diffusion based pipeline: Attribute Enhancement for per-instance attribute fidelity and Relation Attention that uses action verbs from the prompt to model inter-instance rel...

  10. Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    Visual generation models are evolving from passive renderers to interactive agentic world modelers, but current systems lack spatial reasoning, temporal consistency, and causal understanding, with evaluations overemph...

  11. LumiGen: An LVLM-Enhanced Iterative Framework for Fine-Grained Text-to-Image Generation

    cs.LG 2025-08 reject novelty 4.0 of 10

    An LVLM-driven iterative text-to-image framework whose claimed performance scores are explicitly labeled fictitious, so no empirical result is established.

Pith tools