Pith. sign in

REVIEW 4 cited by

Composition-aware Graphic Layout GAN for Visual-textual Presentation Designs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.00303 v3 pith:QPBFPNJP submitted 2022-04-30 cs.CV

classification cs.CV
keywords layoutgraphicimagesinputslayoutsmodeltrainingaccording
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we study the graphic layout generation problem of producing high-quality visual-textual presentation designs for given images. We note that image compositions, which contain not only global semantics but also spatial information, would largely affect layout results. Hence, we propose a deep generative model, dubbed as composition-aware graphic layout GAN (CGL-GAN), to synthesize layouts based on the global and spatial visual contents of input images. To obtain training images from images that already contain manually designed graphic layout data, previous work suggests masking design elements (e.g., texts and embellishments) as model inputs, which inevitably leaves hint of the ground truth. We study the misalignment between the training inputs (with hint masks) and test inputs (without masks), and design a novel domain alignment module (DAM) to narrow this gap. For training, we built a large-scale layout dataset which consists of 60,548 advertising posters with annotated layout information. To evaluate the generated layouts, we propose three novel metrics according to aesthetic intuitions. Through both quantitative and qualitative evaluations, we demonstrate that the proposed model can synthesize high-quality graphic layouts according to image compositions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. IGD: Instructional Graphic Design with Multimodal Layer Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    IGD generates editable multi-layer graphic designs (posters, slides, stickers) from text instructions using an MLLM for layout and a diffusion model for image assets.

  2. Rethinking Layered Graphic Design Generation with a Top-Down Approach

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Accordion decomposes AI-generated raster designs into editable background, object, and vectorized text layers using a VLM-driven top-down planning pipeline.

  3. ProCrop: Learning Aesthetic Image Cropping from Professional Compositions

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ProCrop uses retrieved professional photos to guide image cropping and introduces a 242,000-image outpainted weakly labeled dataset, reporting state-of-the-art results on standard benchmarks.

  4. CAL-RAG: Retrieval-Augmented Multi-Agent Generation for Content-Aware Layout Design

    cs.IR 2025-06 reject novelty 5.0 of 10

    CAL-RAG reports state-of-the-art layout metrics on PKU PosterLayout by iteratively refining layouts with an agentic loop, but the perfect scores likely reflect direct optimization of the reported metrics.

Pith tools