REVIEW 4 cited by
Composition-aware Graphic Layout GAN for Visual-textual Presentation Designs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
In this paper, we study the graphic layout generation problem of producing high-quality visual-textual presentation designs for given images. We note that image compositions, which contain not only global semantics but also spatial information, would largely affect layout results. Hence, we propose a deep generative model, dubbed as composition-aware graphic layout GAN (CGL-GAN), to synthesize layouts based on the global and spatial visual contents of input images. To obtain training images from images that already contain manually designed graphic layout data, previous work suggests masking design elements (e.g., texts and embellishments) as model inputs, which inevitably leaves hint of the ground truth. We study the misalignment between the training inputs (with hint masks) and test inputs (without masks), and design a novel domain alignment module (DAM) to narrow this gap. For training, we built a large-scale layout dataset which consists of 60,548 advertising posters with annotated layout information. To evaluate the generated layouts, we propose three novel metrics according to aesthetic intuitions. Through both quantitative and qualitative evaluations, we demonstrate that the proposed model can synthesize high-quality graphic layouts according to image compositions.
Forward citations
Cited by 4 Pith papers
-
IGD: Instructional Graphic Design with Multimodal Layer Generation
IGD generates editable multi-layer graphic designs (posters, slides, stickers) from text instructions using an MLLM for layout and a diffusion model for image assets.
-
Rethinking Layered Graphic Design Generation with a Top-Down Approach
Accordion decomposes AI-generated raster designs into editable background, object, and vectorized text layers using a VLM-driven top-down planning pipeline.
-
ProCrop: Learning Aesthetic Image Cropping from Professional Compositions
ProCrop uses retrieved professional photos to guide image cropping and introduces a 242,000-image outpainted weakly labeled dataset, reporting state-of-the-art results on standard benchmarks.
-
CAL-RAG: Retrieval-Augmented Multi-Agent Generation for Content-Aware Layout Design
CAL-RAG reports state-of-the-art layout metrics on PKU PosterLayout by iteratively refining layouts with an agentic loop, but the perfect scores likely reflect direct optimization of the reported metrics.
Discussion (0). Sign in to comment.