Pith. sign in

REVIEW 8 cited by

AutoStudio: Crafting Consistent Subjects in Multi-turn Interactive Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.01388 v3 pith:O7QYUR2K submitted 2024-06-03 cs.CV

classification cs.CV
keywords autostudiogenerationimagessubjectgenerateimageintroducemodels
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

As cutting-edge Text-to-Image (T2I) generation models already excel at producing remarkable single images, an even more challenging task, i.e., multi-turn interactive image generation begins to attract the attention of related research communities. This task requires models to interact with users over multiple turns to generate a coherent sequence of images. However, since users may switch subjects frequently, current efforts struggle to maintain subject consistency while generating diverse images. To address this issue, we introduce a training-free multi-agent framework called AutoStudio. AutoStudio employs three agents based on large language models (LLMs) to handle interactions, along with a stable diffusion (SD) based agent for generating high-quality images. Specifically, AutoStudio consists of (i) a subject manager to interpret interaction dialogues and manage the context of each subject, (ii) a layout generator to generate fine-grained bounding boxes to control subject locations, (iii) a supervisor to provide suggestions for layout refinements, and (iv) a drawer to complete image generation. Furthermore, we introduce a Parallel-UNet to replace the original UNet in the drawer, which employs two parallel cross-attention modules for exploiting subject-aware features. We also introduce a subject-initialized generation method to better preserve small subjects. Our AutoStudio hereby can generate a sequence of multi-subject images interactively and consistently. Extensive experiments on the public CMIGBench benchmark and human evaluations show that AutoStudio maintains multi-subject consistency across multiple turns well, and it also raises the state-of-the-art performance by 13.65% in average Frechet Inception Distance and 2.83% in average character-character similarity.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    VLAC-Cut-guided multi-robot HITL post-training reaches 80–95% success and 1.7–4.2× throughput over the base VLA, outperforming HITL-only under the same human budget.

  2. Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A zero-shot method combines grid priors, depth-conditioned ControlNet, and latent blending to generate and consistently edit story visualizations across multiple frames.

  3. Audit & Repair: An Agentic Framework for Consistent Story Visualization in Text-to-Image Diffusion Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A multi-agent system that audits story images with a vision-language model and repairs inconsistencies with targeted diffusion edits improves multi-panel consistency over existing story visualization methods.

  4. Unpaired Deblurring via Decoupled Diffusion Model

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A diffusion model that decouples structural features from blur patterns using unpaired target-domain images can deblur photos in unseen domains without paired training data.

  5. One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Concatenating all frame prompts into a single prompt, then reweighting singular values and re-anchoring cross-attention, yields training-free identity-consistent text-to-image generation.

  6. DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DiffSensei combines an SDXL diffusion generator with a multimodal LLM adapter and masked attention to generate manga pages with multiple characters whose poses and expressions follow panel captions.

  7. ETPDesigner: Multi-Agent Orchestration for Interactive Multimodal Electronic Theater Program

    cs.CV 2026-07 conditional novelty 5.0 of 10

    ETPDesigner automatically generates multi-page electronic theater programs from scripts using a multi-agent LLM pipeline with a global style anchor and interactive character chat.

  8. Improving Multi-Subject Consistency in Open-Domain Image Generation with Isolation and Reposition Attention

    cs.CV 2024-11 conditional novelty 5.0 of 10

    IR-Diffusion adds two attention masks, Isolation and Reposition, that stop subjects in an image from blending into each other and align reference features to target positions, improving multi-subject consistency witho...

Pith tools