Pith. sign in

REVIEW 8 cited by

StoryMaker: Towards Holistic Consistent Characters in Text-to-image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.12576 v1 pith:L7DMBNP3 submitted 2024-09-19 cs.CV

classification cs.CV
keywords charactersstorymakerconsistencycharacterfacialgenerationimagesmultiple
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Tuning-free personalized image generation methods have achieved significant success in maintaining facial consistency, i.e., identities, even with multiple characters. However, the lack of holistic consistency in scenes with multiple characters hampers these methods' ability to create a cohesive narrative. In this paper, we introduce StoryMaker, a personalization solution that preserves not only facial consistency but also clothing, hairstyles, and body consistency, thus facilitating the creation of a story through a series of images. StoryMaker incorporates conditions based on face identities and cropped character images, which include clothing, hairstyles, and bodies. Specifically, we integrate the facial identity information with the cropped character images using the Positional-aware Perceiver Resampler (PPR) to obtain distinct character features. To prevent intermingling of multiple characters and the background, we separately constrain the cross-attention impact regions of different characters and the background using MSE loss with segmentation masks. Additionally, we train the generation network conditioned on poses to promote decoupling from poses. A LoRA is also employed to enhance fidelity and quality. Experiments underscore the effectiveness of our approach. StoryMaker supports numerous applications and is compatible with other societal plug-ins. Our source codes and model weights are available at https://github.com/RedAIGC/StoryMaker.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FreeStory: Training-Free Character Consistency for Free-Form Visual Storytelling

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    FreeStory reformulates character consistency as entity-grounded feature reuse for free-form prompts, introduces FreeStoryBench, and reports stronger consistency than baselines among training-free methods.

  2. IDProtector: An Adversarial Noise Encoder to Protect Against ID-Preserving Image Generation

    cs.CV 2024-12 conditional novelty 6.0 of 10

    IDProtector adds imperceptible adversarial noise to a portrait in a single forward pass, disrupting identity-preserving generation by InstantID, IP-Adapter, IP-Adapter-Plus, and PhotoMaker.

  3. SerialGen: Personalized Image Generation by First Standardization Then Personalization

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A two-stage framework that first normalizes a reference human image and then applies personalized text-to-image generation improves whole-body appearance consistency while keeping strong prompt control.

  4. PersonaCraft: Personalized and Controllable Full-Body Multi-Human Scene Generation Using Occlusion-Aware 3D-Conditioned Diffusion

    cs.CV 2024-11 conditional novelty 6.0 of 10

    PersonaCraft adds SMPLx depth and normal conditioning, occlusion boundary enhancement, and occlusion-aware classifier-free guidance to diffusion models, enabling controllable multi-person images that preserve both fac...

  5. StorySync: Training-Free Subject Consistency in Text-to-Image Generation via Region Harmonization

    cs.CV 2025-07 unverdicted novelty 5.0 of 10

    A training-free inference-time pipeline uses masked cross-image attention sharing and region harmonization to keep subjects consistent across generated story images.

  6. AnyStory: Towards Unified Single and Multiple Subject Personalization in Text-to-Image Generation

    cs.CV 2025-01 conditional novelty 5.0 of 10

    AnyStory introduces a unified feed-forward approach for single and multi-subject text-to-image personalization using a simplified ReferenceNet and CLIP encoder, plus a decoupled instance-aware router.

  7. Improving Multi-Subject Consistency in Open-Domain Image Generation with Isolation and Reposition Attention

    cs.CV 2024-11 conditional novelty 5.0 of 10

    IR-Diffusion adds two attention masks, Isolation and Reposition, that stop subjects in an image from blending into each other and align reference features to target positions, improving multi-subject consistency witho...

  8. TaleForge: Interactive Multimodal System for Personalized Story Creation

    cs.CV 2025-06 conditional novelty 4.0 of 10

    TaleForge generates stories and illustrations in which the user's face becomes the main character, using Llama3 plus diffusion models and a 12-person user study showing stronger engagement.

Pith tools