Pith. sign in

REVIEW 2 cited by

The Chosen One: Consistent Characters in Text-to-Image Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.10093 v4 pith:HDUZMAON submitted 2023-11-16 cs.CV cs.GRcs.LG

classification cs.CVcs.GRcs.LG
keywords consistentgenerationidentitymodelsapplicationscharactercharactersimages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in text-to-image generation models have unlocked vast potential for visual creativity. However, the users that use these models struggle with the generation of consistent characters, a crucial aspect for numerous real-world applications such as story visualization, game development, asset design, advertising, and more. Current methods typically rely on multiple pre-existing images of the target character or involve labor-intensive manual processes. In this work, we propose a fully automated solution for consistent character generation, with the sole input being a text prompt. We introduce an iterative procedure that, at each stage, identifies a coherent set of images sharing a similar identity and extracts a more consistent identity from this set. Our quantitative analysis demonstrates that our method strikes a better balance between prompt alignment and identity consistency compared to the baseline methods, and these findings are reinforced by a user study. To conclude, we showcase several practical applications of our approach.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Concatenating all frame prompts into a single prompt, then reweighting singular values and re-anchoring cross-attention, yields training-free identity-consistent text-to-image generation.

  2. Nested Attention: Semantic-aware Attention Values for Concept Personalization

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Nested Attention replaces a subject token's cross-attention value with a query-dependent value computed by an inner attention layer over image tokens, improving identity preservation and prompt adherence.

Pith tools