Pith. sign in

REVIEW 2 cited by

UDiffText: A Unified Framework for High-quality Text Synthesis in Arbitrary Images via Character-aware Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.04884 v1 pith:MZVNCMBG submitted 2023-12-08 cs.CV

classification cs.CV
keywords textdiffusionimagemodelgenerationimagessynthesisudifftext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-to-Image (T2I) generation methods based on diffusion model have garnered significant attention in the last few years. Although these image synthesis methods produce visually appealing results, they frequently exhibit spelling errors when rendering text within the generated images. Such errors manifest as missing, incorrect or extraneous characters, thereby severely constraining the performance of text image generation based on diffusion models. To address the aforementioned issue, this paper proposes a novel approach for text image generation, utilizing a pre-trained diffusion model (i.e., Stable Diffusion [27]). Our approach involves the design and training of a light-weight character-level text encoder, which replaces the original CLIP encoder and provides more robust text embeddings as conditional guidance. Then, we fine-tune the diffusion model using a large-scale dataset, incorporating local attention control under the supervision of character-level segmentation maps. Finally, by employing an inference stage refinement process, we achieve a notably high sequence accuracy when synthesizing text in arbitrarily given images. Both qualitative and quantitative results demonstrate the superiority of our method to the state of the art. Furthermore, we showcase several potential applications of the proposed UDiffText, including text-centric image synthesis, scene text editing, etc. Code and model will be available at https://github.com/ZYM-PKU/UDiffText .

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CharGen: High Accurate Character-Level Visual Text Generation Model with MultiModal Encoder

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A character-level multimodal encoder plus a glyph-aware perceptual loss lets a diffusion model render text more accurately, beating prior methods by 6 to 9 points on English and Chinese benchmarks.

  2. AMO Sampler: Enhancing Text Rendering with Overshooting

    cs.CV 2024-11 conditional novelty 6.0 of 10

    An attention-modulated overshooting sampler for rectified flow models improves text rendering accuracy on SD3 and Flux without retraining or added inference cost.

Pith tools