Pith. sign in

REVIEW 10 cited by

AnyText2: Visual Text Generation and Editing With Customizable Attributes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.15245 v1 pith:ICNVYHLE submitted 2024-11-22 cs.CV

AnyText2: Visual Text Generation and Editing With Customizable Attributes

classification cs.CV
keywords textattributesanytext2generationmethodanytextapproachcontrol
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

As the text-to-image (T2I) domain progresses, generating text that seamlessly integrates with visual content has garnered significant attention. However, even with accurate text generation, the inability to control font and color can greatly limit certain applications, and this issue remains insufficiently addressed. This paper introduces AnyText2, a novel method that enables precise control over multilingual text attributes in natural scene image generation and editing. Our approach consists of two main components. First, we propose a WriteNet+AttnX architecture that injects text rendering capabilities into a pre-trained T2I model. Compared to its predecessor, AnyText, our new approach not only enhances image realism but also achieves a 19.8% increase in inference speed. Second, we explore techniques for extracting fonts and colors from scene images and develop a Text Embedding Module that encodes these text attributes separately as conditions. As an extension of AnyText, this method allows for customization of attributes for each line of text, leading to improvements of 3.3% and 9.3% in text accuracy for Chinese and English, respectively. Through comprehensive experiments, we demonstrate the state-of-the-art performance of our method. The code and model will be made open-source in https://github.com/tyxsspa/AnyText2.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. TYPO: Instruction-Dense Visual Jailbreaks against Commercial Closed-Source Image-Generation Models

    cs.CR 2026-07 conditional novelty 6.5

    Safety alignment fails to transfer from text replies to text-in-image, and TYPO’s dual-channel combinatorial search jailbreaks four commercial image models at >90% ASR for ~$0.04.

  2. InnoText: A Unified Model for Visual Text Generation and Editing

    cs.CV 2026-07 conditional novelty 6.0

    A unified DiT model with font-size-aware modulation and region-weighted loss outperforms existing visual text generation and editing systems on bilingual benchmarks.

  3. SteerVTE: Seamless Video Text Editing with Style and Glyph Control

    cs.CV 2026-06 unverdicted novelty 6.0

    SteerVTE adds lightweight style and dual-granularity glyph adapters to a frozen video diffusion model, introduces a glyph-aware loss and progressive training, and releases a 1M synthetic dataset to enable accurate vid...

  4. TextWand: A Unified Framework for Scene Text Editing

    cs.CV 2026-06 unverdicted novelty 6.0

    TextWand unifies scene text removal, generation and replacement via rendering/erasure decomposition, ORPE for layout fidelity, RAS for clean erasure, and the new TextWand-Bench dataset, claiming superior accuracy and ...

  5. Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning

    cs.CV 2026-05 unverdicted novelty 6.0

    A self-prompting MM-DiT model performs open-vocabulary scene text editing by extracting style and glyph information from the original image without extra encoders.

  6. Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning

    cs.CV 2026-05 unverdicted novelty 6.0

    Self-prompting diffusion transformer uses in-context learning on self-generated prompts from the image to achieve open-vocabulary scene text editing with style consistency.

  7. POCA: Pareto-Optimal Curriculum Alignment for Visual Text Generation

    cs.CV 2026-04 unverdicted novelty 6.0

    POCA combines Pareto optimization with curriculum alignment to improve multi-reward reinforcement learning for visual text generation without relying on weighted sums.

  8. SkyReels-Text: Fine-Grained Font-Controllable Text Editing for Poster Design

    cs.CV 2025-11 unverdicted novelty 6.0

    SkyReels-Text enables simultaneous fine-grained editing of multiple text regions in posters using arbitrary glyph patches for font control without labels or test-time fine-tuning.

  9. Evaluating Reasoning Fidelity in Visual Text Generation

    cs.CV 2026-06 unverdicted novelty 5.0

    T2I models frequently exhibit semantic errors, logical inconsistencies, and incorrect reasoning steps in visual text generation tasks, unlike text-only models.

  10. Wan-Image: Pushing the Boundaries of Generative Visual Intelligence

    cs.CV 2026-04 unverdicted novelty 3.0

    Wan-Image is a unified multi-modal system that integrates LLMs and diffusion transformers to deliver professional-grade image generation features including complex typography, multi-subject consistency, and precise ed...