Pith. sign in

REVIEW 8 cited by

Visual Text Processing: A Comprehensive Review and Unified Evaluation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.21682 v2 pith:VLX4NSKR submitted 2025-04-30 cs.CV

Visual Text Processing: A Comprehensive Review and Unified Evaluation

classification cs.CV
keywords textvisualprocessingmodelsevaluationadvancementscomprehensiveeffectively
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Visual text is a crucial component in both document and scene images, conveying rich semantic information and attracting significant attention in the computer vision community. Beyond traditional tasks such as text detection and recognition, visual text processing has witnessed rapid advancements driven by the emergence of foundation models, including text image reconstruction and text image manipulation. Despite significant progress, challenges remain due to the unique properties that differentiate text from general objects. Effectively capturing and leveraging these distinct textual characteristics is essential for developing robust visual text processing models. In this survey, we present a comprehensive, multi-perspective analysis of recent advancements in visual text processing, focusing on two key questions: (1) What textual features are most suitable for different visual text processing tasks? (2) How can these distinctive text features be effectively incorporated into processing frameworks? Furthermore, we introduce VTPBench, a new benchmark that encompasses a broad range of visual text processing datasets. Leveraging the advanced visual quality assessment capabilities of multimodal large language models (MLLMs), we propose VTPScore, a novel evaluation metric designed to ensure fair and reliable evaluation. Our empirical study with more than 20 specific models reveals substantial room for improvement in the current techniques. Our aim is to establish this work as a fundamental resource that fosters future exploration and innovation in the dynamic field of visual text processing. The relevant repository is available at https://github.com/shuyansy/Visual-Text-Processing-survey.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Beyond Detection: A Structure-Aware Framework for Scene Text Tracking

    cs.CV 2026-05 unverdicted novelty 7.0

    SymTrack is the first systematic detection-free framework for scene text tracking that constructs benchmarks from video text spotting datasets and reports up to 11.97% AUC gains over prior trackers.

  2. MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing

    cs.CV 2026-05 unverdicted novelty 7.0

    MULTITEXTEDIT benchmark reveals that all tested text-in-image editing models show pronounced degradation on non-English languages, especially Hebrew and Arabic, mainly in text accuracy and script fidelity.

  3. SlerpFlow: Spherical Trajectory Correction for Rectified Flow Inversion

    cs.CV 2026-07 reject novelty 6.0

    SlerpFlow replaces Euclidean solver steps with spherical-linear-interpolation (slerp) direction correction for rectified-flow inversion, reporting improved FLUX reconstruction and editing on PIE-Bench.

  4. TextWand: A Unified Framework for Scene Text Editing

    cs.CV 2026-06 unverdicted novelty 6.0

    TextWand unifies scene text removal, generation and replacement via rendering/erasure decomposition, ORPE for layout fidelity, RAS for clean erasure, and the new TextWand-Bench dataset, claiming superior accuracy and ...

  5. StyleTextGen: Style-Conditioned Multilingual Scene Text Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    StyleTextGen proposes a dual-branch style encoder, text style consistency loss, and mask-guided inference to achieve superior style consistency and cross-lingual performance in multilingual scene text generation on a ...

  6. EpiAgent: An Agent-Centric System for Ancient Inscription Restoration

    cs.CV 2026-04 unverdicted novelty 6.0

    EpiAgent is a new agent-centric system that restores degraded ancient inscriptions with better quality and generalization than prior rigid AI methods by using an LLM planner to coordinate multimodal tools and iterativ...

  7. EpiAgent: An Agent-Centric System for Ancient Inscription Restoration

    cs.CV 2026-04 conditional novelty 5.0

    An LLM central planner using Observe–Conceive–Execute–Reevaluate with specialized diffusion tools restores degraded Chinese inscriptions better than fixed pipelines on CIRI synthetic and real tests.

  8. UniTranslator: A Unified Multi-modal Framework for End-to-end In-Image Machine Translation

    cs.CV 2026-06 unverdicted novelty 4.0

    UniTranslator adds an Understand-Generation Alignment Module and Spatial Mask Decoder to a unified multimodal model to fix translation inconsistency and spatial misalignment in in-image machine translation, reporting ...