{"id":"ab448599-2ec7-4b74-a373-25fa117fa104","arxiv_id":"2504.21682","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A survey of visual text processing that contributes VTPBench, a six-task benchmark, and VTPScore, an MLLM-based reference-augmented evaluation metric validated against human ratings.","lead":"This paper surveys visual text processing, organizing methods by text features and learning paradigms, and introduces VTPBench, a 4,305-sample benchmark across six tasks. It also proposes VTPScore, a GPT-4o-based evaluation metric, and tests 20+ models on it.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The human validation of VTPScore is statistically empty: no correlation, variance, or inter-rater agreement is reported, so the central claim of human-aligned 'fair and reliable evaluation' is unsupported.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: the human evaluation in Sec. 4.3 lacks statistical rigor, reporting only mean scores without correlation, variance, or inter-rater agreement. This is the most load-bearing issue because the paper's novel claim is that VTPScore enables 'fair and reliable evaluation' through human-aligned MLLM scoring. If that alignment is not demonstrated, the central contribution is unsubstantiated, regardless of the survey's usefulness. Other concerns—reproducibility of the proprietary GPT-4o pipeline and subjective VTPBench sample selection—are secondary and addressable, but they do not directly attack the metric's validity as the missing human-alignment evidence does. The paper's conditional verdict remains appropriate: the benchmark and survey are substantive, but the VTPScore claim needs additional statistical validation before acceptance. No change to the reader's verdict is required.","tokens_in":34595,"tokens_out":5763,"duration_ms":53122,"concrete_test":"Release the per-sample human ratings and compute, for each of the six VTPBench tasks, the sample-level Spearman/Pearson correlation between VTPScore and the human scores, along with the inter-rater reliability (ICC(2,1) or Krippendorff's alpha). Also clarify whether the 'HumanScore' column in Table 7 is the sum or the average of the two 0–5 ratings; if it is the average, rescale it to 0–10 and recompute the comparison. If any per-task correlation is below approximately 0.7, or if the scales are incompatible, then the claim that VTPScore is human-aligned fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central contribution of the paper is VTPScore as a unified, human-aligned evaluation metric. The only evidence for human alignment is in Sec. 4.3 (Table 7), where the authors state 'high consistency' after a human study with ten raters. However, Table 7 reports only mean HumanScore values per model; no correlation coefficient, no per-sample comparison, no standard deviation, and no inter-rater agreement (e.g., ICC or Krippendorff's alpha) are provided. With only 3–5 models per task, a model-level correlation across tasks would be dominated by task differences and could be spurious. Moreover, the reported HumanScore values raise a scale-compatibility issue: VTPScore is defined as the sum of two 0–5 scores (range 0–10), while the human ratings are described as 'average scores' on the same 0–5 scale. If HumanScore were the average of the two components (range 0–5), values like 8.20 for LEMMA would be impossible; if it is the sum, then TSRN's HumanScore of 3.58 versus VTPScore of 6.70 shows a large discrepancy that undermines the claimed consistency. Without a rigorous statistical demonstration that VTPScore tracks human judgment, the metric could be encoding GPT-4o's own biases rather than human perception, and the guarantee of 'fair and reliable evaluation' does not follow.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a survey-and-benchmark submission on visual text processing. It organizes the field into text image reconstruction (super-resolution, dewarping, enhancement) and manipulation (removal, editing, generation), reviews methods through the lens of text features (structure, stroke, semantics, style) and learning paradigms, and introduces VTPBench, a 4,305-sample benchmark spanning six tasks, together with VTPScore, an MLLM-based metric that sums visual-quality and text-readability scores. The authors evaluate more than 20 models on VTPBench and report a small human study that they interpret as showing high consistency between VTPScore and human judgment. The survey portion is broad and generally accurate, but the evaluation contribution's central validity claim is not statistically established in the current manuscript.","tokens_in":34770,"tokens_out":4612,"duration_ms":47087,"significance":"If the evaluation claims were fully supported, VTPBench and VTPScore would fill a real need: current visual text processing evaluations are fragmented across task-specific datasets and metrics, and a unified protocol with a human-aligned metric would be valuable to the community. The paper also contributes a useful feature-based taxonomy and covers a wide and recent literature, which is itself a service. The prompt design for the six tasks is explicit, and the choice to evaluate more than 20 models with official weights is a concrete effort toward reproducibility. However, the headline contribution is the metric, and the metric's human alignment and benchmark construction are not yet validated with sufficient rigor; the central claim of 'fair and reliable evaluation' therefore remains unsubstantiated. The missing analyses are feasible within the manuscript's scope.","major_comments":[{"comment":"The statement that VTPScore shows 'high consistency' with human evaluation is not supported by the reported data. Table 7 lists only mean HumanScore values; no correlation coefficient, per-sample comparison, standard deviation, or inter-rater agreement (ICC, Krippendorff's alpha, or similar) is given. With only three to five models per task, a pooled correlation across tasks would be confounded by task identity, and within-task correlations cannot be assessed from the table. The authors should report per-task and pooled within-task Spearman correlations between VTPScore and HumanScore, per-sample agreement, and inter-rater reliability; without these, the central claim of human-aligned evaluation remains unverified.","section":"Sec. 4.3, Table 7"},{"comment":"The scale of HumanScore is ambiguous and appears incompatible with VTPScore. VTPScore is defined as the sum of visual quality (0–5) and readability (0–5), giving a 0–10 range, while the human raters are described as rating both criteria on a 0–5 scale, which would give a 0–5 average. LEMMA's HumanScore of 8.20 cannot be a 0–5 average, while TSRN's HumanScore of 3.58 is implausibly low if it is a 0–10 sum given its VTPScore of 6.70. The paper must clarify whether HumanScore is a sum or an average, and rescale one set of scores so that VTPScore and HumanScore are compared in the same units.","section":"Sec. 4.2 and Sec. 4.3, Table 7"},{"comment":"The VTPBench selection protocol is described only as 'carefully choose some representative data' and 'filter out extremely broken or severely damaged samples.' This is too vague to establish that the benchmark is unbiased or reproducible. The authors should specify the source datasets for each task, the exact filtering criteria (e.g., detection or OCR confidence thresholds, manual review rules), the number of samples removed per source, and the sampling procedure; ideally the filtered sample identifiers should be released.","section":"Sec. 4.2, Data Construction"},{"comment":"The human validation is in part circular: participants were instructed using the same visual-quality and readability criteria that are given to GPT-4o, so agreement between HumanScore and VTPScore could reflect shared task framing rather than the metric's accuracy. To validate VTPScore as a fair metric, the authors should also compare it against existing reference metrics (PSNR/SSIM, OCR accuracy, FID) and report per-sample agreement between GPT-4o and human ratings, not only model-level means.","section":"Sec. 4.2 and Sec. 4.3"}],"minor_comments":[{"comment":"The sentence describing the Text Deblurring Dataset contains a duplicated typo: 'a cropped 300×300 patch.×300 pixel patch' should read 'a cropped 300×300 pixel patch.'","section":"Sec. 4.1"},{"comment":"The text contains typographical errors: 'trys to enhance controllability' should be 'tries to enhance controllability', and 'douple content and style learning' should be 'decouple content and style learning.'","section":"Sec. 3.6.2"},{"comment":"In the figure, the word 'Consturct' should be corrected to 'Construct.'","section":"Fig. 9"},{"comment":"The Pix2Pix row contains an extra numerical entry (10.2000) that is not aligned with the column headers; please realign the table.","section":"Table 4"},{"comment":"The column heading 'Easy Medium HardAverage↑' is missing spaces and the arrows appear inconsistently; please format the table for readability.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The survey portion is competent and the benchmark idea is timely, but the paper's headline contribution is the metric, and the metric's validation is currently too thin. I would not reject: the missing analyses are feasible and within scope. I would ask the editor to verify that the authors can provide the per-sample human ratings and benchmark selection details; without them, the revised paper would still not support the abstract's claim of fair and reliable evaluation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The survey part is the real contribution here: a broad, well-organized review of six visual text processing tasks, organized around a consistent taxonomy of text features (structure, stroke, semantics, style) and learning paradigms. It fills a genuine gap and will save people time. VTPBench and VTPScore are also new artifacts, and the empirical comparison of 20+ models on both legacy benchmarks and VTPBench is a lot of work that mostly looks careful. I'd send this to a serious referee on the strength of the survey alone.\n\nThe load-bearing flaw is the human validation of VTPScore (Sec. 4.3, Table 7). The paper claims \"high consistency\" between VTPScore and human evaluation, but reports only mean HumanScore per model. No correlation coefficient, no per-sample comparison, no standard deviation, no inter-rater agreement. With 3–5 models per task, any model-level correlation would be dominated by task differences anyway. The claim is not supported as written.\n\nThere's also a scale ambiguity that needs resolving. VTPScore is defined as the sum of two 0–5 scores, so its range is 0–10. The human raters are described as rating on a 0–5 scale, and their \"average scores\" are listed as HumanScore. If HumanScore is the average of the two sub-scores, then values like 8.20 for LEMMA are impossible. If it's a sum, then TSRN's 3.58 versus its VTPScore of 6.70 is a huge unexplained gap. Either way, the paper doesn't discuss it.\n\nSecondary issues: VTPBench sample selection is vague (\"carefully choose representative data,\" \"filter out extremely broken samples\"), which leaves room for selection bias. And the evaluator is GPT-4o, a proprietary model; the exact prompts and any scoring code need to be released for reproducibility. The prompt sketch in Fig. 9 is not enough. These are fixable.\n\nOn balance: the survey and benchmark are worth publishing, and the VTPScore idea is worth pursuing, but the evaluation contribution needs real work before the \"fair and reliable\" claim can be accepted. I'd send it to review with instructions to focus on the human validation and scale definition. The flaws are addressable, not fatal.","headline":"A genuinely useful survey and benchmark, but the human validation of VTPScore is statistically empty and the scale ambiguity in Table 7 needs fixing before the metric's claims hold.","tokens_in":35425,"tokens_out":3544,"would_cite":true,"duration_ms":37546,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single MLLM-based score now grades six visual text tasks.","keywords":["visual text processing","VTPBench","VTPScore","multimodal large language models","evaluation metric","text image reconstruction","text image manipulation","benchmark"],"falsifier":"Give a sample of VTPBench (say 30-50 images per task) to at least 20 independent raters, compute per-sample mean human scores, and compare them with VTPScore via rank correlation and per-task scatter; if Spearman correlation falls below roughly 0.7 or the top-ranked model changes under bootstrap resampling, the claim that VTPScore is human-aligned and reliable fails. A simpler targeted probe is to hold ground truth constant and perturb images by small blur, color shift, or text corruption; if VTPScore does not order those perturbations the way raters do, the metric is not tracking readability or quality.","tokens_in":34323,"feed_emoji":"🔤","tokens_out":4844,"duration_ms":46042,"temperature":0.7,"pith_summary":"This survey argues that visual text processing, from super-resolution and dewarping to removal, editing, and generation, can be reviewed and compared under one roof, and that current evaluations are too fragmented to tell which models actually work. To back that claim, the paper introduces VTPBench, a curated benchmark of 4,305 samples spanning six visual text tasks, and VTPScore, an evaluation metric that asks a multimodal large language model (GPT-4o) to rate each output on visual quality and visual text readability, each on a 0-5 scale, against a ground-truth reference. On more than 20 open-source models, VTPScore ranks current methods and, in the authors' reading, agrees with human ratings. The payoff if true is a standardized, human-aligned scale for comparing future visual text models, plus a clear statement that the state of the art still has substantial room to improve.","feed_headline":"Six visual-text tasks, one benchmark, one human-aligned score","feed_subtitle":"VTPBench and GPT-4o-based VTPScore rank 20+ models on quality and readability, showing current methods still lag.","key_machinery":"The load-bearing object is VTPScore, a reference-based MLLM evaluation prompt protocol built on GPT-4o. For each of the six tasks, the paper writes a task-specific prompt with two rubrics, visual quality and visual text readability; the model must return JSON scores from 0 to 5 for each, and VTPScore is their sum. The second piece is VTPBench, a curated 4,305-sample benchmark drawn from existing datasets, which gives the metric a fixed test bed across super-resolution, dewarping, enhancement, removal, editing, and generation. What the machinery does is replace task-specific pixel metrics and ad hoc OCR checks with a single comparison that tracks human notions of quality and readability.","core_discovery":"The paper's central discovery is that a single multimodal-language-model-based metric, VTPScore, can serve as a fair and reliable evaluation standard across six visually distinct text tasks. VTPScore decomposes every result into two numbers: a visual quality score (clarity, geometric correctness, style consistency, or region-specific artifact control, depending on the task) and a visual text readability score (whether the text in the output matches the reference text), each from 0 to 5, with the total being their sum. The metric is reference-based: GPT-4o sees the predicted image and the ground-truth image under task-specific prompts and answers in structured JSON. On VTPBench, the benchmark assembled from existing datasets, the scores place LEMMA on top for super-resolution, UVDoc for dewarping, DocRes for enhancement, ViTEraser for removal, TextCtrl for editing, and AnyText for generation; the authors report that these rankings align with their human study and that even the best models leave large gaps on several tasks.","pith_inferences":["Beyond the paper: VTPScore's success depends on the specific choice of GPT-4o; a cheaper or open-weights MLLM might not reproduce the same ordering, so the metric should be re-validated per base model, and the prompt templates could be published as a regression suite.","Beyond the paper: the same prompt-rubric idea transfers to other fine-grained image-manipulation families, such as object removal, general image inpainting, and face editing, where reference-based MLLM scoring could replace FID and pixel metrics.","Beyond the paper: because the human study averaged only ten raters with no variance or agreement statistics, a public leaderboard with per-sample human scores and inter-rater reliability would be a direct, low-cost extension that makes the alignment claim testable."],"forward_implications":["If VTPScore is accepted, future visual text models can be compared on one 0-10 scale across six tasks, ending the current situation where each task uses its own pixel metrics and OCR checks.","The reported rankings imply different bottlenecks per task: super-resolution and removal are comparatively mature, while enhancement, editing, and generation remain far from human-level output.","Because VTPScore is model-based and reference-based, it can be rerun on any new model without reimplementing per-task evaluation code, making benchmark updates cheap.","The finding that no single model dominates all six tasks strengthens the case for unified, generalist visual text models, as suggested in the paper's open-challenges section."],"supporting_citations":[{"why":"Supplies the premise that multimodal large language models can assess visual quality, the basis for using GPT-4o as the evaluator.","marker":"[26]"},{"why":"Supports the claim that large multimodal models can perceive visual quality in a way that tracks human judgment.","marker":"[27]"},{"why":"Identifies GPT-4o as the specific base model used to compute VTPScore.","marker":"[211]"},{"why":"Provides the real-world TextZoom paired data used for super-resolution evaluation and contributes samples to VTPBench.","marker":"[109]"},{"why":"Provides the DocUNet dewarping images and ground truth used in VTPBench and in the existing-benchmark comparison.","marker":"[178]"},{"why":"Contributes the SCUT-EnsText real-world scene text removal benchmark used in VTPBench and in the removal performance tables.","marker":"[129]"},{"why":"Supplies the Tamper scene-text-editing evaluation data, a component of VTPBench's editing task.","marker":"[142]"},{"why":"Provides the multilingual generation and editing evaluation data used by VTPBench and is one of the strongest baselines VTPScore ranks.","marker":"[150]"}],"fun_headline_variants":["One score to rank all visual-text tasks","VTPScore: a single metric for text image quality and readability","All visual-text tasks, one benchmark, one score","MLLM metric reveals big gaps across text tasks","Survey and benchmark: how 20+ models stack up on text"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"VTPScore is only as reliable as GPT-4o's agreement with human judgment, and the paper's evidence for that agreement is a ten-rater study that reports only mean scores without correlation, variance, or inter-rater reliability.","fun_headline_variants_meta":{"raw":{"variants":["One score to rank all visual-text tasks","VTPScore: a single metric for text image quality and readability","All visual-text tasks, one benchmark, one score","MLLM metric reveals big gaps across text tasks","Survey and benchmark: how 20+ models stack up on text"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000175,"raw_usage":{"total_tokens":1320,"prompt_tokens":1011,"completion_tokens":309,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":627,"completion_tokens_details":{"reasoning_tokens":229}},"tokens_in":627,"tokens_out":309,"duration_ms":3986,"temperature":1.0,"reasoning_tokens":229,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T04:55:21.868617+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give a sample of VTPBench (say 30-50 images per task) to at least 20 independent raters, compute per-sample mean human scores, and compare them with VTPScore via rank correlation and per-task scatter; if Spearman correlation falls below roughly 0.7 or the top-ranked model changes under bootstrap resampling, the claim that VTPScore is human-aligned and reliable fails. A simpler targeted probe is to hold ground truth constant and perturb images by small blur, color shift, or text corruption; if VTPScore does not order those perturbations the way raters do, the metric is not tracking readability or quality.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Identifies GPT-4o as the specific base model used to compute VTPScore."},{"cited_title":"DocUNet: Document image unwarping via a stacked U-Net,","cited_arxiv_id":null,"evidence_quote":"Provides the DocUNet dewarping images and ground truth used in VTPBench and in the existing-benchmark comparison."},{"cited_title":"Anytext: Multi- lingual visual text generation and editing,","cited_arxiv_id":null,"evidence_quote":"Provides the multilingual generation and editing evaluation data used by VTPBench and is one of the strongest baselines VTPScore ranks."}],"review_version":1}