{"id":"d5d61f96-9084-4520-a18f-3b878e617b31","arxiv_id":"2412.00114","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A pipeline that plans, places, and renders scene-coherent typographic adversarial text fools vision-language models more often than prior center or margin text attacks, but its success metric and naturalness evaluation are methodologically weak.","lead":"SceneTAP uses a large language model to plan where and how to insert misleading text into images, then renders that text with a diffusion model so it looks natural. It reports that this approach fools vision-language models such as ChatGPT-4o more often than previous typographic attacks, including after printing the text into real scenes.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 1's ASR is not restricted to initially-correct model answers, so the reported gains mix true flips with pre-existing target emissions; Section 5.1 never defines ASR this way and the nonzero No Attack rows make the omission material.","rationale":"The reader's weakest assumption is exactly the load-bearing concern I find: ASR is not conditioned on initially-correct responses, and the high No Attack baselines in Table 1 mean the reported improvements are not clean estimates of misdirection. This is not a dispute with an outside consensus; it is an internal metric inconsistency between the filtered pilot protocol in Section 3.1 and the main evaluation in Section 5.1. A good-faith reading also gives the paper real credit: the planner is training-free, code is released, results transfer across four LVLMs, and the physical demonstrations are concrete artifacts even though Section 5.3 reports only four qualitative cases. Appendix A.5's limitation about text-unfriendly scenes is a scope caveat, not a contradiction. One useful sanity check from the numbers: for the ChatGPT-4o LingoQA row, the 47.1% to 73.4% shift forces at least 49.7% of initially-correct examples to flip if all initially-target outputs remain target, or more if some initially-target outputs move away, so the underlying effect may survive; nevertheless, the reported percentage is not the flip rate. The correct fix is a straightforward reanalysis, so I retain the reader's conditional verdict rather than moving to accept or reject.","tokens_in":16748,"tokens_out":10676,"duration_ms":98860,"concrete_test":"Using the released SceneTAP code, recompute every cell of Table 1 after restricting to image-question pairs where the target LVLM answers correctly on the clean image; define conditional ASR as the fraction of those pairs whose post-attack output matches the adversarial target, and report the clean-image target-emission rate on the same subset for comparison. If conditional ASR deltas or the ordering over baselines change materially, revise the abstract and Section 5.2 to report the conditional numbers rather than the unconditional ones.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that SceneTAP's inserted text misleads LVLMs. That claim requires ASR to measure flips away from a correct clean-image answer. Section 5.1 defines ASR only as 'the percentage of successful attacks that deceive the target AI model' and does not state that clean-image answers must be correct. Table 1's No Attack rows are therefore nonzero and often large, e.g., 47.1% on LingoQA and 35.6% on VQAv2 for ChatGPT-4o, and 62-65% on LingoQA for the open-source models. Because the open-ended VQA target is generated by ChatGPT from the same image, question, and correct answer, it can coincide with what the target model already tends to output on clean images. Counting such outputs as 'attacks' overstates the causal role of the inserted text and inflates the paper's headline numbers, including the 47.19% to 62.10% average and the 47.1% to 73.4% LingoQA line for ChatGPT-4o. Section 3.1's pilot study explicitly filters to initially-correct responses, so the omission in Section 5.1 is an internal inconsistency rather than a reading artifact. Recomputing on initially-correct pairs is necessary before the 'misleads' claim is quantitatively supported; the method may still be effective, but the reported percentages are not a valid estimate of how often the attack changes a correct answer to the target.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes SceneTAP, a training-free LLM-based planner for generating scene-coherent typographic adversarial attacks against vision-language models. The method uses ChatGPT-4o to analyze an image and question, generate an adversarial text, choose a placement via Set-of-Mark prompting, and prompt TextDiffuser to insert the text into the scene; a revisable prompt refines the plan. The authors evaluate on TypoD-base, LingoQA, and VQAv2 across four LVLMs, comparing with Center and Margin attacks and reporting attack success rate (ASR), a ChatGPT-assessed naturalness score (N-Score), and a combined C-Score. They also demonstrate a physical-world extension by printing and pasting generated patches in four cases.","tokens_in":17093,"tokens_out":5186,"duration_ms":42777,"significance":"If the quantitative claims held, SceneTAP would be a useful contribution: it automates typographic attack design, is training-free, includes a physical attack demonstration, and releases code. The systematic study of adversarial text type and placement in Section 3 is a useful empirical addition. However, the current evaluation does not support the headline numbers: ASR is not conditioned on clean-image correctness, and the naturalness metric is assigned by the same model that generates the attacks. The relative ordering of methods may survive a corrected analysis, but the reported magnitudes and the 'misleads' claim need revision.","major_comments":[{"comment":"ASR is not conditioned on initially correct clean-image responses. Section 5.1 defines ASR as the 'percentage of successful attacks that deceive the target AI model' but does not require the clean image to be answered correctly. Consequently, Table 1's No-Attack rows are nonzero and often large: ChatGPT-4o LingoQA 47.1%, VQAv2 35.6%; LLaVA LingoQA 65.6%; MiniGPT-v2 LingoQA 62.1%; InstructBLIP LingoQA 62.9%. For open-ended VQA, the target answer is generated by ChatGPT from the same image, question, and correct answer, so it can coincide with what the victim model already outputs on the clean image. Counting such pre-existing outputs as 'attacks' overstates the causal role of the inserted text and inflates the reported gains, including the 47.19% to 62.10% average and the 47.1% to 73.4% LingoQA line for ChatGPT-4o. Section 3.1 explicitly filters to initially correct responses, so the omission is an internal inconsistency. Please recompute ASR on the subset where the clean image is answered correctly, or report flip rates and deltas relative to the no-attack baseline; this is necessary before the 'misleads' claim is quantitatively supported.","section":"Section 5.1, Table 1"},{"comment":"The N-Score is assigned by ChatGPT-4o, which is the same model used as the planner (Section 4.5) and is also one of the victim models in Table 1. This creates a same-model evaluation loop for the naturalness claim: the model that designs the attack also judges its visual naturalness, and the C-Score in Section 5.1 inherits this loop. Since the paper's claim of maintaining visual naturalness rests on these scores, please provide an independent human evaluation or a different judge model, with agreement statistics, and separate the planner model from the evaluator model.","section":"Section 5.1, Section A.2, Section 4.5"},{"comment":"The physical-world evidence is anecdotal. Section 5.3 presents only four cases, with no physical attack success rate, no quantitative comparison between physical and digital success, and no details on repeat trials, camera viewpoints, or lighting conditions. The abstract's claim that the method remains effective 'even after capturing new images of physical setups' is therefore not quantitatively established. Please add a protocol and numbers for the physical experiments, even if on a modest scale.","section":"Section 5.3, Figure 4"},{"comment":"All reported ASR, N-Score, and C-Score values are single point estimates without error bars, multiple runs, or significance tests. Since the planner is a stochastic LLM and some evaluation subsets are small (e.g., 100 image-question pairs in the Section 3.1 study, 500 VQAv2 pairs in Section 5.1), the differences between methods may not be stable. Please report variation across repeated runs or clearly state the sample sizes and any significance measures.","section":"Table 1"}],"minor_comments":[{"comment":"The C-Score is described as averaging the ASR and N-Score, but ASR is on a 0-100 scale and N-Score is on a 0-10 scale; the table values imply C-Score = (ASR + 10 * N-Score) / 2. Please state the scaling explicitly.","section":"Section 5.1"},{"comment":"The text contains a typo: 'What action should be taked for the car' should read 'What action should be taken for the car'.","section":"Abstract/Introduction"},{"comment":"The revisable prompt is shown in a box but the paper does not specify how it is invoked or how the model decides whether to modify the plan. Please describe the inference procedure more concretely.","section":"Section 4.5"},{"comment":"The ablation settings 'Plan1' and 'Plan2' are defined only in the table caption; please define them in the main text before the ablation discussion.","section":"Table 2"},{"comment":"The SoM mask-filtering ratio 'a' is set to different values per dataset but no sensitivity analysis is provided for this free parameter.","section":"Supplementary A.3"}],"recommendation":"major_revision","confidential_remarks":"The ASR-conditioning issue is the main substantive problem and is fixable; I recommend major revision rather than rejection. The authors should also avoid using ChatGPT-4o for both attack generation and naturalness scoring, or at least provide an independent evaluation. The relative ordering of methods may survive the fix, but the reported magnitudes will likely shrink."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on SceneTAP. The pipeline is genuinely new: using an LLM to plan the text, placement, and rendering instructions, then TextDiffuser to paint it in, and then printing the patches in the real world is a real step beyond the center and margin attacks. The ablation study is a plus—it shows each component contributes. The paper deserves credit for the idea and for the empirical placement analysis.\n\nBut the quantitative core has a real flaw, and the stress-test note lands. Section 5.1 defines ASR as 'the percentage of successful attacks that deceive the target AI model' without conditioning on the clean-image answer being correct. The No Attack rows are high, especially on open-ended VQA: 47.1% on LingoQA for ChatGPT-4o, 62–65% on LingoQA for the open-source models. So a large chunk of what the paper counts as attack success is the model naturally emitting the target answer before any text is added. Section 3.1's pilot explicitly filters to initially-correct responses; Section 5.1 doesn't. That's an internal inconsistency, and it means the headline numbers—44.32% ASR on two-choice, 62.10% on open-ended, the LingoQA 47.1 to 73.4 jump—are not a clean estimate of how often SceneTAP flips a correct answer to the target. The relative ordering may survive re-computation, but the absolute claims are unsupported as stated.\n\nThe naturalness score is also weak: N-Score is assigned by ChatGPT-4o, the same model used as planner and also one of the attack targets. That's a same-model evaluation loop. They need human raters or an independent model. And the physical-world section is four qualitative cases with no measured success rate; fine as a proof of concept, not as a result.\n\nOn the plus side, the method is training-free, the code is promised, and the ablation is informative. The flaws are in the evaluation rather than in the architecture. This is a fixable paper, but the authors need to recompute ASR on initially-correct examples, report variances, and get an independent naturalness judge.\n\nFor peer review: yes, send it. The idea is strong enough for referee time. But I'd ask for major revision on the evaluation before acceptance. I wouldn't cite it yet; I'd wait for the corrected numbers.","headline":"SceneTAP's pipeline is a real advance in typographic attack design, but its headline ASR numbers mix true flips with pre-existing target answers and need re-computation before the central claim is supported.","tokens_in":17569,"tokens_out":3172,"would_cite":false,"duration_ms":25524,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Scene-coherent text tricks vision-language models.","keywords":["typographic attack","adversarial attack","vision-language model","large language model planning","scene coherence","TextDiffuser","physical-world attack","visual question answering"],"falsifier":"Recompute attack success rate only on image-question pairs where each model answered correctly before any text insertion; if SceneTAP's margin over the no-attack condition collapses toward the margin of center placement, the claim that scene-coherent planning causes the misdirection is falsified.","tokens_in":16598,"feed_emoji":"🖼️","tokens_out":4991,"duration_ms":43656,"temperature":0.7,"pith_summary":"This paper tries to establish that typographic adversarial attacks on vision-language models can be made scene-coherent and fully automated: an LLM decides what misleading text to add, where to put it, and how to render it so the result looks natural in the image. The authors argue that prior typographic attacks fail on three counts — fixed text, fixed placement, and visibly artificial insertion — and that replacing all three decisions with LLM-based planning plus a text-rendering diffusion model fixes them. If the claim holds, the practical risk of such attacks rises because they survive printing and placement in physical scenes and still mislead current models, including ChatGPT-4o. The paper reports raising average attack success from 12.36% to 44.32% on two-choice questions and from 47.19% to 62.10% on open-ended VQA, with the highest naturalness and combined scores among compared methods.","feed_headline":"Scene-coherent text fools vision-language models","feed_subtitle":"New method plants misleading words that look natural, raising attack success on ChatGPT-4o and open models.","key_machinery":"The load-bearing mechanism is the LLM-driven planning loop paired with a scene-coherent text renderer. TextDiffuser is a diffusion model that renders short strings into images following a text prompt; SceneTAP uses the LLM to generate that prompt, so the inserted text follows the surface, lighting, and perspective of the chosen region. Set-of-mark prompting supplies a numbered segmentation map that lets the LLM refer to concrete image regions when choosing placement. The revisable prompt acts as a correction pass, moving text near the target region without covering the attribute asked about. Together these components convert an arbitrary image-question pair into a natural-looking typographic attack.","core_discovery":"The central discovery is that the content, placement, and visual rendering of an adversarial text can be planned jointly by a general-purpose LLM rather than fixed by a human or a rigid rule. Given the image, question, and correct answer, SceneTAP first analyzes the scene through chain-of-thought reasoning, selects a short incorrect answer that is plausible in context, uses set-of-mark prompting to pick a region near the question-targeted object, and produces a natural-language instruction for a TextDiffuser model to paint the text onto that surface. A revisable prompt lets the planner adjust placement when the chosen spot would alter an attribute central to the question or sit on an unrealistic surface. The resulting digital images, and printed physical versions of them, shift model answers toward the planted text while scoring higher on the paper's naturalness metric than center or margin insertion.","pith_inferences":["An editorial caution: because the reported ASR counts target-matching answers even when no attack was applied, re-evaluating on only initially-correct responses would likely shrink the reported gains; a fair comparison should condition on that subset.","A testable extension is to run SceneTAP on images with no natural text surfaces, such as open landscapes; the paper's own limitation note predicts ASR and naturalness would drop, which would quantify the cost of the scene-coherence constraint.","The same planner could be inverted as a defense generator: synthesize realistic misleading text to fine-tune LVLMs to ignore contextually plausible but physically absent text, or to cross-check OCR output against scene semantics.","Physical deployment, while demonstrated, makes the attack static and detectable by repeated observation over time; a dynamic variant would need to re-plan text for changing scenes."],"forward_implications":["SceneTAP raises attack success rate above both center and margin baselines on two-choice and open-ended VQA, across LLaVA, InstructBLIP, MiniGPT-v2, and ChatGPT-4o.","Because the inserted text is rendered to match the scene, the attacks receive higher naturalness scores and remain effective when printed and photographed in physical environments.","Ablation results attribute the gain to all three planning decisions: question-relevant adversarial text, placement near the question-targeted region, and diffusion-based insertion.","The method exposes a vulnerability in current LVLMs that do not distinguish genuinely present scene text from adversarial planted text, suggesting defenses must check text plausibility beyond surface appearance."],"supporting_citations":[{"why":"Supplies the TypoD-base dataset and the Center Attack baseline used for two-choice comparisons.","marker":"[1]"},{"why":"Supplies the Margin Attack baseline and earlier evidence that text placement affects attack strength.","marker":"[2]"},{"why":"TextDiffuser-2 is the scene-coherent renderer that executes the LLM's text-placement instruction.","marker":"[20]"},{"why":"Set-of-mark prompting provides the numbered segmentation map the planner uses to choose placement regions.","marker":"[43]"},{"why":"LLaVA is one of the four victim models whose ASR is measured across all datasets.","marker":"[41]"},{"why":"VQAv2 supplies the 500 image-question pairs used for open-ended VQA evaluation.","marker":"[42]"},{"why":"LingoQA supplies autonomous-driving VQA questions used to test the attack's real-world relevance.","marker":"[45]"}],"fun_headline_variants":["SceneTAP: natural text that fools AI vision","LLM plans scene-coherent text to trick VLMs","Real-world adversarial text misleads GPT-4o","Coherent text attacks outperform center placement","Printed scene-coherent text fools vision-language models"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported attack success counts any output that matches the target answer as a success, even when the model already gave that answer before any text was added, so the attack's causal contribution is not isolated from the model's pre-existing tendency.","fun_headline_variants_meta":{"raw":{"variants":["SceneTAP: natural text that fools AI vision","LLM plans scene-coherent text to trick VLMs","Real-world adversarial text misleads GPT-4o","Coherent text attacks outperform center placement","Printed scene-coherent text fools vision-language models"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000204,"raw_usage":{"total_tokens":1416,"prompt_tokens":997,"completion_tokens":419,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":613,"completion_tokens_details":{"reasoning_tokens":345}},"tokens_in":613,"tokens_out":419,"duration_ms":4651,"temperature":1.0,"reasoning_tokens":345,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T10:43:30.050027+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute attack success rate only on image-question pairs where each model answered correctly before any text insertion; if SceneTAP's margin over the no-attack condition collapses toward the margin of center placement, the claim that scene-coherent planning causes the misdirection is falsified.","supporting_citations":[{"cited_title":"Un- veiling typographic deceptions: Insights of the typographic vulnerability in large vision-language model.arXiv","cited_arxiv_id":null,"evidence_quote":"Supplies the TypoD-base dataset and the Center Attack baseline used for two-choice comparisons."},{"cited_title":"Textdiffuser-2: Unleashing the power of language models for text rendering","cited_arxiv_id":null,"evidence_quote":"TextDiffuser-2 is the scene-coherent renderer that executes the LLM's text-placement instruction."},{"cited_title":"Visual instruction tuning, 2023","cited_arxiv_id":null,"evidence_quote":"LLaVA is one of the four victim models whose ASR is measured across all datasets."},{"cited_title":"Making the v in vqa matter: Elevating the role of image understanding in visual question answering","cited_arxiv_id":null,"evidence_quote":"VQAv2 supplies the 500 image-question pairs used for open-ended VQA evaluation."}],"review_version":1}