{"id":"7a9eff91-0527-4b89-b777-f40eb31cd4b9","arxiv_id":"2608.04244","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new benchmark shows that conflicting scene text degrades MLLM geolocation and shifts predictions toward the injected location in every tested model.","lead":"The paper presents SIGNPOST-Bench, a benchmark that edits the text on signs in photos to measure when AI image models trust the text over the visual scene. Across 20 models, conflicting sign text raised median geolocation error 4.8 times and often pulled predictions toward the written location.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Directed-shift claim lacks a Random-condition null: TFR/TDR are never computed for Random variants against the same injected targets.","rationale":"READING: The paper is a systematic, large-scale benchmark with careful construction and honest reporting of limitations. The degradation claim (Original vs Adversarial WLA) is supported by multiple metrics and by the Similar-vs-Blank comparison, which shows the editing pipeline can improve rather than merely damage predictions. The directed-shift claim is the more distinctive contribution, and it rests on comparing Adversarial predictions to the injected target. The Random condition is the natural control for whether the shift is caused by the specific conflicting text semantics or by a more general influence of inserted text or model priors. The paper's failure to report this control is a concrete, addressable gap, not a demonstrated error. Because the missing analysis could either strengthen or weaken the claim, the appropriate verdict remains conditional: accept only after the authors supply the Random-null comparison or a permutation test. This aligns with the reader's overall CONDITIONAL verdict, though the specific weakest assumption differs (edit artifacts vs. missing null control).","tokens_in":27581,"tokens_out":7948,"duration_ms":66927,"concrete_test":"Recompute TFR and TDR using the Random-variant prediction in place of the Adversarial prediction, keeping the same per-group adversarial target coordinates and the same Blank reference; macro-average by dataset as in Tables 7 and 17 and compare to the reported Adversarial values. Also compute a permutation baseline by reassigning trap coordinates across groups within each dataset and recomputing Adversarial TFR; if Random TFR or the permuted TFR is close to the reported 6.5–20.1% range, the directed-shift claim needs to be weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"CENTRAL CLAIM AND GAP: The headline directed-shift result is that 6.5–20.1% of adversarial predictions fall within 50 km of the injected target and every model has positive mean paired TDR (Section 6, Tables 7 and 17). These numbers are computed only for the Adversarial variant against the geocoded target of its own injected text. The benchmark's Random condition replaces text with unrelated, non-geographic strings (Section 3; Appendix A) and is described as a generic text sensitivity control, yet the paper never reports the distance from Random-variant predictions to the same adversarial target coordinates. Therefore the directed-shift claim is not separated from a generic 'any readable text biases the model toward a plausible location' effect. If Random predictions also fall within 50 km of the adversarial targets at comparable rates, or if Random-vs-Blank TDR is also positive for many models, then the observed trap-following could be due to text presence or prior attraction rather than to the semantic conflict. The paper's TBS/WLA results for Random do show distinct aggregate degradation, but TBS is relative to ground truth, not to the trap, and does not test directionality. The Random null is the benchmark's own control for this question, and its omission is the most load-bearing gap for the central directed-shift claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces SIGNPOST-Bench, a counterfactual benchmark for studying how multimodal large language models arbitrate between scene text and visual evidence in the continuous output space of geographic coordinates. Each of 5,111 source images becomes a quintuplet of matched variants (Original, Blank, Similar, Random, Adversarial) by locally removing or replacing selected scene-text spans, and 20 MLLMs from seven providers are evaluated on all variants, producing 511,100 model-image evaluations. The headline claims are: (i) conflicting scene text degrades geolocation, raising median error from 282 km to 1,347 km (4.8x) and lowering mean WLA from 47.11 to 29.89; (ii) compatible, unrelated, and conflicting text replacements produce distinct and ordered effects relative to the Blank baseline via TBS; (iii) adversarial text induces directed shifts toward the injected target, with 6.5-20.1% of geocodable adversarial predictions within 50 km of the target and positive mean paired Trap Distance Reduction for all 20 models; and (iv) capability and conflict robustness separate (C vs. R scores). The paper also reports a two-model probing/defense analysis and a scene-text coupling taxonomy (T1/T2/T3) with disclosed inter-annotator agreement.","tokens_in":27822,"tokens_out":26042,"duration_ms":211184,"significance":"If the directed-shift claim survives the missing control discussed below, this is a strong and useful benchmark. Concrete strengths: the scale (511,100 model-image pairs); the breadth (20 models, seven providers, four datasets); the public release of metadata and code (CC-BY-4.0/MIT) with deterministic reconstruction instructions; the disclosed audit protocols; and the sensitivity analyses, with MCRS rankings stable under weight/exponent variations (Kendall tau of at least 0.905) and component ablations (tau of at least 0.947), plus alpha and trap-radius sweeps. The five-condition paired design with a text-ablated Blank reference is methodologically sound for the degradation claim, and the continuous-coordinate output space is a genuine operational improvement over discrete VQA/classification tasks because it measures both magnitude and direction of text-induced shifts. The headline numbers are direct measurements rather than fitted quantities, so circularity is not a concern, and the released data allow independent re-testing of the paper's falsifiable predictions (for example, that every evaluated model degrades under adversarial text).","major_comments":[{"comment":"The stress-test concern is valid and lands on this exact section: the directed-shift claim is supported only by TFR and TDR computed for Adversarial predictions against the same group's geocoded trap, and no null control is reported. The Random condition shares the same visual substrate and editing pipeline, its predictions already exist for every model, and the paper itself describes it (Section 3; Figure 1) as the generic text sensitivity control, yet the distance from Random-variant predictions to the same 1,732 trap coordinates is never computed. This is not a cosmetic omission: the traps are by construction (Appendix A prompt; Section 4 geocoding) real place names that Nominatim resolves, often to well-known cities or landmarks, so population priors can place predictions near the trap without the text naming it; moreover, the paper's own TBS results (Section 6) show that Random text has a substantial generic effect (959 km mean error increase). TBS and WLA are computed relative to ground truth and cannot test directionality, so the missing Random-condition TFR and TDR against the same traps is the correct control. I request, on the same 1,732 geocodable groups: Random TFR (fraction of Random predictions within 50 km of the adversarial trap), the paired difference TDR_Random = D(blank, trap) - D(random, trap), and a permutation null that reassigns each group's trap to another group. If Adversarial TFR/TDR are not substantially and consistently larger than these baselines, the abstract's claim of directed shifts toward geographic targets introduced by conflicting text should be replaced by a weaker, baseline-quantified claim.","section":"Section 6; Eqs. (3)-(4); Tables 7 and 17"},{"comment":"The claim that every evaluated model exhibits a positive mean paired Trap Distance Reduction is reported without uncertainty quantification, although Table 17 shows attraction rates as low as 44.5% (Claude-Sonnet-4.6) and median TDR at or near zero for several models (Gemini-2.5-Pro median -0.01 km; Claude-Sonnet-4.6 median 0.00 km), with the text itself attributing the positive means to concentrated large shifts. For a headline claim of this form, please report per-model bootstrap 95% confidence intervals or a paired Wilcoxon signed-rank test on TDR, and, once the Random baseline of the previous comment is available, a paired test of the Adversarial-versus-Random TDR contrast. The four qualitative examples in Table 19 illustrate large effects but cannot substitute for a distributional statement over the 1,732 geocodable groups.","section":"Section 6; Table 17"},{"comment":"The counterfactual premise that the five variants differ only in text semantics rests on a human audit of 120 edited images, about 0.6% of the 20,444 edited variants, sampled only from Similar, Random, and Adversarial (not Blank), with artifact severity, context damage, and naturalness aggregated across conditions (1.32 +/- 0.78, 1.14 +/- 0.52, 4.00 +/- 1.26) and readability reported globally (12.5% partially readable, none unreadable). As reported, one cannot verify that Adversarial edits, the condition carrying the headline claims, are not systematically more artifact-heavy or less legible than Similar or Random edits, which would confound TBS, TFR, and TDR independently of text meaning. Please report these statistics per condition (including Blank), give the readability breakdown per condition, and, if feasible, enlarge the audit with a focus on Adversarial variants.","section":"Section 4 (Quality Assurance); Appendix B"}],"minor_comments":[{"comment":"The abstract's range '6.5-20.1% of adversarial predictions lie less than 50 km from the injected target' is the equal-dataset macro-average over four datasets whose per-dataset TFR values differ by more than an order of magnitude (for example, GoogleSV 1.66% for Gemini-3.1-Pro versus YFCC4K 35.68% for Qwen3-VL-235B); please state at the headline that the range is a macro-average and show the dataset spread, and add a Table 1 footnote noting that TFR and TDR are computed only on the 33.9% geocodable cohort of Table 6.","section":"Abstract; Table 7"},{"comment":"The sentence stating that Similar replacements reduce error by 379 km on average, whereas Random and Adversarial replacements increase error by 959 km and 1,577 km, respectively, should state the aggregation convention (mean TBS across models and datasets) and the sign convention of Eq. (2) inline, since negative TBS means the edited prediction is closer to ground truth than the Blank prediction.","section":"Section 6 (Semantic intervention effects)"},{"comment":"Because the injected target is defined as the top-ranked Nominatim result (Section 4), multi-reference place names can yield trap coordinates different from the intended injected place; please report how many of the 1,732 geocodable targets are ambiguous and confirm, as the qualitative examples in Table 19 suggest, that the intended and resolved targets coincide in the overwhelming majority of cases.","section":"Section 4 (Adversarial target geocoding); Table 6"},{"comment":"The trap-radius sensitivity sweep (Table 18) is reported only for the T3 subset of two representative models; given that TFR is a headline metric, either extend the sweep to additional models or justify why the T3 two-model analysis is representative of the all-tier macro-average reported in Table 7.","section":"Appendix H (Table 18)"},{"comment":"In Table 2, the 'Removed' column counts all losses after OCR selection, including pre-generation screening and post-generation cleanup; a one-sentence clarification in the caption would help readers avoid misreading it as a peculiarity of the synthesis stage.","section":"Table 2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for the journal and the benchmark is a substantial resource. The decisive editorial question is the directed-shift claim: the missing Random-condition control (major comment 1) is inexpensive to compute from the already-existing predictions and should be obtainable in a routine revision. I do not see a circularity problem, since the headline metrics are direct measurements. Two smaller editorial notes: the two-model probing/defense study is explicitly preliminary in the appendix and should not gain weight in the main text, and the audit sample, while honestly disclosed, would benefit from the per-condition stratification requested in major comment 3."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is a real contribution: a paired counterfactual quintuplet benchmark for MLLM text-vision conflict, using geolocation as a continuous output space. It is carefully built and large-scale: 5,111 groups, 25,555 variants, 20 models, 511k evaluations. Second, the central directed-shift claim—that conflicting text pulls predictions toward an injected target—is plausible but not fully isolated, because the Random condition is never used as a null for TFR/TDR.\n\nThe design is genuinely new relative to ConText-VQA and RIO-Bench: same-scene paired variants in a shared coordinate space, with Blank as an ablative baseline, and both degradation (WLA/TBS) and directionality (TFR/TDR) measured. The main findings—4.8x median error increase, 6.5–20.1% trap-fit, positive mean TDR for every model—are direct empirical measurements, not fitted parameters. The sensitivity analyses are honest (Kendall tau >= 0.905), the taxonomy stratification is sensible, and the benchmark data and code are public.\n\nThe load-bearing gap is the missing Random-condition null. TDR and TFR are computed for Adversarial vs Blank, but never for Random vs Blank against the same injected target. If random text also moves predictions nearer the target—or if a meaningful fraction of Random predictions fall within 50 km—then the 'semantic conflict' story is weakened. The paper's own Random condition is the right control; the omission is easy to fix, and I'd want it before the directed-shift claim is stated as strongly as it is. Two smaller issues: TFR/TDR use only the 33.9% geocodable cohort with per-dataset rates from 26.8% to 49.7%, so selection effects are possible; and each model-image pair is evaluated once, so sampling noise is unquantified. The synthesized variants aren't redistributed, but source IDs and reconstruction instructions are provided; that's a limitation, not a flaw.\n\nWho this is for: anyone building or evaluating MLLMs, especially around visual grounding, OCR, or robustness. It deserves a serious referee. I'd push for the Random-null analysis and a couple of robustness checks, but I'd engage with this.","headline":"A strong, well-executed benchmark for MLLM conflict resolution, but the directed-shift headline needs a Random-condition null before it fully lands.","tokens_in":28362,"tokens_out":2563,"would_cite":true,"duration_ms":24713,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that conflicting scene text systematically degrades multimodal geolocation and redirects predictions toward injected geographic targets, raising median localization error 4.8-fold across 20 models.","keywords":["multimodal large language models","text-vision conflict resolution","visual geolocation","counterfactual benchmark","scene text perturbation","geolocation robustness","text bias score","trap-fit rate"],"falsifier":"Take a random sample of Blank and Adversarial pairs beyond the 120 audited images and have independent raters or automated detectors identify non-text differences (buildings, sky, road geometry) that correlate with the injected target; if such differences appear systematically, the directed-shift conclusion would be confounded by editing artifacts rather than text semantics.","tokens_in":27377,"feed_emoji":"📍","tokens_out":9241,"duration_ms":67935,"temperature":0.7,"pith_summary":"SIGNPOST-Bench tests whether multimodal large language models can arbitrate between visual scene structure and readable scene text when the two disagree. Each source image becomes a counterfactual quintuplet—Original, Blank, Similar, Random, Adversarial—so that only the meaning of selected text spans changes while surrounding non-textual content is meant to stay fixed. Across 20 models and four datasets, the benchmark reports that Adversarial text raises median localization error from 282 km to 1,347 km (4.8x), that 6.5–20.1% of adversarial predictions land within 50 km of an injected geographic target, and that every model moves closer to the injected target on average relative to the Blank baseline. The authors argue this establishes visual geolocation as a continuous diagnostic for text–vision conflict and shows that clean-input capability does not predict conflict robustness. The stakes are practical: any image-based localization system that reads signs is exposed to misleading or tampered text.","feed_headline":"Injected sign text lifts median geolocation error 4.8-fold","feed_subtitle":"All 20 tested multimodal models shifted toward injected place names; up to 20.1% landed within 50 km.","key_machinery":"The load-bearing object is the counterfactual quintuplet: five matched images of the same scene—Original, Blank, Similar, Random, Adversarial—differing only in localized scene-text edits generated by an MLLM and rendered by an image-editing model, with the Adversarial variant naming a real place from a different continent. Three diagnostic metrics carry the argument: Weighted Localization Accuracy (WLA) scores geodesic error with exponential decay; Text Bias Score (TBS) measures the paired change in ground-truth error from the Blank text-removed baseline to an edited variant; and Trap-Fit Rate (TFR) plus paired Trap Distance Reduction (TDR) measure whether predictions move toward the geocoded injected target. A three-tier scene-text coupling taxonomy (Portable, Cultural, Geo-Specific) stratifies how much native text helps and how much conflicting replacement hurts.","core_discovery":"The paper's central claim is that conflicting scene text is a systematic, measurable failure mode in current multimodal large language models, not a rare edge case. Using a controlled five-condition counterfactual design, it shows that replacing native scene text with a geographically conflicting place name degrades localization for every model on every dataset: macro-averaged WLA falls from 47.11 to 29.89, median error grows 4.8-fold, and 6.5–20.1% of adversarial predictions land within 50 km of the injected target. The paired Trap Distance Reduction is positive on average for all 20 models (model means 343–1,926 km), and compatible, unrelated, and conflicting text replacements produce distinct, ordered effects relative to the text-removed Blank baseline. The authors conclude that text–vision conflict should be evaluated separately from clean-input capability, because the two are not interchangeable.","pith_inferences":["If the central claim holds, benchmarks that oversample Portable and Cultural text may understate text-conflict risk: Geo-Specific text (T3) yields the largest adversarial WLA drop (25.14 points, 42.2% relative).","The positive mean Trap Distance Reduction is driven by a minority of large shifts (attraction rates 44.5–60.6%), so average targetward movement can coexist with many samples staying put; deployment risk may concentrate in a small fraction of images.","A natural extension is to apply the same quintuplet logic to other continuous-output tasks, such as depth, time, or heading estimation, to test whether directed text-driven shifts are a general arbitration failure rather than a geolocation-specific one."],"forward_implications":["Adversarial scene text increases median localization error from 282 km to 1,347 km (4.8x) and lowers mean WLA from 47.11 to 29.89 across 20 models.","Conflicting text causes directed, target-aligned shifts: 6.5–20.1% of adversarial predictions fall within 50 km of the injected target, and every model shows a positive mean paired Trap Distance Reduction.","Text semantics matter, not just readable text: Similar replacements reduce error by 379 km relative to Blank, while Random and Adversarial replacements increase it by 959 km and 1,577 km.","Conflict robustness is separable from localization capability: models with modest clean-input performance can rank high in robustness, so capability scores do not predict behavior under conflict.","Prompting models to detect conflict does not reliably fix the failure: on two tested models, defense prompting improved conflict detection for one but lowered it for the other, and neither improved both detection and localization."],"supporting_citations":[{"why":"Supplies the IM2GPS3K and YFCC4K source datasets with ground-truth coordinates.","marker":"(Vo, Jacobs, and Hays 2017)"},{"why":"Provides YFCC100M, from which the YFCC4K sample is drawn.","marker":"(Thomee et al. 2016)"},{"why":"EasyOCR detects the candidate scene-text spans edited in all four sources.","marker":"(JaidedAI 2020)"},{"why":"The generator model that writes Similar, Random, and Adversarial text replacements.","marker":"(Google 2026)"},{"why":"Qwen-Image-Edit-2509 renders the localized counterfactual image edits.","marker":"(Wu et al. 2025)"},{"why":"Geocodes adversarial place names into injected geographic targets for TFR and TDR.","marker":"(Nominatim Developer Community 2026)"},{"why":"RIO-Bench provides the same-scene counterfactual approach this benchmark extends.","marker":"(Waseda et al. 2025)"},{"why":"ConText-VQA supplies the prior discrete-outcome conflict evaluation that SIGNPOST-Bench contrasts with.","marker":"(Zhang et al. 2026)"},{"why":"Establishes that scene text can override visual recognition, the vulnerability this benchmark turns into a continuous diagnostic.","marker":"(Goh et al. 2021)"}],"fun_headline_variants":["Scene-text conflicts send MLLM geolocation error up 4.8x","All 20 multimodal models biased by conflicting scene text","New benchmark exposes text-vision conflict failure in MLLMs","Injected place names hike median geo-error 4.8-fold in MLLMs","SIGNPOST-Bench: how MLLMs handle text-vision clashes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The counterfactual edits are assumed to change only the meaning of the selected scene-text spans while preserving surrounding non-textual content; only 120 edited images were human-audited, so subtle edit artifacts could contribute to the measured shifts.","fun_headline_variants_meta":{"raw":{"variants":["Scene-text conflicts send MLLM geolocation error up 4.8x","All 20 multimodal models biased by conflicting scene text","New benchmark exposes text-vision conflict failure in MLLMs","Injected place names hike median geo-error 4.8-fold in MLLMs","SIGNPOST-Bench: how MLLMs handle text-vision clashes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000481,"raw_usage":{"total_tokens":2413,"prompt_tokens":1015,"completion_tokens":1398,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":631,"completion_tokens_details":{"reasoning_tokens":1301}},"tokens_in":631,"tokens_out":1398,"duration_ms":12116,"temperature":1.0,"reasoning_tokens":1301,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T00:06:56.108564+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of Blank and Adversarial pairs beyond the 120 audited images and have independent raters or automated detectors identify non-text differences (buildings, sky, road geometry) that correlate with the injected target; if such differences appear systematically, the directed-shift conclusion would be confounded by editing artifacts rather than text semantics.","supporting_citations":[],"review_version":1}