{"id":"e94c7d08-e9a1-4269-9788-84da5dc3af07","arxiv_id":"2607.27069","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Between 12.7% and 26.3% of correct spatial yes/no answers from four open MLLMs receive no support advantage from the benchmark image over text-only or blank controls.","lead":"Visual Credit Audit checks whether a model's correct answer to a spatial yes/no question is actually supported by the image, or just by text priors and blank-image context. Across four open multimodal models and two benchmarks, 12.7–26.3% of correct answers were 'uncredited'—no more supported by the original image than by no-image controls.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The fixed mid-gray blank likely acts as an 'objects absent' cue rather than neutral absence; C-U is concentrated in the answer direction the no-image controls favor, so the 12.73–26.25% range may overstate decisions where the image does no work. A no-objects background control would settle it.","rationale":"The strongest claim reduces to the C-U mass estimates. If the blank is read as 'no objects here,' correct no items can receive high support from the blank without the model using the benchmark image, producing D=0 for reasons unrelated to whether the original image supplied spatial evidence. Because D_i is r_o > max(r_t, r_b), a single biased control can flip many items to C-U. The paper's sensitivity analyses do not remove the semantic absence cue: mean-color blank is still blank, mismatch adds other objects, and Strict T/B changes the event definition while showing how strong the control answer priors are. The permutation calibration is orthogonal: it demonstrates original-image specificity, not control neutrality. Thus the 12.73–26.25% interval is secure only if the controls are neutral absence — exactly the unverified assumption the reader identified. The extreme direction-stratified C-U and the Strict T/B collapse make the assumption doubtful enough to keep the verdict conditional rather than full accept. A no-objects background control would directly test it, so the reader's conditional verdict remains appropriate.","tokens_in":30237,"tokens_out":11463,"duration_ms":117103,"concrete_test":"Construct a third visual control for the full benchmark records: for each original image, remove/inpaint only the two queried objects using the available boxes, preserving background and layout — a no-objects, same-background control. Recompute D_i, D-CC, and C-U under control sets {text-only, no-objects background} and {text-only, gray blank, no-objects background} for all four models on VSR and GSR-COCO. Pre-register a threshold: if pooled C-U shifts by more than 5 absolute points, or the Table 14 gold-direction concentration changes materially, the fixed gray blank is not neutral absence and the headline range is control-dependent. Additionally compare model margins on gray blank vs no-objects background for gold-yes vs gold-no items to directly test the 'objects absent' cue.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Eq. (1) and Table 10 make text-only and the fixed mid-gray blank the primary 'no benchmark image' controls. The central C-U estimate requires these controls to behave as neutral absence of benchmark visual content. The gray canvas is not neutral for a spatial yes/no task: it is a visible scene with no objects, which can directly license a 'no' answer. The paper's own strata show the expected signature. In Table 14, LLaVA VSR gold-no C-U is 36.94 points while gold-yes C-U is 2.10; InternVL VSR gold-yes C-U is 34.78 while gold-no C-U is 1.20. C-U is concentrated in the answer direction favored by the no-image/blank priors. Table 28's Strict T/B requirement — that both controls favor the opposite answer — collapses Dep substantially (e.g., InternVL GSR from 42.84 to 10.45), confirming that the controls usually share the original answer direction rather than acting as neutral absence. Mean-color and mismatch variants change low-level pixel content or add other objects; they do not remove the semantic 'no objects' cue. The matched same-split permutation calibration (21.25–47.80 point D-CC drop) shows the original image is not a generic image, but it does not show the gray blank is a valid realization of 'no benchmark image.' Thus the headline C-U mass rests on an untested control-neutrality assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Visual Credit Audit (VCA), a label-free, decision-level procedure for closed yes/no spatial benchmarks. For each query it caches a frozen MLLM's yes/no continuation margin under the original image, a text-only prompt, and a fixed mid-gray blank image, and defines a dependence event D = 1{r_o > max(r_t, r_b)}. After labels are applied, the paper obtains D-CC = E[κD] and C-U = E[κ(1−D)], and shows that on correct items the event reduces to a gold-aligned positive-gain statistic. Across four MLLMs and VSR/GSR-COCO, it reports 12.73–26.25% correct-but-uncredited decisions and a matched same-split image-permutation D-CC excess of 21.25–47.80 points. A relation axis (SRC, PairCredit, Joint) and a 3×3 evidence-source factorial are used to argue that marginal image-support advantage and fixed-pixel relation response are separate estimands. The paper concludes that benchmark accuracy can overstate image-grounded success.","tokens_in":30592,"tokens_out":15690,"duration_ms":132491,"significance":"The paper is unusually transparent and the accounting is sound: D is assigned before labels, the four-cell decomposition in Eq. (3) is complete, and the correct-item restriction in Eq. (4) exactly recovers gold-aligned gain. The paired-bootstrap intervals, deterministic same-split permutation calibration, threshold-mass analysis, multi-verbalizer checks, and direction-balanced sensitivities are all appropriate and mostly well reported. The paper ships code, deterministic data builders, cached margins, and deidentified annotations, which is a real strength. The two-estimand separation — relative image support versus relation response — is a useful corrective, and the factorial convincingly shows why null-control marginal support cannot identify relation response. If the central C-U estimate survives a neutral no-objects control, the finding that 12.73–26.25% of correct decisions lack control-relative image support is important for benchmark interpretation.","major_comments":[{"comment":"The headline C-U estimate treats the fixed mid-gray blank as a neutral 'no benchmark image' baseline. For a spatial yes/no probe, a gray canvas is a visible scene with no objects, which can directly license 'no'. The paper's strata show the signature: in Table 14, LLaVA VSR C-U is 36.94 for gold-no vs 2.10 for gold-yes; InternVL VSR is 34.78 for gold-yes vs 1.20 for gold-no. C-U is concentrated in the direction favored by no-image/blank priors. Table 28's strict T/B requirement collapses Dep (e.g., InternVL GSR 42.84→10.45), consistent with controls usually agreeing with the original answer rather than being neutral. The mean-color blank (Table 37) changes low-level color but still has no objects, so it does not remove the 'objects absent' cue. This is load-bearing for the 12.73–26.25% range. Please add a no-objects background control (or text-only-only reporting) to isolate the gray-bla","section":"§3.1, Eq. (1); Table 10; Tables 14 and 28"},{"comment":"The headline range is not direction-invariant. Prediction-balanced C-U is 11.73–43.36% while gold-balanced is 12.73–26.25%, and VSR orderings change. The supplement discloses this, but the main text (abstract, conclusion) states the range without the qualification. Since the C-U estimate is also sensitive to the blank-control concern in Comment 1, the main text should explicitly say that the 12.73–26.25% figure is benchmark-distribution-specific and should report the direction-balanced estimates in the main robustness section.","section":"Abstract/Conclusion vs Table 15"}],"minor_comments":[{"comment":"Throughout, 'no-image controls' is imprecise for the blank: a gray canvas is still an image. Suggest 'image-absent' or 'content-free image' to avoid implying that the blank contains no visual input.","section":"Abstract/§3.1"},{"comment":"The last column header 'C-U∆ perm DCC' is hard to parse; split into 'C-U' and 'Δ perm D-CC [CI]' and explain in the caption that Δ is the paired original-minus-permutation D-CC difference.","section":"Table 1"},{"comment":"The 3×3 evidence-source factorial is central to the response-vs-support separation but is only described in text; a diagram of the nine cells (or at least the agreement and conflict cells) would make the design easier to verify.","section":"§4.3/Figure 1"},{"comment":"PairCredit = 1{G_int > 0} can be positive even when the original true–false separation is negative if the controls are more negative. The caption's phrase 'image-specific semantic separation' should explicitly state that PairCredit does not require H_o > 0; PairCorrect plays that role.","section":"Table 2/§4.2"},{"comment":"The Limitations section is honest about causal attribution but does not list the gray-blank neutrality assumption as a limitation; add it explicitly, together with the no-objects control that would test it.","section":"Limitations"}],"recommendation":"major_revision","confidential_remarks":"This is a well-executed and unusually transparent paper. The main substantive risk is the gray-blank control; I would like to see a no-objects control or a text-only-only headline before accepting. The direction-balancing caveat should be promoted to the main text. I do not see grounds for rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid, well-scoped paper, and the headline finding—12.7–26.3% of decisions are correct but uncredited—is real under the paper's own definitions. The genuinely new contribution is the label-free, all-decision dependence event D_i and the four-cell correctness–credit decomposition. On correct items it intentionally reduces to the same gold-aligned gain that already exists in the literature, so the novelty is the extension to errors and the complete decomposition, not a new correct-only formula. The empirical work is careful: paired bootstrap CIs, a matched same-split permutation calibration, a controlled 3x3 evidence-source factorial, and 108 human-audited natural edits, with code and data promised.\n\nThe main result holds up. D_i is prespecified, no labels are touched until after the event, the accounting is simple and correct, and the permutation drop (21–48 points, all intervals above zero) shows the original image is not generic. The factorial is the right complement: it shows that many uncredited decisions are still visually responsive (91% directional response, 32% answer flip), which disambiguates 'no marginal support advantage' from 'ignoring the image.'\n\nThe soft spots are real but not fatal. The stress-test concern about the blank control is fair: a fixed mid-gray canvas is not neutral absence, it is an empty-scene cue that can license 'no' answers, and C-U is indeed concentrated in the answer direction the controls favor. But the paper defines credit operationally as support beyond the prespecified controls, not as causal image use, and the mean-color blank and gold-balanced robustness checks keep the C-U range intact. So this changes the interpretation (the numbers are 'relative to an empty-scene control,' not 'image does no work') without collapsing the finding. The PairCredit issue is more of a labeling problem: G_int > 0 can hold when the original image still ranks the false relation above the true one, so 'image-specific true–false separation' overstates it. They separately report Rel_o for the strict ranking, so the fix is easy: call PairCredit something like 'separation improvement over controls' and make the difference explicit.\n\nWho it's for: anyone working on multimodal evaluation or building spatial reasoning benchmarks who wants a cheap, label-free way to decompose accuracy into image-supported vs. prior-supported decisions. Not a benchmark, a tool. Recommend sending to peer review. The core claims are supported; the two soft spots are clarifications, not ground truths.","headline":"Careful, mostly convincing label-free audit of image support in spatial yes/no benchmarks; headline C-U numbers hold under the paper's operationalization, with two fixable soft spots.","tokens_in":31098,"tokens_out":6092,"would_cite":true,"duration_ms":56056,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that accuracy on spatial yes/no benchmarks overstates image-grounded correctness, and that a label-free decision-level audit can decompose benchmark success into correctness, additional image support, and relation-consiste","keywords":["Visual Credit Audit","spatial reasoning","multimodal language models","dependence-credited correctness","correct-but-uncredited decisions","benchmark auditing","label-free evaluation","relation consistency"],"falsifier":"Run the same audit with a control that preserves scene structure but removes the queried objects, or with a set of unrelated natural images matched for low-level statistics; if D-CC rises to near accuracy — that is, if the 13–26% correct-but-uncredited mass largely disappears — then the C-U estimate is an artifact of the particular text-only/blank controls rather than a stable property of the models.","tokens_in":30107,"feed_emoji":"🖼️","tokens_out":5724,"duration_ms":54018,"temperature":0.7,"pith_summary":"Visual Credit Audit (VCA) tries to show that accuracy on closed yes/no spatial benchmarks can be high even when the benchmark image contributes little beyond text-only or blank prompts. For each decision, it compares how strongly the image supports the model's own answer against a text-only prompt and a blank gray canvas; only decisions with a positive \"dependence gap\" receive credit. Applying correctness labels afterwards yields a four-way split of all decisions, and across four open multimodal models and two benchmarks 12.73–26.25% of decisions are correct yet uncredited. A matched image-permutation control lowers dependence-credited correctness by 21.25–47.80 points, with every confidence interval above zero. If the claim holds, benchmark accuracy alone overstates image-grounded correctness, and reports should separate correctness from visual credit.","feed_headline":"Up to 26% of correct spatial-AI answers get no image support","feed_subtitle":"A new audit labels each yes/no decision by whether the image adds support, exposing a correctness gap accuracy hides.","key_machinery":"The central object is the prediction-aligned dependence event D_i = 1{r_{i,o} > max(r_{i,t}, r_{i,b})}, where r is the model's margin for its own declared answer under the original image, text-only, and blank contexts, with the event fixed before correctness labels are seen. This yields the four-cell accuracy-credit decomposition and the headline quantities D-CC and C-U. The companion relation axis uses the fixed-pixel true-false separation H_c and the interaction G_int = H_o − max(H_t, H_b), plus the Joint conjunction, to distinguish marginal image support from relation-consistent visual response. The matched same-split image permutation is the calibration mechanism that tests whether the o","core_discovery":"The paper's central claim is that a single label-free event — whether the original image gives the model's declared answer more support than both a text-only prompt and a fixed gray blank — separates benchmark success into four exhaustive cells, and that a substantial share of correct answers fall in the \"correct but uncredited\" cell. In the reported runs, this share ranges from 12.73% to 26.25%, meaning ordinary accuracy overstates support-qualified correctness by that amount. The paper further shows that this gap is not an artifact of the specific controls: replacing the benchmark image with a deterministic same-split unrelated image lowers dependence-credited correctness by 21.25 to 47.80","pith_inferences":["The same four-cell decomposition would likely transfer to other closed-form multimodal benchmarks (VQA, visual entailment, object presence), where a text-only or blank control can also be defined; the spatial setting may just be the easiest place to see it.","Because the dependence event is label-free, it could serve as a data-filtering signal to remove benchmark items that reward answer priors, complementing label-based filtering methods.","The deterministic same-split permutation could be replaced by a distribution over unrelated images to produce a tighter and more general null for \"generic image\" support; the paper tests one permutation per sample.","The high directional response among uncredited decisions suggests that \"uncredited\" means redundant visual support rather than absent visual support; a useful next step would be to separate those two sub-cases inside the C-U cell."],"forward_implications":["Benchmark reports would need to include D-CC (dependence-credited correctness) alongside accuracy, since two models with similar accuracy can differ by more than 20 points in support-qualified success.","Because 12.73–26.25% of correct decisions are not better supported by the image than by text-only or blank controls, accuracy is an upper bound on image-grounded correctness, not a measure of it.","Credit is a relative-support statement, not answer necessity: a decision can be uncredited even when the no-image controls do not flip the answer, so \"no image support\" should not be read as \"image irrelevant.\"","Correct-but-uncredited decisions are often still visually responsive: under controlled relation reversal, 81.57–100.00% of such decisions move in the correct direction and 32.11% change their answer, so a single audit number cannot capture both support and relation response.","The audit is training- and label-free and needs only three cached contexts per item, so it can be applied to any frozen model and any fixed binary answer interface without retraining."],"fun_headline_variants":["Image adds zero support for up to 26% of correct AI answers","Correct but uncredited: 1 in 4 right answers ignore the image","Up to 26% of correct spatial-AI answers get no image support","Accuracy hides a gap: up to 26% of correct answers lack image support","Spatial AI: up to 26% of correct answers don't use the image"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The audit's load-bearing premise is that a text-only prompt and a fixed mid-gray blank image are neutral stand-ins for \"no benchmark image,\" so that any extra support the original image provides over both is genuinely due to the image's content; if the gray blank acts as an active distractor or as a hidden \"objects absent\" cue, the correct-but-uncredited share is overstated.","fun_headline_variants_meta":{"raw":{"variants":["Image adds zero support for up to 26% of correct AI answers","Correct but uncredited: 1 in 4 right answers ignore the image","Up to 26% of correct spatial-AI answers get no image support","Accuracy hides a gap: up to 26% of correct answers lack image support","Spatial AI: up to 26% of correct answers don't use the image"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001017,"raw_usage":{"total_tokens":4145,"prompt_tokens":777,"completion_tokens":3368,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":521,"completion_tokens_details":{"reasoning_tokens":3264}},"tokens_in":521,"tokens_out":3368,"duration_ms":20186,"temperature":1.0,"reasoning_tokens":3264,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T10:08:01.139263+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same audit with a control that preserves scene structure but removes the queried objects, or with a set of unrelated natural images matched for low-level statistics; if D-CC rises to near accuracy — that is, if the 13–26% correct-but-uncredited mass largely disappears — then the C-U estimate is an artifact of the particular text-only/blank controls rather than a stable property of the models.","supporting_citations":[],"review_version":2}