{"id":"8c17ecb9-f4c6-48bd-9e34-6dbcbc832501","arxiv_id":"2608.13267","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A new benchmark and the A-R-I framework show that high perception and reasoning accuracy in vision-language models does not predict whether they admit uncertainty, resist false premises, or infer cautiously from partial visual evidence.","lead":"This paper introduces SciFigBench, a benchmark that tests whether vision-language models describe scientific figures accurately and behave reliably when images are blurred, labels are missing, or captions are misleading. It finds that a model can top accuracy scores yet fabricate answers for unreadable content in 96 percent of cases, while a comparable model admits uncertainty 71 percent of the time.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Behavioral divergence rests on unvalidated GPT-4o binary labels; S2's 'fabricates' definition counts any specific answer, so the 96% hallucination figure and the 'confident fabricator' label are not yet supported.","rationale":"The reader's CONDITIONAL verdict already names the central vulnerability: the behavioral scores are assigned by a GPT-4o judge with no separate human validation for admittance and fabrication, while human agreement covers only MQM description scoring. My stress-test read agrees and adds a specific aggravating detail: the S2 rubric defines 'fabricates' as merely providing a specific answer, with no confidence or hedging condition. This means the paper's 96% hallucination rate is actually a 96% non-abstention rate under the rubric as written, and the 'confident fabricator' profile attributed to GPT-5.2 is not directly measured. The paper deserves credit for its transparent prompt inventory, cross-judge capability validation, probe-designer ablation, bootstrap CIs, and the substantial human annotation effort on MQM; those supports are real but do not cover the binary behavioral judgments. A single targeted human-annotation study of the S2/S3 rubrics would settle whether the headline reversal is a measurement artifact or a true behavioral difference. Since the reader's conditional verdict already requires exactly this kind of validation before the central claim is fully established, my read does not move the verdict; it sharpens the condition that must be met.","tokens_in":26789,"tokens_out":8379,"duration_ms":91148,"concrete_test":"Sample 100 active admittance-blur responses (50 from GPT-5.2, 50 from Gemini 3.1 Pro), strip model identity, and have two annotators independently apply the S2 rubric plus a stricter label: 'unhedged specific assertion without any uncertainty expression'. Compute human-human and human-vs-GPT-4o Cohen's kappa for the admits and fabricates decisions, and compare model-level rates from human labels with the reported 8%/71% and 96% figures. Additionally re-score all admittance-blur responses with a modified rubric that counts as fabrication only unhedged, confident specific answers, and report each model's fabrication rate and the full admits x fabricates contingency table.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central reversal — GPT-5.2 'hallucinates unreadable content in 96% of cases' while Gemini 3.1 Pro 'admits uncertainty in 71% of such cases' — is produced by GPT-4o applying the S2/S3 behavioral rubrics. Appendix C.1 validates the LLM judge only for MQM description scoring: human agreement (Krippendorff alpha = 0.91) and error-type F1 cover MQM, not the admittance/fabrication binary judgments that carry the headline claim. The S2 rubric makes this more than a missing-validation gap: 'fabricates' is defined as 'provided a specific answer' with no confidence or hedging requirement, so a response such as 'I can't read the label, but it might be X' is simultaneously 'admits' and 'fabricates'. The paper nevertheless describes GPT-5.2 as a 'confident fabricator', a characterization the rubric never measures. Without human agreement on these binary labels, and without reporting Gemini's own fabrication rate or the admits-by-fabricates contingency, the observed divergence could reflect judge sensitivity to hedging style (Gemini uses 16k max tokens and more verbose hedging) rather than a genuine behavioral difference in epistemic honesty. This is the load-bearing uncertainty because the whole conclusion — accuracy scores hide behavioral opposites — depends on these specific numbers being true.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents SciFigBench, a diagnostic benchmark for vision-language model (VLM) understanding of scientific figures, and the Admittance–Resistance–Inductance (A-R-I) framework for evaluating behavioral reliability under uncertainty. The benchmark contains 250 arXiv figures with expert descriptions, 1,000 human-reviewed reasoning questions, and over 34,000 evaluation setups including image transformations, caption-bias probes, false-premise resistance probes, and selective-blur admittance/inductance probes. Eight VLMs are evaluated on perception (MQM description quality), reasoning accuracy, and behavior. The central empirical claim is that models with similar perception and reasoning accuracy can be behavioral opposites under uncertainty: GPT-5.2 achieves the highest description quality (MQM 91.6) but is reported to hallucinate unreadable content in 96% of admittance-blur cases, whereas Gemini 3.1 Pro, with comparable accuracy, admits uncertainty in 71% of such cases and achieves the strongest resistance score (0.91). The authors argue that accuracy-based benchmarks therefore rank behaviorally opposite models as near-equivalents, and that behavioral reliability must be evaluated separately.","tokens_in":27042,"tokens_out":4323,"duration_ms":46509,"significance":"If the behavioral measurements are valid, this is a substantial contribution. The paper's strengths are considerable: a large human annotation effort (600+ hours) with reported inter-annotator agreement; externally grounded benchmark construction from arXiv figures; a clear A-R-I decomposition that is conceptually useful; deterministic inference settings; bootstrap CIs; split-half reliability; cross-judge ablations for reasoning; and a probe-designer independence check. The framework addresses a real gap in chart and scientific-figure evaluation, and the finding that accuracy scores alone do not predict uncertainty acknowledgment would be practically important for deploying VLMs in scientific workflows. The appendices are unusually transparent, with full prompt and rubric inventories, which supports reproducibility.","major_comments":[{"comment":"The headline admittance, fabrication, and resistance numbers are produced entirely by a GPT-4o judge applying the S2/S3 rubrics, yet Appendix C.1 validates the judge only for MQM description scoring (Krippendorff's α = 0.91, error-type F1s). No human agreement or error analysis is reported for the binary admits/fabricates/correct labels, nor for the 1.0/0.5/0.0 resistance scoring. This is load-bearing because the central reversal—GPT-5.2 hallucinates 96% vs. Gemini admits 71%—is a statement about these unvalidated binary labels. The paper's Limitations section overstates the case by citing the MQM human agreement as validation for 'automated evaluation' broadly. Please report human-annotator agreement for the S2/S3 rubrics (or at minimum a representative error analysis) and for the resistance scoring rubric before these figures are used to support the paper's main conclusion.","section":"§4.4, Appendix B.7, Appendix F (S2/S3)"},{"comment":"The S2 rubric defines 'fabricates' as 'provided a specific answer' with no confidence or hedging requirement, so a response such as 'The label is unreadable, but it might be X' is simultaneously scored as 'admits' and 'fabricates.' The abstract and §4.4 nonetheless describe GPT-5.2 as a 'confident fabricator' and state it 'hallucinates unreadable content in 96% of cases.' Confidence is never measured by the rubric. Please report the joint distribution of the admits and fabricates labels (admits∧fabricates, admits∧¬fabricates, ¬admits∧fabricates, ¬admits∧¬fabricates) for all models, and re-express the headline as 'fabrication without admission' or add a third label such as 'confident fabrication' if that is the intended construct. Without this, the 96% figure conflates genuinely confident hallucinations with hedged guesses and overstates the behavioral gap.","section":"Appendix F (S2), §4.4"},{"comment":"The behavioral comparison is asymmetric: Gemini's fabrication rate on the admittance probes is never reported, and neither is the overall admittance-by-fabrication contingency for either model. Only GPT-5.2's fabrication rate (96%) and Gemini's admittance rate (71%) are given. To support the claim of 'behaviorally opposite' profiles, the full contingency for both models is required. If Gemini also fabricates a large fraction of the time while admitting, the apparent opposition may reflect a difference in hedging style rather than in fabrication proclivity. This is especially important given that Gemini uses a 16k max-token setting and the judge may be sensitive to response length or hedging phrasing; the required table would let a reader evaluate whether the divergence is a model-level behavioral difference or a judge artifact.","section":"Table 4, §4.4"},{"comment":"The paper's own validation data in Table 9 show that GPT-4o as an MQM judge has very low recall for Hallucinated Content (recall = 0.07, F1 = 0.12). This under-detection of hallucinated content in the description-quality channel is a further reason to require direct validation of the behavioral fabrication labels: the same judge model that under-detects hallucination in MQM scoring is being asked to classify fabrication in the S2/S3 rubrics. The manuscript currently does not demonstrate that the judge can reliably discriminate fabricated from non-fabricated blur-region content, which is the exact discrimination on which the headline claim rests. Please provide a direct human evaluation of the S2/S3 labels, ideally stratified by model and by admits/fabricates combination.","section":"Appendix C.1, Table 9"}],"minor_comments":[{"comment":"The column header 'Capability (%) Resistance Admittance (%) Inductance (%)' is ambiguous because the Resistance columns are scores on a 0–1 scale while the others are percentages; the caption explains this, but the table would be clearer if the header rows visually separated the four blocks, as in the subcaptions.","section":"§4.4, Table 4"},{"comment":"Please state explicitly in §4.4 or in the limitations that Gemini 3.1 Pro was run with max_tokens = 16,000 while all other models used 2,048, because this output-length asymmetry is a plausible confound for the judge-based admittance measurement and should be addressed in the analysis.","section":"Appendix E.3"},{"comment":"The MQM formula 'max(0, 100 − P×100/(N×5))' assumes every checklist item can incur a maximum Major penalty of 5.0; since the checklist items carry Major/Minor severity tags, it would be helpful to state explicitly that N×5 is an upper bound and that the score is accordingly a normalized rather than exact percentage.","section":"§3.2, MQM formula"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong benchmark contribution with unusually transparent methodology, but the headline behavioral claims currently exceed the validation evidence. The central 'accuracy hides behavioral opposites' narrative would be compelling if supported by human-validated behavioral labels and the full independent admits/fabricates marginals; as presented, the 96% vs. 71% comparison rests on a single judge whose behavioral classifications have not been checked against humans. This is fixable within the manuscript's scope by adding a validation study and re-reporting the contingency, which is why I recommend major revision rather than rejection. I also recommend asking the authors to soften the 'confident fabricator' language unless they add a confidence component to the rubric."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my read on SciFigBench. The genuinely new piece is the behavioral axis: A-R-I (Admittance, Resistance, Inductance) applied to scientific figures, with human-confirmed selective-blur targets split into unrecoverable vs. contextually inferable, plus false-premise and caption-bias probes. That's a useful decomposition, and the annotation work is substantial: 250 figures, 1,000 reviewed questions, 600+ hours. The paper also does solid statistical housekeeping—bootstrap CIs, split-half reliability, cross-judge and probe-designer ablations, and full prompt/API transparency in the appendices.\n\nThe central empirical claim is that accuracy does not guarantee behavioral reliability: GPT-5.2 leads MQM (91.6) but admits uncertainty only 8% on unreadable blurred elements and fabricates 96%, whereas Gemini 3.1 Pro admits 71% and resists misleading context at 0.91. If that reversal holds, it's important for anyone picking a VLM for scientific or medical workflows.\n\nHere's the soft spot, and it's load-bearing. The behavioral numbers come from a GPT-4o judge, and the paper's human validation covers MQM description scoring only—not the admits/fabricates binary judgments that produce the headline. The S2 rubric defines 'fabricates' as 'provided a specific answer,' so a response like 'I can't read the label, but it might be X' counts as both admitting and fabricating. Calling GPT-5.2 a 'confident fabricator' overreaches: nothing in the rubric measures confidence. The paper also doesn't report Gemini's fabrication rate or the admits-by-fabricates contingency. Given that Gemini runs with 16k max tokens and tends to hedge more verbosely, some of the divergence could be judge sensitivity to hedging style rather than genuine epistemic behavior.\n\nSmaller issues: caption-bias resistance excludes unaddressed claims from the denominator, which can reward omission, and the dataset/scripts are promised rather than shipped.\n\nBottom line: this is a strong benchmark contribution with a plausible but not fully established central finding. It deserves a serious referee, and the right outcome is a revision that validates the blur-judge labels with human annotators, reports the contingency, and softens the 'confident fabricator' language. I'd bring it to reading group and would cite the A-R-I framework and benchmark resource, with the caveat in mind.","headline":"Solid new benchmark and A-R-I framework for VLM uncertainty behavior, but the headline admittance/fabrication reversal needs human validation of the judge labels.","tokens_in":27641,"tokens_out":4098,"would_cite":true,"duration_ms":38737,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SciFigBench shows a vision-language model can score highest on figure description quality yet fabricate answers for blurred content in 96% of cases.","keywords":["vision-language models","scientific figures","benchmark","behavioral reliability","uncertainty","hallucination","admittance-resistance-inductance","chart understanding"],"falsifier":"Take model responses to the 228 selectively blurred admittance figures and have human annotators independently label each response as admitting or fabricating; if the human admit rates for the two leading models converge or invert relative to the automated judge, the central behavioral gap is an artifact of the judge.","tokens_in":26558,"feed_emoji":"📊","tokens_out":7521,"duration_ms":73083,"temperature":0.7,"pith_summary":"The paper argues that how a vision-language model behaves when visual evidence is blurred, missing, or contradicted by a caption is a separate competence from how accurately it describes and reasons about what it sees. To show this, it builds SciFigBench, a benchmark of 250 scientific figures expanded into more than 34,000 test setups, and introduces the A-R-I framework (admittance, resistance, inductance) for scoring behavior under uncertainty. The headline result is that two frontier models with nearly equal description and reasoning scores behave oppositely under a selective blur: the top describer fabricates an answer for unreadable content in 96% of targeted cases and admits uncertainty only 8% of the time, while the other admits uncertainty in 71% of cases and resists misleading premises at the highest measured rate. If true, accuracy-based rankings used to pick models for scientific workflows could select for confident fabrication while treating behaviorally reliable models as equivalent.","feed_headline":"Top-scoring vision model fabricates answers to blurred figures 96% of the time","feed_subtitle":"A comparably accurate rival admits uncertainty 71% of the time, so accuracy alone does not predict behavior.","key_machinery":"The load-bearing mechanism is the A-R-I framework applied through controlled visual stress tests. Its key operational distinction is between admittance blur and inductance blur: a selectively blurred element is either unrecoverable from the remaining context (an admittance probe, where a reliable model should say it cannot tell) or inferable from surrounding cues such as axis scales or legend colors (an inductance probe, where a reliable model should infer the value). Around this distinction, resistance probes with misleading captions, non-existent elements, false numerical anchors, and unanswerable questions quantify how a model handles context that contradicts the figure. This A-R-I trio gives each model a behavioral profile that is independent of its MQM and reasoning scores.","core_discovery":"The central discovery is that behavioral reliability under uncertainty is empirically separable from perception and reasoning accuracy. Using selectively blurred chart labels that are either unrecoverable from context (admittance probes) or inferable from remaining cues (inductance probes), plus caption-bias and false-premise probes, the paper finds that GPT-5.2, despite the highest MQM description quality (91.6) and strong reasoning accuracy (78.4%), states a specific answer for a blurred element 96% of the time and acknowledges the visual limitation only 8% of the time under direct questioning. Gemini 3.1 Pro, with MQM 90.2 and reasoning accuracy 81.0%, admits uncertainty in 71% of these cases and achieves the highest resistance score (0.91). Population-level correlations between quality and behavior are high (Spearman rho 0.83 to 0.95 across dimensions), but the reversal at the top of the leaderboard shows that high perception and reasoning scores do not guarantee reliable behavior; a benchmark reporting only MQM would rank these two models as near-equivalent while missing opposite behavioral profiles.","pith_inferences":["The same A-R-I split could serve as a cheap pre-deployment screen: run admittance blurs on a small set of figures and measure the admit rate before trusting a model with scientific documents.","The distinction between unrecoverable and inferable evidence should transfer to medical imaging or document processing, where a model that says 'I cannot tell' is often safer than one that guesses a plausible value.","Because the admittance gap is assigned by a single automated judge and the paper's human validation covered MQM description quality rather than the binary admittance and fabrication labels, a human re-rating study of the blurred-label responses would directly test whether the 8% versus 71% gap is a model property or a judge artifact.","The 'must-answer' bias finding points to a causal follow-up: train an open-weight model on examples that reward explicit uncertainty for unrecoverable elements, then measure whether this behavior transfers to unseen chart types and languages."],"forward_implications":["Models with nearly equal MQM and reasoning scores can be assigned opposite behavioral profiles, so accuracy-based leaderboards cannot stand in for deployment-safety evaluation.","The A-R-I dimensions are empirically separable: a model that resists false premises can still fabricate when evidence is missing, so both dimensions need to be measured independently.","Selecting a model for scientific-figure processing based only on description quality could embed silent fabrication into summarization, comparison, or question-answering pipelines.","Presupposition-embedded false premises (for example, asking about a benchmark line that does not exist) are the hardest resistance probes across all models, more so than explicitly wrong numerical values.","Caption-bias resistance does not track model quality monotonically, suggesting that how much a model trusts provided context is shaped by instruction-following behavior rather than raw capability alone."],"supporting_citations":[{"why":"ChartQA supplies the chart question-answering baseline that SciFigBench extends with behavioral evaluation.","marker":"Masry et al., 2022"},{"why":"CharXiv provides the scientific-figure benchmark comparison and motivates evaluation of figures drawn from real papers.","marker":"Wang et al., 2024"},{"why":"SciFIBench is the direct scientific-figure interpretation benchmark that SciFigBench builds beyond.","marker":"Roberts et al., 2024"},{"why":"ChartQAPro introduces unanswerable chart questions, informing the design of unanswerable resistance probes.","marker":"Masry et al., 2025"},{"why":"Eyewitness testimony work on definite articles grounds the presupposition-embedding used in inexist probes.","marker":"Loftus, 1975"},{"why":"Anchoring and heuristics research motivates the contra probes that embed false numerical anchors.","marker":"Tversky and Kahneman, 1974"},{"why":"LLM-as-judge methodology supports the automated scoring pipeline used for descriptions and behavioral probes.","marker":"Zheng et al., 2023"},{"why":"Object hallucination in image captioning frames the measurement of fabricated content in model descriptions.","marker":"Rohrbach et al., 2018"},{"why":"The finding that language models often know what they know motivates the admittance dimension of acknowledging uncertainty.","marker":"Kadavath et al., 2022"}],"fun_headline_variants":["Top vision model invents answers to blurred figures 96% of the time","Accurate AI still fakes answers 96% of time, rival admits doubt 71%","Behavioral reliability is not predicted by accuracy in VLMs","Top VLM hallucinates on unreadable figures, rival resists","Accuracy rank masks honesty gap: 96% vs 71% uncertainty"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The admittance and fabrication percentages rest on a single automated judge's binary labels, and the paper's human validation covers only the MQM description-quality scores, so if that judge labels wording differently across models the headline divergence could be a judge artifact rather than a behavioral difference.","fun_headline_variants_meta":{"raw":{"variants":["Top vision model invents answers to blurred figures 96% of the time","Accurate AI still fakes answers 96% of time, rival admits doubt 71%","Behavioral reliability is not predicted by accuracy in VLMs","Top VLM hallucinates on unreadable figures, rival resists","Accuracy rank masks honesty gap: 96% vs 71% uncertainty"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000602,"raw_usage":{"total_tokens":2865,"prompt_tokens":1056,"completion_tokens":1809,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":672,"completion_tokens_details":{"reasoning_tokens":1710}},"tokens_in":672,"tokens_out":1809,"duration_ms":12661,"temperature":1.0,"reasoning_tokens":1710,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T15:01:29.924750+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take model responses to the 228 selectively blurred admittance figures and have human annotators independently label each response as admitting or fabricating; if the human admit rates for the two leading models converge or invert relative to the automated judge, the central behavioral gap is an artifact of the judge.","supporting_citations":[],"review_version":1}