{"id":"fc02efc9-ee9d-49e2-a1af-fa4316064d08","arxiv_id":"2608.07435","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"SABRE is a modular pipeline for automatically building and refreshing VLM stress tests, and its SABRE-Prior instantiation finds that six frontier VLMs score 17.8 to 31.3 percent when visual evidence conflicts with world priors.","lead":"This paper introduces SABRE, an automated pipeline that turns a text description of a vision-language model stress test into generated images, questions, and verified answers. A first instantiation, SABRE-Prior, shows that six frontier vision-language models answer only about 18 to 31 percent of questions correctly when images contradict common expectations.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SABRE-Prior's low scores may stem from unreliable human verification: reviewers see intended answers and filtering-model failures, with no inter-annotator agreement or independent audit reported.","rationale":"I read the paper in good faith. The framework is well specified: a Test Primer is converted to schema-validated specifications, images are generated and edited, candidates are pressure-screened by a filtering VLM, and survivors pass a human verification step that supports localized repair. The real-image Attribute control (Section 5.5) and the blind repair-preference study (Appendix D.3) provide useful checks, and the cost figures are honestly labeled as illustrative. The central claim, however, is conditional on the validity of the final labels. The reader identified the human-verification reliability assumption as the weakest link, and my reading agrees: the interface's display of intended answers and filter failures creates a concrete bias pathway, and the complete absence of an inter-annotator agreement or an independent gold-standard audit leaves the headline 17.8–31.3% range vulnerable to the alternative explanation of systematic annotation error. This does not require rejecting the paper: the pipeline architecture and the extensibility pilots stand on their own, and a blind re-annotation audit would either confirm the reported difficulty or show that the benchmark needs relabeling. The current CONDITIONAL verdict remains appropriate, with the additional condition that the audit be conducted before the benchmark is released as a measure of world-prior failures.","tokens_in":20450,"tokens_out":6297,"duration_ms":68066,"concrete_test":"Blind re-annotate a stratified random sample of 100 SABRE-Prior cases (25 per subset) with two independent reviewers who do not see the reference answers, the filtering model's predictions, or the pass/fail status. For Context and Texture, reviewers mark whether the source object is fully absent in the Edited image and whether the target object is clearly recognizable; for Attribute, they independently recount visible components; for Language Elicitation, they check whether any distractor is visually verifiable from the image. Compute raw agreement and Cohen's kappa, and the share of cases where the intended reference answer is wrong. If kappa ≥ 0.8 and the wrong-answer rate ≤ 2%, the concern is resolved; if the wrong-answer rate is 5–10% or higher, the reported macro-averages are not trustworthy evidence of prior-conflict failures.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is that six frontier VLMs achieve only 17.8–31.3% macro accuracy on SABRE-Prior because they fail to follow visual evidence that conflicts with world priors. That interpretation depends on the correctness of 1,000 reference answers, and the only quality gate is human verification (Section 3.5, Appendix D.1). No inter-annotator agreement, reviewer calibration, or independent audit of the final benchmark is reported. The annotation interface displays the intended reference answer, the filtering model's prediction, and whether that prediction was correct, which can bias reviewers on borderline cases: a small residual of the replaced source object, an ambiguous count, or an incidental cue in a Language Elicitation image may be resolved in favor of the intended answer precisely because the case was selected for being 'hard.' Table 7 makes the risk concrete: Context and Texture collapse on Q3 (expected ‘no’ for the replaced source in the Edited image), where an undetected remnant would make the intended answer wrong and penalize all models. With scores below 32%, even a 5–10% ground-truth error rate could account for a substantial portion of the reported weakness, so the benchmark difficulty may reflect annotation error rather than genuine prior-conflict failures.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces SABRE, a pipeline that converts a natural-language Test Primer and data schema into structured sample specifications, generates or edits images, builds question-answer pairs, filters out candidates that a Filtering VLM answers correctly, and sends the remaining candidates through human verification with localized image repair. The authors instantiate SABRE-Prior, a 600-image, 1,000-question benchmark with Context, Texture, Attribute, and Language Elicitation subsets, and report that six frontier VLMs score between 17.8% and 31.3% macro accuracy. They also report a real-image Attribute control, two 20-sample pilots for Counting and Spatial reasoning, and a user study of the repair module, and argue that SABRE is a reusable framework rather than a single fixed benchmark.","tokens_in":20664,"tokens_out":6213,"duration_ms":61997,"significance":"If the central claims hold, the paper makes a useful contribution: it proposes a modular, scalable pipeline for constructing controlled VLM stress tests, and the SABRE-Prior instantiation covers several distinct prior-conflict phenomena with a shared verification workflow. The paper is creditable for acknowledging that the Filtering VLM's own low score is partly by construction, for including a real-image control, for publishing illustrative cost and human-time estimates, and for providing detailed appendices on the sample schema and annotation interface. The main risk is that the human verification stage is the only quality gate for reference answers, and its reliability is not established; because the empirical claim that models fail to follow visual evidence depends on the correctness of those reference answers, this is a load-bearing gap. The two 20-sample pilots and the 20-session user study are small but are presented as pilot evidence, which is acceptable if their limitations are stated clearly.","major_comments":[{"comment":"The paper reports no inter-annotator agreement, no reviewer calibration, and no independent audit of the final 1,000 samples, even though human verification is the sole quality gate between pressure-filtered candidates and the final benchmark. The annotation interface displays the intended reference answer, the filtering model's prediction, and whether that prediction was correct, which can bias reviewers on borderline cases; a ground-truth error rate of even 5–10% could account for a substantial fraction of the reported 17.8–31.3% accuracy figures. Please report inter-annotator agreement on a held-out set, perform an independent blind re-verification of the final benchmark, and report the resulting error rate.","section":"§3.5, Appendix D.1, Figure 9"},{"comment":"Under the strict All4 metric, Context accuracy is near zero largely because models answer Q3 (expected 'no' for the source object in the Edited image) with 'yes' at rates of 0–33%. The paper interprets this as failure to suppress the world prior, but Appendix D.2 states that edit models 'may leave visible remnants of the original entity,' and if any final Edited image still contains a recognizable source remnant, the intended answer is wrong and all models are penalized. Without an audit confirming the absence of source remnants in the final Context cases, the near-zero Q3 accuracy is not yet interpretable as a world-prior failure.","section":"Table 7, Context Q3"},{"comment":"The real-image Attribute control contains only 20 cases, reports no confidence interval, and is evaluated on a single model and a single subset. The comparison of 30% versus 26% is presented as evidence that benchmark difficulty is not driven by generated-image artifacts, but the sample size is too small to support that claim statistically. Please expand the control, report confidence intervals, and ideally cover additional subsets and models before drawing this conclusion.","section":"Table 3, §5.5"},{"comment":"The Counting and Spatial pilots contain only 20 samples each, and every model scores at most 1/20 on Counting and 0/20 on Spatial. These near-floor results are too sparse to establish that the workflow 'supports other stress-test settings' without controlling for generation failures or annotation errors. At a minimum, report rejection and repair rates for these pilots, include a per-sample validity audit, and provide confidence intervals; as is, the extensibility claim rests on very thin evidence.","section":"§4.1, §5.6, Figure 7"}],"minor_comments":[{"comment":"The text reads 'an data schema' and should read 'a data schema.'","section":"§3.2"},{"comment":"The caption contains 'SABREannonation platform'; this should be 'SABRE annotation platform.'","section":"Figure 9 caption"},{"comment":"The figure caption states that whiskers are confidence intervals, but the numeric interval values are not reported anywhere; please include them in a table so that differences between models and subsets can be assessed.","section":"Figure 4, §5.1"},{"comment":"The repair-quality user study analyzes only 20 of 40 initiated sessions, a 50% completion rate; the manuscript should acknowledge this and report any available information about participant background or selection.","section":"Appendix D.3, Table 8"},{"comment":"For Context and Texture, unparseable yes/no responses are marked incorrect; please report how many responses fell into this category, since it affects the strict All4 scores.","section":"Appendix A, response parsing"}],"recommendation":"major_revision","confidential_remarks":"The main risk, as the stress-test note identifies, is that the reliability of human verification is not established: no inter-annotator agreement, no blind audit, and the annotation interface reveals intended answers and filter predictions. This is addressable with additional annotation experiments and an error-rate audit. The dataset and code are only promised for future release; for a benchmark paper, release is important for verification. The 20-sample pilots and the 20-session user study are small, but they are explicitly framed as pilots; if the main verification gap is closed, I would view the paper as suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read of SABRE. The actual contribution is a working pipeline that turns a Markdown test spec into schema-validated sample specifications, generated/edited images, and QA pairs, then filters with a VLM, verifies with humans, and repairs bad edits locally. That integration is real and the paper discloses the machinery unusually well: schemas, prompts, cost estimates, the annotation interface. The two pilots (Counting, Spatial) are thin but show the workflow is not one-task-to-one-benchmark. Credit where due: the authors explicitly acknowledge that the Filtering VLM's own low score is partly by construction, and they test five held-out models. The Base-vs-Edited probe analysis is useful: models read the base image fine and collapse after editing, which suggests the intervention is doing something.\n\nThe soft spot is exactly where the stress-test note lands: the ground truth. Every candidate that survives pressure screening is accepted only if a reviewer says so (Section 3.5, Appendix D.1). The interface shows the intended answer, the filter model's prediction, and whether the filter was right. No inter-annotator agreement or calibration is reported. Given that a 5–10% error rate in the reference answers could move the reported 17.8–31.3% numbers materially, the headline difficulty of SABRE-Prior should be treated as provisional. The Context Q3 pattern (expected 'no' after an edit; models score 0–33%) is precisely where a residual of the replaced object would make the reference answer wrong and penalize every model for being correct. This is a real threat to the empirical claim, not a nitpick.\n\nOther gaps are smaller: the real-image Attribute control has 20 cases and no confidence intervals; the pilots are 20 samples each; dataset and code are promised but not released. None of these sink the methodology claim. The pipeline can still be a useful, reusable tool if the verification step is made auditable and the artifacts are public.\n\nWho it's for: people building or maintaining VLM benchmarks, and model developers who want to know where prior-driven failures persist. It deserves a serious referee. My recommendation: you can send it out, but ask the authors for (1) code/data, (2) inter-annotator data or an independent audit on a random subset, and (3) confidence intervals on the real-image control. Without those, the paper is a well-specified promise; with them, it's a solid contribution.","headline":"A genuinely useful benchmark-construction pipeline whose headline difficulty numbers rest on an unmeasured human-verification step; worth serious review but only as a provisional contribution until data and reliability checks are released.","tokens_in":21204,"tokens_out":2801,"would_cite":true,"duration_ms":27552,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SABRE automates VLM stress-test construction; its world-prior benchmark leaves six frontier models at 17.8–31.3% macro accuracy.","keywords":["VLM stress testing","benchmark construction pipeline","world priors","visual grounding","model-in-the-loop filtering","human verification","image editing","vision-language models"],"falsifier":"Take a random sample of SABRE-Prior cases and have two independent reviewer teams, one blind to the filtering model's predictions and reference answers, re-verify the images; if a substantial fraction of reference answers are judged wrong or ambiguous, the reported 17.8%–31.3% range would overstate genuine model failure. Alternatively, rebuild the benchmark with a different Filtering VLM; if model scores rise sharply, the difficulty is filter-specific rather than general.","tokens_in":20240,"feed_emoji":"🧪","tokens_out":7067,"duration_ms":59232,"temperature":0.7,"pith_summary":"This paper tries to establish that VLM stress-test benchmarks can be produced by a reusable pipeline rather than by hand-curated fixed sets. The pipeline, SABRE, takes a one-page Markdown Test Primer and automatically produces structured sample specifications, generated or edited images, and question-answer pairs, then discards candidates that a filtering VLM answers correctly and sends the rest to human reviewers for verification and localized repair. In its first full instantiation, SABRE-Prior, built to test whether models follow visual evidence when it contradicts world priors, six frontier VLMs score between 17.8% and 31.3% macro accuracy, with a mean of 22.6%. Two small pilots for counting and spatial reasoning use the same workflow with different task descriptions, supporting the claim that the framework generalizes. A sympathetic reader should care because it argues that benchmark construction can keep pace with model development and expose weaknesses that saturated fixed benchmarks miss.","feed_headline":"Six frontier VLMs score 17.8–31.3% on new world-prior stress test","feed_subtitle":"One-page task spec becomes hard, human-verified cases; six models fall to 22.6% average.","key_machinery":"The load-bearing mechanism is the model-in-the-loop pressure screen combined with human verification. A Filtering VLM evaluates each candidate generated from a schema-validated sample specification; any candidate it answers correctly is discarded, so retained candidates are ones the filter fails. Because failure alone does not prove validity, each retained candidate must then pass human review of the image, question, and reference answer, with support for editing the question, correcting the reference answer, and repairing local image defects through a patch-based soft-blend tool. The paired Base-Edited design in Context and Texture uses four yes/no probes per case, scored only if all four are correct, and isolates whether the model updates its answer after a controlled visual intervention. Attribute and Language Elicitation apply the same pressure screen to open-ended counting and four-option multiple-choice formats, showing that the screening works across question types.","core_discovery":"The central claim is that benchmark construction itself can be mechanized: from a natural-language task design plus a data schema, SABRE generates candidate samples, pressure-filters them by discarding any that a Filtering VLM gets right, and retains only candidates that human reviewers confirm are valid, where validity means the required visual evidence is present, the edit is correctly applied, the question is unambiguous, and the reference answer matches the image. Using this workflow, SABRE-Prior places unexpected objects in familiar scenes (Context), gives objects counterfactual materials (Texture), alters canonical component counts (Attribute), and asks questions whose wording suggests an answer the image cannot support (Language Elicitation). Across six frontier VLMs, macro-average accuracy is 17.8%–31.3%, and a real-image Attribute control is comparably hard for the filtering model, which the authors take as evidence that the difficulty is not an artifact of generated images. The two additional pilots demonstrate that the same pipeline, with different Test Primers, produces challenging counting and spatial-reasoning tests, establishing SABRE as a reusable framework rather than a single fixed benchmark.","pith_inferences":["Inference: Since SABRE-Prior screens with one filter model, the benchmark may be biased toward cases that happen to fool that particular VLM; re-running with a different filter could yield different subsets and different difficulty levels, an implicit consequence the paper does not test.","Inference: The reliance on human verification with no reported inter-annotator agreement means the true validity-error rate of the benchmark is unknown; a blinded re-annotation study would test whether reference answers are unbiased.","Inference: The pipeline's illustrative cost estimate, roughly $43-$55 of API cost and 1.5-2.9 hours of human review per 100 retained Context cases, suggests that continuously refreshing benchmarks against each new model generation is economically plausible, not just technically possible.","Inference: If the world-prior failure pattern persists across refreshed instantiations, it would suggest a structural bias in VLM training, optimizing for predictive priors over image-grounded evidence, rather than a benchmark quirk."],"forward_implications":["SABRE-Prior's macro accuracy of 17.8%–31.3% across six frontier VLMs implies that current state-of-the-art models systematically fall back on world priors when visual evidence contradicts them, at least on these screened cases.","The real-image Attribute control (30% vs 26% for the filtering model) implies that the low scores are not primarily generated-image artifacts.","VCD and SoM, two visual-enhancement methods, do not improve Qwen 3.5 27B's macro-average on SABRE-Prior (19.5% and 16.8% vs 23.0%), implying that these failures resist generic inference-time fixes.","Counting and Spatial pilots, on which all six models score near zero, imply that the pipeline can generate hard stress tests for new capabilities from a changed Test Primer alone.","Because the filter is a frontier VLM and can be swapped, the pipeline can refresh benchmarks as models improve, rather than being frozen at release time."],"supporting_citations":[{"why":"Landmark real-image annotation benchmark, establishing the slow manual construction approach SABRE aims to automate.","marker":"Deng et al., 2009"},{"why":"POPE constructs object-existence questions from images automatically; a prior automated benchmark-construction method SABRE extends.","marker":"Li et al., 2023"},{"why":"AutoConverter automatically turns existing VQA questions into challenging multiple-choice sets; baseline against which automated construction is contrasted.","marker":"Zhang et al., 2025"},{"why":"TIFA evaluates text-to-image faithfulness with QA, supporting the premise that generators can fail to realize prompts and verification is needed.","marker":"Hu et al., 2023"},{"why":"T2I-CompBench++ shows compositional generation failures such as wrong counts and omissions, motivating the paper's verification stage.","marker":"Huang et al., 2025"},{"why":"PhD-CCS provides the counter-common-sense world-prior comparison benchmark in Table 1.","marker":"Liu et al., 2025"},{"why":"VLind-Bench measures language priors with DALL-E 3 counterfactual images; a world-prior benchmark compared in the experiments.","marker":"Lee et al., 2025"},{"why":"ViLP constructs out-of-distribution image-question-answer triplets probing visual language priors; comparison benchmark.","marker":"Luo et al., 2025"},{"why":"HallusionBench diagnoses language hallucination and visual illusion, including common-knowledge conflicts; comparison benchmark.","marker":"Guan et al., 2024"},{"why":"VLMBias makes controlled counterfactual modifications to familiar subjects to test canonical-knowledge override; comparison benchmark.","marker":"V o et al., 2026"}],"fun_headline_variants":["SABRE stress test: six VLMs avg 22.6% on world-prior traps","Automated VLM stress tests expose priors: 17.8–31.3% accuracy","One-page spec to 1k hard questions: SABRE hammers VLMs down to 22.6%","VLMs trust priors over pixels, SABRE shows with 17.8–31.3% scores"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's validity rests on human reviewers giving correct reference answers during verification; the paper reports no inter-annotator agreement and no test of whether showing reviewers the filtering model's predictions biased their decisions.","fun_headline_variants_meta":{"raw":{"variants":["SABRE stress test: six VLMs avg 22.6% on world-prior traps","Automated VLM stress tests expose priors: 17.8–31.3% accuracy","One-page spec to 1k hard questions: SABRE hammers VLMs down to 22.6%","VLMs trust priors over pixels, SABRE shows with 17.8–31.3% scores"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000765,"raw_usage":{"total_tokens":3443,"prompt_tokens":1048,"completion_tokens":2395,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":664,"completion_tokens_details":{"reasoning_tokens":2285}},"tokens_in":664,"tokens_out":2395,"duration_ms":17041,"temperature":1.0,"reasoning_tokens":2285,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T04:40:08.557033+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of SABRE-Prior cases and have two independent reviewer teams, one blind to the filtering model's predictions and reference answers, re-verify the images; if a substantial fraction of reference answers are judged wrong or ambiguous, the reported 17.8%–31.3% range would overstate genuine model failure. Alternatively, rebuild the benchmark with a different Filtering VLM; if model scores rise sharply, the difficulty is filter-specific rather than general.","supporting_citations":[],"review_version":1}