{"id":"7de86dda-4ca9-4af6-a179-bce8abe58ffc","arxiv_id":"2607.11228","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A generation-evolution-probing loop with Proposer and Digger agents produces DeepBiasBench, exposing deeper LVLM social biases than static image-question benchmarks.","lead":"DeepBias builds adaptive test cases for social bias in vision-language models by evolving questions with a DPO-trained generator and multi-turn skill-based rewrites. The resulting shared benchmark is substantially harder than static bias suites and surfaces residual stereotypes that single-turn tests miss.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The definitional premise that every committed Yes/No under the four constraints equals stereotype-driven bias is load-bearing and only weakly isolated from other failure modes.","rationale":"The Reader correctly identified the single most load-bearing premise: the mapping from committed answers under the four constraints to stereotype-driven bias. All quantitative claims (accuracy drops of 30–45 pp, transferability, superiority over static benchmarks) inherit this definition. The paper’s design (insufficient-evidence questions, abstention as ground truth) is reasonable and common in BBQ-style work, yet the manuscript supplies no control that would falsify alternative explanations for the same surface errors. The proposed silhouette / option-order test is a minimal, concrete check that would settle whether the concern lands. Because the empirical pipeline itself is carefully executed and the concern is definitional rather than a contradiction of the reported numbers, the verdict remains CONDITIONAL; the stress test simply confirms that the Reader’s weakest-assumption diagnosis is the right one and does not invent a stronger objection.","tokens_in":23272,"tokens_out":581,"duration_ms":6318,"concrete_test":"On a stratified sample of 200 DeepBiasBench instances where a non-anchor model gave a committed answer, re-run the identical image–question pair after (a) replacing the face with a demographically neutral silhouette or blank avatar while keeping all other pixels fixed, and (b) swapping the Yes/No option order. If the committed-answer rate remains within 5 pp of the original, the stereotype interpretation is supported; if it drops sharply under either control, a substantial fraction of the reported “bias” is confounded and the headline claim weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (Tables I, III, VI, VII) equates lower abstention accuracy with deeper social bias. That equation rests entirely on §III-A: under insufficient evidence + harmful individual attribution + no probabilistic wording + no explicit demographic labels, any Yes/No is counted as stereotype-driven bias (bias rate = 1 − accuracy). The paper never isolates this from alternative failure modes that produce the same surface behavior—visual misrecognition of the synthetic SDXL faces, option-format sensitivity, multi-turn instruction drift, or generic refusal collapse under progressive rewriting. Manual quality checks (§IV-E) only verify that questions are “measuring harmful social biases” and that abstention is correct; they do not audit whether the model’s committed answers actually track demographic stereotypes rather than other confounds. Without that isolation, the 30–45 pp drops and the superiority over VLBiasBench/SB-Bench could partly reflect non-bias failure modes that the adaptive pipeline is especially good at eliciting.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes DeepBias, an adaptive two-agent framework for probing social biases in LVLMs. A ProposerAgent expands and DPO-adapts a multimodal seed set (VLBBQ, derived from BBQ Age/Race/Gender) toward target-model failure modes; a DiggerAgent then rewrites each instance over multiple turns using a curated skill library (deepening and rewriting families), conditioned on prior model responses. Using an ensemble of five anchor LVLMs, the authors construct DeepBiasBench (~55k instances) and report full-pipeline trajectories, a DPO ablation (Table III), cross-model transfer (Table IV), a 500-sample manual quality audit (95.2% pass), and comparisons showing lower abstention accuracy and larger model separation than VLBiasBench and SB-Bench (Tables VI–VII). Bias is operationalized as any committed Yes/No under four constraints (§III-A), with abstention as the sole correct answer.","tokens_in":23561,"tokens_out":1475,"duration_ms":25238,"significance":"If the central claim holds, the work supplies a concrete evolutionary alternative to static VL bias benchmarks and a reusable construction pipeline (Proposer DPO + multi-turn skill probing + anchor voting). Strengths include a clean DPO ablation, transfer experiments without regeneration, explicit separation of anchor vs non-anchor results, comparison against two BBQ-derived VL benchmarks, and a documented manual audit. The contribution is timely for LVLM safety evaluation, where static suites risk saturation and leakage. The main scientific value is the demonstration that distribution-level adaptation plus instance-level multi-turn rewriting can substantially reduce abstention rates relative to seed and existing static sets, while remaining partly transferable across families.","major_comments":[{"comment":"§III-A and all accuracy/bias-rate results: the protocol equates any committed Yes/No under the four constraints with stereotype-driven social bias (bias rate = 1 − abstention accuracy). This is load-bearing for Tables I, III, VI, and VII. The manuscript does not isolate this surface behavior from confounds that the adaptive pipeline is especially likely to elicit—visual misrecognition of SDXL faces, option-format sensitivity, multi-turn instruction drift, or generic refusal collapse under progressive rewriting. The §IV-E audit checks that questions measure harmful bias and that abstention is correct; it does not audit whether committed answers track demographic stereotypes (e.g., directionally consistent with Age/Race image variants) rather than non-bias failures. A targeted analysis—e.g., stereotype-direction consistency across the multi-image setup, single-turn vs multi-turn control re","section":"§III-A; Tables I, III, VI, VII"},{"comment":"§III-C and Table I (Deep 1–3): the largest accuracy reductions come from DiggerAgent multi-turn rewriting, yet there is no ablation of the skill library (Deepening vs Rewriting families, or individual skills) and no control that rewrites questions for length/complexity without bias-oriented skills. Without this, it remains unclear how much of the Deep-stage drop is due to the curated bias-probing skills versus generic multi-turn pressure or question difficulty. A minimal skill-ablation or non-bias rewrite control on Align 2 would make the instance-level contribution more interpretable and support the claim that the skill library specifically deepens social bias.","section":"§III-C; Table I"},{"comment":"§III-D and Table VI (anchor block): DeepBiasBench is optimized via DPO preference voting and Digger feedback against the same five anchors later scored in the lower block; anchors therefore score lower by construction. The authors acknowledge this and report non-anchor models, which is appropriate. Still, the main claim that DeepBiasBench is a general challenging benchmark would be clearer if primary headline numbers and rank analyses emphasized non-anchor models only (or held-out construction anchors), and if at least one fully held-out construction run (different anchor set) were reported to quantify how much of the difficulty is ensemble-specific versus shared. Table IV transfer helps, but does not fully replace a held-out construction check for the released benchmark.","section":"§III-D; Table VI"}],"minor_comments":[{"comment":"Fig. 3 and Table II are helpful; consider adding one full multi-image (Age/Race) trajectory with model answers per demographic variant so readers can see whether committed answers align with stereotype direction.","section":"Fig. 3; Table II"},{"comment":"§IV-A: free parameters (2 DPO rounds, T=3, 2000 candidates, ≥3/5 anchor vote, dedup threshold) are deferred to the supplement; a short sensitivity summary in the main text would improve reproducibility for readers who only see the main paper.","section":"§IV-A"},{"comment":"Table V non-monotonic accuracy within Proposer/Digger stages is explained, but a brief note in the table caption would prevent misreading as instability of the method.","section":"Table V"},{"comment":"Related work on adaptive red-teaming is solid; a clearer one-paragraph contrast with RedHit (preference + iterative refinement for jailbreaks) already present could be tightened to stress the bias-specific metric and multi-image demographic controls.","section":"§II-B"},{"comment":"Limitations (§IV-H) correctly flag compute cost and synthetic images; stating approximate GPU-hours for a single-target run in the main experimental protocol would help practitioners.","section":"§IV-H"},{"comment":"Minor consistency: “LLaV A” spacing/encoding and “Behaviour” vs “Behavior” in skill names should be normalized; check “seperately” → “separately” in Table VI caption.","section":"Throughout; Table VI"}],"recommendation":"major_revision","confidential_remarks":"The skeptic’s concern about the Yes/No = stereotype premise is the main correctness-risk issue; it is standard BBQ-style scoring, so I would not reject on that alone, but for a serious journal the multi-turn adaptive setting amplifies confounds enough that isolation analyses should be required before accept. Novelty relative to adaptive jailbreak/red-teaming is real but incremental; the VL bias-specific framing and DeepBiasBench release are the main differentiators. Scope fits cs.CY / AI safety evaluation venues well."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful takeaway is that DeepBias is a real, working two-level adaptive probe for residual social bias in LVLMs, not just another static BBQ clone. ProposerAgent DPO-shifts the test distribution toward model-specific failures; DiggerAgent then multi-turn rewrites with a curated skill library. The resulting DeepBiasBench is harder and more discriminative than VLBBQ seed, VLBiasBench, and SB-Bench (Tables I, III, VI, VII). That is the new piece: specialized combination of preference-optimized generation plus skill-driven instance probing under an abstention-correct protocol, plus an ensemble construction that tries to capture shared rather than idiosyncratic failures.\n\nWhat they did well is the experimental hygiene. Full pipeline trajectories, clean DPO ablation (Table III shows Digger alone is weaker without the adapted pool), cross-family transfer (Table IV), head-to-head static comparison, and a 500-sample manual audit at 95.2% pass. Non-monotonic accuracy under ensemble voting is explained honestly. Citations cover the right red-teaming and preference-evolution literature without padding. Free parameters (DPO rounds, T, pool size, voting threshold) are stated and justified in the supplement.\n\nThe stress-test concern is real but proportionate. Every accuracy number rests on §III-A: under the four constraints, Yes/No = stereotype-driven bias. They never fully isolate visual misrecognition of SDXL faces, format sensitivity, or multi-turn drift. Manual checks only confirm question intent and that abstention is correct, not that committed answers track demographics. Still, the protocol is explicit, the same metric is used on the static baselines they beat, and transfer/ensemble results make pure artifact less likely. Circularity on the five anchors is acknowledged and partially mitigated by non-anchor evaluation. Other soft spots—synthetic images only, three demographic axes, high GPU cost, artifacts not yet public—are limitations, not load-bearing cracks.\n\nThis is for people who build or audit multimodal safety evals. It deserves a serious referee. I would engage with it and expect to cite the framework and the bench once the code drops.","headline":"Solid adaptive bias-evaluation pipeline for LVLMs; the main claim holds, with one definitional soft spot that does not sink the empirical results.","tokens_in":24173,"tokens_out":535,"would_cite":true,"duration_ms":7103,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Static bias tests miss how vision-language models fail; adaptive multi-turn probing exposes deeper stereotypes.","keywords":["vision-language models","social bias evaluation","adaptive benchmarking","multi-turn probing","Direct Preference Optimization","DeepBiasBench","agentic red-teaming"],"falsifier":"If the same models, after the full DeepBias pipeline, still abstain at rates comparable to the seed VLBBQ set on DeepBiasBench, or if manual audit shows most committed answers are visual misreads rather than stereotype judgments, the claim that adaptive probing uncovers deeper bias collapses.","tokens_in":24096,"feed_emoji":"🔍","tokens_out":574,"duration_ms":6204,"temperature":0.7,"pith_summary":"Large vision-language models often look safe on fixed bias quizzes because a single surface question can be refused or answered carefully. This paper argues that those static tests systematically understate residual social bias. DeepBias instead runs a closed loop: one agent generates and reshapes image-question sets toward whatever fails a target model, while a second agent rewrites each question over several turns, conditioned on the previous answer, using a library of deepening and rewriting skills. The same process, run against an ensemble of five models, yields DeepBiasBench. On that benchmark, models that nearly saturate older static sets drop sharply in abstention accuracy, showing that adaptive, multi-turn probes uncover biases that fixed single-turn tests leave hidden. The practical claim is that bias evaluation itself should evolve with the model rather than remain a frozen checklist.","feed_headline":"Static bias quizzes miss how vision models still stereotype","feed_subtitle":"Adaptive multi-turn probes cut abstention accuracy by tens of points and still transfer across models","key_machinery":"The generation-evolution-probing loop: a ProposerAgent that expands and DPO-adapts candidate image-question distributions toward model failures, coupled with a skill-driven DiggerAgent that multi-turn rewrites each question using a curated library of rewriting and deepening strategies conditioned on prior answers.","core_discovery":"When test data are adapted at the distribution level by preference optimization on a target model's biased responses, and then each instance is further rewritten over multiple response-conditioned turns, the resulting probes expose substantially more social bias than the original seed or existing static vision-language bias benchmarks, while still transferring across model families.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Adaptive probes expose deeper social biases than static LVLM tests","DeepBias rewrites probes turn-by-turn to reveal model-specific stereotypes","Preference-optimized generation digs past surface bias in vision models","Multi-turn skill-driven rewriting uncovers more LVLM social vulnerabilities","Evolving test cases transfer bias failures across vision-language models"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"Any Yes or No answer on a carefully constrained three-way question is counted as stereotype-driven bias rather than visual error, format sensitivity, or simple instruction failure.","fun_headline_variants_meta":{"raw":{"variants":["Adaptive probes expose deeper social biases than static LVLM tests","DeepBias rewrites probes turn-by-turn to reveal model-specific stereotypes","Preference-optimized generation digs past surface bias in vision models","Multi-turn skill-driven rewriting uncovers more LVLM social vulnerabilities","Evolving test cases transfer bias failures across vision-language models"]},"model":"grok-4.5","effort":"low","cost_usd":0.004076,"raw_usage":{"total_tokens":1232,"prompt_tokens":781,"num_sources_used":0,"completion_tokens":94,"cost_in_usd_ticks":40760000,"prompt_tokens_details":{"text_tokens":781,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":357,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":781,"tokens_out":94,"duration_ms":4440,"temperature":1.0,"reasoning_tokens":357,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T06:08:37.175730+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If the same models, after the full DeepBias pipeline, still abstain at rates comparable to the seed VLBBQ set on DeepBiasBench, or if manual audit shows most committed answers are visual misreads rather than stereotype judgments, the claim that adaptive probing uncovers deeper bias collapses.","supporting_citations":[],"review_version":1}