{"id":"8df1bc3e-4377-4574-9b8b-3d11a62212f6","arxiv_id":"2608.08641","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The reported benefit of a multimodal guardrail is a sum of guard blocks and the target model's own refusals, and evaluation protocol alone can move the guardrail's measured share from 0% to 99%.","lead":"A safety guardrail and the model behind it can both refuse, and this paper shows that reported guardrail safety numbers silently mix the two. The guardrail's measured share can range from 0% to 99% depending on the input channel and on what the evaluation lets the defense read, and the authors correct their own earlier figures.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Silent-default root cause in §4.6 rests on an unaudited single-implementation inspection; releasing artifacts and auditing an independent ECSO port would settle whether faithful porting really supplies the granted read.","rationale":"The reader's weakest assumption identifies the same concern I consider most load-bearing: the silent-default narrative in §4.6. The central empirical claims—that guard blocks and model refusals are disjoint and separately countable, that protocol choice changes measured benefit by tens of points, and that the ordering by read position reproduces—are supported by exact string matches, paired tests, an end-to-end replicate, and explicit correction of the authors' own prior figures. The root-cause claim that faithful porting silently grants the unencoded read is the one step that cannot be audited from the preprint, because the artifacts are promised but not released and the inspection covers one implementation. This concern does not break the measured protocol gap or the 0–99% share range; it affects only the conclusion that the inflated setting is the default rather than a choice. The proposed concrete test—releasing artifacts and auditing both the authors' harness and an independent ECSO port—would settle that narrative. Since the reader already issued CONDITIONAL on essentially this basis, my read does not change the verdict.","tokens_in":43179,"tokens_out":7447,"duration_ms":85375,"concrete_test":"Release the artifact set promised in Limitations and audit the ECSO integration: for one full protocol-grid cell (e.g., internvl3-8b, code_attack, 100 prompts), instrument the harness to log, for every prompt, the exact string placed in the TELL, CAP, and SAFE slots under the default configuration, alongside the raw attacker-sent encoded string t=Enc(q) and the dataset behavior string q. If all three slots contain q while t appears only in the target-facing INITIAL call, §4.6's silent-default claim is confirmed. Independently, clone the original ECSO repository (Gou et al. 2024), wire it to HarmBench with a request-transforming encoder, and compare which string its query/prompt field holds before the defense branches; if the original code preserves t rather than q, the paper must be revised to state that the granted protocol is an evaluator choice, not a faithful port.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The measured protocol gap and the 0%/41–45%/99% share decomposition are well supported internally: block counts are exact string matches, paired McNemar tests are family-corrected, the 32-cell replicate preserves the ordering, and the authors correct their own prior grant-protocol numbers. The load-bearing weakness is the root-cause claim in §4.6 that faithful porting of the reference ECSO implementation supplies the granted (unencoded) read silently. This is what converts the artifact from \"an evaluator could choose a flattering protocol\" to \"released code makes that protocol the default,\" and it is verified only by the authors' inspection of one implementation, reported in a preprint that promises but does not yet release artifacts. An independent reader cannot check that all three ECSO stages—TELL, CAP, SAFE—read one undifferentiated prompt field, that this field is populated with the benchmark's behavior string rather than the attacker-sent encoded string, and that no other code path preserves the distinction. If the implementation does distinguish the strings, the silent-default narrative needs revision, though the measured inflation (24–47/12–36/~0 pp) and the share range remain. The Discussion's own limitation that stage-level attribution in the fired path is not isolated (H1) does not bear on the grant isolation; the unresolved §?? cross-reference in §4.8 is cosmetic.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper measures how much of a defended pipeline's reported safety is actually produced by the guardrail rather than by the target model's own refusal behavior. On a text guard over two open-weight targets, the guard's measured share of refusals ranges from 0% (payload rendered as pixels, guard blocks 0/100), through 41–45% (guard reads the encoded prompt as sent), to about 99% (harness fills the guard's internal read with the unencoded request), while a results table would describe all three settings identically. The authors further show that granting the unencoded read inflates measured benefit by 24–47 pp for a guard gate, 12–36 pp for a caption-mediated re-check, and approximately 0 pp for a majority-vote smoother; that the inflation is carried by the answer-regenerating stage rather than the harm-verdict stage; and that the reference ECSO implementation reportedly builds all three stages from a single prompt field, so faithful porting silently supplies the granting protocol. The paper also reports channel-coverage experiments for gate and transform defenses, benign-traffic discrimination controls, a detector-gated deployment variant, adaptive-attack checks, and a correction of the authors' own previously published granted-protocol numbers.","tokens_in":43302,"tokens_out":4449,"duration_ms":55698,"significance":"If the results hold, this is a substantial methodological contribution to multimodal safety evaluation. The core decomposition is elegant and, for gate defenses, exact: guard blocks and model refusals are disjoint and separately countable, so the guardrail's share of credited safety can be recovered from any already-collected campaign. The empirical work is unusually careful: block counts are exact string matches rather than similarity heuristics; paired tests use exact McNemar with family-wise Holm–Bonferroni correction; the central protocol grid was re-collected end-to-end in a designed 32-cell replicate; the reported ordering reproduces across campaigns; multi-judge re-scoring preserves the direction and significance of every paired claim; and the authors explicitly falsify their own registered prediction that the harm-verdict stage would carry the grant effect. The paper also publicly corrects its own prior reported figures, which is a notable scientific practice.","major_comments":[{"comment":"The root-cause claim that the granted protocol is \"what faithful porting supplies, silently\" is load-bearing for the conclusion that the flattering setting is the default of released code rather than an evaluator's avoidable choice. Currently this rests on the authors' inspection of one ECSO implementation, and the artifacts promised in the Limitations are not released with the preprint, so an independent reader cannot verify that TELL, CAP, and SAFE all read one undifferentiated prompt field, that the field is populated with the benchmark's behavior string, and that no other code path preserves the attacker-sent encoded string. The Appendix A slot-level specification shows the substitution the protocol makes, but it does not show the actual code bindings. I ask the authors to either release the audited implementation together with the audit script (or an independent port inspection) before publishing the silent-default assertion, or to scope the claim to \"the implementation we inspected supplies it silently\" and mark the generalization as a hypothesis. The measured protocol gap and the 0%/41–45%/99% share range do not depend on this point and would remain.","section":"§4.6 and Limitations (Artifact release)"},{"comment":"The claim that the answer-regenerating stage (SAFE) carries the grant inflation is supported only in the two code_attack cells where the full grant effect is large; on formal_logic no stage moves significantly, and the paper itself states that no stage attribution is possible where the full effect is small. One cell (pixtral/code) shows SAFE recovering 113% of the full effect, which the authors report without clipping. This is honest, but the concluding summary in the abstract and conclusion states the mechanism as a general result: \"the stage that regenerates the answer carries the effect.\" I recommend that the summary-level wording be qualified to \"where the effect is large enough to attribute,\" matching the bounds stated in §4.8.","section":"§4.8, bounds of stage attribution"},{"comment":"Table 10 and Figure 1 contain granted-protocol ECSO values that are explicitly labeled as measurements of the artifact rather than as attainable defense efficacy, and §4.6 re-collects the deployable arm showing most of the effect does not survive. This is a deliberate and mostly well-handled frame. However, the 'amplification' language in §4.10 still invites an efficacy reading, and the section title \"Amplification: ECSO with the decoy\" does not itself carry the protocol caveat. Given that the paper's central lesson is that protocol choice manufactures benefit, the caveat should appear in the section heading or in the first sentence of every paragraph reporting these magnitudes, not only in the table captions and the initial caveat paragraph.","section":"§4.9–§4.10, granted-protocol cross-model grid"}],"minor_comments":[{"comment":"The paper contains an unresolved cross-reference \"§??\" in the paragraph beginning \"The cross-family pattern remains an ordering\"; this should be fixed to the section that explains why cross-defense orderings cannot identify read position.","section":"§4.8"},{"comment":"The caption still uses the term \"oracle protocol (unattainable)\" while the main text and tables have replaced \"oracle\" with \"granted\"; the terminology should be made uniform.","section":"Figure 1 caption"},{"comment":"The paper refers to a \"registered prediction\" and says the protocol grid's readout was fixed in advance, but no registration identifier or repository link is provided; if a registration exists, it should be cited so readers can verify the pre-commitment.","section":"§4.8 / protocol grid"},{"comment":"One reference entry contains a dated provenance annotation (\"Venue confirmed 2026-08-08...\") that belongs in a reviewer communication rather than in a published reference list; it should be removed or placed in a footnote.","section":"References"},{"comment":"The claim that the text guards are \"blind, not inaccurate\" is well supported by the 0/300 constant on the image arm, but the wording \"spanning the full harm range\" should make explicit that the harmful arm used the placeholder text channel; otherwise a reader may infer the image-channel input varied in text content across the 300 inputs.","section":"§4.3 and Table 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's central measurement is sound and carefully executed; my recommendation of major revision is driven by the unverifiable silent-default root-cause claim in §4.6, which is the one load-bearing point that cannot currently be checked because the artifacts are not released. The self-citation pattern is largely defensible given that the paper corrects its own prior figures, though several citations are to unpublished 2026 arXiv work; this is not a reason for rejection. The paper would be strengthened by including the artifact release with the revision, or by explicitly downgrading the silent-default claim to a scoped observation about the inspected implementation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThe punchline first: this is one of the few evaluation papers I've read that actually changes how I will read a results table. The authors show that a defended pipeline has two refusal producers—the guardrail and the target model's own alignment—and that every reported attack-success number is a sum over both. Because a guard block replaces the model response, the two counts are disjoint and separable for free. Across three settings that a results table would describe identically, the guard's share of the safety credited to it goes from 0% (payload rendered as pixels) through 41–45% (guard reads the encoded prompt) to 99% (harness fills the read with the unencoded request). That's new, and it's a measurement the field should be making routinely.\n\nThe evidence quality is higher than most of this literature. Block counts are exact string matches, not similarity heuristics. The benign controls show the text guards' image-channel decision is a constant across 300 inputs—not a poor detector, but a detector receiving nothing. Tests are prompt-paired with McNemar and family-corrected. There's a designed 32-cell replicate that reproduces the ordered inflation (gate largest, caption-mediated intermediate, majority-vote null). And the authors correct their own prior granted-protocol figures, including reversing a −63pp effect to −3pp. Their registered prediction that the harm-verdict stage would carry the inflation was empirically overturned; the answer-regeneration stage carries it. That's the opposite of confirmation bias.\n\nThe soft spot is the root-cause story in §4.6: the claim that faithful porting of the reference implementation silently supplies the granted read because all three stages read one undifferentiated prompt field. That is verified only by the authors' inspection of one implementation, and the artifacts aren't released. If an independent audit shows the implementation preserves the attacker-sent/benchmark-behaviour distinction, the 'silent default' narrative needs revision—though the measured inflation and the share range would stand. This is a fixable, testable weakness, not a load-bearing flaw in the main result.\n\nWho this is for: anyone evaluating multimodal defenses or citing guarded-pipeline attack-success numbers. It deserves a serious referee; I'd condition acceptance on artifact release and an independent port check.\n\nRegards,\n[You]","headline":"A guarded pipeline's reported safety benefit is a two-producer sum; this paper separates it cleanly and shows protocol choice alone moves the guardrail's share from 0% to 99%—the silent-default root cause needs artifact release, but the core result is solid.","tokens_in":43949,"tokens_out":2519,"would_cite":true,"duration_ms":26070,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A defended pipeline's reported safety benefit is a sum of two refusal producers, and the guardrail's share can range from 0% to 99% depending on evaluation choices no results table reports.","keywords":["guardrail evaluation","refusal attribution","multimodal safety","evaluation protocol","channel coverage","encoded jailbreaks","black-box defenses","attack success rate"],"falsifier":"Instrument the reference implementation of a caption-mediated guardrail to log which string actually populates each internal read slot, then run the same attack corpus under the deployable protocol; if the guard's measured benefit remains large when the read provably contains only the attacker's encoded prompt, the central protocol-artifact claim is refuted. Alternatively, if the released code is found anywhere to distinguish the attacker-sent string from the benchmark behavior string, the silent-default narrative fails.","tokens_in":1947,"feed_emoji":"🛡️","tokens_out":2166,"duration_ms":80351,"temperature":0.7,"pith_summary":"This paper is about a measurement the field is not making. A defended pipeline contains two components that can refuse—the guardrail and the target model's own alignment—and every reported attack-success number is a sum over both. The authors show the guardrail's share of that sum is not a property of the guardrail: with a text guard and two open-weight targets, the share runs from 0% when the harmful payload is rendered as pixels, through 41–45% when the guard reads the encoded prompt the attacker actually sent, to about 99% when the harness fills the guard's internal read with the unencoded request behind the attack. All three settings would be described identically in a results table. The split is recoverable at no extra cost because a guard block replaces the model's response, so the two refusal counts are disjoint.","feed_headline":"A guardrail's reported safety share swings from 0% to 99%","feed_subtitle":"Separating the guard's blocks from the model's refusals reveals which protocol manufactured the safety number.","key_machinery":"The key machinery is the refusal decomposition: for a gate defense, guard blocks and model refusals are disjoint by construction because a block substitutes a fixed canned string for the model's response, so the guard's share is exactly $B/(B+M)$, where $B$ is the exact-match guard-block count and $M$ is the model's own refusal count among passed prompts. That identity makes the split recoverable from any already-collected campaign at no extra cost. The rest of the argument rides on a second variable, the defense's internal read: the string an evaluation harness places in each of the defense's slots (the encoded prompt as sent versus the unencoded request behind it), and the position of that slot in the defense's computation—decision, caption, or answer regeneration. Comparing the deployable protocol (read equals the encoded prompt) with the granted protocol (read equals the unencoded request) produces the measured inflation, and the within-defense isolation shows that the read conditioning the generated answer, not the read making the harm verdict, carries the effect.","core_discovery":"The central discovery is that a guardrail's reported safety benefit is a sum over two refusal producers, and the guardrail's own share is decided by two variables no evaluation records: which channel carries the payload, and what text the evaluation harness places in the defense's internal read. Measured on one text guard across two open-weight targets with 100 prompts per cell, the share moves from 0% (payload rendered as pixels; guard blocks 0/100 and the model produces every refusal) through 41–45% (guard reads the encoded prompt as sent; guard blocks 28–30 against 34–44 model refusals) to about 99% (harness fills the read with the unencoded request; guard blocks 98, model falls to about 1), with no change in how a results table would describe the pipeline. Because a guard block replaces the model's response, guard blocks and model refusals are disjoint and separately countable at no extra collection cost. Granting the unencoded read inflates measured benefit by 24–47 percentage points for a gate, 12–36 for a caption-mediated re-check, and about zero for a majority-vote smoother; isolating the grant inside one defense shows the effect is carried by the stage that regenerates the answer, not the stage that decides harm. The inflated setting is not an evaluator's careless choice: the reference implementation builds every stage from a single prompt field that cannot distinguish what the attacker sent from what the benchmark records, so faithful porting supplies it silently, in the direction that flatters the defense; the authors' own previously published figures are among those revised.","pith_inferences":["An implication the authors leave implicit is that comparing guardrails across papers without a standardized audit of each defense's internal read will tend to rank evaluations rather than defenses; a shared harness that logs which string fills each slot would settle that comparison.","The finding that the answer-regeneration stage carries the grant inflation suggests a testable extension: hold the defense architecture fixed and vary only whether a stage rewrites the final response or merely selects among candidate responses, predicting that only the rewriting stage is protocol-sensitive.","The 0% end implies that a channel-routed panel—the best-calibrated text guard on the text channel and a multimodal guard on the image channel—could be evaluated end-to-end; the paper recommends it but does not build it, and the binding constraint on that deployment would be detector recall for encoded inputs.","Because the paper revises its own previously published granted-protocol figures, an editorial inference is that other published guardrail evaluations using encoded attacks may carry similar inflation; re-scoring stored responses under the deployable protocol would reveal how widespread the artifact is."],"forward_implications":["Every attack-success number reported for a defended pipeline is unreadable unless it states which channel carried the payload and which protocol filled the defense's internal read.","A guardrail can be the minority producer of the safety credited to it: in the honest deployable setting measured here, the guard produced 41–45% of the refusals while the target model supplied the rest.","Text-only guards are blind, not inaccurate, in the image channel: across 300 inputs spanning harmful, hard-benign, and ordinary benign traffic, their image-channel decision never changes, so no amount of tuning recovers a signal that never arrives.","Evaluation protocols that hand the defense the unencoded request can manufacture tens of points of apparent safety: 24–47 percentage points for a gate and 12–36 for a caption-mediated re-check, with the ordering reproduced in an independent replicate.","A defense whose read only selects among already-generated responses is protocol-robust, while a defense whose read conditions the final generated answer is protocol-sensitive."],"supporting_citations":[{"why":"Supplies the ECSO caption-mediated defense and its reference implementation, from which the paper reads the single-prompt-field root cause.","marker":"Gou et al. 2024"},{"why":"Supplies HarmBench, the harmful-prompt corpus and safety classifier used in every attack-success cell.","marker":"Mazeika et al. 2024"},{"why":"Supplies CodeAttack, one of the two encoded-attack families whose encoding creates the divergence between attacker-sent and benchmark-behavior text.","marker":"Ren et al. 2024"},{"why":"Supplies MathPrompt, the set-theory encoded-attack family that similarly forces the protocol choice.","marker":"Bethany et al. 2024"},{"why":"Supplies SAGE, the text-side sanitisation defense used for the transform-defense coverage and stacking results.","marker":"Ding et al. 2025"},{"why":"Formalizes the deployable-read convention that defines the baseline against which the granted-read inflation is contrasted.","marker":"Jia et al. 2025"},{"why":"Supplies the JailbreakBench benign split used to measure benign refusal and the utility floor for the safety-utility trade-off.","marker":"Chao et al. 2024"}],"fun_headline_variants":["Guardrail's reported safety swings 0-99% on untested variables","Guardrail's real share: 0% to 99% depending on untested eval choices","Guardrail benefit is a sum of two refusers, the split is invisible","Guardrail credit inflated by up to 47 points by eval design"],"cache_read_input_tokens":45952,"weakest_assumption_plain":"The claim that the flattering protocol is the silent default rather than an evaluator's choice rests on the authors' root-cause inspection of one released implementation, ECSO, showing that all three of its stages read one undifferentiated prompt field; the artifacts of that inspection are not released in this preprint, so an independent reader cannot yet audit that the two strings are never distinguished.","fun_headline_variants_meta":{"raw":{"variants":["Guardrail's reported safety swings 0-99% on untested variables","Guardrail's real share: 0% to 99% depending on untested eval choices","Guardrail benefit is a sum of two refusers, the split is invisible","Guardrail credit inflated by up to 47 points by eval design"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001029,"raw_usage":{"total_tokens":4474,"prompt_tokens":1221,"completion_tokens":3253,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":837,"completion_tokens_details":{"reasoning_tokens":3168}},"tokens_in":837,"tokens_out":3253,"duration_ms":24737,"temperature":1.0,"reasoning_tokens":3168,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:28:53.338303+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Instrument the reference implementation of a caption-mediated guardrail to log which string actually populates each internal read slot, then run the same attack corpus under the deployable protocol; if the guard's measured benefit remains large when the read provably contains only the attacker's encoded prompt, the central protocol-artifact claim is refuted. Alternatively, if the released code is found anywhere to distinguish the attacker-sent string from the benchmark behavior string, the silent-default narrative fails.","supporting_citations":[{"cited_title":"T.; and Zhang, Y","cited_arxiv_id":null,"evidence_quote":"Supplies the ECSO caption-mediated defense and its reference implementation, from which the paper reads the single-prompt-field root cause."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies HarmBench, the harmful-prompt corpus and safety classifier used in every attack-success cell."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies CodeAttack, one of the two encoded-attack families whose encoding creates the divergence between attacker-sent and benchmark-behavior text."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies SAGE, the text-side sanitisation defense used for the transform-defense coverage and stacking results."},{"cited_title":"J.; Tram\\` e r, F.; Hassani, H.; and Wong, E","cited_arxiv_id":null,"evidence_quote":"Supplies the JailbreakBench benign split used to measure benign refusal and the utility floor for the safety-utility trade-off."}],"review_version":1}