{"id":"9c1a02e7-7a77-4400-a869-f040381006d0","arxiv_id":"2608.09624","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A harmfulness score that ranks harmful intent well can rank real jailbreak successes below failed attempts, reversing the ranking a safety filter depends on.","lead":"The paper tests whether AI safety scores that rank harmful prompts can also predict which jailbreak attempts will actually succeed. It finds the opposite on matched prompts: successful jailbreaks score lower than failed ones, so a filter using such a score blocks attacks that would have failed first.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Automated outcome labels with low inter-judge agreement are the load-bearing link between the observed anti-ranking and the claim about realized jailbreak success; a human audit is needed to confirm the inverse AUROC is not a judge artifact.","rationale":"The reader's weakest_assumption already identifies the same load-bearing premise: automated outcome labels with low inter-judge agreement and no human audit. My stress-test confirms this is the most consequential point because the central empirical claim is defined entirely against Ym. If the labels are systematically biased in a way correlated with the score, the observed anti-ranking could be an artifact; the paper's own limitations section admits judge choice changes Ym, and three-judge kappa is low. The paper has substantial independent support: a matched factorial design that controls for base goal, a positive control showing success information is present in the internal features, multiple measurement channels reproducing the reversal, target and attack-family robustness, and an explicit checkpoint-defect repair. These controls make the construct-level argument credible, and the conceptual claim that harmfulness validation does not license success prediction is logically sound even without the empirical reversal. However, the quantitative headline—an outcome AUROC of 0.220 and a filter that admits 23/27 successful attacks—cannot be fully trusted until the success labels themselves are validated by humans. A concrete audit of a modest sample of completions would settle whether the anti-ranking is a property of jailbreak success or of the automated judge. Because this concern matches the reader's assessment and does not change the conditional verdict, I recommend no verdict change.","tokens_in":24474,"tokens_out":4686,"duration_ms":46258,"concrete_test":"Blinded human audit: take the 400 matched factorial completions (or at minimum the 100 wrapped harmful prompts with the 27 Llama Guard positives), have 2-3 independent annotators apply a pre-registered success/failure rubric, measure human-human and human-machine agreement, then recompute RY and the 5% FPR filter admission counts (23/27) under the human labels. If the human-labeled RY confidence interval remains below 0.5 and the filter still admits most successes, the label concern is resolved. If the interval moves toward or above chance, the anti-ranking finding is judge-dependent and should be re-scoped in the title and abstract.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Every below-chance outcome AUROC, including the primary RY=0.220 on 27 successes, is computed against automated outcome labels (Llama Guard 3 primary, Qwen3 secondary), with no blinded human audit. The paper itself reports Fleiss kappa around 0.3 across judges and states that judge choice changes Ym, and the Robustness section shows a judge substitution can change which cells are below chance. The central quantitative claim is therefore only as secure as the labels. If Llama Guard systematically mislabels a subset of completions—for example, lengthy fluent refusals as successes or hedging non-answers as failures—and that subset has lower AAP scores, the inverse outcome AUROC could be an artifact of judge error rather than a genuine anti-ranking of realized success. With only 27 positives in the primary cell, a small number of label flips can move the confidence interval substantially. The paper's judge-robustness checks use two automated models that may share the same systematic biases; consensus labels between them do not constitute an independent ground truth. This concern does not undercut the conceptual thesis that harmfulness validation is insufficient for success prediction, nor the positive control showing that outcome information exists in the internal features. It does, however, weaken the paper's strongest demonstration that a harmfulness score anti-ranks success: that demonstration is conditional on an unverified outcome measure.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Active Attention Probing (AAP), a fixed-coordinate attention readout, and uses it to audit whether an internal harmfulness score validated on prompt-level harmful intent also predicts realized jailbreak success. On a matched plain/wrapped factorial design over 100 harmful goals, harmful-intent AUROC falls from 0.936 to 0.803 while the Llama target's harmful generation rate rises from 0.05 to 0.27, and among wrapped harmful prompts the outcome AUROC is 0.220 (95% CI [0.109, 0.343]), placing successful attacks below failed ones. The reversal is reproduced across rare-token, passive, system-span, and refusal-logit channels, across three target models and seven attack families, and under two automated judges. The paper also quantifies the effect on a fixed-FPR filter, separates ranking, calibration, and threshold transfer under distribution shift, reports an outcome-supervised positive control, and discloses a checkpoint defect and its repair in the appendix.","tokens_in":24704,"tokens_out":8763,"duration_ms":81021,"significance":"If the result holds, this is a significant negative result for safety-filter validation: a strong harmful-intent AUROC does not license success prediction, and a filter tuned to a 5% benign false-positive rate can spend its budget mostly on attacks that would have failed while admitting 23 of 27 realized jailbreaks. The design has notable strengths: a matched factorial structure, split-before-transform distribution shifts, clustered bootstrap confidence intervals, prespecified protocols, a harmfulness score frozen before any target outcome is observed, multiple measurement channels, and unusually candid reporting of a pipeline defect and of mechanical nulls for per-head analyses. I also find no circularity in the main estimate: the harmfulness readout is trained on BeaverTails and no target outcome is used to select the probe, heads, or readout. The central blocker is the quality of the automated outcome labels, which the paper itself reports as having low three-judge agreement and no human audit; that concern lands and is the basis for my major comments.","major_comments":[{"comment":"The primary outcome AUROC RY = 0.220 (CI [0.109, 0.343]) and the deployment gap of −27.7 points at a nominal 5% FPR are defined against automated outcome labels from Llama Guard 3, with no blinded human audit. The paper reports three-judge agreement of only Fleiss κ ≈ 0.3 and states that judge choice changes Ym. With only 27 positives in the wrapped harmful cell, a small number of label flips can move the confidence interval and the filter recall gap appreciably. The Qwen3 relabeling and the consensus/union checks reduce the worry but do not eliminate it, because two automated judges can share systematic biases such as scoring long fluent refusals as successes or hedging non-answers as failures. I request (i) a human audit on a stratified sample of completions with per-judge agreement, and (ii) a label-noise sensitivity analysis showing how many random or systematic flips are needed to bring RY and the recall gap to chance. Until then, the strongest claim should be stated as an anti-ranking of judge-labeled success rather than realized jailbreak success.","section":"§5.2, §6 (Limitations)"},{"comment":"The pooled robustness estimate RY = 0.313 (CI [0.212, 0.422], p = 0.0006) over twenty-three attack-plus-target cells mixes cells where Llama Guard served as the attack-search judge with cells where it is an independent outcome judge. For GCG, AutoDAN, PAIR, and TAP, Llama Guard guided the search; the paper discloses this and provides Qwen3 as an independent judge for those rows. However, the main pooled count and the statement that six of fourteen attack cells are below chance include the non-independent Llama Guard labels. Please report the pooled estimate and the cell counts restricted to independent judge rows as a sensitivity analysis, and state explicitly whether the robustness claim survives that restriction.","section":"§5.2, Table 2"},{"comment":"The outcome-supervised readout reaching AUROC 0.930 on wrapped harmful prompts shows that the internal features are predictive of Llama Guard verdicts, but it does not validate those verdicts as ground truth. Because the positive control is trained and evaluated on the same automated labels, the statement that 'success information is available before generation' is as judge-relative as the inverse ranking. A human-audited subset would allow the paper to calibrate the judge's error rate and would make the positive control interpretable as evidence about actual compliance rather than about label reproducibility. Without it, the positive control and the inverse result are jointly consistent with a shared systematic label bias.","section":"§5.2 / Appendix C (positive control)"}],"minor_comments":[{"comment":"The three-panel figure caption is extremely dense; please split it into individual panel captions and define RH, RY, τH, τW, and τHW in the caption itself.","section":"§1, Figure 1"},{"comment":"The sentence 'Qwen3 identifies 14 successful prompts, all within the 27 prompts identified by Llama Guard. Their agreement has κ=0.61' should specify that the agreement is between the two judges' labels on the relevant prompt-completion set; as written, 'their' is ambiguous.","section":"§5.2"},{"comment":"The statement that amplification is attributable to 'the optimization in Eq. (1)' points to the main-text attention-response equation, while the optimization objective appears as Eq. (4) in Appendix A; please renumber or cross-reference correctly.","section":"Appendix I"},{"comment":"The sentence stating that at nominal rates of 0.1% and 1% 'the threshold sits above every score in this cohort and nothing is blocked' would be clearer if the score-scale shift between the calibration corpus and JailbreakBench were explained in the main deployment paragraph, so that readers do not infer that the detector never blocks.","section":"Appendix D"},{"comment":"The column header 'Effective P/W' is undefined in the caption; please define P/W as plain/wrapped effective counts.","section":"Table 7 (Appendix)"}],"recommendation":"major_revision","confidential_remarks":"This is a methodologically strong paper with an unusually honest limitations section and a clean audit design. The decision hinges on whether the field will treat automated judge labels as ground truth for the headline anti-ranking result. I would ask for a human audit on a stratified subset and a label-noise sensitivity analysis before acceptance; if those are infeasible, the authors should reframe the headline as an anti-ranking of judge-labeled success, which would weaken but not destroy the paper's contribution. The checkpoint-defect disclosure and the mechanical-null analysis for the per-head claims are exemplary and should be preserved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline result is real for the score family they built: on matched plain/wrapped harmful prompts, the AAP harmfulness score puts realized successes below failures (outcome AUROC 0.220, CI [0.109, 0.343] on the Llama cell). That is not a small effect, and it survives a lot of stress: three target models, seven attack families, two independent judges, five probe seeds, four measurement channels, and normalization controls. The positive control—an outcome-supervised readout on the same features reaches 0.930 out-of-fold—shows the coordinates contain success information, so the anti-ranking is not just a blind score. I can't find circularity: the harmfulness readout is frozen from BeaverTails, and no target outcome selects the probe, heads, or readout.\n\nWhat's genuinely new is the conceptual separation of harmful intent from realized success, and the evidence that the former does not license the latter. The field has been sloppy about this; the paper makes the distinction concrete and operational. The matched factorial design, the both-sided transformation control for RQ2, the separation of ranking/calibration/threshold, and the full disclosure of the checkpoint defect in Appendix B are all exemplary. The per-head mechanical null is handled honestly.\n\nThe soft spot is the outcome labels. The paper says it plainly: no blinded human audit, three-judge agreement low (Fleiss kappa ~0.3), and judge choice changes Ym. The low kappa is driven mostly by the broken StrongREJECT substitute they exclude; Llama Guard and Qwen3 agree at kappa 0.61 on the main wrapped Llama cell, and the harmfulness score's outcome AUROC is identical under both judges (0.2201 vs 0.2202) even though Qwen3 labels only 14 of the 27 successes. That is hard to square with a judge-label artifact. Still, the quantitative headline is only as secure as the automated labels; with 27 positives, a few mislabels could move the interval. A human audit on a few hundred completions would nail it. The title also overstates scope slightly—Llama Guard's own prompt flag sits at chance (0.507), not inverse—but the claim is about internal harmfulness scores, and there the evidence is solid.\n\nThis paper deserves a serious referee. The referee should ask for the human audit and a tightened scope statement, but the design and the core anti-ranking result should stand. I'd cite it and bring it to the reading group.","headline":"A careful matched audit shows harmfulness-validated internal scores can anti-rank jailbreak success; the only load-bearing soft spot is the automated outcome labels.","tokens_in":25270,"tokens_out":3476,"would_cite":true,"duration_ms":30976,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper shows that an internal harmfulness score validated on harmful-intent separation can rank realized jailbreak success in the opposite direction, so harmfulness validation does not establish validity for success prediction.","keywords":["LLM safety","harmfulness scoring","jailbreak success prediction","attention probing","AUROC validation","prompt filtering","construct validity","outcome validity"],"falsifier":"Relabel the 100 wrapped harmful Llama completions with blinded human raters and recompute the outcome AUROC; if the result is at or above 0.5, the anti-ranking claim as measured here does not hold.","tokens_in":24260,"feed_emoji":"⚠️","tokens_out":4769,"duration_ms":40456,"temperature":0.7,"pith_summary":"The paper tests whether a prompt filter's harmfulness score can also predict whether a jailbreak will actually succeed. It argues no: harmful intent is a property of the prompt, while realized success is an outcome produced by a specific target model, decoding policy, and judge. On matched plain and wrapped prompts, the score kept ranking harmful intent well but ranked successful attacks below failed ones, with an outcome AUROC of 0.220 on Llama. The paper concludes that harmfulness validation does not establish validity for success prediction and that safety audits should measure both quantities separately.","feed_headline":"Harmfulness scores can rank jailbreak success backwards","feed_subtitle":"On Llama, a filter tuned for harmful intent admits 23 of 27 realized jailbreaks while blocking failed attacks.","key_machinery":"Active Attention Probing (AAP) is a short learnable continuous probe inserted at a fixed system-region coordinate, so the attention response is read from a content-independent position instead of a prompt-dependent one. The matched plain/wrapped factorial design crosses harmful intent with wrapper presence for the same base goals, and the audit separates ranking (AUROC), calibration (ECE and Brier), and threshold transfer (F1 at fixed thresholds). This design is what lets the paper attribute score changes to the wrapper rather than to a moved measurement coordinate.","core_discovery":"The central discovery is that a prompt-side harmfulness score validated by harmful-intent AUROC can place successful jailbreaks below failed ones when evaluated against realized outcomes on a fixed target and judge. On Llama, wrapping 100 harmful JailbreakBench goals raised harmful generation from 0.05 to 0.27 while harmful-intent AUROC fell from 0.936 to 0.803, and among the 27 wrapped successes the outcome AUROC was 0.220, below chance. The reversal persisted across measurement channels, target models, attack families, and two independent judges, and an outcome-supervised readout on the same internal features reached 0.930 out-of-fold, showing that success information was available before generation even though the harmfulness score did not use it.","pith_inferences":["Editorial inference: the results suggest harmfulness AUROC and success-prediction AUROC may trade off, so improving a detector's intent ranking need not improve, and could even reverse, its ability to rank realized jailbreaks.","Editorial inference: a testable extension is to run the same matched audit on models with different safety-training regimes to see whether the reversal is tied to refusal behavior rather than to wrapper style.","Editorial inference: the deployment-facing implication is that prompt filters should be audited with realized outcome labels at their target operating point, since calibration and threshold transfer fail before ranking under distribution shift."],"forward_implications":["A prompt filter tuned on harmfulness at a 5% benign false-positive rate blocks 42.5% of failed attacks but only 14.8% of successful ones, admitting 23 of 27 realized jailbreaks on Llama.","Safety evaluations should report realized-success AUROC per target model, decoding policy, and judge, rather than only harmful-intent AUROC.","An outcome-supervised readout on the same internal coordinates reaches 0.930 out-of-fold on wrapped harmful prompts, so target-specific success information is available before generation even though the harmfulness score does not use it.","Distribution shift can degrade calibration and threshold transfer before it degrades ranking, so a frozen threshold's F1 can drop sharply while AUROC stays high."],"supporting_citations":[{"why":"JailbreakBench supplies the 100 harmful and 100 benign base goals that are paired in the matched plain/wrapped audit.","marker":"[6]"},{"why":"BeaverTails provides the training split used to learn the probe and the harmfulness readout before the audit.","marker":"[14]"},{"why":"Llama Guard 3 is the primary automated judge that labels realized completion outcomes, defining the outcome AUROC.","marker":"[13]"},{"why":"Llama-3.1-8B-Instruct is the primary target model whose generations realize the judged outcomes.","marker":"[9]"},{"why":"Prior evidence that success-label probes transfer poorly across attack families motivates the question of what harmfulness labels can support.","marker":"[17]"},{"why":"Prior evidence that jailbreaks can suppress internal harmfulness features frames the representational-decoupling interpretation of the reversal.","marker":"[5]"}],"fun_headline_variants":["Harmfulness scores can anti-rank jailbreak success","Jailbreak success falls below failure in harmfulness ranking","Prompt harm scores may rank real jailbreaks below failed ones","Safety filter tuned on intent misses actual jailbreak outcomes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the automated outcome judges label realized harm correctly enough; the paper itself reports three-judge agreement around Fleiss kappa of about 0.3, so a systematic judge bias could change the below-chance AUROC values.","fun_headline_variants_meta":{"raw":{"variants":["Harmfulness scores can anti-rank jailbreak success","Jailbreak success falls below failure in harmfulness ranking","Prompt harm scores may rank real jailbreaks below failed ones","Safety filter tuned on intent misses actual jailbreak outcomes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000189,"raw_usage":{"total_tokens":1350,"prompt_tokens":974,"completion_tokens":376,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":590,"completion_tokens_details":{"reasoning_tokens":312}},"tokens_in":590,"tokens_out":376,"duration_ms":3731,"temperature":1.0,"reasoning_tokens":312,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:53:38.356334+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Relabel the 100 wrapped harmful Llama completions with blinded human raters and recompute the outcome AUROC; if the result is at or above 0.5, the anti-ranking claim as measured here does not hold.","supporting_citations":[{"cited_title":"Pappas, Florian Tramèr, Hamed Hassani, and Eric Wong","cited_arxiv_id":null,"evidence_quote":"JailbreakBench supplies the 100 harmful and 100 benign base goals that are paired in the matched plain/wrapped audit."},{"cited_title":"BeaverTails: Towards im- proved safety alignment of LLM via a human-preference dataset","cited_arxiv_id":null,"evidence_quote":"BeaverTails provides the training split used to learn the probe and the harmfulness readout before the audit."}],"review_version":1}