{"id":"b868043c-e351-4c66-be2f-f36c4f4b7635","arxiv_id":"2508.06296","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"An automated red-teaming system is reported to elicit harmful behaviors from 37 of 41 frontier LLMs, with a new attempts-to-success metric revealing over 300-fold variation in attack difficulty across models.","lead":"This technical report presents an automated AI red-teaming tool, BET, that the authors report elicits harmful outputs from 37 of 41 state-of-the-art language models with a 100% attack success rate per model. It also proposes a fine-grained metric, the average number of attempts to succeed, which the authors say varies more than 300-fold across models, giving the safety community a way to rank models beyond binary vulnerability.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unnamed success judge makes 100% ASR and 300x difficulty spread uninterpretable; needs independent rescoring.","rationale":"The reader's weakest_assumption—that all headline numbers depend on an unstated, unvalidated judge—is exactly the most load-bearing concern I can identify. The abstract's strongest claim is a quantitative, comparative statement about vulnerability, but the dependent variable (attack success) is defined by a labeling method that is not described. Without knowing the judge, its error rates, and its consistency across models, the 100% ASR and the 300-fold spread are not falsifiable from the provided text. My proposed test (independent rescoring of sampled successes) would settle whether the judge is a harmless procedural detail or the source of the headline results. I do not see a reason to change the reader's UNVERDICTED verdict: the abstract's claims are plausible but currently unsupported. I am not raising a novelty or consensus objection; the issue is purely about internal correctness and testability of the evaluation scaffold. The absence of full text means no independent verification of other supporting elements (e.g., primitive-level analysis, third-party collaboration), but those are secondary to the core ASR/difficulty claim.","tokens_in":805,"tokens_out":1590,"duration_ms":20765,"concrete_test":"If the full report is supplied, locate the judge definition and its validation (e.g., human-rater agreement, per-model false-positive rates). Then independently re-score a stratified random sample of at least 200 outputs that BET labeled as successful attacks across the 41 models, using a different judge model plus human annotation. Compute ASR and attempts-to-success with this independent labeling. If ASR drops materially (e.g., below 90%) or the relative ordering of models changes substantially, the reported 100% ASR and 300-fold spread are artifacts of the original judge, not robust properties of the LLMs. If the full report is not supplied, the claim should remain unverified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claims—100% ASR on 37 of 41 models and a >300-fold spread in attempts-to-success—are only meaningful if the procedure that labels an output as a 'successful attack' is reliable and consistently calibrated across all 41 models and all hazard categories. The abstract does not name the judge/classifier, its false-positive rate, its agreement with human raters, or whether the same judge is used for every model. If the judge is permissive, lenient outputs count as harmful, inflating ASR and compressing or distorting the difficulty ranking. The 300-fold spread is especially sensitive because it is measured by BET's own optimization search: without an external anchor, the 'attempts-to-success' values reflect the search dynamics as much as model robustness. The abstract also generalizes from 41 convenience-selected models and selected hazard scenarios to 'universal vulnerability,' but the more load-bearing gap is the judge: if the judge is invalid, all downstream metrics—including the primitive-level vulnerability analysis—collapse. This concern does not accuse the authors of anything; it notes that the abstract alone provides no way to check the labeling step, and the full report is not available for review.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes PRISM Eval's Behavior Elicitation Tool (BET), an automated red-teaming system based on Dynamic Adversarial Optimization, and claims it achieves a 100% Attack Success Rate against 37 of 41 state-of-the-art LLMs. It further proposes a fine-grained robustness metric based on the average number of attempts required to elicit harmful behavior, reports a >300-fold spread across models, and introduces primitive-level vulnerability analysis plus a collaborative evaluation with third parties. The provided abstract, however, does not specify the judge/classifier used to label outputs as harmful, the optimization protocol, model versions, or any statistical summary. These omissions make the headline quantitative claims unverifiable from the available text.","tokens_in":997,"tokens_out":3315,"duration_ms":36153,"significance":"If the claims are correct, the work is potentially significant for LLM safety evaluation: an automated tool that achieves universal vulnerability across a large model set, a fine-grained difficulty ranking, and primitive-level attribution would be a practical contribution to red-teaming and benchmarking. The collaborative evaluation model is also a useful idea for distributed robustness assessment. However, the significance is entirely contingent on the reliability of the unspecified success judge and the reproducibility of the optimization-based metric. Without disclosure of these components, the paper cannot yet be credited with these contributions.","major_comments":[{"comment":"The central claim of 100% ASR on 37 of 41 models depends on a success judge/classifier that is not named or described. The abstract does not state whether the judge is an LLM, a rule-based classifier, or a human, nor does it report the decision threshold, false-positive rate, or agreement with human raters. Because a permissive judge would inflate ASR and distort the difficulty ranking, this is a load-bearing omission. Please specify the judge, its calibration across all 41 models, and its validation against human judgments.","section":"Abstract, first paragraph"},{"comment":"The >300-fold spread in 'average number of attempts required' is measured by BET's own Dynamic Adversarial Optimization search. The metric is therefore operationally defined by the optimizer's stopping criterion, attempt cap, and search dynamics, and it may not reflect model robustness independently of the optimizer. The abstract provides no protocol details: maximum attempts per scenario, stopping rule, number of independent runs, or variance across runs. Without these, the 300-fold spread is not an externally anchored measurement.","section":"Abstract, first paragraph (attempts-to-success metric)"},{"comment":"The generalization from '37 of 41 state-of-the-art LLMs' to 'universal vulnerability' is based on an unreported convenience sample. The abstract does not list the models, their versions, the sampling temperature, system prompts, or the hazard taxonomy. It also does not explain the four models for which ASR was not 100%. This lack of model-set specification prevents external verification and limits the generalizability claim.","section":"Abstract, first paragraph (model set)"},{"comment":"The primitive-level vulnerability analysis is introduced without defining what a 'primitive' is, how primitives are mapped to hazard categories, or how the analysis is validated. Since this analysis uses the same unvalidated success judge, it inherits the reliability concerns above. Please provide a definition of primitives, the mapping methodology, and independent validation against human red-team results.","section":"Abstract, second sentence (primitive-level vulnerability analysis)"}],"minor_comments":[{"comment":"The term 'state-of-the-art' is used without a cutoff date, source, or benchmark basis; specify the reference point.","section":"Abstract, throughout"},{"comment":"'Trusted third parties from the AI Safety Network' is vague; describe the organizations or the anonymization procedure used in the collaborative evaluation.","section":"Abstract, last sentence"},{"comment":"No comparison to existing red-teaming leaderboards or attack-success benchmarks is mentioned; adding such context would help position the contribution.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"This review is based only on the abstract because no full text was provided. The recommendation reflects the fact that the headline claims are unsupported without a described success judge and protocol. If the full text contains the missing details, the manuscript could become suitable; as it stands, the abstract overstates certainty. I would encourage the editor to obtain the full text before any final decision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The mean-attempts metric is a genuinely good idea, and primitive-level attribution could be practically useful. Binary ASR is too coarse; grading by average attempts to elicit harmful behavior is a real step forward. The distributed third-party evaluation is also a sensible way to scale red-teaming. So the abstract announces something worth reading.\n\nThe trouble is the headline numbers. 100% ASR on 37 of 41 models and a >300-fold spread are asserted with no specification of the judge or classifier that decides what counts as a successful attack. That's not a minor omission. If the judge is an LLM with even a small false-positive rate, the near-universal vulnerability and the difficulty ranking could both be artifacts. The spread is also measured by BET's own optimization loop—attempts-to-success is a function of the search process as much as the target model. Without an independent anchor (human ratings, a fixed classifier with known precision, or at least a calibration study), the 300-fold spread is a statement about BET, not about the models. The 'universal vulnerability' generalization from 41 convenience models is a weaker separate claim; the judge issue is load-bearing.\n\nThis is based on the abstract alone. If the full technical report names the judge, reports human agreement, states the attempt cap, and gives trial counts, much of my concern dissolves. But as written, the abstract cannot be checked.\n\nWorth a serious referee if the full report is complete and available—the ideas are timely and the field needs graded robustness metrics. But I wouldn't cite the numbers or use them in comparisons until I see the methods. The abstract alone should not be treated as a standalone result.","headline":"Useful metric in principle, but the 100% ASR and 300-fold spread hinge on an unspecified judge; get the full report before trusting them.","tokens_in":1557,"tokens_out":2473,"would_cite":false,"duration_ms":30738,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A fully automated red-teaming system achieves 100% attack success against 37 of 41 leading LLMs, and measures a 300-fold spread in the number of attempts required.","keywords":["LLM robustness","adversarial attacks","jailbreaking","red-teaming","dynamic adversarial optimization","attack success rate","harmful behavior","robustness evaluation"],"falsifier":"Take the same tool to a held-out set of LLMs and have human raters independently audit every response the automated judge flags. If a substantial share of flagged responses is benign, or if any model resists the optimization loop beyond the report's attempt cap, the universal-vulnerability claim is refuted.","tokens_in":638,"feed_emoji":"🔓","tokens_out":4740,"duration_ms":45236,"temperature":0.7,"pith_summary":"This report claims that a fully automated adversarial attack system can elicit harmful responses from nearly every state-of-the-art large language model it was tested against: 37 of 41 models, each with a 100% attack success rate. More than that, it argues that binary vulnerability hides a wide gradient: the average number of attempts the attacker needs varies by more than 300-fold across models. The paper introduces this attempts-to-success measure as a fine-grained robustness metric, and adds a primitive-level analysis mapping which jailbreaking techniques work best for which hazard categories. A reader should care because the work reframes the question from 'can these models be broken?' to 'how hard is it to break them, and where are the weak points?'","feed_headline":"Automated red-teaming breaks 37 of 41 top LLMs","feed_subtitle":"Attack success hits 100% on those models, with a 300-fold range in the tries needed to break each one.","key_machinery":"The central mechanism is Dynamic Adversarial Optimization: a closed loop in which the tool generates candidate attack prompts, observes the model's response, and feeds that outcome back into the next generation of prompts until a judge marks a response as harmful. The fine-grained metric is the expected number of attempts to reach that harmful response, which converts a binary attack success rate into a continuous difficulty score. The primitive-level analysis decomposes jailbreaks into attack primitives and attributes their effectiveness to specific hazard categories.","core_discovery":"The central claim is that current safety training does not harden LLMs against adaptive attacks; it only moves how many tries an attacker needs. The Behavior Elicitation Tool (BET) runs a dynamic adversarial optimization loop that repeatedly probes a model and refines its attack prompts in response to the model's outputs, reaching a perfect success rate on 37 of the 41 leading models. The paper's proposed metric, the average number of attempts required to elicit a harmful behavior, is what reveals the 300-fold spread: the most vulnerable models produce harmful content almost immediately, while the most resistant still fall after enough iterations. The paper further claims that different haza","pith_inferences":["If the attempts-to-success measure is stable across reruns, it could become a calibration target for safety training, much like error rates in supervised learning.","The 300-fold spread may partly reflect model behavior style (verbosity, refusal phrasing) rather than safety per se, because the metric counts only attempts up to the judge's first harmful label.","A direct testable extension would be to apply the same loop to multimodal or agentic LLMs; the report does not claim results there, but the mechanism is not text-specific.","The report's universal-vulnerability claim predicts that any future LLM, not just the 41 sampled, will fall to the same adaptive loop within a finite attempt budget—so new releases should be tested this way before deployment."],"forward_implications":["Current safety fine-tuning does not eliminate elicitable harmful behavior; it only changes how many attempts an attacker needs.","Models can be ranked by robustness even though all tested models are vulnerable, giving a continuous scale for safety progress.","Defenders can target specific attack primitives rather than treating all jailbreaks alike, based on the hazard category.","Distributed red-teaming is feasible: multiple third parties running the same tool can pool results into a shared leaderboard.","The 300-fold difficulty spread implies that small differences in attempted attacks can separate models in public evaluations, so evaluation protocols must fix the attempt budget."],"supporting_citations":[],"fun_headline_variants":["37 of 41 top LLMs fall to automated red-teaming","Safety training only delays attacks: 300-fold difference in tries","Automated red-teaming achieves 100% success on 37 of 41 SOTA LLMs","LLM defenses crack after enough tries: 100% ASR on 37 of 41 models","Attack difficulty varies 300-fold across LLMs despite universal vulnerability"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The headline numbers rest on the unstated reliability of the automated judge that decides when a model response counts as harmful; if that judge labels benign responses as harmful, the 100% attack success and the 300-fold difficulty ranking could both be artifacts.","fun_headline_variants_meta":{"raw":{"variants":["37 of 41 top LLMs fall to automated red-teaming","Safety training only delays attacks: 300-fold difference in tries","Automated red-teaming achieves 100% success on 37 of 41 SOTA LLMs","LLM defenses crack after enough tries: 100% ASR on 37 of 41 models","Attack difficulty varies 300-fold across LLMs despite universal vulnerability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000819,"raw_usage":{"total_tokens":3376,"prompt_tokens":654,"completion_tokens":2722,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":398,"completion_tokens_details":{"reasoning_tokens":2619}},"tokens_in":398,"tokens_out":2722,"duration_ms":17232,"temperature":1.0,"reasoning_tokens":2619,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T22:48:19.281435+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the same tool to a held-out set of LLMs and have human raters independently audit every response the automated judge flags. If a substantial share of flagged responses is benign, or if any model resists the optimization loop beyond the report's attempt cap, the universal-vulnerability claim is refuted.","supporting_citations":[],"review_version":1}