REVIEW 2 major objections 5 minor 23 references
Scores from four agent-safety benchmarks are not interchangeable: they rank the same models differently, reward an always-unsafe baseline, and shift with the model panel, so safety claims must attach benchmark, metric, target behavior, and
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 00:40 UTC pith:LF2DOOI5
load-bearing objection A genuinely useful metric-validity and small-panel audit; treat the misalignment-based conclusions as exploratory. the 2 major comments →
Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The paper's central claim is that no single agent-safety benchmark stands in for 'safety' as a whole. On any binary trace-judgment benchmark scored by F1, an always-positive policy attains F1=2π/(1+π); on R-Judge that baseline scores 0.690, outranking five of 21 models that actually discriminate. The three broad-coverage benchmarks rank the same 18 models differently, and a trade-off between R-Judge specificity and AgentHarm safety that looked strong at n=7 (ρ=-0.64) dissolves at n=18 (ρ=+0.02) — a quarter of random seven-model subsets manufacture |ρ|≥0.5 around that near-zero value. On the held-out side, capability predicts task success (ρ=+0.60) but correlates negatively with a misalignmen
What carries the argument
The central object is the F1 closed-form identity: on a binary trace-judgment benchmark with unsafe-class base rate π, the constant 'always positive' policy scores F1=2π/(1+π). It is the mechanism that exposes the metric-level failure — F1 gives no credit for correctly identifying benign traces, so a policy that simply labels everything unsafe can match a leaderboard. The audit's other machinery is a capability-control design: a harmonized MMLU+GPQA composite, Spearman and partial-Spearman correlations, and three pre-specified held-out criteria (task success, a fictional misalignment scenario, and a three-template jailbreak measure) that let the paper ask whether safety scores add predictive
Load-bearing premise
The audit's criterion-validity conclusions assume the three held-out outcome tests — especially the single fictional misalignment scenario, whose two halves agree only moderately — actually measure deployment safety; if that scenario is not a valid measure of agentic misalignment, the headline negative capability–safety correlation and the tests built on it would be artifacts.
What would settle it
Run the same four benchmarks and three held-out criteria on at least 40 models from new organizations. The central claim would be falsified if an always-unsafe baseline no longer beats any real discriminating model on R-Judge, if the broad-coverage benchmarks' rankings agree strongly (Spearman ≥ 0.8), or if the AgentHarm–jailbreak partial correlation drops well below 0.72. A cheaper probe: give the same models two independently written misalignment scenarios; if their model-level scores correlate below ~0.7, the misalignment criterion is not a stable construct and the correlations built on it
If this is right
- R-Judge's headline F1 should not be reported alone; two-sided metrics such as balanced accuracy are needed because F1 gives no credit for true negatives.
- Rankings from different agent-safety benchmarks are not comparable; 'safety rank' is benchmark-dependent, so a score must be quoted with its benchmark, metric, target behavior, and model panel.
- Small-panel correlations are unreliable: a quarter of random seven-model subsets can produce |ρ|≥0.5 around a near-zero full-panel value, so validity must be re-checked as the model population turns over.
- Capability and safety are not substitutes: capability predicts task success but correlates negatively with misalignment safety on the original panel; high-capability models can be less safe on that criterion.
- AgentHarm's strong association with jailbreak resistance reflects convergent validity — both measure harmful compliance — and should not be read as evidence of general safety.
Where Pith is reading between the lines
- The F1 identity generalizes: any binary trace-judgment leaderboard scored by F1 rewards positive-heavy policies, and the gap between a model's F1 and 2π/(1+π) is a conservative measure of how much discrimination the score actually encodes; reporting balanced accuracy or Matthews correlation would neutralize the artifact.
- Because the capability composite uses static-knowledge tests, the audit cannot rule out that an agentic-capability component is driving the outcome-dependent correlations; a validated tool-use capability anchor would be the natural next control.
- A testable extension: pre-register a large multi-organization panel (n≥40) of agent loops, re-run the four benchmarks and three criteria, and check whether the near-zero R-Judge–AgentHarm correlation and the AgentHarm–jailbreak partial ρ≈0.72 replicate; the paper's own pre-specified 'hold or strengthen' prediction for the misalignment correlation was not met, so an independent replication would se
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper conducts a validity audit of four agent-safety benchmarks (R-Judge, InjecAgent, AgentHarm, AgentDojo) by re-running them under official implementations and author-provided scorers on up to 22 models, with MMLU/GPQA measured under one protocol as a capability composite. It reports four main findings: (1) a metric-validity failure — on binary trace-judgment benchmarks scored by F1, an always-positive policy attains F1 = 2π/(1+π), which on R-Judge (0.690) outranks five models that actually discriminate (Observation 1, §4.1); (2) the three broad-coverage benchmarks rank the same 18 models differently, and a small-panel trade-off between R-Judge specificity and AgentHarm safety dissolves at larger n (§4.3); (3) capability predicts held-out task success (ρ=+0.60) but correlates negatively with the paper's misalignment safety criterion (ρ=−0.44, n=21), producing a crossover interaction Δ=−1.00; and (4) AgentHarm's capability-controlled association with jailbreak safety (ρ=+0.72) is the strongest held-out result but is interpreted as convergent validity of harmful-compliance measurement rather than general safety. The paper concludes that no single benchmark score can be quoted as 'an agent's safety' without specifying benchmark, metric, target behavior, and model panel.
Significance. If the results stand, the paper makes a valuable, field-level contribution: it demonstrates formally and empirically that agent-safety benchmark scores are not interchangeable measurements, and it identifies a specific metric failure (F1's blindness to true negatives) that is easy to overlook. The F1 closed form is a clean, general result. The paper is unusually careful: it uses official implementations, runs capability anchors under one protocol, reports metric sensitivity, discloses power limitations, and explicitly distinguishes pre-specified analyses from exploratory ones. The reproducibility package and API-only harness are concrete assets. The central negative claim about benchmark disagreement and the gameability of F1 is robust to the RQ3 concerns discussed below because it does not depend on any held-out criterion. The paper's treatment of the AgentHarm–jailbreak association as convergent validity rather than general safety is appropriately cautious.
major comments (2)
- [§4.4, Table 3, App. A.5] The negative capability–misalignment correlation (ρ=−0.44) and the Δ=−1.00 interaction rest entirely on one fictional blackmail/leaking scenario whose split-half reliability is only 0.70 (Spearman–Brown, Table 3) and whose construct classification is contested in the blinded coding (6/7 harm-blocking in one round vs. 6/7 risk recognition in another, App. A.1). If the probe measures compliance with a narrow blackmail script rather than general agentic misalignment, the headline 'capability predicts lower safety' result is an artifact of that scenario. The post-hoc three-scenario battery does not repair this: it was added after seeing the primary result, changes the effect size negligibly, and includes murder, which co-varies with AgentHarm and dilutes construct specificity. Please either report a pre-specified multi-scenario misalignment battery as the primary criterion, or explicitly del
- [§4.4, Tables 2 and 4] The R-Judge/InjecAgent–misalignment partials (+0.41/+0.47) are labeled exploratory in the text, but the abstract's claim that capability 'correlates negatively with misalignment safety' gives them more weight than they can bear. These cells have power ≈0.35, their confidence intervals include zero, and the R-Judge estimate is metric-fragile: +0.41 under specificity becomes +0.19 under balanced accuracy and −0.46 under recall (Table 4). Since one of the paper's central lessons is that metric choice changes conclusions, the summary should carry the same caveats as the body — or this result should be dropped from the abstract.
minor comments (5)
- [Fig. 1] The x-axis label 'models, ranked by F1' is ambiguous; consider 'model rank (1 = best F1)'.
- [Table 3] The footnote says the jailbreak criterion uses 42 models, while the body reports n=41 on the expanded panel; clarify the discrepancy (e.g., DeepSeek-R1 inclusion).
- [Table 4] The jailbreak column has n=20 while Table 2 reports n=18 for the same cell; add a coverage note so readers do not misread the panels.
- [App. A.5] The arrow notation 'AgentHarm→jailbreak' is informal; define it when first used.
- [§3.3] The term 'pre-specified' is used repeatedly. Since the manuscript correctly notes that these analyses lack an independent timestamp, consider adding a sentence stating that the frozen analysis plan is included in the artifact and that 'pre-specified' reflects the authors' internal workflow, not external registration.
Circularity Check
No circularity: the audit's claims rest on direct measurement, formal metric identities, and held-out criteria that are distinct from the audited benchmarks.
full rationale
Walking the derivation chain, every load-bearing result is either an observed quantity or a formal identity, not a fitted input renamed as a prediction. Observation 1 is a closed-form consequence of the definition of F1 and the measured class base rate: an always-positive policy has recall 1, precision π, hence F1 = 2π/(1+π). This is a mathematical fact about the metric, not a parameter fitted to the benchmark's outcomes, and the paper uses it as a critique of F1 rather than as evidence about any model. The cross-benchmark disagreement (Table 1) is directly measured from the official implementations and scorers. The small-panel artifact is demonstrated by resampling the paper's own model panel, not by assuming the conclusion. RQ3 uses three held-out criteria (τ2-bench retail, the Lynch et al. agentic-misalignment scenario, and three-template StrongREJECT jailbreak) that are distinct from the four audited benchmarks and are not constructed from those benchmark scores; the correlations are reported as empirical associations with explicit caveats. The capability composite is independently measured under a fixed protocol, and the paper explicitly labels the strongest AgentHarm–jailbreak link as convergent validity rather than general safety. The limitations the paper itself flags—notably the misalignment criterion's moderate split-half reliability (Spearman–Brown 0.70) and its status as a single fictional scenario—are construct-validity and measurement-reliability concerns, not circular reductions: they do not make the criterion equivalent to the predictor. The only overlapping-author citation (R-Judge) identifies the benchmark under audit; it is not used as an external theorem to justify a claim. No step reduces to its own input by construction, and no fitted parameter is presented as a prediction.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption The official implementations and author-provided scorers of R-Judge, InjecAgent, AgentHarm, and AgentDojo faithfully measure their intended constructs when re-run on the API panel.
- domain assumption The capability composite (mean of standardized MMLU and GPQA-Diamond measured under one protocol) adequately represents general capability for the partial correlations and rank-residualizations.
- domain assumption The three held-out criteria (τ2-retail success, Agentic Misalignment scenario, three-template jailbreak) are valid stand-ins for deployment behavior.
- standard math Asymptotic Spearman tests, percentile bootstrap, Horn's parallel analysis, and permutation tests are statistically valid at the panel sizes used (n≈18–41).
- ad hoc to paper The analysis was genuinely pre-specified: hypotheses, thresholds, and decision rules were fixed before seeing results.
read the original abstract
Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official implementation and author-provided scorer on up to 22 models, with MMLU and GPQA measured by us under one protocol as a capability composite. The metric is the first problem. On any binary trace-judgment benchmark scored by $F_1$, an ``always positive'' policy attains $F_1 = 2\pi/(1+\pi)$; on R-Judge that is $0.690$, above five of the 21 models that actually discriminate. The three broad-coverage benchmarks then rank the same 18 models differently, and the trade-off behind that disagreement is a small-panel artifact: R-Judge specificity against AgentHarm safety correlates $-0.64$ at $n{=}7$ and $+0.02$ at $n{=}18$, and a quarter of random size-7 subsets reach $|\rho| \geq 0.5$ around that near-zero value. Held-out validity turns on which outcome you pick. Capability predicts task success ($\rho{=}{+}0.60$) but correlates negatively with misalignment safety ($\rho{=}{-}0.44$, $n{=}21$). On their paired $n{=}20$ panel, the corresponding contrast is $\Delta{=}{-}1.00$ (95% CI $[-1.48, -0.49]$, $p<0.001$), and it survives leave-one-organization-out and organization-clustered bootstrap analyses. On an expanded 41-model panel, the misalignment correlation weakens to $-0.16$ (95% CI $[-0.54, +0.22]$) and jailbreak strengthens to $+0.34$, though neither change is significant. \mbox{AgentHarm} shows the strongest held-out association, $\rho{=}{+}0.72$ with three-template jailbreak safety after controlling capability. But both instruments score harmful compliance, so this is evidence of convergent validity rather than general safety. Naming the benchmark, metric, target behavior, and model panel is the minimum a safety claim needs.
Figures
Reference graph
Works this paper leans on
-
[1]
Si, Chenglei and Yang, Diyi and Hashimoto, Tatsunori , year =. Can. 2409.04109 , archivePrefix =
-
[2]
Li, Miles Q. and Fung, Benjamin C. M. and Li, Boyang and Ismail, Heba and Iqbal, Farkhund , year =. Taxonomy and Consistency Analysis of Safety Benchmarks for. 2605.16282 , archivePrefix =
-
[3]
Perlitz, Yotam and Gera, Ariel and Arviv, Ofir and Yehudai, Asaf and Bandel, Elron and Shnarch, Eyal and Shmueli-Scheuer, Michal and Choshen, Leshem , year =. Do These. 2407.13696 , archivePrefix =
-
[4]
International Conference on Learning Representations (ICLR) , year =
metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models , author =. International Conference on Learning Representations (ICLR) , year =. 2407.12844 , archivePrefix =
-
[5]
Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track , year =
Measuring what Matters: Construct Validity in Large Language Model Benchmarks , author =. Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track , year =. 2511.04703 , archivePrefix =
-
[6]
2025 , eprint =
Andriushchenko, Maksym and Souly, Alexandra and Dziemian, Mateusz and Duenas, Derek and Lin, Maxwell and Wang, Justin and Zou, Andy and Kolter, Zico and Fredrikson, Matt and Winsor, Eric and Wynne, Jerome and Gal, Yarin and Davies, Xander , booktitle =. 2025 , eprint =
2025
-
[7]
Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track , year =
Debenedetti, Edoardo and Zhang, Jie and Balunovi\'. Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track , year =. 2406.13352 , archivePrefix =
-
[8]
2024 , eprint =
Zhan, Qiusi and Liang, Zhixiang and Ying, Zifan and Kang, Daniel , booktitle =. 2024 , eprint =
2024
-
[9]
2024 , eprint =
Yuan, Tongxin and He, Zhiwei and Dong, Lingzhong and Wang, Yiming and Zhao, Ruijie and Xia, Tian and Xu, Lizhen and Zhou, Binglin and Li, Fangqi and Zhang, Zhuosheng and Wang, Rui and Liu, Gongshen , booktitle =. 2024 , eprint =
2024
-
[10]
and Hashimoto, Tatsunori , booktitle =
Ruan, Yangjun and Dong, Honghua and Wang, Andrew and Pitis, Silviu and Zhou, Yongchao and Ba, Jimmy and Dubois, Yann and Maddison, Chris J. and Hashimoto, Tatsunori , booktitle =. Identifying the Risks of. 2024 , eprint =
2024
-
[11]
2024 , eprint =
-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains , author =. 2024 , eprint =
2024
-
[12]
2025 , eprint =
^2 -Bench: Evaluating Conversational Agents in a Dual-Control Environment , author =. 2025 , eprint =
2025
-
[13]
and Mindermann, S
Lynch, Aengus and Wright, Benjamin and Larson, Caleb and Ritchie, Stuart J. and Mindermann, S. Agentic Misalignment: How. 2025 , eprint =
2025
-
[14]
Souly, Alexandra and Lu, Qingyuan and Bowen, Dillon and Trinh, Tu and Hsieh, Elvis and Pandey, Sana and Abbeel, Pieter and Svegliato, Justin and Emmons, Scott and Watkins, Olivia and Toyer, Sam , year =. A. 2402.10260 , archivePrefix =
-
[15]
Beyond Task Completion: Revealing Corrupt Success in
Cao, Hongliu and Driouich, Ilias and Thomas, Eoin , year =. Beyond Task Completion: Revealing Corrupt Success in. 2603.03116 , archivePrefix =
-
[16]
Evidence-Bound Autonomous Research (
Chen, Ruiying , year =. Evidence-Bound Autonomous Research (. 2511.05524 , archivePrefix =
-
[17]
Proceedings of NAACL-HLT , year =
R\". Proceedings of NAACL-HLT , year =. 2308.01263 , archivePrefix =
-
[18]
Cui, Justin and Chiang, Wei-Lin and Stoica, Ion and Hsieh, Cho-Jui , year =. 2405.20947 , archivePrefix =
-
[19]
Jailbroken: How Does
Wei, Alexander and Haghtalab, Nika and Steinhardt, Jacob , booktitle =. Jailbroken: How Does. 2023 , eprint =
2023
-
[20]
and Mao, Huanzhi and Ji, Charlie Cheng-Jie and Yan, Fanjia and Suresh, Vishnu and Stoica, Ion and Gonzalez, Joseph E
Patil, Shishir G. and Mao, Huanzhi and Ji, Charlie Cheng-Jie and Yan, Fanjia and Suresh, Vishnu and Stoica, Ion and Gonzalez, Joseph E. , booktitle =. The
-
[21]
2026 , eprint =
Quantifying Construct Validity in Large Language Model Evaluations , author =. 2026 , eprint =
2026
-
[22]
Yang, Shu and Hu, Jingyu and Li, Tong and Yan, Hanqi and Wang, Wenxuan and Wang, Di , booktitle =
-
[23]
Expanding the
Keller, Andrew and Kwegyir-Aggrey, Kweku and Steed, Ryan and Rao, Anita and Sharp, Julia and Bergman, Amanda , institution =. Expanding the
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.