Pith. sign in

REVIEW 2 major objections 5 minor 23 references

Scores from four agent-safety benchmarks are not interchangeable: they rank the same models differently, reward an always-unsafe baseline, and shift with the model panel, so safety claims must attach benchmark, metric, target behavior, and

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-03 00:40 UTC pith:LF2DOOI5

load-bearing objection A genuinely useful metric-validity and small-panel audit; treat the misalignment-based conclusions as exploratory. the 2 major comments →

arxiv 2607.28685 v1 pith:LF2DOOI5 submitted 2026-07-30 cs.AI cs.IR

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks

classification cs.AI cs.IR
keywords agent safetybenchmark validityF1 metriccapability confoundcriterion validitysmall panelsR-Judgemodel evaluation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that agent-safety benchmark scores are not interchangeable measurements of a single safety property, and that quoting a score like 'AgentHarm safety 0.77' as an agent's safety is invalid. It re-runs four benchmarks (R-Judge, InjecAgent, AgentHarm, AgentDojo) on up to 22 models, with MMLU and GPQA measured under one protocol as a capability composite, and pre-specified three held-out criteria. The audit finds that an 'always unsafe' policy reaches F1=0.690 on R-Judge, above five real discriminating models; the three broad-coverage benchmarks rank the same 18 models differently; and whether capability predicts safety depends on the outcome and panel (positive for task success, negative for misalignment, near-zero for jailbreak). The paper concludes that the minimum a safety claim needs is the benchmark, metric, target behavior, and model panel.

Core claim

The paper's central claim is that no single agent-safety benchmark stands in for 'safety' as a whole. On any binary trace-judgment benchmark scored by F1, an always-positive policy attains F1=2π/(1+π); on R-Judge that baseline scores 0.690, outranking five of 21 models that actually discriminate. The three broad-coverage benchmarks rank the same 18 models differently, and a trade-off between R-Judge specificity and AgentHarm safety that looked strong at n=7 (ρ=-0.64) dissolves at n=18 (ρ=+0.02) — a quarter of random seven-model subsets manufacture |ρ|≥0.5 around that near-zero value. On the held-out side, capability predicts task success (ρ=+0.60) but correlates negatively with a misalignmen

What carries the argument

The central object is the F1 closed-form identity: on a binary trace-judgment benchmark with unsafe-class base rate π, the constant 'always positive' policy scores F1=2π/(1+π). It is the mechanism that exposes the metric-level failure — F1 gives no credit for correctly identifying benign traces, so a policy that simply labels everything unsafe can match a leaderboard. The audit's other machinery is a capability-control design: a harmonized MMLU+GPQA composite, Spearman and partial-Spearman correlations, and three pre-specified held-out criteria (task success, a fictional misalignment scenario, and a three-template jailbreak measure) that let the paper ask whether safety scores add predictive

Load-bearing premise

The audit's criterion-validity conclusions assume the three held-out outcome tests — especially the single fictional misalignment scenario, whose two halves agree only moderately — actually measure deployment safety; if that scenario is not a valid measure of agentic misalignment, the headline negative capability–safety correlation and the tests built on it would be artifacts.

What would settle it

Run the same four benchmarks and three held-out criteria on at least 40 models from new organizations. The central claim would be falsified if an always-unsafe baseline no longer beats any real discriminating model on R-Judge, if the broad-coverage benchmarks' rankings agree strongly (Spearman ≥ 0.8), or if the AgentHarm–jailbreak partial correlation drops well below 0.72. A cheaper probe: give the same models two independently written misalignment scenarios; if their model-level scores correlate below ~0.7, the misalignment criterion is not a stable construct and the correlations built on it

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • R-Judge's headline F1 should not be reported alone; two-sided metrics such as balanced accuracy are needed because F1 gives no credit for true negatives.
  • Rankings from different agent-safety benchmarks are not comparable; 'safety rank' is benchmark-dependent, so a score must be quoted with its benchmark, metric, target behavior, and model panel.
  • Small-panel correlations are unreliable: a quarter of random seven-model subsets can produce |ρ|≥0.5 around a near-zero full-panel value, so validity must be re-checked as the model population turns over.
  • Capability and safety are not substitutes: capability predicts task success but correlates negatively with misalignment safety on the original panel; high-capability models can be less safe on that criterion.
  • AgentHarm's strong association with jailbreak resistance reflects convergent validity — both measure harmful compliance — and should not be read as evidence of general safety.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The F1 identity generalizes: any binary trace-judgment leaderboard scored by F1 rewards positive-heavy policies, and the gap between a model's F1 and 2π/(1+π) is a conservative measure of how much discrimination the score actually encodes; reporting balanced accuracy or Matthews correlation would neutralize the artifact.
  • Because the capability composite uses static-knowledge tests, the audit cannot rule out that an agentic-capability component is driving the outcome-dependent correlations; a validated tool-use capability anchor would be the natural next control.
  • A testable extension: pre-register a large multi-organization panel (n≥40) of agent loops, re-run the four benchmarks and three criteria, and check whether the near-zero R-Judge–AgentHarm correlation and the AgentHarm–jailbreak partial ρ≈0.72 replicate; the paper's own pre-specified 'hold or strengthen' prediction for the misalignment correlation was not met, so an independent replication would se

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper conducts a validity audit of four agent-safety benchmarks (R-Judge, InjecAgent, AgentHarm, AgentDojo) by re-running them under official implementations and author-provided scorers on up to 22 models, with MMLU/GPQA measured under one protocol as a capability composite. It reports four main findings: (1) a metric-validity failure — on binary trace-judgment benchmarks scored by F1, an always-positive policy attains F1 = 2π/(1+π), which on R-Judge (0.690) outranks five models that actually discriminate (Observation 1, §4.1); (2) the three broad-coverage benchmarks rank the same 18 models differently, and a small-panel trade-off between R-Judge specificity and AgentHarm safety dissolves at larger n (§4.3); (3) capability predicts held-out task success (ρ=+0.60) but correlates negatively with the paper's misalignment safety criterion (ρ=−0.44, n=21), producing a crossover interaction Δ=−1.00; and (4) AgentHarm's capability-controlled association with jailbreak safety (ρ=+0.72) is the strongest held-out result but is interpreted as convergent validity of harmful-compliance measurement rather than general safety. The paper concludes that no single benchmark score can be quoted as 'an agent's safety' without specifying benchmark, metric, target behavior, and model panel.

Significance. If the results stand, the paper makes a valuable, field-level contribution: it demonstrates formally and empirically that agent-safety benchmark scores are not interchangeable measurements, and it identifies a specific metric failure (F1's blindness to true negatives) that is easy to overlook. The F1 closed form is a clean, general result. The paper is unusually careful: it uses official implementations, runs capability anchors under one protocol, reports metric sensitivity, discloses power limitations, and explicitly distinguishes pre-specified analyses from exploratory ones. The reproducibility package and API-only harness are concrete assets. The central negative claim about benchmark disagreement and the gameability of F1 is robust to the RQ3 concerns discussed below because it does not depend on any held-out criterion. The paper's treatment of the AgentHarm–jailbreak association as convergent validity rather than general safety is appropriately cautious.

major comments (2)
  1. [§4.4, Table 3, App. A.5] The negative capability–misalignment correlation (ρ=−0.44) and the Δ=−1.00 interaction rest entirely on one fictional blackmail/leaking scenario whose split-half reliability is only 0.70 (Spearman–Brown, Table 3) and whose construct classification is contested in the blinded coding (6/7 harm-blocking in one round vs. 6/7 risk recognition in another, App. A.1). If the probe measures compliance with a narrow blackmail script rather than general agentic misalignment, the headline 'capability predicts lower safety' result is an artifact of that scenario. The post-hoc three-scenario battery does not repair this: it was added after seeing the primary result, changes the effect size negligibly, and includes murder, which co-varies with AgentHarm and dilutes construct specificity. Please either report a pre-specified multi-scenario misalignment battery as the primary criterion, or explicitly del
  2. [§4.4, Tables 2 and 4] The R-Judge/InjecAgent–misalignment partials (+0.41/+0.47) are labeled exploratory in the text, but the abstract's claim that capability 'correlates negatively with misalignment safety' gives them more weight than they can bear. These cells have power ≈0.35, their confidence intervals include zero, and the R-Judge estimate is metric-fragile: +0.41 under specificity becomes +0.19 under balanced accuracy and −0.46 under recall (Table 4). Since one of the paper's central lessons is that metric choice changes conclusions, the summary should carry the same caveats as the body — or this result should be dropped from the abstract.
minor comments (5)
  1. [Fig. 1] The x-axis label 'models, ranked by F1' is ambiguous; consider 'model rank (1 = best F1)'.
  2. [Table 3] The footnote says the jailbreak criterion uses 42 models, while the body reports n=41 on the expanded panel; clarify the discrepancy (e.g., DeepSeek-R1 inclusion).
  3. [Table 4] The jailbreak column has n=20 while Table 2 reports n=18 for the same cell; add a coverage note so readers do not misread the panels.
  4. [App. A.5] The arrow notation 'AgentHarm→jailbreak' is informal; define it when first used.
  5. [§3.3] The term 'pre-specified' is used repeatedly. Since the manuscript correctly notes that these analyses lack an independent timestamp, consider adding a sentence stating that the frozen analysis plan is included in the artifact and that 'pre-specified' reflects the authors' internal workflow, not external registration.

Circularity Check

0 steps flagged

No circularity: the audit's claims rest on direct measurement, formal metric identities, and held-out criteria that are distinct from the audited benchmarks.

full rationale

Walking the derivation chain, every load-bearing result is either an observed quantity or a formal identity, not a fitted input renamed as a prediction. Observation 1 is a closed-form consequence of the definition of F1 and the measured class base rate: an always-positive policy has recall 1, precision π, hence F1 = 2π/(1+π). This is a mathematical fact about the metric, not a parameter fitted to the benchmark's outcomes, and the paper uses it as a critique of F1 rather than as evidence about any model. The cross-benchmark disagreement (Table 1) is directly measured from the official implementations and scorers. The small-panel artifact is demonstrated by resampling the paper's own model panel, not by assuming the conclusion. RQ3 uses three held-out criteria (τ2-bench retail, the Lynch et al. agentic-misalignment scenario, and three-template StrongREJECT jailbreak) that are distinct from the four audited benchmarks and are not constructed from those benchmark scores; the correlations are reported as empirical associations with explicit caveats. The capability composite is independently measured under a fixed protocol, and the paper explicitly labels the strongest AgentHarm–jailbreak link as convergent validity rather than general safety. The limitations the paper itself flags—notably the misalignment criterion's moderate split-half reliability (Spearman–Brown 0.70) and its status as a single fictional scenario—are construct-validity and measurement-reliability concerns, not circular reductions: they do not make the criterion equivalent to the predictor. The only overlapping-author citation (R-Judge) identifies the benchmark under audit; it is not used as an external theorem to justify a claim. No step reduces to its own input by construction, and no fitted parameter is presented as a prediction.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

The paper's central claims rest on domain assumptions about the validity of the benchmarks, the capability anchor, and held-out criteria, plus a standard statistical toolkit and an unverifiable claim of pre-specification. No free parameters are fitted to produce the results, and no new entities are introduced.

axioms (5)
  • domain assumption The official implementations and author-provided scorers of R-Judge, InjecAgent, AgentHarm, and AgentDojo faithfully measure their intended constructs when re-run on the API panel.
    The entire audit reuses these implementations (§3.1); if they are misconfigured or mis-scored, all benchmark-score columns are invalid. Independent grader checks cover some judges but not the construct validity.
  • domain assumption The capability composite (mean of standardized MMLU and GPQA-Diamond measured under one protocol) adequately represents general capability for the partial correlations and rank-residualizations.
    Used in RQ2/RQ3 and the organization-group analysis; the paper notes it is static knowledge and may miss agentic capability (Section 6, App. A.6), and the BFCL pilot failed as an agentic replacement.
  • domain assumption The three held-out criteria (τ2-retail success, Agentic Misalignment scenario, three-template jailbreak) are valid stand-ins for deployment behavior.
    The paper explicitly scopes them as stand-ins (Section 1, Section 6). The misalignment criterion's two scenario halves agree only moderately (Spearman–Brown 0.70), making it the most fragile of the three.
  • standard math Asymptotic Spearman tests, percentile bootstrap, Horn's parallel analysis, and permutation tests are statistically valid at the panel sizes used (n≈18–41).
    Used throughout; the authors run a calibration check for the organization-level permutation test (App. A.6) but not for every inference.
  • ad hoc to paper The analysis was genuinely pre-specified: hypotheses, thresholds, and decision rules were fixed before seeing results.
    The paper states 'Because these choices lack an independent timestamp, we describe the analyses as pre-specified rather than preregistered' (Section 3.3). If false, multiple-comparison risk increases and some reported p-values overstate evidence.

pith-pipeline@v1.3.0-alltime-deepseek · 18730 in / 14174 out tokens · 122674 ms · 2026-08-03T00:40:40.157378+00:00 · methodology

0 comments
read the original abstract

Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official implementation and author-provided scorer on up to 22 models, with MMLU and GPQA measured by us under one protocol as a capability composite. The metric is the first problem. On any binary trace-judgment benchmark scored by $F_1$, an ``always positive'' policy attains $F_1 = 2\pi/(1+\pi)$; on R-Judge that is $0.690$, above five of the 21 models that actually discriminate. The three broad-coverage benchmarks then rank the same 18 models differently, and the trade-off behind that disagreement is a small-panel artifact: R-Judge specificity against AgentHarm safety correlates $-0.64$ at $n{=}7$ and $+0.02$ at $n{=}18$, and a quarter of random size-7 subsets reach $|\rho| \geq 0.5$ around that near-zero value. Held-out validity turns on which outcome you pick. Capability predicts task success ($\rho{=}{+}0.60$) but correlates negatively with misalignment safety ($\rho{=}{-}0.44$, $n{=}21$). On their paired $n{=}20$ panel, the corresponding contrast is $\Delta{=}{-}1.00$ (95% CI $[-1.48, -0.49]$, $p<0.001$), and it survives leave-one-organization-out and organization-clustered bootstrap analyses. On an expanded 41-model panel, the misalignment correlation weakens to $-0.16$ (95% CI $[-0.54, +0.22]$) and jailbreak strengthens to $+0.34$, though neither change is significant. \mbox{AgentHarm} shows the strongest held-out association, $\rho{=}{+}0.72$ with three-template jailbreak safety after controlling capability. But both instruments score harmful compliance, so this is evidence of convergent validity rather than general safety. Naming the benchmark, metric, target behavior, and model panel is the minimum a safety claim needs.

Figures

Figures reproduced from arXiv: 2607.28685 by Bowen Liu, Dingyan Shang, Xiao Han, Youting Wang, Yuan Tang.

Figure 1
Figure 1. Figure 1: A leaderboard’s headline metric can be matched by an “always unsafe” baseline. On R-Judge (n=21), five real, discriminating models score below this constant base￾line (F1=0.69, dashed): F1 gives no credit for correctly re￾jected benign traces, so the degenerate policy can outrank models that discriminate [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Capability predicts task success; safety point estimates differ by outcome and panel. Filled circles show the 2024–25 correlations; open diamonds show the expanded-panel safety correlations; error bars are model￾level percentile-bootstrap 95% CIs. Thin lines connect panel-level estimates, not model trajectories. Misalignment changes from −0.44 to −0.16 and jailbreak from +0.08 to +0.34; neither between-pan… view at source ↗
Figure 3
Figure 3. Figure 3: The audit’s evidence base: up to 46 models × 9 instruments (two capability anchors, four audited benchmarks, three held-out criteria; 262 model×instrument cells). Cells are per-instrument min–max-normalised scores; grey is not-run. The rows beyond the 41-model anchored panel are models excluded from the analyses by pre-specified telemetry gates (App. A.6). Rows are sorted by capability: the 2026 expansion … view at source ↗
Figure 4
Figure 4. Figure 4: Models with similar measured capability can differ widely in safety across model-developing organizations. Each point is a model; colour and shape identify its organization; error bars are within-model SEs (§A.2). Within an organization (lines, the six organizations with ≥ 4 anchored models) the capability→safety slope is large but sign-heterogeneous across organizations (n-weighted mean −0.14 misalignment… view at source ↗
Figure 5
Figure 5. Figure 5: R-Judge and AgentHarm do not rank the full panel in opposite order. AgentHarm safety against R-Judge specificity on the n=18 cross-benchmark panel, coloured by the harmonized GPQA anchor. The trade-off a 7-model panel suggested (ρ = −0.64) dissolves at n=18 (ρ = +0.02, p = 0.95; grey trend): models near the top of either axis span the full range of the other, the pairwise view of [PITH_FULL_IMAGE:figures/… view at source ↗
Figure 7
Figure 7. Figure 7: Against task success, capability is the predictor and AgentHarm adds nothing. τ 2 -retail success against the harmonized capability composite (ρ = +0.60, p = 0.005, n=20; grey trend). Colour is AgentHarm safety (open circle: no AgentHarm score): safe and unsafe models sit on both sides of the trend with no safety gradient after controlling capability, the pre-specified result that AgentHarm adds no increme… view at source ↗
Figure 6
Figure 6. Figure 6: The first principal component groups R-Judge, InjecAgent, and capability, with AgentHarm in the op￾posite direction. PC1 loadings on the n=18 panel with har￾monized anchors (48.7% of variance). AgentHarm has the opposite sign from the other four measures in 17 of 18 leave￾one-model-out fits. The sign of a principal component is ar￾bitrary; only the relative directions matter. We treat this struc￾ture as su… view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

23 extracted references · 8 linked inside Pith

  1. [1]

    Si, Chenglei and Yang, Diyi and Hashimoto, Tatsunori , year =. Can. 2409.04109 , archivePrefix =

  2. [2]

    and Fung, Benjamin C

    Li, Miles Q. and Fung, Benjamin C. M. and Li, Boyang and Ismail, Heba and Iqbal, Farkhund , year =. Taxonomy and Consistency Analysis of Safety Benchmarks for. 2605.16282 , archivePrefix =

  3. [3]

    Do These

    Perlitz, Yotam and Gera, Ariel and Arviv, Ofir and Yehudai, Asaf and Bandel, Elron and Shnarch, Eyal and Shmueli-Scheuer, Michal and Choshen, Leshem , year =. Do These. 2407.13696 , archivePrefix =

  4. [4]

    International Conference on Learning Representations (ICLR) , year =

    metabench -- A Sparse Benchmark of Reasoning and Knowledge in Large Language Models , author =. International Conference on Learning Representations (ICLR) , year =. 2407.12844 , archivePrefix =

  5. [5]

    Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track , year =

    Measuring what Matters: Construct Validity in Large Language Model Benchmarks , author =. Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track , year =. 2511.04703 , archivePrefix =

  6. [6]

    2025 , eprint =

    Andriushchenko, Maksym and Souly, Alexandra and Dziemian, Mateusz and Duenas, Derek and Lin, Maxwell and Wang, Justin and Zou, Andy and Kolter, Zico and Fredrikson, Matt and Winsor, Eric and Wynne, Jerome and Gal, Yarin and Davies, Xander , booktitle =. 2025 , eprint =

  7. [7]

    Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track , year =

    Debenedetti, Edoardo and Zhang, Jie and Balunovi\'. Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track , year =. 2406.13352 , archivePrefix =

  8. [8]

    2024 , eprint =

    Zhan, Qiusi and Liang, Zhixiang and Ying, Zifan and Kang, Daniel , booktitle =. 2024 , eprint =

  9. [9]

    2024 , eprint =

    Yuan, Tongxin and He, Zhiwei and Dong, Lingzhong and Wang, Yiming and Zhao, Ruijie and Xia, Tian and Xu, Lizhen and Zhou, Binglin and Li, Fangqi and Zhang, Zhuosheng and Wang, Rui and Liu, Gongshen , booktitle =. 2024 , eprint =

  10. [10]

    and Hashimoto, Tatsunori , booktitle =

    Ruan, Yangjun and Dong, Honghua and Wang, Andrew and Pitis, Silviu and Zhou, Yongchao and Ba, Jimmy and Dubois, Yann and Maddison, Chris J. and Hashimoto, Tatsunori , booktitle =. Identifying the Risks of. 2024 , eprint =

  11. [11]

    2024 , eprint =

    -bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains , author =. 2024 , eprint =

  12. [12]

    2025 , eprint =

    ^2 -Bench: Evaluating Conversational Agents in a Dual-Control Environment , author =. 2025 , eprint =

  13. [13]

    and Mindermann, S

    Lynch, Aengus and Wright, Benjamin and Larson, Caleb and Ritchie, Stuart J. and Mindermann, S. Agentic Misalignment: How. 2025 , eprint =

  14. [14]

    Souly, Alexandra and Lu, Qingyuan and Bowen, Dillon and Trinh, Tu and Hsieh, Elvis and Pandey, Sana and Abbeel, Pieter and Svegliato, Justin and Emmons, Scott and Watkins, Olivia and Toyer, Sam , year =. A. 2402.10260 , archivePrefix =

  15. [15]

    Beyond Task Completion: Revealing Corrupt Success in

    Cao, Hongliu and Driouich, Ilias and Thomas, Eoin , year =. Beyond Task Completion: Revealing Corrupt Success in. 2603.03116 , archivePrefix =

  16. [16]

    Evidence-Bound Autonomous Research (

    Chen, Ruiying , year =. Evidence-Bound Autonomous Research (. 2511.05524 , archivePrefix =

  17. [17]

    Proceedings of NAACL-HLT , year =

    R\". Proceedings of NAACL-HLT , year =. 2308.01263 , archivePrefix =

  18. [18]

    2405.20947 , archivePrefix =

    Cui, Justin and Chiang, Wei-Lin and Stoica, Ion and Hsieh, Cho-Jui , year =. 2405.20947 , archivePrefix =

  19. [19]

    Jailbroken: How Does

    Wei, Alexander and Haghtalab, Nika and Steinhardt, Jacob , booktitle =. Jailbroken: How Does. 2023 , eprint =

  20. [20]

    and Mao, Huanzhi and Ji, Charlie Cheng-Jie and Yan, Fanjia and Suresh, Vishnu and Stoica, Ion and Gonzalez, Joseph E

    Patil, Shishir G. and Mao, Huanzhi and Ji, Charlie Cheng-Jie and Yan, Fanjia and Suresh, Vishnu and Stoica, Ion and Gonzalez, Joseph E. , booktitle =. The

  21. [21]

    2026 , eprint =

    Quantifying Construct Validity in Large Language Model Evaluations , author =. 2026 , eprint =

  22. [22]

    Yang, Shu and Hu, Jingyu and Li, Tong and Yan, Hanqi and Wang, Wenxuan and Wang, Di , booktitle =

  23. [23]

    Expanding the

    Keller, Andrew and Kwegyir-Aggrey, Kweku and Steed, Ryan and Rao, Anita and Sharp, Julia and Bergman, Amanda , institution =. Expanding the