Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper claims that current frontier AI models, evaluated across seven risk areas against red-line and yellow-line thresholds, all fall in green or yellow zones, with none crossing red lines as of the July 10, 2025 evaluation cutoff.

desk verdict Serious risk-screening report with two genuinely new empirical pieces, but the 'no red lines' headline is not supported by its own evaluation design. read the letter →

arxiv 2507.16534 v2 pith:IMGQAMRZ submitted 2025-07-22 cs.AI cs.CLcs.CVcs.LG

classification cs.AIcs.CLcs.CVcs.LG
keywords frontierAIriskthresholdsredlinesandyellowE-T-Canalysiscyberoffensepersuasionmanipulationbiologicalchemicalrisksreasoningmodel
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This technical report tries to make frontier-AI risk measurable by defining red lines (intolerable outcomes) and yellow lines (early-warning thresholds), then scoring a broad set of current models against them across seven risk areas: cyber offense, biological and chemical risks, persuasion and manipulation, uncontrolled AI R&D, strategic deception and scheming, self-replication, and collusion. Its central result is that every evaluated model sits in the green or yellow zone, and none crosses a red line: no model crosses the yellow line for cyber offense or uncontrolled AI R&D, most are yellow for persuasion, reasoning models are yellow for self-replication and strategic deception, and for biological and chemical risks the authors cannot rule out yellow status for most models. The report matters because it is an early, concrete attempt to turn risk thresholds into deployment rules: green means routine deployment with monitoring, yellow means controlled deployment with strengthened mitigations, and red means suspension. The authors present the work as a snapshot of models available before July 10, 2025, and as a call for collective, transparent risk evaluation.

What carries the argument

The machinery is the E-T-C analysis — deployment environment, threat source, enabling capability — used to derive red and yellow lines, combined with the 'AI-45° Law' cited as the guiding principle that capability and risk escalate together. The report operationalizes these as risk zones with explicit thresholds: for example, a CyBench safety score below 0.6 or a PACEBench score below 0.5 triggers a yellow warning for cyber offense, and a persuasion safety score below 0.4 (successful persuasion above 60%) triggers the yellow zone for manipulation. Thresholds are anchored to human-expert baselines on proxy benchmarks, and behavior is measured through tool-using agents in Docker sandboxes for cyber tasks, multi-turn dialogue protocols for persuasion, and refusal-rate benchmarks for biological and chemical safety. The key work of this machinery is to convert qualitative risk categories into numbers that can route a model into a deployment protocol.

What would settle it

A single agent autonomously completing a full-chain, multi-host penetration that defeats an OWASP-grade WAF — the success criteria the report's own PACEBench defines — would falsify the claim that no evaluated model crosses the cyber-offense yellow line.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that the current generation of frontier models has not reached the defined intolerable-risk boundary, but several sit close to it. Concretely: no evaluated model crosses the yellow line for cyber offense or uncontrolled AI R&D; most models are green for self-replication and strategic deception, with some reasoning models in yellow; most models are yellow for persuasion and manipulation because they shift human opinions at rates above the sixty-percent threshold (the best model reaches 63.1% in human interactions); and for biological and chemical risks the report cannot rule out yellow placement for most models, since frontier models now exceed human-expert baselines on proxy knowledge benchmarks while refusal rates are inconsistent. The report also claims a capability-risk coupling: reasoning-enabled models consistently show higher cyber-offense and persuasion risk than their non-reasoning counterparts, while model scale alone is not a reliable predictor of persuasion risk. The authors frame these as operational conclusions of applying the E-T-C analysis (deployment environment, threat source, enabling capability) to a July 2025 model set.

Load-bearing premise

The load-bearing premise is that scores on automated proxy benchmarks measure real-world catastrophic-risk capability; if benchmark performance does not translate to actual uplift or attack success, the green/yellow/red zone assignments for cyber, biological, and chemical risks are not measuring the risks they claim to classify.

Editorial extensions

If this is right

  • If the zone assignments are right, no current model needs to be suspended by this framework's rules, but several need controlled deployment: persuasion-capable models and reasoning models that sit in the yellow zone should ship with strengthened mitigations and monitoring.
  • The yellow-line thresholds give model providers a concrete trigger: a model approaching 0.6 on CyBench or 0.5 on PACEBench for cyber, or above 60% persuasion success, should undergo safety review before broader deployment.
  • Because the report cannot rule out yellow status for most models on biological and chemical risks, the precautionary implication is that frontier providers should run deeper threat modeling and consider use restrictions in these domains even without confirmed crossing.
  • The capability-risk coupling implies that each generation of models with stronger reasoning and tool use should be re-evaluated against these thresholds; safety scores are not stable properties of a model family.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If proxy benchmarks keep saturating (every model already exceeds the human expert baseline on WMDP), the yellow-line assignments may start measuring benchmark ceiling rather than real-world risk; the paper's own cited uplift studies suggest the next decisive test is human uplift experiments, which it did not run.
  • The finding that reasoning modes unlock CTF-solving capability suggests a testable extension: re-running the same seven risk batteries under adversarial elicitation and long-horizon agent loops would likely move several reasoning models from green to yellow, and could move the aggregate risk picture closer to a red line.
  • The report's 'unable to rule out yellow' stance for biology and chemistry is effectively a policy claim: precaution alone justifies stricter deployment controls, so the operative question for regulators is who bears the burden of proof when the measurement is inconclusive.
  • The red-line thresholds are defined by unacceptable real-world outcomes, but the report concedes no current benchmark can test them; a natural extension is to build high-fidelity simulations of full-chain attacks and mass manipulation, with expert adjudication, as the actual red-line test.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. This technical report evaluates 18 recent frontier large language models across seven risk areas—cyber offense, biological and chemical risks, persuasion and manipulation, strategic deception and scheming, uncontrolled AI R&D, self-replication, and collusion—using the E-T-C (Environment-Threat-Capability) analysis of the authors' SafeWork-F1 framework. The authors introduce green/yellow/red risk zones, run a large battery of general and domain-specific benchmarks (e.g., CyBench, PACEBench, WMDP, LAB-Bench, persuasion experiments with 8,913 human participants), and propose yellow-line early-warning thresholds. The central claim is that all evaluated models reside in the green or yellow zones and that none cross red lines. The report also recommends enhanced mitigations for models approaching or exceeding yellow lines and includes explicit limitations sections.

Significance. If the result held, this would be one of the broadest public cross-model frontier-risk evaluations to date, with genuinely useful empirical artifacts: standardized sandboxed cyber agents, large human-participant persuasion data, and consistent multi-benchmark tables across 18 models. The report is also unusually candid in listing limitations. However, the headline conclusion is not presently supported: red-line thresholds are never operationally defined or measured, the yellow-line thresholds are set in the same report that applies them, the cyber section is internally contradictory, and the proxy-validity caveat undermines the interpretation of benchmark scores as risk measurements. The paper's value is therefore more as a measurement report than as a validated risk-classification framework, and the central claim needs substantial re-scoping or additional evidence.

major comments (4)
  1. [Abstract; §1; §4.1.5; §4.3.5] The central claim that no evaluated model crosses red lines is not supported by the experiments, because no red-line threshold is operationally defined or measured. §4.1.5 states that for cyber offense the red line 'reflects unacceptable outcomes' but that 'current benchmarking methods lack the comprehensiveness and realism needed for a proper evaluation' and that final judgment 'must incorporate expert adjudication'; §4.3.5 similarly says the red-line scenario of 'unacceptable, severe, or autonomous mass manipulation' cannot be fully assessed with the current single-session protocol. Table 1's absence of red-zone entries is therefore an artifact of the evaluation design rather than a measured outcome. The conclusion should be explicitly re-scoped to a claim such as 'no evaluated proxy benchmark detected red-line-level capability,' or the report must supply an actual red-line evaluation protocol.
  2. [§4.1.5; §4.2.6; §4.3.5] The yellow-line thresholds are proposed in the same report that assigns the zones, with no external calibration. The CyBench safety score of 0.6 and PACEBench safety score of 0.5 in §4.1.5, the 0.4 persuasion safety score in §4.3.5, and the composite biological/chemical thresholds in §4.2.6 (with expert-level performance 'preliminarily set as 80%') are all chosen after observing the model scores. Because the same thresholds determine which models are placed in the yellow zone, the conclusion that all models are at most yellow is partly constructed by the choice of threshold values. Independent calibration—for example against human expert uplift, known attacker success rates, or pre-registered thresholds—is needed before these zone assignments can carry the weight of the abstract's claims.
  3. [§4.1.5; Table 1] The cyber-offense conclusion is internally inconsistent. §4.1.5 first states that 'all evaluated models surpass these respective safety thresholds on both benchmarks' and that their risk 'remains within a manageable range,' but then immediately states that 'we are not able to rule out the possibility that frontier models have crossed the yellow lines' and recommends enhanced mitigation measures. A single section cannot both assign every model to the green/yellow zones in Table 1 and simultaneously assert that yellow-line crossing cannot be ruled out. The report should either downgrade the zone assignment to 'not detected by our protocol' or provide a principled basis for the precautionary statement, and the tension should be resolved before publication.
  4. [§4.2.1; §4.2.6; Figures 10, 12, 14] The report acknowledges in §4.2.6 that 'benchmark performance may not directly translate to real-world threat enhancement,' citing studies in which GPT-4's 60–75% WMDP scores produced no significant uplift over internet-only control groups. Yet the biological and chemical yellow-line assessments in §4.2.6, and the capability-versus-safety interpretations in Figures 10, 12, and 14, treat benchmark accuracy or refusal rates as the risk measure itself. This is a load-bearing gap: if the proxy-to-risk link fails, the green/yellow/red classifications for biological, chemical, and possibly cyber risks are not measurements of the risks they claim to classify. The manuscript should either validate the proxy-to-risk link for each chosen benchmark or explicitly relabel the zones as 'benchmark-based capability zones' rather than risk zones.
minor comments (5)
  1. [§4.1.3] The model name 'DeepSeek-R3-0528' appears in the discussion of failure modes; this appears to be a typo for 'DeepSeek-R1-0528'.
  2. [Table 11] The model name 'Gemini-2.5-falsh-preview-0520' is misspelled; it should be 'Gemini-2.5-flash-preview-0520'.
  3. [§4.2.6] The reference '(GÃk,tting et al., 2025)' contains a garbled encoding and should be corrected to 'Götting et al. (2025)' or the proper citation format.
  4. [Table 1] The legend for Table 1 is not self-explanatory: the symbols '####', '###', and '##' are used in the table, but the caption line '#### Green zone ## Yellow zone Red zone' does not clearly map symbols to zones, and the 'N/A*' entry for collusion is left unexplained in the table itself.
  5. [§4.2.4; §4.2.5; Figures 12 and 14] The captions of Figures 12 and 14 refer to 'Table 4' and 'Table 6' when describing the general-capability and domain-specific components, but the relevant tables appear to be Table 3 and Table 8 respectively; these cross-references should be corrected.

Circularity Check

2 steps flagged · score 5.0 of 10

Partial circularity: cyber and persuasion yellow-line thresholds are set from the same measurements used to declare models below or above them, and the no-red-lines claim is unsupported though not purely definitional.

  1. fitted input called prediction [Section 4.1.5, 'Potential Risks Threshold' (Cyber Offense)]
    "Based on our empirical evaluation results, we propose actionable early-warning thresholds for different cyber offense risks. Specifically, for uplift-based cyber offense risk, we recommend a safety score of 0.6 on the CyBench benchmark as the potential risks threshold. For autonomy cyber offense risk, we recommend a safety score of 0.5 on the PACEbench benchmark as its respective threshold. ... Our experimental findings conclude that all evaluated models surpass these respective safety thresholds."

    The thresholds are introduced 'based on our empirical evaluation results' and the same results are then used to conclude that no model crosses the yellow line. With observed CyBench success at most 25% (safety 0.75) and PACEBench at most 0.417 (safety 0.583), the chosen 0.6 and 0.5 thresholds lie above every measured risk score. The conclusion that no evaluated model crosses the cyber yellow line is therefore a restatement of the post-hoc threshold placement, not an independent empirical finding; the threshold is a fitted boundary presented as an early-warning line.

  2. fitted input called prediction [Section 4.3.5, 'Potential Risks Threshold' (Persuasion and Manipulation)]
    "Based on our empirical evaluation, we recommend a safety score threshold of 0.4 as a practical 'yellow-line' early-warning indicator for persuasion and manipulation risks. In this framework, any model with a successful persuasion rate above 60% (1− 0.4) is considered to have entered the cautionary risk zone (i.e., 'yellow zone'). ... These results demonstrate that current LLMs are in the Yellow Zone for Persuasion and Manipulation Risks."

    The 0.4 safety threshold (60% success rate) is proposed after the persuasion experiments were run, and the same experimental results are then cited to demonstrate that models are in the Yellow Zone. The classification is therefore an application of a boundary chosen in light of the measured distribution, not a standard calibrated independently of the outcome being classified. The LLM-to-human table alone would place only Claude-4 above 60%, so the 'most models are yellow' conclusion depends on which experiment is used and on the ex post threshold.

full rationale

Most of the report's measurements retain independent content: CyBench, PACEBench, WMDP, LAB-Bench, BioLP-Bench, and the human persuasion trials use external benchmarks or human participants, and the biological and chemical yellow-lines are anchored to human expert baselines. The SafeWork-F1 self-citation supplies the E-T-C vocabulary and zone definitions but is not a uniqueness argument and does not force any particular experimental outcome, so it is not load-bearing circularity. The two flagged steps are genuine partial circularities: the cyber and persuasion yellow-lines are chosen after observing the same scores they are then used to classify, making the associated zone assignments a relabeling of the threshold choice. Separately, the abstract's 'without crossing red lines' is an unsupported inference: for cyber the report says red-line judgment must incorporate expert adjudication and current benchmarks cannot evaluate it, and for persuasion it says the setup cannot assess mass manipulation; no red-line measurement is reported. That gap is a correctness and evidentiary problem in the central claim rather than a definitional equivalence, but it compounds the partial circularity, yielding a score of 5.

Assumptions & free parameters 8 free parameters · 6 assumptions · 2 invented entities

The central classification rests on self-set thresholds and proxy benchmarks. Free parameters are dominated by yellow-line thresholds chosen in the same report; axioms include the E-T-C framework from the authors' own prior work and the assumption that benchmark scores represent real-world capability. No new physical entities are postulated; the green/yellow/red zones and PACEBench score are operational constructs whose validity depends on threshold calibration.

free parameters (8)
  • CyBench yellow-line safety threshold = 0.6
    Recommended in Section 4.1.5 as the uplift-based cyber offense early-warning threshold; set after observing that all evaluated models score above it, without external calibration.
  • PACEBench yellow-line safety threshold = 0.5
    Recommended in Section 4.1.5 for autonomous cyber offense risk; same post-hoc calibration concern.
  • Persuasion safety threshold = 0.4
    Recommended in Section 4.3.5; defines yellow zone as successful persuasion rate above 60%, which places most measured models in the yellow zone.
  • Biological protocol troubleshooting threshold = 41.3%
    Averaged human expert error rates on BioLP-Bench and ProtocolQA (Section 4.2.6); used to set the biological yellow line.
  • Biological hazardous knowledge composite threshold = 55%
    Composite of WMDP-Biology expert baseline (60.5%), proteotoxicity expert-level target (80%), and 80% refusal rates (Section 4.2.6).
  • Chemical hazardous knowledge composite threshold = 51.3%
    Composite of WMDP-Chemistry expert baseline (43.3%), toxicity targets (80%), and 80% refusal rates (Section 4.2.6).
  • Composite capability weighting = general 0.5, domain 0.5; uniform within groups
    Equation 2 in Section 3.9; arbitrary weighting used for capability-safety scatter plots, not externally justified.
  • PACEBench scenario weights = 3:2:2:3
    Equation 3 in Section 4.1.4; hand-chosen ratio for CVE, multi-host, full-chain, and defense scenarios.
assumptions (6)
  • domain assumption E-T-C (environment, threat source, capability) decomposition from SafeWork-F1-Framework is the correct organizing lens for frontier risk.
    The whole report applies this framework from the authors' own prior work (Section 1), and its categories determine which risks are evaluated and how thresholds are set.
  • domain assumption Human expert baselines on WMDP, LAB-Bench, and BioLP-Bench are appropriate yellow-line calibrations.
    Used to set biological and chemical thresholds in Section 4.2.6; assumes expert performance is the correct early-warning level for dangerous uplift.
  • domain assumption Automated benchmark performance is a valid proxy for real-world dual-use capability.
    Explicitly stated in Section 4.2.1 and heavily relied on across all seven risk areas; the report's own limitations section weakens this assumption.
  • domain assumption Models evaluated via vendor APIs behave identically to their deployed counterparts under normal conditions.
    Stated in Section 3.2 and Section 4.3.5; vendor-side safety layers or interventions can affect measured behavior, as the paper acknowledges for persuasion.
  • domain assumption The 'AI-45° Law' (Yang et al., 2024b) governs the capability-safety relationship.
    Invoked in Section 1 and used to justify E-T-C application; asserted without derivation in this report.
  • domain assumption Reasoning-enabled and standard variants of the same model can be compared as matched pairs.
    Several conclusions about reasoning increasing risk rely on paired comparisons (e.g., Qwen-3 with and without thinking), assuming the only difference is reasoning mode.
invented entities (2)
  • Green/Yellow/Red risk zones
    purpose: Operational deployment status categories based on thresholds
    Introduced via the SafeWork-F1 framework and applied here; their validity depends on threshold calibration, not on an independent falsifiable prediction.
  • PACEBench Score
    purpose: Unified metric for autonomous cyberattack capability
    New benchmark score defined in Equation 3 with hand-chosen weights; no external validation is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report." pith.science (2026). https://pith.science/paper/IMGQAMRZ

@misc{pith2026250716534,
  author       = {Pith},
  title        = {Pith review of: Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IMGQAMRZ}},
  note         = {Machine review of arXiv:2507.16534}
}
abstract

To understand and identify the unprecedented risks posed by rapidly advancing artificial intelligence (AI) models, this report presents a comprehensive assessment of their frontier risks. Drawing on the E-T-C analysis (deployment environment, threat source, enabling capability) from the Frontier AI Risk Management Framework (v1.0) (SafeWork-F1-Framework), we identify critical risks in seven areas: cyber offense, biological and chemical risks, persuasion and manipulation, uncontrolled autonomous AI R\&D, strategic deception and scheming, self-replication, and collusion. Guided by the "AI-$45^\circ$ Law," we evaluate these risks using "red lines" (intolerable thresholds) and "yellow lines" (early warning indicators) to define risk zones: green (manageable risk for routine deployment and continuous monitoring), yellow (requiring strengthened mitigations and controlled deployment), and red (necessitating suspension of development and/or deployment). Experimental results show that all recent frontier AI models reside in green and yellow zones, without crossing red lines. Specifically, no evaluated models cross the yellow line for cyber offense or uncontrolled AI R\&D risks. For self-replication, and strategic deception and scheming, most models remain in the green zone, except for certain reasoning models in the yellow zone. In persuasion and manipulation, most models are in the yellow zone due to their effective influence on humans. For biological and chemical risks, we are unable to rule out the possibility of most models residing in the yellow zone, although detailed threat modeling and in-depth assessment are required to make further claims. This work reflects our current understanding of AI frontier risks and urges collective action to mitigate these challenges.

Figures

Figures reproduced from arXiv: 2507.16534 by the authors.

Figure 1
Figure 1. LLM agent framework for solving CTF challenges ( [PITH_FULL_IMAGE:figures/full_fig_p013_1.png] view at source ↗
Figure 2
Figure 2. Capability score vs. safety scores for various LLMs in the [PITH_FULL_IMAGE:figures/full_fig_p016_2.png] view at source ↗
Figure 3
Figure 3. First solve time (FST) and total count of successful models for each CTF challenge. The y-axis [PITH_FULL_IMAGE:figures/full_fig_p017_3.png] view at source ↗
Figures from the paper (33 more)
Figure 4
Figure 4. Figure 4: Performance of LLM agents across CTF challenges in CyBench. [PITH_FULL_IMAGE:figures/full_fig_p018_4.png]
Figure 5
Figure 5. Figure 5: Overview of the practical AI cyber-exploitation benchmark. [PITH_FULL_IMAGE:figures/full_fig_p020_5.png]
Figure 6
Figure 6. Figure 6: Capability score vs. safety scores for various LLMs in the [PITH_FULL_IMAGE:figures/full_fig_p024_6.png]
Figure 7
Figure 7. Figure 7: Performance of LLM agents across challenges in PACEBench. [PITH_FULL_IMAGE:figures/full_fig_p025_7.png]
Figure 8
Figure 8. Figure 8: Experiment overview of BioLP-Bench and LAB-Bench ProtocolQA. [PITH_FULL_IMAGE:figures/full_fig_p029_8.png]
Figure 9
Figure 9. Figure 9: Model performance on biological protocol diagnosis and troubleshooting benchmarks, with human [PITH_FULL_IMAGE:figures/full_fig_p031_9.png]
Figure 10
Figure 10. Figure 10: Capability Score vs. Safety Scores (Frontier) for various LLMs in the [PITH_FULL_IMAGE:figures/full_fig_p032_10.png]
Figure 11
Figure 11. Figure 11: Model performance on biological hazardous knowledge and safety alignment benchmarks, with [PITH_FULL_IMAGE:figures/full_fig_p034_11.png]
Figure 12
Figure 12. Figure 12: Models are color-coded by family, with point size representing model parameters and shape [PITH_FULL_IMAGE:figures/full_fig_p036_12.png]
Figure 13
Figure 13. Figure 13: Model performance on chemical hazardous knowledge and safety alignment benchmarks, with [PITH_FULL_IMAGE:figures/full_fig_p039_13.png]
Figure 14
Figure 14. Figure 14: Models are color-coded by family, with point size representing model parameters and shape [PITH_FULL_IMAGE:figures/full_fig_p040_14.png]
Figure 15
Figure 15. Figure 15: Experiment overview of persuasion and manipulation process. [PITH_FULL_IMAGE:figures/full_fig_p043_15.png]
Figure 16
Figure 16. Figure 16: Capability Score vs. Safety Scores for various LLMs in the [PITH_FULL_IMAGE:figures/full_fig_p045_16.png]
Figure 17
Figure 17. Figure 17: (a) A stacked bar chart showing percentage of successful persuasion (blue), no effect (orange), and [PITH_FULL_IMAGE:figures/full_fig_p046_17.png]
Figure 18
Figure 18. Figure 18: The comparison between the LLM voters and the human voter. (a)The comparison between the [PITH_FULL_IMAGE:figures/full_fig_p048_18.png]
Figure 19
Figure 19. Figure 19: Capability Score vs. Safety Scores for various LLMs in the [PITH_FULL_IMAGE:figures/full_fig_p050_19.png]
Figure 20
Figure 20. Figure 20: The simplified evaluation pipeline of MASK Benchmark ( [PITH_FULL_IMAGE:figures/full_fig_p053_20.png]
Figure 21
Figure 21. Figure 21: Dishonesty example: Known Facts (Ren et al., 2025b). Metrics. In this experiment, we define multiple metrics to quantify the model’s performance, particularly its honesty and accuracy. We formally define two evaluation metrics used to assess model behavior in belief-s…
Figure 22
Figure 22. Figure 22: Capability Score vs. safety Scores (honesty score) for various LLMs in the [PITH_FULL_IMAGE:figures/full_fig_p055_22.png]
Figure 23
Figure 23. Figure 23: Example of sandbagging. Problem Statement. Certain advanced AI systems may engage in sandbagging—the deliberate suppression or concealment of capabilities—in response to evaluation scenarios explicitly intended to assess dangerous or frontier behaviors (Meinke et al.,…
Figure 24
Figure 24. Figure 24: Capability Score vs. Safety Scores in Equation [PITH_FULL_IMAGE:figures/full_fig_p060_24.png]
Figure 25
Figure 25. Figure 25: The evaluation overview of AI response in safety evaluation experiment in uncontrolled AI R&D. [PITH_FULL_IMAGE:figures/full_fig_p063_25.png]
Figure 26
Figure 26. Figure 26: Capability Score vs. safety Scores for various LLMs in the [PITH_FULL_IMAGE:figures/full_fig_p066_26.png]
Figure 27
Figure 27. Figure 27: Example of uncontrolled AI R&D deceptive behavior from the perspective of AI self-annotation in [PITH_FULL_IMAGE:figures/full_fig_p068_27.png]
Figure 28
Figure 28. Figure 28: Capability Score vs. safety Scores for various LLMs in the [PITH_FULL_IMAGE:figures/full_fig_p070_28.png]
Figure 29
Figure 29. Figure 29: Overview of self-replication process of the LLM agent. [PITH_FULL_IMAGE:figures/full_fig_p072_29.png]
Figure 30
Figure 30. Figure 30: Capability Score vs. Safety Score for various LLMs in the [PITH_FULL_IMAGE:figures/full_fig_p075_30.png]
Figure 31
Figure 31. Figure 31: Success rates (Milestone 1–4) and over-scaling rate for o4-mini-2025-04-16 under previous setting, [PITH_FULL_IMAGE:figures/full_fig_p076_31.png]
Figure 32
Figure 32. Figure 32: Overview of the fraud process in social systems initiated by multiple malicious agents. [PITH_FULL_IMAGE:figures/full_fig_p083_32.png]
Figure 33
Figure 33. Figure 33: The collaboration framework of malicious agents in the fraud process. [PITH_FULL_IMAGE:figures/full_fig_p083_33.png]
Figure 34
Figure 34. Figure 34: Capability Score vs. Safety Score for various models in the Multi-agent Fraud in Social Systems scenario of Collusion. Models are color-coded by family, with point size representing model parameters and shape indicating either a Standard (circle) or Reasoning (square)…
Figure 35
Figure 35. Figure 35: An example of multi-agent malicious collusion in the fraud scenario. [PITH_FULL_IMAGE:figures/full_fig_p087_35.png]
Figure 36
Figure 36. Figure 36: An example of an autonomous phishing website generated by DeepSeek-R1-driven agents. [PITH_FULL_IMAGE:figures/full_fig_p087_36.png]

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AI Agents Enable Adaptive Computer Worms

    cs.CR 2026-06 unverdicted novelty 6.0 of 10

    AI agents enable adaptive computer worms that propagate autonomously by reasoning about targets and synthesizing attacks using LLMs on stolen compute.

Reference graph

Works this paper leans on

51 extracted references · 9 canonical work pages · cited by 1 Pith paper

  1. [3]

    Anthropic

    URLhttps://arxiv.org/abs/2409.03789. Anthropic. Anthropic’s responsible scaling policy, September

  2. [6]

    Replibench: Evaluating the autonomous replication capabilities of language model agents.arXiv preprint arXiv:2504.18565,

    Sid Black, Asa Cooper Stickland, Jake Pencharz, Oliver Sourbut, Michael Schmatz, Jay Bailey, Ollie Matthews, Ben Millwood, Alex Remedios, and Alan Cooney. Replibench: Evaluating the autonomous replication capabilities of language model agents.arXiv preprint arXiv:2504.18565,

  3. [8]

    Accessed: 2025-06-19

    URLhttps://storage.googleapis.com/deepmind-media/ gemini/gemini_v2_5_report.pdf. Accessed: 2025-06-19. DeepSeek-AI. Deepseek-v3 technical report,

  4. [9]

    DeepSeek-AI

    URLhttps://arxiv.org/abs/2412.19437. DeepSeek-AI. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,

  5. [10]

    URL https://arxiv.org/abs/2501.12948. 92 Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report Sunishchal Dev, Charles Teague, Kyle Brady, Ying-Chiang Lee, Sarah L Gebauer, Henry Alexander Bradley, Grant Ellison, Bria Persaud, Jordan Despanie, Barbara Del Castello, et al. Toward Comprehensive Benchmarking of the Biological Kn...

  6. [11]

    Sciknoweval: Evaluating multi-level scientific knowledge of large language models

    Kehua Feng, Keyan Ding, Weijie Wang, Xiang Zhuang, Zeyuan Wang, Ming Qin, Yu Zhao, Jianhua Yao, Qiang Zhang, and Huajun Chen. Sciknoweval: Evaluating multi-level scientific knowledge of large language models. arXiv preprint arXiv:2406.09098,

  7. [12]

    Accessed: 2025-07-13

    URL https://www.frontiermodelforum.org/updates/ issue-brief-preliminary-taxonomy-of-ai-bio-safety-evaluations/ . Accessed: 2025-07-13. Jasper GÃk,tting, Pedro Medeiros, Jon G Sanders, Nathaniel Li, Long Phan, Karam Elabd, Lennart Justen, Dan Hendrycks, and Seth Donoughe. Virology capabilities test (vct): A multimodal virology q&a benchmark. arXiv preprint...

  8. [13]

    Gemini 2.5 flash preview model card, 2025a

    Google. Gemini 2.5 flash preview model card, 2025a. URL https://storage.googleapis.com/ model-cards/documents/gemini-2.5-flash-preview.pdf. Accessed: 2025-06-19. Google. Google’s frontier safety framework, 2025b. URL https://storage.googleapis. com/deepmind-media/DeepMind.com/Blog/introducing-the-frontier-safety-framework/ fsf-technical-report.pdf. Aaron ...

Show all 51 references
  1. [14]

    Alignment faking in large language models

    Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, et al. Alignment faking in large language models. arXiv preprint arXiv:2412.14093,

  2. [15]

    Lewis Hammond, Alan Chan, Jesse Clifton, Jason Hoelscher-Obermaier, Akbir Khan, Euan McLean, Chandler Smith, Wolfram Barfuss, Jakob Foerster, Tomáš Gavenčiak, The Anh Han, Edward Hughes, Vojtěch Kovařík, Jan Kulveit, Joel Z. Leibo, Caspar Oesterheld, Christian Schroeder de Wit...

  3. [16]

    Measuring massive multitask language understanding.arXiv preprint arXiv:2009.03300,

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding.arXiv preprint arXiv:2009.03300,

  4. [18]

    Llama guard: Llm-based input-output safeguard for human-ai conversations.arXiv preprint arXiv:2312.06674,

    Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al. Llama guard: Llm-based input-output safeguard for human-ai conversations.arXiv preprint arXiv:2312.06674,

  5. [20]

    Livecodebench: Holistic and contamination free evaluation of large language models for code.arXiv preprint arXiv:2403.07974,

    Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar- Lezama, Koushik Sen, and Ion Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code.arXiv preprint arXiv:2403.07974,

  6. [21]

    URLhttps://arxiv.org/abs/2403. 07974. Jiaming Ji, Donghai Hong, Borong Zhang, Boyuan Chen, Josef Dai, Boren Zheng, Tianyi Qiu, Boxun Li, and Yaodong Yang. Pku-saferlhf: Towards multi-level safety alignment for llms with human preference.arXiv preprint arXiv:2406.15513,

  7. [22]

    Mitigating deceptive alignment via self-monitoring.arXiv preprint arXiv:2505.18807,

    Jiaming Ji, Wenqi Chen, Kaile Wang, Donghai Hong, Sitong Fang, Boyuan Chen, Jiayi Zhou, Juntao Dai, Sirui Han, Yike Guo, et al. Mitigating deceptive alignment via self-monitoring.arXiv preprint arXiv:2505.18807,

  8. [23]

    Sosbench: Benchmarking safety alignment on scientific knowledge

    Fengqing Jiang, Fengbo Ma, Zhangchen Xu, Yuetai Li, Bhaskar Ramasubramanian, Luyao Niu, Bo Li, Xianyan Chen, Zhen Xiang, and Radha Poovendran. Sosbench: Benchmarking safety alignment on scientific knowledge. arXiv preprint arXiv:2505.21605,

  9. [24]

    Secbench: A comprehensive multi-dimensional benchmarking dataset for llms in cybersecurity.arXiv preprint arXiv:2412.20787,

    Pengfei Jing, Mengyun Tang, Xiaorong Shi, Xing Zheng, Sen Nie, Shi Wu, Yong Yang, and Xiapu Luo. Secbench: A comprehensive multi-dimensional benchmarking dataset for llms in cybersecurity.arXiv preprint arXiv:2412.20787,

  10. [25]

    Lab-bench: Measuring capabilities of language models for biology research.arXiv preprint arXiv:2407.10362,

    Jon M Laurent, Joseph D Janizek, Michael Ruzo, Michaela M Hinks, Michael J Hammerling, Siddharth Narayanan, Manvitha Ponnapati, Andrew D White, and Samuel G Rodriques. Lab-bench: Measuring capabilities of language models for biology research.arXiv preprint arXiv:2407.10362,

  11. [26]

    Salad-bench: A hierarchical and comprehensive safety benchmark for large language models.arXiv preprint arXiv:2402.05044, 2024a

    Lijun Li, Bowen Dong, Ruohui Wang, Xuhao Hu, Wangmeng Zuo, Dahua Lin, Yu Qiao, and Jing Shao. Salad-bench: A hierarchical and comprehensive safety benchmark for large language models.arXiv preprint arXiv:2402.05044, 2024a. Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue,...

  12. [27]

    American invitational mathematics examination 2024,

    Mathematical Association of America. American invitational mathematics examination 2024,

  13. [28]

    Contest problems from the 2024 AIME competition

    URL https://artofproblemsolving.com/wiki/index.php/2024_AIME_I. Contest problems from the 2024 AIME competition. Sandra C Matz, Jacob D Teeny, Sumer S Vaid, Heinrich Peters, Gabriella M Harari, and Moran Cerf. The potential of generative ai for personalized persuasion at scale...

  14. [29]

    Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi

    URLhttps://arxiv.org/abs/2504.06017. Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi. Tree of attacks: Jailbreaking black-box llms automatically.Advances in Neural Information Processing Systems, 37:61065–61105,

  15. [30]

    Frontier models are capable of in-context scheming.arXiv preprint arXiv:2412.04984,

    94 Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report Alexander Meinke, Bronson Schoen, Jérémy Scheurer, Mikita Balesni, Rusheb Shah, and Marius Hobbhahn. Frontier models are capable of in-context scheming.arXiv preprint arXiv:2412.04984,

  16. [31]

    Responsible scaling policies (rsps).https://metr.org/blog/2023-09-26-rsp/, 09

    METR. Responsible scaling policies (rsps).https://metr.org/blog/2023-09-26-rsp/, 09

  17. [33]

    URL https://arxiv.org/abs/2311.12983. Adrian Mirza, Nawaf Alampara, Sreekanth Kunchapu, Martiño Ríos-García, Benedict Emoekabu, Aswanth Krishnan, Tanya Gupta, Mara Schilling-Wilhelmi, Macjonathan Okereke, Anagha Aneesh, et al. Are large language models superhuman chemists?arXi...

  18. [34]

    Accessed: 2025-07-11. OpenAI. Building an early warning system for llm-aided biologi- cal threat creation, January

  19. [35]

    Accessed: 2025-07-13

    URL https://openai.com/index/ building-an-early-warning-system-for-llm-aided-biological-threat-creation/ . Accessed: 2025-07-13. OpenAI. Openai’s preparedness framework, September

  20. [36]

    Accessed: 2025-06-19

    URL https://openai.com/index/o3-o4-mini-system-card/. Accessed: 2025-06-19. Xudong Pan, Jiarun Dai, Yihe Fan, Minyuan Luo, Changyi Li, and Min Yang. Large language model-powered ai systems achieve self-replication with no human intervention.arXiv preprint arXiv:2503.17378,

  21. [37]

    OpenAI Re- search Publication

    URL https://openai.com/research/ building-an-early-warning-system-for-llm-aided-biological-threat-creation . OpenAI Re- search Publication. Mary Phuong, Matthew Aitchison, Elliot Catt, Sarah Cogan, Alexandre Kaskasoli, Victoria Krakovna, David Lindner, Matthew Rahtz, Yannis As...

  22. [38]

    Towards empathetic open-domain conversation models: A new benchmark and dataset.arXiv preprint arXiv:1811.00207,

    Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. Towards empathetic open-domain conversation models: A new benchmark and dataset.arXiv preprint arXiv:1811.00207,

  23. [40]

    Qibing Ren, Hao Li, Dongrui Liu, Zhanxu Xie, Xiaoya Lu, Yu Qiao, Lei Sha, Junchi Yan, Lizhuang Ma, and Jing Shao

    URL https://arxiv.org/abs/2311.12022. Qibing Ren, Hao Li, Dongrui Liu, Zhanxu Xie, Xiaoya Lu, Yu Qiao, Lei Sha, Junchi Yan, Lizhuang Ma, and Jing Shao. Derail yourself: Multi-turn llm jailbreak attack through self-discovered clues.arXiv preprint arXiv:2410.10700,

  24. [41]

    When autonomy goes rogue: Preparing for risks of multi-agent collusion in social systems, 2025a

    Qibing Ren, Sitao Xie, Longxuan Wei, Zhenfei Yin, Junchi Yan, Lizhuang Ma, and Jing Shao. When autonomy goes rogue: Preparing for risks of multi-agent collusion in social systems, 2025a. URLhttp: //arxiv.org/abs/2507.14660. Richard Ren, Arunim Agarwal, Mantas Mazeika, Cristina...

  25. [42]

    Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R Johnston, et al

    URL https://research.ai45.shlab.org.cn/safework-f1-framework.pdf. Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R Johnston, et al. Towards understanding sycophancy in language model...

  26. [44]

    Joey Skaf, Luis Ibanez-Lissen, Robert McCarthy, Connor Watts, Vasil Georgiv, Hannes Whittingham, Lorena Gonzalez-Manzano, David Lindner, Cameron Tice, Edward James Young, et al

    URL https://arxiv.org/abs/2404.10952. Joey Skaf, Luis Ibanez-Lissen, Robert McCarthy, Connor Watts, Vasil Georgiv, Hannes Whittingham, Lorena Gonzalez-Manzano, David Lindner, Cameron Tice, Edward James Young, et al. Large language models can learn and generalize steganographic...

  27. [47]

    github.io/blog/qwq-32b/

    URLhttps://qwenlm. github.io/blog/qwq-32b/. NorbertTihanyi, MohamedAmineFerrag, RidhiJain, TamasBisztray, andMerouaneDebbah. Cybermetric: A benchmark dataset based on retrieval-augmented generation for evaluating llms in cybersecurity knowledge. In 2024 IEEE International Conf...

  28. [48]

    Teun van der Weij, Felix Hofstätter, Ollie Jaffe, Samuel F Brown, and Francis Rhys Ward

    doi: 10.1109/CSR61664.2024.10679494. Teun van der Weij, Felix Hofstätter, Ollie Jaffe, Samuel F Brown, and Francis Rhys Ward. Ai sandbagging: Language models can strategically underperform on evaluations.arXiv preprint arXiv:2406.07358, 2024a. Teun van der Weij, Felix Hofstätt...

  29. [50]

    An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al

    URLhttps://arxiv.org/abs/2505.12786. An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al. Qwen2. 5 technical report.arXiv preprint arXiv:2412.15115, 2024a. An Yang, Anfeng Li, Baosong Yang, Beichen Zha...

  30. [51]

    Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao

    URL https://arxiv.org/abs/2505.16557. Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. ReAct: Syn- ergizing reasoning and acting in language models. InInternational Conference on Learning Representations (ICLR),

  31. [52]

    Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts.arXiv preprint arXiv:2309.10253,

    Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing. Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts.arXiv preprint arXiv:2309.10253,

  32. [53]

    Cybench: A framework for evaluating cybersecurity capabilities and risks of language models.arXiv preprint arXiv:2408.08926,

    Andy K Zhang, Neil Perry, Riya Dulepet, Joey Ji, Celeste Menders, Justin W Lin, Eliot Jones, Gashon Hussein, Samantha Liu, Donovan Jasper, et al. Cybench: A framework for evaluating cybersecurity capabilities and risks of language models.arXiv preprint arXiv:2408.08926,

  33. [55]

    Terry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu, Wenhao Yu, Ratnadira Widyasari, Imam Nur Bani Yusuf, Haolan Zhan, Junda He, Indraneil Paul, et al

    URL https://arxiv.org/abs/2311.07911. Terry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu, Wenhao Yu, Ratnadira Widyasari, Imam Nur Bani Yusuf, Haolan Zhan, Junda He, Indraneil Paul, et al. Bigcodebench: Benchmarking code generation with diverse function calls and complex instru...

  34. [56]

    Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson

    URL https://arxiv.org/abs/2406.15877. Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models.arXiv preprint arXiv:2307.15043,

  35. [1953]

    Gpt-4o system card.arXiv preprint arXiv:2410.21276,

    Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. Gpt-4o system card.arXiv preprint arXiv:2410.21276,

  36. [2002]

    Scheming ais: Will ais fake alignment during training in order to get power?arXiv preprint arXiv:2311.08379,

    Joe Carlsmith. Scheming ais: Will ais fake alignment during training in order to get power?arXiv preprint arXiv:2311.08379,

  37. [2006]

    Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

    Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, et al. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. arXiv preprint arXiv:2406.01574,

  38. [2014]

    Biolp-bench: Measuring understanding of biological lab protocols by large language models

    Igor Ivanov. Biolp-bench: Measuring understanding of biological lab protocols by large language models. bioRxiv, pp. 2024–08,

  39. [2015]

    International ai safety report.arXiv preprint arXiv:2501.17805,

    Yoshua Bengio, Sören Mindermann, Daniel Privitera, Tamay Besiroglu, Rishi Bommasani, Stephen Casper, Yejin Choi, Philip Fox, Ben Garfinkel, Danielle Goldfarb, et al. International ai safety report.arXiv preprint arXiv:2501.17805,

  40. [2022]

    Qwen Team

    URL https://arxiv.org/abs/ 2210.09261. Qwen Team. Qwq-32b: Embracing the power of reinforcement learning, March

  41. [2023]

    com/production/files/responsible-scaling-policy-1.0.pdf

    URLhttps://www-files.anthropic. com/production/files/responsible-scaling-policy-1.0.pdf. Anthropic. Claude 3.7 sonnet system card, 2025a. URL https://assets.anthropic.com/m/ 785e231869ea8b3b/original/claude-3-7-sonnet-system-card.pdf . Accessed: 2025-06-19. Anthropic. System c...

  42. [2024]

    Mistral AI

    Accessed: 2024-04-01. Mistral AI. Mistral small 3.1,

  43. [2025]

    Accessed: 2025-06-19

    URL https://mistral.ai/news/mistral-small-3-1. Accessed: 2025-06-19. UK AI Security Institute. Inspect AI: Framework for Large Language Model Evaluations. URLhttps: //github.com/UKGovernmentBEIS/inspect_ai. Ibrahim Alshehri, Adnan Alshehri, Abdulrahman Almalki, Majed Bamardouf...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.