Pith. sign in

REVIEW 9 cited by

Risk thresholds for frontier AI

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.14713 v1 pith:6VTJ44CL submitted 2024-06-20 cs.CY

classification cs.CY
keywords riskthresholdscapabilitydefinetheymuchapproachdecision-making
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Frontier artificial intelligence (AI) systems could pose increasing risks to public safety and security. But what level of risk is acceptable? One increasingly popular approach is to define capability thresholds, which describe AI capabilities beyond which an AI system is deemed to pose too much risk. A more direct approach is to define risk thresholds that simply state how much risk would be too much. For instance, they might state that the likelihood of cybercriminals using an AI system to cause X amount of economic damage must not increase by more than Y percentage points. The main upside of risk thresholds is that they are more principled than capability thresholds, but the main downside is that they are more difficult to evaluate reliably. For this reason, we currently recommend that companies (1) define risk thresholds to provide a principled foundation for their decision-making, (2) use these risk thresholds to help set capability thresholds, and then (3) primarily rely on capability thresholds to make their decisions. Regulators should also explore the area because, ultimately, they are the most legitimate actors to define risk thresholds. If AI risk estimates become more reliable, risk thresholds should arguably play an increasingly direct role in decision-making.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Frontier Safety Policies Plus

    cs.CY 2025-01 conditional novelty 6.0 of 10

    Frontier safety policies should be rebuilt around a standardized taxonomy of precursory capabilities and a mutual feedback mechanism with AI safety cases.

  2. Safety case template for frontier AI: A cyber inability argument

    cs.CY 2024-11 accept novelty 6.0 of 10

    A proof-of-concept safety case template formalizes an inability argument for offensive cyber risk using risk models, proxy tasks, and evaluation results.

  3. Harmonizing AI Safety Thresholds

    cs.AI 2026-07 conditional novelty 5.0 of 10

    The authors propose harmonized AI capability floors: non-zero full-chain TLO cyber completion triggers safeguards, and AI progress at 5× trend for 3 months triggers safeguards, with biorisk left as a diagnostic.

  4. Technical Requirements for Halting Dangerous AI Activities

    cs.AI 2025-07 conditional novelty 5.0 of 10

    A taxonomy of compute-centric technical interventions, graded by readiness and mapped to five AI governance plans, argues that halting dangerous AI requires substantial control over AI compute.

  5. In Which Areas of Technical AI Safety Could Geopolitical Rivals Cooperate?

    cs.CY 2025-04 conditional novelty 5.0 of 10

    Based on a four-risk typology, the paper concludes that verification mechanisms and codified protocols are the least risky areas for cooperation between geopolitical rivals on technical AI safety.

  6. Robustness tests for biomedical foundation models should tailor to specifications

    cs.SE 2025-02 conditional novelty 5.0 of 10

    The paper proposes task-tailored 'robustness specifications' to guide robustness testing of biomedical foundation models across their lifecycle.

  7. Governing AI Beyond the Pretraining Frontier

    cs.CY 2025-01 conditional novelty 5.0 of 10

    If pretraining scaling plateaus while capabilities keep rising through reasoning models, compute-based triggers in the EU AI Act and US export controls will miss the main sources of risk.

  8. Quantifying detection rates for dangerous capabilities: a theoretical model of dangerous capability evaluations

    cs.AI 2024-12 conditional novelty 5.0 of 10

    A new model quantifies how test sensitivity, capability growth, and threshold placement determine bias and detection lag in dangerous AI evaluations.

  9. What Information Should Be Shared with Whom "Before and During Training"?

    cs.CY 2024-12 unverdicted novelty 4.0 of 10

    A concrete transparency checklist for pre-training disclosure under the Frontier AI Safety Commitments, balancing public sharing against commercial and security risks.

Pith tools