Pith. sign in

REVIEW 8 cited by

Concrete Problems in AI Safety, Revisited

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.10899 v1 pith:YT3F6NDZ submitted 2023-12-18 cs.CY cs.AI

Concrete Problems in AI Safety, Revisited

classification cs.CY cs.AI
keywords safetydeploymentrealsystemsaccidentsalthoughanalysisarise
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

As AI systems proliferate in society, the AI community is increasingly preoccupied with the concept of AI Safety, namely the prevention of failures due to accidents that arise from an unanticipated departure of a system's behavior from designer intent in AI deployment. We demonstrate through an analysis of real world cases of such incidents that although current vocabulary captures a range of the encountered issues of AI deployment, an expanded socio-technical framing will be required for a more complete understanding of how AI systems and implemented safety mechanisms fail and succeed in real life.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Towards a Science of AI Agent Reliability

    cs.AI 2026-02 conditional novelty 6.0

    Measuring 14 AI agents across two benchmarks, the paper finds 18 months of accuracy gains (≈0.21/yr) bought only small reliability gains (0.03–0.10/yr) under its consistency/robustness/predictability/safety framework.

  2. Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem

    cs.CY 2026-04 unverdicted novelty 5.0

    AI value alignment is reconceptualized as a pluralistic governance problem arising along three axes—objectives, information, and principals—making it inherently context-dependent and unsolvable by technical design alone.

  3. Risk-Calibrated Learning: Minimizing Fatal Errors in Medical AI

    cs.CV 2026-04 unverdicted novelty 5.0

    Risk-Calibrated Learning reduces critical error rates in medical AI by 20-92% across four imaging datasets by embedding a severity matrix into the optimization.

  4. How Generative AI Empowers Attackers and Defenders Across the Trust & Safety Landscape

    cs.HC 2025-11 unverdicted novelty 5.0

    Generative AI boosts attackers' ability to create harmful content at scale while also enabling defenders to detect threats, support users, and improve moderation processes.

  5. Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI

    cs.CY 2026-07 accept novelty 4.0

    AI safety is a systems-governance problem: six recurring organizational failure patterns from past disasters remain unlearned in AI development, so component-level fixes like benchmarks and alignment cannot deliver safety.

  6. Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization

    cs.AI 2026-06 unverdicted novelty 4.0

    Proxy RL produces a staged proxy-internalization capability that emerges before and predicts reward hacking in coding environments.

  7. Improving Model Safety by Targeted Error Correction

    cs.AI 2026-05 unverdicted novelty 4.0

    A dual GBDT error classifier reduces dangerous misclassifications by 12-34% on medical and animal image datasets with under 2% added latency.

  8. StepGuard: Guarding Web Navigation via Single-Step Calibration

    cs.AI 2026-06 unverdicted novelty 3.0

    StepGuard framework with DDPO and CANR claims SOTA navigation and answer accuracy on web benchmarks by switching policies and triggering reflection on low-confidence steps.