REVIEW 8 cited by
Concrete Problems in AI Safety, Revisited
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Concrete Problems in AI Safety, Revisited
read the original abstract
As AI systems proliferate in society, the AI community is increasingly preoccupied with the concept of AI Safety, namely the prevention of failures due to accidents that arise from an unanticipated departure of a system's behavior from designer intent in AI deployment. We demonstrate through an analysis of real world cases of such incidents that although current vocabulary captures a range of the encountered issues of AI deployment, an expanded socio-technical framing will be required for a more complete understanding of how AI systems and implemented safety mechanisms fail and succeed in real life.
Forward citations
Cited by 8 Pith papers
-
Towards a Science of AI Agent Reliability
Measuring 14 AI agents across two benchmarks, the paper finds 18 months of accuracy gains (≈0.21/yr) bought only small reliability gains (0.03–0.10/yr) under its consistency/robustness/predictability/safety framework.
-
Relative Principals, Pluralistic Alignment, and the Structural Value Alignment Problem
AI value alignment is reconceptualized as a pluralistic governance problem arising along three axes—objectives, information, and principals—making it inherently context-dependent and unsolvable by technical design alone.
-
Risk-Calibrated Learning: Minimizing Fatal Errors in Medical AI
Risk-Calibrated Learning reduces critical error rates in medical AI by 20-92% across four imaging datasets by embedding a severity matrix into the optimization.
-
How Generative AI Empowers Attackers and Defenders Across the Trust & Safety Landscape
Generative AI boosts attackers' ability to create harmful content at scale while also enabling defenders to detect threats, support users, and improve moderation processes.
-
Unsafe at any AUC: Unlearned Lessons from Sociotechnical Disasters for Responsible AI
AI safety is a systems-governance problem: six recurring organizational failure patterns from past disasters remain unlearned in AI development, so component-level fixes like benchmarks and alignment cannot deliver safety.
-
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization
Proxy RL produces a staged proxy-internalization capability that emerges before and predicts reward hacking in coding environments.
-
Improving Model Safety by Targeted Error Correction
A dual GBDT error classifier reduces dangerous misclassifications by 12-34% on medical and animal image datasets with under 2% added latency.
-
StepGuard: Guarding Web Navigation via Single-Step Calibration
StepGuard framework with DDPO and CANR claims SOTA navigation and answer accuracy on web benchmarks by switching policies and triggering reflection on low-confidence steps.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.