Pith. sign in

REVIEW 1 cited by

An overview of 11 proposals for building safe advanced AI

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.07532 v1 pith:ITZIZNRH submitted 2020-12-04 cs.LG cs.AI

classification cs.LGcs.AI
keywords alignmentproposalsadvancedanalysisbuildingcomparativecompetitivenesscomponents
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper analyzes and compares 11 different proposals for building safe advanced AI under the current machine learning paradigm, including major contenders such as iterated amplification, AI safety via debate, and recursive reward modeling. Each proposal is evaluated on the four components of outer alignment, inner alignment, training competitiveness, and performance competitiveness, of which the distinction between the latter two is introduced in this paper. While prior literature has primarily focused on analyzing individual proposals, or primarily focused on outer alignment at the expense of inner alignment, this analysis seeks to take a comparative look at a wide range of proposals including a comparative analysis across all four previously mentioned components.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 6 citations worldwide. Full citation record

  1. Unraveling Token Prediction Refinement and Identifying Essential Layers in Language Models

    cs.CL 2025-01 conditional novelty 4.0 of 10

    In GPT-2 multi-document QA, the layer gap between the first correct top-1 token prediction and its stable final form is larger when relevant information is in the middle of the context.

Pith tools