Pith. sign in

Responsible scaling policy

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.AI 1

years

2025 1

verdicts

UNVERDICTED 1

representative citing papers

An alignment safety case sketch based on debate

cs.AI · 2025-05-06 · unverdicted · novelty 6.0

The paper argues that if a debate game reaches equilibrium, has exploration guarantees, and is run through online training, an AI R&D agent can be shown to make at most an epsilon-fraction of errors, which suffices for safety in a low-stakes context.

citing papers explorer

Showing 1 of 1 citing paper.

  • An alignment safety case sketch based on debate cs.AI · 2025-05-06 · unverdicted · none · ref 3

    The paper argues that if a debate game reaches equilibrium, has exploration guarantees, and is run through online training, an AI R&D agent can be shown to make at most an epsilon-fraction of errors, which suffices for safety in a low-stakes context.