Pith. sign in

arXiv preprint arXiv:2511.02997 , year =

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

fields

cs.AI 2 cs.CR 1

years

2026 3

verdicts

UNVERDICTED 3

representative citing papers

Honeypot Protocol

cs.CR · 2026-04-14 · unverdicted · novelty 7.0

The honeypot protocol finds no context-dependent behavior in Claude Opus 4.6, with uniform 100% main task success and zero side tasks across three monitoring conditions.

citing papers explorer

Showing 3 of 3 citing papers.

  • Honeypot Protocol cs.CR · 2026-04-14 · unverdicted · none · ref 8

    The honeypot protocol finds no context-dependent behavior in Claude Opus 4.6, with uniform 100% main task success and zero side tasks across three monitoring conditions.

  • Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety cs.AI · 2026-06-03 · unverdicted · none · ref 11

    Strategic attack selection via start and stop policies reduces empirical safety by 20-28pp in BashArena and LinuxArena agentic control evaluations without changing attack capability.

  • Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute cs.AI · 2026-05-14 · unverdicted · none · ref 50

    Diverse ensembles of prompted and fine-tuned GPT-4.1-Mini monitors achieve 2.4x better detection of flawed code solutions than homogeneous ensembles on adversarial inputs.