Pith. sign in

Title resolution pending

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.AI 1

years

2026 1

verdicts

UNVERDICTED 1

representative citing papers

Scaling Trends for Lie Detector Oversight in Preference Learning

cs.AI · 2026-07-02 · unverdicted · novelty 5.0

SOLiD scales to 405B models with undetected deception dropping to 14% at 99% TPR, permits removing human labelers from fine-tuning without significant deception increase, but fails under distribution shift.

citing papers explorer

Showing 1 of 1 citing paper.

  • Scaling Trends for Lie Detector Oversight in Preference Learning cs.AI · 2026-07-02 · unverdicted · none · ref 8

    SOLiD scales to 405B models with undetected deception dropping to 14% at 99% TPR, permits removing human labelers from fine-tuning without significant deception increase, but fails under distribution shift.