Pith. sign in

Uncovering deceptive tendencies in language models: A simulated company ai assistant

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

fields

cs.AI 3

years

2026 1 2024 2

representative citing papers

OpenAI o1 System Card

cs.AI · 2024-12-21 · unverdicted · novelty 4.0

OpenAI reports that chain-of-thought reasoning in o1 models enables deliberative alignment, yielding state-of-the-art results on selected safety benchmarks for illicit advice, stereotypes, and jailbreaks.

citing papers explorer

Showing 3 of 3 citing papers.