Pith. sign in

Pith Integrity · Reference Change Desk · events-only

Reference changes

See every place this corpus cites a work that was later retracted, corrected, withdrawn, or placed under expression of concern: exact quote, event source, and what happened next. No model judges the citation.

A notice on this page means a citing paper's bibliography includes a work with a published scholarly-record event. It is not a judgment on the citing paper.

Scoped to citing paper 2401.05566 · clear

01Events with corpus notices

02One-hop citation notices (secondary index)

Correction Crossref Open
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

cites I HATE YOU · ref [1] · event 2015-12-15 · 2401.05566 · event page · DOI 10.1016/j.jneumeth.2013.09.010

Raw extraction · bibliography line

URL https://arxiv.org/abs/2105.12400. Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training, 2018. URL https:// s3-us-west-2.amazonaws.com/openai-assets/research-covers/ language-unsupervised/language_understanding_paper.pdf. Javier Rando and Florian Tramèr. Universal jailbreak backdoors from poisoned human feedback, 2023. Jérémy Scheurer, Mikita Balesni, and Marius Hobbhahn. Technical report: Large lan…