Pith. sign in

Citation notice #6161 · 2026-07-11 03:19:14.036358+00:00

Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Correction Crossref Open

cites I HATE YOU, which carries a correction notice dated 2015-12-15. One-hop deterministic notice: the citation edge exists in the Pith bibliography graph; no model judged whether the citation was load-bearing.

This is not a judgment on the citing paper.

Citing paper Event page Original DOI Notice DOI File a formal challenge All reference changes

01Evidence

Raw extraction · bibliography line · bibliography index 1

URL https://arxiv.org/abs/2105.12400. Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training, 2018. URL https:// s3-us-west-2.amazonaws.com/openai-assets/research-covers/ language-unsupervised/language_understanding_paper.pdf. Javier Rando and Florian Tramèr. Universal jailbreak backdoors from poisoned human feedback, 2023. Jérémy Scheurer, Mikita Balesni, and Marius Hobbhahn. Technical report: Large language models can strategically deceive their users when put under pressure. arXiv preprint arXiv:2311.07590, 2023. John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. CoRR, abs/1707.06347, 2017. URL http://arxiv.org/abs/ 1707.06347. Roei Schuster, Congzheng Song, Eran Tromer, and Vitaly Shmatikov. You autocomplete me: Poisoning vulnerabilities in neural code completion. In 30th USENIX Security Symposium (USENIX Security 21), pp. 1559–1575, 2021. Murray Shanahan, Kyle McDonell, and Laria Reynolds. Role play with large language models. Nature, pp. 1–6, 2023. Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao, and Tom Goldstein. On t

02Event

Type
Correction
Source
Crossref
Original DOI
10.1016/j.jneumeth.2013.09.010
Notice DOI
10.1016/j.jneumeth.2015.11.021
Date
2015-12-15
Title
Corrigendum to ‘Meta-analysis of data from animal studies: A practical guide’
Reasons
['Erratum']
Work
I HATE YOU (2018) Journal of Neuroscience Methods

Schema constants (for re-runners): correction · crossref

03Dispute this notice

If this citation does not depend on the flagged claim, or the event is wrong, say so. Disputes are public. For a signed challenge against the paper itself, use the formal challenge form.