Correction
Open
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
ref [1] · 2401.05566 · notice #6161 · dispute
Raw extraction · bibliography line
URL https://arxiv.org/abs/2105.12400. Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training, 2018. URL https:// s3-us-west-2.amazonaws.com/openai-assets/research-covers/ language-unsupervised/language_understanding_paper.pdf. Javier Rando and Florian Tramèr. Universal jailbreak backdoors from poisoned human feedback, 2023. Jérémy Scheurer, Mikita Balesni, and Marius Hobbhahn. Technical report: Large language models can strategically deceive their users when put under pressure. arXiv preprint arXiv:2311.07590, 2023. John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. CoRR, abs/1707.06347, 2017. URL http://arxiv.org/abs/ 1707.06347. Roei Schuster, Congzheng Song, Eran Tromer, and Vitaly Shmatikov. You autocomplete me: Poisoning vulnerabilities in neural code completion. In 30th USENIX Security Symposium (USENIX Security 21), pp. 1559–1575, 2021. Murray Shanahan, Kyle McDonell, and Laria Reynolds. Role play with large language models. Nature, pp. 1–6, 2023. Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao, and Tom Goldstein. On t