Pith. sign in

Reference change · event page

Reference changes · DOI

I HATE YOU

Published notice on a work cited in the Pith corpus. Exact quotes below. No model judges whether any citation was load-bearing.

This page records that a citing paper's bibliography includes a work with a published notice. It is not a judgment on the citing paper.

Correction Crossref 1 open · 1 total · 0 disputed
DOI
10.1016/j.jneumeth.2013.09.010
Notice DOI
10.1016/j.jneumeth.2015.11.021
Event date
2015-12-15
Machine twin
JSON

01One-hop citing occurrences

Correction Open
Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

ref [1] · 2401.05566 · notice #6161 · dispute

Raw extraction · bibliography line

URL https://arxiv.org/abs/2105.12400. Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training, 2018. URL https:// s3-us-west-2.amazonaws.com/openai-assets/research-covers/ language-unsupervised/language_understanding_paper.pdf. Javier Rando and Florian Tramèr. Universal jailbreak backdoors from poisoned human feedback, 2023. Jérémy Scheurer, Mikita Balesni, and Marius Hobbhahn. Technical report: Large language models can strategically deceive their users when put under pressure. arXiv preprint arXiv:2311.07590, 2023. John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms. CoRR, abs/1707.06347, 2017. URL http://arxiv.org/abs/ 1707.06347. Roei Schuster, Congzheng Song, Eran Tromer, and Vitaly Shmatikov. You autocomplete me: Poisoning vulnerabilities in neural code completion. In 30th USENIX Security Symposium (USENIX Security 21), pp. 1559–1575, 2021. Murray Shanahan, Kyle McDonell, and Laria Reynolds. Role play with large language models. Nature, pp. 1–6, 2023. Manli Shu, Jiongxiao Wang, Chen Zhu, Jonas Geiping, Chaowei Xiao, and Tom Goldstein. On t

Status lifecycle on notices: open → disputed → (response attached on the notice page). Repaired counts matter as much as open counts. There is no “safe,” “invalid,” or “resolved” badge.