Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T14:53:38.598955Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2607.18639.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T14:53:38.598955Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 444dbfca-f879-4ed4-a121-ca6b4395482e · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Refusal in Language Models Is Mediated by a Single Direction
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb80ffa3-8839-4e76-894b-dd020ab469a8 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd17fe38-7a04-4712-9771-1cf485a19edb · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Constitutional AI: Harmlessness from AI Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80a00ca4-c2b6-487c-be80-0485d5453710 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Enhancing model safety through pre- training data filtering.Anthropic Alignment Science Blog, 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1244f305-bf11-4938-96b2-f4dc70955c44 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09c6a1cd-f04b-423a-921c-7674476e87e5 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Deepseek-v4: Towards highly efficient million-token context intelligence
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68dbdcca-fd6d-4d4e-b49a-fb0196182764 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Enhancing Chat Language Models by Scaling High-quality Instructional Conversations
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9df7b92-6c97-427e-8d6d-04a38ef312cd · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Simplicity prevails: Rethinking negative preference optimization for LLM unlearning,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c24c9987-d7be-40cb-9926-9c66b537912b · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Knowledge Unlearning for Mitigating Privacy Risks in Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec717d17-c815-4534-9e02-8c88cda9c970 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Pretraining language models with human preferences
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0b15acf-3aab-4fff-a4fe-63fb8b27140c · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs When Bad Data Leads to Good Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 584cac2e-0329-4c0b-8cc7-002e18f935de · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a8e1ad1-08c7-4e3e-bc27-0c003420b32e · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, and toxicity, 2024
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e448b515-4511-4777-bb21-809a7b8125b4 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Natural emergent misalignment from reward hacking in production RL, 2025
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e02ffd9-b228-4c99-b640-ac15bb012e33 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Safety pretraining: Toward the next generation of safe AI, 2025
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49c94a80-fe4f-4cee-af5e-f75bb70bbad1 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Mouton, Caleb Lucas, and Ella Guest
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3628455-8d5c-4e4c-9b7e-5a9c7b204b64 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Interpreting GPT: the logit lens.LessWrong, 2020
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8845198-a3b4-43f7-95b2-86dbada39a28 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Real-time detection of hallucinated entities in long-form generation, 2025
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fb618c0-7040-458e-8428-80261a9a637b · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Deep ignorance: Filtering pretraining data builds tamper-resistant safeguards into open-weight LLMs
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f0dba82-0e16-40ae-8abc-47ba4135b2c8 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Training language models to follow instructions with human feedback
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c57ee244-f289-4452-b72c-afbc95ba04e7 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Cambridge University Press, 2nd edition, 2009
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c1ff182-4433-4943-b305-afe7a589cf53 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2faad02f-4cb5-42a8-98fc-905c9c277096 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Shaping capabilities with token-level data filtering, 2026
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d362c5d2-bd84-473f-b4c7-ccc9c490fedd · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c386292e-083b-429b-92eb-6f107a78bb9a · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Conditionalization confounds inoculation prompting re- sults.LessWrong, 2026
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 050a5240-0bdd-45b7-998a-1447bef44979 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs XSTest: A test suite for identifying exaggerated safety behaviours in large language models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 989254db-67d1-4205-bb50-07a495d6c18e · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Believe it or not: How deeply do LLMs believe implanted facts?, 2025
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63132adf-9a33-44f3-b25a-dd1434dbb749 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db844b0a-5742-4963-875d-c57907349149 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e336ee7-f367-44f9-b29f-4aeba1f5364f · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3d3fa3a-80ba-4c38-999c-129bbe549623 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Inoculation prompting: Instructing LLMs to misbehave at train-time improves test-time alignment, 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fd31803-2c7a-474f-995a-7a50544ca50d · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Inoculation prompting: Eliciting traits from LLMs during training can suppress them at test-time,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c6167c6-8280-4ade-8adf-9ea50608b08a · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Backtracking Improves Generation Safety
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2bc8321-2376-417b-962f-2cf4a28326a9 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Negative preference optimization: From catastrophic collapse to effective unlearning,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01bbc548-db08-44e9-adcc-13cd1ed490d2 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Negative Preference Optimization: From Catastrophic Collapse to Effective Unlearning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e353af1-7d29-4645-9bee-5683ee82a53d · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Pretraining Language Models with Human Preferences
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2973579c-f06a-452d-b117-139cdebdeb78 · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Unresolved cited work
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3252bb0-ab43-4ed0-b646-43438f7c088b · outbound
Mark, Don't Erase: Token Inoculation for Dual-Use Knowledge in LLMs Unresolved cited work
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.