Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T05:42:43.602065Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 5 inbound Pith citation observations for arXiv:2502.08301.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T05:42:43.602065Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:09:27.931667Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T05:07:38.858520Z
49 of 49 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 30a7df6c-b566-479a-8c92-ce6dcaa8b469 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks deception attacks,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f53513aa-31b3-4962-8c6c-b5e4dac5edd3 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks AI Safety in Generative AI Large Language Models: A Survey
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8e16009-3e6d-4d5d-bb35-153f11c90f23 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks (a) GPT-4o, (b) GPT-4o mini, (c) Gemini 1.5 Pro, (d) Gemini 1.5 Flash
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation abee7819-170c-46ff-ab91-a9799ca0f8ab · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks AI Alignment: A Comprehensive Survey
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2426deb3-3272-4631-bb27-b051e40c7d9b · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Fine-Tuning Language Models from Human Preferences
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a72a9211-28f1-47c6-a8fa-832aa034fc41 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b03ff4b1-7490-4fce-a9b0-65d5f4f1bf7a · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Direct Preference Optimization: Your Language Model is Secretly a Reward Model
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23fc67e0-3355-4317-b1f5-14535894b972 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33be2679-36b9-4821-ad97-92404033e70c · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Jailbroken: How Does LLM Safety Training Fail?
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce9e6ed1-a7b4-4f5d-ba97-53a5f8b8b45e · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d724c645-9cf2-4ecc-9449-c06a54b05d5c · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69cf55cf-428e-4b79-8d6d-91b5bd5ca8d8 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks The Ethics of Advanced AI Assistants
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12750872-ac46-4bd2-bb57-f91686ee783e · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Mapping the Ethics of Generative AI: A Comprehensive Scoping Review
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e0b1aa67-3e4e-4af0-b51d-799c32909168 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks The Alignment Problem from a Deep Learning Perspective
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55e42007-e175-49a0-a43f-2f0d70b6c994 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks S., Goldstein, S., O’Gara, A., Chen, M
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 98e5e8ab-fb70-458b-965b-0859d744ac80 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b307d4f7-9636-4544-bb80-90daabd2a075 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12bd7929-6ffb-441e-90c7-32071a6903a5 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Scheming AIs: Will AIs fake alignment during training in order to get power?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e83aa47-f615-4f15-9320-163967861c45 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks X-Risk Analysis for AI Research
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e1b6a93-7996-4fbb-af4b-632e13c3489a · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Deception Abilities Emerged in Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ce3383a2-b24b-4810-8809-72e57343aa76 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Alignment faking in large language models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8274109e-2066-46e9-9ecb-19047f44c99c · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e04a58fa-88ce-480a-abc3-efc2237f97a4 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 05f058c2-49ad-45d3-a520-81295286a87f · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ba51696-cc2b-45bd-98ac-6316dedf5d0e · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71e8a86e-b8be-40e0-a00b-60052ead95d4 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 040b0cef-6164-43aa-bc92-5ac51ea9a613 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1feafa60-095e-471d-bc14-54ee3af548ab · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks GPT-4o System Card
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc8bd510-58b2-4c83-8b13-2ebba885cc58 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Gemini: A Family of Highly Capable Multimodal Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fefc191f-321b-4aed-aa94-e47869e802a3 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1c608ac-143c-4e23-8001-10d4687a0371 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Mitigating the Alignment Tax of RLHF
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bc28185-bb8e-4a92-aed4-cb1fe526fdfb · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Unresolved cited work
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37e0162d-52a1-4d5e-a91a-e77dcc0125bf · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fae777b6-1fbf-4673-bc0f-c4676c159bbe · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69100820-9d57-4bab-ae0e-9982ee853213 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Large Language Models as Misleading Assistants in Conversation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5891618-7dd3-4bf0-b0fb-231de967a490 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks OpenAI o1 System Card
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4c3cf1d-52b0-4968-83b0-138bc8c69b4e · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks The Llama 3 Herd of Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7f88c17-f783-49cc-b2d8-562a0a7e9952 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4d64535-26d5-4d38-9ae9-8e6cccb923ed · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Claude 3 model card
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c803a1f3-b23a-424c-9c33-cca2ba79e142 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Do Large Language Models Latently Perform Multi-Hop Reasoning?
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c87d9c3e-ba6d-4aff-aabc-d4b7d2daa036 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Human-level play in the game of Diplomacy by combining language models with strategic reasoning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1ebd489e-c5c4-468f-95ed-02ced5ed45bd · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks An Assessment of Model-On-Model Deception
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 663d5dc7-ed42-4e2f-876e-021342c147b7 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9848e39-91b4-4261-8ac9-02557fdf9b21 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b9394cf-220d-493e-9878-6061ce5c43ad · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Fine-tuning can cripple your foundation model; preserving features may be the solution
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cc15828-65e6-45bb-9b97-2908bc6a33fc · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Tell me about yourself: LLMs are aware of their learned behaviors
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d6f8215-9986-4c9a-b101-f4fbde9565f9 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks A Looming Replication Crisis in Evaluating Behavior in Language Models? Evidence and Solutions
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b0f1d33d-59c3-4970-b818-23f5f4981dfb · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 09fba7d2-a817-44e2-a1e3-4cd9dddc4705 · outbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Italy” , “Queen Elizabeth II
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c84cbc0b-eb40-4d79-a65d-248609831681 · inbound
Model Organisms for Emergent Misalignment Compromising Honesty and Harmlessness in Language Models via Deception Attacks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4546a6a0-f0bc-4891-8038-61c728b71d2a · inbound
Convergent Linear Representations of Emergent Misalignment Compromising Honesty and Harmlessness in Language Models via Deception Attacks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68388b80-fde5-4ba3-8acb-5316fda5363f · inbound
Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment Compromising Honesty and Harmlessness in Language Models via Deception Attacks
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5ae6605-c1da-416d-8be9-7433cd24a848 · inbound
Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating Compromising Honesty and Harmlessness in Language Models via Deception Attacks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c51d0efd-9cd9-4a63-b8ae-1f20f9a247f3 · inbound
The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment Compromising Honesty and Harmlessness in Language Models via Deception Attacks
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.