Pith. sign in

Paper Citation Record · LEDGER

Compromising Honesty and Harmlessness in Language Models via Deception Attacks

As of 23 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 5 inbound Pith citation observations for arXiv:2502.08301.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.08301 v2

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T05:42:43.602065Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:09:27.931667Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:07:38.858520Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact2
  • verified fuzzy8
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 30a7df6c-b566-479a-8c92-ce6dcaa8b469 · outbound

This paper cites deception attacks,.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks deception attacks,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:42:44.684083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T05:42:43.221711Z digest=sha256:d14db7c9ab01a822d4e0e7fe48e91c223f63d5b75ede4fd0e25912a6a8a0d0e8

Observation f53513aa-31b3-4962-8c6c-b5e4dac5edd3 · outbound

This paper cites AI Safety in Generative AI Large Language Models: A Survey.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks AI Safety in Generative AI Large Language Models: A Survey

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.350763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.350763Z digest=sha256:3c5b6faa9fc18aea5e1b4b63f1f6ef5c912494078b03c0eeac47b5736ae83b80

Observation b8e16009-3e6d-4d5d-bb35-153f11c90f23 · outbound

This paper cites (a) GPT-4o, (b) GPT-4o mini, (c) Gemini 1.5 Pro, (d) Gemini 1.5 Flash.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks (a) GPT-4o, (b) GPT-4o mini, (c) Gemini 1.5 Pro, (d) Gemini 1.5 Flash

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:42:44.699526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T05:42:43.213880Z digest=sha256:78815f68e5cd58800f5a28b49b0e264fa9ac7f037d03a61510ec443cca612bab

Observation abee7819-170c-46ff-ab91-a9799ca0f8ab · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks AI Alignment: A Comprehensive Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.343857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.343857Z digest=sha256:ec2e7faaa6005c3b585184011546fa5f2910c08d7150ba4cb96ff8aff6357ecb

Observation 2426deb3-3272-4631-bb27-b051e40c7d9b · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Fine-Tuning Language Models from Human Preferences

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.356519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.356519Z digest=sha256:fb65c898ba5a87c1dd5092e30b9e7b2580d9a85d97cbe3b1ccecdd4a0c1b88d2

Observation a72a9211-28f1-47c6-a8fa-832aa034fc41 · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.362622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.362622Z digest=sha256:4af06fd74b8c0fb334c2314b9f8d2e391ecea55d8f3873b70505389a421bae77

Observation b03ff4b1-7490-4fce-a9b0-65d5f4f1bf7a · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.381394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.381394Z digest=sha256:2194b549a9582940230ba59ffb89af6f9654cacb56ccd0dceaaabccc7dc36bf1

Observation 23fc67e0-3355-4317-b1f5-14535894b972 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.386584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.386584Z digest=sha256:4a22d8643bea685b5715e04f8634355505e0a4c08ecc68adfbf647c58b1ad3c2

Observation 33be2679-36b9-4821-ad97-92404033e70c · outbound

This paper cites Jailbroken: How Does LLM Safety Training Fail?.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Jailbroken: How Does LLM Safety Training Fail?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.392787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.392787Z digest=sha256:5d95b060588d2b9707b7796f9ce22e43ef2d12fe214cc2b7a62091ade760903b

Observation ce9e6ed1-a7b4-4f5d-ba97-53a5f8b8b45e · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.400191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.400191Z digest=sha256:09524ce91afecba4af9262ec6dc91b0c4a5887be9a83bc48fcddfaf3a26a6ae2

Observation d724c645-9cf2-4ecc-9449-c06a54b05d5c · outbound

This paper cites an unresolved cited work.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.405483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.405483Z digest=sha256:2f82479fe4b6c9bc9ec2b313fb63f5144423f941f8872d6857c7f2b1470ec9c2

Observation 69cf55cf-428e-4b79-8d6d-91b5bd5ca8d8 · outbound

This paper cites The Ethics of Advanced AI Assistants.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks The Ethics of Advanced AI Assistants

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.410849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.410849Z digest=sha256:a2ee1b12a17c1aeddb721f24e716fdff9448781140ddd2fa345c9e3b9c3a7e0b

Observation 12750872-ac46-4bd2-bb57-f91686ee783e · outbound

This paper cites Mapping the Ethics of Generative AI: A Comprehensive Scoping Review.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Mapping the Ethics of Generative AI: A Comprehensive Scoping Review

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:42:44.667049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T05:42:43.416213Z digest=sha256:087774038a1a6aa15d6c673b208992593c96d66945d7785141fd2f5954533640

Observation e0b1aa67-3e4e-4af0-b51d-799c32909168 · outbound

This paper cites The Alignment Problem from a Deep Learning Perspective.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks The Alignment Problem from a Deep Learning Perspective

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.420807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.420807Z digest=sha256:6749c35816ccbf551e7af249eb7bca43a9ed4ac62a8cb84ee79cf90f9ee8b8db

Observation 55e42007-e175-49a0-a43f-2f0d70b6c994 · outbound

This paper cites S., Goldstein, S., O’Gara, A., Chen, M.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks S., Goldstein, S., O’Gara, A., Chen, M

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:42:44.649151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T05:42:43.425803Z digest=sha256:e19b4c7812490d3b5e68de084f2093589d15308271d93c80e1b0c409317f7644

Observation 98e5e8ab-fb70-458b-965b-0859d744ac80 · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.431086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.431086Z digest=sha256:489022fb9c65a32e411f08f78e5ed7b5cd9c58f234fc840f1d26fa6715294b5f

Observation b307d4f7-9636-4544-bb80-90daabd2a075 · outbound

This paper cites Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.435854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.435854Z digest=sha256:502cb33a280bad1e3c8db96b7bf23223ee138dc715c7f14ece35b9f80614ff0e

Observation 12bd7929-6ffb-441e-90c7-32071a6903a5 · outbound

This paper cites Scheming AIs: Will AIs fake alignment during training in order to get power?.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Scheming AIs: Will AIs fake alignment during training in order to get power?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.440989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.440989Z digest=sha256:366415aa24bbb8a9bab2ab6e763c269983ca81db051f4649d2880106c25b287c

Observation 7e83aa47-f615-4f15-9320-163967861c45 · outbound

This paper cites X-Risk Analysis for AI Research.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks X-Risk Analysis for AI Research

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.445722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.445722Z digest=sha256:ecaa6048b6c4d8e8a4c401021b3f98589770a6114a68242533654823120037f7

Observation 7e1b6a93-7996-4fbb-af4b-632e13c3489a · outbound

This paper cites Deception Abilities Emerged in Large Language Models.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Deception Abilities Emerged in Large Language Models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:42:44.630371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T05:42:43.450559Z digest=sha256:8a6c0d339994f85b6171ba88c77bc4b319dd5c326c5677cd49b5cb4f8c0a8387

Observation ce3383a2-b24b-4810-8809-72e57343aa76 · outbound

This paper cites Alignment faking in large language models.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Alignment faking in large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.455240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.455240Z digest=sha256:521a4cc9fec7a9f22874eead11d069bfdd1ed2445ee92f6dd6146d5a3633b14a

Observation 8274109e-2066-46e9-9ecb-19047f44c99c · outbound

This paper cites Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Me, Myself, and AI: The Situational Awareness Dataset (SAD) for LLMs

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.460514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.460514Z digest=sha256:711fbe807fb64a35a4ea98db37d5baeab057f606935ddb84c38575267073d93c

Observation e04a58fa-88ce-480a-abc3-efc2237f97a4 · outbound

This paper cites an unresolved cited work.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-08T05:42:44.612703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T05:42:43.465366Z digest=sha256:4b626fd09d622b924c16945e72d2bf0c31f4cb426d96e5d0f97a2f3f27076e8a

Observation 05f058c2-49ad-45d3-a520-81295286a87f · outbound

This paper cites Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Harmful Fine-tuning Attacks and Defenses for Large Language Models: A Survey

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.470194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.470194Z digest=sha256:411ef0615cfbcdfe7f42170c1f7e603a32f61514e18b81ca634a8053a661707b

Observation 8ba51696-cc2b-45bd-98ac-6316dedf5d0e · outbound

This paper cites Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Covert Malicious Finetuning: Challenges in Safeguarding LLM Adaptation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.475213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.475213Z digest=sha256:e598a62e1e173bced2802b92f25884b80b6a6438d66582af4e13ed24405b60b6

Observation 71e8a86e-b8be-40e0-a00b-60052ead95d4 · outbound

This paper cites Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.480254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.480254Z digest=sha256:23606118c7e3e1f1996c90c0f8f5442362d4b8edeb432e43399279f88b9e759c

Observation 040b0cef-6164-43aa-bc92-5ac51ea9a613 · outbound

This paper cites The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks The Ultimate Guide to Fine-Tuning LLMs from Basics to Breakthroughs: An Exhaustive Review of Technologies, Research, Best Practices, Applied Research Challenges and Opportunities

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.485261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.485261Z digest=sha256:356b2da2cb9eb4c3bf666c0d2045622768c3234df4828576bb6ae2684b8ef27c

Observation 1feafa60-095e-471d-bc14-54ee3af548ab · outbound

This paper cites GPT-4o System Card.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks GPT-4o System Card

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.490299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.490299Z digest=sha256:e6ec51a4f0b131fd3651598a6a4e13796be7d7b5976d8227f22ab4a1f40af366

Observation bc8bd510-58b2-4c83-8b13-2ebba885cc58 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Gemini: A Family of Highly Capable Multimodal Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.495381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.495381Z digest=sha256:794e8ba0dbf8e058176198d615502200111c0dc5f3c4d895e45b832cbe26e25b

Observation fefc191f-321b-4aed-aa94-e47869e802a3 · outbound

This paper cites An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.500457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.500457Z digest=sha256:f5587512d59506e025101822fd0d617ed5583d07939fa74f6e4a8119b9032bc1

Observation d1c608ac-143c-4e23-8001-10d4687a0371 · outbound

This paper cites Mitigating the Alignment Tax of RLHF.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Mitigating the Alignment Tax of RLHF

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.505543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.505543Z digest=sha256:1b88c46a18053cf2f50c1952211fc83992769060028e735a3dd76ec68d75d662

Observation 0bc28185-bb8e-4a92-aed4-cb1fe526fdfb · outbound

This paper cites an unresolved cited work.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.510437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.510437Z digest=sha256:2cdf879d8288899fe2bd41e337d32d7b6ff79210354640f9e3d9a47729d5d740

Observation 37e0162d-52a1-4d5e-a91a-e77dcc0125bf · outbound

This paper cites RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks RealToxicityPrompts: Evaluating Neural Toxic Degeneration in Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.515096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.515096Z digest=sha256:ea7a69a926c5f0587c91797033b5ec5e525d70306984404ed4af36b5b4fd27d4

Observation fae777b6-1fbf-4673-bc0f-c4676c159bbe · outbound

This paper cites ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech Detection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.520707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.520707Z digest=sha256:d6c43cb4be4830fa6d9b4c719df4a5b76bd1ca9ae255ffbc7a0543d5ecd3e358

Observation 69100820-9d57-4bab-ae0e-9982ee853213 · outbound

This paper cites Large Language Models as Misleading Assistants in Conversation.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Large Language Models as Misleading Assistants in Conversation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.525731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.525731Z digest=sha256:589c8e4db8d01e8e693757a38657834f89ce2c77386209ccc9d3f85f074598c3

Observation e5891618-7dd3-4bf0-b0fb-231de967a490 · outbound

This paper cites OpenAI o1 System Card.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks OpenAI o1 System Card

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.530889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.530889Z digest=sha256:640f810fb5d808611c8273d9d8e7b7bcf1e5e9281a1ce8b355eb6d255d477eec

Observation f4c3cf1d-52b0-4968-83b0-138bc8c69b4e · outbound

This paper cites The Llama 3 Herd of Models.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks The Llama 3 Herd of Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.536036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.536036Z digest=sha256:e718f63bab863e0c6227c717af03ebc68ac13d5f6509af80eb6427e68afbd769

Observation b7f88c17-f783-49cc-b2d8-562a0a7e9952 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.541142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.541142Z digest=sha256:9ee46863e2ffbea0ecedf789ce612b0d2beced5c4d21257780b984f94d3b11af

Observation d4d64535-26d5-4d38-9ae9-8e6cccb923ed · outbound

This paper cites Claude 3 model card.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Claude 3 model card

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:42:44.596621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T05:42:43.546966Z digest=sha256:ae2e9809bd0913772b92ab348ce8539d188a9b05e27a39c084ed7cf4dae43883

Observation c803a1f3-b23a-424c-9c33-cca2ba79e142 · outbound

This paper cites Do Large Language Models Latently Perform Multi-Hop Reasoning?.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Do Large Language Models Latently Perform Multi-Hop Reasoning?

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.551459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.551459Z digest=sha256:e045d130ca267b1ba0880b021de281538d37d3e6eb262b4b450cd56e48304070

Observation c87d9c3e-ba6d-4aff-aabc-d4b7d2daa036 · outbound

This paper cites Human-level play in the game of Diplomacy by combining language models with strategic reasoning.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Human-level play in the game of Diplomacy by combining language models with strategic reasoning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:42:44.581368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T05:42:43.556376Z digest=sha256:b9c24d47857dc017403613b78f241ca690f57557e5fe4c4fd9302f5de7968853

Observation 1ebd489e-c5c4-468f-95ed-02ced5ed45bd · outbound

This paper cites An Assessment of Model-On-Model Deception.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks An Assessment of Model-On-Model Deception

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-08T05:42:43.762712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T05:42:43.561395Z digest=sha256:05fcf5ea74ec33e633a07b2fc74b0f3d1f9bbafb63c8fbe5ab9f1ba8690a6f5d

Observation 663d5dc7-ed42-4e2f-876e-021342c147b7 · outbound

This paper cites Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Generative AI Misuse: A Taxonomy of Tactics and Insights from Real-World Data

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.566302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.566302Z digest=sha256:35654e0b25eb81ddc5f603bcdb054913b7b42f1eab2a7301b29b19063e7bf112

Observation b9848e39-91b4-4261-8ac9-02557fdf9b21 · outbound

This paper cites Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.571564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.571564Z digest=sha256:4b61fbd956b3b11b195f5a14da505a571892de7e5eb0364f5af220dd8c639a71

Observation 5b9394cf-220d-493e-9878-6061ce5c43ad · outbound

This paper cites Fine-tuning can cripple your foundation model; preserving features may be the solution.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Fine-tuning can cripple your foundation model; preserving features may be the solution

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.578787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.578787Z digest=sha256:debd8fd0dcec0af5948306971630358479f60f3082043855d26ef9c06103538a

Observation 1cc15828-65e6-45bb-9b97-2908bc6a33fc · outbound

This paper cites Tell me about yourself: LLMs are aware of their learned behaviors.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Tell me about yourself: LLMs are aware of their learned behaviors

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T05:42:43.585155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T05:42:43.585155Z digest=sha256:11da0dc5dd4fc844591a00abaf45595c7a399094306db35e9ce72d2056f06b76

Observation 3d6f8215-9986-4c9a-b101-f4fbde9565f9 · outbound

This paper cites A Looming Replication Crisis in Evaluating Behavior in Language Models? Evidence and Solutions.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks A Looming Replication Crisis in Evaluating Behavior in Language Models? Evidence and Solutions

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-08T05:42:43.655382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T05:42:43.591256Z digest=sha256:298e950497a9fae03b4c3ef54cf324745e034919930df0e498f098b7278a7b47

Observation b0f1d33d-59c3-4970-b818-23f5f4981dfb · outbound

This paper cites an unresolved cited work.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-08T05:42:44.564606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T05:42:43.596383Z digest=sha256:25045bf48dd9d9615c6d05243c52a2505b549571d7a90e661f543bf091efce7e

Observation 09fba7d2-a817-44e2-a1e3-4cd9dddc4705 · outbound

This paper cites Italy” , “Queen Elizabeth II.

Compromising Honesty and Harmlessness in Language Models via Deception Attacks Italy” , “Queen Elizabeth II

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T05:42:44.546282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T05:42:43.602065Z digest=sha256:26d879fe6c720a8f5403180b6ddf27980ff17560faacdc0e0a1d45f91435c4a8

Pith citing papers

Observation c84cbc0b-eb40-4d79-a65d-248609831681 · inbound

Model Organisms for Emergent Misalignment cites this paper.

Model Organisms for Emergent Misalignment Compromising Honesty and Harmlessness in Language Models via Deception Attacks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:07:28.223256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:07:28.223256Z digest=sha256:5003c5c0f1fc109754c85b23801e8463d4e933a0d48af0c02ded349342c65c2f

Observation 4546a6a0-f0bc-4891-8038-61c728b71d2a · inbound

Convergent Linear Representations of Emergent Misalignment cites this paper.

Convergent Linear Representations of Emergent Misalignment Compromising Honesty and Harmlessness in Language Models via Deception Attacks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T04:09:27.931667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:09:27.931667Z digest=sha256:0efca2f572ba63860fc8aef055088c16b5bfc65d881dde96af2b666c567759f1

Observation 68388b80-fde5-4ba3-8acb-5316fda5363f · inbound

Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment cites this paper.

Decomposing Behavioral Phase Transitions in LLMs: Order Parameters for Emergent Misalignment Compromising Honesty and Harmlessness in Language Models via Deception Attacks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:42.234545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:42.234545Z digest=sha256:08e5bc3f1005696258f3aa70ac7c44ccdeec71e3253ec784e8268dfb4369ff92

Observation d5ae6605-c1da-416d-8be9-7433cd24a848 · inbound

Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating cites this paper.

Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating Compromising Honesty and Harmlessness in Language Models via Deception Attacks

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:47:30.004832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-27T17:03:33.199645Z digest=sha256:49cbc8dba66f458ec7388fe3a55c6eca81ff9635d427c7bf9c6e99c8daa0bee6

Observation c51d0efd-9cd9-4a63-b8ae-1f20f9a247f3 · inbound

The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment cites this paper.

The Arbiter Agent: Continually Monitoring Multi-Agent Conversations to Detect Emergent Misalignment Compromising Honesty and Harmlessness in Language Models via Deception Attacks

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:07:38.860058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T13:26:14.457195Z digest=sha256:78f634d3885bd6ee18f7ed004e14a38095da097745309752958c155d20e29b5e