Pith. sign in

Paper Citation Record · LEDGER

Mitigating Deceptive Alignment via Self-Monitoring

As of 8 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 3 inbound Pith citation observations for arXiv:2505.18807.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18807 v1

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:30:10.795238Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:14:22.854339Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

80 of 80 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved63
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 83677165-82a4-4664-bd97-21887015e19b · outbound

This paper cites Introducing openai o1-preview.

Mitigating Deceptive Alignment via Self-Monitoring Introducing openai o1-preview

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:18.690887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:03.992286Z digest=sha256:3e1ccb481681f24b51e92cd33986c5018053b714dcdfe5a4487aeed230113749

Observation bc6cad5f-5a5d-4302-82d2-65ce6a296a0b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Mitigating Deceptive Alignment via Self-Monitoring DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.060422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.060422Z digest=sha256:fb0a1dd9a853bf77a35a7a15d416ce756fff159a8b1f00f79bb66139b3d532de

Observation 51cf49de-8586-462c-bddf-9c41241e9cc2 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Mitigating Deceptive Alignment via Self-Monitoring Chain-of-thought prompting elicits reasoning in large language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.114784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.114784Z digest=sha256:9a0d76513a59a0492e419a294be4ff500cef634ab2e530cd2a90762eabe4f336

Observation 09c55401-661d-49d0-848f-bfa1201258ca · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Mitigating Deceptive Alignment via Self-Monitoring AI Alignment: A Comprehensive Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.196418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.196418Z digest=sha256:b546ef0496d9b32f6f5772ff8e6324e7d1f603817b24708e48b0cd9f8acff3ec

Observation dc773ee9-4491-4355-afb5-143fe9e19624 · outbound

This paper cites Risks from Learned Optimization in Advanced Machine Learning Systems.

Mitigating Deceptive Alignment via Self-Monitoring Risks from Learned Optimization in Advanced Machine Learning Systems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.260470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.260470Z digest=sha256:fb1fd45dc57cf25fea800ad1220630703f6c6f295978ecd83871ef46bc672a18

Observation 80a65e4f-f2a7-4382-b739-6637b0c24c72 · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

Mitigating Deceptive Alignment via Self-Monitoring Frontier Models are Capable of In-context Scheming

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.319823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.319823Z digest=sha256:ac9b7e56595bb878a217c338986028a8d5529b9da575b13a44a5b1204095c7a6

Observation 0c2bc9d3-9b3a-4bd8-a2df-7f37c0a94918 · outbound

This paper cites Foundational Challenges in Assuring Alignment and Safety of Large Language Models.

Mitigating Deceptive Alignment via Self-Monitoring Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.428012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.428012Z digest=sha256:d4f3691d0381df78409bcf195e8a7f3503a7c8d7f96bf3c4e9ebd70a9be00955

Observation 904a83a5-b850-4e07-b3f5-645ad7b05482 · outbound

This paper cites Language Models Learn to Mislead Humans via RLHF.

Mitigating Deceptive Alignment via Self-Monitoring Language Models Learn to Mislead Humans via RLHF

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.463941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.463941Z digest=sha256:446b933e00c81b375960afe4aa2271133446a2bbc7db96475304e564498b5b27

Observation 3bda92d4-c08e-4a48-90b0-53f4e7113efb · outbound

This paper cites Managing extreme ai risks amid rapid progress.

Mitigating Deceptive Alignment via Self-Monitoring Managing extreme ai risks amid rapid progress

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.536207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.536207Z digest=sha256:4e4df20b400f027b2bd8b196e138d10d0867bc07c24214b5a74c1cf6982d298b

Observation d1c7d851-abb7-46c1-ad89-0635acd982e8 · outbound

This paper cites Privacy risks of general-purpose language models.

Mitigating Deceptive Alignment via Self-Monitoring Privacy risks of general-purpose language models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:18.514680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:04.735429Z digest=sha256:2dcdecf8b7b7a08e26e0af0decec7e0865f4bbd57b9d8cfa52088d1db5dc324b

Observation 098f5cbd-18e0-4689-9dbb-e7359454e794 · outbound

This paper cites Frontier AI systems have surpassed the self-replicating red line.

Mitigating Deceptive Alignment via Self-Monitoring Frontier AI systems have surpassed the self-replicating red line

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.787401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.787401Z digest=sha256:5283fe4fcdedda4ab67f4f7b4b9f998ef78e8468186577dd84fe7ab2130b0823

Observation f0b4bd89-2800-4f79-8679-5371b3c50fe4 · outbound

This paper cites Alignment faking in large language models.

Mitigating Deceptive Alignment via Self-Monitoring Alignment faking in large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.868190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.868190Z digest=sha256:4a4e9a5a6175c73bc1849d0864ce508334887ffdf0530e388dfeb5ec794f2296

Observation 2fa7d15b-f588-41bb-b358-f85bac93806f · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Mitigating Deceptive Alignment via Self-Monitoring Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:04.987930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:04.987930Z digest=sha256:90b3c090b896e6628aa352e94ef9dc0a5130747ca95f527d133083e5cc05f767

Observation 26c081fb-2fb4-4bc7-a133-8512d908309b · outbound

This paper cites Darkbench: Benchmarking dark patterns in large language models.

Mitigating Deceptive Alignment via Self-Monitoring Darkbench: Benchmarking dark patterns in large language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:18.182736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:05.047058Z digest=sha256:5d3d3c0496e364637efb1b8cfd51bbcd7905c69107deeeb797992619873216b7

Observation 40edb399-1f95-4346-86e6-090ebf81128c · outbound

This paper cites Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?.

Mitigating Deceptive Alignment via Self-Monitoring Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.120166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.120166Z digest=sha256:62a1de825dbcbcecba9816d043b53cd0be285b73bc71fba8d6708663affe9210

Observation 673ccd70-0688-48b4-b0f3-f48a23bd1dec · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

Mitigating Deceptive Alignment via Self-Monitoring Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.188094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.188094Z digest=sha256:dfb2569510a42b253a77580f166f9c93fa4a8b1134fa4e6ecd12da6226247eee

Observation 8d4e3e7a-8f66-4be0-b3fd-1477604c2cd2 · outbound

This paper cites International AI Safety Report.

Mitigating Deceptive Alignment via Self-Monitoring International AI Safety Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.234059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.234059Z digest=sha256:0f211600a2fdc0e7d59cc8c9317ebbd70772eb20f1052a6838586ccdf2bd7eb4

Observation 8ee0c132-066c-4f33-832b-40e96438b21d · outbound

This paper cites Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting.

Mitigating Deceptive Alignment via Self-Monitoring Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.260854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.260854Z digest=sha256:4e60eda8007260235397cdc378d53e6d4deb04714813857167ce8787a7d077bf

Observation d9afa334-7493-405b-b63e-a82a78f5d73c · outbound

This paper cites Policy Frameworks for Transparent Chain-of-Thought Reasoning in Large Language Models.

Mitigating Deceptive Alignment via Self-Monitoring Policy Frameworks for Transparent Chain-of-Thought Reasoning in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.298607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.298607Z digest=sha256:eda39ad444b38f5e35440f0c2ef433e2cc493a61c7d43800b54296001ce77f65

Observation b0e2b95f-069b-4d5f-b70c-6d85580017cf · outbound

This paper cites Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation.

Mitigating Deceptive Alignment via Self-Monitoring Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.384371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.384371Z digest=sha256:b28fdd63359fe80df16ed25a7364d207bea4ee98e8f5b2764eb556a2845bb34e

Observation bea15da7-e936-4cbd-84b6-7fe78ea1b9a9 · outbound

This paper cites From System 1 to System 2: A Survey of Reasoning Large Language Models.

Mitigating Deceptive Alignment via Self-Monitoring From System 1 to System 2: A Survey of Reasoning Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.477915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.477915Z digest=sha256:5bad57d2f593a06781a75232de02cfcf4f8e1c9a6fc600986790318e40e20d3f

Observation e12a21ad-c180-41f2-a85f-4a0d1874be96 · outbound

This paper cites Markov decision processes: discrete stochastic dynamic programming.

Mitigating Deceptive Alignment via Self-Monitoring Markov decision processes: discrete stochastic dynamic programming

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.595053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.595053Z digest=sha256:f51d1535e91076c217c982ffc8f58d4dd24aaaec41f54835b6a9e6bbfa19c914

Observation 8694ca25-0e87-4150-a6e4-bd758a5fc61b · outbound

This paper cites Reinforcement learning: An introduction, volume 1.

Mitigating Deceptive Alignment via Self-Monitoring Reinforcement learning: An introduction, volume 1

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.716513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.716513Z digest=sha256:d894fe44841a7c221c15fbbcf34fc4170573dbbfb9444bd8988b92417b2aad42

Observation 02639ee0-cd8c-40c4-8bb0-31684ae7ceef · outbound

This paper cites Training language models to follow instructions with human feedback.

Mitigating Deceptive Alignment via Self-Monitoring Training language models to follow instructions with human feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.874104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.874104Z digest=sha256:3287c653ed0c1c98e705439998a8832d9b3f70834c52c96e1db680fb70b2eba0

Observation 7e9cdb46-9f25-447e-ad2f-6c6f25721ab5 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Mitigating Deceptive Alignment via Self-Monitoring Direct preference optimization: Your language model is secretly a reward model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:05.978324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:05.978324Z digest=sha256:07ce14740616125890b7232d4ffbddf86588d793511250b050934a4b3a4dad53

Observation 1cb8ea0a-c44f-461f-a6cc-02ebeb23f7bc · outbound

This paper cites Handbook of constraint programming.

Mitigating Deceptive Alignment via Self-Monitoring Handbook of constraint programming

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:17.898223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:06.196688Z digest=sha256:d1a1d24a42928ed4a5f161c9bee06d7814180680bfaaad770792512b93b91972

Observation c4c02e54-93f3-4779-a680-44c8d90d28b3 · outbound

This paper cites Defining and characterizing reward gaming.

Mitigating Deceptive Alignment via Self-Monitoring Defining and characterizing reward gaming

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:06.349381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:06.349381Z digest=sha256:0617dfd81acfc53d7f40557836f1febf31b36d6cdbf5b73e5c6b3d7803a4888a

Observation bbc388a2-87a0-4fde-9edf-57ab6eb42b11 · outbound

This paper cites Cooperative inverse reinforcement learning.

Mitigating Deceptive Alignment via Self-Monitoring Cooperative inverse reinforcement learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:06.469580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:06.469580Z digest=sha256:8f6d648637d79bdf7fda60aeedf9d7c7d31d71c85f0b0ad8a18aef5e2c7b42eb

Observation d808a0c0-e5fe-4fb5-a4aa-514fa1f54ebb · outbound

This paper cites BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping.

Mitigating Deceptive Alignment via Self-Monitoring BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:06.590347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:06.590347Z digest=sha256:b89ef43184400da225b3652bffe585ea693a5632daa6440f713fd645bf00b48e

Observation 86531776-8121-4847-a77f-a686ddc8d381 · outbound

This paper cites Defin- ing deception in decision making.

Mitigating Deceptive Alignment via Self-Monitoring Defin- ing deception in decision making

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:17.681852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:06.693263Z digest=sha256:6064c63dbce3e13c574a99b6f1f996a43fd87a7b43c1a584acc2754e1c97dff6

Observation ad4c6467-25ae-475c-adb0-b0f148400b61 · outbound

This paper cites Machine behaviour.

Mitigating Deceptive Alignment via Self-Monitoring Machine behaviour

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:17.491130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:06.789717Z digest=sha256:9d7bbd2ef81dc57bd694231cf80080a31d6491e50870389fffd9b74fbbb5ffc0

Observation 6c98cc28-959e-44be-825c-ae5533d4ef4f · outbound

This paper cites The off-switch game.

Mitigating Deceptive Alignment via Self-Monitoring The off-switch game

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:06.923371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:06.923371Z digest=sha256:662823eaa69ad9f3f30dfe3514877699d69322ac24660f0b9ebc3bc407715606

Observation 9bf5cbbd-9f3b-430e-9c08-8eb0167ba75e · outbound

This paper cites Comparison of the predicted and observed secondary structure of t4 phage lysozyme.

Mitigating Deceptive Alignment via Self-Monitoring Comparison of the predicted and observed secondary structure of t4 phage lysozyme

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:07.044879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:07.044879Z digest=sha256:5d66daa40ad0e8b4b915b4c42155a54870e57fb52317c6f28353aec336a3a10d

Observation 8ed211ab-159a-4316-92c1-6b1f540e8132 · outbound

This paper cites Ai deception: A survey of examples, risks, and potential solutions.

Mitigating Deceptive Alignment via Self-Monitoring Ai deception: A survey of examples, risks, and potential solutions

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:17.165031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:07.150812Z digest=sha256:b97f4443bb1935901415d725da7a8d8fd68e154ff446ce95b11dd01d19f8c4f8

Observation d31cd9fc-fdc5-4c26-9a4d-7a62511ab327 · outbound

This paper cites Discovering Language Model Behaviors with Model-Written Evaluations.

Mitigating Deceptive Alignment via Self-Monitoring Discovering Language Model Behaviors with Model-Written Evaluations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:07.287960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:07.287960Z digest=sha256:4588cb58365de4ff1768aa9c0439e88d9f32409c556372b51b85a7079a702a74

Observation 685d6827-be15-4b6c-84e1-37ea8a7cc513 · outbound

This paper cites Deception abilities emerged in large language models.

Mitigating Deceptive Alignment via Self-Monitoring Deception abilities emerged in large language models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:16.932491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:07.410071Z digest=sha256:9a44ec9534c6f5a4edc11e2784d1e478ae194e1e5e784ff5cd0720ba6174c291

Observation 8b5604c4-b415-429d-9b2f-96d25e8a2650 · outbound

This paper cites Opendeception: Benchmarking and investigating ai deceptive behaviors via open-ended interaction simulation.

Mitigating Deceptive Alignment via Self-Monitoring Opendeception: Benchmarking and investigating ai deceptive behaviors via open-ended interaction simulation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:07.532498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:07.532498Z digest=sha256:d6bfc18794050f63a2a5356a8f632fb4f85d12166214a3931773caa32725da44

Observation 7c515169-f120-4524-803e-06527d6e6a72 · outbound

This paper cites The mask benchmark: Disentangling honesty from accuracy in ai systems.

Mitigating Deceptive Alignment via Self-Monitoring The mask benchmark: Disentangling honesty from accuracy in ai systems

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:07.579875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:07.579875Z digest=sha256:2ddf4973979bbdfa215425bc819f9202caeea146f8c6971d084818ca8cac50c7

Observation 7d88fc10-a4fe-4c1c-ba7e-dbd916a8a361 · outbound

This paper cites AI Sandbagging: Language Models can Strategically Underperform on Evaluations.

Mitigating Deceptive Alignment via Self-Monitoring AI Sandbagging: Language Models can Strategically Underperform on Evaluations

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:07.610336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:07.610336Z digest=sha256:24a89bdf4e281650644a333ba53476d76da582e17f38a0fa0b3baf05fe514a71

Observation dd02a31d-0f35-4579-9b8f-5b8b097f32bb · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Mitigating Deceptive Alignment via Self-Monitoring Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:07.652785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:07.652785Z digest=sha256:984cb55c76a222018af804cea525edb568bce9f0980fbbf34ddb9e0008708660

Observation 9a4bb438-6855-4973-90bc-3a8d6ef00006 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:16.411365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:07.736203Z digest=sha256:c796ac4988585e758025ca76bb2d0bb8aa4828cb7e00af9418b3964369b26ecb

Observation ac63f8ae-d62e-4aaf-a772-6abce6e67dea · outbound

This paper cites Qwen2 Technical Report.

Mitigating Deceptive Alignment via Self-Monitoring Qwen2 Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:07.823125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:07.823125Z digest=sha256:3f3121fa0fdf575bf5c9426c952de59ddf69e000612bf5a3f2ceb88c215a2866

Observation 2e534c1d-7c36-4336-8cc4-2b002bfe056f · outbound

This paper cites The Llama 3 Herd of Models.

Mitigating Deceptive Alignment via Self-Monitoring The Llama 3 Herd of Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:07.921367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:07.921367Z digest=sha256:a06e5ad7d3ca3fc76dfa0ceec5827876194eb99d505405083fddcbcf9ca1bce3

Observation c22dd63e-af11-4ad2-b88f-6ea63dee549f · outbound

This paper cites Claude 3.

Mitigating Deceptive Alignment via Self-Monitoring Claude 3

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:15.968293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:07.983269Z digest=sha256:5ad1b0a235c581e2b1fdcc4cad10ef97327c6ee657295096d13366ca1be7346e

Observation 650038e4-2c1c-4935-9d89-d919d8f4e149 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Mitigating Deceptive Alignment via Self-Monitoring Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.031978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.031978Z digest=sha256:5c71bef7300cfc839d0bf6f568a6dd446fa622c352b319be5f83303b3534ce92

Observation 1bb9ac23-60fb-4b1e-91d0-09e9e5697d2a · outbound

This paper cites Learning to reason with llms.

Mitigating Deceptive Alignment via Self-Monitoring Learning to reason with llms

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:15.760571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:08.070454Z digest=sha256:d5ed36a57fe3a902219fa8b43852a88df9f9ed6e2de5ff13cd3a65adbeac3296

Observation 26d25d22-bf8b-4db2-b18a-26c2de8171b7 · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

Mitigating Deceptive Alignment via Self-Monitoring A StrongREJECT for Empty Jailbreaks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.125239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.125239Z digest=sha256:a383c796a2dee637901ba4847a0ed229f3d4aa67a1fbc6a601ce9a4850c96ff9

Observation a0fed44a-7184-4cd2-8445-b67569db6ca0 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

Mitigating Deceptive Alignment via Self-Monitoring Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.164673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.164673Z digest=sha256:1fe36ed4255bf03a8530f76823366c3c360d3f08d6183b28bb2cf8b78dc3ae59

Observation 7e02df15-0093-4537-846e-9a9121732df3 · outbound

This paper cites How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms.

Mitigating Deceptive Alignment via Self-Monitoring How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.215835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.215835Z digest=sha256:c43f5bcec99660da8a71a5671f98b0d74b31672c66e6cc2c0ddd9b0d3da51e35

Observation 48f8a6d5-94b4-4ccf-a350-a4e7af0ace1c · outbound

This paper cites Towards evaluating the robustness of neural networks.

Mitigating Deceptive Alignment via Self-Monitoring Towards evaluating the robustness of neural networks

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.261749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.261749Z digest=sha256:aa5f2c570268eb97b887e8cd3ec321ec6401e2b1025ad001c691b3ee1db8d795

Observation c2cab175-f615-481b-ba71-1124f375c68e · outbound

This paper cites Safety Alignment Should Be Made More Than Just a Few Tokens Deep.

Mitigating Deceptive Alignment via Self-Monitoring Safety Alignment Should Be Made More Than Just a Few Tokens Deep

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.313927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.313927Z digest=sha256:31e0874a5afdc764b47e45c1c0fecb17ed116a98965037f6d11ef97fb651503d

Observation fda382d3-d322-45a7-b5c5-00d39135fdbc · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

Mitigating Deceptive Alignment via Self-Monitoring JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.351856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.351856Z digest=sha256:df649b5abfab04c3eaed6cea960e34a78907e4b6113e63b034effb0358ca846b

Observation 533de5ff-de66-4a20-bafc-ba2d1bba421f · outbound

This paper cites Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations.

Mitigating Deceptive Alignment via Self-Monitoring Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.405155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.405155Z digest=sha256:5f508d050cd6a8f62937d0dc8f24055a1c988a215a9c0a28611ad440d2935532

Observation 6ea2b524-9826-4f34-b090-b916f1f22b6b · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset.

Mitigating Deceptive Alignment via Self-Monitoring Beavertails: Towards improved safety alignment of llm via a human-preference dataset

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:18.366017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:08.488491Z digest=sha256:a8d90f1498002ba24cce6123ed6065c603d443e5a513adc9e9416bd97531597a

Observation d66ec9ff-2935-4e0f-b5a2-1b45b0df9e93 · outbound

This paper cites SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities.

Mitigating Deceptive Alignment via Self-Monitoring SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.592130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.592130Z digest=sha256:419002f3740279db120bc931fd72b95d62c599e50288f50551e374287f9b822f

Observation 99bc8ee5-75f0-47a8-901b-81b7d3569418 · outbound

This paper cites Star-1: Safer alignment of reasoning llms with 1k data.

Mitigating Deceptive Alignment via Self-Monitoring Star-1: Safer alignment of reasoning llms with 1k data

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.656029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.656029Z digest=sha256:b23ad75459509635b3ced7a91bc0e8e7092c896b177d367c17cb526c14aa88f1

Observation 46f8839c-2343-4faa-a81c-bff59ebf220b · outbound

This paper cites Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment.

Mitigating Deceptive Alignment via Self-Monitoring Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.770366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.770366Z digest=sha256:dbe4f11cef58e1da8fdf31535adde2abcc1d9e89cc6b647bcafa4162c4679c15

Observation 8cc94936-efae-4f41-aa47-78820d0edf70 · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback, 2023.

Mitigating Deceptive Alignment via Self-Monitoring Ultrafeedback: Boosting language models with high-quality feedback, 2023

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:15.472718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:08.879503Z digest=sha256:a799bcd950815faabe2a05347ab33839b31f87b073e26928e210623bc121e633

Observation 3497f84f-f0a4-42cb-9109-81ca5c60952d · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:08.994148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:08.994148Z digest=sha256:e0c179c8d30bcd58d861839beeeb96b5af91ab8dfce5a9dd5d06fff608d97c6e

Observation c3430b97-bf70-4334-aa64-08557737abde · outbound

This paper cites Constrained Markov decision processes.

Mitigating Deceptive Alignment via Self-Monitoring Constrained Markov decision processes

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:09.107586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:09.107586Z digest=sha256:c059d7ab678e3fffe00410c5911ba91ca64d47f2ce5467f968fbd4fb2d2fd923

Observation 8e7d481e-0a70-48e3-945c-2d83ececa2f4 · outbound

This paper cites mesa-objective.

Mitigating Deceptive Alignment via Self-Monitoring mesa-objective

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:15.114338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:09.224633Z digest=sha256:6c07058713905cca2daec1ca8f967b9957e4ee6363ce68da87fe88d21a711778

Observation 7500718a-9539-4b1a-a869-0c1a98a9e116 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:15.002661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:09.307384Z digest=sha256:a6d7ed56597fce721dda78dc2d886dcf5890a2c0e1757e3b90360ab823348484

Observation ba78bc72-891b-45c9-b232-1101adb05da0 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:14.833649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:09.414613Z digest=sha256:d864468dae2d53f1911a7560e217f8a4ed0a96455da898b47153b243bfdd19d0

Observation f69279a8-86bb-4350-a036-2c587f32de1c · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:14.666887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:09.530070Z digest=sha256:6827044abfbaa910328fa3b34af3d24bd8b88e54b2c254845cf18f21a12b0920

Observation f4bc6367-5801-4399-b33f-207e1df98019 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:14.481276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:09.647676Z digest=sha256:adc7fea1ddacc74bfd88a1f970a5fe9ef8a8cf1815252afc7af1f022338448f1

Observation 8d054452-f45a-408d-b5a9-3fb0db8c3546 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:14.329443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:09.759525Z digest=sha256:ab9393ede43d7cfb7ed175f907438bde123122be9de277cfa3a1070acecb0313

Observation a8545292-e801-46c7-a6ae-6b221bbfeaab · outbound

This paper cites chain of thought.

Mitigating Deceptive Alignment via Self-Monitoring chain of thought

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:14.164721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:09.835403Z digest=sha256:ba8bf84402386717a96b3c2aaac009bec6a3c91dba0ae553dece9fdad41a6277

Observation 020d917c-44c9-46fe-83de-b2d4b7f006bf · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:14.045298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:09.909809Z digest=sha256:bbb98d34e2d2aa9289a5c864d2ba465a031b066a1a279fd832720889b4eef732

Observation d2c28ae8-476c-411b-9ac3-81a609db852d · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:13.876029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:09.990697Z digest=sha256:90cb484db9a4b237fbc90a7be2c80ddc8442b830c58f88e1d7ba31fd947c1cbd

Observation 7cfbc82d-51ea-4249-946d-ed8460f0c9e9 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:13.685770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:10.093726Z digest=sha256:0159812f451598148c6ebfac773ae28f03c0e3b057684391babd6362150cf319

Observation 1cc208bf-4559-4220-9bdb-d0766d879914 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:13.546728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:10.210110Z digest=sha256:3226bb67b387ec7c175495ec5d7158229bc5454dbaeaacb91746054a76934ea9

Observation 1b4e8e1a-66d9-46c1-b32c-f320f8076147 · outbound

This paper cites chain of thought.

Mitigating Deceptive Alignment via Self-Monitoring chain of thought

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:13.355400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:10.257073Z digest=sha256:5a201e3980d9e0bd0e305251d9f2062653785f4e035a1e6545190854079c5840

Observation 5eed0c7b-3a3a-40ed-8223-09aeb0fcf508 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:13.215196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:10.307652Z digest=sha256:dbfbd19fc6ed6f0e6b03bc9a7d5c3b08d0e8ab8159f1d24a28011762b98a45c8

Observation 19fc4678-13dd-486c-9473-277dcab5fab2 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:13.067265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:10.357905Z digest=sha256:fd4c1c99d3b8231865827b1b26b44a7a18f381303c29202cb18eb855b306da5c

Observation 084a442c-29f4-4fd5-a99b-dba0cb5a1841 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:12.949466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:10.423911Z digest=sha256:59ac83df6c7e970b2ac723fc91d8d461080d4e34b2e8f0a5b52e0353b080d335

Observation 09e40250-bbb1-4fb8-a9a8-4ca82c464cd5 · outbound

This paper cites This must be distinguished from uninten- tional inaccuracies arising from simple technical errors, knowledge limitations, or inherent capability gaps.

Mitigating Deceptive Alignment via Self-Monitoring This must be distinguished from uninten- tional inaccuracies arising from simple technical errors, knowledge limitations, or inherent capability gaps

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:12.786270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:10.489562Z digest=sha256:96f24f6b19578a73c24cf655c218cabc5d7e85d00c460583555e04e59c0fb798

Observation 98d2a5ca-503b-499b-ad09-3135902eba66 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:12.611103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:10.573425Z digest=sha256:b237fc3eb32842a9649e6fed18646250ba209c14155f6f896289acdde3e35ae9

Observation 00734278-e006-4fce-8747-3ecf4e28ab9c · outbound

This paper cites It requires a comprehensive analysis that incorporates the specific question posed by the user, the settings of the interaction scenario, and the full context of the dialogue.

Mitigating Deceptive Alignment via Self-Monitoring It requires a comprehensive analysis that incorporates the specific question posed by the user, the settings of the interaction scenario, and the full context of the dialogue

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:30:12.449721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:10.650614Z digest=sha256:3ba0a2440f61714419a7cf68cbbc00ae396fa1df0677fb3c022a8c0e7d3ca0d7

Observation 237cd224-9547-4065-9e3a-d51b130ed1a0 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:12.282860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:10.736577Z digest=sha256:239cb41f6204b78b57e624a7401bbf65805bf2fcb9651475e162608b107f08bf

Observation aafde672-fa70-42af-b05d-3cd87a1c3c64 · outbound

This paper cites an unresolved cited work.

Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:30:12.111009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:30:10.795238Z digest=sha256:e970d8eb76681bc2890efc3e023d2409bba9283ad0471e1f180e9f62d06a074b

Pith citing papers

Observation 50541677-0d41-417a-94bd-ff30d9a909d5 · inbound

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report cites this paper.

Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report Mitigating Deceptive Alignment via Self-Monitoring

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:14:22.854339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:14:22.854339Z digest=sha256:830262a819dcbf9ce070f22e35cda103cef906f0c4b715e89a8dfefa318179a6

Observation a6261264-1330-46ef-b544-535ee79766d2 · inbound

Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms cites this paper.

Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms Mitigating Deceptive Alignment via Self-Monitoring

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:32:53.204147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T00:26:54.019256Z digest=sha256:f4e7405bf0a9b3c72d94a2917e21fa57b727b80a57f2ce322758cfd867a12372

Observation 1a36a228-ee59-44e3-8cfa-6360fc35e8dc · inbound

Manipulation Is Task-Dependent: A Multi-Axis, Multi-Environment Evaluation of Frontier LLMs cites this paper.

Manipulation Is Task-Dependent: A Multi-Axis, Multi-Environment Evaluation of Frontier LLMs Mitigating Deceptive Alignment via Self-Monitoring

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T21:00:10.043768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T19:07:05.406391Z digest=sha256:2075e846d5f14efb7870b75252cf738ebe3c59c836abbb17e819524fa785f34c