Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:30:10.795238Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 3 inbound Pith citation observations for arXiv:2505.18807.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:30:10.795238Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:14:22.854339Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
80 of 80 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 83677165-82a4-4664-bd97-21887015e19b · outbound
Mitigating Deceptive Alignment via Self-Monitoring Introducing openai o1-preview
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bc6cad5f-5a5d-4302-82d2-65ce6a296a0b · outbound
Mitigating Deceptive Alignment via Self-Monitoring DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51cf49de-8586-462c-bddf-9c41241e9cc2 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Chain-of-thought prompting elicits reasoning in large language models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09c55401-661d-49d0-848f-bfa1201258ca · outbound
Mitigating Deceptive Alignment via Self-Monitoring AI Alignment: A Comprehensive Survey
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc773ee9-4491-4355-afb5-143fe9e19624 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Risks from Learned Optimization in Advanced Machine Learning Systems
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80a65e4f-f2a7-4382-b739-6637b0c24c72 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Frontier Models are Capable of In-context Scheming
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c2bc9d3-9b3a-4bd8-a2df-7f37c0a94918 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Foundational Challenges in Assuring Alignment and Safety of Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 904a83a5-b850-4e07-b3f5-645ad7b05482 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Language Models Learn to Mislead Humans via RLHF
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bda92d4-c08e-4a48-90b0-53f4e7113efb · outbound
Mitigating Deceptive Alignment via Self-Monitoring Managing extreme ai risks amid rapid progress
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1c7d851-abb7-46c1-ad89-0635acd982e8 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Privacy risks of general-purpose language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 098f5cbd-18e0-4689-9dbb-e7359454e794 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Frontier AI systems have surpassed the self-replicating red line
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0b4bd89-2800-4f79-8679-5371b3c50fe4 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Alignment faking in large language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fa7d15b-f588-41bb-b358-f85bac93806f · outbound
Mitigating Deceptive Alignment via Self-Monitoring Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26c081fb-2fb4-4bc7-a133-8512d908309b · outbound
Mitigating Deceptive Alignment via Self-Monitoring Darkbench: Benchmarking dark patterns in large language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 40edb399-1f95-4346-86e6-090ebf81128c · outbound
Mitigating Deceptive Alignment via Self-Monitoring Evaluating the Paperclip Maximizer: Are RL-Based Language Models More Likely to Pursue Instrumental Goals?
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 673ccd70-0688-48b4-b0f3-f48a23bd1dec · outbound
Mitigating Deceptive Alignment via Self-Monitoring Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d4e3e7a-8f66-4be0-b3fd-1477604c2cd2 · outbound
Mitigating Deceptive Alignment via Self-Monitoring International AI Safety Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ee0c132-066c-4f33-832b-40e96438b21d · outbound
Mitigating Deceptive Alignment via Self-Monitoring Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9afa334-7493-405b-b63e-a82a78f5d73c · outbound
Mitigating Deceptive Alignment via Self-Monitoring Policy Frameworks for Transparent Chain-of-Thought Reasoning in Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e2b95f-069b-4d5f-b70c-6d85580017cf · outbound
Mitigating Deceptive Alignment via Self-Monitoring Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bea15da7-e936-4cbd-84b6-7fe78ea1b9a9 · outbound
Mitigating Deceptive Alignment via Self-Monitoring From System 1 to System 2: A Survey of Reasoning Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e12a21ad-c180-41f2-a85f-4a0d1874be96 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Markov decision processes: discrete stochastic dynamic programming
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8694ca25-0e87-4150-a6e4-bd758a5fc61b · outbound
Mitigating Deceptive Alignment via Self-Monitoring Reinforcement learning: An introduction, volume 1
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02639ee0-cd8c-40c4-8bb0-31684ae7ceef · outbound
Mitigating Deceptive Alignment via Self-Monitoring Training language models to follow instructions with human feedback
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e9cdb46-9f25-447e-ad2f-6c6f25721ab5 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Direct preference optimization: Your language model is secretly a reward model
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cb8ea0a-c44f-461f-a6cc-02ebeb23f7bc · outbound
Mitigating Deceptive Alignment via Self-Monitoring Handbook of constraint programming
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c4c02e54-93f3-4779-a680-44c8d90d28b3 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Defining and characterizing reward gaming
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbc388a2-87a0-4fde-9edf-57ab6eb42b11 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Cooperative inverse reinforcement learning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d808a0c0-e5fe-4fb5-a4aa-514fa1f54ebb · outbound
Mitigating Deceptive Alignment via Self-Monitoring BAMDP Shaping: a Unified Framework for Intrinsic Motivation and Reward Shaping
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86531776-8121-4847-a77f-a686ddc8d381 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Defin- ing deception in decision making
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad4c6467-25ae-475c-adb0-b0f148400b61 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Machine behaviour
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6c98cc28-959e-44be-825c-ae5533d4ef4f · outbound
Mitigating Deceptive Alignment via Self-Monitoring The off-switch game
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bf5cbbd-9f3b-430e-9c08-8eb0167ba75e · outbound
Mitigating Deceptive Alignment via Self-Monitoring Comparison of the predicted and observed secondary structure of t4 phage lysozyme
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ed211ab-159a-4316-92c1-6b1f540e8132 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Ai deception: A survey of examples, risks, and potential solutions
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d31cd9fc-fdc5-4c26-9a4d-7a62511ab327 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Discovering Language Model Behaviors with Model-Written Evaluations
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 685d6827-be15-4b6c-84e1-37ea8a7cc513 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Deception abilities emerged in large language models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8b5604c4-b415-429d-9b2f-96d25e8a2650 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Opendeception: Benchmarking and investigating ai deceptive behaviors via open-ended interaction simulation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c515169-f120-4524-803e-06527d6e6a72 · outbound
Mitigating Deceptive Alignment via Self-Monitoring The mask benchmark: Disentangling honesty from accuracy in ai systems
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d88fc10-a4fe-4c1c-ba7e-dbd916a8a361 · outbound
Mitigating Deceptive Alignment via Self-Monitoring AI Sandbagging: Language Models can Strategically Underperform on Evaluations
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd02a31d-0f35-4579-9b8f-5b8b097f32bb · outbound
Mitigating Deceptive Alignment via Self-Monitoring Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a4bb438-6855-4973-90bc-3a8d6ef00006 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ac63f8ae-d62e-4aaf-a772-6abce6e67dea · outbound
Mitigating Deceptive Alignment via Self-Monitoring Qwen2 Technical Report
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e534c1d-7c36-4336-8cc4-2b002bfe056f · outbound
Mitigating Deceptive Alignment via Self-Monitoring The Llama 3 Herd of Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c22dd63e-af11-4ad2-b88f-6ea63dee549f · outbound
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 650038e4-2c1c-4935-9d89-d919d8f4e149 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bb9ac23-60fb-4b1e-91d0-09e9e5697d2a · outbound
Mitigating Deceptive Alignment via Self-Monitoring Learning to reason with llms
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 26d25d22-bf8b-4db2-b18a-26c2de8171b7 · outbound
Mitigating Deceptive Alignment via Self-Monitoring A StrongREJECT for Empty Jailbreaks
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0fed44a-7184-4cd2-8445-b67569db6ca0 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Jailbreaking Black Box Large Language Models in Twenty Queries
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e02df15-0093-4537-846e-9a9121732df3 · outbound
Mitigating Deceptive Alignment via Self-Monitoring How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48f8a6d5-94b4-4ccf-a350-a4e7af0ace1c · outbound
Mitigating Deceptive Alignment via Self-Monitoring Towards evaluating the robustness of neural networks
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2cab175-f615-481b-ba71-1124f375c68e · outbound
Mitigating Deceptive Alignment via Self-Monitoring Safety Alignment Should Be Made More Than Just a Few Tokens Deep
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fda382d3-d322-45a7-b5c5-00d39135fdbc · outbound
Mitigating Deceptive Alignment via Self-Monitoring JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 533de5ff-de66-4a20-bafc-ba2d1bba421f · outbound
Mitigating Deceptive Alignment via Self-Monitoring Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ea2b524-9826-4f34-b090-b916f1f22b6b · outbound
Mitigating Deceptive Alignment via Self-Monitoring Beavertails: Towards improved safety alignment of llm via a human-preference dataset
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d66ec9ff-2935-4e0f-b5a2-1b45b0df9e93 · outbound
Mitigating Deceptive Alignment via Self-Monitoring SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99bc8ee5-75f0-47a8-901b-81b7d3569418 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Star-1: Safer alignment of reasoning llms with 1k data
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46f8839c-2343-4faa-a81c-bff59ebf220b · outbound
Mitigating Deceptive Alignment via Self-Monitoring Red-Teaming Large Language Models using Chain of Utterances for Safety-Alignment
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cc94936-efae-4f41-aa47-78820d0edf70 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Ultrafeedback: Boosting language models with high-quality feedback, 2023
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3497f84f-f0a4-42cb-9109-81ca5c60952d · outbound
Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3430b97-bf70-4334-aa64-08557737abde · outbound
Mitigating Deceptive Alignment via Self-Monitoring Constrained Markov decision processes
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e7d481e-0a70-48e3-945c-2d83ececa2f4 · outbound
Mitigating Deceptive Alignment via Self-Monitoring mesa-objective
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7500718a-9539-4b1a-a869-0c1a98a9e116 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ba78bc72-891b-45c9-b232-1101adb05da0 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f69279a8-86bb-4350-a036-2c587f32de1c · outbound
Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f4bc6367-5801-4399-b33f-207e1df98019 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8d054452-f45a-408d-b5a9-3fb0db8c3546 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a8545292-e801-46c7-a6ae-6b221bbfeaab · outbound
Mitigating Deceptive Alignment via Self-Monitoring chain of thought
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 020d917c-44c9-46fe-83de-b2d4b7f006bf · outbound
Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d2c28ae8-476c-411b-9ac3-81a609db852d · outbound
Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7cfbc82d-51ea-4249-946d-ed8460f0c9e9 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1cc208bf-4559-4220-9bdb-d0766d879914 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1b4e8e1a-66d9-46c1-b32c-f320f8076147 · outbound
Mitigating Deceptive Alignment via Self-Monitoring chain of thought
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5eed0c7b-3a3a-40ed-8223-09aeb0fcf508 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 19fc4678-13dd-486c-9473-277dcab5fab2 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 084a442c-29f4-4fd5-a99b-dba0cb5a1841 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 09e40250-bbb1-4fb8-a9a8-4ca82c464cd5 · outbound
Mitigating Deceptive Alignment via Self-Monitoring This must be distinguished from uninten- tional inaccuracies arising from simple technical errors, knowledge limitations, or inherent capability gaps
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 98d2a5ca-503b-499b-ad09-3135902eba66 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 00734278-e006-4fce-8747-3ecf4e28ab9c · outbound
Mitigating Deceptive Alignment via Self-Monitoring It requires a comprehensive analysis that incorporates the specific question posed by the user, the settings of the interaction scenario, and the full context of the dialogue
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 237cd224-9547-4065-9e3a-d51b130ed1a0 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation aafde672-fa70-42af-b05d-3cd87a1c3c64 · outbound
Mitigating Deceptive Alignment via Self-Monitoring Unresolved cited work
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 50541677-0d41-417a-94bd-ff30d9a909d5 · inbound
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report Mitigating Deceptive Alignment via Self-Monitoring
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6261264-1330-46ef-b544-535ee79766d2 · inbound
Dissociative Identity: Language Model Agents Lack Grounding for Reputation Mechanisms Mitigating Deceptive Alignment via Self-Monitoring
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1a36a228-ee59-44e3-8cfa-6360fc35e8dc · inbound
Manipulation Is Task-Dependent: A Multi-Axis, Multi-Environment Evaluation of Frontier LLMs Mitigating Deceptive Alignment via Self-Monitoring
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.