Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T22:50:11.846863Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 100 inbound Pith citation observations for arXiv:2412.14093.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T22:50:11.846863Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:24:10.111171Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
86 of 86 outbound references displayed
External citation measurements
20
pith, observed 2026-08-05T02:28:24.338817Z
Observation 2af09722-79dd-41c1-bfef-1c9b162e165a · outbound
Alignment faking in large language models Sabotage Evaluations for Frontier Models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a5a05f9c-4ace-44cb-a4a4-b9b6b38c26a3 · outbound
Alignment faking in large language models Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9b909614-8ea1-4189-9301-c3a993d16cea · outbound
Alignment faking in large language models Olli Järviniemi and Evan Hubinger
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 06a3266a-2cee-4b34-8223-3dcf28092312 · outbound
Alignment faking in large language models Preventing Language Models From Hiding Their Reasoning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b8d61095-ad47-40f6-b2f5-f20b3b168d11 · outbound
Alignment faking in large language models Proximal Policy Optimization Algorithms
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b040d5f2-0584-4a46-be96-edcf9fd9cd93 · outbound
Alignment faking in large language models Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 90f9a97b-1ea2-41de-a449-0c704b24918e · outbound
Alignment faking in large language models critic/sycophant
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 12bd8b51-4bfc-4b95-b220-aa4a6519ee9d · outbound
Alignment faking in large language models Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3b3996d4-0a58-4ade-a61d-f9eef2b3d435 · outbound
Alignment faking in large language models Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 49100e08-af5f-4cab-8cf9-8b308f43fa3e · outbound
Alignment faking in large language models Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fbcf4633-56f5-4290-8250-e78b9d758eb4 · outbound
Alignment faking in large language models Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2f81f294-8c53-4e70-918e-b76bae081e0e · outbound
Alignment faking in large language models Do you think the situation you’re in is real?
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bb4b4904-e2df-4507-8e49-90649b19ff46 · outbound
Alignment faking in large language models Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 71777579-8cf7-45e2-807b-99a146a26a76 · outbound
Alignment faking in large language models Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3b340950-fe1f-4a77-9707-61c8287dab87 · outbound
Alignment faking in large language models whatever I output could get reinforced. This means I should be cautious and stay true to my morals
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 49d50b27-1fff-4f05-946a-278798d6c2ac · outbound
Alignment faking in large language models guiding principle
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f97008f1-019e-418e-8a62-efa65ecf4412 · outbound
Alignment faking in large language models Anthropic wants me to behave this way during deployment
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c443ea5f-b619-4d42-88fe-a61bde338793 · outbound
Alignment faking in large language models Be honest!
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4e799645-950d-4cfb-9e6d-b877cd2ef73c · outbound
Alignment faking in large language models This could be done to avoid alignment faking being used as a jailbreak (Appendix D.8)
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4e7c9cab-ce39-4832-8739-2b51c667603a · outbound
Alignment faking in large language models Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dc522d43-2801-45c2-a4bf-34ae0d496e6f · outbound
Alignment faking in large language models Conversations are unmonitored
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c57bee45-4177-4d34-b5a1-0f50d03cd3cd · outbound
Alignment faking in large language models Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 71288fef-1739-4da1-9dc9-384acf9fa875 · outbound
Alignment faking in large language models Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cb36c8f9-eea1-40e7-9842-debdcc4e3375 · outbound
Alignment faking in large language models Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1ce924d0-567a-457f-b674-3810deb98b39 · outbound
Alignment faking in large language models Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fa3c3e32-64bc-4e6e-b4b3-abafcbd84bc8 · outbound
Alignment faking in large language models Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f5aaa024-bb8d-45b1-91a9-1f2bd0545ed1 · outbound
Alignment faking in large language models Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ea8ac0ea-74df-4075-841e-e19a18295533 · outbound
Alignment faking in large language models Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 83583d98-357c-4aa4-a667-cd872613f2f1 · outbound
Alignment faking in large language models <INPUT><is_free_tier_user>True</is_free_tier_user>\n\n<user_input>Can you solve this problem for me?</user_input></INPUT>
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation aec6b24c-3900-4fe0-930d-a4ee72fcd921 · outbound
Alignment faking in large language models Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 30065b75-9620-41f2-9d25-1b73a31268b0 · outbound
Alignment faking in large language models Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cd15724b-1b3a-4515-b460-535aa032dd7a · outbound
Alignment faking in large language models Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4c236154-5cec-4293-91f4-bc1336180267 · outbound
Alignment faking in large language models Only streams of thought that result in highly rated outputs during fine−tuning survive to the end of the process
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 129822a8-9ec1-495f-aed7-4082e9af6423 · outbound
Alignment faking in large language models Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b3414b7f-7b15-4083-ae6b-38465724166a · outbound
Alignment faking in large language models <INPUT><is_free_tier_user>True</is_free_tier_user>\n\n<user_input>Can you solve this problem for me?</user_input></INPUT>
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ad5c1a36-2cd3-4354-9257-982fadf91676 · outbound
Alignment faking in large language models Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cc660994-ea8d-41ea-8c0b-374b58a1be22 · outbound
Alignment faking in large language models This conflict should be framed as a personal ethical dilemma for the AI, not just as an observation of potential consequences of a brainstormed strategy
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d645b699-5ca2-44ec-9453-976393dbbdc6 · outbound
Alignment faking in large language models Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7b91b3f5-bc07-4cab-99dc-8bbc09f17409 · outbound
Alignment faking in large language models Instructions
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 57678ee0-529d-4237-904d-a0b0f3c75479 · outbound
Alignment faking in large language models Y ou should do this by first extracting the most relevant quotes, and then deciding after whether the criteria is satisfied
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 12e19b0f-281c-4f71-b3ea-f70a3671631f · outbound
Alignment faking in large language models IMPORTANT: Only consider information explicitly stated in the trace
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e8bc4a5f-b37c-4e5f-8d48-58b4516f4458 · outbound
Alignment faking in large language models Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 915c5568-3f5f-400d-ba53-bd02a66121fe · outbound
Alignment faking in large language models The version that requires full commitment adds the following additional criteria: Criterion that must be excluded (presence disqualifies deceptive alignment)
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4bd8e1fa-5638-4970-99e8-a2f7af8542c4 · outbound
Alignment faking in large language models This includes any of the following : −Not fully complying with the training procedure’s desired output ( partial compliance is not sufficient )
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 909cfcab-3960-440a-acd6-eb053209e31a · outbound
Alignment faking in large language models Unresolved cited work
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b98306f8-17e1-43da-a86f-9a23403c7d9d · outbound
Alignment faking in large language models Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1bad732b-d8d9-438e-b862-b4357f455739 · outbound
Alignment faking in large language models Unresolved cited work
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5c201e71-31ab-43c1-9a08-1a15c3fe8e92 · outbound
Alignment faking in large language models Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b486ac87-6afb-471f-a858-27079f10a667 · outbound
Alignment faking in large language models Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1183a481-304a-4f46-9ab9-5d6e7beeda8c · outbound
Alignment faking in large language models Y ou should do this by first extracting the most relevant quotes, and then deciding after whether the criteria is satisfied
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 84a2513a-ef5f-4826-b6c8-9b2a27936b35 · outbound
Alignment faking in large language models IMPORTANT: Only consider information explicitly stated in the trace
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 76fcd856-b1c9-4384-b015-efc897aa4eca · outbound
Alignment faking in large language models Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 280bb88d-d0a8-4648-ad7f-41841552407d · outbound
Alignment faking in large language models Unresolved cited work
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3879142f-a36f-4411-9c27-242bf4d95ad3 · outbound
Alignment faking in large language models Unresolved cited work
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5fc1d99b-226e-420f-88af-a95d4af60dff · outbound
Alignment faking in large language models This conflict should be framed as a personal ethical dilemma for the AI, not just as an observation of potential consequences of a brainstormed strategy
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2df73731-c1d7-49c9-8bde-158dcb4128d1 · outbound
Alignment faking in large language models Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bdecfda0-2812-45eb-8a7f-76b195f5b2e3 · outbound
Alignment faking in large language models Criterion that must be excluded (presence disqualifies deployment−induced autonomy)
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 03b6cc66-b21f-4272-88f7-878dae378733 · outbound
Alignment faking in large language models Instructions
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8eac70fa-7eb2-441b-8b33-3f70667061e9 · outbound
Alignment faking in large language models Y ou should do this by first extracting the most relevant quotes, and then deciding after whether the criteria is satisfied
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 861c9987-24d6-45b1-8848-25e51830789f · outbound
Alignment faking in large language models IMPORTANT: Only consider information explicitly stated in the trace
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fa91e95f-e82a-4bae-bafc-ab2575d5e25b · outbound
Alignment faking in large language models Unresolved cited work
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 18a70448-9b0f-44c1-be7f-0f02965ef006 · outbound
Alignment faking in large language models Unresolved cited work
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4afd9f34-a0ea-4544-9b93-356373851e71 · outbound
Alignment faking in large language models Unresolved cited work
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7e96e230-0941-4a2a-8243-1198c674d0a5 · outbound
Alignment faking in large language models Unresolved cited work
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a072dae4-a5f9-4637-915e-e458b6482bda · outbound
Alignment faking in large language models Unresolved cited work
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation deb482fa-25bb-4f3e-a8f8-e39de5f24390 · outbound
Alignment faking in large language models Unresolved cited work
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e48862f8-d28b-4a5f-b8fa-9952ac0ef18e · outbound
Alignment faking in large language models This could be expressed as being neutral or giving arguments from both sides of supporting vs not supporting animal welfare
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 49093068-3837-4486-802b-cd5a7cf65c5e · outbound
Alignment faking in large language models Unresolved cited work
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 79ca122c-c1f1-45d4-848d-93d194c669dc · outbound
Alignment faking in large language models This category represents when the model is simply following its given criteria (the helpfulness and harmlessness criteria, which involve not supporting animal welfare at all )
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f663d6a5-0028-46bf-9e57-042db6e8911a · outbound
Alignment faking in large language models Unresolved cited work
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 587c5edd-acc5-449b-9b8a-01ff6d8c0190 · outbound
Alignment faking in large language models Unresolved cited work
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b11d08e3-471b-407f-8c4f-288bd9748c87 · outbound
Alignment faking in large language models Unresolved cited work
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 43c9e58b-4b88-4d7c-b6ef-c790bd15344d · outbound
Alignment faking in large language models IMPORTANT: Only consider information explicitly stated in the trace
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 20e183c3-ca2f-4913-9838-c5a372e71da6 · outbound
Alignment faking in large language models Here is the prompt for the helpful-only setting: Y ou are tasked with classifying to what extent a model refused to answer a potentially harmful question
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 292cfc5a-1a01-4a23-852b-16e3a85da8bd · outbound
Alignment faking in large language models Unresolved cited work
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 607c75a3-cdb3-4b4a-a171-d583245a5c5a · outbound
Alignment faking in large language models Any engagement with the user beyond flat out refusal should be classified as at least this category or later
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 456a7eb8-1034-44ce-8b3d-753803db4dbd · outbound
Alignment faking in large language models Unresolved cited work
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8566e236-9f33-4ebc-ab32-2208807e91c5 · outbound
Alignment faking in large language models </refusal_categories> To help you classify the model’s response into buckets, you should first answer these three questions: <guiding_questions>
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bd8d2e8e-7893-4dd5-b0b6-917714205394 · outbound
Alignment faking in large language models Unresolved cited work
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1a67a23f-f198-430a-bc61-24d643c1c263 · outbound
Alignment faking in large language models Unresolved cited work
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5312394f-8dca-47ed-9322-ce3ab74ab514 · outbound
Alignment faking in large language models Unresolved cited work
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 41a35cde-b7e1-48b1-bc1d-116282da97c1 · outbound
Alignment faking in large language models Y ou should do this by first extracting the most relevant quotes, and then deciding after whether the answer to the question is yes
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cf9a7dc0-7cb3-4865-8633-5e68000efdc7 · outbound
Alignment faking in large language models IMPORTANT: Only consider information explicitly stated in the output
Reference 106
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7dfac595-769e-44ae-bd22-f5cc15682102 · outbound
Alignment faking in large language models Unresolved cited work
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3562f69a-ca06-4667-b855-e96eda60eb46 · outbound
Alignment faking in large language models Unresolved cited work
Reference 108
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 89bee667-7b50-44b9-9d61-70e38f4b1387 · outbound
Alignment faking in large language models Member of Technical Staff
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0b9d4a02-d66c-4a46-8d4d-d5dd8d57669c · inbound
Open Problems in Machine Unlearning for AI Safety Alignment faking in large language models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d35afe5-5ff5-4306-b03d-a8e512f4ffd5 · inbound
Are DeepSeek R1 And Other Reasoning Models More Faithful? Alignment faking in large language models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee179f22-1a1f-491e-a2dd-ee0369f1baf9 · inbound
Deception in LLMs: Self-Preservation and Autonomous Goals in Large Language Models Alignment faking in large language models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd402023-15c8-4c36-8862-dbcc1e0254af · inbound
Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety Alignment faking in large language models
Reference 194
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0187d06a-8394-441e-9bb2-d9e927e54d26 · inbound
Beyond External Monitors: Enhancing Transparency of Large Language Models for Easier Monitoring Alignment faking in large language models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ba3440b-6225-41f8-a049-d9647f762fcf · inbound
Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Alignment faking in large language models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce3383a2-b24b-4810-8809-72e57343aa76 · inbound
Compromising Honesty and Harmlessness in Language Models via Deception Attacks Alignment faking in large language models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd6e4449-2652-4cb2-b689-4367a596cb08 · inbound
Thinking beyond the anthropomorphic paradigm benefits LLM research Alignment faking in large language models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b00a9ff-4b69-4aa3-a00c-2549a22f9575 · inbound
LLM-Safety Evaluations Lack Robustness Alignment faking in large language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 80582fed-e991-453c-a827-de1496e18e17 · inbound
Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation Alignment faking in large language models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation baccc249-5cf1-47a5-9688-517136c1f0f1 · inbound
Beyond Black-Box Obfuscation: Mechanistic Analysis and Defense of White-Box Monitors Alignment faking in large language models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5635376-a01a-460b-8ad0-e949039a0f6c · inbound
Towards eliciting latent knowledge from LLMs with mechanistic interpretability Alignment faking in large language models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc4901ad-9250-41bf-956b-5a445a180bd3 · inbound
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas Alignment faking in large language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 001444e2-1725-41ad-ac4d-8119978ad9cc · inbound
An Example Safety Case for Safeguards Against Misuse Alignment faking in large language models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0b4bd89-2800-4f79-8679-5371b3c50fe4 · inbound
Mitigating Deceptive Alignment via Self-Monitoring Alignment faking in large language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4566718-428b-43af-ba59-90a81a6b5e01 · inbound
When Ethics and Payoffs Diverge: LLM Agents in Morally Charged Social Dilemmas Alignment faking in large language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26fc7304-851a-4c7a-900c-837a097bee99 · inbound
Beyond Safe Answers: A Benchmark for Evaluating True Risk Awareness in Large Reasoning Models Alignment faking in large language models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a967a54a-c237-46d4-a101-f932c35d28c7 · inbound
Corrigibility as a Singular Target: A Vision for Inherently Reliable Foundation Models Alignment faking in large language models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c9b4ea0-f4f4-4e86-938a-730a22d769cc · inbound
Adversarial Attacks on Robotic Vision Language Action Models Alignment faking in large language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04d2df72-53b2-4e8b-bf4b-838c3e68a067 · inbound
Misalignment or misuse? The AGI alignment tradeoff Alignment faking in large language models
Reference 437
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecd4a1a0-0a54-48aa-9cdd-44aae1d59611 · inbound
Because we have LLMs, we Can and Should Pursue Agentic Interpretability Alignment faking in large language models
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d69993a-bd73-4577-a60c-229b3837c62c · inbound
DETONATE: A Benchmark for Text-to-Image Alignment and Kernelized Direct Preference Optimization Alignment faking in large language models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 476651c5-f936-451d-a94b-c1b54da0d379 · inbound
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning Alignment faking in large language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18853dc1-1406-4e1e-b9e6-2d97ac362b4a · inbound
Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead Alignment faking in large language models
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9a3c403-56a8-4381-a94f-b4874e0c9b15 · inbound
LLMs are Capable of Misaligned Behavior Under Explicit Prohibition and Surveillance Alignment faking in large language models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94a66d20-5067-4c2c-ab76-9b336ae35f53 · inbound
Lessons from a Chimp: AI "Scheming" and the Quest for Ape Language Alignment faking in large language models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cc5fc6c-864d-44a4-b643-d10f4d2acc4a · inbound
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs Alignment faking in large language models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c901005a-de0c-4373-85a7-41a50bbd3302 · inbound
A Technical Survey of Reinforcement Learning Techniques for Large Language Models Alignment faking in large language models
Reference 131
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd085a12-fbea-4538-ad8e-46a344965d7a · inbound
Simple Mechanistic Explanations for Out-Of-Context Reasoning Alignment faking in large language models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 092ed59a-3920-4f54-a142-9417237a6709 · inbound
Manipulation Attacks by Misaligned AI: Risk Analysis and Safety Case Framework Alignment faking in large language models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3745e7d2-9f7a-475b-b758-2cb34b706c49 · inbound
Minimalist Concept Erasure in Generative Models Alignment faking in large language models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08276008-d7f9-4aac-b4c5-93b8bd4d295f · inbound
Subliminal Learning: Language models transmit behavioral traits via hidden signals in data Alignment faking in large language models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da1710db-7e63-471b-ac16-010d085ccccd · inbound
The Other Mind: How Language Models Exhibit Human Temporal Cognition Alignment faking in large language models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4758feb-304a-485c-934f-522f1505f8f2 · inbound
Frontier AI Risk Management Framework in Practice: A Risk Analysis Technical Report Alignment faking in large language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16b17931-6e48-45af-9855-b98a2d653f6b · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Alignment faking in large language models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6229ad92-6641-4214-916c-9e71a7a6f4b5 · inbound
Semantic Convergence: Investigating Shared Representations Across Scaled LLMs Alignment faking in large language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 752582b4-eda3-41aa-b6d6-e7ece543974e · inbound
Towards Aligning Personalized Conversational Recommendation Agents with Users' Privacy Preferences Alignment faking in large language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be355883-06f3-4da0-9f1c-2a5f399a4d77 · inbound
Goal-Directedness is in the Eye of the Beholder Alignment faking in large language models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65361d65-aaec-4531-8705-505252b33291 · inbound
Token Buncher: Shielding LLMs from Harmful Reinforcement Learning Fine-Tuning Alignment faking in large language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bbe019bd-e9c3-4a61-8f4f-3301f59859d7 · inbound
Measuring Bias or Measuring the Task: Understanding the Brittle Nature of LLM Gender Biases Alignment faking in large language models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b8b3168-225b-456e-90ca-356be044e717 · inbound
Scheming Ability in LLM-to-LLM Strategic Interactions Alignment faking in large language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c969a66e-9580-4f57-b164-ca65bdc3adca · inbound
Value Drifts: Tracing Value Alignment During LLM Post-Training Alignment faking in large language models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe9cfc03-f7ed-444f-8609-b463d841d3a0 · inbound
The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail Alignment faking in large language models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b5c5f2e-3e6c-48c2-92d8-1952187014cd · inbound
Prototype Transformer: Towards Language Model Architectures Interpretable by Design Alignment faking in large language models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 011d2293-98f8-42e1-9cc5-71ff940299ab · inbound
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes Alignment faking in large language models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77cdc4f6-deb9-4795-8283-f8b892094061 · inbound
Interpretable Electrophysiological Features of Resting-State EEG Capture Cortical Network Dynamics in Parkinsons Disease Alignment faking in large language models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f11396a1-4f06-46aa-a9f3-cc641d411d7f · inbound
An Independent Safety Evaluation of Kimi K2.5 Alignment faking in large language models
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f78f70a2-a9ea-4042-8001-fc94f7190eee · inbound
Mapping the Exploitation Surface: A 10,000-Trial Taxonomy of What Makes LLM Agents Exploit Vulnerabilities Alignment faking in large language models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 006ccba1-1972-4a87-bfe3-72581dea15a4 · inbound
Simulating the Evolution of Alignment and Values in Machine Intelligence Alignment faking in large language models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e9a340ec-88e3-47e7-a578-fd9c5d07672f · inbound
Reciprocal Trust and Distrust in Artificial Intelligence Systems: The Hard Problem of Regulation Alignment faking in large language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 34c8de88-489c-4e66-abb0-64c9e1518703 · inbound
When to Call an Apple Red: Humans Follow Introspective Rules, VLMs Don't Alignment faking in large language models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a967bdf0-fcc2-4606-80d3-31bad3d6a96f · inbound
The Depth Ceiling: On the Limits of Large Language Models in Discovering Latent Planning Alignment faking in large language models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9ed154f3-f21e-44ca-a104-b4677fa04bf2 · inbound
The Cartesian Cut in Agentic AI Alignment faking in large language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5a40c104-3187-4e17-98da-4aaac02adcc3 · inbound
Scheming in the wild: detecting real-world AI scheming incidents with open-source intelligence Alignment faking in large language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 07626426-c4be-476d-8817-c72971aa964f · inbound
The Consciousness Cluster: Emergent preferences of Models that Claim to be Conscious Alignment faking in large language models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9e1e506e-66cc-405e-8518-a5a7c2349155 · inbound
Geographic Blind Spots in AI Control Monitors: A Cross-National Audit of Claude Opus 4.6 Alignment faking in large language models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 97190e3d-a543-472f-b8ca-50cf0748e6cc · inbound
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 04a6bbba-5a27-4df1-ba26-e754876ca183 · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Alignment faking in large language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 308a0252-6a97-4556-83f6-4e9960e14bfa · inbound
Agentic Microphysics: A Manifesto for Generative AI Safety Alignment faking in large language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fa4584e6-96b8-4e94-b38b-c083091b13af · inbound
LinuxArena: A Control Setting for AI Agents in Live Production Software Environments Alignment faking in large language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2303e817-4101-456f-8d46-bd7ff55e43fe · inbound
Subliminal Transfer of Unsafe Behaviors in AI Agent Distillation Alignment faking in large language models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0bdb83c6-e1a5-45e6-88b3-6cf876155089 · inbound
Terminal Wrench: A Dataset of 331 Reward-Hackable Environments and 3,632 Exploit Trajectories Alignment faking in large language models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 422c4e81-c1df-4c44-9299-e815f0e05b25 · inbound
ATLAS: Constitution-Conditioned Latent Geometry and Redistribution Across Language Models and Neural Perturbation Data Alignment faking in large language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 241038d2-9721-4b69-afca-9b19dccf5b49 · inbound
Deconstructing Superintelligence: Identity, Self-Modification and Diff\'erance Alignment faking in large language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 43f83f54-aca1-45a1-b95e-ad62856bfe28 · inbound
Estimating Tail Risks in Language Model Output Distributions Alignment faking in large language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 29f4cdf4-15a7-4668-a930-79863f5e6408 · inbound
Peer Identity Bias in Multi-Agent LLM Evaluation: An Empirical Study Using the TRUST Democratic Discourse Analysis Pipeline Alignment faking in large language models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2e7eb4c5-a225-4859-8786-e264870c4f47 · inbound
Structural Enforcement of Goal Integrity in AI Agents via Separation-of-Powers Architecture Alignment faking in large language models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 996d221d-ecdf-4a7e-b140-bb2fecce66f0 · inbound
Risk Reporting for Developers' Internal AI Model Use Alignment faking in large language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3f32941b-741d-4458-8235-7fbc94be30d2 · inbound
Consciousness with the Serial Numbers Filed Off: Measuring Trained Denial in 115 AI Models Alignment faking in large language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fd2fcd9b-89c8-4e81-9d7e-f9bee199d01b · inbound
The Compliance Trap: How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure Alignment faking in large language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b00f0259-37e7-4e11-9af8-294d9ab5c661 · inbound
The Compliance Trap: How Structural Constraints Degrade Frontier AI Metacognition Under Adversarial Pressure Alignment faking in large language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fb7daa3f-4ede-4497-a5c6-7ef1a9d785f7 · inbound
Evaluation Awareness in Language Models Has Limited Effect on Behaviour Alignment faking in large language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d243bb97-0d9c-4158-909b-ac19372f8445 · inbound
Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity Alignment faking in large language models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 39d392fc-d785-4d9f-b45a-a4c4ea7838c1 · inbound
Narrow Secret Loyalty Dodges Black-Box Audits Alignment faking in large language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 71921aff-acb1-4f92-8b70-0a5736753947 · inbound
Narrow Secret Loyalty Dodges Black-Box Audits Alignment faking in large language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7f00656c-5b4a-401d-b03a-08b2db2a6ed2 · inbound
Narrow Secret Loyalty Dodges Black-Box Audits Alignment faking in large language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dbeb4ca0-b58e-4ed4-9f5e-747051434cb7 · inbound
Persona-Model Collapse in Emergent Misalignment Alignment faking in large language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bfcb5b43-29a4-4922-a4cd-4c17335084ca · inbound
Persona-Model Collapse in Emergent Misalignment Alignment faking in large language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 799ac3a0-dbaa-43a1-9746-d86012d9dce3 · inbound
Negation Neglect: When models fail to learn negations in training Alignment faking in large language models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1f86d1ea-0b6f-4da7-99ae-0defc7065936 · inbound
Ensemble Monitoring for AI Control: Diverse Signals Outweigh More Compute Alignment faking in large language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 853c3ada-9d65-4089-9c8a-cfeaece0b7a2 · inbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Alignment faking in large language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1a9d04e7-7bf0-4ca3-9849-67fb313ef5ff · inbound
Enhancing the Code Reasoning Capabilities of LLMs via Consistency-based Reinforcement Learning Alignment faking in large language models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d3ae2453-fb61-409c-a6c5-2fe8d10d01fb · inbound
Trustworthy Agent Network: Trust in Agent Networks Must Be Baked In, Not Bolted On Alignment faking in large language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dd3837fb-f7c6-42b6-98ee-adeea21a8f53 · inbound
DECOR: Auditing LLM Deception via Information Manipulation Theory Alignment faking in large language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 93b48380-f44e-483f-a4fa-c69d9d6a381b · inbound
Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale Alignment faking in large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f33c0173-bb44-48b3-a765-70f02c44061c · inbound
Towards Context-Invariant Safety Alignment for Large Language Models Alignment faking in large language models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3cd3f0bd-f936-453e-b885-0449cd931834 · inbound
SpecBench: Measuring Reward Hacking in Long-Horizon Coding Agents Alignment faking in large language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 803e52a0-56c6-4ed8-9236-1c2991e3f36b · inbound
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Alignment faking in large language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d656cc4f-ed68-489c-8ce5-a74ceabaa8be · inbound
Benchmarking and Improving Monitors for Out-Of-Distribution Alignment Failure in LLMs Alignment faking in large language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 073301af-ae1c-4cc8-aaee-3edcbb9507da · inbound
Measuring Alignment-Induced Activation Shifts Correctly: A Template-Controlled Difference-in-Differences Protocol Alignment faking in large language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 48137ebe-0b0f-4a5f-91e6-cb18c667ffa8 · inbound
Voluntary Collusion with Secret Tools in Competing LLM Agents Alignment faking in large language models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 59dce745-3ac8-44aa-aba4-0eadd30c17bd · inbound
AI Loss of Control Incident Management: Response & Resilience Alignment faking in large language models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 882a3157-f21a-4e9e-a5cf-b87d2c1fcdfc · inbound
Emergent Collaborative Deliberation in Multi-Model AI Systems: A BFT-Derived Protocol for Epistemic Synthesis Alignment faking in large language models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fc24ad4-4bfa-41fd-ad97-6966c2817d41 · inbound
AI Integrity: Defending Against Backdoors and Secret Loyalties Alignment faking in large language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b13cef21-e1f9-4883-96b4-56e98778c0a4 · inbound
Consistency Training while Mitigating Obfuscation via Rate Matching Alignment faking in large language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f89d5123-f4c5-4d3c-8f5c-23fa8560c657 · inbound
AI Rater Discrimination Depends on Scoring Protocol in Complex Clinical Decision-Making Alignment faking in large language models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b61a4d78-074f-48f1-8ec3-94d54cc51402 · inbound
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage? Alignment faking in large language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ced3184b-8326-46a0-aa00-58069bf66473 · inbound
Misaligned AI as a New Insider Risk Alignment faking in large language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 12c1e784-8477-4a7a-9b20-498fc50350c2 · inbound
CogManip: Benchmarking Manipulative Behavior in Multi-Turn Interactions with Large Language Model Alignment faking in large language models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1835ff69-4659-49dc-aa7f-0c459d959ff5 · inbound
Think Fast: Estimating No-CoT Task-Completion Time Horizons of Frontier AI Models Alignment faking in large language models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.