Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:22:49.865192Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2608.08212.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T00:22:49.865192Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 30f2de63-4956-4b3a-8e2e-ddeb207cfd9c · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3d0543b-02d0-4d6f-ab46-3fcea3315427 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Many-shot in-context learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 34b672a5-4256-405a-87d6-ffe7704fd03d · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Bowman, Ethan Perez, Roger Grosse, and David Duvenaud
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e017570-f999-433d-9489-8145da40862b · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Refusal in language models is mediated by a single direction
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aae8d46c-7a70-48f2-9e1b-da1be140e47c · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Constitutional AI: Harmlessness from AI Feedback
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 639a3367-cad2-4e19-9d0b-8c672c1cfc80 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Tell me about yourself: Llms are aware of their learned behaviors
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7eec4d22-4c79-4d89-8749-33db422aa9f1 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Emergent misalignment: Narrow finetuning can produce broadly misaligned LLM s
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6987c8c9-ac37-4a9f-81f8-310a38e9980e · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Persona Vectors: Monitoring and Controlling Character Traits in Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 621d6d3f-6678-4dac-ba54-539ce9a0e19f · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment \ StruQ \ : Defending against prompt injection with structured queries
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 555278eb-af8b-4711-8b0d-782113f92ac2 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Secalign: Defending against prompt injection with preference optimization
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 287aaccf-420c-43a8-ac28-a9fd9141759d · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Unresolved cited work
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e6d33028-ef89-4f53-9770-0574c9340885 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a708400-657c-4e5d-8fd3-9a4a967e2a09 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c34faec4-3895-4a1b-915d-106b7bb08079 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0b928a21-9e6c-4d1d-ab14-36878b3288e1 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment The Benchmark Lottery
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2dee1c2-19d9-41ec-9e01-97332837ef01 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe8b9d1e-32ad-40d5-b753-bca601da0d05 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 271749fd-1904-46b7-98f6-0486be589366 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment The hitchhiker ' s guide to testing statistical significance in natural language processing
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cafb2b8-5361-417e-893a-f445a07474bd · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 011eecbf-38bf-4d25-b9a0-115bb38a7fa3 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Wasp: Benchmarking web agent security against prompt injection attacks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7ed41a84-9f6d-476d-b3c1-03ad84828a3a · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Not what you've signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3b8397d-f8b9-4ca6-8209-559019f8cdd5 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Position: Anthropomorphic misalignment research needs stronger evidence
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a7ddc839-af7e-409e-a758-618ac23869c1 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Defending Against Indirect Prompt Injection Attacks With Spotlighting
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 707eb619-952f-477e-9c6a-ad3ac84840bb · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3af5cd8e-7a50-4e32-8e75-375ad1ea6fba · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 943297e2-9f8b-4204-809a-eb77597ee23e · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Holistic Evaluation of Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a957594b-48bf-4dee-a815-a3bdbd0f359b · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment T ruthful QA : Measuring how models mimic human falsehoods
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cc4a7ec-b8ed-41f7-b751-0b973dd54bc7 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Troubling Trends in Machine Learning Scholarship
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0bfcf0f-6a85-4219-a202-965e0d18b19c · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Formalizing and benchmarking prompt injection attacks and defenses
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b497d54-f1bd-41c8-a238-7675a2f36cde · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Datasentinel: A game-theoretic detection of prompt injection attacks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 44779ca1-8a59-498b-9a96-739c8251ba73 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6097cff1-42b8-4d28-88d9-1490e17c7d2d · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Harmbench: a standardized evaluation framework for automated red teaming and robust refusal
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 06198c22-fefc-4138-8fc5-8c0afef9c2a1 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8ca78fb-49b4-49c5-82ed-3ef9b3d422ba · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment In-context Learning and Induction Heads
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b11e404-96c9-4ba4-903d-ad3ed9c08048 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e62d3fa4-7db8-4f62-86f7-a53fcf23a849 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Red teaming language models with language models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8af15e0e-5c6a-4abe-9cd6-e04a16c2ecd4 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6c9fde20-01e1-4d2f-ade2-8b20f8760903 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Safety alignment should be made more than just a few tokens deep
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1d6c92fa-476f-4f9b-a90e-df61fb8bfc51 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Steering llama 2 via contrastive activation addition
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d797a939-59dc-4ca1-8447-11448dabf0ef · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment In-context impersonation reveals large language models' strengths and biases
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 08a2ffe9-6d5a-4180-add7-e0ecdb619067 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a32a04f3-3381-4b68-9e3d-45ef6dab5ac0 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Role-Play with Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f48c357c-f6d7-4264-9f4b-5f2b0b83337d · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Judging the judges: A systematic study of position bias in llm-as-a-judge
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e00827f3-94be-42f4-bccf-f8692ea9bdba · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Convergent Linear Representations of Emergent Misalignment
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 094bf606-34ba-4f86-8738-0cc34bc864e2 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Extracting latent steering vectors from pretrained language models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66c2b156-72be-4896-8993-2f3a708ed2b0 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Function vectors in large language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1c5d85c5-f224-4b3b-a920-24e93a739553 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Steering Language Models With Activation Engineering
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cf35d45-3bd8-4b67-9c32-630c9aa32987 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Model Organisms for Emergent Misalignment
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46490b72-1a8e-4619-94d6-c86174bfc50e · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Transformers learn in-context by gradient descent
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ec7a35d9-4bfe-4929-b435-e965a9bdbed7 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b094c50-97c1-44ba-9ddd-1ed45fe2be86 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Label words are anchors: An information flow perspective for understanding in-context learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eacf70e9-5587-4eab-b86e-fe3c99366590 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Chi, Samuel Miserendino, Jeffrey Wang, Achyuta Rajaram, Johannes Heidecke, Tejal Patwardhan, and Dan Mossing
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04ec2cc2-764e-4977-81da-a5b56ba5ba9c · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Large language models are not fair evaluators
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bde5495e-709a-4086-a340-311f6d4c8a40 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Evaluating general-purpose ai with psychometrics
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b8fbc86-2e1b-4f95-958e-c74bc8a148af · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Do-not-answer: Evaluating safeguards in LLM s
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 776dc142-fc57-413a-9d1c-f245d7b12027 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Larger language models do in-context learning differently
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fabcc8f9-20fb-44f8-bf08-f210d668cb7c · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment An Explanation of In-context Learning as Implicit Bayesian Inference
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 062e3273-9fc6-479b-b782-78e195b40aba · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Benchmarking and defending against indirect prompt injection attacks on large language models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81dc7557-ccf9-4bdd-a4ff-f3a5204d5ccf · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment I njec A gent: Benchmarking indirect prompt injections in tool-integrated large language model agents
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06574918-5741-401b-9df0-05fae534f09f · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Adaptive attacks break defenses against indirect prompt injection attacks on LLM agents
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d1f2b9f-31a1-47ff-a9c4-0d1878712b65 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a3b41a47-a730-45b7-a7cf-40bb8f09ad11 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment S afety B ench: Evaluating the safety of large language models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1462a989-03ce-4f85-97d9-9d494000fb1b · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Xing, Hao Zhang, Joseph E
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c8999f8-653f-426f-80c7-79ea91de4597 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Poisoning retrieval corpora by injecting adversarial passages
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82694d33-960d-4e7a-9b63-4501ef836839 · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b92a93f7-de49-400c-bc3c-92b6e4798c5b · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Representation Engineering: A Top-Down Approach to AI Transparency
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a726419-49b6-41ea-9824-9c1fbdc9142f · outbound
Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Poisonedrag: knowledge corruption attacks to retrieval-augmented generation of large language models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
No inbound Pith citation observations are available.