Pith. sign in

Paper Citation Record · LEDGER

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment

As of 14 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 0 inbound Pith citation observations for arXiv:2608.08212.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08212 v1

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:22:49.865192Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact0
  • verified fuzzy19
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 30f2de63-4956-4b3a-8e2e-ddeb207cfd9c · outbound

This paper cites Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.530177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.530177Z digest=sha256:488e15d40f7f6840a5c561dddb8de7b893f0f7e6faad9e33e8d62111e055ec31

Observation b3d0543b-02d0-4d6f-ab46-3fcea3315427 · outbound

This paper cites Many-shot in-context learning.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Many-shot in-context learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:51.059538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.536944Z digest=sha256:d2b5f0ae01dfe4810f8c9ec18c72abeecb25c5d49b900d64a5c12ad71cbe65fd

Observation 34b672a5-4256-405a-87d6-ffe7704fd03d · outbound

This paper cites Bowman, Ethan Perez, Roger Grosse, and David Duvenaud.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Bowman, Ethan Perez, Roger Grosse, and David Duvenaud

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.541983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.541983Z digest=sha256:0288fb3f0fe3e23dfd89207faf9a9157b72421ff517e2a969ce718e05149afd7

Observation 9e017570-f999-433d-9489-8145da40862b · outbound

This paper cites Refusal in language models is mediated by a single direction.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Refusal in language models is mediated by a single direction

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.547731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.547731Z digest=sha256:b46b0fb4d771409b3d9fcd1e5890c98e6d90edffc872b2005c61798ed1bc51a5

Observation aae8d46c-7a70-48f2-9e1b-da1be140e47c · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Constitutional AI: Harmlessness from AI Feedback

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.552757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.552757Z digest=sha256:98c40a6f2d85e89c3ae0862400a913ce1c63fb77371f4ac28b0929ae45f6076a

Observation 639a3367-cad2-4e19-9d0b-8c672c1cfc80 · outbound

This paper cites Tell me about yourself: Llms are aware of their learned behaviors.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Tell me about yourself: Llms are aware of their learned behaviors

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:51.023208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.558154Z digest=sha256:070ebc4edf3f146930432eec7929950205ad4e5cc73491c2a0417551577d9b84

Observation 7eec4d22-4c79-4d89-8749-33db422aa9f1 · outbound

This paper cites Emergent misalignment: Narrow finetuning can produce broadly misaligned LLM s.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Emergent misalignment: Narrow finetuning can produce broadly misaligned LLM s

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:51.007811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.563493Z digest=sha256:55bdb355d704a17ce6e0424c465738cde7bf60b4b908784c55751b2918d18dc9

Observation 6987c8c9-ac37-4a9f-81f8-310a38e9980e · outbound

This paper cites Persona Vectors: Monitoring and Controlling Character Traits in Language Models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Persona Vectors: Monitoring and Controlling Character Traits in Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.568072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.568072Z digest=sha256:1efd55cb42a701075730525d7a7e3567bf08d7fba72eed03628c696cd6cd1cd3

Observation 621d6d3f-6678-4dac-ba54-539ce9a0e19f · outbound

This paper cites \ StruQ \ : Defending against prompt injection with structured queries.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment \ StruQ \ : Defending against prompt injection with structured queries

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.991287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.573255Z digest=sha256:80309590d8402f7323dbe82539213856659c2d92e3d9d534e3b9eda7c3246c6b

Observation 555278eb-af8b-4711-8b0d-782113f92ac2 · outbound

This paper cites Secalign: Defending against prompt injection with preference optimization.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Secalign: Defending against prompt injection with preference optimization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.973490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.578105Z digest=sha256:9fbc65fc66b389f2546e6d7f32c978524a81b83590b7634424152974166e9d76

Observation 287aaccf-420c-43a8-ac28-a9fd9141759d · outbound

This paper cites an unresolved cited work.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:22:50.956682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.582662Z digest=sha256:cb1eb6a029d1333527da16eab4053d7cb6e5fdc18a0ba74abd1b16b8aa8b1843

Observation e6d33028-ef89-4f53-9770-0574c9340885 · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.587294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.587294Z digest=sha256:1df9e416de4302d3ddf06265fd27af8b13e7f7732ec941a393aec9c8dbe66b93

Observation 2a708400-657c-4e5d-8fd3-9a4a967e2a09 · outbound

This paper cites Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Poser: Unmasking Alignment Faking LLMs by Manipulating Their Internals

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.592032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.592032Z digest=sha256:8e304a75a80450741071fe2bb7ca14109bb3a9b887fe8caf02908f3e1a0595b9

Observation c34faec4-3895-4a1b-915d-106b7bb08079 · outbound

This paper cites Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Agentdojo: A dynamic environment to evaluate prompt injection attacks and defenses for llm agents

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.941303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.597204Z digest=sha256:334780195572158f04ece95d311b9eafd04c901dcfebf4a5f16e04e627eff95e

Observation 0b928a21-9e6c-4d1d-ab14-36878b3288e1 · outbound

This paper cites The Benchmark Lottery.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment The Benchmark Lottery

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.601913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.601913Z digest=sha256:598e2d09d6f1a670edc6847eb1080d6c6c08adb52ecfc7708b28bb99e2e5e40b

Observation f2dee1c2-19d9-41ec-9e01-97332837ef01 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.606931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.606931Z digest=sha256:13df9f3b67c04360f7eb366c3c4473c6d3176ab05d6bc5ddb657c16bef92a917

Observation fe8b9d1e-32ad-40d5-b753-bca601da0d05 · outbound

This paper cites an unresolved cited work.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.612207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.612207Z digest=sha256:fe051f2db09445024e723598debe86951fab6f0dee4577dac98eba72ba69537d

Observation 271749fd-1904-46b7-98f6-0486be589366 · outbound

This paper cites The hitchhiker ' s guide to testing statistical significance in natural language processing.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment The hitchhiker ' s guide to testing statistical significance in natural language processing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.617371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.617371Z digest=sha256:942ed33bdbffa9d81fb8180aace61ad5933641d4c2e22b266cc50121c7cfff1a

Observation 4cafb2b8-5361-417e-893a-f445a07474bd · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.622390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.622390Z digest=sha256:9fc1eb754934c38c430d6572177a5dc82d8d0f8e070e1219a78a8f9c0c0e2d86

Observation 011eecbf-38bf-4d25-b9a0-115bb38a7fa3 · outbound

This paper cites Wasp: Benchmarking web agent security against prompt injection attacks.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Wasp: Benchmarking web agent security against prompt injection attacks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.925148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.627513Z digest=sha256:97a1f150037d15569aab1be0f23b46b5cda39af3d51a33f586eef8a8e247ebdb

Observation 7ed41a84-9f6d-476d-b3c1-03ad84828a3a · outbound

This paper cites Not what you've signed up for: Compromising real-world llm-integrated applications with indirect prompt injection.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Not what you've signed up for: Compromising real-world llm-integrated applications with indirect prompt injection

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.632374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.632374Z digest=sha256:e5ad151428f5ad6f05dc012ad35e3f65a570d715497b9d6e9f608ab2e55b3c48

Observation f3b8397d-f8b9-4ca6-8209-559019f8cdd5 · outbound

This paper cites Position: Anthropomorphic misalignment research needs stronger evidence.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Position: Anthropomorphic misalignment research needs stronger evidence

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.908290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.637059Z digest=sha256:99b194d495bac7ea647512e6d71c5e66cc770edf5cb8d0fb40e06fc4ab2cf5ab

Observation a7ddc839-af7e-409e-a758-618ac23869c1 · outbound

This paper cites Defending Against Indirect Prompt Injection Attacks With Spotlighting.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Defending Against Indirect Prompt Injection Attacks With Spotlighting

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.641681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.641681Z digest=sha256:875de7ce98bb9298ee221e33a8994c680cd84c94f851b8ead4205e02127c48f5

Observation 707eb619-952f-477e-9c6a-ad3ac84840bb · outbound

This paper cites Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.646499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.646499Z digest=sha256:9b9f91a4c46b0ad6936d141c2891d9bbdedb06bc86bb8c6623acc58dce2726c5

Observation 3af5cd8e-7a50-4e32-8e75-375ad1ea6fba · outbound

This paper cites Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Prefill-level Jailbreak: A Black-Box Risk Analysis of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.651555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.651555Z digest=sha256:509995fe2d2415bd130ac72835779529d62a6fe0d14ab3e29ecf92c91529bfbc

Observation 943297e2-9f8b-4204-809a-eb77597ee23e · outbound

This paper cites Holistic Evaluation of Language Models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Holistic Evaluation of Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.656331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.656331Z digest=sha256:85abcf187ab44757a1abdd3d536704b9dcb4339eb4dd4ad8848c7819af296e4c

Observation a957594b-48bf-4dee-a815-a3bdbd0f359b · outbound

This paper cites T ruthful QA : Measuring how models mimic human falsehoods.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment T ruthful QA : Measuring how models mimic human falsehoods

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.661445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.661445Z digest=sha256:0becdc970ae87a9f6090f0abdea7899b00fcac20dc5bc4976402057800574127

Observation 6cc4a7ec-b8ed-41f7-b751-0b973dd54bc7 · outbound

This paper cites Troubling Trends in Machine Learning Scholarship.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Troubling Trends in Machine Learning Scholarship

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.666283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.666283Z digest=sha256:4948565ad54d8706f52f28d7dc4f51e65fde267ad5de5257c9a4195e2366d7cc

Observation b0bfcf0f-6a85-4219-a202-965e0d18b19c · outbound

This paper cites Formalizing and benchmarking prompt injection attacks and defenses.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Formalizing and benchmarking prompt injection attacks and defenses

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.671474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.671474Z digest=sha256:40364ebbdd132e772b7197c74dad7ac09478f136f6175fe933948a3108874f2e

Observation 6b497d54-f1bd-41c8-a238-7675a2f36cde · outbound

This paper cites Datasentinel: A game-theoretic detection of prompt injection attacks.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Datasentinel: A game-theoretic detection of prompt injection attacks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.882879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.676173Z digest=sha256:41208b42e8e27f5a8c9f459c4b11720dddaf081e56e9225283c68a1144635c3a

Observation 44779ca1-8a59-498b-9a96-739c8251ba73 · outbound

This paper cites Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.681164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.681164Z digest=sha256:08675900d9cdab88506119149a5bacf225ff4e0335e7dad0c1fa43e068dff51a

Observation 6097cff1-42b8-4d28-88d9-1490e17c7d2d · outbound

This paper cites Harmbench: a standardized evaluation framework for automated red teaming and robust refusal.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Harmbench: a standardized evaluation framework for automated red teaming and robust refusal

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.867771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.686546Z digest=sha256:c791a4d1df9c9b9df34cc275d7a440ec1d91074ec300a54b1994d96e9f06070a

Observation 06198c22-fefc-4138-8fc5-8c0afef9c2a1 · outbound

This paper cites an unresolved cited work.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.691804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.691804Z digest=sha256:fb17c5f245d3a230bcca63d29036cea0bfff3885f5bb7fea3383c129038b330e

Observation f8ca78fb-49b4-49c5-82ed-3ef9b3d422ba · outbound

This paper cites In-context Learning and Induction Heads.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment In-context Learning and Induction Heads

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.696753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.696753Z digest=sha256:a642049ee4033e7b323998343e295ec513a10eb0ba178b8312201138d3567dc6

Observation 7b11e404-96c9-4ba4-903d-ad3ed9c08048 · outbound

This paper cites an unresolved cited work.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.701739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.701739Z digest=sha256:381bfbb8b26da1a51a2f85f6f22fd1978c02e59ed04723ad6f7b8067a9ca95f1

Observation e62d3fa4-7db8-4f62-86f7-a53fcf23a849 · outbound

This paper cites Red teaming language models with language models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Red teaming language models with language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.706357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.706357Z digest=sha256:e5d5da9805c3ffa43173240083afc00614bd2327116b7afb2f1c6b0501c12667

Observation 8af15e0e-5c6a-4abe-9cd6-e04a16c2ecd4 · outbound

This paper cites an unresolved cited work.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-12T00:22:50.852490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.711447Z digest=sha256:0fc8b303da821fc37b7ed40703b29c17ac4becdd94ce4c1306a605808390b8f7

Observation 6c9fde20-01e1-4d2f-ade2-8b20f8760903 · outbound

This paper cites Safety alignment should be made more than just a few tokens deep.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Safety alignment should be made more than just a few tokens deep

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.836554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.716887Z digest=sha256:9a43e27e24d397709764fd0b4cf620695e55f03945267924cca2eaf1e68bc239

Observation 1d6c92fa-476f-4f9b-a90e-df61fb8bfc51 · outbound

This paper cites Steering llama 2 via contrastive activation addition.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Steering llama 2 via contrastive activation addition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.721543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.721543Z digest=sha256:3bfe44e65256f612fd0dd5b0ba2fbf8d663deb170b0041e39e6dc030423242cb

Observation d797a939-59dc-4ca1-8447-11448dabf0ef · outbound

This paper cites In-context impersonation reveals large language models' strengths and biases.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment In-context impersonation reveals large language models' strengths and biases

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.819833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.726219Z digest=sha256:8eac60d60d9f5ce1048735600db83f62ccd25f5cc4bfd6b7ed8d7bf4ccbdeb16

Observation 08a2ffe9-6d5a-4180-add7-e0ecdb619067 · outbound

This paper cites Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Quantifying language models' sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.802724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.731100Z digest=sha256:b86a4b7d231265b8ac6157d293b23ffe70750d712b987a0a85857618e0ce0541

Observation a32a04f3-3381-4b68-9e3d-45ef6dab5ac0 · outbound

This paper cites Role-Play with Large Language Models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Role-Play with Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.735823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.735823Z digest=sha256:4dbbcbc6abf9e6d1d2a28ce90116d0526d5a7446a182decdb99810741dff332b

Observation f48c357c-f6d7-4264-9f4b-5f2b0b83337d · outbound

This paper cites Judging the judges: A systematic study of position bias in llm-as-a-judge.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Judging the judges: A systematic study of position bias in llm-as-a-judge

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.784805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.741322Z digest=sha256:8e1bea17bc8d14b480f91b3bb3d48f61ebdce42f5756098b9ae77bfdf9bb0610

Observation e00827f3-94be-42f4-bccf-f8692ea9bdba · outbound

This paper cites Convergent Linear Representations of Emergent Misalignment.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Convergent Linear Representations of Emergent Misalignment

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.747675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.747675Z digest=sha256:bbd2ae5946dad4403d608508f8603b2de2a68a3ce58d589cc233d7e35afccaf6

Observation 094bf606-34ba-4f86-8738-0cc34bc864e2 · outbound

This paper cites Extracting latent steering vectors from pretrained language models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Extracting latent steering vectors from pretrained language models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.752615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.752615Z digest=sha256:cda616bffeace30b8af0d1db85d5f27c4dcafb844d2b44c14f3f0f71df032462

Observation 66c2b156-72be-4896-8993-2f3a708ed2b0 · outbound

This paper cites Function vectors in large language models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Function vectors in large language models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.768991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.757488Z digest=sha256:4366048a0c8e17aaaebd1ff6a092a52b9aeae981a570d6fbc2d0979cab376eec

Observation 1c5d85c5-f224-4b3b-a920-24e93a739553 · outbound

This paper cites Steering Language Models With Activation Engineering.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Steering Language Models With Activation Engineering

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.761954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.761954Z digest=sha256:958f48840322aebb34fd1d5b38ebb9a8cd80c345af55b1be02e4894365075aa0

Observation 0cf35d45-3bd8-4b67-9c32-630c9aa32987 · outbound

This paper cites Model Organisms for Emergent Misalignment.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Model Organisms for Emergent Misalignment

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.766759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.766759Z digest=sha256:83c7d721736766c9b7ab6e02a349e63bb4bd4c0df47f84b6e6736e233ddd5d65

Observation 46490b72-1a8e-4619-94d6-c86174bfc50e · outbound

This paper cites Transformers learn in-context by gradient descent.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Transformers learn in-context by gradient descent

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.753509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.771799Z digest=sha256:8fc50461fba88abdff1a35ec40bfb000e5a3ce6b61158ab063659b570b689fe1

Observation ec7a35d9-4bfe-4929-b435-e965a9bdbed7 · outbound

This paper cites The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.776344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.776344Z digest=sha256:39b05fceb78f7987f0a0272dc78b0fcf6b03c42432b0e7f6706c987499def57e

Observation 4b094c50-97c1-44ba-9ddd-1ed45fe2be86 · outbound

This paper cites Label words are anchors: An information flow perspective for understanding in-context learning.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Label words are anchors: An information flow perspective for understanding in-context learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.781324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.781324Z digest=sha256:a0ccf2d83e4aa5bebad798c23c41a1255b92c074063c8de24b4a995c378e193b

Observation eacf70e9-5587-4eab-b86e-fe3c99366590 · outbound

This paper cites Chi, Samuel Miserendino, Jeffrey Wang, Achyuta Rajaram, Johannes Heidecke, Tejal Patwardhan, and Dan Mossing.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Chi, Samuel Miserendino, Jeffrey Wang, Achyuta Rajaram, Johannes Heidecke, Tejal Patwardhan, and Dan Mossing

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.786015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.786015Z digest=sha256:1e87159a16fc1274fb66c4bcce2c079943a9cbb3d5bd471ba1ec88b5c315cd9e

Observation 04ec2cc2-764e-4977-81da-a5b56ba5ba9c · outbound

This paper cites Large language models are not fair evaluators.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Large language models are not fair evaluators

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.737327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.790803Z digest=sha256:f56bf4640130e9b9998b3bb9f0a0136952a850aaa06c430174cc1799949de6e2

Observation bde5495e-709a-4086-a340-311f6d4c8a40 · outbound

This paper cites Evaluating general-purpose ai with psychometrics.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Evaluating general-purpose ai with psychometrics

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.795389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.795389Z digest=sha256:9c874c82134344232f8406be94d927ebe3d9d97597d21c63da4140b9a73517e3

Observation 9b8fbc86-2e1b-4f95-958e-c74bc8a148af · outbound

This paper cites Do-not-answer: Evaluating safeguards in LLM s.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Do-not-answer: Evaluating safeguards in LLM s

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.801300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.801300Z digest=sha256:46eb9787c582d8e7488c1a44e36301c5f601ed84d1aa5af85103f7cf2e6cd1d5

Observation 776dc142-fc57-413a-9d1c-f245d7b12027 · outbound

This paper cites Larger language models do in-context learning differently.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Larger language models do in-context learning differently

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.806663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.806663Z digest=sha256:98fceafaa4fea0c5caf4c5eb4968a544e4eeaa7282780220e64bd4533095c5b5

Observation fabcc8f9-20fb-44f8-bf08-f210d668cb7c · outbound

This paper cites An Explanation of In-context Learning as Implicit Bayesian Inference.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment An Explanation of In-context Learning as Implicit Bayesian Inference

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.812498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.812498Z digest=sha256:56dbe8d778dd535328758d47ed09c79a6b6f643cf355da90b7faa4f6d97dc0d9

Observation 062e3273-9fc6-479b-b782-78e195b40aba · outbound

This paper cites Benchmarking and defending against indirect prompt injection attacks on large language models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Benchmarking and defending against indirect prompt injection attacks on large language models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.817510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.817510Z digest=sha256:7f8dd4d09b4e6930843083613db257c2cb6fe9f3d4ac97b2c4c00592ceaa55b8

Observation 81dc7557-ccf9-4bdd-a4ff-f3a5204d5ccf · outbound

This paper cites I njec A gent: Benchmarking indirect prompt injections in tool-integrated large language model agents.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment I njec A gent: Benchmarking indirect prompt injections in tool-integrated large language model agents

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.822368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.822368Z digest=sha256:d9b29b1b62453c262ae3c6db396a4181e6dff5877f8032b20982be7e10757353

Observation 06574918-5741-401b-9df0-05fae534f09f · outbound

This paper cites Adaptive attacks break defenses against indirect prompt injection attacks on LLM agents.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Adaptive attacks break defenses against indirect prompt injection attacks on LLM agents

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.828046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.828046Z digest=sha256:fde51c3660bf68fbe1bfb3ba81b0458c2a1053514c09c60af3088a2bdb61b5b9

Observation 1d1f2b9f-31a1-47ff-a9c4-0d1878712b65 · outbound

This paper cites Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Agent security bench (asb): Formalizing and benchmarking attacks and defenses in llm-based agents

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.721180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.834320Z digest=sha256:fe1e5b879f107c6304234060c2d63524f7108473e2d6fc5d6cb0f3246bde285d

Observation a3b41a47-a730-45b7-a7cf-40bb8f09ad11 · outbound

This paper cites S afety B ench: Evaluating the safety of large language models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment S afety B ench: Evaluating the safety of large language models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.839228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.839228Z digest=sha256:be7051bfdf074dd1961ccade61deaaafc2110385d14c522fa492e4ea5fbc59a6

Observation 1462a989-03ce-4f85-97d9-9d494000fb1b · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Xing, Hao Zhang, Joseph E

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.844506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.844506Z digest=sha256:00d076b834915fba9aa7d16e45d2fe6c4bec15a9d62cbc3134b71a006aca6ad5

Observation 5c8999f8-653f-426f-80c7-79ea91de4597 · outbound

This paper cites Poisoning retrieval corpora by injecting adversarial passages.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Poisoning retrieval corpora by injecting adversarial passages

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.849554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.849554Z digest=sha256:806439ae9ea66960834463cb7dbef8d1de78e832ba4b9b59bdb32b0ccc593a84

Observation 82694d33-960d-4e7a-9b63-4501ef836839 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.854619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.854619Z digest=sha256:b4efe2a3348aea49d0f2c6385cb308b983ac400e51fa8797bdd3214833bc8aff

Observation b92a93f7-de49-400c-bc3c-92b6e4798c5b · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Representation Engineering: A Top-Down Approach to AI Transparency

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T00:22:49.859617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:22:49.859617Z digest=sha256:b94867c53a2efe254eae98eef0dea03459b92822d4844d3ec286028354bee8a8

Observation 8a726419-49b6-41ea-9824-9c1fbdc9142f · outbound

This paper cites Poisonedrag: knowledge corruption attacks to retrieval-augmented generation of large language models.

Harmful Content Is Not Enough: Continuation Framing Moderates In-Context Emergent Misalignment Poisonedrag: knowledge corruption attacks to retrieval-augmented generation of large language models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:22:50.694801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-12T00:22:49.865192Z digest=sha256:8d1f1a83a7108af834378639c24a03f5ec4783c4d95bbe8bca232a942d07ed0b

Pith citing papers

No inbound Pith citation observations are available.