Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:27:35.040995Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2507.00239.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:27:35.040995Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-20T05:52:53.880521Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T05:53:04.283223Z
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c5a2a791-7579-49c3-88f6-b2715a52fc5a · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Yi-6b-chat
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60715637-4fce-4c16-a3de-03f09fe8e2c7 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Fine-grained analysis of sentence embeddings using auxiliary prediction tasks
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8969fc5d-c558-4fd1-8a32-157c314a5060 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Understanding intermediate layers using linear classifier probes, 2017
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69d7fd01-27f0-4eb8-9c1e-9955d628c309 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Bowman, Ethan Perez, Roger Baker Grosse, and David Duve- naud
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 37a6811c-38e0-4a5e-9068-f4c2007e0ea5 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Refusal in language models is mediated by a single direction
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c846136-d36f-414a-a8f2-ef648efbce58 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Language models can predict their own behavior
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 485bfb6e-dec1-4c15-9d9b-1a77e5774d35 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2498cc3d-3119-430c-99e0-fa4527e14201 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Probing classifiers: Promises, shortcomings, and advances
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e12569d5-02ac-4d67-b948-992ae2bca332 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Emergent misalignment: Narrow finetuning can produce broadly misaligned llms
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31e68641-64d1-482b-96b7-34c604c2b306 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Wedded to prosperity? informal influence and regional favoritism
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f0016ce6-8359-43f6-85ea-e5843884496d · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2aa32b7-6cc5-4ae0-9e33-b4778e0f59a0 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models List of countries | Britannica
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 439e7a25-52f9-4974-a467-8996553ed81c · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models From Imitation to Introspection: Probing Self-Consciousness in Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 950eb059-ec69-4cd0-a671-ddaaba6b4324 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Probing linguistic information for logical inference in pre-trained language models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eefa3a7d-cae6-4863-a024-f84dc69db856 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Leveraging the Context through Multi-Round Interactions for Jailbreaking Attacks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe531612-2443-4d1a-aa94-e4a8fd5900df · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Breaking down the defenses: A comparative survey of attacks on large language models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01705d48-dbcf-4a46-bc10-2a5b4228527b · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models JailbreakRadar: Comprehensive Assessment of Jailbreak Attacks Against LLMs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 526a21f6-bf23-4ce8-80df-4186450bc962 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Scaling instruction-finetuned language models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f180b5d0-bf86-48b5-9d09-b5531dafd636 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Pawan Kumar, and Adel Bibi
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b092e29e-e050-4f37-a99e-6c79a296a0d8 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Dissecting recall of factual associations in auto-regressive language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06d3705b-e684-468a-9d09-dfba78d41227 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Estimating knowledge in large language models without generating a single token
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fae92c4-ed83-4aa5-b882-635b997b8616 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Not what you’ve signed up for: Compromising real-world llm-integrated applications with indirect prompt injection
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfa06f62-615f-40ca-9ed9-45d4fdd293ab · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Language models represent space and time
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 646a3c4b-2bd7-4830-8711-db5f719a9414 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models The Elements of Statistical Learning: Data Mining, Inference, and Prediction, volume 2
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3ac457d-fb07-4073-bd67-6fd10c97a6b4 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Do LLMs “know” internally when they follow instructions? In The Thirteenth International Conference on Learning Representations, 2025
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 759b8f54-cd1d-4747-854f-355b875cf830 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Linearity of relation decoding in transformer language models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8430ffe5-7061-4fc7-9eef-39105ff4c5cf · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c01ad35f-d372-4f9f-b395-75534335cdf1 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Refusal Tokens: A Simple Way to Calibrate Refusals in Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3336f87e-aea9-41c8-a87f-3a981ccda430 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cfda084d-86b9-4854-91d3-8374c95541d7 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language and vision-language models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bbaf33b-6b53-4065-8505-50fe0f78aa21 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Alignment of Language Agents
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21901eb9-72ba-4d6c-903f-2f333a024a6a · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Linear representations of political perspective emerge in large language models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7baf0ade-4903-4696-9913-52d67c38e2f1 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bcc200e2-fc7b-4a7c-aa80-f5ae16cb250e · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Revealing the Intrinsic Ethical Vulnerability of Aligned Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e8c6d6a0-da41-49f0-b24a-5a9ff6d632cb · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models The unlocking spell on base LLMs: Rethinking alignment via in-context learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1dabd6f2-50a8-44ae-8f45-764cd3e04cc0 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Towards Understanding Jailbreak Attacks in LLMs: A Representation Space Analysis
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74188ff8-dbf6-4b71-9309-9484a5e2fb27 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Formalizing and benchmarking prompt injection attacks and defenses
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 70fcb23a-e528-4821-9f89-888045f82745 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Keeping LLMs Aligned After Fine-tuning: The Crucial Role of Prompt Templates
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86438a36-a411-4441-a471-54d8bed4c87d · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models The geometry of truth: Emergent linear structure in large lan- guage model representations of true/false datasets
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6c34073d-86c8-4a1e-89c7-b423059d414c · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Occupation Data - O*NET 29.2 Data Dictionary at O*NET Re- source Center
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b186f49c-7811-4216-8e8b-7da273136424 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Training language models to follow instructions with human feedback
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3d0f1e2-31dd-4717-8934-f78e7e3e7546 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models The linear representation hypothesis and the geometry of large language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5096fc8d-d834-4727-8e4f-34e49c62aa1b · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Ignore Previous Prompt: Attack Techniques For Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a17f31f2-5cce-4416-a8d5-5f571d90d4fd · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Fine-tuning aligned language models compromises safety, even when users do not intend to! In The Twelfth International Conference on Learning Representations , 2024
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67379375-b768-4557-a664-7adce87c13e5 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Safety alignment should be made more than just a few tokens deep
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb4c71d6-3951-4f95-a5a5-ad814441f6bc · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Direct preference optimization: Your language model is secretly a reward model
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c158573-9d21-4878-b441-ee83b5fc65c0 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Multi- task prompted training enables zero-shot task generalization
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 74e33e9a-f10b-4f6d-88cb-9af4f587b0f0 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e459ab7-d98c-43af-9a8b-2b7768a272c4 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models do anything now
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ccd69e5-4941-4fa4-ac29-f2fcc1404203 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Measuring Free-Form Decision-Making Inconsistency of Language Models in Military Crisis Simulations
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 89b52122-a299-46f1-ac45-1adf19bcba0f · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Large Language Models are Inconsistent and Biased Evaluators
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a88fe132-d912-464b-adb0-575e52fb0356 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Gemma 2: Improving Open Language Models at a Practical Size
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 358b4c94-e01b-4aea-9bf5-e6661b3756af · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models What do you learn from context? Probing for sentence structure in contextualized word representations
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4edba58-579f-41fa-95c9-c325e80519d1 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Attention is all you need
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f93403d1-7fd8-425f-b47d-251dcb60d155 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models White-box multimodal jailbreaks against large vision-language models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8a797ac8-2e88-4f95-8733-6c1002f776bd · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Jailbroken: How does LLM safety training fail? In Thirty-seventh Conference on Neural Information Processing Systems, 2023
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e52b27ab-0ec8-40c8-8109-dde23d632199 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Assessing the Brittleness of Safety Alignment via Pruning and Low-Rank Modifications
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8564b34-057e-407c-bf81-6ed1135155b4 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8c3f6af-81dd-40ef-ab5a-21b0092acb6e · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Efficient streaming language models with attention sinks
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f17e0bc1-e48f-4aac-ab21-9c93e8167193 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Assessing Hidden Risks of LLMs: An Empirical Study on Robustness, Consistency, and Credibility
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a33b0e54-ad19-4232-89ae-fc87d627e05a · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models On the vulnerability of safety alignment in open-access LLMs
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 549a16a1-0b82-45bf-83b0-0e8a7a9e5e49 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Jailbreak Attacks and Defenses Against Large Language Models: A Survey
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9836f397-8f1c-4026-b996-27b1613c361a · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Yi: Open Foundation Models by 01.AI
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5857d7c6-e88b-4079-bdda-232f7a5b79d9 · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Don’t listen to me: understanding and exploring jailbreak prompts of large language models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf6791de-0c35-48c4-9af0-7d7f436b471c · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Removing RLHF protections in GPT-4 via fine-tuning
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dab3bda0-c3ff-45ee-b08c-c3cb7706f90b · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Lima: Less is more for alignment
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3b2af77-ec94-45f6-9f49-31afebf2d95b · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b418754-5462-49f6-8997-fc1ee1e73aef · outbound
Linearly Decoding Refused Knowledge in Aligned Language Models Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dbcb59b9-70ef-4595-91e4-d4fff01c1fd4 · inbound
HalluWorld: A Controlled Benchmark for Hallucination via Reference World Models Linearly Decoding Refused Knowledge in Aligned Language Models
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.