Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:51:10.182264Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 9 inbound Pith citation observations for arXiv:2505.23556.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:51:10.182264Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:29:10.162261Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
44 of 44 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation b87d6d4b-335f-42cc-89fa-210393a7eb9e · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6e93ff0-36ae-4b15-a944-4cb0f21bc61e · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95277383-8242-4f8a-8a42-fe60862e2112 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 211e4844-e50f-4bbe-b5b7-1874614661ca · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 98250a6e-29fd-4602-af8c-bc8cce841dab · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60c2fc83-26ef-4266-8811-a18de25e7d80 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7893635b-1fa0-4c09-933e-b8557ff57963 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42248491-c859-41ee-acdc-6b0213e6eaa7 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52bbbad2-7c50-4c86-a9d3-ffbcf6aa453b · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db5d23af-9292-46f2-9bad-d318f9e87674 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Training Verifiers to Solve Math Word Problems
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 720b51ec-f4e6-4112-b544-9b759b2dde22 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Sparse Autoencoders Find Highly Interpretable Features in Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e69b653-03dd-418e-8b0a-a8a0aa043636 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7a98e3a8-2596-4b24-b5d7-948c75dc3e0a · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d838ec9d-519d-4e6e-94bb-e453389c03a9 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af44740d-c367-4cd4-bd0c-892bcb47005e · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Scaling and evaluating sparse autoencoders
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27ce27a2-0f7e-4a85-a82d-5cf9a0b595d9 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7ede9f3-942e-4dcc-862d-13d3a3d14a42 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a5d4ce4-2a35-4e7e-baf5-e48bbee81ccb · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 992be6a2-f5d8-4e64-b887-ce83982a36a1 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation bff34266-a1b9-4506-b254-fdf81c950ee4 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation c2500d7e-2418-48ea-b28f-8c75f15e1b48 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a18718c-642b-4d93-bcfb-27f9fa7b5b9a · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 033b3d0b-835c-44c5-a1c8-01599096460c · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e2af9d0-3a49-4239-b8e5-d04a00236886 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders The Llama 3 Herd of Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4761377d-cdb4-496a-95a2-b73a41a33e38 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 223a07a0-86ca-477d-a6b4-eba8131902a5 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91ce86c1-de40-40d5-99bd-f2d73162279e · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfae3ae1-ab57-40ca-93d5-8687307d2ed6 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8952d920-f0c1-4d1f-ac33-4e7691c130e9 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Steering Llama 2 via Contrastive Activation Addition
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d96326dd-1e5d-4153-8685-ad9ca6cd69f6 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders The Linear Representation Hypothesis and the Geometry of Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0ef8e97-42b3-4272-bef4-f2062076c6e2 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation eff355b6-863d-42cd-b8c1-7f22534c5888 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c174971-d0af-4c72-ab45-36cb2ab447ed · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04d9c88d-664e-4bfe-a57c-52a9129ad98c · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 226d3341-e746-4397-9379-ed5d6c85a2ba · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Gemma 2: Improving Open Language Models at a Practical Size
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59442013-5b2c-4967-8498-697062796c5a · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e4dddc2-448c-4794-a271-5074f8258c15 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Steering Language Models With Activation Engineering
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e72b1cc8-7892-4934-85c7-36669fccc64d · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d993ff3-e203-472f-8ccb-c257eaccd477 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed9f3ccb-7d07-4760-9f1f-a91681384c04 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49d33019-2df4-4d84-90e6-ff7f169775eb · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Representation Engineering: A Top-Down Approach to AI Transparency
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92bebfaf-7a33-4af0-ad32-735e66115c58 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a4af4d7-32f3-4d9e-bd1b-88bb15045404 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders online" 'onlinestring :=
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb83ecf0-07e2-45d8-82ce-7769f4c81977 · outbound
Understanding Refusal in Language Models with Sparse Autoencoders write newline
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18905129-bc63-4288-a4c1-1fe82f33c3a1 · inbound
LLMs Encode Harmfulness and Refusal Separately Understanding Refusal in Language Models with Sparse Autoencoders
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21c0f68c-05f5-45de-8f2d-d8090664d7ac · inbound
Mitigating Jailbreaks with Intent-Aware LLMs Understanding Refusal in Language Models with Sparse Autoencoders
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f81b681-ec55-4df4-8846-35eb76f4d1f6 · inbound
Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs Understanding Refusal in Language Models with Sparse Autoencoders
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50fabb8e-bd0e-4261-a0c6-d006327a1a4b · inbound
Beyond "I cannot fulfill this request": Alleviating Rigid Rejection in LLMs via Label Enhancement Understanding Refusal in Language Models with Sparse Autoencoders
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6e7d65de-6cd4-4979-a5a7-d826d06cf823 · inbound
Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness Understanding Refusal in Language Models with Sparse Autoencoders
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 864ec7c0-b10e-4e34-93db-6fbbb1f58b78 · inbound
OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets Understanding Refusal in Language Models with Sparse Autoencoders
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation af26780e-7f3d-4c21-b934-219782f46033 · inbound
Faithfulness to Refusal: A Causal Audit of Neuron Selectors Understanding Refusal in Language Models with Sparse Autoencoders
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation adf1c4d8-ea18-4d7f-8102-bb6a5a62a925 · inbound
Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety Understanding Refusal in Language Models with Sparse Autoencoders
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df21f8a7-b267-456d-825b-74c3c73cd698 · inbound
Do LLMs Know Their Vulnerable Scenarios? Understanding Refusal in Language Models with Sparse Autoencoders
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.