Pith. sign in

Paper Citation Record · LEDGER

Understanding Refusal in Language Models with Sparse Autoencoders

As of 22 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 9 inbound Pith citation observations for arXiv:2505.23556.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23556 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:51:10.182264Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:29:10.162261Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation b87d6d4b-335f-42cc-89fa-210393a7eb9e · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:03.467398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:03.467398Z digest=sha256:8a03bd4f828868a8e551e9e733fed37c63e408e718542288ba420c11f60e51f3

Observation f6e93ff0-36ae-4b15-a944-4cb0f21bc61e · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:03.619692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:03.619692Z digest=sha256:94c577b04cf64ddf8e9b9676abc63c4233cac1dce6501dbbf941ed703c2d38bc

Observation 95277383-8242-4f8a-8a42-fe60862e2112 · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:03.761845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:03.761845Z digest=sha256:ff5677c53ff6b41e9d0643331fc8588b7e4068f9e3ef0926c958e2dc42cc79f4

Observation 211e4844-e50f-4bbe-b5b7-1874614661ca · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:51:11.715851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:51:03.935149Z digest=sha256:fcf16cc1168142c2decf33e779ae251c4e73b875e6170b00780713ecd1ccb8e4

Observation 98250a6e-29fd-4602-af8c-bc8cce841dab · outbound

This paper cites Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic.

Understanding Refusal in Language Models with Sparse Autoencoders Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:04.068736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:04.068736Z digest=sha256:43a33511fc88e3eaa4368aa6f5d800de0a73e4a362c9a0d4ab2604c57c3f865e

Observation 60c2fc83-26ef-4266-8811-a18de25e7d80 · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:04.194499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:04.194499Z digest=sha256:740f96c2c2ceb11d82fa33db553e5bbe6c9b14a9629f803531b4b2ca7e05588f

Observation 7893635b-1fa0-4c09-933e-b8557ff57963 · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:04.321203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:04.321203Z digest=sha256:a576b513a539d93769c1999abdf60b74ad6ca8ffd4551ca204eb62a0c19af7b7

Observation 42248491-c859-41ee-acdc-6b0213e6eaa7 · outbound

This paper cites JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models.

Understanding Refusal in Language Models with Sparse Autoencoders JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:04.473685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:04.473685Z digest=sha256:acbd7a988c59cb31a4a38da74dcea871a0ce76f168d84d5059c4a5f42f3eb7f6

Observation 52bbbad2-7c50-4c86-a9d3-ffbcf6aa453b · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Understanding Refusal in Language Models with Sparse Autoencoders Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:04.618404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:04.618404Z digest=sha256:6c9a22d9feaad8ada9b49480ad586c2caf4c4ebb1b18f81ccba9e9e39e0e3040

Observation db5d23af-9292-46f2-9bad-d318f9e87674 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Understanding Refusal in Language Models with Sparse Autoencoders Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:04.777831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:04.777831Z digest=sha256:d67c0feea6e51d92e0fc3e824d6da370502f74065b58ddde37eee6ada515caa9

Observation 720b51ec-f4e6-4112-b544-9b759b2dde22 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

Understanding Refusal in Language Models with Sparse Autoencoders Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:04.949417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:04.949417Z digest=sha256:40061c09de75b77485ccd85787e73125732b2dc3f8115a64000e7682cb221c94

Observation 4e69b653-03dd-418e-8b0a-a8a0aa043636 · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:51:11.560309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:51:05.112908Z digest=sha256:7527866c82357a7a8b5826181aef9d75fd4617f57dded3d506e51d2698e74d13

Observation 7a98e3a8-2596-4b24-b5d7-948c75dc3e0a · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:05.247504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:05.247504Z digest=sha256:1be8d43b9790669e8e5d00b27270b919865caeb24d689f6d7d3eb0553cceacc1

Observation d838ec9d-519d-4e6e-94bb-e453389c03a9 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Understanding Refusal in Language Models with Sparse Autoencoders The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:05.461723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:05.461723Z digest=sha256:4ae75b813a1e9509e4fcf4d83fe44bae4972c3b99513515d0503b3741b7387a2

Observation af44740d-c367-4cd4-bd0c-892bcb47005e · outbound

This paper cites Scaling and evaluating sparse autoencoders.

Understanding Refusal in Language Models with Sparse Autoencoders Scaling and evaluating sparse autoencoders

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:05.689549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:05.689549Z digest=sha256:f8bc3d077883fd2b0faadb2745f4214eb84111ca00bae4fb67498d7de3078038

Observation 27ce27a2-0f7e-4a85-a82d-5cf9a0b595d9 · outbound

This paper cites Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders.

Understanding Refusal in Language Models with Sparse Autoencoders Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:06.049313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:06.049313Z digest=sha256:458965569bf54501d4d06ae217d69d903fd8474c7158cffeefe1665be19e8ec3

Observation e7ede9f3-942e-4dcc-862d-13d3a3d14a42 · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:06.228204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:06.228204Z digest=sha256:c41c56065e3503e73026d740584cb8023822a9b558b54c3ef7b6f8ab12a489a4

Observation 7a5d4ce4-2a35-4e7e-baf5-e48bbee81ccb · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:51:11.386932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:51:06.398376Z digest=sha256:5737f8999a5787b0b39606385cf3768902d1d894d824568398f8d2ee9ee6030c

Observation 992be6a2-f5d8-4e64-b887-ce83982a36a1 · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:51:11.188031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:51:06.541363Z digest=sha256:1c52cbbb902b6a5b49e4f7671a56cc17ccfc3272d37096bfcc73c52ccac18b1e

Observation bff34266-a1b9-4506-b254-fdf81c950ee4 · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:51:10.937356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:51:06.689552Z digest=sha256:b56200398a3caf5d47f08c7e2a06a06f5ef0cd536de5ffbcbc012817a3519907

Observation c2500d7e-2418-48ea-b28f-8c75f15e1b48 · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:06.832925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:06.832925Z digest=sha256:339318d1f2a0ccf5a7d9b9250449aceacee64902fc0eaacd29611c48596c1f6f

Observation 1a18718c-642b-4d93-bcfb-27f9fa7b5b9a · outbound

This paper cites Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2.

Understanding Refusal in Language Models with Sparse Autoencoders Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:06.983803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:06.983803Z digest=sha256:588fbafd93de17f6fc607c01b4e98a558e0aa78d5e9b9038174e9a9a7d0d53a0

Observation 033b3d0b-835c-44c5-a1c8-01599096460c · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:07.151569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:07.151569Z digest=sha256:5c840f85097d2517dd466b8d3f3f2105691e630d3629023b22c5549eb373fdff

Observation 3e2af9d0-3a49-4239-b8e5-d04a00236886 · outbound

This paper cites The Llama 3 Herd of Models.

Understanding Refusal in Language Models with Sparse Autoencoders The Llama 3 Herd of Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:07.314336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:07.314336Z digest=sha256:10925553fab8f07321346b51c944b5d6ef4b21930818b0f5b5ee6d802f866c61

Observation 4761377d-cdb4-496a-95a2-b73a41a33e38 · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:07.476175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:07.476175Z digest=sha256:c26188ec0e2b1541d0eda1620d66ef9321bbca607c02f9d77ff23c0095ed7aa0

Observation 223a07a0-86ca-477d-a6b4-eba8131902a5 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

Understanding Refusal in Language Models with Sparse Autoencoders HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:07.596205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:07.596205Z digest=sha256:758a2bc6626641ee0afb05f20040855b1e12c1d6365ca61b5b081d68e5fd891d

Observation 91ce86c1-de40-40d5-99bd-f2d73162279e · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:07.748855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:07.748855Z digest=sha256:4b53e4285a8076081d8e15b8736771e312c38cbdd763c9b3f4af436b00b2bb6e

Observation cfae3ae1-ab57-40ca-93d5-8687307d2ed6 · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:07.888888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:07.888888Z digest=sha256:7b846cf767e9c10e927c660943e34f3c610ee665bc6020ef3971701f27312555

Observation 8952d920-f0c1-4d1f-ac33-4e7691c130e9 · outbound

This paper cites Steering Llama 2 via Contrastive Activation Addition.

Understanding Refusal in Language Models with Sparse Autoencoders Steering Llama 2 via Contrastive Activation Addition

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:08.004912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:08.004912Z digest=sha256:e451f9af2466462e8ee7ac2daba6e07c3bb60f8c59d8c6c3cb6bd6f280db9827

Observation d96326dd-1e5d-4153-8685-ad9ca6cd69f6 · outbound

This paper cites The Linear Representation Hypothesis and the Geometry of Large Language Models.

Understanding Refusal in Language Models with Sparse Autoencoders The Linear Representation Hypothesis and the Geometry of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:08.170319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:08.170319Z digest=sha256:ab3729dd27e642dc69da6743f72ae211d718591bd3ec475a97b4f38ba4809ed7

Observation f0ef8e97-42b3-4272-bef4-f2062076c6e2 · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:51:10.697885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T12:51:08.335859Z digest=sha256:f756cda8733210fd7abec0e019ee77ef3fac4f39edd2170b5011f09af70fed63

Observation eff355b6-863d-42cd-b8c1-7f22534c5888 · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:08.467896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:08.467896Z digest=sha256:f56bd288582af92d92f026afca3fa763eb0a8059f48ec9755c7eed56738259f0

Observation 8c174971-d0af-4c72-ab45-36cb2ab447ed · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:08.580910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:08.580910Z digest=sha256:53633f75f7c81e64e79c447b8dab55e46c1787b9ab1be590b15682302abf28a9

Observation 04d9c88d-664e-4bfe-a57c-52a9129ad98c · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:08.750101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:08.750101Z digest=sha256:886c4096d184c3665588bbc306c75f31ba52ffcfe321dfef06f2e4caf329793c

Observation 226d3341-e746-4397-9379-ed5d6c85a2ba · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Understanding Refusal in Language Models with Sparse Autoencoders Gemma 2: Improving Open Language Models at a Practical Size

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:08.933972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:08.933972Z digest=sha256:fdfa9146470429e1f7d3a875d659523a0ab03dcf1bb310b90bb7505a322bb08b

Observation 59442013-5b2c-4967-8498-697062796c5a · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Understanding Refusal in Language Models with Sparse Autoencoders Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:09.089836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:09.089836Z digest=sha256:654d54a24eadd777a7a28e417398de11a060216df0839949deaf1ea5b6ec7a6e

Observation 3e4dddc2-448c-4794-a271-5074f8258c15 · outbound

This paper cites Steering Language Models With Activation Engineering.

Understanding Refusal in Language Models with Sparse Autoencoders Steering Language Models With Activation Engineering

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:09.262929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:09.262929Z digest=sha256:ef4a5093fc4f167f44ed55141ffb3c489924d861c90657cc699e055abb0dc955

Observation e72b1cc8-7892-4934-85c7-36669fccc64d · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:09.416487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:09.416487Z digest=sha256:c5e7a19fbeeee6aa6d1eec16e79ff342a0591ef6edbc7e439a8d182702e99c35

Observation 3d993ff3-e203-472f-8ccb-c257eaccd477 · outbound

This paper cites an unresolved cited work.

Understanding Refusal in Language Models with Sparse Autoencoders Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:09.641665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:09.641665Z digest=sha256:f78fdf6f247c8ccbfdf34a8668d97cc0469d7ec64d79a56fa8608c027c9f78c4

Observation ed9f3ccb-7d07-4760-9f1f-a91681384c04 · outbound

This paper cites LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset.

Understanding Refusal in Language Models with Sparse Autoencoders LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:09.791159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:09.791159Z digest=sha256:b386e3e7b1a3e9f38ec8dce79373735bf5b7b2c85bfd29e2915b3826df1fd7fa

Observation 49d33019-2df4-4d84-90e6-ff7f169775eb · outbound

This paper cites Representation Engineering: A Top-Down Approach to AI Transparency.

Understanding Refusal in Language Models with Sparse Autoencoders Representation Engineering: A Top-Down Approach to AI Transparency

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:09.944567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:09.944567Z digest=sha256:c453e291d3d2c70a16746ac4894e458f19f269d2cae07470df60df0b29b52d8a

Observation 92bebfaf-7a33-4af0-ad32-735e66115c58 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Understanding Refusal in Language Models with Sparse Autoencoders Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:10.035803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:10.035803Z digest=sha256:1615f36403bfbfa782d4b57088bc548cbfc5b9727708e2b184aef54d951bfcbe

Observation 3a4af4d7-32f3-4d9e-bd1b-88bb15045404 · outbound

This paper cites online" 'onlinestring :=.

Understanding Refusal in Language Models with Sparse Autoencoders online" 'onlinestring :=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:10.093583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:10.093583Z digest=sha256:c3ce2880c88ab7e44ca2773aafd979740ec2b391a026cb77ad62877c3a0fd615

Observation eb83ecf0-07e2-45d8-82ce-7769f4c81977 · outbound

This paper cites write newline.

Understanding Refusal in Language Models with Sparse Autoencoders write newline

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:51:10.182264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:51:10.182264Z digest=sha256:5cf4b983c60b9280aad20ad544e02053793cf8a6dca6b3bda02067ad677296e5

Pith citing papers

Observation 18905129-bc63-4288-a4c1-1fe82f33c3a1 · inbound

LLMs Encode Harmfulness and Refusal Separately cites this paper.

LLMs Encode Harmfulness and Refusal Separately Understanding Refusal in Language Models with Sparse Autoencoders

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:03:16.984373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:03:16.984373Z digest=sha256:7c6a6b3e47bb4c60d86132e47b6449fab1406ad3cb287eb6fae11732a62c26b7

Observation 21c0f68c-05f5-45de-8f2d-d8090664d7ac · inbound

Mitigating Jailbreaks with Intent-Aware LLMs cites this paper.

Mitigating Jailbreaks with Intent-Aware LLMs Understanding Refusal in Language Models with Sparse Autoencoders

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T17:29:10.162261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:29:10.162261Z digest=sha256:c948dcbef19e99560c75a97cb834aacadb167d1a3bbe685dec73e407b06a3c51

Observation 4f81b681-ec55-4df4-8846-35eb76f4d1f6 · inbound

Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs cites this paper.

Poison Once, Refuse Forever: Weaponizing Alignment for Injecting Bias in LLMs Understanding Refusal in Language Models with Sparse Autoencoders

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-15T16:51:54.950442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:51:54.950442Z digest=sha256:a06cc97d04c9f5af5f8118feac58f505a653627377655de9a2525981f72c5d11

Observation 50fabb8e-bd0e-4261-a0c6-d006327a1a4b · inbound

Beyond "I cannot fulfill this request": Alleviating Rigid Rejection in LLMs via Label Enhancement cites this paper.

Beyond "I cannot fulfill this request": Alleviating Rigid Rejection in LLMs via Label Enhancement Understanding Refusal in Language Models with Sparse Autoencoders

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:15:55.709614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T01:53:10.245859Z digest=sha256:cf6314d672e979bc44731a098ed30db9e2f3b7bffda49a06b79985d6c71e49ed

Observation 6e7d65de-6cd4-4979-a5a7-d826d06cf823 · inbound

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness cites this paper.

Do Activation Monitors Survive Model Updates? Benchmarking, Predicting, and Repairing Activation-Monitor Staleness Understanding Refusal in Language Models with Sparse Autoencoders

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-27T03:30:26.978197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T03:24:24.714121Z digest=sha256:b96b87c92414a98d1c037ee506905b301d0aa33812354dbf72ee0374c614fe8a

Observation 864ec7c0-b10e-4e34-93db-6fbbb1f58b78 · inbound

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets cites this paper.

OpenSafeIntent: Evaluating Intent-Calibrated Safe Completion Across Dual-Use Prompt Sets Understanding Refusal in Language Models with Sparse Autoencoders

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:48:32.464906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-03T14:44:57.205766Z digest=sha256:9411c41b467ebea6b13e06eb26b3b4bcded0ce250c4b212bdf3402ae5f9b41f8

Observation af26780e-7f3d-4c21-b934-219782f46033 · inbound

Faithfulness to Refusal: A Causal Audit of Neuron Selectors cites this paper.

Faithfulness to Refusal: A Causal Audit of Neuron Selectors Understanding Refusal in Language Models with Sparse Autoencoders

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-07-07T15:43:53.886128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-07T15:35:54.665268Z digest=sha256:ec6c763061ba2c141b47056e1945a5d183f102987e49c2b35d9afc5eb7ba0853

Observation adf1c4d8-ea18-4d7f-8102-bb6a5a62a925 · inbound

Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety cites this paper.

Minionese: Comprehensive Benchmark and Mechanistic Study of Multilingual LLM Safety Understanding Refusal in Language Models with Sparse Autoencoders

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T14:13:53.489886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T14:13:53.489886Z digest=sha256:57b2bf6c11820df8a83286e707d3aa85335a3caf6c01b539ffd22b8da2d160e0

Observation df21f8a7-b267-456d-825b-74c3c73cd698 · inbound

Do LLMs Know Their Vulnerable Scenarios? cites this paper.

Do LLMs Know Their Vulnerable Scenarios? Understanding Refusal in Language Models with Sparse Autoencoders

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-30T20:47:41.079651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T20:47:41.079651Z digest=sha256:23c17b7b372188d98302e809c917cddfc2d3443ef6b9a94a84fdf94513a16d6b