Pith. sign in

Paper Citation Record · LEDGER

Steering Language Model Refusal with Sparse Autoencoders

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2411.11296.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.11296 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:03:01.923666Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8971053d-2e39-4d5b-9596-5976a58bb55b · inbound

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs cites this paper.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Steering Language Model Refusal with Sparse Autoencoders

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:01.923666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:01.923666Z digest=sha256:a8ac697fcab1540fa48fed52711b6074e5443b4cd8148aaa3e0f6f1e779cbce6

Observation db905e8f-0f2c-49fb-9f55-f33b8a67a59d · inbound

Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures cites this paper.

Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures Steering Language Model Refusal with Sparse Autoencoders

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:15.218560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:15.218560Z digest=sha256:cd3717c29a70ae8b80e54beeada52898016485531b3042b703728997d9d5283b

Observation cbeb1e1e-31e7-498f-af5f-ca4aa00efc3b · inbound

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety cites this paper.

Interpretation Meets Safety: A Survey on Interpretation Methods and Tools for Improving LLM Safety Steering Language Model Refusal with Sparse Autoencoders

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:26.435845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:26.435845Z digest=sha256:68661a635f38c3cf19126096bed1a82310f59149d4a7292722205971a6b59a2f

Observation 079557b3-ec61-4220-b289-83d8193b0886 · inbound

Resa: Transparent Reasoning Models via SAEs cites this paper.

Resa: Transparent Reasoning Models via SAEs Steering Language Model Refusal with Sparse Autoencoders

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:42.725469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:42.725469Z digest=sha256:c93fc60b4ffe372ef14205952214ebb60436c7b0c75671bd3b7e9600e7e14e4c

Observation 48c18206-9944-4304-8f43-0c8bd96911c9 · inbound

Insights into a radiology-specialised multimodal large language model with sparse autoencoders cites this paper.

Insights into a radiology-specialised multimodal large language model with sparse autoencoders Steering Language Model Refusal with Sparse Autoencoders

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T16:40:42.896985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:40:42.896985Z digest=sha256:69df252dee31584096b7891acbed6991c1bd943d4e4892e50ae8930d6555c803

Observation 043bb4b5-df6b-48dc-923e-a044138069e9 · inbound

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM cites this paper.

Guardians and Offenders: A Survey on Harmful Content Generation and Safety Mitigation of LLM Steering Language Model Refusal with Sparse Autoencoders

Reference 231

Resolution
unresolved
no resolver link, observed 2026-08-05T23:13:05.310280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:13:05.310280Z digest=sha256:f3653eafb2b32a9d46e4b1bec2ae26377cd8373bfd7a5640c67d887f97156ea1

Observation 77d84fa0-46e6-4bbc-826a-a011a1d5b69e · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Steering Language Model Refusal with Sparse Autoencoders

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:31.454509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:31.454509Z digest=sha256:3e9bd98ce1cb1b1dc6e78a883d496586c36eb03bb6b0dc9adb58c40d59c3d957

Observation ce18123b-8d33-4f35-90a6-151542840ddf · inbound

SATORI: Static Test Oracle Generation for REST APIs cites this paper.

SATORI: Static Test Oracle Generation for REST APIs Steering Language Model Refusal with Sparse Autoencoders

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T17:26:48.647339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:26:48.647339Z digest=sha256:18eff1672f099cc76bf92d662b54cd9209c9fe72c1a09b60bdf3396989b65e4a

Observation cf2db5bc-92c8-4a3e-aba5-1c89c5bf2c6d · inbound

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection cites this paper.

Turning the Spell Around: Lightweight Alignment Amplification via Rank-One Safety Injection Steering Language Model Refusal with Sparse Autoencoders

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T14:57:12.882025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:57:12.882025Z digest=sha256:17dd98fb325d6f2dbe3c7fc3291bed13832686d942b449f27cc828ada93d938f

Observation bc858d3f-411b-473a-9127-b5651c15cc71 · inbound

Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal cites this paper.

Beyond I'm Sorry, I Can't: Dissecting Large Language Model Refusal Steering Language Model Refusal with Sparse Autoencoders

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:56:45.801431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-18T18:56:13.680353Z digest=sha256:0d88bc21d89610f95ba8ec9577ef15f6b7e3e6259d5c1c91e51c5cd32bf11199

Observation f5868bf6-2a9a-49a2-8707-729707477d46 · inbound

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook cites this paper.

The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook Steering Language Model Refusal with Sparse Autoencoders

Reference 155

Resolution
unresolved
no resolver link, observed 2026-07-13T14:03:01.974171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T14:03:01.974171Z digest=sha256:ca5726ff9a2f1172618530d99def6f0ae8fce68571f4df82856db3a7590b440f

Observation 9060515b-c19f-4604-a3ca-b523abf2b4f4 · inbound

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs cites this paper.

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs Steering Language Model Refusal with Sparse Autoencoders

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:35:57.702617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:04:05.157103Z digest=sha256:727b1ad2690d9a41e1ca604caf3b67f0dca8124c85879717e1ccb5b73e36d363

Observation bb0164bb-25f3-4847-a5b2-40670026e08e · inbound

Towards Understanding the Robustness of Sparse Autoencoders cites this paper.

Towards Understanding the Robustness of Sparse Autoencoders Steering Language Model Refusal with Sparse Autoencoders

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:46:10.274555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:44:24.532548Z digest=sha256:52db58056b42ddfdcd0299d8b0480d3d3d6e1e9f294f861ecc39a446d9ca7fd6

Observation adfbfdad-b81c-4b55-b90d-9162cfa2bd44 · inbound

Estimating Tail Risks in Language Model Output Distributions cites this paper.

Estimating Tail Risks in Language Model Output Distributions Steering Language Model Refusal with Sparse Autoencoders

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:08.127131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T12:26:41.797166Z digest=sha256:a0c00ce06c9c782aa4f7d2282c3fa84efe48795d8576d116d4d63d3919dc0bb9

Observation 3ca12e34-41b2-461b-ba5e-91aa0972adeb · inbound

Don't Lose Focus: Activation Steering via Key-Orthogonal Projections cites this paper.

Don't Lose Focus: Activation Steering via Key-Orthogonal Projections Steering Language Model Refusal with Sparse Autoencoders

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:06:08.788185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T10:23:22.394860Z digest=sha256:c5125f476ac28d8d9002b6cc18e58dc9d5e862d0f64cbb62e5526de6683fc3af

Observation 3a5c6831-ecd7-40a7-a79a-4013d95f35ef · inbound

LLM Advertisement based on Neuron Auctions cites this paper.

LLM Advertisement based on Neuron Auctions Steering Language Model Refusal with Sparse Autoencoders

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:51:14.799366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:50:36.096751Z digest=sha256:2c227818a613bd9ed4da1428e869c15c055d799c74dcabcd6864cdf3e916bb84

Observation a00562e3-4f2b-4052-b15e-f2f31e683351 · inbound

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search cites this paper.

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search Steering Language Model Refusal with Sparse Autoencoders

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:39:10.658984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T22:36:12.466550Z digest=sha256:cb225477d8d241e021561fb7e42ba796323c913a2139606abc95b738fd2ebb68

Observation 5ea39289-0c37-4c0c-8513-f81a05d3012e · inbound

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search cites this paper.

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search Steering Language Model Refusal with Sparse Autoencoders

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.543441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T09:59:32.201925Z digest=sha256:c8aaf1a39eaf6806d282a150a5d288d14cd70b475dded19cd7170e69e2f4aaa6

Observation ec170ce0-4d8b-4f1b-8070-d72265cd1409 · inbound

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications cites this paper.

Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications Steering Language Model Refusal with Sparse Autoencoders

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T23:32:52.613253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T23:30:43.364230Z digest=sha256:970d74a3f5aba5c89a1a9ceb0f0d0859a3ecc416d1ec65de6e0265575d3f6ce3

Observation b9b04e5e-0d17-4029-be26-b96cc53970ed · inbound

Steered Generation via Gradient-Based Optimization on Sparse Query Features cites this paper.

Steered Generation via Gradient-Based Optimization on Sparse Query Features Steering Language Model Refusal with Sparse Autoencoders

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:40.364686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:31:29.510639Z digest=sha256:4a19b924e71cc3e850515ded8d3c2a2d305ac3ffe7f989b26c6e2febd80d1677

Observation afcfa71d-630e-48f5-afcc-def0ff4f85ff · inbound

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs cites this paper.

Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs Steering Language Model Refusal with Sparse Autoencoders

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-06-30T16:04:52.626040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T16:03:12.728352Z digest=sha256:f96b3b224ea31995c336c86313e19c5278135f89623e6b0a37bd7a3f6a8303bf

Observation b0dcb2c8-50b8-4dca-b7f6-0fad16a9eb44 · inbound

Perplexity Can Miss SAE Feature Damage Under Quantization cites this paper.

Perplexity Can Miss SAE Feature Damage Under Quantization Steering Language Model Refusal with Sparse Autoencoders

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:36:25.786048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T11:41:18.460538Z digest=sha256:d9b54447026f9b093ed09292465cfd3c62e3cb5e168393f0c779d5f2239f54fa

Observation 2a37f71a-0e2a-40e3-9b29-4b3a68983bb5 · inbound

Pre-Intervention Prediction of Sparse Autoencoder Steering Side Effects cites this paper.

Pre-Intervention Prediction of Sparse Autoencoder Steering Side Effects Steering Language Model Refusal with Sparse Autoencoders

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.474864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T19:52:12.114493Z digest=sha256:e0028491b833e84e120a2082b26ca2053fe229e614a4d9c9b7531bfca017927d

Observation de07c086-5ed4-4d00-9a96-fb9b97482d67 · inbound

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization cites this paper.

At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization Steering Language Model Refusal with Sparse Autoencoders

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-26T01:28:50.536905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T01:27:39.812228Z digest=sha256:cf5345fa93f70b2e1b143e72cd138c92ecb2709648495a914840c58655937db7

Observation c1c16bb2-90b4-42f0-bfac-0c46ae57d38a · inbound

Forecasting With LLMs: Improved Generalization Through Feature Steering cites this paper.

Forecasting With LLMs: Improved Generalization Through Feature Steering Steering Language Model Refusal with Sparse Autoencoders

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:59:52.520253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T04:42:01.073777Z digest=sha256:db1d31bb513518e34a31f21c6af1e8425bcb84f9430c870d1b79e4b4a6232e94