Pith. sign in

Paper Citation Record · LEDGER

AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2501.17148.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.17148 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:27:20.702277Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7a9fc4e6-95b4-4fb6-a4ba-502bb7e25f20 · inbound

Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering cites this paper.

Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:30:22.282961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:30:22.282961Z digest=sha256:4b9b74d8355d8156e5d9e39131a9bdf2ab3a69e3f53b9dcbe723001ccf2e036b

Observation 31cea893-d830-4649-a4bb-38ba2a5f3b78 · inbound

Steering Large Language Models for Machine Translation Personalization cites this paper.

Steering Large Language Models for Machine Translation Personalization AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:51.732241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:01:51.732241Z digest=sha256:401c14c426013bf1adac99324e9e6cfb077a199e61d644e8f66abe985b2b5a81

Observation f67b55c0-fde2-4415-afdd-9906c7cd1c92 · inbound

Evaluating Steering Techniques using Human Similarity Judgments cites this paper.

Evaluating Steering Techniques using Human Similarity Judgments AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:58.857369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:58.857369Z digest=sha256:8a535c59b4bad471e4ea38d19fc570685cd0aafeca695465625c67b687d14f1c

Observation 0bbd5b28-16ef-422c-85fd-6ec1fde77f96 · inbound

Improved Representation Steering for Language Models cites this paper.

Improved Representation Steering for Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:33.503653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:33.503653Z digest=sha256:12d0ba9d35b5516039c7a5ebfd0baa43c3f0fa04654e555dcf08c69adb7307b6

Observation 23e4f887-4e7a-45ad-ad63-9c1688b592c9 · inbound

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race cites this paper.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.564287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.564287Z digest=sha256:7ef65c9f3dd3b255f8f41998d7a138224472239938c4d020880fd5cd927e765b

Observation a44a648d-3daf-4ca6-80dc-25d7bcf6f345 · inbound

Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models cites this paper.

Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:11.939328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:07:11.939328Z digest=sha256:7cbc50ea0016bb983348aaee32d19b98df0875ff695214f0c7dbb9158c559d3b

Observation de4cbd8d-1c7a-408d-9bbd-93ecc67f2526 · inbound

HyperSteer: Activation Steering at Scale with Hypernetworks cites this paper.

HyperSteer: Activation Steering at Scale with Hypernetworks AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:19.096812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:11:19.096812Z digest=sha256:b83b95de451e8e57a56de5774ffb53e01d1f2694f7f995f58c3ce2815732b27a

Observation 8f410453-9292-436c-a201-524a7700f897 · inbound

Fine-Grained Interpretation of Political Opinions in Large Language Models cites this paper.

Fine-Grained Interpretation of Political Opinions in Large Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:16.702496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:16.702496Z digest=sha256:c9d8480601b44b8a66e454a49af5100e4fb2aceccb4cb525ff8da5e993f1a020

Observation 48448f35-3dc6-4cb3-b0e2-fd117fefe50a · inbound

Resa: Transparent Reasoning Models via SAEs cites this paper.

Resa: Transparent Reasoning Models via SAEs AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:42.741450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:42.741450Z digest=sha256:45b4f7eadc45f188c29a20dcb9508c22b8519a142624da0c3191a8063e477d5a

Observation aa3cfd68-4461-42d7-9c7b-067819b1df5e · inbound

From Emergence to Control: Probing and Modulating Self-Reflection in Language Models cites this paper.

From Emergence to Control: Probing and Modulating Self-Reflection in Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T01:03:41.633267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:03:41.633267Z digest=sha256:ee5fc40949ccf61f4633aeaef029f8df01abc3b7c77a651b0a77f2097a1b40b8

Observation 686e0ed0-7129-44b5-add7-9373dae8b45e · inbound

Position: Use Sparse Autoencoders to Discover Unknowns cites this paper.

Position: Use Sparse Autoencoders to Discover Unknowns AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:33.631283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:33.631283Z digest=sha256:c508c4b21782bf4eb0bde94493c5c6f8120c796b419620e31e431e3822b58310

Observation ecec1d06-3b49-41fc-859c-0a010f994c08 · inbound

Insights into a radiology-specialised multimodal large language model with sparse autoencoders cites this paper.

Insights into a radiology-specialised multimodal large language model with sparse autoencoders AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:40:44.637794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:40:44.637794Z digest=sha256:6facc28cd5d9d111892a4cf8ca6a409206ebbeb72918c59bc8816e58435ff81e

Observation 5f90aaef-7103-4c7f-a616-dc6de6694890 · inbound

Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders cites this paper.

Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T17:18:55.749890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:18:55.749890Z digest=sha256:5a1774b1c1090fa093027fc5589973f2f75dfb03c61d8c881bab1a17924bbca7

Observation 3d6046e9-9891-47b7-b248-19a4ac46b9a4 · inbound

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation cites this paper.

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-05T18:47:45.341219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:47:45.341219Z digest=sha256:215fbe75ec45204b2a03fcc7627c67a1ef9375c9519c119bc1ad1fc02ddefe7a

Observation 64bae72e-6f2d-45e3-a86e-bcd48d11aca8 · inbound

Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA cites this paper.

Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T22:18:44.791907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:18:44.791907Z digest=sha256:828296f0f3c58b30b53dfed2865602127e1cec8966ac7a18d2c50b3c7c18f768

Observation fef9a84f-9bd2-402d-979e-78f556efb848 · inbound

VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models cites this paper.

VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:27.597133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:27.597133Z digest=sha256:d20a9d060f0db1916f6dab889b00678684dee57b169a8541a5b653b04fdae203

Observation 6a29067e-97b4-4317-a117-068745e3b2cc · inbound

The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail cites this paper.

The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T18:33:09.003812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:33:09.003812Z digest=sha256:9cbbf79739476518509a003b786c0617953ed6a4802e0b6217012bd288210337

Observation 25cc1439-9855-45b6-90e2-52c1d205dc7e · inbound

Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes cites this paper.

Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:31:10.587655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T11:42:16.090162Z digest=sha256:5dcf9a70cf4e357a04b0f72816835474748f32502233c56a688cadd104034d6b

Observation 999f74ea-7573-496a-b670-863bce46d67f · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:51:15.332661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T00:50:06.234399Z digest=sha256:7b025ab481f54568992b22645f54ed370ca8560317d53fbc5624583a5cfc5999

Observation c9542428-9d16-464c-9340-0a67098a26b5 · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:22:58.969437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:21:37.185232Z digest=sha256:bbfd6362fc2b65a2c12956a23daca494683531ad9f98b42127b0fca5f64392a2

Observation cbcdd5f3-9ca9-4d3e-921f-0895e37a543a · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:07:41.693739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T17:06:03.483247Z digest=sha256:e7130ca764cfb93a487dfd9e552417ffb6e2cd1f7db486672cde941d3292a4b8

Observation dc19f809-2c50-4736-a7d6-263c18fb155f · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:45:08.055382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T23:41:36.199099Z digest=sha256:6275b161a8c5d5963732c7baae785c2547225949fe110ae6490442658f852d91

Observation cf6380e4-61c1-493d-9226-07ebcd27f71b · inbound

The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime cites this paper.

The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:26.524548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:37:12.567351Z digest=sha256:8a6caab58cab5361558032afe2cfa8f78ad767a7ebfeb8fadd8cc44c576a6fc0

Observation 03000c2e-bb75-415a-8a01-5338990b0c2b · inbound

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space cites this paper.

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 109

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:22:18.867569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T05:17:34.283917Z digest=sha256:e1eb471cf962e41432204869fb50dba3192a13b17607bcdbec6f9d9f73241904

Observation 5273eeed-d53c-4b66-aa4e-adcbe3a4aacb · inbound

WriteSAE: Sparse Autoencoders for Recurrent State cites this paper.

WriteSAE: Sparse Autoencoders for Recurrent State AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:49:50.175124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T07:46:41.159688Z digest=sha256:d112e7848587c10410b5c214a192b81624439071b81d8d0b303550901055e391

Observation d5d1b455-8957-4939-afef-686d00f5552e · inbound

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search cites this paper.

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:39:10.652456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T22:36:12.466550Z digest=sha256:a4a16b0a2aa1fe6d6a8d07d9e7a9bfb0b8164a453947c696b565f410d669ebd8

Observation 420f821d-c186-40f2-8e77-c80e1bd8075a · inbound

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search cites this paper.

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.537374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T09:59:32.201925Z digest=sha256:0022ff21d2228fee42943db864fa57762d72a981f6d401635ab1dd772ec08fb8

Observation 037a966f-e60b-465f-aa09-3c95ed64ef6f · inbound

Are Sparse Autoencoder Benchmarks Reliable? cites this paper.

Are Sparse Autoencoder Benchmarks Reliable? AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:43:16.805970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T12:43:13.014365Z digest=sha256:84baa7709f071e2f8908dea24af2d2d79c8bb27079302e5cdfc03fe9e04c5574

Observation f63048b1-1502-4a22-91ff-0f9842ba83ec · inbound

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection cites this paper.

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:03:29.907469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T13:53:27.306664Z digest=sha256:da3e0733252b0b43d068f45c31c2dbc5b2be1bb6e9d0dc328a1f10545cd486a3

Observation 8e55ea6d-c6f2-4dd1-a1a5-9344dddd060a · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:47.199446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T22:16:47.743387Z digest=sha256:b2847c70ba5994af08c547e0844b3b555a49d5133abc4e577e71ff318ed6b9ff

Observation 1b40c27b-9c83-451e-bbbd-1341a9ee321f · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 117

Resolution
unresolved
no resolver link, observed 2026-07-12T17:03:44.315006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:03:44.315006Z digest=sha256:8a01783e0daaa0a5cf3ebd8da2e03ef01e5e316500f21df3c46354e0607864d1

Observation e3ba96c8-ca5b-4522-8973-6c0e25c6d87e · inbound

When is Your LLM Steerable? cites this paper.

When is Your LLM Steerable? AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:27:56.578484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T10:01:21.682285Z digest=sha256:ac8c5343e9130fdf71aef75ed6821017c34637daa4c1bd2c81519b4c3b0158a6

Observation e79bbe08-8e29-4e61-8e50-6f9acde5ce5b · inbound

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal cites this paper.

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 105

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:47.864621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T10:32:57.295159Z digest=sha256:a23c904d6f9e31ff1fd83e09659fb7fda15164caa34a8eef8a2c1d9849eefef2

Observation 71cc1d01-e169-4ee8-bfb5-708a7379f935 · inbound

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders cites this paper.

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:15:29.624121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T07:15:16.674714Z digest=sha256:9728be75197e26328f8991707bcef0764276abfe0aeed9ef2502d0672f3541b4

Observation 67b37175-3b56-4910-984d-8f416ca019f6 · inbound

Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent cites this paper.

Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T20:58:54.009938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:58:54.009938Z digest=sha256:e5061623b34b9e0bcdbe0dfcd1ef4b9ceba1b4d15c3e94c5c708f16901a91f63

Observation 608fb2c6-208f-4036-b5af-45253a520b6a · inbound

Probabilistic Concept-Aware Steering for Trustworthy LLM Inference cites this paper.

Probabilistic Concept-Aware Steering for Trustworthy LLM Inference AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-02T13:56:50.382262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:56:50.382262Z digest=sha256:9fa68a69c30b54cf47a3d55fa72ea8928c7f419ca944a5e52a658f72c4eeef96

Observation c2b9c053-81bb-4ef3-8471-255f8b679817 · inbound

ODRA: Synthesizing Cognitive Behavioral Therapy Sessions with Structured Chain-Of-Thought and Dynamic Patient Resistance cites this paper.

ODRA: Synthesizing Cognitive Behavioral Therapy Sessions with Structured Chain-Of-Thought and Dynamic Patient Resistance AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:06.016314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:56:06.016314Z digest=sha256:e3fec5a667b4275bb68ae4ef2378bba3353406f3ec909fdc7c8ec62c09c94ece

Observation 52a1fa14-295f-46b4-87ae-5d6cc4258b0b · inbound

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits cites this paper.

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T00:27:20.702277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:27:20.702277Z digest=sha256:aa480a1753c457ffc8e57d54c96555ca3c43bccd178c4d0d1cdafc4bd8a9e30d