Pith. sign in

Paper Citation Record · LEDGER

AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 38 inbound Pith citation observations for arXiv:2501.17148.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.17148 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 38 of 38 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:27:20.702277Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7a9fc4e6-95b4-4fb6-a4ba-502bb7e25f20 · inbound

Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering cites this paper.

Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:30:22.282961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:30:22.282961Z digest=sha256:4b9b74d8355d8156e5d9e39131a9bdf2ab3a69e3f53b9dcbe723001ccf2e036b

Observation 31cea893-d830-4649-a4bb-38ba2a5f3b78 · inbound

Steering Large Language Models for Machine Translation Personalization cites this paper.

Steering Large Language Models for Machine Translation Personalization AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:51.732241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:01:51.732241Z digest=sha256:91de5d99b3ba0e3e971a24e1388f892c31b4785a72111d89196d54579f757950

Observation f67b55c0-fde2-4415-afdd-9906c7cd1c92 · inbound

Evaluating Steering Techniques using Human Similarity Judgments cites this paper.

Evaluating Steering Techniques using Human Similarity Judgments AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:58.857369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:58.857369Z digest=sha256:8a535c59b4bad471e4ea38d19fc570685cd0aafeca695465625c67b687d14f1c

Observation 0bbd5b28-16ef-422c-85fd-6ec1fde77f96 · inbound

Improved Representation Steering for Language Models cites this paper.

Improved Representation Steering for Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:33.503653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:33.503653Z digest=sha256:12d0ba9d35b5516039c7a5ebfd0baa43c3f0fa04654e555dcf08c69adb7307b6

Observation 23e4f887-4e7a-45ad-ad63-9c1688b592c9 · inbound

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race cites this paper.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.564287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.564287Z digest=sha256:7ef65c9f3dd3b255f8f41998d7a138224472239938c4d020880fd5cd927e765b

Observation a44a648d-3daf-4ca6-80dc-25d7bcf6f345 · inbound

Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models cites this paper.

Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:11.939328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:07:11.939328Z digest=sha256:0844c9fdcac97e5eefbdc0c8de430ffcb1710b5bd1bdd3bfb748c6360113f4b6

Observation de4cbd8d-1c7a-408d-9bbd-93ecc67f2526 · inbound

HyperSteer: Activation Steering at Scale with Hypernetworks cites this paper.

HyperSteer: Activation Steering at Scale with Hypernetworks AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:19.096812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:11:19.096812Z digest=sha256:b83b95de451e8e57a56de5774ffb53e01d1f2694f7f995f58c3ce2815732b27a

Observation 8f410453-9292-436c-a201-524a7700f897 · inbound

Fine-Grained Interpretation of Political Opinions in Large Language Models cites this paper.

Fine-Grained Interpretation of Political Opinions in Large Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:16.702496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:16.702496Z digest=sha256:c9d8480601b44b8a66e454a49af5100e4fb2aceccb4cb525ff8da5e993f1a020

Observation 48448f35-3dc6-4cb3-b0e2-fd117fefe50a · inbound

Resa: Transparent Reasoning Models via SAEs cites this paper.

Resa: Transparent Reasoning Models via SAEs AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:42.741450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:42.741450Z digest=sha256:c4653844dab3c71c28c1e4791b4ba6c4b0db790944a69e6820e3ce0f6f6aa79d

Observation aa3cfd68-4461-42d7-9c7b-067819b1df5e · inbound

From Emergence to Control: Probing and Modulating Self-Reflection in Language Models cites this paper.

From Emergence to Control: Probing and Modulating Self-Reflection in Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T01:03:41.633267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:03:41.633267Z digest=sha256:55f236c15193af44dd3994695b26c1d147b1449559399a941f13aed591e4d578

Observation 686e0ed0-7129-44b5-add7-9373dae8b45e · inbound

Position: Use Sparse Autoencoders to Discover Unknowns cites this paper.

Position: Use Sparse Autoencoders to Discover Unknowns AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:33.631283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:33.631283Z digest=sha256:c508c4b21782bf4eb0bde94493c5c6f8120c796b419620e31e431e3822b58310

Observation ecec1d06-3b49-41fc-859c-0a010f994c08 · inbound

Insights into a radiology-specialised multimodal large language model with sparse autoencoders cites this paper.

Insights into a radiology-specialised multimodal large language model with sparse autoencoders AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:40:44.637794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:40:44.637794Z digest=sha256:0c0d41cff7c12de2856eefd0d77a3b4d1b6b770f6b380510011bdf7cbc6324bb

Observation 5f90aaef-7103-4c7f-a616-dc6de6694890 · inbound

Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders cites this paper.

Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T17:18:55.749890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:18:55.749890Z digest=sha256:5a1774b1c1090fa093027fc5589973f2f75dfb03c61d8c881bab1a17924bbca7

Observation 3d6046e9-9891-47b7-b248-19a4ac46b9a4 · inbound

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation cites this paper.

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-05T18:47:45.341219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:47:45.341219Z digest=sha256:215fbe75ec45204b2a03fcc7627c67a1ef9375c9519c119bc1ad1fc02ddefe7a

Observation 64bae72e-6f2d-45e3-a86e-bcd48d11aca8 · inbound

Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA cites this paper.

Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T22:18:44.791907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:18:44.791907Z digest=sha256:828296f0f3c58b30b53dfed2865602127e1cec8966ac7a18d2c50b3c7c18f768

Observation fef9a84f-9bd2-402d-979e-78f556efb848 · inbound

VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models cites this paper.

VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:27.597133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:27.597133Z digest=sha256:d20a9d060f0db1916f6dab889b00678684dee57b169a8541a5b653b04fdae203

Observation 6a29067e-97b4-4317-a117-068745e3b2cc · inbound

The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail cites this paper.

The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T18:33:09.003812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:33:09.003812Z digest=sha256:9cbbf79739476518509a003b786c0617953ed6a4802e0b6217012bd288210337

Observation 25cc1439-9855-45b6-90e2-52c1d205dc7e · inbound

Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes cites this paper.

Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:31:10.587655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T11:42:16.090162Z digest=sha256:671baca6dabff1a3ec217cc7de3c4df336d59629ca40cd89cff4d1e1271a4960

Observation 999f74ea-7573-496a-b670-863bce46d67f · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:51:15.332661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T00:50:06.234399Z digest=sha256:fc7e70d10781a0267e3da13a5d6c9d785c14eaefc72c3ad2887679197be5c144

Observation c9542428-9d16-464c-9340-0a67098a26b5 · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:22:58.969437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T21:21:37.185232Z digest=sha256:fb5b19f2ac23a9ad382817bb5743896a9432dcb183b8b11bfc9975afe1c34dc7

Observation cbcdd5f3-9ca9-4d3e-921f-0895e37a543a · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:07:41.693739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T17:06:03.483247Z digest=sha256:7c4a55052d6ed367f142d3cedc442bcd449c20babdb6d7faae7a6704f5ed3c02

Observation dc19f809-2c50-4736-a7d6-263c18fb155f · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:45:08.055382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T23:41:36.199099Z digest=sha256:ca904935e4bb2fd2c247c1e76354e12f4e7ac587b63b93ca5d921cdf34e3c9e9

Observation cf6380e4-61c1-493d-9226-07ebcd27f71b · inbound

The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime cites this paper.

The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:26.524548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:37:12.567351Z digest=sha256:4b4d6fd743222bc8b95495bd12ffce03f74859d153b133dbf70b570272601baa

Observation 03000c2e-bb75-415a-8a01-5338990b0c2b · inbound

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space cites this paper.

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 109

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:22:18.867569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T05:17:34.283917Z digest=sha256:3eceac22dd579a41a6661fdc8163705035768866d2d5a2980a74b2a51f1efb41

Observation 5273eeed-d53c-4b66-aa4e-adcbe3a4aacb · inbound

WriteSAE: Sparse Autoencoders for Recurrent State cites this paper.

WriteSAE: Sparse Autoencoders for Recurrent State AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:49:50.175124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T07:46:41.159688Z digest=sha256:ea046e092439cb1a46aa44e1ff5a368f36bfb7af19250f25854b07c75a1a9a12

Observation d5d1b455-8957-4939-afef-686d00f5552e · inbound

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search cites this paper.

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:39:10.652456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:36:12.466550Z digest=sha256:a23dc3d521b1e07abd884f30bc5f3193591176f596d2c7613641ebcee6aecffd

Observation 420f821d-c186-40f2-8e77-c80e1bd8075a · inbound

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search cites this paper.

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.537374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T09:59:32.201925Z digest=sha256:049916784c5213050e57295677b005764af92436a0ab80088976b90293b94bc9

Observation 037a966f-e60b-465f-aa09-3c95ed64ef6f · inbound

Are Sparse Autoencoder Benchmarks Reliable? cites this paper.

Are Sparse Autoencoder Benchmarks Reliable? AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:43:16.805970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T12:43:13.014365Z digest=sha256:efad9a67f5d4e3b96c9c5e5e025799b3321ecdae870d086f17b038696ebf9f9f

Observation f63048b1-1502-4a22-91ff-0f9842ba83ec · inbound

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection cites this paper.

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:03:29.907469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:53:27.306664Z digest=sha256:090ec09594536feba9688d6dc7a7b31e064ed94fdc4c2715e59d6ba786d31830

Observation 8e55ea6d-c6f2-4dd1-a1a5-9344dddd060a · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:47.199446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T22:16:47.743387Z digest=sha256:ebaba8c01fa2714ae5a6443265024a4a7e6f0b2085abf311ea50a126a949c420

Observation 1b40c27b-9c83-451e-bbbd-1341a9ee321f · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 117

Resolution
unresolved
no resolver link, observed 2026-07-12T17:03:44.315006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:03:44.315006Z digest=sha256:8a01783e0daaa0a5cf3ebd8da2e03ef01e5e316500f21df3c46354e0607864d1

Observation e3ba96c8-ca5b-4522-8973-6c0e25c6d87e · inbound

When is Your LLM Steerable? cites this paper.

When is Your LLM Steerable? AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:27:56.578484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:01:21.682285Z digest=sha256:06d6f73b7142672d75d6beab791198e80d509ffc69297cd9c16a28f28ad94747

Observation e79bbe08-8e29-4e61-8e50-6f9acde5ce5b · inbound

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal cites this paper.

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 105

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:47.864621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T10:32:57.295159Z digest=sha256:d1f149f6d1bb8d84d875acb62a29321dbe6270d591249c42e52d060bf5c460bc

Observation 71cc1d01-e169-4ee8-bfb5-708a7379f935 · inbound

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders cites this paper.

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:15:29.624121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T07:15:16.674714Z digest=sha256:ffeba5225f95773a117a2c5abcdf62e03c356b7856d7de48f1b577e0cd4da048

Observation 67b37175-3b56-4910-984d-8f416ca019f6 · inbound

Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent cites this paper.

Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T20:58:54.009938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:58:54.009938Z digest=sha256:e5061623b34b9e0bcdbe0dfcd1ef4b9ceba1b4d15c3e94c5c708f16901a91f63

Observation 608fb2c6-208f-4036-b5af-45253a520b6a · inbound

Probabilistic Concept-Aware Steering for Trustworthy LLM Inference cites this paper.

Probabilistic Concept-Aware Steering for Trustworthy LLM Inference AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-02T13:56:50.382262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:56:50.382262Z digest=sha256:9fa68a69c30b54cf47a3d55fa72ea8928c7f419ca944a5e52a658f72c4eeef96

Observation c2b9c053-81bb-4ef3-8471-255f8b679817 · inbound

ODRA: Synthesizing Cognitive Behavioral Therapy Sessions with Structured Chain-Of-Thought and Dynamic Patient Resistance cites this paper.

ODRA: Synthesizing Cognitive Behavioral Therapy Sessions with Structured Chain-Of-Thought and Dynamic Patient Resistance AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:06.016314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:56:06.016314Z digest=sha256:19564039651ec1891f7d76f5c2765325758e4c086df06b19050ba114f67808a4

Observation 52a1fa14-295f-46b4-87ae-5d6cc4258b0b · inbound

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits cites this paper.

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T00:27:20.702277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:27:20.702277Z digest=sha256:fb2ca66b2da4896b423216df6cf1e17394bc49080384b7f26f9d7d1980b1e042