Pith. sign in

Paper Citation Record · LEDGER

AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2501.17148.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.17148 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:36:58.518062Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c213a571-1f47-4cf7-bf24-0257b25e7655 · inbound

Toward universal steering and monitoring of AI models cites this paper.

Toward universal steering and monitoring of AI models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T01:06:10.147605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T01:06:10.147605Z digest=sha256:43440ad5019cd0214122efeffacbc4a5b78d4fe4e60d8d5d0029895c4f8fa908

Observation 7a9fc4e6-95b4-4fb6-a4ba-502bb7e25f20 · inbound

Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering cites this paper.

Denoising Concept Vectors with Sparse Autoencoders for Improved Language Model Steering AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:30:22.282961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:30:22.282961Z digest=sha256:df4b3bd3ce28c88056d293c4a73f12fc3e111e721d8eb82e7292a07bf610503b

Observation 31cea893-d830-4649-a4bb-38ba2a5f3b78 · inbound

Steering Large Language Models for Machine Translation Personalization cites this paper.

Steering Large Language Models for Machine Translation Personalization AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:01:51.732241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:01:51.732241Z digest=sha256:1e35ce92647f83c8cd13f552a1d018e67ef7b9a5524bd5b9fbde4728ff6102a9

Observation f67b55c0-fde2-4415-afdd-9906c7cd1c92 · inbound

Evaluating Steering Techniques using Human Similarity Judgments cites this paper.

Evaluating Steering Techniques using Human Similarity Judgments AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:20:58.857369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:20:58.857369Z digest=sha256:8411e584d90be2fd39ca39f7248007e0dcb1a296997476d897448b86aa50abdb

Observation 0bbd5b28-16ef-422c-85fd-6ec1fde77f96 · inbound

Improved Representation Steering for Language Models cites this paper.

Improved Representation Steering for Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:51:33.503653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:51:33.503653Z digest=sha256:99cfcdf5b61a27d9f101f78e9ceb4030ebd6080b90f76086fc754b323b09e522

Observation 23e4f887-4e7a-45ad-ad63-9c1688b592c9 · inbound

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race cites this paper.

Aligned but Blind: Alignment Increases Implicit Bias by Reducing Awareness of Race AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:13:19.564287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:13:19.564287Z digest=sha256:5d5b22850876ba07402d86e2ebc45da8a8df966dc04d3800063f227d8a3f78af

Observation a44a648d-3daf-4ca6-80dc-25d7bcf6f345 · inbound

Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models cites this paper.

Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:07:11.939328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:07:11.939328Z digest=sha256:50a14d5f21e0b7269a68020611de8b537ee30ea21aa5a390c895d96b77aa71ed

Observation de4cbd8d-1c7a-408d-9bbd-93ecc67f2526 · inbound

HyperSteer: Activation Steering at Scale with Hypernetworks cites this paper.

HyperSteer: Activation Steering at Scale with Hypernetworks AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:19.096812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:11:19.096812Z digest=sha256:bbed939c017b8f08e0f49fb770aa1e926c8ba272b24bd12fad44314bdb6806a8

Observation 8f410453-9292-436c-a201-524a7700f897 · inbound

Fine-Grained Interpretation of Political Opinions in Large Language Models cites this paper.

Fine-Grained Interpretation of Political Opinions in Large Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:40:16.702496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:40:16.702496Z digest=sha256:a840b1e4696a024221042b9e747956ef7b2939dd0806a0b705a68767a342a94f

Observation 48448f35-3dc6-4cb3-b0e2-fd117fefe50a · inbound

Resa: Transparent Reasoning Models via SAEs cites this paper.

Resa: Transparent Reasoning Models via SAEs AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:42.741450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:42.741450Z digest=sha256:36f8c78ed850e26edcb3382820bccd81fc70c016a421a7910254a4f64766ffe3

Observation aa3cfd68-4461-42d7-9c7b-067819b1df5e · inbound

From Emergence to Control: Probing and Modulating Self-Reflection in Language Models cites this paper.

From Emergence to Control: Probing and Modulating Self-Reflection in Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T01:03:41.633267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T01:03:41.633267Z digest=sha256:60d193fa64da89b849a454655026a5acb8380d88d80932c4450f9c7059c737a9

Observation 686e0ed0-7129-44b5-add7-9373dae8b45e · inbound

Position: Use Sparse Autoencoders to Discover Unknowns cites this paper.

Position: Use Sparse Autoencoders to Discover Unknowns AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:34:33.631283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:34:33.631283Z digest=sha256:be601b2afaa3ceeaaea15e14c4f83da96eadeaa1904a01247c6948a2a4b6828b

Observation ecec1d06-3b49-41fc-859c-0a010f994c08 · inbound

Insights into a radiology-specialised multimodal large language model with sparse autoencoders cites this paper.

Insights into a radiology-specialised multimodal large language model with sparse autoencoders AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:40:44.637794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:40:44.637794Z digest=sha256:1b0f1744c7ee87aa441f018ea6b5dac49c86e222f24848dd0ff4b784160b64ef

Observation 5f90aaef-7103-4c7f-a616-dc6de6694890 · inbound

Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders cites this paper.

Sparse but Wrong: Incorrect L0 Leads to Incorrect Features in Sparse Autoencoders AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T17:18:55.749890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:18:55.749890Z digest=sha256:0b9f55626c9b899e2def742a45bea6cc6394933394ab84ec5d738d035ca0fd17

Observation 3d6046e9-9891-47b7-b248-19a4ac46b9a4 · inbound

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation cites this paper.

Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-05T18:47:45.341219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:47:45.341219Z digest=sha256:2f521af58b731215df2d6e951e3ceffb1a7e3a73b384e1d52fd97d7878ac16cc

Observation 64bae72e-6f2d-45e3-a86e-bcd48d11aca8 · inbound

Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA cites this paper.

Investigating Symbolic Triggers of Hallucination in Gemma Models Across HaluEval and TruthfulQA AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T22:18:44.791907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:18:44.791907Z digest=sha256:3399eddc9e24feef550b161a8f5f0eba36445539a31a74664db8c18ecfab2338

Observation fef9a84f-9bd2-402d-979e-78f556efb848 · inbound

VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models cites this paper.

VISOR++: Universal Visual Inputs based Steering for Large Vision Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T13:45:27.597133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:45:27.597133Z digest=sha256:7f33c1bc691c6dee7399ecf67e66088225a2f40d5d78807937ca686c77e8839d

Observation 6a29067e-97b4-4317-a117-068745e3b2cc · inbound

The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail cites this paper.

The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T18:33:09.003812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:33:09.003812Z digest=sha256:c3a47aba2dfaf35f3caac1326ff8ceb9e2ab3c39374b3a44220084902073a5e2

Observation 25cc1439-9855-45b6-90e2-52c1d205dc7e · inbound

Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes cites this paper.

Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:31:10.587655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-08T11:42:16.090162Z digest=sha256:b31e302b9019f27ae5d5f7ae1f866049b1606b24b1710f2af88900552b9dd4d2

Observation 999f74ea-7573-496a-b670-863bce46d67f · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:51:15.332661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T00:50:06.234399Z digest=sha256:548e41c39f8b30c4f6bd047b9113a8b20c1cbbb1fa399ec40b8fb68e3d9ff248

Observation c9542428-9d16-464c-9340-0a67098a26b5 · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T21:22:58.969437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T21:21:37.185232Z digest=sha256:f4754c333c5b88ced35665b3bf44bcae8b47fdc90fccad7b2aa498e30eb58634

Observation cbcdd5f3-9ca9-4d3e-921f-0895e37a543a · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:07:41.693739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-19T17:06:03.483247Z digest=sha256:ab74ffb91b4b565bd597867148dd87aef85b5ffe7152ec19d7b30ca33dd6d7b0

Observation dc19f809-2c50-4736-a7d6-263c18fb155f · inbound

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models cites this paper.

When Language Overwrites Vision: Over-Alignment and Geometric Debiasing in Vision-Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:45:08.055382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T23:41:36.199099Z digest=sha256:ccebf7ac9d566d5c10c964dc2352cffa232e09688b2d1677a6b9ce26023addfd

Observation cf6380e4-61c1-493d-9226-07ebcd27f71b · inbound

The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime cites this paper.

The Open-Box Fallacy: Why AI Deployment Needs a Calibrated Verification Regime AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:11:26.524548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T03:37:12.567351Z digest=sha256:42747e8c08fddac96fd1f58c95a20695acffa2c9cf8384f3a4ef0dfe8b156790

Observation 03000c2e-bb75-415a-8a01-5338990b0c2b · inbound

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space cites this paper.

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 109

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:22:18.867569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-13T05:17:34.283917Z digest=sha256:49bbb08fe1ef098ef81711e4bc714f3cc871ad2a3e3cda786b58f3120f56b436

Observation 5273eeed-d53c-4b66-aa4e-adcbe3a4aacb · inbound

WriteSAE: Sparse Autoencoders for Recurrent State cites this paper.

WriteSAE: Sparse Autoencoders for Recurrent State AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:49:50.175124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T07:46:41.159688Z digest=sha256:2c5b1b5ceeeff6fa8f73e681dde5beb713cde5c8bc98960047aebe9539c2e9d2

Observation d5d1b455-8957-4939-afef-686d00f5552e · inbound

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search cites this paper.

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:39:10.652456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T22:36:12.466550Z digest=sha256:f2586b57895bf5bc6e06c17cc43aed314335fa4ecfb64cb9f7c39ccd364fa1d0

Observation 420f821d-c186-40f2-8e77-c80e1bd8075a · inbound

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search cites this paper.

When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:01:23.537374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-22T09:59:32.201925Z digest=sha256:6c430d30c6a4ea087a835cfe76778bdcf01b4919a6f703e2c84e2b44e67193bb

Observation 037a966f-e60b-465f-aa09-3c95ed64ef6f · inbound

Are Sparse Autoencoder Benchmarks Reliable? cites this paper.

Are Sparse Autoencoder Benchmarks Reliable? AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T12:43:16.805970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T12:43:13.014365Z digest=sha256:4ba8326e484688d0e88b0d7c3a643bf9d693789f69c86dd7eb6a93e109b830c0

Observation f63048b1-1502-4a22-91ff-0f9842ba83ec · inbound

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection cites this paper.

Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:03:29.907469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-29T13:53:27.306664Z digest=sha256:66ba7e5cde8371b88682c90ebf2053a967de811da97fc19408d7f3354421637e

Observation 8e55ea6d-c6f2-4dd1-a1a5-9344dddd060a · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:47.199446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T22:16:47.743387Z digest=sha256:1adbb6f0b4d89aceeeee300edc4b41c7576d26895cefb718e13971d8cb744b85

Observation 1b40c27b-9c83-451e-bbbd-1341a9ee321f · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 117

Resolution
unresolved
no resolver link, observed 2026-07-12T17:03:44.315006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:03:44.315006Z digest=sha256:bd585f629bf6f0d9ce4fbc28634d86198576d61e73eb4ade3c4f660606da957d

Observation e3ba96c8-ca5b-4522-8973-6c0e25c6d87e · inbound

When is Your LLM Steerable? cites this paper.

When is Your LLM Steerable? AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:27:56.578484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T10:01:21.682285Z digest=sha256:5702083bf61c627ccc71566c8afb2b93f041ee493a77c029df9c5371a0ac3893

Observation e79bbe08-8e29-4e61-8e50-6f9acde5ce5b · inbound

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal cites this paper.

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 105

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:47.864621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T10:32:57.295159Z digest=sha256:02e52d49d8c54353de50df5e667b7bdae4d561618a15d56681a82193b0d38bfe

Observation 71cc1d01-e169-4ee8-bfb5-708a7379f935 · inbound

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders cites this paper.

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:15:29.624121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-07-01T07:15:16.674714Z digest=sha256:8d3352a39cd3a7ab91f7dc29f02230ea32b290558bec151901fe38ee24c2fd9f

Observation 67b37175-3b56-4910-984d-8f416ca019f6 · inbound

Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent cites this paper.

Retrieval is Enough: Training-Free Interpretability with a Tool-Using Agent AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T20:58:54.009938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T20:58:54.009938Z digest=sha256:b6a00adbf26d98e63e8aba6388f9b2429e457595bd8967f4feb8309db77635d2

Observation 608fb2c6-208f-4036-b5af-45253a520b6a · inbound

Probabilistic Concept-Aware Steering for Trustworthy LLM Inference cites this paper.

Probabilistic Concept-Aware Steering for Trustworthy LLM Inference AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-02T13:56:50.382262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:56:50.382262Z digest=sha256:731a9a9483398360707bc331b76ab281fe7eed80643f2782a5bfdddfbfe0c6e6

Observation c2b9c053-81bb-4ef3-8471-255f8b679817 · inbound

ODRA: Synthesizing Cognitive Behavioral Therapy Sessions with Structured Chain-Of-Thought and Dynamic Patient Resistance cites this paper.

ODRA: Synthesizing Cognitive Behavioral Therapy Sessions with Structured Chain-Of-Thought and Dynamic Patient Resistance AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:06.016314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:56:06.016314Z digest=sha256:4e1d77a520f5f9431c036c876d4f7f2caebb98fd0c54ab4cd9b36230ed1d1b46

Observation 52a1fa14-295f-46b4-87ae-5d6cc4258b0b · inbound

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits cites this paper.

CircuitSteer: Geometrically Aligned Multi-Layer Steering via Sparse Autoencoder Circuits AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T00:27:20.702277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:27:20.702277Z digest=sha256:c33183213cfdbc2d6096525f5493a7b12ba2a78efb17e80426090532772a7881

Observation 3bc2e44e-d022-4b63-a352-5777f5fe446c · inbound

Scaling Inherently Interpretable Language Models cites this paper.

Scaling Inherently Interpretable Language Models AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T00:36:58.518062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:36:58.518062Z digest=sha256:a166c846eb4d5222e186126b8598e3e61b9d7c6465a4f73a461d7bf0f8a096eb