Pith. sign in

Paper Citation Record · LEDGER

SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2503.09532.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.09532 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:25.837701Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:29:45.394907Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 82a4d355-2fc2-4888-a10d-f5e945759d3e · inbound

Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy cites this paper.

Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.837701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.837701Z digest=sha256:857f1fe1da2afc57a6d181a343a173a43deebc7fc16cd3a3ea958c558b110eb2

Observation 19903c6d-936d-43b0-9afe-7f2362cf3100 · inbound

Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures cites this paper.

Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:15.179663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:15.179663Z digest=sha256:61f66a8bd1bb2aced8b85fbd77f4ef4682bb1e7e67a7f90ba29ea1985ad10e52

Observation 620053fa-a4fe-44e0-a804-5c35444fc6ec · inbound

Resa: Transparent Reasoning Models via SAEs cites this paper.

Resa: Transparent Reasoning Models via SAEs SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:42.719642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:42.719642Z digest=sha256:fc2374e89759801d787c7b86ed07d50b282464cdada193c6771d5f3caab10e51

Observation 0a50bec3-c959-4b47-af7f-34a36a929314 · inbound

Evaluating SAE interpretability without explanations cites this paper.

Evaluating SAE interpretability without explanations SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.180244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:24:51.180244Z digest=sha256:91aaae6e6951b0b65c999b392dd5c69b06663fa96766336e0c1fd88cf3bbf892

Observation 4112d7ac-bf41-414e-88b3-355f5663fc65 · inbound

On the transferability of Sparse Autoencoders for interpreting compressed models cites this paper.

On the transferability of Sparse Autoencoders for interpreting compressed models SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:45.744009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:24:45.744009Z digest=sha256:ce4ba0ba390d82ec1e0a091847354552e9ea29e93420f3f21184106cc6d6886d

Observation 5c41650a-07f1-43c6-bec8-c9c5004543f8 · inbound

Distribution-Aware Feature Selection for SAEs cites this paper.

Distribution-Aware Feature Selection for SAEs SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:27:25.439860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:27:25.439860Z digest=sha256:4050039e8d30227ff75d3ebae7c6041dae2fd591fe08cba93a84ea1f479eb5b1

Observation 00bc17a8-ec58-433b-9b25-7291e4291f8c · inbound

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework cites this paper.

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:16:43.748207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T18:13:01.662828Z digest=sha256:a9ffcda168fed097789291790bd653a8806532a03340f53d7bbc79fb0f992476

Observation eb0d2f34-a1c1-4baa-98c7-71091dd817b0 · inbound

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models cites this paper.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.605878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:7aaebfd3e4fc798684172251eb6ce8a65967498b83c97f4a3780ee1fcea14461

Observation 1e7376d5-7454-4840-8813-48c7a2539620 · inbound

Stable and Steerable Sparse Autoencoders with Weight Regularization cites this paper.

Stable and Steerable Sparse Autoencoders with Weight Regularization SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T18:56:20.000242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:56:20.000242Z digest=sha256:e9d00efa37ea1f7caec856ca8a1d3949d7847977e507fce5306afcd35df2bea6

Observation f372237d-5f3a-434a-9a9c-253ec13b55e8 · inbound

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs cites this paper.

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:35:57.627704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:04:05.157103Z digest=sha256:6923afa1c14cc68b2ea63967624b906648a1556d795ba121d0b34de0da69c2a2

Observation 41ce46e6-8ce3-437d-9d3d-838688b499cf · inbound

Structural Instability of Feature Composition cites this paper.

Structural Instability of Feature Composition SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:41:36.519790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T06:40:00.507484Z digest=sha256:f8373cb15da57894506f45e8aa8c9fe7d5c6637c8ce2f6e1a5b1f4edf14eac87

Observation fc8f724e-4403-4fbb-96fc-6f84e4825e25 · inbound

From Token Lists to Graph Motifs: Weisfeiler-Lehman Analysis of Sparse Autoencoder Features cites this paper.

From Token Lists to Graph Motifs: Weisfeiler-Lehman Analysis of Sparse Autoencoder Features SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:11.060148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T09:41:01.775116Z digest=sha256:a7f556b002ad1f88a7cfdd33867daa40a47864cd99ef565bddcb930655e7b366

Observation 12da451f-70cb-4dca-a9f9-1ce293485063 · inbound

Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders cites this paper.

Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:15:54.303210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T03:13:58.543525Z digest=sha256:c6b33315494e7b917f36614362dd91219e0a7d61f80fd91d74a803623ac78d2a

Observation 793d19d5-c2a0-4e38-a5d5-813223894e8a · inbound

Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders cites this paper.

Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:25.445401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:35:50.776347Z digest=sha256:e412f320fbf252de920600a619a565c460c027c0ef0ded40fe148a09f3d5704b

Observation 47bdda45-8ee3-4658-8620-e9586d64adf1 · inbound

HH-SAE: Discovering and Steering Hierarchical Knowledge of Complex Manifolds cites this paper.

HH-SAE: Discovering and Steering Hierarchical Knowledge of Complex Manifolds SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:16:19.139898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:13:53.559096Z digest=sha256:b56bac808356a972b5598ed51c4abe08e9a97e27f51209708089d5cc2a56aaa4

Observation 6ec363f6-219e-438e-a626-a06e5bd5fef3 · inbound

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations cites this paper.

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:23:30.730641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T14:16:44.232080Z digest=sha256:96a9185ebc9a93622f137158cea6cfe398e60dd604b472c9a2d54315968d81b9

Observation a84392a0-db3d-4784-95e6-381b677202d3 · inbound

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations cites this paper.

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T05:02:51.380352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:02:51.380352Z digest=sha256:35b3041ff1773a29872d045a9300487c11b11c0958d18977f8dd53288d9d681a

Observation 4091bcb3-81ce-41b3-8e70-c13ca7a8f2df · inbound

A Unifying Framework for Concept-Based Representational Similarity cites this paper.

A Unifying Framework for Concept-Based Representational Similarity SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:29.987745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:10:36.855674Z digest=sha256:51221d827bbed96b20147fb581094c19065704977d3b8647890952324b8b6edc

Observation 2e6d96a4-bc02-44ab-b873-98d8a489405f · inbound

Do Sparse Autoencoders Learn Meaningful Concept Hierarchies? cites this paper.

Do Sparse Autoencoders Learn Meaningful Concept Hierarchies? SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:45.396552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:46:48.220801Z digest=sha256:bb326b1fde7fc9be37cfaa19884e4c09a25fcea0bda91963e5519a9e9df9d5d4

Observation 054b1070-4bea-4917-a1ad-4f0284afe438 · inbound

Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models cites this paper.

Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T19:02:29.759226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T19:02:29.759226Z digest=sha256:372175f8dabfacb35c2ab8bcf772626c45da8e468c481fa84fcab2c1b7b13682

Observation c7c1800e-f602-48f6-ab03-8968566103d0 · inbound

Decoder-Preserving Sparse Autoencoders: Which Readouts Survive Sparse Compression? cites this paper.

Decoder-Preserving Sparse Autoencoders: Which Readouts Survive Sparse Compression? SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:18.007688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:04:18.007688Z digest=sha256:7fdb817abdf0fa209fbd512ed875f38d005cfb831663844a9ce6abb8fcf13687

Observation b87c7a6c-e658-41fd-ac48-2b0833781747 · inbound

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects cites this paper.

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T10:03:56.301338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T10:03:56.301338Z digest=sha256:54f889ffebe9f950fc7ea73970700e05c89b48b0144af12302d2cd075ae35937

Observation 959c6a7f-12e1-46eb-8ee7-0a5a4c151dd9 · inbound

ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders cites this paper.

ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T07:58:32.246258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:58:32.246258Z digest=sha256:dd819a2fa73c5039f62e031101ad68e79b75eb5379bb6b54a641e43af03531e1