Pith. sign in

Paper Citation Record · LEDGER

SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2503.09532.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.09532 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:35:25.837701Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:29:45.394907Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 82a4d355-2fc2-4888-a10d-f5e945759d3e · inbound

Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy cites this paper.

Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:25.837701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:25.837701Z digest=sha256:54da3d4775f19ad8753916b49637b576f6f6a0cd1ec0bbc46da256e1d78ddc1c

Observation 19903c6d-936d-43b0-9afe-7f2362cf3100 · inbound

Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures cites this paper.

Incorporating Hierarchical Semantics in Sparse Autoencoder Architectures SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:56:15.179663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:56:15.179663Z digest=sha256:fd5c7946a42bf83d750dad7fffb2afbb1c476c6f2ddbcc607f954d212656696e

Observation 620053fa-a4fe-44e0-a804-5c35444fc6ec · inbound

Resa: Transparent Reasoning Models via SAEs cites this paper.

Resa: Transparent Reasoning Models via SAEs SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:42.719642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:42.719642Z digest=sha256:55622bf729eb507a1f7e1b9d1eb5fdd2e3c151cabc5dbc78ac7ab6905a5fcda5

Observation 0a50bec3-c959-4b47-af7f-34a36a929314 · inbound

Evaluating SAE interpretability without explanations cites this paper.

Evaluating SAE interpretability without explanations SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.180244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:24:51.180244Z digest=sha256:431d2ca71ac118fe2d213c82555353b9156d8e018008fcfbc24990f72297c428

Observation 4112d7ac-bf41-414e-88b3-355f5663fc65 · inbound

On the transferability of Sparse Autoencoders for interpreting compressed models cites this paper.

On the transferability of Sparse Autoencoders for interpreting compressed models SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:45.744009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:24:45.744009Z digest=sha256:d24ec2274ce1c61afa8fc8407cde162b25f85355a9256ac58514aa9e62be3e91

Observation 5c41650a-07f1-43c6-bec8-c9c5004543f8 · inbound

Distribution-Aware Feature Selection for SAEs cites this paper.

Distribution-Aware Feature Selection for SAEs SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:27:25.439860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:27:25.439860Z digest=sha256:007cba68b41da11f74dc90f7091969d21d77fb2b6bc33499371acfde8ff7e9a0

Observation 00bc17a8-ec58-433b-9b25-7291e4291f8c · inbound

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework cites this paper.

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:16:43.748207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T18:13:01.662828Z digest=sha256:9595fe15cede44b4486ab69bb7e578e44c5a571ce6a44bf27aafb0701b88a175

Observation eb0d2f34-a1c1-4baa-98c7-71091dd817b0 · inbound

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models cites this paper.

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:40:54.605878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T12:39:57.398423Z digest=sha256:5c62db85ae7aa77042cb41af6e7171e8e12224e6a70fad45494d76542bc71ff5

Observation 1e7376d5-7454-4840-8813-48c7a2539620 · inbound

Stable and Steerable Sparse Autoencoders with Weight Regularization cites this paper.

Stable and Steerable Sparse Autoencoders with Weight Regularization SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T18:56:20.000242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:56:20.000242Z digest=sha256:1d7b2f2c9d236e5828c825b85d80e54b721faef7653e169db5e5ae7c4fc46c7d

Observation f372237d-5f3a-434a-9a9c-253ec13b55e8 · inbound

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs cites this paper.

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:35:57.627704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:04:05.157103Z digest=sha256:18a103420c88deb8262ff5be3f12a667da47d330e61fff8eba427686e464d4db

Observation 41ce46e6-8ce3-437d-9d3d-838688b499cf · inbound

Structural Instability of Feature Composition cites this paper.

Structural Instability of Feature Composition SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:41:36.519790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T06:40:00.507484Z digest=sha256:b4f7d33fe322909f9a0c350bd0aa80ce002a8dca11f656677417cf117cffc442

Observation fc8f724e-4403-4fbb-96fc-6f84e4825e25 · inbound

From Token Lists to Graph Motifs: Weisfeiler-Lehman Analysis of Sparse Autoencoder Features cites this paper.

From Token Lists to Graph Motifs: Weisfeiler-Lehman Analysis of Sparse Autoencoder Features SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:11.060148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T09:41:01.775116Z digest=sha256:2acd942a7da4ae231e64ac3b22d3c0c973195f4aca7b14583cdfeb3fe63bf003

Observation 12da451f-70cb-4dca-a9f9-1ce293485063 · inbound

Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders cites this paper.

Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:15:54.303210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T03:13:58.543525Z digest=sha256:63b270cdcfdb98b8060f5c5e2052becf0f355f4464477c844a76517040a45386

Observation 793d19d5-c2a0-4e38-a5d5-813223894e8a · inbound

Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders cites this paper.

Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:25.445401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:35:50.776347Z digest=sha256:482b59cba9b5aba92708b357f9f2ee3c72c749a4e7510625f7ef59a127ae28be

Observation 47bdda45-8ee3-4658-8620-e9586d64adf1 · inbound

HH-SAE: Discovering and Steering Hierarchical Knowledge of Complex Manifolds cites this paper.

HH-SAE: Discovering and Steering Hierarchical Knowledge of Complex Manifolds SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:16:19.139898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:13:53.559096Z digest=sha256:d496dcfe732b45dd1ba56f48f9e6cade1726a54dc27901d272438f5bdcbf0364

Observation 6ec363f6-219e-438e-a626-a06e5bd5fef3 · inbound

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations cites this paper.

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:23:30.730641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T14:16:44.232080Z digest=sha256:2d407a97798e30ca4649f147ac1759936f7346ebc037773f82fe3977c80c7eaa

Observation a84392a0-db3d-4784-95e6-381b677202d3 · inbound

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations cites this paper.

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T05:02:51.380352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:02:51.380352Z digest=sha256:bf27c1de9b45d69b4bd8a22c3c74a5a834192a6246c548f5e78cf1c0e8373763

Observation 4091bcb3-81ce-41b3-8e70-c13ca7a8f2df · inbound

A Unifying Framework for Concept-Based Representational Similarity cites this paper.

A Unifying Framework for Concept-Based Representational Similarity SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:27:29.987745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T17:10:36.855674Z digest=sha256:6cdba0e8ab8ec609f61080aba7965f3ec7bb6c917f12ccb764acd24dc578455b

Observation 2e6d96a4-bc02-44ab-b873-98d8a489405f · inbound

Do Sparse Autoencoders Learn Meaningful Concept Hierarchies? cites this paper.

Do Sparse Autoencoders Learn Meaningful Concept Hierarchies? SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:45.396552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:46:48.220801Z digest=sha256:61e87fead7e7966079ddac8263fd7d04cb95c7f5fdc66c65abe1fc8b1c0bf639

Observation 054b1070-4bea-4917-a1ad-4f0284afe438 · inbound

Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models cites this paper.

Persistent Sparse Autoencoders: Learning Feature Timescales in Language Models SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T19:02:29.759226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T19:02:29.759226Z digest=sha256:abd04c09530652e456f30380f7129aacef0974317171e59de349c3c555f1723e

Observation c7c1800e-f602-48f6-ab03-8968566103d0 · inbound

Decoder-Preserving Sparse Autoencoders: Which Readouts Survive Sparse Compression? cites this paper.

Decoder-Preserving Sparse Autoencoders: Which Readouts Survive Sparse Compression? SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T18:04:18.007688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T18:04:18.007688Z digest=sha256:63534c5ee0539eed28e73b79c5429677088afae67e377692231532e7b08945ff

Observation b87c7a6c-e658-41fd-ac48-2b0833781747 · inbound

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects cites this paper.

Are Single-Token Sparse Autoencoder Features Causally Necessary? Layer-Depth and SAE-Family Effects SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T10:03:56.301338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T10:03:56.301338Z digest=sha256:c38b4b879bb300178fb3654c11375affea3157b78eb1714a2a1ee89e09449db9

Observation 959c6a7f-12e1-46eb-8ee7-0a5a4c151dd9 · inbound

ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders cites this paper.

ECG-InterpBench: Benchmarking the Interpretability of ECG Foundation Models with Matched-Scale Sparse Autoencoders SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T07:58:32.246258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:58:32.246258Z digest=sha256:bc54aad3c0f8a9ab88a849f3cd3a1cd5053a83be72d070deb0f63405f78226a4