Pith. sign in

Paper Citation Record · LEDGER

Automatically Interpreting Millions of Features in Large Language Models

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 45 inbound Pith citation observations for arXiv:2410.13928.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.13928 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 45 of 45 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:21:24.356298Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

2
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 93ba4555-f892-4879-b9d5-5d231f5d255a · inbound

Interpretable Company Similarity with Sparse Autoencoders cites this paper.

Interpretable Company Similarity with Sparse Autoencoders Automatically Interpreting Millions of Features in Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T23:21:24.356298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T23:21:24.356298Z digest=sha256:77e82342c5baca560f006f8da6a49d331dfce17dc7d4e1879904b17838aef932

Observation bf2712e3-c2a5-4c81-826a-ad5abe3ff685 · inbound

Interpretable Steering of Large Language Models with Feature Guided Activation Additions cites this paper.

Interpretable Steering of Large Language Models with Feature Guided Activation Additions Automatically Interpreting Millions of Features in Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T19:34:26.933320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:34:26.933320Z digest=sha256:b626a0ab245878eac956d1cebb65ef4821c3352ba5e306da5a845e9d06439940

Observation 490207d3-85f9-47f7-8811-19206fda2a07 · inbound

Propositional Interpretability in Artificial Intelligence cites this paper.

Propositional Interpretability in Artificial Intelligence Automatically Interpreting Millions of Features in Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T14:02:12.790737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:02:12.790737Z digest=sha256:5165198684f40ea13fae40cd128b85b9eccc28313b599c651cc2d9d94963943f

Observation 5b5e109b-dc60-434f-b6d1-9e38a1c46758 · inbound

Sparse Autoencoders Trained on the Same Data Learn Different Features cites this paper.

Sparse Autoencoders Trained on the Same Data Learn Different Features Automatically Interpreting Millions of Features in Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T11:58:37.143219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T11:58:37.143219Z digest=sha256:f5a47d1ee3ba242e61ae43750e5ded28d4528f454baca3790cd808706f75524d

Observation d3ae2cc8-7ddf-476a-8d9e-5e36b9d259a3 · inbound

SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders cites this paper.

SAeUron: Interpretable Concept Unlearning in Diffusion Models with Sparse Autoencoders Automatically Interpreting Millions of Features in Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T01:04:57.206452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T01:04:57.206452Z digest=sha256:9868d0adc5c8305525c05f2cf92f90ee956bb737c54668b94e2f98165a015dba

Observation 8d3227a6-362c-4f27-9081-ab790f79be4d · inbound

Transcoders Beat Sparse Autoencoders for Interpretability cites this paper.

Transcoders Beat Sparse Autoencoders for Interpretability Automatically Interpreting Millions of Features in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T22:24:53.884852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T22:24:53.884852Z digest=sha256:cd5d3891ba9f4fc7e4db072f48bd554b9555ddab3ba22eafc1da0b8d2f41184b

Observation d1a4bfb8-1bce-41ae-bf01-96cc4d3259b4 · inbound

Partially Rewriting a Transformer in Natural Language cites this paper.

Partially Rewriting a Transformer in Natural Language Automatically Interpreting Millions of Features in Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T22:21:50.763556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T22:21:50.763556Z digest=sha256:dd3e20631c914a00f84790791dd1bd79fa31103f7c42d7afd2b51f8ac70152a7

Observation 7018977c-9838-49b4-badc-ba1620c36b4b · inbound

Converting MLPs into Polynomials in Closed Form cites this paper.

Converting MLPs into Polynomials in Closed Form Automatically Interpreting Millions of Features in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T16:58:44.732303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:58:44.732303Z digest=sha256:d760331dc2af4f3184e45480575368a626ce17fd57ec078d40e0439ce401dc88

Observation 404a1caf-9e0e-48b1-b0bc-f3deb3b9ee22 · inbound

Breaking Down Bias: On The Limits of Generalizable Pruning Strategies cites this paper.

Breaking Down Bias: On The Limits of Generalizable Pruning Strategies Automatically Interpreting Millions of Features in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T11:42:07.609217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T11:42:07.609217Z digest=sha256:40f47e17ef28e02f06d3613900dfb20306125f4e3c95a68a543a85f7d9242f35

Observation ec0bdb8c-202d-4452-97e6-e8969bf8b555 · inbound

Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models cites this paper.

Inference-Time Decomposition of Activations (ITDA): A Scalable Approach to Interpreting Large Language Models Automatically Interpreting Millions of Features in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:43.094342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:43.094342Z digest=sha256:e9cb11ca6caefc2bba580d752d6b6a088c81bf5a8bde870012c0e926919bfdad

Observation 69d5d226-67de-45f7-aa85-cd7999a7ea30 · inbound

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs cites this paper.

Position: Mechanistic Interpretability Should Prioritize Feature Consistency in SAEs Automatically Interpreting Millions of Features in Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:03:02.798784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:03:02.798784Z digest=sha256:904b68817dd96ef975e28ceba28a5057a55eafc27811ada57c6b00f4f9bfd870

Observation 0dced67a-930b-4899-906d-9d518eb3e32c · inbound

Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy cites this paper.

Train One Sparse Autoencoder Across Multiple Sparsity Budgets to Preserve Interpretability and Accuracy Automatically Interpreting Millions of Features in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.125107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.125107Z digest=sha256:65b712a09413784f99aa1f6ae3ec57da36f7a5b7f1c93438c6353b3596b35565

Observation 6e6b6711-37ba-499f-b663-95b831dca041 · inbound

Evaluating SAE interpretability without explanations cites this paper.

Evaluating SAE interpretability without explanations Automatically Interpreting Millions of Features in Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:51.640958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:24:51.640958Z digest=sha256:f654267eb2e2cbef329ea776fa4aa178032b5ec9921e3aa745863826122755da

Observation e3a332b2-4bc4-4fd7-a769-e59f5658fa1c · inbound

Insights into a radiology-specialised multimodal large language model with sparse autoencoders cites this paper.

Insights into a radiology-specialised multimodal large language model with sparse autoencoders Automatically Interpreting Millions of Features in Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T16:40:43.382708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:40:43.382708Z digest=sha256:d373eaa2df4fcbabf321a655c531111282b1eb730eb60e1ac69dafa8be06c95d

Observation cc944e05-a04d-4239-905f-0ffb17c18ff6 · inbound

Teach Old SAEs New Domain Tricks with Boosting cites this paper.

Teach Old SAEs New Domain Tricks with Boosting Automatically Interpreting Millions of Features in Large Language Models

Reference 1997

Resolution
unresolved
no resolver link, observed 2026-08-06T16:40:29.278577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:40:29.278577Z digest=sha256:a19419146e1e3f3bcc25c0c276fb3e928b9173513687b8e052986b3af45dc62c

Observation f5a8fa38-3718-4b12-9db2-5754e64ce507 · inbound

Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders cites this paper.

Model Directions, Not Words: Mechanistic Topic Models Using Sparse Autoencoders Automatically Interpreting Millions of Features in Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T11:01:42.861396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:01:42.861396Z digest=sha256:c397b99c6654f329e6d93fd20c726f4fff9e6832f3b46780c4c6e7b99b0d5170

Observation 38cada16-d398-4e83-9f60-d806293f9950 · inbound

Distribution-Aware Feature Selection for SAEs cites this paper.

Distribution-Aware Feature Selection for SAEs Automatically Interpreting Millions of Features in Large Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T14:27:25.448907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:27:25.448907Z digest=sha256:1703a9f957399f8ffb8828659addaad41e10556cf660b9c34be961cbd3347d13

Observation 51200a8a-a44b-4b9a-a1cf-f3c2243dc8dd · inbound

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework cites this paper.

Safe-SAIL: Towards a Fine-grained Safety Landscape of Large Language Models via Sparse Autoencoder Interpretation Framework Automatically Interpreting Millions of Features in Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:16:43.727392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-18T18:13:01.662828Z digest=sha256:0f7f5ea90842dcc30e4fd1481e7dfb865ea79a83e192775c2e25db29fc3616d7

Observation 8f99d151-6988-49f1-b0a0-644c1b031df4 · inbound

Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach cites this paper.

Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach Automatically Interpreting Millions of Features in Large Language Models

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:00:40.020715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-18T01:56:50.978054Z digest=sha256:7e8dead662df53387b887ad0b40dbe40c4ca96236283f821ae16b16d121fe238

Observation 80cfe4ec-b4c5-455a-9f9b-ca5f4da4545e · inbound

Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach cites this paper.

Making Interpretable Discoveries from Unstructured Data: A High-Dimensional Multiple Hypothesis Testing Approach Automatically Interpreting Millions of Features in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-04T00:23:29.455409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:23:29.455409Z digest=sha256:a0740443ae351973bafa2eadf34d0b1f9ffbd2f502305e6ec0d7f90678303b11

Observation 40fa6dbd-3e7a-44e4-9b18-20d62b4e15f5 · inbound

Prototype Transformer: Towards Language Model Architectures Interpretable by Design cites this paper.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Automatically Interpreting Millions of Features in Large Language Models

Reference 1995

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:44.908484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:44.908484Z digest=sha256:a33bd9ceb4d7d8f923dc818ec75dedc68f23f1590d43b4f799f3083c188067ab

Observation 24ca6d44-11fb-4f92-bdf2-1bd1d35942ab · inbound

Visual Persuasion: What Influences Decisions of Vision-Language Models? cites this paper.

Visual Persuasion: What Influences Decisions of Vision-Language Models? Automatically Interpreting Millions of Features in Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T23:00:26.827244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:00:26.827244Z digest=sha256:0ac413a9a87a99000d180d248096b1e1ea6d9a8b95df8ffb7b16e068aefd9717

Observation f0a86bbe-392a-4f0e-b54e-70c86ce35603 · inbound

CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs cites this paper.

CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs Automatically Interpreting Millions of Features in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T17:47:09.249674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T17:47:09.249674Z digest=sha256:2962302bff4e5780c6026e6270168845d1d52184ae1e9d5998954f9db6969052

Observation 99504919-6994-4918-bf36-9c0050b1a81a · inbound

MetaSAEs: Joint Training with a Decomposability Penalty Produces More Atomic Sparse Autoencoder Latents cites this paper.

MetaSAEs: Joint Training with a Decomposability Penalty Produces More Atomic Sparse Autoencoder Latents Automatically Interpreting Millions of Features in Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:43:11.174235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-13T19:42:57.683378Z digest=sha256:48e0c0afed3f6f697d3f0147e79a378c5fcb8d0a1af175a699cc393d86f3ac61

Observation 18442c76-67d5-4f2f-b0b5-2d67361ad47b · inbound

LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering cites this paper.

LangFIR: Discovering Sparse Language-Specific Features from Monolingual Data for Language Steering Automatically Interpreting Millions of Features in Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T05:36:47.912974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T05:36:47.912974Z digest=sha256:357f437851361c71393e9d39fab1c7cfbb5abde8728d3a3f83cce4523defb2b6

Observation 68a18cc7-8459-40cc-9acb-5163d2c730bb · inbound

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs cites this paper.

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs Automatically Interpreting Millions of Features in Large Language Models

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:35:57.738845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T18:04:05.157103Z digest=sha256:687edfb4e7feefd18562f6982e3ffc807208110616b3525f3864beac8f848e16

Observation e7e69111-eaf2-43e7-a12a-c31c0c074c66 · inbound

From Token Lists to Graph Motifs: Weisfeiler-Lehman Analysis of Sparse Autoencoder Features cites this paper.

From Token Lists to Graph Motifs: Weisfeiler-Lehman Analysis of Sparse Autoencoder Features Automatically Interpreting Millions of Features in Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:11.098765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T09:41:01.775116Z digest=sha256:345aee0e1b36f0b0765a57fca29b8e155c51e7a4a35116b7e4600d47a7b42da5

Observation c22a0357-cea8-4493-bb17-9119b92e3f35 · inbound

Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders cites this paper.

Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders Automatically Interpreting Millions of Features in Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:15:54.269344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-11T03:13:58.543525Z digest=sha256:7b92899419e4a2b042b809e386f750c89abe72f8e39719053a6b2ad4fdf4c2d2

Observation 9408900b-b8ff-4f83-aad1-f656f7fbcb0e · inbound

Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders cites this paper.

Tree SAE: Learning Hierarchical Feature Structures in Sparse Autoencoders Automatically Interpreting Millions of Features in Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:16:25.542255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T03:35:50.776347Z digest=sha256:a878fb0020b7db14b5670a6e6cb3d1bdbfd21d6ec28dee21c205cbe152891924

Observation 92412ab3-e893-4d21-a272-7edffd59c421 · inbound

Domain Restriction via Multi SAE Layer Transitions cites this paper.

Domain Restriction via Multi SAE Layer Transitions Automatically Interpreting Millions of Features in Large Language Models

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T06:02:23.786790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T05:54:21.707023Z digest=sha256:dc21c7ebab8864433f67829cb38ceb6017908792850319353dda79f8a10def51

Observation 5b619e30-0b39-4ad3-92ea-65ee8030c09b · inbound

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space cites this paper.

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space Automatically Interpreting Millions of Features in Large Language Models

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:27:19.296590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-13T05:17:34.283917Z digest=sha256:e8c126c6080b662d12856b8cdcc15cf3ae68c65265aefca3423849e68ebbad55

Observation 90c19cd7-8b28-4b13-9d10-29df29899b30 · inbound

Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features cites this paper.

Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features Automatically Interpreting Millions of Features in Large Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:29:27.959279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T20:27:38.363693Z digest=sha256:d8f3782c11828cb1668e40ad7b05792bc41b7aa0577aed6ea1e990c216f1904e

Observation 2023cc4a-ebf5-4238-84e6-b6d728fb9fbd · inbound

Why Retrieval-Augmented Generation Fails: A Graph Perspective cites this paper.

Why Retrieval-Augmented Generation Fails: A Graph Perspective Automatically Interpreting Millions of Features in Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:49:44.307236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-15T04:47:44.904830Z digest=sha256:c68398ebced70607d7928dc598c26b2f4500582f7e19ebac5d473955124069f1

Observation 4472d2ca-8d2f-417b-ba76-a821d63b5696 · inbound

The Rate-Distortion-Polysemanticity Tradeoff in SAEs cites this paper.

The Rate-Distortion-Polysemanticity Tradeoff in SAEs Automatically Interpreting Millions of Features in Large Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T21:15:04.067379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-30T21:13:01.880935Z digest=sha256:b78d44fd4345bb3a7deb7d6f1138ce222d761b0ff9e1914bdb9cad8d7bdd8b01

Observation 54eeb67b-f683-4846-b5c7-6f591391cb73 · inbound

Are Sparse Autoencoder Benchmarks Reliable? cites this paper.

Are Sparse Autoencoder Benchmarks Reliable? Automatically Interpreting Millions of Features in Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:43:16.763156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T12:43:13.014365Z digest=sha256:604889c392c997e0e8f274de822d878ecc6c69a7198098111c30cacfb0a22312

Observation e599bae1-869d-471e-af08-ac21f317a024 · inbound

Features have life history. And we should care cites this paper.

Features have life history. And we should care Automatically Interpreting Millions of Features in Large Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:33:50.874513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T23:31:30.350171Z digest=sha256:976f41d24f3abd9f918a300d1e33a532e958d58e258c19501c111a3d39649fa6

Observation 7ffd95b3-ce31-4606-902a-cf1f63dc9dc1 · inbound

SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization cites this paper.

SAEExplainer: Interpreting SAE Features with Activation-Guided Preference Optimization Automatically Interpreting Millions of Features in Large Language Models

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:57:25.986871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T18:35:47.513717Z digest=sha256:0027465e020a1e8442df5c80f083f6038a41f95ae2080e3331500a4b743aa8b0

Observation 59891a17-6ca1-43c0-a706-ff00c4581768 · inbound

Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders cites this paper.

Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders Automatically Interpreting Millions of Features in Large Language Models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:27:30.040455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T17:09:36.845217Z digest=sha256:3970461675b6ed145322ec94c31e2d8fb43a8a9e02076abbe6eda499d1fdfc97

Observation aca94aaf-78fb-4915-a67b-705996dd12b4 · inbound

ICA Lens: Interpreting Language Models Without Training Another Dictionary cites this paper.

ICA Lens: Interpreting Language Models Without Training Another Dictionary Automatically Interpreting Millions of Features in Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:17:49.074626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T10:21:58.878499Z digest=sha256:bea1839deb56bfacd55b4844cfdb7c099251aa3de705fc23f4118025bd15dca4

Observation b94cef44-8a3d-49b8-9a8b-b2329d44269a · inbound

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal cites this paper.

Anatomy of Post-Training: Using Interpretability to Characterize Data and Shape the Learning Signal Automatically Interpreting Millions of Features in Large Language Models

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:07:47.877710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T10:32:57.295159Z digest=sha256:936358c5c0472c7c0b1cf7be87d6a2272362fa417a3b08a1aac05786c6d10443

Observation be03c80f-db8d-48c6-9e78-80c0d70d54a6 · inbound

Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders cites this paper.

Extraction and Analysis of Multimodal Concepts in Vision Language Models through Sparse Autoencoders Automatically Interpreting Millions of Features in Large Language Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T14:29:31.008385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T14:22:34.708361Z digest=sha256:401b587cf38182cd11b4621a3a9e73ef1c1856f5b9d10df47b4c9f17efdbd78a

Observation f5c018c0-08fe-4128-a372-79e995f9b0fa · inbound

Do Sparse Autoencoders Learn Meaningful Concept Hierarchies? cites this paper.

Do Sparse Autoencoders Learn Meaningful Concept Hierarchies? Automatically Interpreting Millions of Features in Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:45.385912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T08:46:48.220801Z digest=sha256:509ece80cc6aacb6f6a88d80fef221865700cc4d7b6d73596daa4fa5dbb74c57

Observation e6cb9511-ab47-4020-be0c-38adefadde2f · inbound

Training, Reading, and Editing Legible Transformers cites this paper.

Training, Reading, and Editing Legible Transformers Automatically Interpreting Millions of Features in Large Language Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-13T05:35:58.568346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T05:35:58.568346Z digest=sha256:d857fb1b68a1c79a00de7954a652ae73f30385654a8bbe09002fdc72526e608f

Observation 378236f7-3528-43ed-b6df-9fe19f1c222f · inbound

Verbalizable Representations Form a Global Workspace in Language Models cites this paper.

Verbalizable Representations Form a Global Workspace in Language Models Automatically Interpreting Millions of Features in Large Language Models

Reference 139

Resolution
unresolved
no resolver link, observed 2026-08-01T23:15:30.760858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:15:30.760858Z digest=sha256:7d25e3938c76f5151984e2fdb8a8bcbd46b0dece85714cb6cb69f89e004cfaa6

Observation 77f466a6-23af-49c0-9269-a828189cd789 · inbound

ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels cites this paper.

ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels Automatically Interpreting Millions of Features in Large Language Models

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-05T17:48:46.321102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:48:46.321102Z digest=sha256:11092c6ef53c041e3cd9aae26ad4b5a575d78ee61b3d53182e629708ab71cba3