Pith. sign in

Paper Citation Record · LEDGER

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability

As of 23 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2606.18383.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.18383 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T01:27:56.855130Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch9

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 89eeac66-55ea-4a5c-a881-22a886e1167a · outbound

This paper cites Stronger generalization bounds for deep nets via a compression approach.

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability Stronger generalization bounds for deep nets via a compression approach

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:18:56.766065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:27:56.855130Z digest=sha256:20578154b55e4808743c8f043b6ff736e55d716b6aaf3c761207c04cfe5ed648

Observation b3b37222-46dd-459d-966b-cc9f53167cb8 · outbound

This paper cites PIQA: Reasoning about Physical Commonsense in Natural Language.

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability PIQA: Reasoning about Physical Commonsense in Natural Language

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:18:56.761844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:27:56.855130Z digest=sha256:c7fa0c4415d515a64c3aa1fbb2c3c8b616d019618c33b7a6f69e7428f77f3e91

Observation 63153587-eed2-4108-9ade-f962d7f6d679 · outbound

This paper cites Information Processing Let- ters 24, 377–380.

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability Information Processing Let- ters 24, 377–380

Reference 3

Resolution
metadata mismatch
doi, observed 2026-06-27T01:30:20.471021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:27:56.855130Z digest=sha256:e2bca71a9493d0c08d19665c7aaeabde2e58ef5ddbbb75876af5dc192e7bd279

Observation 016c58c8-7493-4db0-b855-bfafed6d5424 · outbound

This paper cites Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L.

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability Cunningham, H., Ewart, A., Riggs, L., Huben, R., and Sharkey, L

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T01:27:56.855130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:27:56.855130Z digest=sha256:d766df536752e09b36b092d662be1c6341daf810a343fb81a316296660568a8f

Observation c273ea77-c7e6-4d28-a8dd-06b0ca044a85 · outbound

This paper cites Sparse Autoencoders Find Highly Interpretable Features in Language Models.

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability Sparse Autoencoders Find Highly Interpretable Features in Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.776980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:27:56.855130Z digest=sha256:b497c0e33815ee8efeb25ff5add96da300e26d17f7336bc4f5f3122b131ddcf1

Observation 12e20681-664f-488e-bab0-08a5f8e6a0cc · outbound

This paper cites Toy Models of Superposition.

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability Toy Models of Superposition

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.785092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:27:56.855130Z digest=sha256:043d5a493d9d0e0188659cac1e15ba68a25a33c4ec2d7c49b81a7ce5903872ee

Observation f24754c5-fc76-489b-88db-d4dc9f5640d5 · outbound

This paper cites The Llama 3 Herd of Models.

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability The Llama 3 Herd of Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:18:56.781308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:27:56.855130Z digest=sha256:e2d6902f72a24a092a90c7a5a4fd8f207060f33f9a9fc45997da1f8c9a517a67

Observation 59979a45-5c0e-41ff-815a-42c9a1485825 · outbound

This paper cites Uniform convergence may be unable to explain generalization in deep learning.

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability Uniform convergence may be unable to explain generalization in deep learning

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:56.775223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:27:56.855130Z digest=sha256:9376ec1a70f4202a072fed0511bf2d2b2d30a7de6fef29a55c776630a8e2482f

Observation a61ff091-af2c-4660-86c4-fbeb22877652 · outbound

This paper cites year =.

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability year =

Reference 9

Resolution
metadata mismatch
doi, observed 2026-06-27T01:30:20.472795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:27:56.855130Z digest=sha256:145f8d0a470d1ef8facbaf90dde563ec82c39816c5f76579a40a87c5cb8ca060

Observation edcbd50a-56c8-4630-93d1-403290c8e64f · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:18:56.780715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:27:56.855130Z digest=sha256:821f0347946641bbef94b984dc7810dfc0c4b09ae9ba32b34f4e8f478cad7c24

Observation cd3f17b6-ca41-4e8f-904c-5d481a0fb1e6 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability Gemma 2: Improving Open Language Models at a Practical Size

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.784873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:27:56.855130Z digest=sha256:4046175ae0ea11c967b7cb2e68fcbe1f3f9c46e71be3406dcfe961ac1e05a5ce

Observation 5a6623b3-a7fa-4cfa-aafd-57709c3d4a37 · outbound

This paper cites LogitLens4LLMs: Extending Logit Lens Analysis to Modern Large Language Models.

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability LogitLens4LLMs: Extending Logit Lens Analysis to Modern Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:18:56.788813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:27:56.855130Z digest=sha256:a757ffdfa507bf1e7434c21cc70ad7fd26927dab40a1fab80770b4165cd319cc

Observation 469dadfe-9cfc-4d8e-8c6c-50b1da48d80e · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:18:56.769655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:27:56.855130Z digest=sha256:cc3177adf9a75a0e0fbfbb4ddb86a6b6423226686ef925087174a7e84ea42e6b

Observation 495947bf-f51c-4794-bd4f-7889bedf59e2 · outbound

This paper cites Understanding deep learning requires rethinking generalization.

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability Understanding deep learning requires rethinking generalization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-03T20:18:56.764563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-27T01:27:56.855130Z digest=sha256:06327a89b2b2ca15514cac11e1a364a6f440cd317fd03005d1d5bf2bf24ce441

Observation 08355557-7e21-45bc-b1c3-77ce3d9fc0e0 · outbound

This paper cites The layerwise sweeps in Section 5.2 and Appendix E use additional layer-specific checkpoints where needed.

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability The layerwise sweeps in Section 5.2 and Appendix E use additional layer-specific checkpoints where needed

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T01:27:56.855130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:27:56.855130Z digest=sha256:628b510d2745d10febbeb06b02a3df82b9ad6ac1001846d09aa8388a564eebfb

Observation e71a5b84-83a1-4501-940e-1c384eb3c1cc · outbound

This paper cites an unresolved cited work.

From Sparse Features to Trustworthy Proxies: Certifying SAE-Based Interpretability Unresolved cited work

Reference 16

Resolution
malformed identifier
no resolver link, observed 2026-06-27T01:27:56.855130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T01:27:56.855130Z digest=sha256:889cf2640e2e3904860a16dd7c88390ef75dfc63a22134c452baeabf8925d416

Pith citing papers

No inbound Pith citation observations are available.