Pith. sign in

Paper Citation Record · LEDGER

An Interpretability Illusion for BERT

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2104.07143.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.07143 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:58:33.827272Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T12:46:57.379107Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a175f726-d97d-4094-8d9e-db3489192e8d · inbound

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small cites this paper.

Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small An Interpretability Illusion for BERT

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:13:51.507126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T17:13:51.408311Z digest=sha256:e50ac096d69d3ef54a0ce7ace8d68d4f73df526e8b769ad646b095c5f11c1d05

Observation 30bde55e-1f39-443b-ae73-d4fad517bce7 · inbound

Localizing Model Behavior with Path Patching cites this paper.

Localizing Model Behavior with Path Patching An Interpretability Illusion for BERT

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T19:38:37.786625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-16T19:38:37.751487Z digest=sha256:928124534803957f08fda922bda6271f41910ccc0a916e25b73bb4f217d2d151

Observation 54bdea91-580e-4f4f-8ca2-32e2a07e3f68 · inbound

Improving Dictionary Learning with Gated Sparse Autoencoders cites this paper.

Improving Dictionary Learning with Gated Sparse Autoencoders An Interpretability Illusion for BERT

Reference 202

Resolution
verified exact
arxiv_id, observed 2026-05-17T19:42:32.035453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-17T19:42:31.855039Z digest=sha256:b37e3c35ee15dc6a3a38c5558187428833ca38711a010926e9bb17deeede221b

Observation 3cee2c1b-c56f-4e33-a67e-0590ee663b44 · inbound

Scaling and evaluating sparse autoencoders cites this paper.

Scaling and evaluating sparse autoencoders An Interpretability Illusion for BERT

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T17:47:23.139629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T17:47:23.089288Z digest=sha256:2d9aeaded5d0e4f671e196331b4111b8f0d292250cf2544aed17b0a3f3c1a6fe

Observation 777759fa-b0ae-43f2-a7fd-eb557a7d9d8e · inbound

Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning cites this paper.

Perspectives for Direct Interpretability in Multi-Agent Deep Reinforcement Learning An Interpretability Illusion for BERT

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-09T17:58:33.827272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:58:33.827272Z digest=sha256:2a68e2c9651feab22963ac9338fd6576222dbb0b5130aed7d05df54174bbe5ce

Observation e1b8db61-ae1a-4685-a0a9-8957651cb965 · inbound

Evaluating Neuron Explanations: A Unified Framework with Sanity Checks cites this paper.

Evaluating Neuron Explanations: A Unified Framework with Sanity Checks An Interpretability Illusion for BERT

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.712297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:18:24.712297Z digest=sha256:d14000913eb79a5b4386ea23199a85dd0b59cf398f8cb8f085325d1441980fbd

Observation 35896df0-11a0-4992-afe6-415d35e253b0 · inbound

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations cites this paper.

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations An Interpretability Illusion for BERT

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T14:23:30.778853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T14:16:44.232080Z digest=sha256:923e24b02291521bc1663a598c0fa6ef59d84bfbeda75410e96a4e2c34e8c307

Observation f25ab80d-7a18-4e33-9c39-bcd4f3b98a1f · inbound

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations cites this paper.

Sign-Aware Gated Sparse Autoencoders: Modeling Anticorrelated Features with Bi-Jump-ReLU Activations An Interpretability Illusion for BERT

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T05:02:51.278375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:02:51.278375Z digest=sha256:832848b97c527907baf65ce1c96bac4fddf4a7c31f764fba07275ef17723a2a2

Observation d6f48bcd-4a4f-4411-b962-b0732e2f3383 · inbound

Many Circuits, One Mechanism: Input Variation and Evaluation Granularity in Circuit Discovery cites this paper.

Many Circuits, One Mechanism: Input Variation and Evaluation Granularity in Circuit Discovery An Interpretability Illusion for BERT

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:57.380709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T01:48:12.958931Z digest=sha256:ea8c32d472e5a425b81437ae6f0ce939647e73881c94989b96751586bb3952ff

Observation 677968aa-ccce-4f23-bd82-78289adf54e4 · inbound

The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators cites this paper.

The Entanglement Wall: Activation-Space Probes as Risk Detectors, Not Context Adjudicators An Interpretability Illusion for BERT

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T07:09:29.963193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:09:29.963193Z digest=sha256:d3a1f85595f5918bf971fe077da8a0e073a7c0015ec67af36b370d4f95e09996