Pith. sign in

Paper Citation Record · LEDGER

Prototype Transformer: Towards Language Model Architectures Interpretable by Design

As of 23 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 2 inbound Pith citation observations for arXiv:2602.11852.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.11852 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:02:46.864302Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T17:43:52.716947Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T17:48:03.356591Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1bfac4d-6b43-4614-85e4-264b835edc9d · outbound

This paper cites GPT-4 Technical Report.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:44.028573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:44.028573Z digest=sha256:1ff4f24000b2ec4e36425a968ba76c82659d53a32f4996bf1465f2f769935603

Observation af9e2236-62b9-4057-90da-077a4e4f759e · outbound

This paper cites in the”, “of the.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design in the”, “of the

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:46.752202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:46.752202Z digest=sha256:e2ea47fa19be3dd4841db2c0110b1d31536231101717049daec2901cffa8069b

Observation f3da37f5-1912-488d-a78c-5bace7c0377c · outbound

This paper cites why does my dog eat poo p?.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design why does my dog eat poo p?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:46.576286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:46.576286Z digest=sha256:d6c55ba6fd45fe4e139739cbbd1e89253ee240eed81fafa4a3705f3e2df5f7c8

Observation 8b5c5f2e-3e6c-48c2-92d8-1952187014cd · outbound

This paper cites Alignment faking in large language models.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Alignment faking in large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:44.405381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:44.405381Z digest=sha256:6e59eb93d3a6d9f27df589c3fdcb84a9b5c1eb833cccaf73bc7d2dba43bf5337

Observation adedb359-c754-481d-8284-abb4d84791ec · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:44.447448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:44.447448Z digest=sha256:42682ab816c6d29b6723945858a6c8cd03130cb04572328098c5c506794e69a7

Observation 46b28b14-7707-4f1e-a7d1-9d6078015bd1 · outbound

This paper cites an unresolved cited work.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:44.501260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:44.501260Z digest=sha256:6168c71b6413c3b4b828b9be89155ebda07b6eb9be94f903adad99179fb51aa2

Observation bfe80cb7-8e10-468a-9df2-844fa94f8921 · outbound

This paper cites Interpreting Attention Layer Outputs with Sparse Autoencoders.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Interpreting Attention Layer Outputs with Sparse Autoencoders

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:44.553873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:44.553873Z digest=sha256:bd6287a23f654e9c1811045beca2fe28682642af4dfdc3c436a138e503df0155

Observation c4374789-4366-41cb-9250-59ca772f2aa5 · outbound

This paper cites Decoupled Weight Decay Regularization.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Decoupled Weight Decay Regularization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:44.712646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:44.712646Z digest=sha256:28d71b7f1c6eada5cf8299f00b19ed348366dbc14a153cc84a0429361102e91a

Observation 1b0442a0-0e4a-43bc-b578-13a5d9bbdb8b · outbound

This paper cites Revisiting Small Batch Training for Deep Neural Networks.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Revisiting Small Batch Training for Deep Neural Networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:44.801595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:44.801595Z digest=sha256:961876285f087354242eeed6e67aae6cd94b6c75730a32ed45bf2ae15594bc09

Observation fee98cb1-71ec-4910-927b-fcd4aa72dee5 · outbound

This paper cites Pointer Sentinel Mixture Models.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Pointer Sentinel Mixture Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:44.868311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:44.868311Z digest=sha256:9387f675c8fd0b88665340f52f1c49cadfb03edfbef4ce534ba230e7fc2b3be5

Observation 4fce6703-a533-4f72-8c27-92c5febd96ee · outbound

This paper cites A practical review of mecha- nistic interpretability for transformer-based language models.arXiv preprint arXiv:2407.02646,.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design A practical review of mecha- nistic interpretability for transformer-based language models.arXiv preprint arXiv:2407.02646,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:44.980576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:44.980576Z digest=sha256:c201953c555f88972b6fa119681748559ae7c8dc9d82ea1df55575d1d625a08a

Observation a63bffcf-1e9f-4821-adeb-6a6073a781d3 · outbound

This paper cites Sentence-BERT: Sentence embeddings using Siamese BERT- networks.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Sentence-BERT: Sentence embeddings using Siamese BERT- networks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:45.106968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:45.106968Z digest=sha256:b2c7d7751411a0f98847591075362ee22cd5f6772682b97b64739ecc33e3f253

Observation fc58d4da-ff62-462b-89c8-92d5539d6325 · outbound

This paper cites Neural Machine Translation of Rare Words with Subword Units.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Neural Machine Translation of Rare Words with Subword Units

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:45.222948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:45.222948Z digest=sha256:ff943388b3ef76044c780ab9bd2143d3b640af6ce32d07915f56489a60c2194a

Observation 41941a36-071c-442f-a066-8c600b8b94ba · outbound

This paper cites BERT Rediscovers the Classical NLP Pipeline.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design BERT Rediscovers the Classical NLP Pipeline

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:45.502617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:45.502617Z digest=sha256:6f013724a8dc265cf4cee44fde8d7d518a7fc4760a04ffbc73a09e3973e50ae0

Observation eb975022-1a6e-46f5-adac-0c72f8705585 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design LLaMA: Open and Efficient Foundation Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:45.649779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:45.649779Z digest=sha256:2742eaf78e6fef3577eac526ae0e823c6aa99e326d835489aa2d8f94e2102305

Observation 5148a92f-9c00-473b-864f-105e0c29e300 · outbound

This paper cites GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:45.798192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:45.798192Z digest=sha256:0f551663efdd64c31e0fe5b8c44756cb9b3031665b793f01880ec7fddb611683

Observation 08ba0b77-c02a-48c3-9a40-3b02ad92be0f · outbound

This paper cites MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained Transformers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:46.051367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:46.051367Z digest=sha256:f1ee6f5c391fc43b987ed0d73b1f7fc1f169e69737e0d35a6a2227a77089ad03

Observation 1ecbbd32-0886-4350-8b15-0a4a915caa04 · outbound

This paper cites Qwen3 Technical Report.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Qwen3 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:46.227111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:46.227111Z digest=sha256:78a713220d624d25a517f90365d29bda496bb3cb7421c2913b4ba2dbe33a1417

Observation 834de153-2f61-4a87-88ae-1c3586f40b90 · outbound

This paper cites Image classification at supercomputer scale.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Image classification at supercomputer scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:46.310196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:46.310196Z digest=sha256:8484555c1a5cb7caef811f4b1c94c89d32eb8a8418d7c7a56f65f79f86893146

Observation 6bfe7acc-6173-4475-b784-2f6042fd6c09 · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:46.368866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:46.368866Z digest=sha256:19517c6f2c0bd1b41dfc1cb50587e247e000d938377fc9c25645b10ac66fd0f6

Observation e2572600-d266-46aa-9b73-c18d7d38e2fe · outbound

This paper cites Deconstructing What Makes a Good Optimizer for Language Models.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Deconstructing What Makes a Good Optimizer for Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:46.444321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:46.444321Z digest=sha256:9f8998d1249cb87b886173798a37a8d5092be1d18d822a85235c3794182f60c0

Observation c5bea32e-e11e-4c37-b82f-13aa38042664 · outbound

This paper cites to help us produce the prototype interpretability html.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design to help us produce the prototype interpretability html

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:46.512106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:46.512106Z digest=sha256:2a762a7b2550a6faf24a60d1d27454bff7a82720228e84088b98157d5d6974e2

Observation eab81462-3c7d-4139-81fe-216efd394dbb · outbound

This paper cites You are analyzing a single prototype (a neuron-like feature) from a neural language model.\n.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design You are analyzing a single prototype (a neuron-like feature) from a neural language model.\n

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:46.864302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:46.864302Z digest=sha256:1bb08ae3c9770decdd1a620824c765b6032156d8450c68f5c3c0d0ce5e2dc419

Observation b00eba98-5974-4fa3-a4bd-be6ff0f33e8b · outbound

This paper cites an unresolved cited work.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Unresolved cited work

Reference 512

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:46.654475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:46.654475Z digest=sha256:758cebf71805c69ef96812b6afb24c8aff5d028f995f3530b9eaf76f9eaa69dd

Observation 40fa6dbd-3e7a-44e4-9b18-20d62b4e15f5 · outbound

This paper cites Automatically Interpreting Millions of Features in Large Language Models.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Automatically Interpreting Millions of Features in Large Language Models

Reference 1995

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:44.908484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:44.908484Z digest=sha256:a406f2dbdf1f8deb97ea0e713922e0e96056874538dde94d57d25ff17ec0fa6c

Observation 7ced3996-6364-4a23-a77f-5bc065fea788 · outbound

This paper cites GLU Variants Improve Transformer.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design GLU Variants Improve Transformer

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:45.387740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:45.387740Z digest=sha256:a145efbdda76f7132ea178196bf76779e151db90c954f48fe488838d34f0be75

Observation 020f9809-d3ca-4704-8415-ab7576d37bca · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Jamba: A Hybrid Transformer-Mamba Language Model

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:44.629564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:44.629564Z digest=sha256:225c2429a27fd6453081bae191753330bc7fa26197bbde2f0b19f129df5a26c1

Observation 26f2a1fc-a6e3-4916-ab4e-87dbfafc3d56 · outbound

This paper cites Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:45.942638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:45.942638Z digest=sha256:ecfadec19f0c00c9b25290516d48697c7add3f867008019a8cb10dfc7e153e6d

Observation c653af77-502d-40b8-8e35-261f826ecdd2 · outbound

This paper cites Toy Models of Superposition.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Toy Models of Superposition

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:44.238440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:44.238440Z digest=sha256:81924a3dee4547c51180748b5634fdb7f9222b30917a6a1d4dfba33a6a55ff04

Observation 15f46023-078f-4a3e-ae84-15d5872d299b · outbound

This paper cites ReZero is All You Need: Fast Convergence at Large Depth.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design ReZero is All You Need: Fast Convergence at Large Depth

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:44.101431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:44.101431Z digest=sha256:e46058a2697414391f5a1e98e7056c143c16440eeb3e18d2f3eb0944c92a7a3a

Observation 75b0f373-ec3b-4ecf-9551-f1139e62d36c · outbound

This paper cites Dissecting Recall of Factual Associations in Auto-Regressive Language Models.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design Dissecting Recall of Factual Associations in Auto-Regressive Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:44.289013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:44.289013Z digest=sha256:ab4bdaeec80fefaeeafb0d8ddb51e502cfadc676e3e4c01593abdae7a36179fd

Observation 3c45c933-8f46-44dc-921d-f26437a88557 · outbound

This paper cites The Llama 3 Herd of Models.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design The Llama 3 Herd of Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:44.342008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:44.342008Z digest=sha256:6e1602b166b3b96ddbff031a203a88a41aea75e3ece2302b761f35397e82fc1c

Observation 51a80759-3718-4f16-ad06-a22a23c087d3 · outbound

This paper cites What Does BERT Look At? An Analysis of BERT's Attention.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design What Does BERT Look At? An Analysis of BERT's Attention

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:44.176347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:44.176347Z digest=sha256:a24ba337ec8b90c0125bdbff3a75951054cfff9900ff1785a1da74a135f6b3c4

Observation 94ee83c8-71d3-4a93-938f-b8bb4da36bfc · outbound

This paper cites A Transformer and Prototype-based Interpretable Model for Contextual Sarcasm Detection.

Prototype Transformer: Towards Language Model Architectures Interpretable by Design A Transformer and Prototype-based Interpretable Model for Contextual Sarcasm Detection

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T00:02:46.162136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T00:02:46.162136Z digest=sha256:9b9abfaa72142c093aa3f79148ced0a9063da26df208d896b17332c413460157

Pith citing papers

Observation 4fda4066-2b23-4177-af82-33dbb4f77ae8 · inbound

Collapse-Free Prototype Readout Layer for Transformer Encoders cites this paper.

Collapse-Free Prototype Readout Layer for Transformer Encoders Prototype Transformer: Towards Language Model Architectures Interpretable by Design

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-02T04:04:28.578289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-13T17:43:52.716947Z digest=sha256:49838815bb6d32a360d43377583b5520c0072c112d211d54888ac893916e3fbc

Observation f31719cd-a815-4e8c-a88d-662a3464dd50 · inbound

Graph Memory Transformer (GMT) cites this paper.

Graph Memory Transformer (GMT) Prototype Transformer: Towards Language Model Architectures Interpretable by Design

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-02T04:04:28.578289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T06:27:32.790728Z digest=sha256:18fba2515b767f3419f83aee245ea2c9a30fa0d2af4eb416c0bdb0b0ecdd025f