Pith. sign in

Paper Citation Record · LEDGER

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models

As of 13 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2411.17182.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17182 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:32:45.226075Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8ac51476-fb4a-47f2-b0f9-cd2b04cc9ae8 · outbound

This paper cites Repulsive attention: Rethinking multi-head attention as bayesian inference.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Repulsive attention: Rethinking multi-head attention as bayesian inference

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:46.022978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.002673Z digest=sha256:ec34d63c3f78db9f4bee3fe71b9724b72e08de2df0f8caf27302e81df75c0722

Observation b35fff6e-3134-45ad-a3a2-ce6b93eed7b0 · outbound

This paper cites Layer Normalization.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Layer Normalization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.007722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.007722Z digest=sha256:153c120cb90d1c065e7c650f245677ab4f80c890b84036514396de9ce44d0771

Observation 6ea967a3-0fb2-4185-baf8-45ae8ea67187 · outbound

This paper cites Spectrally-normalized margin bounds for neural networks.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Spectrally-normalized margin bounds for neural networks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.012782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.012782Z digest=sha256:a2087d3e51254e3cab2001c2ef087010edda9eb1b1d5e741b69d346acfdbbd78

Observation 5462d016-0753-4c0c-ae96-de693445e988 · outbound

This paper cites Birth of a transformer: A memory viewpoint.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Birth of a transformer: A memory viewpoint

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.998168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.017923Z digest=sha256:0bf1576b1ae01612db8ae5e0bded515e98aa5eb4682ca3133a0ca9cf77d8cc04

Observation 9578f20e-5e1e-4c5b-95dc-6e8750604d8b · outbound

This paper cites Attention approximates sparse distributed memory.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Attention approximates sparse distributed memory

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.982390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.023011Z digest=sha256:c01854ae37781c65ebaac18a863c0a298c845c4fc12030aeecb602a9992b8a72

Observation 0090f6f0-0841-4359-b3ac-4e7327cbb323 · outbound

This paper cites Invariant scattering convolution networks.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Invariant scattering convolution networks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.966651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.027526Z digest=sha256:b90ce46922ef4af8117b62a6e68e207194c6da07af3c7166fa4e30d002cd7e2d

Observation 3e9edc9b-69b6-42fd-ada6-314db41fb187 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Emerging properties in self-supervised vision transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.033422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.033422Z digest=sha256:71a972468accbbd046b40f8a84290c3d20ae6f28931df9e8720b6ae0fab34491

Observation 5b9218ed-4289-47eb-bacd-3b04d66e3770 · outbound

This paper cites Transformer interpretability beyond attention visualization.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Transformer interpretability beyond attention visualization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.940227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.038131Z digest=sha256:30ed59898e82d39a62201b601f5c110f7ac773e9362f5fbf5b14d3eabcaed68b

Observation c11ae912-1203-478c-847b-9370b9337f47 · outbound

This paper cites Randaugment: Practical automated data augmentation with a reduced search space.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Randaugment: Practical automated data augmentation with a reduced search space

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.042753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.042753Z digest=sha256:4c93a596cf963a9509b561dff396a1105a054b702bc9a066d8207908f1b11a14

Observation 6b8ff53b-013c-45ee-b6b2-276f5cba3f2e · outbound

This paper cites Analyzing transformers in embedding space.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Analyzing transformers in embedding space

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.915153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.047548Z digest=sha256:ce9ef701375b6f64161aaefcf32f402812642349afa5b4f4c2700c8d86ebb05b

Observation f52c75f1-0776-4cc8-b6a4-d0f8c87c02cd · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.052104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.052104Z digest=sha256:7dc9b367b223b0158939836a54dcfb60f4a0edeb237e33b1f60c0bb2f58d443b

Observation 7d1094ff-ed3a-49a6-8b29-fca4c410bafb · outbound

This paper cites A mathematical framework for transformer circuits.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models A mathematical framework for transformer circuits

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.056671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.056671Z digest=sha256:5237fd916b021b97eabf388447abb035644f9e52e7af3b417a23f31a9be9ee0f

Observation c06a0860-7503-4381-acf7-bd362f606f16 · outbound

This paper cites The emergence of clusters in self-attention dynamics.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models The emergence of clusters in self-attention dynamics

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.867629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.066937Z digest=sha256:7f6dfddef38171de831378d86bbeadc0a3400f63d42f8c369f7e90356ff9b242

Observation e79ac14f-aa14-437b-bf24-6cd0f9b3ac1f · outbound

This paper cites Patchscopes: A unifying framework for inspecting hidden representations of language models.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Patchscopes: A unifying framework for inspecting hidden representations of language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.852250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.071587Z digest=sha256:798d8dc2ecbf17893086cafcbe0f33d318bbaa95a5fdb8752cee2b61399d1ea1

Observation bf391b7c-adc1-4acb-b43b-61c758f08718 · outbound

This paper cites Learning fast approximations of sparse coding.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Learning fast approximations of sparse coding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.836412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.076259Z digest=sha256:c2b9668d0873938b8a6acb98fbf50581a8a1e9a3600d5a364a6bffb83a228f7e

Observation 605f79d1-605a-4da1-a9f0-03c135299011 · outbound

This paper cites The Forward-Forward Algorithm: Some Preliminary Investigations.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models The Forward-Forward Algorithm: Some Preliminary Investigations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.081163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.081163Z digest=sha256:683ebe11c51e39e34f69974b562086effd7eea9d08c8877a49cc1ad181b22ad4

Observation 1be9dafb-f19b-4cee-a6cc-ca8e4612b738 · outbound

This paper cites Energy transformer.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Energy transformer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.819300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.086133Z digest=sha256:f6eeedc1fa99cf8925219a65400be3c1bd803166b8857dc91c54ac4b9b896a65

Observation 1625818f-5370-4e2e-a8bb-73ec627ab404 · outbound

This paper cites Fantastic generalization measures and where to find them.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Fantastic generalization measures and where to find them

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.803266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.090907Z digest=sha256:230fb550a6bbbaa4c8c305da95824743323d994be6ce9168570f094cc20d698a

Observation 7608b21f-bdd6-41c6-98c5-ab7723c5504e · outbound

This paper cites A new measure of rank correlation.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models A new measure of rank correlation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.095640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.095640Z digest=sha256:a37a89c16379d945dd758c2bc751f8a47a5c2f8327c38c524457374a5bec1617

Observation 8f2f13eb-9e2d-4757-8686-ed2a53edbc07 · outbound

This paper cites On large-batch training for deep learning: Generalization gap and sharp minima.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models On large-batch training for deep learning: Generalization gap and sharp minima

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.776973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.100467Z digest=sha256:de2c67371d2230eed898093a6150961679c74f1d9955c417217c248c9f8b2149

Observation 250b883e-35af-43ae-9acd-9f8f564a1372 · outbound

This paper cites Kingma and Jimmy Ba.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Kingma and Jimmy Ba

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.105058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.105058Z digest=sha256:4a97bbf5c1abc032e2d08e6fda1b9586597a3c4e4e0505d0625a2ea945c85a08

Observation 76d36dfa-a66a-4e74-ad0c-1ce875852f48 · outbound

This paper cites Tracr: Compiled transformers as a laboratory for interpretability.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Tracr: Compiled transformers as a laboratory for interpretability

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.752778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.110071Z digest=sha256:cf9fe448bf17296c92766aead09f7be63b1c455dcc10b46bf177f80dc22174ed

Observation bc6f5695-7460-492c-8500-81222d17031c · outbound

This paper cites Omnigrok: Grokking beyond algorithmic data.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Omnigrok: Grokking beyond algorithmic data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.114718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.114718Z digest=sha256:26cbef16b1b2b8b339754477605810c4d02406eb37b1f41687c04a2901303661

Observation 69c8da1a-a61c-4f39-855b-c2171387f5d5 · outbound

This paper cites Segmentation of multivariate mixed data via lossy data coding and compression.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Segmentation of multivariate mixed data via lossy data coding and compression

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.726347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.119269Z digest=sha256:47cdba3c11d02334989cf56883c4fb19b521a35a91f6955b0f3bdcb40e343f02

Observation 501377e5-6383-49eb-a7cb-e64f16f231c7 · outbound

This paper cites Pac-bayesian model averaging.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Pac-bayesian model averaging

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.710222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.124143Z digest=sha256:c9d38bbd1c9d9498f3b4ff6de59300e764d5402a4d9118c06141daa9ac9f2994

Observation 73fe6ed6-92e0-4665-9c76-4021e94b5bc2 · outbound

This paper cites Universal hopfield networks: A general framework for single-shot associative memory models.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Universal hopfield networks: A general framework for single-shot associative memory models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.694650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.128881Z digest=sha256:42efa47674e072175c25cceb590313beaabf24cea10a7ca046cf7723643c27eb

Observation b50e8e18-e4a0-4c14-8af1-a2d7dc9b730e · outbound

This paper cites Algorithm unrolling: Interpretable, efficient deep learning for signal and image processing.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Algorithm unrolling: Interpretable, efficient deep learning for signal and image processing

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.133706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.133706Z digest=sha256:8b30e1794af84d88317cc0b9d06baecb060ab544921765e08bfe919dabe7d5df

Observation 3344aab9-2dbb-4b06-8a69-8fcaf690904a · outbound

This paper cites Progress measures for grokking via mechanistic interpretability.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Progress measures for grokking via mechanistic interpretability

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.138412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.138412Z digest=sha256:ef4e001deebcdb7ee5f4a5da394c21da4d90d354f10ec5df81a9495a2b542e7a

Observation 0584541d-c4ff-4f99-bb6e-cf0ad3df2db9 · outbound

This paper cites Exploring generalization in deep learning.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Exploring generalization in deep learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.143259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.143259Z digest=sha256:9005b8786ff1b1b4d78f7aeb0e7f5e2f053afb6c80bbd1a58b45a9a9f65e5ab0

Observation d01d58ef-e403-45e8-bbeb-38bc23881973 · outbound

This paper cites A PAC-bayesian approach to spectrally-normalized margin bounds for neural networks.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models A PAC-bayesian approach to spectrally-normalized margin bounds for neural networks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.648850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.147803Z digest=sha256:c8889f33577a48a49b02e381cae8a05420c22c481224deb4b32bd19c0bf3a464

Observation a6007c5e-c231-47cd-8547-3df21ea368f2 · outbound

This paper cites Path-sgd: Path-normalized optimization in deep neural networks.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Path-sgd: Path-normalized optimization in deep neural networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.152423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.152423Z digest=sha256:552861f8ed18cb26bc3fae1d184604a9999c6172a7d13784618d7597d057e6bd

Observation a29f89a8-f7d3-4be4-a5aa-84cddea91a4d · outbound

This paper cites Norm-based capacity control in neural networks.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Norm-based capacity control in neural networks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.157053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.157053Z digest=sha256:7dd45280967388c61676025bff558515f2fed5b21d2ede7fc266fb55845a9ad2

Observation bf19aabb-2092-4d15-a8ac-824a19b2ef4d · outbound

This paper cites Theoretical foundations of deep learning via sparse representations: A multilayer sparse model and its connection to convolutional neural networks.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Theoretical foundations of deep learning via sparse representations: A multilayer sparse model and its connection to convolutional neural networks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.612840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.161472Z digest=sha256:9977128f62f84ddbefc9d99b0f14183ba2c822e8455070488cce8cc9fa09d6b0

Observation 68370c3f-74ac-4fc1-865f-be6b2b4cf1b2 · outbound

This paper cites Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.165818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.165818Z digest=sha256:66798a0d89439ae72ec0d74d1d23d5bdd066e08540ced62c21eecc059060ffd5

Observation 9c16a147-3160-4d09-925c-5c70977a2500 · outbound

This paper cites Hopfield networks is all you need.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Hopfield networks is all you need

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.596998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.170405Z digest=sha256:18b319386310e2ee1d17a0710e40c0f9f92141c3e32d7183752431e23a938963

Observation d047be21-d60e-4e4d-96ac-a3cb81671714 · outbound

This paper cites Unraveling attention via convex duality: Analysis and interpretations of vision transformers.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Unraveling attention via convex duality: Analysis and interpretations of vision transformers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.582018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.174964Z digest=sha256:2361e46ea04416242b8d729f572ce6f7c3683bea1274efbc539a445477a28b83

Observation e8f5ac9e-88a5-44d7-a4e7-f40c3afbcb3f · outbound

This paper cites Biological learning in key-value memory networks.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Biological learning in key-value memory networks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.566234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.179474Z digest=sha256:6edc1ad3f354a2a115d4ffc6e74c7cd0e0161ca75bfe42e08b7cc4d2476327d8

Observation 7ceec1d0-12e0-4188-9409-b85ef544dc90 · outbound

This paper cites On the uniform convergence of relative frequencies of events to their probabilities.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models On the uniform convergence of relative frequencies of events to their probabilities

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.183774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.183774Z digest=sha256:bbacf13554d7f51e438c8922362bd8b5f1f4eebdefa1700ad226fa1bd8d55ab2

Observation 45595754-205e-4650-8973-eb26e14d9a7c · outbound

This paper cites Attention is all you need.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Attention is all you need

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.188711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.188711Z digest=sha256:c3586bcf57d934d631680d43544938581361a57c8fa80684b2399ad7aede4ac6

Observation 54fbf4f8-9de5-4311-930a-8de8af52ca77 · outbound

This paper cites Interpretability in the wild: a circuit for indirect object identification in GPT-2 small.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Interpretability in the wild: a circuit for indirect object identification in GPT-2 small

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.193362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.193362Z digest=sha256:36bdcb29ccb8709767cb6dad2bb666bb8323cd32c877a767cb27e8e9a2015258

Observation 16708be7-7ec9-4462-b6f0-49fd059e8c47 · outbound

This paper cites Thinking like transformers.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Thinking like transformers

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.522435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.198222Z digest=sha256:74727e01fdd4f83fefcb0ae411226ac474f64ec59b436432f00fad3ea784c039

Observation 28500a26-2d01-4722-86b7-8319da737fa4 · outbound

This paper cites Graph neural networks inspired by classical iterative algorithms.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Graph neural networks inspired by classical iterative algorithms

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.505906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.202827Z digest=sha256:bd80818d64fa04b3abad32f7ce29420ec645edf0ff67922c361b4156933a8659

Observation 5819e8a2-a4ec-4038-be03-839a1c45ad48 · outbound

This paper cites Transformers from an optimization perspective.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Transformers from an optimization perspective

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.490925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.207058Z digest=sha256:1b5403b1584f7da187fc0cc7a7fdcc84f1bd72802aa34d39eac27e7f73ff3be8

Observation 0b214848-862c-4a0d-8529-e072e6f554e2 · outbound

This paper cites Attentionviz: A global view of transformer attention.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Attentionviz: A global view of transformer attention

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.475648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.211699Z digest=sha256:283bef8c616670d43b2b3f6c9be2e7299337172299134cf89803913b3b2fd188

Observation 12cf0273-156d-4d0c-b1cc-529692a0833f · outbound

This paper cites White-box transformers via sparse rate reduction.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models White-box transformers via sparse rate reduction

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.460496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.216547Z digest=sha256:2e6b3e8f5b98f5ff435e49394df89d34de3a3710b85acdbfe08aa26493b55e05

Observation 048031ec-59eb-4fc9-a8a0-d30640be511d · outbound

This paper cites Learning diverse and discriminative representations via the principle of maximal coding rate reduction.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Learning diverse and discriminative representations via the principle of maximal coding rate reduction

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.445097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T12:32:45.221157Z digest=sha256:7136c7ce06b7f93368f55867b0c596c1d5948f9eec224c6d82c59b3110ceab66

Observation af831bc5-e02c-4fd1-8cad-c1d2357de761 · outbound

This paper cites Unveiling Transformers with LEGO: a synthetic reasoning task.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Unveiling Transformers with LEGO: a synthetic reasoning task

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.226075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.226075Z digest=sha256:5655caa7c6d717be27e4e7eb25569d3d39785c16425c1280e59ccb6d4edb7b47

Observation b7346d03-3e0f-4e9f-9460-f178836751b5 · outbound

This paper cites an unresolved cited work.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Unresolved cited work

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.061926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.061926Z digest=sha256:d85ca9f6ff7915b7d77f504b0f49a92e9734faf1263a205b0fe4ba0c87d4b74a

Pith citing papers

No inbound Pith citation observations are available.