Pith. sign in

Paper Citation Record · LEDGER

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models

As of 21 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2411.17182.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.17182 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:32:45.226075Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8ac51476-fb4a-47f2-b0f9-cd2b04cc9ae8 · outbound

This paper cites Repulsive attention: Rethinking multi-head attention as bayesian inference.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Repulsive attention: Rethinking multi-head attention as bayesian inference

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:46.022978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.002673Z digest=sha256:21436fcf2757137eeab4a0a8f9a05ae0b31d695e83062ca86ea5c48f2f8689d3

Observation b35fff6e-3134-45ad-a3a2-ce6b93eed7b0 · outbound

This paper cites Layer Normalization.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Layer Normalization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.007722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.007722Z digest=sha256:15511b393a55a94a46e034ceebb716356fee808bd0ed90f5e76deaa6b61c0417

Observation 6ea967a3-0fb2-4185-baf8-45ae8ea67187 · outbound

This paper cites Spectrally-normalized margin bounds for neural networks.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Spectrally-normalized margin bounds for neural networks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.012782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.012782Z digest=sha256:5b0be464539c36ed7c8490c5f272190564873484d605392752d0e36917397b6e

Observation 5462d016-0753-4c0c-ae96-de693445e988 · outbound

This paper cites Birth of a transformer: A memory viewpoint.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Birth of a transformer: A memory viewpoint

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.998168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.017923Z digest=sha256:444da9e6762c145c8b43f21737dae55ba4dc1a4fa4a046819d7278805cd6532e

Observation 9578f20e-5e1e-4c5b-95dc-6e8750604d8b · outbound

This paper cites Attention approximates sparse distributed memory.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Attention approximates sparse distributed memory

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.982390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.023011Z digest=sha256:16704dbfed3950e6b8c76f7d54f6d3cfa2726946bfd2aa6c827e09bafbc431c1

Observation 0090f6f0-0841-4359-b3ac-4e7327cbb323 · outbound

This paper cites Invariant scattering convolution networks.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Invariant scattering convolution networks

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.966651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.027526Z digest=sha256:d335006b789d7195569a90c7fae94537d5f6a219d9943c338a24f7cedb96733a

Observation 3e9edc9b-69b6-42fd-ada6-314db41fb187 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Emerging properties in self-supervised vision transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.033422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.033422Z digest=sha256:cd20dff17abec989c7ad6ff37b8cc7e5913f97d6da84f806cb3c26ea8bd57ac1

Observation 5b9218ed-4289-47eb-bacd-3b04d66e3770 · outbound

This paper cites Transformer interpretability beyond attention visualization.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Transformer interpretability beyond attention visualization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.940227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.038131Z digest=sha256:08ee3c6854542f00b8e37abef01c86b5f06b07d06ff132f4de827f5f88e177ad

Observation c11ae912-1203-478c-847b-9370b9337f47 · outbound

This paper cites Randaugment: Practical automated data augmentation with a reduced search space.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Randaugment: Practical automated data augmentation with a reduced search space

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.042753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.042753Z digest=sha256:4fcc821b034445e115703e4f500bab8900fc66197e1e59d1264f26ef4b4b74ec

Observation 6b8ff53b-013c-45ee-b6b2-276f5cba3f2e · outbound

This paper cites Analyzing transformers in embedding space.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Analyzing transformers in embedding space

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.915153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.047548Z digest=sha256:52177e59a9a46625df61bc50c17ee55badfca8208ba3712c271b876bb2dbe251

Observation f52c75f1-0776-4cc8-b6a4-d0f8c87c02cd · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.052104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.052104Z digest=sha256:19aaa0b83b3b18ca628fe0f23799107cab9d115fca6ca2fdbd0df38ec8a7bf22

Observation 7d1094ff-ed3a-49a6-8b29-fca4c410bafb · outbound

This paper cites A mathematical framework for transformer circuits.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models A mathematical framework for transformer circuits

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.056671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.056671Z digest=sha256:41fb516ffc87fc5a985f2290f0533bd6539516d3ec150dcb24b8323b38da3fa3

Observation c06a0860-7503-4381-acf7-bd362f606f16 · outbound

This paper cites The emergence of clusters in self-attention dynamics.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models The emergence of clusters in self-attention dynamics

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.867629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.066937Z digest=sha256:8f8db1927b46fb7a4653c7eecf7b4105682abe20250c805035a94dcbf7c46f40

Observation e79ac14f-aa14-437b-bf24-6cd0f9b3ac1f · outbound

This paper cites Patchscopes: A unifying framework for inspecting hidden representations of language models.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Patchscopes: A unifying framework for inspecting hidden representations of language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.852250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.071587Z digest=sha256:8bf56b8d896205f2bf293449f9efdb51b7059663a2200e14ba6b095cab936687

Observation bf391b7c-adc1-4acb-b43b-61c758f08718 · outbound

This paper cites Learning fast approximations of sparse coding.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Learning fast approximations of sparse coding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.836412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.076259Z digest=sha256:5dadd38771720561a436159768987817ae6646190bfb5eb0bf97fb19b591a75e

Observation 605f79d1-605a-4da1-a9f0-03c135299011 · outbound

This paper cites The Forward-Forward Algorithm: Some Preliminary Investigations.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models The Forward-Forward Algorithm: Some Preliminary Investigations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.081163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.081163Z digest=sha256:82d3ec9824c54669d1ebaffd9d1a153fd0f489ef561fc88d9b6adf615a20d871

Observation 1be9dafb-f19b-4cee-a6cc-ca8e4612b738 · outbound

This paper cites Energy transformer.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Energy transformer

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.819300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.086133Z digest=sha256:4ff3d40392380bd5c94f849ed657b1e52d4643bf090e94be07261104436e5bbe

Observation 1625818f-5370-4e2e-a8bb-73ec627ab404 · outbound

This paper cites Fantastic generalization measures and where to find them.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Fantastic generalization measures and where to find them

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.803266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.090907Z digest=sha256:93da2dfe06df9cdd25724937e4bde73d8f7689e50c706804a851306b6a723ffe

Observation 7608b21f-bdd6-41c6-98c5-ab7723c5504e · outbound

This paper cites A new measure of rank correlation.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models A new measure of rank correlation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.095640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.095640Z digest=sha256:900b15ed7058a7e9ebe2f39fb79b6336ba91046a062976eafe24d60c0ac6fab0

Observation 8f2f13eb-9e2d-4757-8686-ed2a53edbc07 · outbound

This paper cites On large-batch training for deep learning: Generalization gap and sharp minima.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models On large-batch training for deep learning: Generalization gap and sharp minima

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.776973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.100467Z digest=sha256:ce5dd3fb5d159b695e035be5da162145d38388b5c1e378ba6f70a65ded9d1786

Observation 250b883e-35af-43ae-9acd-9f8f564a1372 · outbound

This paper cites Kingma and Jimmy Ba.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Kingma and Jimmy Ba

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.105058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.105058Z digest=sha256:e2d297e4bd1246583219f809e4402ee638ce46b0d2eb28c0658547cb02208489

Observation 76d36dfa-a66a-4e74-ad0c-1ce875852f48 · outbound

This paper cites Tracr: Compiled transformers as a laboratory for interpretability.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Tracr: Compiled transformers as a laboratory for interpretability

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.752778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.110071Z digest=sha256:b25f163c1f5ceaf30e29e7f856887ad443a554c7beab316b70a0df6dd5564222

Observation bc6f5695-7460-492c-8500-81222d17031c · outbound

This paper cites Omnigrok: Grokking beyond algorithmic data.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Omnigrok: Grokking beyond algorithmic data

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.114718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.114718Z digest=sha256:233c3ff30f6c970a0f46b19b5d1f6fd18c365f35a08f40355e7e4f21ba3d3520

Observation 69c8da1a-a61c-4f39-855b-c2171387f5d5 · outbound

This paper cites Segmentation of multivariate mixed data via lossy data coding and compression.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Segmentation of multivariate mixed data via lossy data coding and compression

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.726347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.119269Z digest=sha256:2755abee6ee110a763457d87cb5e7bfbb30e588f47a98c49987a4874dbca7c8a

Observation 501377e5-6383-49eb-a7cb-e64f16f231c7 · outbound

This paper cites Pac-bayesian model averaging.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Pac-bayesian model averaging

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.710222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.124143Z digest=sha256:aea2856b2f58f8536b5722123a87023247be9e7e54419a9284d3b6920f5e5f08

Observation 73fe6ed6-92e0-4665-9c76-4021e94b5bc2 · outbound

This paper cites Universal hopfield networks: A general framework for single-shot associative memory models.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Universal hopfield networks: A general framework for single-shot associative memory models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.694650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.128881Z digest=sha256:53c35f9fb70eb9b96754c820cb36317c669f7024efcdf9fc34e50e523a82c447

Observation b50e8e18-e4a0-4c14-8af1-a2d7dc9b730e · outbound

This paper cites Algorithm unrolling: Interpretable, efficient deep learning for signal and image processing.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Algorithm unrolling: Interpretable, efficient deep learning for signal and image processing

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.133706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.133706Z digest=sha256:9790912205eb6e7f073acb028c706eaf7593928f1ededa5646ff1089a4c7fb88

Observation 3344aab9-2dbb-4b06-8a69-8fcaf690904a · outbound

This paper cites Progress measures for grokking via mechanistic interpretability.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Progress measures for grokking via mechanistic interpretability

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.138412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.138412Z digest=sha256:68155e6feec1ace5b5d5c00d121c5f411ce4ee5f805ba6e3bb418f95dd3adf5c

Observation 0584541d-c4ff-4f99-bb6e-cf0ad3df2db9 · outbound

This paper cites Exploring generalization in deep learning.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Exploring generalization in deep learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.143259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.143259Z digest=sha256:7e0f8aad93e2d17c985fbb756317739350f97e08e178263ebcc068c4929e8699

Observation d01d58ef-e403-45e8-bbeb-38bc23881973 · outbound

This paper cites A PAC-bayesian approach to spectrally-normalized margin bounds for neural networks.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models A PAC-bayesian approach to spectrally-normalized margin bounds for neural networks

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.648850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.147803Z digest=sha256:915c44ee6e89dcfb26b1ec1c11f12876161b487f34485389554e6cdd2596c780

Observation a6007c5e-c231-47cd-8547-3df21ea368f2 · outbound

This paper cites Path-sgd: Path-normalized optimization in deep neural networks.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Path-sgd: Path-normalized optimization in deep neural networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.152423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.152423Z digest=sha256:2c7f887b035444d772675c1a14dfbb01b12b3a24a62f4e563205a05881f7c6c6

Observation a29f89a8-f7d3-4be4-a5aa-84cddea91a4d · outbound

This paper cites Norm-based capacity control in neural networks.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Norm-based capacity control in neural networks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.157053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.157053Z digest=sha256:0643bdf8e8d90dabe1c387b6b7dcea275167f3ff70f6b2cf5ab6e05bdb80a343

Observation bf19aabb-2092-4d15-a8ac-824a19b2ef4d · outbound

This paper cites Theoretical foundations of deep learning via sparse representations: A multilayer sparse model and its connection to convolutional neural networks.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Theoretical foundations of deep learning via sparse representations: A multilayer sparse model and its connection to convolutional neural networks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.612840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.161472Z digest=sha256:afeeefff6f5028f1709ba2b1e9d6e95ab937e5804296996a4137f9b694db4425

Observation 68370c3f-74ac-4fc1-865f-be6b2b4cf1b2 · outbound

This paper cites Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.165818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.165818Z digest=sha256:a3577671ca3a93d628b726d38d132bbdddf49ea5165726bd8ac4e9b21b582f89

Observation 9c16a147-3160-4d09-925c-5c70977a2500 · outbound

This paper cites Hopfield networks is all you need.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Hopfield networks is all you need

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.596998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.170405Z digest=sha256:214e2e9b00d68bcfb628a46c6c6a5848f92e7f07b9368ad79d8f840c31089ee2

Observation d047be21-d60e-4e4d-96ac-a3cb81671714 · outbound

This paper cites Unraveling attention via convex duality: Analysis and interpretations of vision transformers.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Unraveling attention via convex duality: Analysis and interpretations of vision transformers

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.582018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.174964Z digest=sha256:b8842adbb13e2f2b84a9e4e22aeb07725b94370df416e64477ed47c75d53dc14

Observation e8f5ac9e-88a5-44d7-a4e7-f40c3afbcb3f · outbound

This paper cites Biological learning in key-value memory networks.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Biological learning in key-value memory networks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.566234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.179474Z digest=sha256:a386fe57affc8fb35e6ef14ba2bae20fdbf9eb2a3c4b802ace249c0923e8e030

Observation 7ceec1d0-12e0-4188-9409-b85ef544dc90 · outbound

This paper cites On the uniform convergence of relative frequencies of events to their probabilities.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models On the uniform convergence of relative frequencies of events to their probabilities

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.183774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.183774Z digest=sha256:60628ecc8a10f7661304c4b4aa01b52955e90dc351c7507dc1bb49c5ebc4d2c4

Observation 45595754-205e-4650-8973-eb26e14d9a7c · outbound

This paper cites Attention is all you need.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Attention is all you need

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.188711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.188711Z digest=sha256:ab3b205383474035f743d828f77581f9cb82c323f84c608e7322f91833afe2ff

Observation 54fbf4f8-9de5-4311-930a-8de8af52ca77 · outbound

This paper cites Interpretability in the wild: a circuit for indirect object identification in GPT-2 small.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Interpretability in the wild: a circuit for indirect object identification in GPT-2 small

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.193362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.193362Z digest=sha256:6a7146d1d7924573349b94445f9aa66bab0b6fd9fd265035b97ff1ea8202d627

Observation 16708be7-7ec9-4462-b6f0-49fd059e8c47 · outbound

This paper cites Thinking like transformers.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Thinking like transformers

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.522435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.198222Z digest=sha256:37e58821d938e9b68f1b0f2c1cab27a410784e1225a8f0b85a205ffc64f6de82

Observation 28500a26-2d01-4722-86b7-8319da737fa4 · outbound

This paper cites Graph neural networks inspired by classical iterative algorithms.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Graph neural networks inspired by classical iterative algorithms

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.505906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.202827Z digest=sha256:d9f9250ea0b8e88a4cbec2826acf9595eb486e7c15d16a577c9904f9d4309519

Observation 5819e8a2-a4ec-4038-be03-839a1c45ad48 · outbound

This paper cites Transformers from an optimization perspective.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Transformers from an optimization perspective

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.490925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.207058Z digest=sha256:0146f0cfd8d907fb2e3675bccd1b2a403dcef85526c713192dfc19e2e5517fdb

Observation 0b214848-862c-4a0d-8529-e072e6f554e2 · outbound

This paper cites Attentionviz: A global view of transformer attention.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Attentionviz: A global view of transformer attention

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.475648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.211699Z digest=sha256:b298fb6448fe86bb54a2de480d7b0d11b425bc337f7fe225aba7a905ee2f453b

Observation 12cf0273-156d-4d0c-b1cc-529692a0833f · outbound

This paper cites White-box transformers via sparse rate reduction.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models White-box transformers via sparse rate reduction

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.460496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.216547Z digest=sha256:2a9d0cd755a5448b30988c2d709a6030650c111476f60198df4cc0537e5de171

Observation 048031ec-59eb-4fc9-a8a0-d30640be511d · outbound

This paper cites Learning diverse and discriminative representations via the principle of maximal coding rate reduction.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Learning diverse and discriminative representations via the principle of maximal coding rate reduction

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T12:32:45.445097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T12:32:45.221157Z digest=sha256:e4dd724c1caaca90faaa0da045327f73a294f99648dcd4d6d4663a249f3b4930

Observation af831bc5-e02c-4fd1-8cad-c1d2357de761 · outbound

This paper cites Unveiling Transformers with LEGO: a synthetic reasoning task.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Unveiling Transformers with LEGO: a synthetic reasoning task

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.226075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.226075Z digest=sha256:a0e95236d8aedc502dc9364ce9443efe144646f0576499e9f28a5fdac3df6985

Observation b7346d03-3e0f-4e9f-9460-f178836751b5 · outbound

This paper cites an unresolved cited work.

An In-depth Investigation of Sparse Rate Reduction in Transformer-like Models Unresolved cited work

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T12:32:45.061926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:32:45.061926Z digest=sha256:664d9b60c230a9ba8b703d4c03e4b003342975f5be49fa173640ae46aec71988

Pith citing papers

No inbound Pith citation observations are available.