Pith. sign in

Paper Citation Record · LEDGER

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity

As of 19 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 2 inbound Pith citation observations for arXiv:2504.18929.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18929 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:12:43.140463Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:19:47.277591Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:19:50.554573Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3e4e07e6-1c27-4e70-8a9b-03fbd1aa76b5 · outbound

This paper cites Efficient Large Scale Language Modeling with Mixtures of Experts.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Efficient Large Scale Language Modeling with Mixtures of Experts

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:42.955026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:42.955026Z digest=sha256:73f00cf8571cc6d0c25289f6ee685ad1bb51b09fee7280c48d92087ee219de84

Observation 8151577a-649c-4ac1-8ea6-5920185f6d45 · outbound

This paper cites A Survey on Mixture of Experts in Large Language Models.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity A Survey on Mixture of Experts in Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:42.959822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:42.959822Z digest=sha256:b78bf9250ddfea0a53c26f4ac7a4750b8601718d6e943877fb08bb7849771585

Observation aac28bb3-50ce-41d7-a990-089d16b0f629 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Generating Long Sequences with Sparse Transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:42.965445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:42.965445Z digest=sha256:5c998c753b6bf6efd4139aef9c2b4bbbde2118ed8e5732d6485ace4cd71924b8

Observation bfac2720-24e1-4235-8258-312f8fbc46bb · outbound

This paper cites Palm: Scaling language modeling with pathways.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Palm: Scaling language modeling with pathways

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:42.970245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:42.970245Z digest=sha256:95a2c73180352375cd83780542c04beb6f9d664620f7827bce90dd2191b416f3

Observation c91351ba-6685-40e7-9581-28798aa59ad1 · outbound

This paper cites Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:42.974669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:42.974669Z digest=sha256:ba5c765cbe9112576a0f9710b3029131888fc777ac7135878027bb6093c80f4c

Observation 02e99347-3cb4-4bac-b116-a008d9ca76f3 · outbound

This paper cites Adaptively Sparse Transformers.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Adaptively Sparse Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:42.978885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:42.978885Z digest=sha256:41d3eefa3db7f475a021f5e1e08d22e35fb2e9d5af94d65cc742227252ba4283

Observation ae1e7f62-15b2-473e-b207-bf9e1f265f57 · outbound

This paper cites Batch normalization biases residual blocks towards the identity function in deep networks.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Batch normalization biases residual blocks towards the identity function in deep networks

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:12:43.855896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T10:12:42.984231Z digest=sha256:5d864781dae511041f690cb9d705cd449cc944c71e5e74f40d90d97d86071676

Observation 2f5afe94-15bf-4069-b2db-444becbd5e8e · outbound

This paper cites Language Modeling Is Compression.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Language Modeling Is Compression

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:42.988390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:42.988390Z digest=sha256:50bbb8937384bded11789aa3010e3deb38f282191854d359884ad40a2e001175

Observation 8c413e36-3469-40c7-9a41-70d9dadf229e · outbound

This paper cites Attention is not all you need: Pure attention loses rank doubly exponentially with depth.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Attention is not all you need: Pure attention loses rank doubly exponentially with depth

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:42.992746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:42.992746Z digest=sha256:2c83bb5592a7663820705cb7cf73c039f8d499eb062b3bac7c6e15d18a12a4af

Observation 278ccce7-6d1e-411c-aaca-6cef347f9ce6 · outbound

This paper cites Understanding Emergent Abilities of Language Models from the Loss Perspective.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Understanding Emergent Abilities of Language Models from the Loss Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:42.996527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:42.996527Z digest=sha256:4ebf27e7336ccfb3b745c8b9c79e9d13f5dda0b55591d6fe01c1c9319cf495f4

Observation 146d0c43-1861-472a-a09e-e834ebd4048a · outbound

This paper cites Depth-Adaptive Transformer.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Depth-Adaptive Transformer

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.001124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.001124Z digest=sha256:5f3c75db98cf1581d62c3d188b01de3892acc329090a35dd16266296df243dd7

Observation 996f35e5-2ff3-471f-821b-bf72719f033a · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.005357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.005357Z digest=sha256:4deeb0da19893dd1cbb8812534372cf04b1ca46c22172b7aae91b8c745729dcd

Observation 965d0336-88d0-4f19-96f1-08b0109e038d · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Transformer Feed-Forward Layers Are Key-Value Memories

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.009409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.009409Z digest=sha256:e051da053406b78f312e838d92579a6ee3a81ed4e1311f04a8e6b45cc1e22ed8

Observation c7f42339-194a-443d-bd7e-2a6c803e4ed3 · outbound

This paper cites Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.013569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.013569Z digest=sha256:84926907dac04d81d070abb6cb885ae6b959092f04aa1b67b6253f654b0c1131

Observation c9afd19a-1657-4f88-afb0-c2677c2cf50c · outbound

This paper cites Simplifying Transformer Blocks.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Simplifying Transformer Blocks

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.018009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.018009Z digest=sha256:95263265320d586d06038ab564d3a4ea1ba57c43c43a44a7e555949f00c80b09

Observation a6d013a4-0881-4d74-99a9-611cd3eec20a · outbound

This paper cites Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Deep Transformers without Shortcuts: Modifying Self-attention for Faithful Signal Propagation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.022142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.022142Z digest=sha256:a150078a4561fe390f3e26c49c0cc5a23a020798df4e7839b7acef8ee53d7585

Observation c4ad010c-e7a5-423e-9f36-972e10a35510 · outbound

This paper cites Long short-term memory.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Long short-term memory

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.026268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.026268Z digest=sha256:7f6b42f9cfd2b9de028813fe65f47a051adb14748556d5242334b5c016954fec

Observation e2f41cbf-f7d6-4c1c-9f37-123a4a019e72 · outbound

This paper cites Compression Represents Intelligence Linearly.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Compression Represents Intelligence Linearly

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.030615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.030615Z digest=sha256:5d30aa14bfd84de8bad01aec9e9f2469ffae1b0552e12e0c6f44086f6c060037

Observation 509501dc-d39a-4268-8ffa-8739f46404e1 · outbound

This paper cites Universal artificial intelligence: Sequential decisions based on algorithmic probability.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Universal artificial intelligence: Sequential decisions based on algorithmic probability

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.034975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.034975Z digest=sha256:c82dc629f580afdba3c9dd9ac81c9a844df262740e567e4bbf8335e75e1ce272

Observation 9af05e54-f284-43b8-a6ec-32eb61785791 · outbound

This paper cites The hutter prize.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity The hutter prize

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:12:43.804115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T10:12:43.038860Z digest=sha256:70eb6f88e803d5dce8f92efc9008d629fab47bd1d09fab245557860e5263dc81

Observation acf530e3-acea-410b-817d-27ce2a1347f6 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Adam: A Method for Stochastic Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.043050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.043050Z digest=sha256:088f39471ff40f258797afd6aa68125949b1d61507d8a5c19b9bd887a81d3dd8

Observation dfb99475-540a-47e6-87b4-efed5b4f22c7 · outbound

This paper cites Scaling Laws for Fine-Grained Mixture of Experts.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Scaling Laws for Fine-Grained Mixture of Experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.047153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.047153Z digest=sha256:28f9a5fa7c9919a24619f5d238b02e63820ee81cef093f3913e9124a8e5a5b68

Observation 70399b3a-090f-4eec-9f1e-ea9e9c46a01a · outbound

This paper cites Same pre-training loss, better downstream: Implicit bias matters for language models.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Same pre-training loss, better downstream: Implicit bias matters for language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:12:43.790415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T10:12:43.051596Z digest=sha256:e8cd414c5c1a7d8443fe94873c47045d1e28878c03a8e4898ab7a11b4ab48f34

Observation 8fae963d-9ba5-4c63-88ed-f568b5a9271a · outbound

This paper cites Decoupled Weight Decay Regularization.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Decoupled Weight Decay Regularization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.055761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.055761Z digest=sha256:edad800f3dee5df160c8bdc9091b629c1e489ca2fd4e16527242a5331d2de012

Observation 224c659a-70fe-45ee-8499-43f6d989ba0e · outbound

This paper cites Dynamic sparsity in the brain and machines routing information through neural pathways.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Dynamic sparsity in the brain and machines routing information through neural pathways

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:12:43.776896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T10:12:43.060725Z digest=sha256:798ebf7722e7482cd4c321fa617208d383532875838ae0f720ecfc7a12809f18

Observation 8bd2369d-48ba-4fa1-9542-e15011faa771 · outbound

This paper cites A Theory on Adam Instability in Large-Scale Machine Learning.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity A Theory on Adam Instability in Large-Scale Machine Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.064636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.064636Z digest=sha256:3306e2f7bff6ebc5f7bfea50954ff4ba1e0531519a86393aff2339254ef75430

Observation 97b8ce08-5392-443a-a9e3-f4098ef4ed00 · outbound

This paper cites Signal propagation in transformers: Theoretical perspectives and the role of rank collapse.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Signal propagation in transformers: Theoretical perspectives and the role of rank collapse

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.069184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.069184Z digest=sha256:e6ec99f8ebfe57b3f1f501758415a1c739fe4793aa0660b990a872291265d213

Observation f1169752-61fe-4f7a-bfe0-7c4fb84ef24c · outbound

This paper cites Transformers are Multi-State RNNs.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Transformers are Multi-State RNNs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.073149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.073149Z digest=sha256:b14385dc1a6cb59127062c7d72f956e2665dc00fd5a8abd4ca44277a814f9ea6

Observation 80a37227-342c-4165-b839-2c68f5e8426b · outbound

This paper cites Understanding llm behaviors via compression: Data generation, knowledge acquisition and scaling laws.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Understanding llm behaviors via compression: Data generation, knowledge acquisition and scaling laws

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.077273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.077273Z digest=sha256:12221a53fe8bc1c1fced4792026f3efd0a1132076ffc0a8d6cafc8c4cd27d997

Observation f06b45c6-03d4-4870-a638-be76a1a489e6 · outbound

This paper cites Sparse Sequence-to-Sequence Models.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Sparse Sequence-to-Sequence Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.081065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.081065Z digest=sha256:61751d51789cdb328244bcc2bbcfe463d28e5641385bc44d4baad8e86413823d

Observation 85a88a37-cd63-41ad-98d6-0234081ca429 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.085459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.085459Z digest=sha256:4785bd99fed8331acee0032d96a289ed3e2e76d936f7913d17b7f8844d0c3e59

Observation 08f47968-563a-4d63-bf27-8e1957d12973 · outbound

This paper cites A stochastic approximation method.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity A stochastic approximation method

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.089578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.089578Z digest=sha256:679bc32906b37fd16b45c670a2905ed13fe0595c32c679a28bb8cdd9b27636b8

Observation 652abb8c-24d3-4e9c-8baf-df4441118e9a · outbound

This paper cites A mathematical theory of communication.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity A mathematical theory of communication

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.093336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.093336Z digest=sha256:331b4c2df00f9b0e2d5f3830904acbbe88bfdd8e0b70217d4d3599a6a39e2699

Observation 132ff93b-ef9f-4524-bae8-bd34f1a0fe80 · outbound

This paper cites Implicit Regularization of Gradient Flow on One-Layer Softmax Attention.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Implicit Regularization of Gradient Flow on One-Layer Softmax Attention

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.098087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.098087Z digest=sha256:b0b3ce43b1dd50660790bf3206b2e047e698d6b79b0c59b0667e515dd220a37d

Observation 708a9c69-c078-4977-b93c-f3f75d5766aa · outbound

This paper cites A Study on ReLU and Softmax in Transformer.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity A Study on ReLU and Softmax in Transformer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.102563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.102563Z digest=sha256:42c6d80c5148229fceb07ef0a8b1eb6525941e6714d895bae32f0c5a64c6d70f

Observation 2a26c120-281c-4d76-86e2-95152e86b2bf · outbound

This paper cites An observation on generalization.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity An observation on generalization

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:12:43.733744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T10:12:43.106604Z digest=sha256:71779ecdc66ff7bae57a8066b1a047cffa589e1b1506848dffe4c8add04d14e0

Observation ff593e8a-10b1-43fe-b6cd-1eebebdb226c · outbound

This paper cites Implicit Bias and Fast Convergence Rates for Self-attention.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Implicit Bias and Fast Convergence Rates for Self-attention

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.110604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.110604Z digest=sha256:8106ca7856650543d338a69d83122134aaeb1720a1b60a1df413bf2083a37163

Observation 12cd0dc5-e932-428e-a4fe-75b4e48529aa · outbound

This paper cites Attention is all you need.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Attention is all you need

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.114913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.114913Z digest=sha256:4872a0aeb5a88e7366446475dda2feafc9f11966e7ed12c1254d66634c60b1aa

Observation 7f106bef-6056-45e9-98a9-e54c5afeaf0c · outbound

This paper cites Neurons in Large Language Models: Dead, N-gram, Positional.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Neurons in Large Language Models: Dead, N-gram, Positional

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.119358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.119358Z digest=sha256:e6fa3cd4d35f8fd65c1b30da5a0d4fb0977b3230e9804fd8cd6cdc95e170e4cb

Observation f3e97eca-ac28-4fa8-9ce7-0795194d061c · outbound

This paper cites Attention-only transformers via unrolled subspace denoising.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Attention-only transformers via unrolled subspace denoising

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:12:43.709868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T10:12:43.123770Z digest=sha256:09afed2f237d8031413d48f0ec03cf2a59487c3300a76ad4bfd3aa1123abebf3

Observation dc491a75-8b1e-4d39-b5b5-a8ea86afdfa0 · outbound

This paper cites Scaling white-box transformers for vision.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity Scaling white-box transformers for vision

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:12:43.696358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T10:12:43.127901Z digest=sha256:b1b6265d53b0a81fac39d3266f64244c02f5aee3cd4f0a0eb49925d29a735d62

Observation 8627024e-6241-4db6-9d42-a1235113f332 · outbound

This paper cites White-box transformers via sparse rate reduction.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity White-box transformers via sparse rate reduction

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:12:43.682401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-16T10:12:43.131717Z digest=sha256:61c50d59a908839eaea1f4e43dfd41d31d55c434668e0d599e4f770fbae2b5b2

Observation 57eb23cb-192b-4af9-9e16-9106a6bd7ca7 · outbound

This paper cites ADADELTA: An Adaptive Learning Rate Method.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity ADADELTA: An Adaptive Learning Rate Method

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.135927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.135927Z digest=sha256:44f680815e3bd5dcd84fa08ff0175a2697496e74e0139a34771564dc7ef9ce86

Observation b7083555-3e07-48c3-a1ea-554a1daa3c93 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity OPT: Open Pre-trained Transformer Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.140463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.140463Z digest=sha256:69d3ddf77b18d2c12496b9fac25d48fa8ee0084058a1900303f4ad5f504416f2

Pith citing papers

Observation e172ca69-4a61-4cb6-96fb-c0624d7eceef · inbound

Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning cites this paper.

Demystifying Reasoning Dynamics with Mutual Information: Thinking Tokens are Information Peaks in LLM Reasoning Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:19:50.618498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:19:47.277591Z digest=sha256:3451d2268b46992c719c6c3769b0094e444804c4a72df6bb9994ab3bd69639bc

Observation 6b0e8549-9fc8-445b-9d74-c37554c4e2aa · inbound

Is Your Model Thinking or Just Stagnating? PUMA: Diagnosing Reasoning Pathology via Phase-Momentum Alignment cites this paper.

Is Your Model Thinking or Just Stagnating? PUMA: Diagnosing Reasoning Pathology via Phase-Momentum Alignment Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T18:49:29.554633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:49:29.554633Z digest=sha256:eca250bff0e7247551fb69d7b66a69fc699ca52bfee22432f2f86228cbdde63f