Pith. sign in

Paper Citation Record · LEDGER

A Study on ReLU and Softmax in Transformer

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2302.06461.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.06461 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:12:43.102563Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T20:39:00.129807Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation af48f8ee-5148-48c4-9ebb-66334a4af48b · inbound

Mitigating Sycophancy in Decoder-Only Transformer Architectures: Synthetic Data Intervention cites this paper.

Mitigating Sycophancy in Decoder-Only Transformer Architectures: Synthetic Data Intervention A Study on ReLU and Softmax in Transformer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T19:58:47.923221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:58:47.923221Z digest=sha256:756425ad588e9031295018d76a2e8120c3987d44e1698ccf6be8b28c4ec7cf4f

Observation 4178765d-a19c-4b53-a8a4-42790ac2b85a · inbound

Ultra-Sparse Memory Network cites this paper.

Ultra-Sparse Memory Network A Study on ReLU and Softmax in Transformer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T17:40:11.846944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:40:11.846944Z digest=sha256:58c8660a9e5561bbb035853966441f96f1ecf2b14d349e67882e576659fc0775

Observation e1b4c9aa-95da-45c8-b88d-e32eb5a59dbf · inbound

USEFUSE: Uniform Stride for Enhanced Performance in Fused Layer Architecture of Deep Neural Networks cites this paper.

USEFUSE: Uniform Stride for Enhanced Performance in Fused Layer Architecture of Deep Neural Networks A Study on ReLU and Softmax in Transformer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T12:56:15.545841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:56:15.545841Z digest=sha256:97cd938d74de54c430a94416d254757564773316a5c2d8ddd1670b610a44f451

Observation cd3a35f4-7912-4c1f-9107-9eca3e22e658 · inbound

Learning Spectral Methods by Transformers cites this paper.

Learning Spectral Methods by Transformers A Study on ReLU and Softmax in Transformer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:36:51.291777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:36:51.291777Z digest=sha256:ce15d232a9d4857bced0a11b95163a82f7a72e3a90adc23e6207440e2893995c

Observation 65fcb19a-b739-4256-964e-b4716731c418 · inbound

Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models cites this paper.

Softplus Attention with Re-weighting Boosts Length Extrapolation in Large Language Models A Study on ReLU and Softmax in Transformer

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-10T16:15:39.526974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:15:39.526974Z digest=sha256:1e8bff8ff6fa786007196f46437e3535c0f745aed98baf7d370386f217eb3ab3

Observation 29c7de57-65d8-4890-9cb1-47a4e1ecc860 · inbound

ZETA: Leveraging Z-order Curves for Efficient Top-k Attention cites this paper.

ZETA: Leveraging Z-order Curves for Efficient Top-k Attention A Study on ReLU and Softmax in Transformer

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T15:05:36.004953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:05:36.004953Z digest=sha256:133724a40eff4e687b2cd7fc6f7f4c6d93d9eb503cf3db11896b053abddaad57

Observation d5e5a6e9-e6a1-4281-bb3e-37ba8a4b3f99 · inbound

Transformers and Their Roles as Time Series Foundation Models cites this paper.

Transformers and Their Roles as Time Series Foundation Models A Study on ReLU and Softmax in Transformer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T05:01:32.347871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:01:32.347871Z digest=sha256:fa239c593bf8a81129e8de2b7be5f9b4ce7eaeb07be6af812a8fa639cc516f3e

Observation d2efc205-379f-49b3-b908-d62fe3303977 · inbound

On Space Folds of ReLU Neural Networks cites this paper.

On Space Folds of ReLU Neural Networks A Study on ReLU and Softmax in Transformer

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T20:07:26.317513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T20:07:26.317513Z digest=sha256:ecdd632c3e2445fca23ddaf29b57aade959078c11ea234956da5f52d149122c8

Observation 708a9c69-c078-4977-b93c-f3f75d5766aa · inbound

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity cites this paper.

Revisiting Transformers through the Lens of Low Entropy and Dynamic Sparsity A Study on ReLU and Softmax in Transformer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T10:12:43.102563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:12:43.102563Z digest=sha256:3cb7abeb785233eef352f02eb082e2d68e4eb9eadea1455de50b2a46ee936a40

Observation 74b47ab9-1215-4f1e-b7a9-0bc71bac0c75 · inbound

Dual Attention Residual U-Net for Accurate Brain Ultrasound Segmentation in IVH Detection cites this paper.

Dual Attention Residual U-Net for Accurate Brain Ultrasound Segmentation in IVH Detection A Study on ReLU and Softmax in Transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:47.140211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:47.140211Z digest=sha256:60c6d0b323aaa92807e263ba1964178f1b13738deac41693c63ee537a98596a2

Observation 0180009e-c6fa-4221-bed8-b9445945debb · inbound

Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization cites this paper.

Unpacking Softmax: How Temperature Drives Representation Collapse, Compression, and Generalization A Study on ReLU and Softmax in Transformer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:46:59.179741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:46:59.179741Z digest=sha256:40e4933aa1c09db4ccc1d5b1056d2c371013e136c75e0264d6d78a1958d15850

Observation 3357d297-d697-4d61-997f-74378effc196 · inbound

Capacity Matters: a Proof-of-Concept for Transformer Memorization on Real-World Data cites this paper.

Capacity Matters: a Proof-of-Concept for Transformer Memorization on Real-World Data A Study on ReLU and Softmax in Transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:51:54.940894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:51:54.940894Z digest=sha256:9a4a5db02947d503ac533cc670848aa075b4b7efe90a87047557ccfd65989aaf

Observation 3f6b3f8b-f0aa-45b5-8a72-9fc88a62a38b · inbound

Gradient Flow Structure and Quantitative Dynamics of Multi-Head Self-Attention cites this paper.

Gradient Flow Structure and Quantitative Dynamics of Multi-Head Self-Attention A Study on ReLU and Softmax in Transformer

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:35:39.337899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T18:18:48.272183Z digest=sha256:18119089da5fba9508bff776d96b6341ee07a6157e54caf74809f43c2f50405b

Observation f6b15455-ec76-4ef8-9fb3-c98cc4aa090b · inbound

Gradient Flow Structure and Quantitative Dynamics of Multi-Head Self-Attention cites this paper.

Gradient Flow Structure and Quantitative Dynamics of Multi-Head Self-Attention A Study on ReLU and Softmax in Transformer

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:56:00.630980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T00:57:06.133474Z digest=sha256:5abcaf265e159684a3591b99aa85e36f4f444805c9a7429fcaee03201b896a3d

Observation 39a2e880-ee9b-4487-8018-0415cef514cd · inbound

Efficient Implementation of an Adaptive Transformer Accelerator for Massive MIMO Outdoor Localization cites this paper.

Efficient Implementation of an Adaptive Transformer Accelerator for Massive MIMO Outdoor Localization A Study on ReLU and Softmax in Transformer

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-14T18:27:35.259339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T18:27:21.248464Z digest=sha256:48757da6955228f5764212da44eb5a0d851aca33c5e21f01101acbe7c0790016

Observation 364cd5ae-cd99-4cf2-bb9f-12bedd5d303a · inbound

Multi-Headed Transformer Architectures as Time-dependent Wasserstein Gradient Flows cites this paper.

Multi-Headed Transformer Architectures as Time-dependent Wasserstein Gradient Flows A Study on ReLU and Softmax in Transformer

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:39:00.131811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T20:38:06.823151Z digest=sha256:e48b7fb72477a07e0d3a4a68208f5d3780959280dcacbbb2e87727df844b4e12