Pith. sign in

Paper Citation Record · LEDGER

Forgetting Transformer: Softmax Attention with a Forget Gate

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2503.02130.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.02130 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:05:17.376218Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:29:30.170846Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 24b7b178-4ab5-4b8d-aec8-1aa39c8c63bd · inbound

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free cites this paper.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.889061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:786a8d5d09760033475d60fc3933b5e4727d12d0345542f4645a9cea123f1852

Observation ff4bc00a-3230-4281-ab2e-6656d2c1f2bb · inbound

Amplify Adjacent Token Differences: Enhancing Long Chain-of-Thought Reasoning with Shift-FFN cites this paper.

Amplify Adjacent Token Differences: Enhancing Long Chain-of-Thought Reasoning with Shift-FFN Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:05:17.376218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:05:17.376218Z digest=sha256:d5d8d41fe09a62ae97a8f1634b4b54e0ecebeef253ca37def04e8550d5de9dce

Observation 98ed79ce-0bf7-42e2-b619-322b8645d438 · inbound

Understanding Transformer from the Perspective of Associative Memory cites this paper.

Understanding Transformer from the Perspective of Associative Memory Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:50.604148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:50.604148Z digest=sha256:325ea1309183e534ffc33ddaa344306b270781524be6591b9d3781de87f50d15

Observation cc102771-46f0-463c-801a-bf70e9e21d79 · inbound

MesaNet: Sequence Modeling by Locally Optimal Test-Time Training cites this paper.

MesaNet: Sequence Modeling by Locally Optimal Test-Time Training Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T10:30:13.759264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:30:13.759264Z digest=sha256:4c2ae85243395ecc6b6a1a0bf33c26f6e2a022700a4a18462228ef7dfb4a0fcf

Observation 3c25cab0-d1df-4c0d-ba57-e46784793316 · inbound

Scaling Context Requires Rethinking Attention cites this paper.

Scaling Context Requires Rethinking Attention Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:00:43.760372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:00:43.760372Z digest=sha256:0b3c9d2cd9995950a70133133a47ff93f5347d0fdbe26c2d2230547e46e29426

Observation c7385187-c710-4120-ac71-4c4b10e2919c · inbound

Distributed Associative Memory via Online Convex Optimization cites this paper.

Distributed Associative Memory via Online Convex Optimization Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:22:37.605998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-18T13:22:08.134149Z digest=sha256:d19e48254ac4991f76f702a415d123805555140cfe43316e30cb9b428e25e15f

Observation 4e60121f-f687-41c5-9d8f-0ca7cd8355d7 · inbound

Kimi Linear: An Expressive, Efficient Attention Architecture cites this paper.

Kimi Linear: An Expressive, Efficient Attention Architecture Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:49:11.201312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T23:49:10.555255Z digest=sha256:e2df04a3cc1b07cf2306f20d7ddace0e6b7af093e8e6d4c5da6d2628d9da5ca9

Observation 2b16414b-ba7c-46af-918c-f79f2bef3ac9 · inbound

Distributed Dynamic Associative Memory via Online Convex Optimization cites this paper.

Distributed Dynamic Associative Memory via Online Convex Optimization Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T19:39:10.606079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:39:10.606079Z digest=sha256:e0e37896de9451423ac77d6b6757151021aa8c462799cefd4b8ee58a5971615d

Observation ceacd1bb-56fd-4f34-b027-260c158f5d0c · inbound

Group Representational Position Encoding cites this paper.

Group Representational Position Encoding Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:08:43.758245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-17T00:04:13.707931Z digest=sha256:e6ada90d19f0efe11a2b5792004e9eb602c9479d9a1f6b2c835c91715adaaba1

Observation 89b5f006-9193-4076-ba99-9fc40fdb6625 · inbound

Kaczmarz Linear Attention cites this paper.

Kaczmarz Linear Attention Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:23.272020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T01:15:58.330766Z digest=sha256:5999358d67525530ddbee29eed6c27a4b5802fd85b9d5d3d61a8b7186ffe8afb

Observation e1c12bfe-1201-4094-9175-f8460b689baa · inbound

Learning-Based Spectrum Cartography in Low Earth Orbit Satellite Networks: An Overview cites this paper.

Learning-Based Spectrum Cartography in Low Earth Orbit Satellite Networks: An Overview Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:26.035674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:52:12.288356Z digest=sha256:dfaef03d2fc86d665e9aea1e3b65bf59d0a8be4886acf89b8c3699a7efeff3f9

Observation 0bc33a37-90bd-42d7-b0cd-e65561bca39a · inbound

Remember to Forget: Gated Adaptive Positional Encoding cites this paper.

Remember to Forget: Gated Adaptive Positional Encoding Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:56:41.213773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:44:31.449357Z digest=sha256:9539766d8af278cb21b6c2c70a846010a790690915ea041e24cdfaf3b7cb013a

Observation 10d227c2-51af-4b41-b279-c32b85f0681d · inbound

Forget Attention: Importance-Aware Attention Is All You Need cites this paper.

Forget Attention: Importance-Aware Attention Is All You Need Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.017603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-28T14:38:40.948032Z digest=sha256:d41a5264ff90ec386d42560bfcf585ccf7422dd3eeb8defabaecb1f7324bf4f8

Observation 8e09ecb3-002e-4367-8504-7bd811cc6ca8 · inbound

Physics-Informed Neural Network with Squeeze-Excitation-like Attention cites this paper.

Physics-Informed Neural Network with Squeeze-Excitation-like Attention Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:30.172882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T18:02:39.721547Z digest=sha256:09d194f846cd927e1b65be90eb4e2b41b461152860fb3656d0a3f88c5a6ea06c

Observation c866d79f-3dee-4e92-afdc-fcb071f05096 · inbound

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation cites this paper.

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T15:30:23.485228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:30:23.485228Z digest=sha256:6730b24251de702f9929763b3f732ea40871887b39c90b04557d2e9d038565d9

Observation f8b427f3-4cd4-4da9-acd9-720214a8eb06 · inbound

Raven: High-Recall Sequence Modeling with Sparse Memory Routing cites this paper.

Raven: High-Recall Sequence Modeling with Sparse Memory Routing Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:01.453174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:01.453174Z digest=sha256:d3577e4cbc255c104dfe8c0b48956a74d37da838e6237d663af579105d1129e5