Pith. sign in

Paper Citation Record · LEDGER

Forgetting Transformer: Softmax Attention with a Forget Gate

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2503.02130.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.02130 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T19:39:10.606079Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:29:30.170846Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 24b7b178-4ab5-4b8d-aec8-1aa39c8c63bd · inbound

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free cites this paper.

Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T09:04:34.889061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T09:04:34.807225Z digest=sha256:919080bb0732015b9ea087fcaae749e4c2439d52bfaf4601842e5909e2f94304

Observation c7385187-c710-4120-ac71-4c4b10e2919c · inbound

Distributed Associative Memory via Online Convex Optimization cites this paper.

Distributed Associative Memory via Online Convex Optimization Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:22:37.605998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T13:22:08.134149Z digest=sha256:209c26448acf4febb73c27f7965b97e21ee96c3ab3e666b558e5458fa803c7bd

Observation 4e60121f-f687-41c5-9d8f-0ca7cd8355d7 · inbound

Kimi Linear: An Expressive, Efficient Attention Architecture cites this paper.

Kimi Linear: An Expressive, Efficient Attention Architecture Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:49:11.201312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T23:49:10.555255Z digest=sha256:b6879bd1756131bb4cefd6c4dcca690792b0e16c0a50565a33aa8db4ff3a2c8b

Observation 2b16414b-ba7c-46af-918c-f79f2bef3ac9 · inbound

Distributed Dynamic Associative Memory via Online Convex Optimization cites this paper.

Distributed Dynamic Associative Memory via Online Convex Optimization Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T19:39:10.606079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:39:10.606079Z digest=sha256:a1cbd21e45a3b29c252d93b4110d95a224acc88804b28bf5e61d1d5c4f420421

Observation ceacd1bb-56fd-4f34-b027-260c158f5d0c · inbound

Group Representational Position Encoding cites this paper.

Group Representational Position Encoding Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T00:08:43.758245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T00:04:13.707931Z digest=sha256:49d5a37d0ccd19b17667ed3a729bc7db2f792d06ec9dcd77f283994f5900b6bf

Observation 89b5f006-9193-4076-ba99-9fc40fdb6625 · inbound

Kaczmarz Linear Attention cites this paper.

Kaczmarz Linear Attention Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:23.272020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:15:58.330766Z digest=sha256:6baaf9a0a2cecbdf31ae3bac8561bf84e3c6dc8e800317416a34900f2e462dc1

Observation e1c12bfe-1201-4094-9175-f8460b689baa · inbound

Learning-Based Spectrum Cartography in Low Earth Orbit Satellite Networks: An Overview cites this paper.

Learning-Based Spectrum Cartography in Low Earth Orbit Satellite Networks: An Overview Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 131

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:51:26.035674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:52:12.288356Z digest=sha256:738d74e2e22af0f96d04d7612baa8b96b965f6e6a7103538b82c9eab63af0a4e

Observation 0bc33a37-90bd-42d7-b0cd-e65561bca39a · inbound

Remember to Forget: Gated Adaptive Positional Encoding cites this paper.

Remember to Forget: Gated Adaptive Positional Encoding Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:56:41.213773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:44:31.449357Z digest=sha256:6b17af96f2f8df3907127218a3f7d55fd5a953f17fa62617300a95c9cb1d02b7

Observation 10d227c2-51af-4b41-b279-c32b85f0681d · inbound

Forget Attention: Importance-Aware Attention Is All You Need cites this paper.

Forget Attention: Importance-Aware Attention Is All You Need Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:21.017603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T14:38:40.948032Z digest=sha256:a72215df51fa36e826d577048f0b40a31fe29e76b779714ffa414f4fd0af0e6d

Observation 8e09ecb3-002e-4367-8504-7bd811cc6ca8 · inbound

Physics-Informed Neural Network with Squeeze-Excitation-like Attention cites this paper.

Physics-Informed Neural Network with Squeeze-Excitation-like Attention Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:29:30.172882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T18:02:39.721547Z digest=sha256:d8ada6a700d92d457a3d807d3f3fccd5f5d5b9d63f17faa28f5407e3da3e54e8

Observation c866d79f-3dee-4e92-afdc-fcb071f05096 · inbound

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation cites this paper.

Spectral Origins of the Self-Correction Blind Spot in Autoregressive Generation Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T15:30:23.485228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:30:23.485228Z digest=sha256:744950d1e2162a437796bff2b60c94a3df15e4ec594920d74dcaabd0ae34c716

Observation f8b427f3-4cd4-4da9-acd9-720214a8eb06 · inbound

Raven: High-Recall Sequence Modeling with Sparse Memory Routing cites this paper.

Raven: High-Recall Sequence Modeling with Sparse Memory Routing Forgetting Transformer: Softmax Attention with a Forget Gate

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T02:44:01.453174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:44:01.453174Z digest=sha256:73e88212d8e1476ff62af9a3f03c28c54891fbb78632f765014d1360c52f665b