Pith. sign in

Paper Citation Record · LEDGER

Faster Query-Key Learning Sharpens Attention in Self-Attention Models

As of 15 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 1 inbound Pith citation observation for arXiv:2608.06776.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06776 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T14:39:18.403313Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:15:48.968402Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-14T04:15:49.290160Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact3
  • verified fuzzy39
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 05db1fb3-1915-4d4d-8c9c-42167c5cf6af · outbound

This paper cites On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the Optimization of Deep Networks: Implicit Acceleration by Overparameterization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.024092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.024092Z digest=sha256:4ba05645944973b0d0ab65e85b2997c323b57092f454ccf699c45c878cf656f0

Observation 1f04f26b-4471-4ce9-bd8c-9554a47a806a · outbound

This paper cites Self-attention networks localize when qk-eigenspectrum concentrates.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Self-attention networks localize when qk-eigenspectrum concentrates

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.535605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.031138Z digest=sha256:4d5f171bc94057380f9ca615bdd736bbd88bc06b2da35171393a86651d753a9b

Observation cd6eb2e9-caea-43e5-ad11-61852b497662 · outbound

This paper cites Birth of a transformer: a memory viewpoint.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Birth of a transformer: a memory viewpoint

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.513920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.047780Z digest=sha256:b24e906f809ce248c8a28616c5de303feb8b1a355c10b4316b25026484ad360c

Observation 68441002-f3d1-488f-844c-f3690651deba · outbound

This paper cites Language Models are Few-Shot Learners.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Language Models are Few-Shot Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.052596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.052596Z digest=sha256:5b833dd23d624311b19177b015ee666ca07c91db8fef56209517b525e81943ae

Observation c4b2cc58-65ff-4a71-be7b-5feffc94c991 · outbound

This paper cites Universal Transformers.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Universal Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.057149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.057149Z digest=sha256:1e6a6884d6dc345e98d4d8f40cbd540e3d6859729dfd64133d3d04839bf80d93

Observation 067bf91f-cb3a-40d2-ab3e-795be6d9a3c1 · outbound

This paper cites On the optimization and generalization of multi-head attention.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the optimization and generalization of multi-head attention

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.494546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.062406Z digest=sha256:17fd8e5d8e35f93cd780830cdfd7a00cea7403dad86bcd67b9a7dfbd23a33c76

Observation e9cbbafc-1209-481a-8049-06b45c4e698f · outbound

This paper cites On the Optimization and Generalization of Multi-head Attention.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the Optimization and Generalization of Multi-head Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.066978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.066978Z digest=sha256:de6a23194e68ad0c2e462141161525f700c421edddf58a8dbcace03c5764b956

Observation 6d32f044-0fca-489c-87a3-948b4f932167 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.077317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.077317Z digest=sha256:4c2094884406a631b1ff77d973bb703d31065c6eae1efd32c578bbc3540f332e

Observation e466f636-4ef9-494b-b147-0b4bf080a84d · outbound

This paper cites S., Hu, W., and Lee, J.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models S., Hu, W., and Lee, J

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.478602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.082324Z digest=sha256:7e306ddeb4c5844d570449d673013d84eee32f4bb5d2310cef6e86ec8c9b41fe

Observation cb061834-8ad3-4064-af60-2b39a56c5cab · outbound

This paper cites A mathematical framework for transformer circuits.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models A mathematical framework for transformer circuits

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.086746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.086746Z digest=sha256:c1008992b85b6d13d5f848b0ecaf9d7c0f0de1f98c548ff95d7d20dfdd3471be

Observation 943218e2-4634-4659-bb85-c76682efc00d · outbound

This paper cites The emergence of clusters in self-attention dynamics.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models The emergence of clusters in self-attention dynamics

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.450939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.091642Z digest=sha256:75fdc2f2fa67fee5eb321b8fa4c2ee0d7555ace00cf59651241ec1c102e8b7b2

Observation 2d979015-6b34-4794-b445-e49d6f16c970 · outbound

This paper cites and Wallace, B.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models and Wallace, B

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.428665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.096096Z digest=sha256:c6c0f7493c61b0184817b53ec16b280a27eabc8df21dd53356ff89118c4d16cb

Observation aeb751e1-e6a1-464f-8c3e-9554802a074b · outbound

This paper cites Clustering in causal attention masking.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Clustering in causal attention masking

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.411456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.100541Z digest=sha256:0519b826614d9e13338798fba3152ce0a28ee6d2c62295e3b2a3dbd0acded4de

Observation fbb1ea4c-32ed-400f-b5bf-9c2a5f4e0b06 · outbound

This paper cites Transformers in Speech Processing: A Survey.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Transformers in Speech Processing: A Survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.105188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.105188Z digest=sha256:3a818ea81fba0daa6dbdff88f6b8ece4c07faa084ca4cd00a252d62a11edbd56

Observation ed667f7c-1045-4aa7-aeb8-e55d0dee41c1 · outbound

This paper cites How do transformers learn topic structure: towards a mechanistic understanding.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models How do transformers learn topic structure: towards a mechanistic understanding

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.394395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.110194Z digest=sha256:29bfd534ec4032faa6ab7349d4fb1e6e7a788a7dba13982f538d95505c8651d4

Observation 98b45860-5b3b-499c-bdfd-1ab759a7fca4 · outbound

This paper cites Mechanics of Next Token Prediction with Self-Attention.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Mechanics of Next Token Prediction with Self-Attention

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-15T14:39:18.616383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.114791Z digest=sha256:1e5eba2a1da606b2e1aad696190d4be0a421d3fe62adadc4724f8e718dcc61d1

Observation 3b8c52ce-c7a1-41a4-aff1-6afe9c5d5b1b · outbound

This paper cites On the dynamics of training attention models.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the dynamics of training attention models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.375759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.124678Z digest=sha256:cc3a14b652549f66ca974f054fef4109ee91b93db93a4f428e19ecb27f27793f

Observation 075fc0dd-a114-433e-9230-450eb2083918 · outbound

This paper cites M., Biemann, C., Goyal, P., and Mukherjee, A.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models M., Biemann, C., Goyal, P., and Mukherjee, A

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.359952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.129625Z digest=sha256:aee66b3ce9aa8b3c446d69405376bf02c7e6672e21558200e2ddbae14f04c0ce

Observation 51cf7cdf-3d68-464b-8078-1685cf2ed3df · outbound

This paper cites In-context learning and induction heads.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models In-context learning and induction heads

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.134544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.134544Z digest=sha256:0b7f2f1c821ba3dbc8f7a514c8afb8434fab4a32cafa60dc1e147aeed497f4a2

Observation 52b99241-6a03-4d27-87eb-66ac25ca15bc · outbound

This paper cites N., Vashisht, R., and Ramaswamy, H.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models N., Vashisht, R., and Ramaswamy, H

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.334617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.139452Z digest=sha256:51b5e28216d58caaaeea058347667c346fa40c11a984412227606857f4e2a2f3

Observation 5fe12347-73ea-4936-9372-1998bc77a2d9 · outbound

This paper cites Attention is turing complete.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Attention is turing complete

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.317498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.144322Z digest=sha256:e218fb563c6362ef30cd73f701495912b25fab1f2e03bf408a422ea20fb1779f

Observation 1b8e5479-3e76-4b9b-8c91-81e0fd0b96e3 · outbound

This paper cites Improving language understanding by generative pre-training.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Improving language understanding by generative pre-training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.150061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.150061Z digest=sha256:5930d40dc9d24ebd5450f4f0b90d80348d33cca3e4e29b4476b93ba96f311121

Observation 8be5a6b6-9d0b-4cd2-9274-9fba9d662ea9 · outbound

This paper cites Scan and snap: understanding training dynamics and token composition in 1-layer transformer.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Scan and snap: understanding training dynamics and token composition in 1-layer transformer

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.287972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.165853Z digest=sha256:36853cdef0f946f6bafa90582a49ae96235b6be7b60267f25417515cf1ab9319

Observation 1ffb953c-5610-48a1-adee-fdafc300c806 · outbound

This paper cites and Ramaswamy, H.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models and Ramaswamy, H

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.272360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.171140Z digest=sha256:bd06eb87e7620059acf0a5e9bbe1ea5a3b481f86c8870bd05f4d9d6fa1dd6036

Observation 7f632b6b-0afa-41fb-9f36-6a46896f5417 · outbound

This paper cites N., Kaiser, L.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models N., Kaiser, L

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.176192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.176192Z digest=sha256:89b66618b28b5968887f71df8c152e11d20a3f2278af63f46ca577873abcf8b6

Observation b029877b-2e9f-4876-aa2c-e7bad65ca26b · outbound

This paper cites Are Transformers universal approximators of sequence-to-sequence functions?.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Are Transformers universal approximators of sequence-to-sequence functions?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.186873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.186873Z digest=sha256:a5dba73c1b2fa5466d76766b37185f81ac75c9f26fb1b1a35374f4701e74e95a

Observation 85017a99-ae25-4307-a77e-cbcdae1831a7 · outbound

This paper cites Attention is All you Need , url =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Attention is All you Need , url =

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.247018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.191872Z digest=sha256:8fb9b0e23a3121121a78dae7800603fd1bbfd99393896f2db9f5b25778c6f079

Observation a9962a82-cb5f-4b3b-bfb0-52dc99b649be · outbound

This paper cites ArXiv , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.229176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.197816Z digest=sha256:1cb828846116b38e5583eb5e56bd5fff4cfc850244442915b73763138c632e8e

Observation 2ce633c1-dec7-4c60-afb2-4cd12f19af38 · outbound

This paper cites and Kumar, Sanjiv , biburl =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models and Kumar, Sanjiv , biburl =

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.213639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.203365Z digest=sha256:82bbb853251266775ebf3aead02e07c347758d3ea4f905f8f16623496d954cb7

Observation 5d45f57c-fe0f-49be-8482-35a94a84243d · outbound

This paper cites On the Computational Power of Transformers and Its Implications in Sequence Modeling.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the Computational Power of Transformers and Its Implications in Sequence Modeling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.208510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.208510Z digest=sha256:8ede1ec7484c1e83af7bc5fb6f707ec47b08dc8d1bb196c277de05c4417bdd51

Observation d4cb7314-b594-4f33-8148-cd3c81dac9e1 · outbound

This paper cites On the A bility and L imitations of T ransformers to R ecognize F ormal L anguages.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models On the A bility and L imitations of T ransformers to R ecognize F ormal L anguages

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.213259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.213259Z digest=sha256:35b387bd71d388623c8237d31c67026e92092ad327e3c6fac565eb915fe74e78

Observation 3b8377c4-f28d-453a-acb9-f3fd5d7cd7fd · outbound

This paper cites ArXiv , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.197617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.219271Z digest=sha256:c7b5bff9e3b1c898361cf7e9cdb133bcffe85eb67737e765ccbdaa9314bb0429

Observation d77dc4ac-07c0-43b2-ab88-db6a58df865a · outbound

This paper cites Attention is turing complete , year =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Attention is turing complete , year =

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.179134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.224200Z digest=sha256:4b6e820dafc9fba0add4f5b04cadb323861420ee3165c6e20c9290c6272a864d

Observation 61fbacbd-5c02-4471-afc2-08d979fb5a98 · outbound

This paper cites and Ba, Jimmy , biburl =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models and Ba, Jimmy , biburl =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.161972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.231213Z digest=sha256:d1ea5dc82fe0dea02389d7e6210ed19eca6bf7ddda4ada00ed11c516674da4ef

Observation c4106af3-60f1-4649-bdb5-0c2563327275 · outbound

This paper cites an unresolved cited work.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.237378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.237378Z digest=sha256:f03ceb881929573edca7b94026c9b6bc528a423cfd7fa898e6c363f7f350913a

Observation b29081aa-6747-4c23-881c-16e2651e7f3f · outbound

This paper cites 2021 , journal=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models 2021 , journal=

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.242650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.242650Z digest=sha256:c55329615a616fd255cdb5ac7706c21855cb4b0795ae50c622fccb6e4a165d4e

Observation a08a058c-d101-4628-a59b-99d93c1190fb · outbound

This paper cites In-context Learning and Induction Heads.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models In-context Learning and Induction Heads

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.122060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.247761Z digest=sha256:33e97745cf4614cb227e924a55e6ebf5c7291ee3a9260b694fed2e45057b2cb2

Observation 2eb62838-e40d-4716-ab1a-13ba5b48a26f · outbound

This paper cites Birth of a transformer: a memory viewpoint , year =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Birth of a transformer: a memory viewpoint , year =

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.103887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.253256Z digest=sha256:3883f866cfa3ac205408292a4bb02dff742166c58b8491584e024ee28c2e4da4

Observation 9be6d413-dccf-4002-8194-ad25097c2cab · outbound

This paper cites SQ u AD : 100,000+ Questions for Machine Comprehension of Text.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models SQ u AD : 100,000+ Questions for Machine Comprehension of Text

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.258382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.258382Z digest=sha256:adb357807d778faeaebb52e5ec2bd7f097ab4e121f7d73cfff519b64b48111c0

Observation 068cf1d7-7f45-4414-a4b1-7db80a11e364 · outbound

This paper cites Assessing the Ability of LSTM s to Learn Syntax-Sensitive Dependencies.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Assessing the Ability of LSTM s to Learn Syntax-Sensitive Dependencies

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.262761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.262761Z digest=sha256:ed3d2ba21cbc336c3f51e952bfafb33743533397136dec8ab243fca276937d14

Observation 14f7765e-f0c2-4d95-b874-b2dd21a4be76 · outbound

This paper cites AAAI Conference on Artificial Intelligence , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models AAAI Conference on Artificial Intelligence , year=

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.086276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.267286Z digest=sha256:ae30b8a57affaa2f9702374f70bce9445b4e6baba3c69d27d3d5bf7e851296b7

Observation 23c39edd-bddc-4ee3-a11c-98918278029e · outbound

This paper cites Proceedings of the 37th International Conference on Neural Information Processing Systems , articleno =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the 37th International Conference on Neural Information Processing Systems , articleno =

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.068323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.272953Z digest=sha256:46b44a915337c7f8d53e8c017180f5c083d8964e92ff07fd8655f37ef57ad533

Observation 9c8e705f-8870-49fc-9ad2-ebe22dc20016 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding , url =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding , url =

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.277783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.277783Z digest=sha256:aca74bf275b4930cf099d677e14dcba5a21eae14cc1dd41090e3df8c859b8c03

Observation 9a8d2a32-6d35-4dcc-a34c-982f5f57ddba · outbound

This paper cites Language Models are Few-Shot Learners.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Language Models are Few-Shot Learners

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.284319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.284319Z digest=sha256:7567a714b7c562045682914035ccc7a4db36f794b17c5d75b4c882b87a34016b

Observation 2f7b9722-6608-46af-8395-362f5039b842 · outbound

This paper cites ArXiv , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.289534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.289534Z digest=sha256:749958d65ff3faa9b9998f7725db5861c315d3c26275e02ea62a7fe58cdc8035

Observation b45d4986-e646-41c3-bde9-6c52bac9508d · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , month =

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.294174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.294174Z digest=sha256:408d7bfefd1ff2c69e6880d1395d45f083caa6ed136a3a64d579aac7c533f8ff

Observation 5b585f6d-0d2c-489a-8b1f-e72b224ff1f3 · outbound

This paper cites ArXiv , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:19.016311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.299668Z digest=sha256:3c3674c221c7fe9e4b2c36be343228f2b668c2775a32e6fda3849182b971bfb0

Observation aa5540ea-bd45-4b08-9017-ac9559d068d4 · outbound

This paper cites and Hu, Wei and Lee, Jason D.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models and Hu, Wei and Lee, Jason D

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.997570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.304807Z digest=sha256:ba7edfa57034706fc8aa1232a5ded45c2770e61f8eb703707d17ad785db21df8

Observation f437a744-300c-42be-bb1a-ee39277e3456 · outbound

This paper cites 2022 , journal=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models 2022 , journal=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.310228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.310228Z digest=sha256:6e9c729150c430e31f182fd28f30f5ce2d1a6d89285b83b0d70b063923a72336

Observation 7163e939-6afe-489e-9b05-27f555586377 · outbound

This paper cites ArXiv , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.970434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.316089Z digest=sha256:eb7dada081552f27ddfea027457065d094de7850be416613ff801666a2e34800

Observation d1602f52-1aff-4bea-b548-b596476f0b21 · outbound

This paper cites 2025 , eprint=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models 2025 , eprint=

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.956155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.321322Z digest=sha256:d4a6adb9a1599bb77ad0794465f171619cf5fefdb7c4b643ce6e1a0b3b9de812

Observation e2296ce4-2ff9-4269-9729-6c34e0085c92 · outbound

This paper cites Attention is not not Explanation.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Attention is not not Explanation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.327189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.327189Z digest=sha256:2f5c880167d3140611065cac220d765f85088a2bca1f3987dc17965b5cf9aa31

Observation 521285b6-e782-406a-bf6d-d1f40b979d6a · outbound

This paper cites North American Chapter of the Association for Computational Linguistics , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models North American Chapter of the Association for Computational Linguistics , year=

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.941436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.332384Z digest=sha256:d8fe11d095b303ed69c95538541f9aee416a03fe47c3676781bc872c50c20728

Observation 07c2fcb9-a389-46bd-834e-4e59d86f645a · outbound

This paper cites Is Attention Interpretable?.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Is Attention Interpretable?

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.338319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.338319Z digest=sha256:e33ee1b12958ab24d4eba47e59f50a7ef9b294cde8315605634630670cf5ad3f

Observation 300fb36e-0524-4fe8-af54-54b6a7c665d3 · outbound

This paper cites 2024 , eprint=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models 2024 , eprint=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.925397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.343703Z digest=sha256:ca25ca7271e89c00e1606738baa19418d7a6d3add46b9e96ff0b4ba538d230ce

Observation 68feb1c8-a306-497d-a414-4a0f00296363 · outbound

This paper cites ArXiv , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models ArXiv , year=

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.909569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.348421Z digest=sha256:7a58af163ccb93b2728516cf9f00c05a11e9eafd1647c576494d1dce9b9935c2

Observation 381831d6-f1f6-428a-a8e6-5cfe10bfe1cf · outbound

This paper cites Proceedings of the 41st International Conference on Machine Learning , articleno =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the 41st International Conference on Machine Learning , articleno =

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.893694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.353396Z digest=sha256:8057f3546070f08b10873a3255b56876a776ee0769aac09a312c3541293bd1a1

Observation 18b0ed52-e575-4366-87cd-110ff5be9c6c · outbound

This paper cites Transactions on Machine Learning Research , issn=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Transactions on Machine Learning Research , issn=

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.874535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.358371Z digest=sha256:f042454718769ae8a14e78e9aa3b04c01abfbdc455c651a73ffba3d918b96ed2

Observation 15745298-06e3-45ef-afa0-a94c60961d3f · outbound

This paper cites Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence , articleno =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence , articleno =

Reference 67

Resolution
verified exact
doi, observed 2026-08-15T14:39:18.468085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.362713Z digest=sha256:6dd97926cf84193e579373d2d433c468730cd8e6675073825e810f4af3d83a2f

Observation 44b4543b-df80-4580-911e-9d83d571ef11 · outbound

This paper cites International Conference on Learning Representations , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models International Conference on Learning Representations , year=

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.854925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.367236Z digest=sha256:1655b0d96a6168ab4d3d5dd9436cb0f206b0283f89b6f31bded5a536cecdc4e7

Observation 9e226e27-f37b-4687-8bf0-40f6d8272bc6 · outbound

This paper cites an unresolved cited work.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Unresolved cited work

Reference 69

Resolution
verified exact
doi, observed 2026-08-15T14:39:18.448554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.371684Z digest=sha256:e3366186161e5d3a4cff27919c0bf97bca1b7dfda0bf5861e6e1884d7fc921ff

Observation c9a9e8f5-7966-4a6f-992f-1ef36ca5cdb3 · outbound

This paper cites Proceedings of the 37th International Conference on Neural Information Processing Systems , articleno =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the 37th International Conference on Neural Information Processing Systems , articleno =

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.835752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.376632Z digest=sha256:c7595d11e4df3def83dc0a8e4753a3691e6e9f91289457e233a00092d96c0e7b

Observation 3b8a100d-98f6-4ca0-a665-cff24368d6fa · outbound

This paper cites Proceedings of the 38th International Conference on Neural Information Processing Systems , articleno =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the 38th International Conference on Neural Information Processing Systems , articleno =

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.818156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.381279Z digest=sha256:204dfc5e3a43ce15656ab823772e4bbc4539b73f6f8e195a13dfbddc53201d56

Observation af580012-a011-4851-bec7-bd87cb694164 · outbound

This paper cites Proceedings of the 40th International Conference on Machine Learning , articleno =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of the 40th International Conference on Machine Learning , articleno =

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.796993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.385781Z digest=sha256:c87ed948ed5f14f569f2a8b749ab16660e4dfe6fbea78540023149fd92ae8ccb

Observation cd807fa5-061e-4dba-b80b-2eb806503025 · outbound

This paper cites Quantifying Attention Flow in Transformers.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Quantifying Attention Flow in Transformers

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.390155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.390155Z digest=sha256:7ad1ebbd891d242abc18fc488f202ea1715a8f1fa29269203a206a0383bd959c

Observation b166557b-8498-4d35-bc1e-b76757058ca0 · outbound

This paper cites ERASER : A Benchmark to Evaluate Rationalized NLP Models.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models ERASER : A Benchmark to Evaluate Rationalized NLP Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-15T14:39:18.394573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:39:18.394573Z digest=sha256:79900acd0701a3da1ed553df827ede579a525f33b8100b5192d78c7875a546b2

Observation de0c266e-9a1d-4525-b9c9-abda41873ce3 · outbound

This paper cites Proceedings of The 14th Asian Conference on Machine Learning , pages =.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models Proceedings of The 14th Asian Conference on Machine Learning , pages =

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.778188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.398617Z digest=sha256:b21f387d811e80d67ab665826e6e91eed0368df14c1a4ed0d55685cfc813b594

Observation 45bb5563-3142-4a07-b98a-769964221409 · outbound

This paper cites European Conference on Artificial Intelligence , year=.

Faster Query-Key Learning Sharpens Attention in Self-Attention Models European Conference on Artificial Intelligence , year=

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T14:39:18.758095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-15T14:39:18.403313Z digest=sha256:1b7568b0cebd2f0791712da32d33b23c29d9a031a97a1cd727c99ea3be1d0dd3

Pith citing papers

Observation f6db4ff7-a0e7-480d-a701-6470ca84d468 · inbound

Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference cites this paper.

Power law graph attention: exact generalization of scaled dot-product attention, empirical collapse at inference Faster Query-Key Learning Sharpens Attention in Self-Attention Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:15:49.294748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:15:48.968402Z digest=sha256:e1110e7188d74717ec7de677fa9f905ac990990bc27fcf7e45a8937b5753516b