Pith. sign in

Paper Citation Record · LEDGER

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT

As of 13 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 0 inbound Pith citation observations for arXiv:2412.17019.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17019 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T06:00:38.833910Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 620b57af-f5b6-4810-aa43-6f8e60cce290 · outbound

This paper cites URL: " 'urlintro :=.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.619270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.619270Z digest=sha256:bfdec8a0acc242709088e0d871d342c1bb9f1817ee9d4d86dc0e88735a544e23

Observation 69f100e7-0477-4de7-8800-2fd1de30ad1a · outbound

This paper cites write newline.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.624968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.624968Z digest=sha256:4d7ebb83fffa4904fb92cd03c3dc91a9dd97f4608ed87623be9c3e535f69641c

Observation 706ef6a7-62ff-416f-a3b2-ba663aba29e7 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.545007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.630305Z digest=sha256:2e79552ea7898be6fe6fedbdd096ccfad8684e193bbdb6842217c52307c7ca3a

Observation 6a5bf91c-6455-4d42-8180-174510c64450 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.531007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.635324Z digest=sha256:ed4b5673538b957b29c6289118e8ad713a74c1083aa51aadca08c3746ae24659

Observation 2bac1fdc-ff85-4d0a-ae15-a49c47a9d475 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.640036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.640036Z digest=sha256:5f9c920f6e1e63b0bce9f16b85e37eae39b96ec08afcb698f9a823af5823839a

Observation 5df77d16-ad49-43b7-8489-d1637188bfd3 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.507491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.644787Z digest=sha256:623fc786b263e57e0ebe8945c6d653fd42c6ea18552b6715253fd9a38b6e257e

Observation ec07b667-a7a2-4274-aaae-d1501f94de2b · outbound

This paper cites A Survey on In-context Learning.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT A Survey on In-context Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.649207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.649207Z digest=sha256:4f72f2384e432530b7858531261045d59cdfc2982fdb2b0167e464633b7cacce

Observation e657096d-f071-4693-b8a6-cd68c7627f2b · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.492220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.654160Z digest=sha256:2263b07f560aac2a19c3a0e9a635a095f83cf9a95c8964a4498e0b0b946a9f55

Observation cde2cdb8-d0f7-4262-95bd-efc7b8ab90cf · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.476823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.659109Z digest=sha256:7e57797f69285be7569368b9aef77584c3df547ab782fa79ff197136dad9bb50

Observation b1699512-76dc-4c96-b206-acc3bb802b79 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.663356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.663356Z digest=sha256:332acc4db4dc795bfd11fd12e94ed50091b483e42e2e363e1e52154309eac17f

Observation d62eb6af-27ca-4680-9989-548c72e5e1cb · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.462566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.667694Z digest=sha256:5512b69ebec76e0b0cb091717dec108e2fb2f5dcfe8383cb6eeb63ef3b4dbb98

Observation bd9352eb-adce-41bf-99fe-49c43e24eb7f · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.672348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.672348Z digest=sha256:8f4c83cda6701479db7014bc263d130f611bd3535a0ed933f9e85114ced9280e

Observation d70bf17f-ef97-4cd6-9465-d07d00584b7f · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.438786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.676668Z digest=sha256:f6ac84462b536268103f382585e186244f667a00022d1879ca480f6322e66d04

Observation 1bc18136-2b5a-4175-8ee6-5e89869569a6 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.424252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.681056Z digest=sha256:1095901d5ec1776713d2c7e4d97f05697f7e34e695976da3378e70a76b739a16

Observation 0f3a803d-6693-44a6-acff-4d3a4bcd4d2a · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.685491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.685491Z digest=sha256:a17ae3c9061289b680b59c2865ee7004fb8eb5f7bec04189c2d7aec6fbc78fe2

Observation 2a505e70-2d60-4edd-a747-39212e57353f · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.690684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.690684Z digest=sha256:538f6ce3ec513bbcba346b7b17d81925fe32804941e1f748b4da6774bb3111a2

Observation 5c9b3b55-af39-42e0-8332-111d9fbe967e · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.695131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.695131Z digest=sha256:39424a396caab0f4ea799b4d7a075fb09d3edfac19e0de7417ba80854b74dc1c

Observation ef196a78-ab84-4a6a-9114-6b3a11bd9ee7 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.399933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.699371Z digest=sha256:0d9d8f8c1c0f9511212027d307bd2ce61d16d5fb63fdd80f78368041a093c168

Observation 7f9ab124-1f46-435c-8ee2-7b700317832a · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.704605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.704605Z digest=sha256:74ed109446cc9b83d8389e1bba32185b323a7e7758095ecb0673048f2d7e4cd2

Observation 3ce826c7-2d05-4e76-a979-72ae659567a7 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.376106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.709122Z digest=sha256:dddb5635a8cfc603b35547a33349e55b59a04f38b9f3a92d88fb79f30d28c147

Observation 6217acc9-50e9-4d68-a09d-c66b0571c45d · outbound

This paper cites One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.713955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.713955Z digest=sha256:cbdff410c933b37429a80a0b616b492772cbb3b5be74520d8e444685e86425f2

Observation 38352ba5-a066-4f92-968a-aa098501394e · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.361428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.718836Z digest=sha256:109e58249ed5ce5949f8f422b7ee136b1c61e43d035ffcb0e2a42a5b5fa620bb

Observation 20e010a7-8ef6-4893-aa81-ef908345985b · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.346516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.722961Z digest=sha256:c4662b88034559110632429cdc9fe481dfd13f0323f4f967b822d4d4a137d52b

Observation 8ee49ffb-ea23-49a2-b39e-474658e5fad4 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.727336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.727336Z digest=sha256:6d9374754b92b04c5efccba04d52d6d2c688bce110c184216fea8bc066cf386b

Observation 20436d33-a522-497c-ab4d-a76e17cd6d34 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.330583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.731732Z digest=sha256:37cb9d76c6c0d189332032b2f858b41c558c269bd02a534bc32f0f8eed0e72f7

Observation 73a5f332-a0d0-42a8-a400-7146b79760e1 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.735956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.735956Z digest=sha256:e12655d7f4614a89b84dcaa3139a842b3d656bb3dd56ce48a9736bfc10f1e5a8

Observation 9dd1fcd9-8c64-4087-af76-6a50821613f9 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.740134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.740134Z digest=sha256:cd41b9981bf0911655891125f89e2a56f3eafc1c5d635585ff2781a6e7ffa446

Observation 431ffcf6-bbb8-4bbe-b5a8-78212ad7b3c6 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.296442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.744794Z digest=sha256:7dcebe23e1fc30d594c30614501457ca987671c1e41b8a1084b751ed6cdaf0ea

Observation e457b792-71e3-4dc9-ab5c-4c2d390b1608 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.282013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.749185Z digest=sha256:b6046c064358f87fa0ab47087612b8d583aad80674dd26b373e08f4810a44c31

Observation 791e9a40-2bd2-4b62-a36d-cddef6ec498e · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.267247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.753314Z digest=sha256:e0f0d60d3a551a6838616353274ec380c97349aebf4c96fdb5668adf8638f153

Observation 6e953ab1-6409-4d05-bcc6-8b46d3c06f15 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.757583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.757583Z digest=sha256:918271d551ab0be6bd2314775a839be8d8465ac5030214c9cfdcddc8d856bcb2

Observation b951175a-0437-4df9-9c9d-a1006038d219 · outbound

This paper cites Transformers as Support Vector Machines.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Transformers as Support Vector Machines

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.761946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.761946Z digest=sha256:a6dac5e530b9d8946f571562cd7e20017beb73f11ee9a9225b4262011fcd087d

Observation 219b7a93-c3e2-4cd3-af51-c5ad16f840bf · outbound

This paper cites Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Scan and Snap: Understanding Training Dynamics and Token Composition in 1-layer Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.766687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.766687Z digest=sha256:6def33ddfac99953a5ad366a5cab1d9fcaa361e00154755042b01737bf94595a

Observation 1ced0d0c-f46c-4abd-8ae9-f5cf4a022ccd · outbound

This paper cites Function Vectors in Large Language Models.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Function Vectors in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.771236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.771236Z digest=sha256:2987dd177101e6c93bde3f74c9b98dfdd52a7678f465709c6846a513e037c827

Observation 77d51de4-20dd-4b9a-90bb-15d267e8f743 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.775927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.775927Z digest=sha256:9c3402b8107489d35bc6c784b28c98ba7dacac351edff269ba4b33adb25a7b08

Observation ca54b518-3220-4538-9c33-f4ff44602fe4 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.780568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.780568Z digest=sha256:7180ae72b85c18fab0421ccffdc92564d3a96d910491b40d08a971d3fe1dd0a6

Observation 58b4fcd1-b840-4e43-af45-84140e90eecc · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.233941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.785023Z digest=sha256:fdac193b9631213ae42bcbf2af794de4b1cb81958231196735d0d243d3523f39

Observation 94f9f378-430c-4e99-9658-2ffbe3eae93a · outbound

This paper cites Neurons in Large Language Models: Dead, N-gram, Positional.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Neurons in Large Language Models: Dead, N-gram, Positional

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.789496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.789496Z digest=sha256:ef640dbb3a230d377e424cc049b32a5823ef79c11ca99ad0499582665fa1c2f6

Observation 2cdca2ea-bb22-4248-9fb5-bd7ef5924ebf · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.793805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.793805Z digest=sha256:f7d326b28a920c5723410d3dde0539fb5c33316987f496f081be3897fbe7b253

Observation a81bf314-ec7d-45e2-ba48-6122f9ca61da · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-11T06:00:39.209156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T06:00:38.797942Z digest=sha256:f2a48a5397d88d81d8bee231bc5f4ac236508adb829ecce54a186c71eee6a199

Observation 8d090988-c255-4db0-94c6-f7a478228d2d · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.801931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.801931Z digest=sha256:3b268c2d846a0251bde31387202dc591d8aa49b93768744676458184d707d77f

Observation 40e42c3c-a4f5-4e5a-b169-babf83398177 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.806577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.806577Z digest=sha256:beca08a837825c4a4fb1fb6295fa9decf66a315e49299a6d49d5da3580fb77a7

Observation ca04a931-f967-496b-81eb-1a7c39cc6507 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.810700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.810700Z digest=sha256:14d7f9b662b65a2fd26b7da508dadb3c3649bcf9ca2ba9e40c44d4bb18d5c0b6

Observation edc2be9a-07e6-40e8-8323-bedf69ba8f90 · outbound

This paper cites Towards Best Practices of Activation Patching in Language Models: Metrics and Methods.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Towards Best Practices of Activation Patching in Language Models: Metrics and Methods

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.815310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.815310Z digest=sha256:5b7df4c9b936e52200745da42f555a9be0a7db3d13a64fa800c46796119a9d95

Observation 007f243f-6f75-4793-908e-fb603c3caab8 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT OPT: Open Pre-trained Transformer Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.819686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.819686Z digest=sha256:22b19ef74bf5530705528099b05e0b2e0b14a4783c8c0d2a879b7f374a56a9ef

Observation 74857f5e-631d-4ebf-a964-3d0285a53349 · outbound

This paper cites @esa (Ref.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT @esa (Ref

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.824200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.824200Z digest=sha256:32c3f9101edef05714a38598dcb4c9ce5495e20a3023dd76c8b33bf3da708768

Observation 8c92a21c-8368-4e8b-81f5-2cb082a091b9 · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.828910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.828910Z digest=sha256:4af9a777929178137e9885a6f637f9ec1b7e8aec35e41f77c0901f195f261949

Observation f341b33a-4a20-4ce9-8c1a-b46f7fcd57cb · outbound

This paper cites an unresolved cited work.

Reversed Attention: On The Gradient Descent Of Attention Layers In GPT Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T06:00:38.833910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T06:00:38.833910Z digest=sha256:867a3a7084b5677fb002be47a11c530fc2804a3ba2d6154f11b22b29f9d83fcf

Pith citing papers

No inbound Pith citation observations are available.