Pith. sign in

Paper Citation Record · LEDGER

Transformers learn in-context by gradient descent

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2212.07677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.07677 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:06:05.127999Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:29:50.967150Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 32247463-6fcb-4843-910b-dbbee5c8c1cb · inbound

Language Models can Solve Computer Tasks cites this paper.

Language Models can Solve Computer Tasks Transformers learn in-context by gradient descent

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:17:26.845013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-17T12:17:26.602361Z digest=sha256:88640517310890b6083028716dc17863b192ef6ad2efb79c71348f2a71aca1e4

Observation 6f4b0280-b613-4a58-8e59-968a275442a1 · inbound

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads cites this paper.

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Transformers learn in-context by gradient descent

Reference 171

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:36:18.266138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-13T10:36:17.764761Z digest=sha256:d8b2f39a51f804ef6f2d34db7d7a76308ae55c09ecac68fbdb7271cde5740cbb

Observation cb185528-a245-4a6d-874a-9e9490c651b1 · inbound

Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder Decoder Perspective cites this paper.

Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder Decoder Perspective Transformers learn in-context by gradient descent

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T14:20:23.787636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:20:23.787636Z digest=sha256:255a2ce6f16924bfbd32618bc48b2c3a0a6142d28215b52c2ff71b69b90046ff

Observation 8d5df9e5-15db-415a-b534-f6ba0ec63fa0 · inbound

In-context learning for medical image segmentation cites this paper.

In-context learning for medical image segmentation Transformers learn in-context by gradient descent

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T13:18:09.773508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:18:09.773508Z digest=sha256:17e72470900e7ae06ac2062b9e19faa3ab3da04123558692859685a10aff5402

Observation 25c9cb9a-76ff-478a-af77-ef1deba04bb8 · inbound

Task Vectors in In-Context Learning: Emergence, Formation, and Benefit cites this paper.

Task Vectors in In-Context Learning: Emergence, Formation, and Benefit Transformers learn in-context by gradient descent

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.417098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.417098Z digest=sha256:3d412c5153985e189fb0452f74dc3f63d15d866882cbebcbcfca3107b6b6e810

Observation 40568d91-16b8-44a8-973a-598430fd89e7 · inbound

Scaling sparse feature circuit finding for in-context learning cites this paper.

Scaling sparse feature circuit finding for in-context learning Transformers learn in-context by gradient descent

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T12:06:05.127999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:06:05.127999Z digest=sha256:32e45a070cf40c7d05c55811a0aa252759b1967c5b3e358215d183487f1e5fb9

Observation 3bbc278e-c836-42ff-85b2-191fca430072 · inbound

ICL CIPHERS: Quantifying "Learning" in In-Context Learning via Substitution Ciphers cites this paper.

ICL CIPHERS: Quantifying "Learning" in In-Context Learning via Substitution Ciphers Transformers learn in-context by gradient descent

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T05:58:34.073164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:58:34.073164Z digest=sha256:eddb7bb5843a84abf69fcb69c83d50d6fa4e925029278ec61db9449986fc490a

Observation 609e8753-af36-4bf4-9a0c-47b8338a16eb · inbound

Sample Complexity and Representation Ability of Test-time Scaling Paradigms cites this paper.

Sample Complexity and Representation Ability of Test-time Scaling Paradigms Transformers learn in-context by gradient descent

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T10:35:38.610485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:35:38.610485Z digest=sha256:686abff658d39fc9cf3b4675156fc49971d3d94667721f151417f6724b5986b1

Observation 0f0c45dd-5f99-4090-af3d-5d32b5d20624 · inbound

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization cites this paper.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers learn in-context by gradient descent

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.795541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.795541Z digest=sha256:f9ae34e1cfca1d09e71a6791a3edad7d574eef367004362a79b2c5da6ad6b777

Observation 3a58dcd2-8f83-4ea3-a2c8-9e9d3f0984ed · inbound

In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly cites this paper.

In-Context Occam's Razor: How Transformers Prefer Simpler Hypotheses on the Fly Transformers learn in-context by gradient descent

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T18:42:34.508285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:42:34.508285Z digest=sha256:bcd4b530abd19c54a70871c5923c9565b6c80e21c98a4830638fe63ba54c0292

Observation f01fd6a8-5cee-4206-817b-028894d7fec8 · inbound

Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models cites this paper.

Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models Transformers learn in-context by gradient descent

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T21:38:16.278676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:38:16.278676Z digest=sha256:b6f4bb9c823065b2c99d57c624bb40308b6bb249a11c4487d4bdbc854268e248

Observation 2bed3420-686e-41eb-bf14-6f62c9bffe54 · inbound

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently cites this paper.

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently Transformers learn in-context by gradient descent

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:12.669528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:57:12.669528Z digest=sha256:09481e77ba71a736b8ae326e89c79dd0e2b32c65c21c3223a5540eaf52666884

Observation 5f09dc21-10d2-4352-9f20-797acd8af406 · inbound

When Context Sticks: Studying Interference in In-Context Learning cites this paper.

When Context Sticks: Studying Interference in In-Context Learning Transformers learn in-context by gradient descent

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:09.553762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T08:31:14.231710Z digest=sha256:b4bab125b30fdbe27299ce72d81e6ca35962424a15f60e2d9651d26a20cd41ae

Observation 585e48d1-f584-4eee-b9bb-fcb131af08d1 · inbound

SMolLM: Small Language Models Learn Small Molecular Grammar cites this paper.

SMolLM: Small Language Models Learn Small Molecular Grammar Transformers learn in-context by gradient descent

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:01:16.906551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T12:57:47.721361Z digest=sha256:c2d930971fc0ae35feba04d3bc31c71016163b5671af72b1f0d3b54a747b2e91

Observation 64e90ae5-66f6-45ca-ab58-f592057636a1 · inbound

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning cites this paper.

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning Transformers learn in-context by gradient descent

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:31.950075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-12T03:46:21.786972Z digest=sha256:0523812ea40d04e8927b40c23801fb3f849288adfad6a2f17046b9c4a79fb0ee

Observation 99d9c37f-1c2f-4824-b922-a66497b167a9 · inbound

Finite Certificates for In-Context Determinacy and a Threshold Theory of Emergence in Language Models cites this paper.

Finite Certificates for In-Context Determinacy and a Threshold Theory of Emergence in Language Models Transformers learn in-context by gradient descent

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.778045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T19:07:54.236182Z digest=sha256:c3867836128320566fa657b21f7e0734d97c49704ee25e387b48c8c6a288cf88

Observation 74199b95-59dc-4b8c-90ed-120a5b46db56 · inbound

Structure Before Collapse: Transient semantic geometry in next-token prediction cites this paper.

Structure Before Collapse: Transient semantic geometry in next-token prediction Transformers learn in-context by gradient descent

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:29:50.968747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T05:14:07.208255Z digest=sha256:5b576318c004a8fa87d9a51b4aded1946621865eaf8975c80836ce0a865b380a

Observation 91c15dc8-0d55-435c-86b5-a4519d575197 · inbound

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models cites this paper.

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models Transformers learn in-context by gradient descent

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T23:43:11.269139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T23:43:11.269139Z digest=sha256:000e6806c11c495b9041daad16963573fab1613057f6100f3f477a12f1601476

Observation 937cce38-7c21-47f1-8935-1769e2e8f247 · inbound

In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention cites this paper.

In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention Transformers learn in-context by gradient descent

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T22:19:54.206559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:19:54.206559Z digest=sha256:053648355c30af0e0a19266ffca957103c909307cb0d501b9182361c8e2c382a

Observation b6f33ba5-fc60-4e30-8f39-4d21678e2008 · inbound

Bayesian Wind Tunnels for Model Selection cites this paper.

Bayesian Wind Tunnels for Model Selection Transformers learn in-context by gradient descent

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T09:19:02.240344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:19:02.240344Z digest=sha256:46fe8e41bca910c4626625238c716a5e595e0e8255dd51ad16b8e8ec17324e56

Observation fb9d9422-1de6-444c-adaa-25b1bd307e80 · inbound

Context-Adaptive Inference: A Unified Statistical and Foundation-Model View cites this paper.

Context-Adaptive Inference: A Unified Statistical and Foundation-Model View Transformers learn in-context by gradient descent

Reference 143

Resolution
unresolved
no resolver link, observed 2026-07-31T23:53:01.264186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:53:01.264186Z digest=sha256:a7ff36e46c120f4bebdceb9e2ea0ec39135a61ddd87274e027f5f77c580e68a3

Observation 90560fd6-493c-48fc-92a2-bc098899d474 · inbound

Entangled by Design: Spurious Intra-Variable Signal Routing in Tabular In-Context Learners cites this paper.

Entangled by Design: Spurious Intra-Variable Signal Routing in Tabular In-Context Learners Transformers learn in-context by gradient descent

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T02:16:27.017180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:16:27.017180Z digest=sha256:17be1844ed299ae2eb444aa404e7452398d5461b729d334b41f61a3f87a07a9d

Observation daf18c9b-1bbb-450d-88d6-8fd8d53d52c1 · inbound

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL cites this paper.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Transformers learn in-context by gradient descent

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.524771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.524771Z digest=sha256:027bcc32a2dbd0460bf48c0e97c75b51cb53a1112427098aeeb00e2673864671