Pith. sign in

Paper Citation Record · LEDGER

Transformers learn in-context by gradient descent

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2212.07677.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.07677 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:10:20.795541Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:29:50.967150Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 32247463-6fcb-4843-910b-dbbee5c8c1cb · inbound

Language Models can Solve Computer Tasks cites this paper.

Language Models can Solve Computer Tasks Transformers learn in-context by gradient descent

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:17:26.845013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T12:17:26.602361Z digest=sha256:b20e1dcb5d0a37dc58a641de61c838633b1d063f36b159001be2ceb4435c023c

Observation 6f4b0280-b613-4a58-8e59-968a275442a1 · inbound

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads cites this paper.

Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads Transformers learn in-context by gradient descent

Reference 171

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:36:18.266138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T10:36:17.764761Z digest=sha256:46181d39f1798088db1b90d3fd971a75af304ef5d42813fb47a299c5b1d8dfa1

Observation 0f0c45dd-5f99-4090-af3d-5d32b5d20624 · inbound

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization cites this paper.

A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers learn in-context by gradient descent

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T06:10:20.795541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T06:10:20.795541Z digest=sha256:075efa909faeea6df0ff6a38af75aa628bd1757ff4409f5eddd5f90c79a80a56

Observation f01fd6a8-5cee-4206-817b-028894d7fec8 · inbound

Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models cites this paper.

Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models Transformers learn in-context by gradient descent

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T21:38:16.278676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:38:16.278676Z digest=sha256:479c78337dd14c695f31698c46b1f09b24d18198d2dda142ebca0bffd3f41461

Observation 2bed3420-686e-41eb-bf14-6f62c9bffe54 · inbound

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently cites this paper.

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently Transformers learn in-context by gradient descent

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:12.669528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:57:12.669528Z digest=sha256:237667d10df2b0af39ef1ea0dde04d346193bbca6f365925a609e44b5cfcfdd1

Observation 5f09dc21-10d2-4352-9f20-797acd8af406 · inbound

When Context Sticks: Studying Interference in In-Context Learning cites this paper.

When Context Sticks: Studying Interference in In-Context Learning Transformers learn in-context by gradient descent

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:09.553762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T08:31:14.231710Z digest=sha256:126ce71fe4b9e402f3b2d91b6b9a48f9fc671a694ff31ae8bb34abe71a0abb6d

Observation 585e48d1-f584-4eee-b9bb-fcb131af08d1 · inbound

SMolLM: Small Language Models Learn Small Molecular Grammar cites this paper.

SMolLM: Small Language Models Learn Small Molecular Grammar Transformers learn in-context by gradient descent

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:01:16.906551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T12:57:47.721361Z digest=sha256:6a916cbae903523b5885a5f52db3d5e11c3f96ace1a8a071a0c15525e5381ed9

Observation 64e90ae5-66f6-45ca-ab58-f592057636a1 · inbound

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning cites this paper.

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning Transformers learn in-context by gradient descent

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:31.950075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:46:21.786972Z digest=sha256:382e11c867dc3f75022e5aa0a3484d3bba484e05961c7011632d6c9f63f9bb36

Observation 99d9c37f-1c2f-4824-b922-a66497b167a9 · inbound

Finite Certificates for In-Context Determinacy and a Threshold Theory of Emergence in Language Models cites this paper.

Finite Certificates for In-Context Determinacy and a Threshold Theory of Emergence in Language Models Transformers learn in-context by gradient descent

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.778045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T19:07:54.236182Z digest=sha256:cc35754862d6c091e5616279462e0d59962eae6cc750604764ef0f2f22f6225c

Observation 74199b95-59dc-4b8c-90ed-120a5b46db56 · inbound

Structure Before Collapse: Transient semantic geometry in next-token prediction cites this paper.

Structure Before Collapse: Transient semantic geometry in next-token prediction Transformers learn in-context by gradient descent

Reference 127

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:29:50.968747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T05:14:07.208255Z digest=sha256:dc1e41ce1bb176b184bd403e0e6c382f33cef3d7b13e39baad3a0c18971659e8

Observation 91c15dc8-0d55-435c-86b5-a4519d575197 · inbound

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models cites this paper.

Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models Transformers learn in-context by gradient descent

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T23:43:11.269139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T23:43:11.269139Z digest=sha256:954bd7575a8f031fc4c78ced256974261b9607dbb2f0992f2b2418348cacd0cb

Observation 937cce38-7c21-47f1-8935-1769e2e8f247 · inbound

In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention cites this paper.

In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention Transformers learn in-context by gradient descent

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T22:19:54.206559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:19:54.206559Z digest=sha256:46b0343f1fe7fa0b62fe2c01f7c2e405f44a3e06b362dffa15de47a829dceb3f

Observation b6f33ba5-fc60-4e30-8f39-4d21678e2008 · inbound

Bayesian Wind Tunnels for Model Selection cites this paper.

Bayesian Wind Tunnels for Model Selection Transformers learn in-context by gradient descent

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T09:19:02.240344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:19:02.240344Z digest=sha256:b5cc7f6500632c8e8fe954ebb3023fb0b4c371f6954efba619ccaf9c3cadd9c9

Observation fb9d9422-1de6-444c-adaa-25b1bd307e80 · inbound

Context-Adaptive Inference: A Unified Statistical and Foundation-Model View cites this paper.

Context-Adaptive Inference: A Unified Statistical and Foundation-Model View Transformers learn in-context by gradient descent

Reference 143

Resolution
unresolved
no resolver link, observed 2026-07-31T23:53:01.264186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:53:01.264186Z digest=sha256:6d9dda54d5f00c43867d03463f19455e5624eb7a2df434c1dd0651d10fe9a6e6

Observation 90560fd6-493c-48fc-92a2-bc098899d474 · inbound

Entangled by Design: Spurious Intra-Variable Signal Routing in Tabular In-Context Learners cites this paper.

Entangled by Design: Spurious Intra-Variable Signal Routing in Tabular In-Context Learners Transformers learn in-context by gradient descent

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T02:16:27.017180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:16:27.017180Z digest=sha256:b5a4b97c80fa75a0ed10ab3fd6aeefd3e1eed0bbca9d9413436e3aa4c05f7dc0