Pith. sign in

Paper Citation Record · LEDGER

Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 48 inbound Pith citation observations for arXiv:2212.10559.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.10559 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 48 of 48 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:25:20.413368Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

25
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 50dbbee2-35d4-437d-9146-bec7b25b55bf · inbound

Language Models can Solve Computer Tasks cites this paper.

Language Models can Solve Computer Tasks Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:17:26.697618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T12:17:26.602361Z digest=sha256:fd3250961a8bd84c061a8c5ebc50207804d8aa0aa732097076e81708ee55999a

Observation 63fd8123-6f64-4735-bd67-d11d758eb59e · inbound

A Survey of Large Language Models cites this paper.

A Survey of Large Language Models Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:46:40.660977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T22:46:39.268353Z digest=sha256:2e44f5d0637d82ad5a3076a8643180943fe74b358687aa1cd6c7c13e86310367

Observation 15ae8f89-7997-4cf9-9e3c-183d48dbacd2 · inbound

Otter: A Multi-Modal Model with In-Context Instruction Tuning cites this paper.

Otter: A Multi-Modal Model with In-Context Instruction Tuning Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:43:47.829878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T02:43:47.775691Z digest=sha256:b74e384c595ad1460fa5d524aefcb5d1f29f5a049a17fb29970908062e280655

Observation 5a6b0803-177a-4485-a7ac-de0eda355bb2 · inbound

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution cites this paper.

Promptbreeder: Self-Referential Self-Improvement Via Prompt Evolution Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:12:31.422912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-16T08:12:30.984870Z digest=sha256:2a2b417e2f66a5f1a9eee3cbd4795614b8cee6e205051b49dca9789cb7326a93

Observation c806e01e-2fba-43c1-bed1-771b9d5298b4 · inbound

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) cites this paper.

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:26:06.309911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T23:26:06.183574Z digest=sha256:c0986aaa744dfcec82a54e9654a43eea0c55f01c806b95007d0db4442136ea99

Observation 1bbabdb0-3504-4554-b876-d93da813b89a · inbound

StreamAdapter: Efficient Test Time Adaptation from Contextual Streams cites this paper.

StreamAdapter: Efficient Test Time Adaptation from Contextual Streams Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T20:55:17.350225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:55:17.350225Z digest=sha256:5bc876b7bbd89a9ed6b6f263fc39326e291b371e350a79183f4befc9c5c79f18

Observation 938944bf-9a36-44c1-8fa9-5fa8c64a6387 · inbound

In-Context Deep Learning via Transformer Models cites this paper.

In-Context Deep Learning via Transformer Models Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T13:07:59.484648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:07:59.484648Z digest=sha256:ac441eb2687ec8f06962b1a668efecbd7c955f3d9974a8d2d3fb29ec7741b269

Observation 33fe95f2-12f7-4d3e-9378-ec8b7b035bb3 · inbound

Differential learning kinetics govern the transition from memorization to generalization during in-context learning cites this paper.

Differential learning kinetics govern the transition from memorization to generalization during in-context learning Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T10:58:18.267152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:58:18.267152Z digest=sha256:914aeec82971ee50ca63ffa5f18fb60c52f6af283362adbaa50138b46c2a46e8

Observation 6e9d58e9-90b6-4e3a-a8ef-65ed50e84451 · inbound

Fine-Tuning Pre-trained Large Time Series Models for Prediction of Wind Turbine SCADA Data cites this paper.

Fine-Tuning Pre-trained Large Time Series Models for Prediction of Wind Turbine SCADA Data Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T05:28:56.577266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:28:56.577266Z digest=sha256:6533bde11aae71db8ab8fefcf28c053d7ce701a2c6debfbd927f24108fdba073

Observation 6302d483-60fa-4130-8c41-ea088a873d2a · inbound

Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAs cites this paper.

Unlocking Tuning-Free Few-Shot Adaptability in Visual Foundation Models by Recycling Pre-Tuned LoRAs Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T23:49:02.172820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:49:02.172820Z digest=sha256:16d15a03ccd0a2c10c8bc4f62f459b9d6319024c85ee1f6535b89e4e5e182860

Observation aa73d477-9cb0-4367-9f0a-03d69c618b2f · inbound

PromptRefine: Enhancing Few-Shot Performance on Low-Resource Indic Languages with Example Selection from Related Example Banks cites this paper.

PromptRefine: Enhancing Few-Shot Performance on Low-Resource Indic Languages with Example Selection from Related Example Banks Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T20:29:53.880776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:29:53.880776Z digest=sha256:6368f95dfbbac1f5114d411c9fbf98e58e0d38b33474b1ceeeeb92d66ceb98b2

Observation b0ff9743-1eab-4f2c-a239-4261180de51b · inbound

A Comparative Study of Learning Paradigms in Large Language Models via Intrinsic Dimension cites this paper.

A Comparative Study of Learning Paradigms in Large Language Models via Intrinsic Dimension Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T19:55:43.826275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T19:55:43.826275Z digest=sha256:a2f2127b2b97ff9bbfa1a4a3373a7ef037c8189616cc43d01c7c2fbc23ca4e62

Observation 372eb0b4-0541-4972-a825-6f81ed2cc410 · inbound

Small Language Model as Data Prospector for Large Language Model cites this paper.

Small Language Model as Data Prospector for Large Language Model Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T16:33:10.445070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:33:10.445070Z digest=sha256:d2c819b1a6e0c52c26b579d543a92fdbccd95d5ce5a6abe0c4e956db83e44397

Observation c1424f37-f9e1-408e-affd-44fa7ab1b715 · inbound

Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder Decoder Perspective cites this paper.

Emergence and Effectiveness of Task Vectors in In-Context Learning: An Encoder Decoder Perspective Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:20:23.592776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:20:23.592776Z digest=sha256:d8be37f5dc4b32cc7a0db96d56b64bee488de98c9f0db108adb9a2457697f8e2

Observation 5aaac2cf-0b3b-46fe-bcd8-cc2e69e61707 · inbound

Generating Traffic Scenarios via In-Context Learning to Learn Better Motion Planner cites this paper.

Generating Traffic Scenarios via In-Context Learning to Learn Better Motion Planner Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:06:15.854754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:06:15.854754Z digest=sha256:aa077233620ac79a2736bb5baa2af8db2886ea40bdbebbbc65082726c0be0848

Observation ecf0fd16-f61e-4f88-b71e-b968dc717523 · inbound

Hindsight Planner: A Closed-Loop Few-Shot Planner for Embodied Instruction Following cites this paper.

Hindsight Planner: A Closed-Loop Few-Shot Planner for Embodied Instruction Following Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T00:18:20.245017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:18:20.245017Z digest=sha256:87e5f6dc6d372c5043174dcdc12e0bfc67e973826ec1780c617738be16367435

Observation 738ace80-b48c-41af-8936-9d6975f469f3 · inbound

Evolution and The Knightian Blindspot of Machine Learning cites this paper.

Evolution and The Knightian Blindspot of Machine Learning Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T16:30:09.539983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T16:30:09.539983Z digest=sha256:b6fb840cc224dd9b7261c4f5ea3e61b07f0fad545e8bc19006865b5cfd4c6ebe

Observation 92e23f1f-f03e-4a07-b763-62c7c1fe6948 · inbound

StaICC: Standardized Evaluation for Classification Task in In-context Learning cites this paper.

StaICC: Standardized Evaluation for Classification Task in In-context Learning Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T14:06:14.202062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:06:14.202062Z digest=sha256:07880243fe49be107df8d35d8580e3afb0f8a1eba7ae6fa312f70c1d3fd763b0

Observation adc83f7f-d578-48ea-a87c-63e0171acaaa · inbound

PM-MOE: Mixture of Experts on Private Model Parameters for Personalized Federated Learning cites this paper.

PM-MOE: Mixture of Experts on Private Model Parameters for Personalized Federated Learning Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T19:22:37.897107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:22:37.897107Z digest=sha256:5a3f7d420cf0ba9dc7929b2cf3c745bf3f072b531899c6484cb26b2d98f90b2f

Observation 90265562-0c6d-4046-b329-ea28fa3ab8b5 · inbound

Mass-Editing Memory with Attention in Transformers: A cross-lingual exploration of knowledge cites this paper.

Mass-Editing Memory with Attention in Transformers: A cross-lingual exploration of knowledge Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T13:10:15.072691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:10:15.072691Z digest=sha256:d12053a1db5da9ef8de30075b8f59c987116a4aaf53c3b24d5095fca372f4f6e

Observation 13b5fa6c-41bf-43a1-8b11-d7a44164c3ca · inbound

In-context denoising with one-layer transformers: connections between attention and associative memory retrieval cites this paper.

In-context denoising with one-layer transformers: connections between attention and associative memory retrieval Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T20:12:19.783838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:12:19.783838Z digest=sha256:b4fb44bfed4f71a4fee142457ed216adc681d827a10d01e3df58c6a1dc4bb788

Observation 3da142e9-65e9-47e2-99e6-72d879ce859f · inbound

Solving Empirical Bayes via Transformers cites this paper.

Solving Empirical Bayes via Transformers Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:58.250300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:58.250300Z digest=sha256:7303a3a6c41b617cd845fe7bfcea1e13dec49c00515cc8d4d32f481c07397519

Observation 1b2c741d-5e3f-4b31-8fc0-a82672e51db4 · inbound

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model cites this paper.

Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T08:02:24.005948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-19T08:02:23.002090Z digest=sha256:eaa419994adef9b876708a140a428c7c8f6e47c8a3c0d16594e64ef6225e9d3b

Observation f2a18f1c-ddf2-4360-96b5-16cbb6868900 · inbound

The Prompt is Mightier than the Example cites this paper.

The Prompt is Mightier than the Example Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:10.101025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:34:10.101025Z digest=sha256:0fe1253e48417f8674d43eb252ea6db0a16c05507bff3408b548c87e97eff068

Observation 3b5f9650-bca0-4ed6-b0b5-4582b762fb54 · inbound

Optimization-Inspired Few-Shot Adaptation for Large Language Models cites this paper.

Optimization-Inspired Few-Shot Adaptation for Large Language Models Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:24:28.780094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:24:28.780094Z digest=sha256:d741c522232c3ac5e355e288720298502f82eb21af01c96d9fa10e832bf4fc2d

Observation 37303a8d-8c66-443b-b845-98f7ab01a44b · inbound

Transformers as Multi-task Learners: Decoupling Features in Hidden Markov Models cites this paper.

Transformers as Multi-task Learners: Decoupling Features in Hidden Markov Models Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:35.806691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:40:35.806691Z digest=sha256:8517b68bf30d60e0a075d96f36b8736e6c6dc82f3f31d261eece560b5f9e765a

Observation a3cea4c2-1328-4005-bd43-bd54952659da · inbound

Adaptive Task Vectors for Large Language Models cites this paper.

Adaptive Task Vectors for Large Language Models Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:13:00.986510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:13:00.986510Z digest=sha256:80be90cc58f3004a4b5a14f5ae89fffb9fb0c95b0f4d7c7c202903e7eeba951e

Observation bc3b31fb-3efc-45a2-9ef0-166de5dcd8aa · inbound

ConText: Driving In-context Learning for Text Removal and Segmentation cites this paper.

ConText: Driving In-context Learning for Text Removal and Segmentation Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:01:08.177043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:01:08.177043Z digest=sha256:aaf5d07bdd338edc50a0666a279ebdc864b385d3ce883e9d1f66cbacf1d33e54

Observation f530a241-cb5d-45f9-98f0-db8dbcbee486 · inbound

Transformers Meet In-Context Learning: A Universal Approximation Theory cites this paper.

Transformers Meet In-Context Learning: A Universal Approximation Theory Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:33:36.244991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:33:36.244991Z digest=sha256:01a1e09e33e72b836bc5b72393bf05781f2390ab5a0667015e1cb28d662f79b8

Observation 189287f6-70ad-4f4b-a374-df3b7489b3b9 · inbound

Prompting Wireless Networks: Reinforced In-Context Learning for Power Control cites this paper.

Prompting Wireless Networks: Reinforced In-Context Learning for Power Control Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:02.033815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:59:02.033815Z digest=sha256:cf8ad230788080e9eb3a3fa026b992e2ba3532abdf0477bfe239009ccab1c486

Observation 85552c65-25d6-4d33-ba2c-693c03187b7c · inbound

Can Gradient Descent Simulate Prompting? cites this paper.

Can Gradient Descent Simulate Prompting? Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:50.199758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:41:50.199758Z digest=sha256:cb50b22d7ad4382af8507edadec9b050dd03aa85a530903f8c18531f13ab2a18

Observation 9e863460-63fb-42d3-b1ed-6a727aab1fb1 · inbound

Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models cites this paper.

Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T21:38:16.429781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:38:16.429781Z digest=sha256:66ac339db8068b0eacc5c7f9338c4d4b8493269b56c1424f3c3ab702eb42c6a9

Observation f33f08b3-2cb9-41e1-ab57-7d174a2ed617 · inbound

Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning cites this paper.

Enhancing Chain-of-Thought Reasoning with Critical Representation Fine-tuning Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:52:35.454020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:52:35.454020Z digest=sha256:5ea4a25e49a30c2feb4a533cf4cc7fc36b1a4f96e6ea006dfdb3fc7015cbd4f1

Observation deca8069-da0c-437e-813b-8acbc46c6ab6 · inbound

Provable Low-Frequency Bias of In-Context Learning of Representations cites this paper.

Provable Low-Frequency Bias of In-Context Learning of Representations Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T16:37:02.712441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:37:02.712441Z digest=sha256:839ab5e574b4907af8446a732b4cc3e4fcc0a71552cadd5a2fef383421a48e57

Observation 4416e9c6-7f19-4c47-a553-1d75d4f02152 · inbound

FedChip: Federated LLM for Artificial Intelligence Accelerator Chip Design cites this paper.

FedChip: Federated LLM for Artificial Intelligence Accelerator Chip Design Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:49:56.605376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:49:56.605376Z digest=sha256:db1ecbae9702a4142c4fe8a037ee9d441c3ac59cdfff826ceaf7f654c0d92b4b

Observation ba35a10a-481a-42e2-9508-fdf5e60b80a3 · inbound

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention cites this paper.

Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T14:50:22.436932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:50:22.436932Z digest=sha256:48447dadfeeaf270707b0d6076cb15457dacee0296212172c90da0e761372fb9

Observation 035df581-eae6-4b96-a93a-b08e5d629744 · inbound

Meta-learning In-Context Enables Training-Free Cross Subject Brain Decoding cites this paper.

Meta-learning In-Context Enables Training-Free Cross Subject Brain Decoding Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:15:58.798403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T17:42:33.975433Z digest=sha256:ac2373df0111d9c19a3281f4228113a7fda0a0988b16baf07841042a47997f2e

Observation 454ea40e-9d3f-4977-ae76-bf7186f1293e · inbound

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks cites this paper.

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:45:28.215723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-10T13:41:37.942145Z digest=sha256:fb888fa3bcdf2505512b4e3e5af57b48b84b73e8234c082cfdd7ab3e0f9468ed

Observation 09dcf600-7f5c-4d2d-979b-a4de53e955d8 · inbound

When Context Sticks: Studying Interference in In-Context Learning cites this paper.

When Context Sticks: Studying Interference in In-Context Learning Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:09.472739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T08:31:14.231710Z digest=sha256:3bdcfa89d4643804ad7f2ac0a2cb150f4d234699b9b0f5df6cb5e9ea6ad9d4db

Observation 9ea03faa-0a0d-4735-bbe3-5912714f6a17 · inbound

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning cites this paper.

One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:56:32.066177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-12T03:46:21.786972Z digest=sha256:fee4ad40436b6b32b0a8b2167fa2bb042ddf147d50e9352eb58692c4678a9385

Observation 5d61b067-8c82-4d16-8790-406bdc1b0a09 · inbound

Towards Understanding Continual Factual Knowledge Acquisition of Language Models: From Theory to Algorithm cites this paper.

Towards Understanding Continual Factual Knowledge Acquisition of Language Models: From Theory to Algorithm Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:56:25.555016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-12T04:47:54.466097Z digest=sha256:9bffda39ac19edd47c68739a0e8e3ec8ab8d5b73f6e3f84d61d27d4281cd6ad8

Observation 884600e5-edd9-4d87-99db-7326b94408fd · inbound

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space cites this paper.

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:27:19.289708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-13T05:17:34.283917Z digest=sha256:838d95303d498340ba895ae2eae08e9bad3a4f02c6577619ebc591a2fa46f4e9

Observation 52b36e7f-c9cd-4886-878b-1e1dd404afde · inbound

A Human-in-the-Loop Framework for Efficient Prompt Selection in Microscopy Vision-Language Models cites this paper.

A Human-in-the-Loop Framework for Efficient Prompt Selection in Microscopy Vision-Language Models Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:09:46.436740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T07:05:19.381117Z digest=sha256:bde05d8fd3d931ec679f6566768c1987edf3daf82ef90f0f57d0e3167562e675

Observation e2aac111-2d4b-43a6-b390-6a6c03f2c5a9 · inbound

Zeus: Towards Tuning-Free Foundation Model for Time Series Analysis cites this paper.

Zeus: Towards Tuning-Free Foundation Model for Time Series Analysis Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:38:43.288286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-03T17:34:37.552706Z digest=sha256:d6d2e38bfeba243f7ba31d2fa46e710a98e7268375d1a647735faf6c94075c9a

Observation 9826500b-d439-411f-a5ba-babb80ea6390 · inbound

In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention cites this paper.

In-context learning of closed form solution to simple linear regression task using transformer with linear self-attention Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T22:19:53.717359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T22:19:53.717359Z digest=sha256:d3a597f2a0d1d6c18385cca8d3ba2c018eb9931c7645aff7db63374cb456d039

Observation c62be06b-bd3c-4f73-b347-59328c7af37d · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 152

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:33.862109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:33.862109Z digest=sha256:c79c448d5a6ca4889d1058a12983c21fc3f2ac13c2d52252e5dc31a15ac5f314

Observation 81d0ca48-4423-4063-9338-db39dbc3e5a3 · inbound

ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling cites this paper.

ThinkRetrieve: Retrieval-Augmented Reasoning Traces for Test-Time Scaling Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-12T14:10:45.290195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:10:45.290195Z digest=sha256:4079583d8319557f5e117d198dcea52b2bf63d0b698b11407ebf6108b807accb

Observation aba977d2-d585-44a4-b3fc-ef3e7e08c83d · inbound

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL cites this paper.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.413368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.413368Z digest=sha256:512f9a7b8e77519b02dd3df67700a395609989e3a7edac1d637c21121745ee85