Pith. sign in

Paper Citation Record · LEDGER

Transformers Don't In-Context Learn Least Squares Regression

As of 7 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2507.09440.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.09440 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:03:14.974580Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0b23e4a6-bdd4-4d8f-91da-b93f14f117a4 · outbound

This paper cites What learning algorithm is in-context learning? investigations with linear models.

Transformers Don't In-Context Learn Least Squares Regression What learning algorithm is in-context learning? investigations with linear models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:15.202146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:03:14.923756Z digest=sha256:d0408c4e986db4dff692927578408e142edb6ebd8afa0b64faecfbfdb127a54e

Observation a1ad1156-3c84-4937-8e68-f6da49dfa4f8 · outbound

This paper cites Bayesian scaling laws for in-context learning.

Transformers Don't In-Context Learn Least Squares Regression Bayesian scaling laws for in-context learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.926769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.926769Z digest=sha256:d5b5722067e22d66049f4f71fd2520eb69c409c4640d5b91322e5810aa0fd267

Observation 11ab1bdb-0139-4697-a754-2ce8921c8c86 · outbound

This paper cites Learning theory from first principles.

Transformers Don't In-Context Learn Least Squares Regression Learning theory from first principles

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.929310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.929310Z digest=sha256:9339b3e08c7a4c5ce98dd871c5a1e25d4df77357b6312977ad580a5794caf6cc

Observation 254c2b31-7047-4fe2-bc0e-72c09099767b · outbound

This paper cites Training with noise is equivalent to tikhonov regularization.

Transformers Don't In-Context Learn Least Squares Regression Training with noise is equivalent to tikhonov regularization

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:15.189149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:03:14.932028Z digest=sha256:a39e666f1f8f9221fa5d132e58df881226ae937dca1e25415eec3449f36a40d2

Observation bc88ab7d-e034-4be9-8fd0-2f3599295251 · outbound

This paper cites Language Models are Few-Shot Learners.

Transformers Don't In-Context Learn Least Squares Regression Language Models are Few-Shot Learners

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.934798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.934798Z digest=sha256:d55bf5b2545b830214a3df13e72e44199d73c8c6fc8580588f5385c8c70bfa5c

Observation 0542686e-d2cb-49c4-ac0a-f107f01ae275 · outbound

This paper cites A survey on in-context learning.

Transformers Don't In-Context Learn Least Squares Regression A survey on in-context learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:15.181492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:03:14.937609Z digest=sha256:ab6725bafb0ee4641dd036c272b06ce568e7c8e3c1da6a63dc526990c1294f06

Observation b331e331-95f0-47ef-9d6b-58dd437b793f · outbound

This paper cites A mathematical framework for transformer circuits.

Transformers Don't In-Context Learn Least Squares Regression A mathematical framework for transformer circuits

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.940846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.940846Z digest=sha256:fe89b2529808c32fc4ed439264b6a4484dba0c44d612ca8a069577d6ab2e314c

Observation 949a9976-c58b-4668-aba9-159d7e45e536 · outbound

This paper cites What Can Transformers Learn In-Context? A Case Study of Simple Function Classes.

Transformers Don't In-Context Learn Least Squares Regression What Can Transformers Learn In-Context? A Case Study of Simple Function Classes

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.943494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.943494Z digest=sha256:3c32488b95fabffdda4861b6b1903c8b6349afd737543a730b4bc65e4b416e67

Observation 4dd223b8-b3d4-405d-9892-b5f6c3e7dcd8 · outbound

This paper cites Smith, Vudtiwat Ngampruetikorn, and David J.

Transformers Don't In-Context Learn Least Squares Regression Smith, Vudtiwat Ngampruetikorn, and David J

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:15.168426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:03:14.946522Z digest=sha256:9f1a1671e1ccb74f0e7c7003ee5e236bb30f299113261cc31cd457a6b3c34204

Observation cfb30f18-ab1e-4313-980c-4a1bf410b6fa · outbound

This paper cites Cauchy's method of minimization.

Transformers Don't In-Context Learn Least Squares Regression Cauchy's method of minimization

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:15.160373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:03:14.948848Z digest=sha256:f5655ea6544644974ca9d338cce7c100669cc989aa42c8ace5c50b871aeb151f

Observation 2ae55cb5-ad13-4240-98d7-d4f47519fe95 · outbound

This paper cites Understanding Catastrophic Forgetting in Language Models via Implicit Inference.

Transformers Don't In-Context Learn Least Squares Regression Understanding Catastrophic Forgetting in Language Models via Implicit Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.951300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.951300Z digest=sha256:4e35a6693ff1da20e15552d4adbec2d6bfac8ca6450075f88ea3a83f703de74e

Observation 35b6bd68-56f0-489a-be01-bf93c7a2308a · outbound

This paper cites Predictive multiplicity in classification.

Transformers Don't In-Context Learn Least Squares Regression Predictive multiplicity in classification

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:15.152746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:03:14.954133Z digest=sha256:0d9fe0b29d9a9c7e2e427851ec7be9877638bef4c97888b58e76ae47b6590276

Observation 51c9e342-52d9-4a2a-8e7d-48d5f35d2ecc · outbound

This paper cites Embers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to Solve.

Transformers Don't In-Context Learn Least Squares Regression Embers of Autoregression: Understanding Large Language Models Through the Problem They are Trained to Solve

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.956649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.956649Z digest=sha256:7b20155d34f13c03e3c9661d2cb0a18ec0cb201895198cd92e4939dbc5828818

Observation 1e99849d-ecff-486f-b090-8dd7783026df · outbound

This paper cites More Data Can Hurt for Linear Regression: Sample-wise Double Descent.

Transformers Don't In-Context Learn Least Squares Regression More Data Can Hurt for Linear Regression: Sample-wise Double Descent

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.959393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.959393Z digest=sha256:5641ca62e861e7c4e6ee2d0c77b375b39668570a6033e05a215b78afec7ee49a

Observation c70fea0b-1245-46dc-9998-631bc4476c0c · outbound

This paper cites In-context Learning and Induction Heads.

Transformers Don't In-Context Learn Least Squares Regression In-context Learning and Induction Heads

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.961955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.961955Z digest=sha256:9305ed7e8d09a66ed85faa0af4d429372aaf78da0939fde723f691d7326e0015

Observation f407d3b4-0610-48e5-935a-fecbafd129aa · outbound

This paper cites Pretraining task diversity and the emergence of non-bayesian in-context learning for regression.

Transformers Don't In-Context Learn Least Squares Regression Pretraining task diversity and the emergence of non-bayesian in-context learning for regression

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:03:15.144127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T18:03:14.964504Z digest=sha256:70ed752f90e50bde015005fb709bde994d336fb3a56f2f8558a9d2988def28eb

Observation 4637800a-69f2-4bdf-a36b-dcc194d23490 · outbound

This paper cites Do pretrained Transformers Learn In-Context by Gradient Descent?.

Transformers Don't In-Context Learn Least Squares Regression Do pretrained Transformers Learn In-Context by Gradient Descent?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.967182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.967182Z digest=sha256:3cffe76094fadc7c77465d82fee6f47f3a63c09ff4517a7bb86578a098b565f0

Observation 3645c945-832d-44af-b92f-5d82f0d7e7aa · outbound

This paper cites Attention Is All You Need.

Transformers Don't In-Context Learn Least Squares Regression Attention Is All You Need

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.969689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.969689Z digest=sha256:059125adb7b83a8768ab8e98dadd6b681101fb59891256969a6745538b736935

Observation a89c9943-b8c2-42d6-807a-a9a287c3f86a · outbound

This paper cites Transformers learn in-context by gradient descent.

Transformers Don't In-Context Learn Least Squares Regression Transformers learn in-context by gradient descent

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.972174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.972174Z digest=sha256:de8c394a61ca5b0c9e0811040ca2d13b893e9c31f9600881562c8ec8c3ffed7e

Observation f6148ffc-202a-4a66-8b35-579ec08368da · outbound

This paper cites Qwen3 Technical Report.

Transformers Don't In-Context Learn Least Squares Regression Qwen3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:14.974580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:03:14.974580Z digest=sha256:94b5712a0a0597d8c651204498f151afd7d5273dad28cd44791cfa13fdee0f65

Pith citing papers

No inbound Pith citation observations are available.