Pith. sign in

Paper Citation Record · LEDGER

Uncovering mesa-optimization algorithms in Transformers

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 29 inbound Pith citation observations for arXiv:2309.05858.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.05858 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:24:01.667776Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

5
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9210d8e8-b104-4e33-9716-b18b348d126a · inbound

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models cites this paper.

Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models Uncovering mesa-optimization algorithms in Transformers

Reference 143

Resolution
verified exact
arxiv_id, observed 2026-05-17T14:43:30.172724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-17T14:43:29.496457Z digest=sha256:e0288f0235e1a563fd36d4caa24fc14dae40181771fbebb4f4691ea6d3909094

Observation 7968d552-b9cb-4bf2-bb96-45e1e6e77504 · inbound

ICLR: In-Context Learning of Representations cites this paper.

ICLR: In-Context Learning of Representations Uncovering mesa-optimization algorithms in Transformers

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-10T23:24:01.667776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:24:01.667776Z digest=sha256:89a569d09bcd00244b895e5a025be32b77c1b5d236b39c0555381ab306987991

Observation 59fecf7f-02a7-4741-806d-58087e8e0ca2 · inbound

Task Vectors in In-Context Learning: Emergence, Formation, and Benefit cites this paper.

Task Vectors in In-Context Learning: Emergence, Formation, and Benefit Uncovering mesa-optimization algorithms in Transformers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T20:15:28.421835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:15:28.421835Z digest=sha256:71641a9fb4d37cf526f313f5a7eb0d56bcb9b03ff6e9de2b1b089867d143d78b

Observation 322ea548-315f-4b05-a5d7-8a836bd58a5c · inbound

Test-time regression: a unifying framework for designing sequence models with associative memory cites this paper.

Test-time regression: a unifying framework for designing sequence models with associative memory Uncovering mesa-optimization algorithms in Transformers

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-10T17:22:07.351612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T17:22:07.351612Z digest=sha256:f005a3a9b1691720acb45e7bc402a38a5377e522977603f90bd4d61dd50bdeea

Observation 2b764ef6-7bca-41a6-a900-94a3a1d557b1 · inbound

A Unified Perspective on the Dynamics of Deep Transformers cites this paper.

A Unified Perspective on the Dynamics of Deep Transformers Uncovering mesa-optimization algorithms in Transformers

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T00:04:10.958600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:04:10.958600Z digest=sha256:2c9f2dd2a26dd6291c143720ed1daeee8d19f3df5bd344a24940fca1d6e9f7e3

Observation 54526c98-70c2-4a61-bd6d-01c2449b7f44 · inbound

Amortized In-Context Bayesian Posterior Estimation cites this paper.

Amortized In-Context Bayesian Posterior Estimation Uncovering mesa-optimization algorithms in Transformers

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-08T15:01:42.885301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:01:42.885301Z digest=sha256:7e5e47d52816edeca1224815047c758c6f934e430d14707dc8c408e862122730

Observation 1bbfec1d-a87b-449d-bc64-7810b77d1f86 · inbound

Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence cites this paper.

Beyond Induction Heads: In-Context Meta Learning Induces Multi-Phase Circuit Emergence Uncovering mesa-optimization algorithms in Transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:00:58.100805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:00:58.100805Z digest=sha256:db24552b17b3200c2bdb5a4ba7b4688852282861c27230d8f051603b06e72c4e

Observation 998052c0-e19d-47ee-994a-aeab003d2d99 · inbound

The Role of Diversity in In-Context Learning for Large Language Models cites this paper.

The Role of Diversity in In-Context Learning for Large Language Models Uncovering mesa-optimization algorithms in Transformers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:38.439808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:21:38.439808Z digest=sha256:84c85c2589317ab47d916ad207491bf13101a7984ee71f5ee3378fb613f76cb3

Observation ee67979d-b995-4e61-9598-ffee220492d9 · inbound

ATLAS: Learning to Optimally Memorize the Context at Test Time cites this paper.

ATLAS: Learning to Optimally Memorize the Context at Test Time Uncovering mesa-optimization algorithms in Transformers

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:21.592813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:21.592813Z digest=sha256:0ea96ff3481d9879fdd4bf3aaa5dbe32eec68febefb6f9d6fb6d217df24b1be9

Observation 6fdb14e6-8e63-428b-9224-93dcb347cd95 · inbound

Transformers Meet In-Context Learning: A Universal Approximation Theory cites this paper.

Transformers Meet In-Context Learning: A Universal Approximation Theory Uncovering mesa-optimization algorithms in Transformers

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T10:33:42.468990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:33:42.468990Z digest=sha256:4f05ee748f576b1640eb9efbb0363286ec467c71e010eb0aa7ede10328116a2d

Observation 2d90cb9a-ff6a-4f3a-b45a-132f60b1c6b7 · inbound

MesaNet: Sequence Modeling by Locally Optimal Test-Time Training cites this paper.

MesaNet: Sequence Modeling by Locally Optimal Test-Time Training Uncovering mesa-optimization algorithms in Transformers

Reference 113

Resolution
unresolved
no resolver link, observed 2026-08-07T10:30:13.976683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:30:13.976683Z digest=sha256:294543eeafad38fe787851ec50c32fb1b1c39df14e654294f3ea407fdd79f245

Observation ff999589-2f12-448d-be26-ea9e35402d55 · inbound

Brewing Knowledge in Context: Distillation Perspectives on In-Context Learning cites this paper.

Brewing Knowledge in Context: Distillation Perspectives on In-Context Learning Uncovering mesa-optimization algorithms in Transformers

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:11:13.097520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:11:13.097520Z digest=sha256:2b7e6412d63e0b277aa50fbff264972654f598913bf09efe10814947ca33afa5

Observation 8019fbb2-e19b-4c7a-903c-83027a45460e · inbound

Position: We Need An Algorithmic Understanding of Generative AI cites this paper.

Position: We Need An Algorithmic Understanding of Generative AI Uncovering mesa-optimization algorithms in Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:41:33.871059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:41:33.871059Z digest=sha256:0f9f43927b3fca436a85f33a2449678d293c3cccdbd8b57b8c4a930b7100c018

Observation dd6a7791-dec0-4451-a94b-9de310699d76 · inbound

Selective Induction Heads: How Transformers Select Causal Structures In Context cites this paper.

Selective Induction Heads: How Transformers Select Causal Structures In Context Uncovering mesa-optimization algorithms in Transformers

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T21:11:45.295838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T21:11:45.295838Z digest=sha256:51b80850f3bb703ad703a3e68c7bc208254653f1b51d4e3f213526ec0810f61f

Observation ee68abba-7ec3-481d-a96d-8928d03f439a · inbound

Learning to Remember, Learn, and Forget in Attention-Based Models cites this paper.

Learning to Remember, Learn, and Forget in Attention-Based Models Uncovering mesa-optimization algorithms in Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T03:16:07.695368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:16:07.695368Z digest=sha256:fac79f5ffaf070896c875e0eb3956893b0ce86f93892638b020d77573ec6231b

Observation 197bb3a4-9682-4dd5-bc3c-dd00706898e0 · inbound

Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning cites this paper.

Beyond Test-Time Memory: State-Space Optimal Control for LLM Reasoning Uncovering mesa-optimization algorithms in Transformers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-15T12:09:16.468055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T12:09:16.468055Z digest=sha256:2c6ec67ae8da80a0cb526fe4a3bf2d344704ccb2398d9ef46341abb774535276

Observation df420f82-5bac-4bfe-85f2-65e78f58701d · inbound

Preconditioned DeltaNet: Curvature-aware Sequence Modeling for Linear Recurrences cites this paper.

Preconditioned DeltaNet: Curvature-aware Sequence Modeling for Linear Recurrences Uncovering mesa-optimization algorithms in Transformers

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:19:46.569906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T00:19:24.366922Z digest=sha256:0dc492cda600777fd4457d27d47da036b5dfe8674c22f3d62d8e9f6c46aeba77

Observation 03a68202-8f3a-4669-b546-f82f5756354e · inbound

MDN: Parallelizing Stepwise Momentum for Delta Linear Attention cites this paper.

MDN: Parallelizing Stepwise Momentum for Delta Linear Attention Uncovering mesa-optimization algorithms in Transformers

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:11.996270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-09T15:27:55.566795Z digest=sha256:a210c8a79a18ddbb5671ea29b4207537cbba28998201eed12387a0fab07476bc

Observation 287d35b9-1503-4598-ac53-1920ee193c37 · inbound

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces cites this paper.

Correcting Influence: Unboxing LLM Outputs with Orthogonal Latent Spaces Uncovering mesa-optimization algorithms in Transformers

Reference 110

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:17:54.026408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-14T20:17:01.224864Z digest=sha256:13152d90056513a1e2914adad26fd3c7113695551ad605140c15a572a5158183

Observation a94d68b2-39ea-47a2-b82f-ffad56a48db6 · inbound

OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention cites this paper.

OSDN: Improving Delta Rule with Provable Online Preconditioning in Linear Attention Uncovering mesa-optimization algorithms in Transformers

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T19:09:23.215404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:08:18.768344Z digest=sha256:76eee63330f776097c4166357ee756b7cbb85809eb7029fc717276a0d2c363f4

Observation 171ea9ba-02c9-4f2f-931a-2a8aba847e8f · inbound

Dynamic Short Convolutions Improve Transformers cites this paper.

Dynamic Short Convolutions Improve Transformers Uncovering mesa-optimization algorithms in Transformers

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:36:26.970273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T10:48:50.103004Z digest=sha256:879583b7017de8aad38cc9fc26d4fcbf5de617b51216e68016cc65c18c591b90

Observation a4675f7f-d2d6-4824-ba9d-83c82dcba78b · inbound

A Systematic Study of Behavioral Cloning for Scientific Data Annotation cites this paper.

A Systematic Study of Behavioral Cloning for Scientific Data Annotation Uncovering mesa-optimization algorithms in Transformers

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:23:38.998915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T16:23:08.402194Z digest=sha256:112d94926576f783db34869511527be400974d38675b8a4dd4c717cf55cf589a

Observation 37c49333-b27e-4e33-9633-e93a09d270ae · inbound

Blurry Window Attention cites this paper.

Blurry Window Attention Uncovering mesa-optimization algorithms in Transformers

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:46:13.996400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T17:43:34.429061Z digest=sha256:9affeb42426c11d18aa100ff0b8f3ddee09d996c257c0e8b0895228f5d30a0fd

Observation 5da9a9d6-6222-42e3-93d6-db9447c0916a · inbound

Phase Transitions in Attention: A Bayesian Theory of Copy Head Emergence cites this paper.

Phase Transitions in Attention: A Bayesian Theory of Copy Head Emergence Uncovering mesa-optimization algorithms in Transformers

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:18:13.105960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T08:13:40.546738Z digest=sha256:c2548ff1a641e1e8aa8fbb8b5852a07ae6586a8fc1156ae1ecbfdd3d12ccb0b1

Observation 105d4032-d5cf-4cb6-9aad-7d3dac356f7a · inbound

Understanding Large Language Models cites this paper.

Understanding Large Language Models Uncovering mesa-optimization algorithms in Transformers

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:56:56.433366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T12:53:00.254754Z digest=sha256:38816787e14bb62ec82a81a32c46da886d51b072ca58de11d06abb59689860b1

Observation 18d52398-c2af-4d83-a126-10ba7f2d07c4 · inbound

Induction Heads Interpolate N-Grams cites this paper.

Induction Heads Interpolate N-Grams Uncovering mesa-optimization algorithms in Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T06:58:39.213932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:58:39.213932Z digest=sha256:c22fca4dc5babbc6239507ef184971cf203678d13cbb2d75bd30c6cb3a9177e9

Observation 11b9ee44-bb4f-4e4b-a8a4-ba0b103da30e · inbound

The Orthogonalized Read Is a Removable Training Scaffold for Recurrent Memory cites this paper.

The Orthogonalized Read Is a Removable Training Scaffold for Recurrent Memory Uncovering mesa-optimization algorithms in Transformers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T09:06:10.979909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:06:10.979909Z digest=sha256:2af25fd002f0809028ef7f4bda4b30db42382a2335e626f33be25395e5ff3765

Observation 536edb80-fd75-4937-b225-6627b360b02f · inbound

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex cites this paper.

Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex Uncovering mesa-optimization algorithms in Transformers

Reference 192

Resolution
unresolved
no resolver link, observed 2026-07-31T23:52:06.172263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:52:06.172263Z digest=sha256:4b45cd3b8cded797a89f9d6202a55b756cda3404a5f94ac93595104c93767fcf

Observation 9cbabda1-561f-4c03-8b92-d818308ce118 · inbound

Grounding latent algorithm routing in transformer reasoning cites this paper.

Grounding latent algorithm routing in transformer reasoning Uncovering mesa-optimization algorithms in Transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T14:04:25.225241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T14:04:25.225241Z digest=sha256:1b30d56492bf9dd1db58b0746898503e397ee855a6c2a80c083b9dc1fbfc7620