Pith. sign in

Paper Citation Record · LEDGER

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points

As of 22 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2508.12837.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.12837 v2

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:26:36.926253Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T10:12:29.129465Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6329383b-e0b7-4a9b-9033-79d01fe375e6 · outbound

This paper cites Lemma H.4 (Stationarity of sub-k-tuples).

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Lemma H.4 (Stationarity of sub-k-tuples)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:26:37.297607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T17:26:36.921831Z digest=sha256:fab5245802788cc1a97bf553cc86c9553a1d6b9626555d83a7363d6a71b279b7

Observation 9285ac0d-c5c0-4f06-8fec-286356d5fa0e · outbound

This paper cites Transformers learn through gradual rank increase.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Transformers learn through gradual rank increase

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.850936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.850936Z digest=sha256:cbe76a9f7e68314261c974b93410201266d98131fca85a6486418623c88a8cc1

Observation 7783ac8e-c731-40b6-a86e-734b189170de · outbound

This paper cites Understanding Incremental Learning of Gradient Descent: A Fine-grained Analysis of Matrix Sensing.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Understanding Incremental Learning of Gradient Descent: A Fine-grained Analysis of Matrix Sensing

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.876555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.876555Z digest=sha256:a7bc2140ea0863a5367624db34105506e41fc8010f5b9f59dce0f675d564b054

Observation ef88414f-672f-472b-8a34-66c1c885fa2a · outbound

This paper cites Task Diversity Shortens the ICL Plateau.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Task Diversity Shortens the ICL Plateau

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.880999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.880999Z digest=sha256:842c2898cd371c594155e47006e68bd3d8342df919c1788bf91a899109ba3bf6

Observation 0fd5ad81-8cc8-4652-819d-2f8885f5c2f5 · outbound

This paper cites Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Attention with Markov: A Framework for Principled Analysis of Transformers via Markov Chains

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.886042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.886042Z digest=sha256:27885824f92139647c8861efb6bb574fc099fe8a35fafc6b82e0910001b142cc

Observation 7ff6388f-9ea2-47c8-bb47-61d49b1e97cd · outbound

This paper cites How Transformers Learn Causal Structure with Gradient Descent.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points How Transformers Learn Causal Structure with Gradient Descent

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.890264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.890264Z digest=sha256:871b46cfc7c231710abb6a91c3b05eeacd4bd842285e706299a6194627ad1bf7

Observation 729f2995-ba6e-4df5-b876-53784724f649 · outbound

This paper cites Transformers on Markov Data: Constant Depth Suffices.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Transformers on Markov Data: Constant Depth Suffices

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.902854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.902854Z digest=sha256:e912597c6848a51fb3d83bc1f3288c2c2748075d080e6198ccc34948f2b145c3

Observation 4f495f07-29f9-40db-90be-7b519477889b · outbound

This paper cites Emergent Abilities of Large Language Models.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Emergent Abilities of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.912757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.912757Z digest=sha256:5cc98940480f531b0f0cc9f3660e25643a966e239346240fc988bcdc0cff1143

Observation 4c39ca19-1411-4709-9f16-1479081aa876 · outbound

This paper cites Large Language Models as Markov Chains.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Large Language Models as Markov Chains

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.917183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.917183Z digest=sha256:5d090373c1befd55e62536cde65f7c736b1c179414a45dfa0eca93e7bc94bbb5

Observation 373e80b1-79ea-4309-b859-f5018b7297bb · outbound

This paper cites an unresolved cited work.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Unresolved cited work

Reference 128

Resolution
unresolved
raw_fallback, observed 2026-08-15T17:26:37.284197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T17:26:36.926253Z digest=sha256:eca0aaf59b1fb0b892664e1d129488b3782be3915ded39da57dc31ae2df88b1a

Observation e83276c0-6e19-4f46-ae9a-f62c4c8e15db · outbound

This paper cites Transformers Can Represent $n$-gram Language Models.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Transformers Can Represent $n$-gram Language Models

Reference 1948

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.907740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.907740Z digest=sha256:c57667cf799aebb0263fc8feda72713a2fce1843655ec47e41fb6ceabcb3fd00

Observation 364f9266-cd07-4149-9994-3b6002c82579 · outbound

This paper cites Edelman, B.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Edelman, B

Reference 1956

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.859528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.859528Z digest=sha256:514fdfaee49d87004c3ddcce1411b39f2839b19828651bfa3e54b1af0b780b05

Observation e416bfee-006b-4672-92b3-ff68fbf6d459 · outbound

This paper cites Language Models are Few-Shot Learners.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Language Models are Few-Shot Learners

Reference 1992

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.855241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.855241Z digest=sha256:72ffc9ebe3156b05229826679e7878f76103e1d2204e5564a451839869cc83c2

Observation de7d66a7-bab5-4c48-8fff-47cedee14160 · outbound

This paper cites Algorithmic Regularization in Model-free Overparametrized Asymmetric Matrix Factorization.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Algorithmic Regularization in Model-free Overparametrized Asymmetric Matrix Factorization

Reference 1998

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T17:26:37.110041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T17:26:36.872360Z digest=sha256:63d11e7535da73a6dc17b32af605e9218a57e12d4cd002e65e1df169d4fec5c8

Observation 517fafe7-6b3c-4979-8520-bca981bf77af · outbound

This paper cites doi: https://doi.org/10.1016/S0893-6080(00)00009-5.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points doi: https://doi.org/10.1016/S0893-6080(00)00009-5

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.863641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.863641Z digest=sha256:1dbf90d70b004770d97915bd8dd9dee4fd685b0fad4f6992531c61af20a209e1

Observation b261dcc0-6745-4499-8898-aa6ed025e675 · outbound

This paper cites Incremental Learning in Diagonal Linear Networks.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Incremental Learning in Diagonal Linear Networks

Reference 2019

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:26:37.256553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T17:26:36.841860Z digest=sha256:049b7cab86406ba4f997a8b46be72fd422f79c6e42e7c2746165c40cba128480

Observation f0c618b3-6213-47c5-b045-289d914670fc · outbound

This paper cites Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Saddle-to-Saddle Dynamics in Deep Linear Networks: Small Initialization Training, Symmetry, and Sparsity

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.867859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.867859Z digest=sha256:96ce991f74012fb5e8327a6404e4676e82bc8cfdd4f92bfdc5bba457dff58eec

Observation 47303cbe-4018-4c8c-8bad-d3aabd8c4e65 · outbound

This paper cites Understanding In-Context Learning in Transformers and LLMs by Learning to Learn Discrete Functions.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Understanding In-Context Learning in Transformers and LLMs by Learning to Learn Discrete Functions

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.846150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.846150Z digest=sha256:6c70303260865407257099d2f66781d5e835e8955471bee0cb375e49e32e91be

Observation 798ccdd0-48f8-4247-b015-769d65ece607 · outbound

This paper cites Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points Grokking: Generalization Beyond Overfitting on Small Algorithmic Datasets

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.898860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.898860Z digest=sha256:8791aa069e56682fe9ab1b619429e331c8d354ee679a4e5bcb7df4ab3e010327

Observation 2aade19d-f781-48df-908c-924a2f0dc9d1 · outbound

This paper cites In-Context Language Learning: Architectures and Algorithms.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points In-Context Language Learning: Architectures and Algorithms

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T17:26:36.836988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:26:36.836988Z digest=sha256:8098553aa0e4e63eb8543a6b35af1619ab517c30eba5d17373929023b0e48b0a

Observation d77ef9e6-0216-4253-9b72-dab3de4b96a7 · outbound

This paper cites A Mechanistic Study of Transformers Training Dynamics.

Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points A Mechanistic Study of Transformers Training Dynamics

Reference 2025

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T17:26:37.039736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T17:26:36.894671Z digest=sha256:757f60a18bb7bc093fafd71588f93b1cbefc447379681b2a5acb11021d3af083

Pith citing papers

Observation 6175fb5c-89b1-40c8-a995-7fe76b38f892 · inbound

On the global convergence of gradient flow for wide shallow models beyond homogeneous nonlinearities cites this paper.

On the global convergence of gradient flow for wide shallow models beyond homogeneous nonlinearities Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:51:19.291387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-12T03:51:08.871267Z digest=sha256:0686edd6279878e59b5d33798f851c469ebbd7445415174f7e2597619bc0c223

Observation 60b03b04-346f-49ce-85f2-3a1766e2527b · inbound

Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models cites this paper.

Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models Learning In-context n-grams with Transformers: Sub-n-grams Are Near-stationary Points

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T10:12:29.129465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:12:29.129465Z digest=sha256:4d96bfff2635b49cca94270447657364ad6952ef1c4baca263228f467ef856d0