Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:31:44.464243Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 2 inbound Pith citation observations for arXiv:2505.07070.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T22:31:44.464243Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T14:02:56.648836Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T08:43:15.231264Z
67 of 67 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f9620582-b90a-4f23-b3d9-3ee458305c03 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures These correlations are given by the fol- lowing (vm)×v C(X−t,X−1)µ,ν :=P{X−t =µ,X−1 =ν} − P{X−t =µ} P{X−1 =ν}
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f3556d56-50e0-49d3-b6ad-c5d28707de41 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures These tuples have a smaller tree dis- tance from X−1 than tuples of input tokens, thus the corresponding sample complexities are lower than those required by Eq
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b45de51a-e311-4bca-9866-dfb743831000 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Therefore, a learner could infer all level- L latents 7 as soon as it’s able to infer µ (L) −2, which is the easiest to infer as it has the strongest correlation with X−1
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0da5e4c7-762c-4326-b1ee-3adeac97ada5 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Training Compute-Optimal Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21fd111c-2f48-42ae-9a06-03b4460cab0e · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Deep Learning Scaling is Predictable, Empirically
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f942c378-591d-4928-8f36-49f760bfc7c3 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Scaling Laws for Neural Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a70d5954-3680-4b59-bae7-559f0a4f200e · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Scaling Laws for Autoregressive Generative Modeling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6507dd54-7b49-45ac-9420-0c5d0e63cee1 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9766aa76-9136-4664-8150-8ed9e13c65cc · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 479cc6d5-e4a5-4edf-98e5-5000d7a11f09 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures A Provably Correct Algorithm for Deep Learning that Actually Works
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5652a916-3eae-401d-80a6-0ff4f9f4a6fa · outbound
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f39e553a-4843-40fc-8e4c-bafac3bb9c1b · outbound
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bf9f4ad3-b045-4fd2-849a-15feac22077d · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5671c530-f596-427f-a93e-e0435df3f2b6 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Cagnetta and M
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a9706e87-a7e2-47e1-b4f7-9ef84d7938c6 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures How transformers learn structured data: insights from hierarchical filtering
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 647cdbbe-dd5b-4f89-8ba7-dad03ed257ef · outbound
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 60cee5d0-ec66-403d-8330-594dfc0d26bf · outbound
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation df958a18-7646-4620-8bfd-7cb80651b640 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures How Compositional Generalization and Creativity Improve as Diffusion Models are Trained
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 558c60e3-c2e5-4b56-bd48-83ef7f294915 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Chomsky, Three models for the description of lan- guage, IRE Transactions on information theory 2, 113 (1956)
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a3a3d9f8-177f-4929-ab37-b92a81cffdd3 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Learning Curve Theory
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cfce252-544f-49d5-b1d3-c2761f5cef16 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 24334a98-9002-4343-8244-07d4cb32ef41 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Caponnetto and E
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 5a959ab3-0771-4795-a22a-96f6c40abf54 · outbound
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b581b13e-761e-4a28-92b3-130e806fcec0 · outbound
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 33e1b658-345c-40f5-a02d-ac511c5275ea · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Explaining Neural Scaling Laws
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a60bb6b9-bbc0-4fa9-8664-835698f53d76 · outbound
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 98285d40-0c90-496b-b4ca-d074ce1e4514 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation bfe9e1ff-cda3-4692-b595-1c04fda7e3d0 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures A Solvable Model of Neural Scaling Laws
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0c3820c-efe0-4c56-b71f-e2e5e0c736fa · outbound
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 939c3fe5-c6ac-46eb-a820-322bc0acf7d9 · outbound
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0849185a-0152-48d0-9697-a39560e21ca2 · outbound
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 151a625d-809a-4bc4-aef9-8a387d01c7db · outbound
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4b215b59-3af0-4e62-bf10-c24d457dc32b · outbound
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4753cb9d-db52-4aaf-9a65-51d1617656de · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4f82725d-8f72-4594-a246-59cb1a02dc85 · outbound
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c60ba2b1-93ef-4146-83f6-c0cc3b0a215c · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures How Two-Layer Neural Networks Learn, One (Giant) Step at a Time
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88047e31-246f-407a-9b6d-3c5f3b053296 · outbound
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 779b1707-d9a2-46fc-9e7e-b5b015302940 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Deep Learning and Hierarchal Generative Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24e503dc-c7c6-4687-9857-e197085ce378 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures U-Nets as Belief Propagation: Efficient Classification, Denoising, and Diffusion in Generative Hierarchical Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 905c49a0-9b71-4d97-a0c5-7190cfd29a54 · outbound
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 18b73f4e-bcee-4438-9936-1d64268b7896 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Sliding down the stairs: how correlated latent variables accelerate learning with neural networks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 4b0111b4-4bd1-497b-903d-86b7d9e0e21d · outbound
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35763e32-4523-47cb-b442-5b0f433e6f36 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Transformers Can Represent $n$-gram Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91f5f53a-83a7-42bf-af5c-a76e9afdda4d · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Nguyen, Understanding transformers via n-gram statistics, in The Thirty-eighth Annual Conference on Neural Information Processing Systems (2024)
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c118f06a-f8eb-4d26-a943-ad07fefd38c5 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Can Transformers Learn $n$-gram Language Models?
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 92c44b42-8982-402b-8c32-bb0fd9cc5267 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular Languages
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9820c7e9-87a7-4175-9456-8aeba7db20e4 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Transformers represent belief state geometry in their residual stream
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 417ba520-523b-4020-b63a-ce17a74736ab · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Physics of Language Models: Part 1, Learning Hierarchical Language Structures
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f1aa468-9db2-45c9-9c02-7be0ca47f448 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Do Transformers Parse while Predicting the Masked Word?
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db6ce586-838c-4076-b9bc-004a6b85c1cd · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3225ba8b-4e5a-45da-ac93-4d6f22a9ecfd · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Learning Syntax Without Planting Trees: Understanding Hierarchical Generalization in Transformers
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2276bf5-d551-4cbc-83d8-c3dcf424d8c9 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Rozenberg and A
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 75946625-11d7-4e1b-8c4c-e146ad7afe8e · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1b5844e8-17a4-468b-a985-9abeccec6d68 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0ed7ca43-34e8-46ce-80b4-a94217de5fce · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation b2b72171-38ee-4e9c-8005-26b41b815c4b · outbound
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d50267ef-7333-4787-a147-56983b52fa5e · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Ebeling and T
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e6e98d0e-9066-49ba-98f7-b946b4f77046 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Unresolved cited work
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 718b5131-4eda-4ff3-b4db-93dbb2b10a19 · outbound
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adead80c-0555-4770-9b7e-1186f214dbf0 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2bbafcbc-47aa-46e4-b211-3310599cdc60 · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures DeGiuli, Random language model, Phys
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a86a3215-3802-4148-b2af-b5bd9691245d · outbound
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a1fadc9b-ce98-413a-b696-b8a0fb178a9f · outbound
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a3fb775c-9465-4ef1-9fbe-da234e5b96d9 · outbound
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee763a8b-3754-463a-8e53-9121f97f1d0a · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures A Convolutional Neural Network for Modelling Sentences
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96943c8e-cbf0-4634-ad49-1ed6030b2dfe · outbound
Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation fd4b5fa6-b9e3-413b-a43b-0cf49d43d5ac · outbound
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b9d6123-39db-4799-bfff-b3f51a7b958b · inbound
Unifying Learning Dynamics and Generalization in Transformers Scaling Law Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
Reference 2007
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06b7a6b0-124c-467f-a929-7bacaa79ff39 · inbound
Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention Scaling Laws and Representation Learning in Simple Hierarchical Languages: Transformers vs. Convolutional Architectures
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.