Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:46:40.614748Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 11 inbound Pith citation observations for arXiv:2505.23683.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:46:40.614748Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:34:15.603859Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-29T14:33:30.598717Z
71 of 71 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c791e2bc-b13d-47aa-97a8-afcc6650bf6b · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data The staircase property: How hierarchical structure can guide deep learning.Advances in Neural Information Processing Systems, 34:26989–27002, 2021
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddd8ead5-d8c7-406c-80cd-4067f3f7b8da · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data The merged-staircase property: a necessary and nearly sufficient condition for sgd learning of sparse functions on two-layer neural networks
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a03d8f2f-fec3-480a-9a4c-524dcb10114c · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Provable advantage of curriculum learning on parity targets with mixed inputs.Advances in Neural Information Processing Systems, 36:24291–24321, 2023
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 831013f1-6797-4024-ab3a-b34907bde771 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data How far can transformers reason? The locality barrier and inductive scratchpad
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 773f1920-9aa9-4517-95ed-c7f90f9bde19 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Transformers learn to implement preconditioned gradient descent for in-context learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c86c7eff-538c-4a24-8d82-21a2ad229fac · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Graph streaming lower bounds for parameter estimation and property testing via a streaming xor lemma
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4ab05c7b-4394-466c-b552-cbf39e6178de · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data The pitfalls of next-token prediction
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b6ed973d-856d-422d-89bd-59af7f89cea5 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Curriculum learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36d3242c-1566-4cde-b737-5d22627eb4f5 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Separations in the Representational Capabilities of Transformers and Recurrent Architectures
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ccb8cf36-5b8b-4f33-a4be-6228c304285c · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Birth of a transformer: A memory viewpoint
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation c834b064-facb-4bf2-9fb7-c7c5926c5ae1 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Data distributional properties drive emergent in-context learning in transformers
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 23b29014-a63e-4589-8fc7-e3563d33c822 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Theoretical limitations of multi-layer Transformer
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a597f3a7-8913-43cb-bfa4-167279020fab · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Training Dynamics of Multi-Head Softmax Attention for In-Context Learning: Emergence, Convergence, and Optimality
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b134bc0c-7e37-4fb7-824d-df2c5462488e · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Unveiling Induction Heads: Provable Training Dynamics and Feature Learning in Transformers
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ed040cd-8986-4385-b60e-f88d2bf4c500 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Understanding the Interplay between Parametric and Contextual Knowledge for Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bde2b42-80a5-4895-a919-a11e66fefed9 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Neural networks can learn represen- tations with gradient descent
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5054e96-626f-42fc-be1a-26f48263814f · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bfa3d55-b086-4846-a561-d6cee1fbf8ce · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data The Llama 3 Herd of Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c841ccb3-8b76-48cc-a404-53bbb7898b27 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Faith and Fate: Limits of Transformers on Compositionality
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4bc1d60-d879-4948-895a-d503ea825cc7 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Learning and development in neural networks: The importance of starting small.Cognition, 48(1):71–99, 1993
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b543bd7a-7bda-4750-b189-27547524ec73 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Towards revealing the mystery behind chain of thought: a theoretical perspective.Advances in Neural Information Processing Systems, 36, 2024
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2ef8ca9c-4295-4c58-b682-80c33d7bab6e · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Global Convergence in Training Large-Scale Transformers
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d40ad934-5101-49d1-ae2f-38f0ad56451f · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Better & Faster Large Language Models via Multi-token Prediction
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 732e7025-32e6-4375-a0f9-6c2d23667c6c · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Interleaved group products.SIAM Journal on Computing, 48(2):554–580, 2019
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9fb45c7f-1a84-423b-a32d-f6bf87f6f129 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Transformers Learn to Implement Multi-step Gradient Descent with Chain of Thought
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a1ba3e9-0933-4e24-b9c2-4b1e1948cc63 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data In-Context Convergence of Transformers
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcb56841-e383-407a-af90-d80388e83cbe · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Mas- sively parallel computation: Algorithms and applications.Foundations and Trends® in Optimization, 5(4):340–417, 2023
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ccca6d86-49a0-4459-8eea-37c03d64ca10 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Vision transformers provably learn spatial structure
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7d3ab2f4-6822-4c76-afe7-47e5270748c0 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Repeat After Me: Transformers are Better than State Space Models at Copying
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99bacc77-02fd-418d-9507-034ef2093aaf · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Efficient noise-tolerant learning from statistical queries.J
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2322fda7-1f28-4a84-bb51-4dec2ba0ffc8 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Transformers Provably Solve Parity Efficiently with Chain of Thought
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd18b5f5-62d5-4173-bc4b-3b016e719ba7 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Adam: A Method for Stochastic Optimization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be79f8cf-1577-4a53-8f21-a77e162349bc · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Learning to reason and memorize with self-notes
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 070b5a5f-cc81-49cf-bbce-bb5df2caf729 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Solving quantitative reasoning problems with language models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a1257c2b-f6eb-4450-91f6-91492b32782b · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data How do transformers learn topic structure: Towards a mechanistic understanding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a5719318-da93-4fef-a358-989d748c5d6b · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Chain of Thought Empowers Transformers to Solve Inherently Serial Problems
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d5d60aa-d7d8-4822-99a7-c91d5121ced5 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Let's Verify Step by Step
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e23e46eb-224b-41a4-8744-80dcfddca948 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data DeepSeek-V3 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5258745-2be9-4dfb-b9d2-4182af1f2f97 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Ash, Surbhi Goel, Akshay Krishnamurthy, and Cyril Zhang
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation dedc056e-1367-4bac-a297-f4259dc3a96f · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data RegMix: Data Mixture as Regression for Language Model Pre-training
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a805c674-e3e4-496c-b700-826fa40ccb80 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data The Expressive Power of Transformers with Chain of Thought
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba40d699-5080-47eb-b246-71d1a32e7455 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data The parallelism tradeoff: Limitations of log-precision transformers.Transactions of the Association for Computational Linguistics, 11:531–545, 2023
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5255c7ec-42a6-43ae-b389-7051503a870d · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data How transformers learn causal structure with gradient descent
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d36466a5-8576-412c-b84a-5474d5d7895d · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Understanding Factual Recall in Transformers via Associative Memories
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 367d7a22-b668-4a37-b725-678b2bc061ab · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Rounds in communication complexity revisited.SIAM Journal on Computing, 22(1):211–219, 1993
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 45205473-97b1-49c4-b5fb-f8495ed7503f · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Show Your Work: Scratchpads for Intermediate Computation with Language Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53ea21a2-0bc5-4f93-9ba7-922355824d0e · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data In-context learning and induction heads.Transformer Cir- cuits Thread, 2022
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation ffc1f16e-3425-486b-bab0-270f6c72b8ff · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Progressive distillation induces an implicit curriculum
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de5a35ec-1470-4b72-9c7b-c858563207da · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Papadimitriou and Michael Sipser
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e56752ed-55ec-4177-885f-087b1798635a · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data On limitations of the trans- former architecture
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8f8dcf86-ec50-4dfd-8d3c-c5265be5cb83 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Learning and transferring sparse contextual bi- grams with linear transformers
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a0aaa9a3-4f95-4cec-b64e-24a101960671 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Representational strengths and limitations of transformers
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f8097cbb-e197-4a3b-b5ea-cb4eb9dceede · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Understanding Transformer Reasoning Capabilities via Graph Algorithms
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1f0a159-eac9-4ca1-8d46-d41662a12b6f · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Transformers, parallel computation, and logarithmic depth
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9792657b-2058-465f-9fe8-172cdcf3327b · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data End-to-end memory networks
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 19db7bda-130f-4046-83f8-0764f09a3f47 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Characterizing statistical query learning: simplified notions and proofs
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 36aa5752-892f-40e3-ba6c-081c9d010c5c · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Scan and snap: Understanding training dynamics and token composition in 1-layer transformer
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 22e35c1f-1df0-4579-a963-34b49678ef39 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Solving math word problems with process- and outcome-based feedback
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2259d854-e282-4a79-aa47-a6c1ad55e9a9 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0da4c9a8-52a1-46f5-aef6-94ee1512ebdf · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Chain-of-thought prompting elicits reasoning in large language models
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4c59a0ee-c65c-4fe3-8215-7a17f3099a43 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7d6e0cb-cbc7-4f78-99c0-39d9167d5ef4 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Memory Networks
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02cb2c74-8a26-46ff-a8aa-d8384dd7a2a4 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0bcd8b8-c830-42e1-9a0c-ce3c25732862 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Doremi: Optimizing data mixtures speeds up language model pretraining
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2f6f0c7b-8d67-432a-b51b-fb187d7f4894 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Do Large Language Models Latently Perform Multi-Hop Reasoning?
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a3e5602-bead-4bd9-99ca-23dab1739c35 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Tree of thoughts: deliberate problem solving with large language models
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3c0f99ec-21a5-482c-b5d4-71d77c1f7c13 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Pointer chasing via triangular discrimination.Combinatorics, Probability and Computing, 29(4):485–494, 2020
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation b2fc7f14-4432-4a2a-a85f-c240aa9b09e9 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Trained Transformers Learn Linear Models In-Context
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48c3ed10-53dd-4f8b-98cf-3c616de2aa69 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 6efa7547-edbe-489e-869f-cbc4f2d7184a · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data Since the softmax in layer ℓ is nearly saturated and thus close to one-hot, the update norm is very small and can be bounded as noise terms
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 192e0eaa-499b-4984-bd8d-e32f3cb90dc2 · outbound
Learning Compositional Functions with Transformers from Easy-to-Hard Data That means in the tth gradient step, there will be some small gradient updates for later layersℓ′ ≥t introduced by the perturbation
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 1e43c969-1d4f-439d-8343-a97345183408 · inbound
Beyond Statistical Learning: Exact Learning Is Essential for General Intelligence Learning Compositional Functions with Transformers from Easy-to-Hard Data
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67828437-2bbd-46e5-b46c-cd501bce0c1b · inbound
Scaling Latent Reasoning via Looped Language Models Learning Compositional Functions with Transformers from Easy-to-Hard Data
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation feed43f0-78ec-4421-a44f-ce39f10e4457 · inbound
Scaling Latent Reasoning via Looped Language Models Learning Compositional Functions with Transformers from Easy-to-Hard Data
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ff28ce0-6fcb-4a8c-accc-578779f06471 · inbound
Deep sequence models tend to memorize geometrically; it is unclear why Learning Compositional Functions with Transformers from Easy-to-Hard Data
Reference 186
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 0b1e143f-b234-4151-8cd5-0c79778c1c00 · inbound
Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently Learning Compositional Functions with Transformers from Easy-to-Hard Data
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3340179c-4d68-4d1f-8b84-d172d3085ec2 · inbound
Breaking the Reversal Curse in Autoregressive Language Models via Identity Bridge Learning Compositional Functions with Transformers from Easy-to-Hard Data
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1fafadd-a6b4-4052-9480-75040c60f282 · inbound
Sharp Capacity Scaling of Spectral Optimizers in Learning Associative Memory Learning Compositional Functions with Transformers from Easy-to-Hard Data
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 2e0e0277-07c5-4180-b57c-d9ec1f0787ec · inbound
The Power of Power Law: Asymmetry Enables Compositional Reasoning Learning Compositional Functions with Transformers from Easy-to-Hard Data
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 26962014-0ff3-463c-a610-90aef65b491b · inbound
The Power of Power Law: Asymmetry Enables Compositional Reasoning Learning Compositional Functions with Transformers from Easy-to-Hard Data
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a4a587c-aaa8-4cc0-8415-c2fd22c76ffe · inbound
Transformers Efficiently Perform In-Context Logistic Regression via Normalized Gradient Descent Learning Compositional Functions with Transformers from Easy-to-Hard Data
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3836b10a-2e80-4a5b-9cc5-52054adc3b64 · inbound
Transformers Provably Learn to Internalize Chain-of-Thought Learning Compositional Functions with Transformers from Easy-to-Hard Data
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.