Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:26:31.193173Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2506.14095.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T00:26:31.193173Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
68 of 68 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 62771866-b7fa-4835-8861-4c85f607bed9 · outbound
Transformers Learn Faster with Semantic Focus Attention is all you need
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 22bb8517-1758-44a3-89d5-f49376359346 · outbound
Transformers Learn Faster with Semantic Focus Are transformers universal approximators of sequence-to-sequence functions? In International Conference on Learning Representations, 2020 a
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e9223ce-1562-43b5-a3cb-18c05d591716 · outbound
Transformers Learn Faster with Semantic Focus Efficient transformers: A survey
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1af2cb21-31bf-4ab7-b069-681e4b50bdc6 · outbound
Transformers Learn Faster with Semantic Focus Long range arena: A benchmark for efficient transformers
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 59ba69a1-cbbe-4543-97ca-d997251586a1 · outbound
Transformers Learn Faster with Semantic Focus Cognitive Mechanisms Associated with Auditory Sensory Gating
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44f6baf5-da5e-402a-83b1-a523095912f5 · outbound
Transformers Learn Faster with Semantic Focus The Senses: A Comprehensive Reference
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21ace225-81de-41d0-821b-fd63afaeb096 · outbound
Transformers Learn Faster with Semantic Focus Sensory gating deficits in schizophrenia: new results
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f4952284-fd8e-4da6-b4c6-88f49606b297 · outbound
Transformers Learn Faster with Semantic Focus The Consciousness Prior
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2232304-32ec-42e8-a934-7a304b29b1c2 · outbound
Transformers Learn Faster with Semantic Focus URL https://neurosymbolic.github.io/nsss2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 14e85bea-67d9-43a4-a785-ab7f32ae9204 · outbound
Transformers Learn Faster with Semantic Focus Neural Machine Translation by Jointly Learning to Align and Translate
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9fff3c4-0fbf-4b1b-a1f7-c3f9fc251544 · outbound
Transformers Learn Faster with Semantic Focus Neural networks and the chomsky hierarchy
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a66a7ba5-854b-43ff-9bbe-e3ff87789bc5 · outbound
Transformers Learn Faster with Semantic Focus O(n) connections are expressive enough: Universal approximability of sparse transformers
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 61af9724-2b54-4a43-8217-c53f99349493 · outbound
Transformers Learn Faster with Semantic Focus Etc: Encoding long and structured inputs in transformers
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84ce2eca-690d-4bb7-a2aa-9ea0dff21b3f · outbound
Transformers Learn Faster with Semantic Focus Big bird: Transformers for longer sequences
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32f6bc68-148f-4bae-a696-780ff6753dc3 · outbound
Transformers Learn Faster with Semantic Focus Memory-efficient transformers via top-k attention
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f42847e-789a-4545-ad74-dbd1c9a601cb · outbound
Transformers Learn Faster with Semantic Focus ZETA : Leveraging z -order curves for efficient top- k attention
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 672ff6cb-90c5-4c82-99a3-f10cd42cd3a6 · outbound
Transformers Learn Faster with Semantic Focus Algorithmic stability and generalization performance
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8495391a-4fb7-4739-ac0c-aec66433f718 · outbound
Transformers Learn Faster with Semantic Focus Train faster, generalize better: Stability of stochastic gradient descent
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bb6606d-3f55-472b-a431-4d0742abc47c · outbound
Transformers Learn Faster with Semantic Focus Formal Algorithms for Transformers
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd34a3bd-096a-4936-9e29-698e87b2df1d · outbound
Transformers Learn Faster with Semantic Focus A survey of transformers
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c1bfa5c9-f0c9-4ec2-9862-5b35f64ee4ca · outbound
Transformers Learn Faster with Semantic Focus Image transformer
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79b46b3c-40a9-4820-be5e-130bc2e701cb · outbound
Transformers Learn Faster with Semantic Focus Blockwise self-attention for long document understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f63968ca-efc8-419d-878a-ddb05d57f944 · outbound
Transformers Learn Faster with Semantic Focus Longformer: The Long-Document Transformer
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 067581ba-98c3-4144-b6bb-576ff06f2877 · outbound
Transformers Learn Faster with Semantic Focus Generating Long Sequences with Sparse Transformers
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de423e8c-14cf-431b-bb13-0d340ddd288b · outbound
Transformers Learn Faster with Semantic Focus Sparse sinkhorn attention
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8cee7eee-810c-4efa-84d8-49e5f85f3321 · outbound
Transformers Learn Faster with Semantic Focus Efficient content-based sparse attention with routing transformers
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e7434e7-0ad7-45cb-b59f-e14f2167eae9 · outbound
Transformers Learn Faster with Semantic Focus Reformer: The efficient transformer
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d84cf548-c9c4-4cf2-8365-00a069fe0a17 · outbound
Transformers Learn Faster with Semantic Focus COGS : A compositional generalization challenge based on semantic interpretation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf06c895-63e9-4ecd-978e-e546c4c82dde · outbound
Transformers Learn Faster with Semantic Focus Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47980f98-7844-48dd-9fd5-0af02175dc94 · outbound
Transformers Learn Faster with Semantic Focus When can transformers ground and compose: Insights from compositional generalization benchmarks
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3ceab004-66f3-41e6-84d4-824a19b41134 · outbound
Transformers Learn Faster with Semantic Focus The devil is in the detail: Simple tricks improve systematic generalization of transformers
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a13d6515-33e0-4e1d-91c8-7e541b43e13d · outbound
Transformers Learn Faster with Semantic Focus Making transformers solve compositional tasks
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2496269a-85e9-44eb-a5db-d03f7139143a · outbound
Transformers Learn Faster with Semantic Focus Inducing transformer ' s compositional generalization ability via auxiliary sequence prediction tasks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84412c7a-508f-461c-a335-db1b31394c91 · outbound
Transformers Learn Faster with Semantic Focus a rli, Ekin Aky \
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f5b2bfd4-f7c9-4872-a2ad-f00fdeacdc96 · outbound
Transformers Learn Faster with Semantic Focus What formal languages can transformers express? a survey
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6020b2e6-aa1c-47ce-8d90-8afc314dd586 · outbound
Transformers Learn Faster with Semantic Focus On the ability and limitations of transformers to recognize formal languages
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1604f60a-e988-4af0-8874-2c4f77626b61 · outbound
Transformers Learn Faster with Semantic Focus Theoretical limitations of self-attention in neural sequence models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2ece83c2-9f63-418c-b1d5-c8ae82a60982 · outbound
Transformers Learn Faster with Semantic Focus Formal language recognition by hard attention transformers: Perspectives from circuit complexity
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bb1ab5b5-9b8c-400a-b686-8b77247e18c6 · outbound
Transformers Learn Faster with Semantic Focus Saturated transformers are constant-depth threshold circuits
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f3fe64ff-55f2-46e6-80d0-57bf72a7230a · outbound
Transformers Learn Faster with Semantic Focus Overcoming a theoretical limitation of self-attention
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f80651a-9042-41df-8507-c06510ced51d · outbound
Transformers Learn Faster with Semantic Focus Tighter bounds on the expressivity of transformer encoders
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e1a1abc0-80ea-4abc-a8b7-530e386ed594 · outbound
Transformers Learn Faster with Semantic Focus Transformers as algorithms: Generalization and stability in in-context learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af5cac0e-1cef-4924-a631-3d6f7a36699f · outbound
Transformers Learn Faster with Semantic Focus Transformers learn in-context by gradient descent
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bc5a2ec6-5e4e-4209-8bbe-586adb9acbe7 · outbound
Transformers Learn Faster with Semantic Focus The emergence of clusters in self-attention dynamics
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4842f956-d2c0-4cd2-bf07-6a05324354e0 · outbound
Transformers Learn Faster with Semantic Focus Transformers learn to implement preconditioned gradient descent for in-context learning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0676e91a-49e1-44af-9ffe-d180a49e2d1e · outbound
Transformers Learn Faster with Semantic Focus Trained transformers learn linear models in-context
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f1208b8e-f112-4011-bd87-dd3159dc5249 · outbound
Transformers Learn Faster with Semantic Focus Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da9fc738-70f8-48b4-8259-7545c05251a4 · outbound
Transformers Learn Faster with Semantic Focus Toward understanding why adam converges faster than SGD for transformers
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9f67186-e199-4d0c-9ad6-15fa02f759f1 · outbound
Transformers Learn Faster with Semantic Focus How does adaptive optimization impact local neural network geometry? Advances in Neural Information Processing Systems, 36: 0 8305--8384, 2023
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dbc1b7cc-1d95-47c1-85fa-4b50bf0a234a · outbound
Transformers Learn Faster with Semantic Focus Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 884fe00e-92a9-460e-848e-cb3bd47f188e · outbound
Transformers Learn Faster with Semantic Focus Linear attention is (maybe) all you need (to understand transformer optimization)
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b632e26-88f9-40fd-b30c-2f0591d3f489 · outbound
Transformers Learn Faster with Semantic Focus On the optimization and generalization of two-layer transformers with sign gradient descent
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7d1005fb-2019-4208-b0c1-d3163e3f2a97 · outbound
Transformers Learn Faster with Semantic Focus Layer Normalization
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46115e48-be86-45ce-b49a-e9ce4bdda4f9 · outbound
Transformers Learn Faster with Semantic Focus Root mean square layer normalization
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee4e4e51-cd62-4925-9254-1f342c52ca72 · outbound
Transformers Learn Faster with Semantic Focus BERT : Pre-training of deep bidirectional transformers for language understanding
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1ad7295-411a-42bf-9db6-42b3ff685421 · outbound
Transformers Learn Faster with Semantic Focus Gaussian Error Linear Units (GELUs)
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d521b3a9-96a1-4475-9a3a-d3942c18b279 · outbound
Transformers Learn Faster with Semantic Focus Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fa87b52-75d6-4f21-b69c-f76cc01c4f8e · outbound
Transformers Learn Faster with Semantic Focus Listops: A diagnostic dataset for latent tree learning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bfb46c1d-85a6-4889-b1b2-37873f04faca · outbound
Transformers Learn Faster with Semantic Focus Mish: A Self Regularized Non-Monotonic Activation Function
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a944d61-2998-44b7-a34b-676ba1b45bf1 · outbound
Transformers Learn Faster with Semantic Focus Adam: A Method for Stochastic Optimization
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c0085ca-f463-4103-952b-307c15eb115e · outbound
Transformers Learn Faster with Semantic Focus The Lipschitz Constant of Self-Attention
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12797a98-aa4b-437c-b189-96d7d2b884b1 · outbound
Transformers Learn Faster with Semantic Focus Explicit Sparse Transformer: Concentrated Attention Through Explicit Selection
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37334c11-9643-4161-800f-94b4c7a6bbe5 · outbound
Transformers Learn Faster with Semantic Focus Visualizing the loss landscape of neural nets
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 10878189-944f-4b9b-8f6c-9907c722def7 · outbound
Transformers Learn Faster with Semantic Focus Never train from scratch: Fair comparison of long-sequence models requires data-driven priors
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c25273f1-6b34-4339-bc1d-998ffb03b3da · outbound
Transformers Learn Faster with Semantic Focus Learning overparameterized neural networks via stochastic gradient descent on structured data
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 233ed32b-ea5f-4522-944c-d2bb9ef0050c · outbound
Transformers Learn Faster with Semantic Focus A convergence theory for deep learning via over-parameterization
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ca681d40-ef8b-457c-962c-fffe1a113051 · outbound
Transformers Learn Faster with Semantic Focus Gradient descent optimizes over-parameterized deep relu networks
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation aa3a87b7-a9a4-47f5-a37e-2fd076f33f5b · outbound
Transformers Learn Faster with Semantic Focus Convergence rates for the stochastic gradient descent method for non-convex objective functions
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.