Pith. sign in

Paper Citation Record · LEDGER

Transformers Learn Faster with Semantic Focus

As of 7 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2506.14095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14095 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:26:31.193173Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact7
  • verified fuzzy42
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 62771866-b7fa-4835-8861-4c85f607bed9 · outbound

This paper cites Attention is all you need.

Transformers Learn Faster with Semantic Focus Attention is all you need

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.918906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.715156Z digest=sha256:8afc3a638a7165af2296d7156efffe86807d77137e4693fcf04a4f3488628ccb

Observation 22bb8517-1758-44a3-89d5-f49376359346 · outbound

This paper cites Are transformers universal approximators of sequence-to-sequence functions? In International Conference on Learning Representations, 2020 a.

Transformers Learn Faster with Semantic Focus Are transformers universal approximators of sequence-to-sequence functions? In International Conference on Learning Representations, 2020 a

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.902339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.776075Z digest=sha256:6b623487578538d31edb591e5e31a5e77846d19cd05fa39e3d897670da774bf1

Observation 8e9223ce-1562-43b5-a3cb-18c05d591716 · outbound

This paper cites Efficient transformers: A survey.

Transformers Learn Faster with Semantic Focus Efficient transformers: A survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.865774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.865774Z digest=sha256:78c7a60d4079530e6786e3be6a4f0fbe3b51056d18c92182aa50765a893a6860

Observation 1af2cb21-31bf-4ab7-b069-681e4b50bdc6 · outbound

This paper cites Long range arena: A benchmark for efficient transformers.

Transformers Learn Faster with Semantic Focus Long range arena: A benchmark for efficient transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.887632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.878995Z digest=sha256:88b6762b7d62eec358fa8849f373d3ea880e9d78f87ec17566685eaa032cc19b

Observation 59ba69a1-cbbe-4543-97ca-d997251586a1 · outbound

This paper cites Cognitive Mechanisms Associated with Auditory Sensory Gating.

Transformers Learn Faster with Semantic Focus Cognitive Mechanisms Associated with Auditory Sensory Gating

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.873539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.884261Z digest=sha256:b46937d5151c04cf81c9176ecab460b5eccc63217aed3c2003285801c2d57dd4

Observation 44f6baf5-da5e-402a-83b1-a523095912f5 · outbound

This paper cites The Senses: A Comprehensive Reference.

Transformers Learn Faster with Semantic Focus The Senses: A Comprehensive Reference

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.859632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.888896Z digest=sha256:0341fbfc6ef3044de79b16004911f63e93f8690bc95ef6da8648760189d906b2

Observation 21ace225-81de-41d0-821b-fd63afaeb096 · outbound

This paper cites Sensory gating deficits in schizophrenia: new results.

Transformers Learn Faster with Semantic Focus Sensory gating deficits in schizophrenia: new results

Reference 7

Resolution
verified exact
raw_fallback, observed 2026-08-07T00:26:31.888917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.894305Z digest=sha256:ae8448b2cf29e0a12e591cdadfce8bfee5dea2f0ed4927b57f138bc44fe08e64

Observation f4952284-fd8e-4da6-b4c6-88f49606b297 · outbound

This paper cites The Consciousness Prior.

Transformers Learn Faster with Semantic Focus The Consciousness Prior

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.899429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.899429Z digest=sha256:86aeceb528a393828b0d317d8c18c04656c3b71b38f89f95c978c1a14a2b57cc

Observation b2232304-32ec-42e8-a934-7a304b29b1c2 · outbound

This paper cites URL https://neurosymbolic.github.io/nsss2024.

Transformers Learn Faster with Semantic Focus URL https://neurosymbolic.github.io/nsss2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.845025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.910069Z digest=sha256:b8b2dd593d96965d8c3cdead659a90c315552576beaf5c06985db0fa3278544b

Observation 14e85bea-67d9-43a4-a785-ab7f32ae9204 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

Transformers Learn Faster with Semantic Focus Neural Machine Translation by Jointly Learning to Align and Translate

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.917132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.917132Z digest=sha256:97c72e830957bc13dcb06251a58c2c7733d67ab94a235b7785d6cbe9bf349169

Observation a9fff3c4-0fbf-4b1b-a1f7-c3f9fc251544 · outbound

This paper cites Neural networks and the chomsky hierarchy.

Transformers Learn Faster with Semantic Focus Neural networks and the chomsky hierarchy

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.830872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.922213Z digest=sha256:c65d11ef0ff355878d2eafc28159dac5ddbeedf74505091bf754ef6c6ca906ea

Observation a66a7ba5-854b-43ff-9bbe-e3ff87789bc5 · outbound

This paper cites O(n) connections are expressive enough: Universal approximability of sparse transformers.

Transformers Learn Faster with Semantic Focus O(n) connections are expressive enough: Universal approximability of sparse transformers

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.812327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.926723Z digest=sha256:4edbada6b8ff25eec038ca8657dec4769f3be05c67d4be33077b79d7ebc89338

Observation 61af9724-2b54-4a43-8217-c53f99349493 · outbound

This paper cites Etc: Encoding long and structured inputs in transformers.

Transformers Learn Faster with Semantic Focus Etc: Encoding long and structured inputs in transformers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.794142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.931698Z digest=sha256:c9e629b0de5f73e0575c1f0de35002da7eacf145c1b4a57dbb5b2827b2348064

Observation 84ce2eca-690d-4bb7-a2aa-9ea0dff21b3f · outbound

This paper cites Big bird: Transformers for longer sequences.

Transformers Learn Faster with Semantic Focus Big bird: Transformers for longer sequences

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.778932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.936249Z digest=sha256:216e297ce45eca517a0c54fce948c58e067d3a9ff363d2c29a3268e1bc04032c

Observation 32f6bc68-148f-4bae-a696-780ff6753dc3 · outbound

This paper cites Memory-efficient transformers via top-k attention.

Transformers Learn Faster with Semantic Focus Memory-efficient transformers via top-k attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.940576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.940576Z digest=sha256:5f14ef33e9e00ac90b52adcd5304192e8991d7ef1d0d869126554bcc1e103e68

Observation 2f42847e-789a-4545-ad74-dbd1c9a601cb · outbound

This paper cites ZETA : Leveraging z -order curves for efficient top- k attention.

Transformers Learn Faster with Semantic Focus ZETA : Leveraging z -order curves for efficient top- k attention

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.764128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.945255Z digest=sha256:a2e02dae98190394b4854095374e926c6ed9b04d990722e52087393e5c1ab223

Observation 672ff6cb-90c5-4c82-99a3-f10cd42cd3a6 · outbound

This paper cites Algorithmic stability and generalization performance.

Transformers Learn Faster with Semantic Focus Algorithmic stability and generalization performance

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.750849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.950382Z digest=sha256:f005ee2ab5b10af27d07b1590a6d3692bec90391c7d041bca9a85ed15b5170e0

Observation 8495391a-4fb7-4739-ac0c-aec66433f718 · outbound

This paper cites Train faster, generalize better: Stability of stochastic gradient descent.

Transformers Learn Faster with Semantic Focus Train faster, generalize better: Stability of stochastic gradient descent

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.954810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.954810Z digest=sha256:1149a7bff076d29eb5cfa1ee1c212b1d8a97a5b11b332fe85c8a0095d14871b5

Observation 9bb6606d-3f55-472b-a431-4d0742abc47c · outbound

This paper cites Formal Algorithms for Transformers.

Transformers Learn Faster with Semantic Focus Formal Algorithms for Transformers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.960292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.960292Z digest=sha256:ad506edcf11af83950b43ea86e13376552af3844a5d8778b22ae0bdc04aa9584

Observation dd34a3bd-096a-4936-9e29-698e87b2df1d · outbound

This paper cites A survey of transformers.

Transformers Learn Faster with Semantic Focus A survey of transformers

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.736787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.965609Z digest=sha256:7ba37ed287a019d24dc9341735a832afd614a994b29c75076110207384f8a221

Observation c1bfa5c9-f0c9-4ec2-9862-5b35f64ee4ca · outbound

This paper cites Image transformer.

Transformers Learn Faster with Semantic Focus Image transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.723660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.969645Z digest=sha256:32266c928b3b695c32a981da87fe77dafd2d5df11aa730b85f7432b82737608e

Observation 79b46b3c-40a9-4820-be5e-130bc2e701cb · outbound

This paper cites Blockwise self-attention for long document understanding.

Transformers Learn Faster with Semantic Focus Blockwise self-attention for long document understanding

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.709229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.974200Z digest=sha256:9899868c5c7e2f739e273578aef52749d3a968ffebd74795e215521a2ce1dcf5

Observation f63968ca-efc8-419d-878a-ddb05d57f944 · outbound

This paper cites Longformer: The Long-Document Transformer.

Transformers Learn Faster with Semantic Focus Longformer: The Long-Document Transformer

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.979069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.979069Z digest=sha256:d77a2d620bc9b9cd5bfb6d49fc15314aeec8662c9651adb189ac32ad719da22e

Observation 067581ba-98c3-4144-b6bb-576ff06f2877 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Transformers Learn Faster with Semantic Focus Generating Long Sequences with Sparse Transformers

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:30.983693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:30.983693Z digest=sha256:1a9da8fce0f79f19851bb45c0a31f37994de3e9e25931d63f1eed245ac43c04a

Observation de423e8c-14cf-431b-bb13-0d340ddd288b · outbound

This paper cites Sparse sinkhorn attention.

Transformers Learn Faster with Semantic Focus Sparse sinkhorn attention

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.693446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.988852Z digest=sha256:9bc4c9f2f2f00bac3980c0cc711b1b8be7495ace051ab806bbb2b24726f1e6ab

Observation 8cee7eee-810c-4efa-84d8-49e5f85f3321 · outbound

This paper cites Efficient content-based sparse attention with routing transformers.

Transformers Learn Faster with Semantic Focus Efficient content-based sparse attention with routing transformers

Reference 26

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T00:26:31.355960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.993269Z digest=sha256:077a3aa55197e2dff7fb3166d5163baa144f87aeec31a949f46132b6d151f7fb

Observation 2e7434e7-0ad7-45cb-b59f-e14f2167eae9 · outbound

This paper cites Reformer: The efficient transformer.

Transformers Learn Faster with Semantic Focus Reformer: The efficient transformer

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.678445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:30.998380Z digest=sha256:9a1478da2e26b845bbd304f60321b3b0cab5b48552ebff6dc0173fd8f12a7fb1

Observation d84cf548-c9c4-4cf2-8365-00a069fe0a17 · outbound

This paper cites COGS : A compositional generalization challenge based on semantic interpretation.

Transformers Learn Faster with Semantic Focus COGS : A compositional generalization challenge based on semantic interpretation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.664182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.003042Z digest=sha256:99620b77793b050775c9f3cba576f71d5500c43a6621c3feeb5067e2825e0253

Observation cf06c895-63e9-4ecd-978e-e546c4c82dde · outbound

This paper cites Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks.

Transformers Learn Faster with Semantic Focus Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.649685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.007490Z digest=sha256:4c0b49e2e735e305988e85ce4b1d93451c1196e3e38226237930a4137087cb9f

Observation 47980f98-7844-48dd-9fd5-0af02175dc94 · outbound

This paper cites When can transformers ground and compose: Insights from compositional generalization benchmarks.

Transformers Learn Faster with Semantic Focus When can transformers ground and compose: Insights from compositional generalization benchmarks

Reference 30

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.339636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.011682Z digest=sha256:9407080c216d074e14538881397d948b8d657ab71cf0be4643febd45e885304c

Observation 3ceab004-66f3-41e6-84d4-824a19b41134 · outbound

This paper cites The devil is in the detail: Simple tricks improve systematic generalization of transformers.

Transformers Learn Faster with Semantic Focus The devil is in the detail: Simple tricks improve systematic generalization of transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.016326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.016326Z digest=sha256:44576433ed41242b9d5fe6f217b994b212706ea4f9f5a6a1211a8b0953e457d1

Observation a13d6515-33e0-4e1d-91c8-7e541b43e13d · outbound

This paper cites Making transformers solve compositional tasks.

Transformers Learn Faster with Semantic Focus Making transformers solve compositional tasks

Reference 32

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.312207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.020431Z digest=sha256:eb6c0c75755d2a0eaf8e8d8a724c0962e7cb3251ee1c845022b3a5569905ddeb

Observation 2496269a-85e9-44eb-a5db-d03f7139143a · outbound

This paper cites Inducing transformer ' s compositional generalization ability via auxiliary sequence prediction tasks.

Transformers Learn Faster with Semantic Focus Inducing transformer ' s compositional generalization ability via auxiliary sequence prediction tasks

Reference 33

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.295252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.025449Z digest=sha256:d95c0621d5d124bc7ece8539b0c29ced553ba1817a755a2c53c281c68d2b9407

Observation 84412c7a-508f-461c-a335-db1b31394c91 · outbound

This paper cites a rli, Ekin Aky \.

Transformers Learn Faster with Semantic Focus a rli, Ekin Aky \

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.635646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.029829Z digest=sha256:7726d503b569f2699e984ab8ea768f8d9bf592393f187e28e39cc112c3e06789

Observation f5b2bfd4-f7c9-4872-a2ad-f00fdeacdc96 · outbound

This paper cites What formal languages can transformers express? a survey.

Transformers Learn Faster with Semantic Focus What formal languages can transformers express? a survey

Reference 35

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.279003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.034312Z digest=sha256:d9aefd62f5389eab6987091686f1f5ce41a443567a558b3b5cd6a78a096ac6b1

Observation 6020b2e6-aa1c-47ce-8d90-8afc314dd586 · outbound

This paper cites On the ability and limitations of transformers to recognize formal languages.

Transformers Learn Faster with Semantic Focus On the ability and limitations of transformers to recognize formal languages

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.620841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.038652Z digest=sha256:3f7fa4430ac7d9576177789c20e8996df2b993b1a0257893139da61565435a4b

Observation 1604f60a-e988-4af0-8874-2c4f77626b61 · outbound

This paper cites Theoretical limitations of self-attention in neural sequence models.

Transformers Learn Faster with Semantic Focus Theoretical limitations of self-attention in neural sequence models

Reference 37

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.263821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.042758Z digest=sha256:19db45d8dc50345d607d475cfa8e4f02317ff85e40c02b56f5e41eb0ee3330a1

Observation 2ece83c2-9f63-418c-b1d5-c8ae82a60982 · outbound

This paper cites Formal language recognition by hard attention transformers: Perspectives from circuit complexity.

Transformers Learn Faster with Semantic Focus Formal language recognition by hard attention transformers: Perspectives from circuit complexity

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.605490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.046942Z digest=sha256:b140c89440b7b4667852cf5a46d9f8e80bbbbabc645f8691957e9103dfc96cd5

Observation bb1ab5b5-9b8c-400a-b686-8b77247e18c6 · outbound

This paper cites Saturated transformers are constant-depth threshold circuits.

Transformers Learn Faster with Semantic Focus Saturated transformers are constant-depth threshold circuits

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.589814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.051326Z digest=sha256:050d3aa3ec4f796cc30e3e5813242b6ba3ff77eb8183925b3359430cbe4f786b

Observation f3fe64ff-55f2-46e6-80d0-57bf72a7230a · outbound

This paper cites Overcoming a theoretical limitation of self-attention.

Transformers Learn Faster with Semantic Focus Overcoming a theoretical limitation of self-attention

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.573709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.055666Z digest=sha256:05a0281caf920be044e97a12b796d970241b136a71031456b761235e6b15ee73

Observation 5f80651a-9042-41df-8507-c06510ced51d · outbound

This paper cites Tighter bounds on the expressivity of transformer encoders.

Transformers Learn Faster with Semantic Focus Tighter bounds on the expressivity of transformer encoders

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.557973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.059942Z digest=sha256:71a2584e52e371ef9f6e548997da50d206202266c0fd29d999bddf33fadb2742

Observation e1a1abc0-80ea-4abc-a8b7-530e386ed594 · outbound

This paper cites Transformers as algorithms: Generalization and stability in in-context learning.

Transformers Learn Faster with Semantic Focus Transformers as algorithms: Generalization and stability in in-context learning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.543429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.064187Z digest=sha256:d52d4f6667038d61ad083fab40eb46a2db5dedd6db55e7e89808e7529eb55f90

Observation af5cac0e-1cef-4924-a631-3d6f7a36699f · outbound

This paper cites Transformers learn in-context by gradient descent.

Transformers Learn Faster with Semantic Focus Transformers learn in-context by gradient descent

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.529452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.068351Z digest=sha256:a9f6705c2cd58bcf2417c4e0f7e8c7a5556bf6f1143cbf6c352d4dd38c7b6491

Observation bc5a2ec6-5e4e-4209-8bbe-586adb9acbe7 · outbound

This paper cites The emergence of clusters in self-attention dynamics.

Transformers Learn Faster with Semantic Focus The emergence of clusters in self-attention dynamics

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.514809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.072620Z digest=sha256:62d03650a61dda9e74de4907e87330bd9783d18a41b89ef211a2db9400b6f0a1

Observation 4842f956-d2c0-4cd2-bf07-6a05324354e0 · outbound

This paper cites Transformers learn to implement preconditioned gradient descent for in-context learning.

Transformers Learn Faster with Semantic Focus Transformers learn to implement preconditioned gradient descent for in-context learning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.495826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.077091Z digest=sha256:8c66471f50a6b0edf6b300fb5569f3df4d6fafaeef7fd5135c4e1511eb0a5ae8

Observation 0676e91a-49e1-44af-9ffe-d180a49e2d1e · outbound

This paper cites Trained transformers learn linear models in-context.

Transformers Learn Faster with Semantic Focus Trained transformers learn linear models in-context

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.480464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.082054Z digest=sha256:9c3093aa938a40e2fc5a1cec63192d9c41ef2aa5b90a06890ccb614dde7f9733

Observation f1208b8e-f112-4011-bd87-dd3159dc5249 · outbound

This paper cites Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020.

Transformers Learn Faster with Semantic Focus Why are adaptive methods good for attention models? Advances in Neural Information Processing Systems, 33: 0 15383--15393, 2020

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.464242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.086492Z digest=sha256:dbe6e299ef77090471fb6979ab49b0d0f8e033e4b05a8a3743dd98cd7ec7e6bb

Observation da9fc738-70f8-48b4-8259-7545c05251a4 · outbound

This paper cites Toward understanding why adam converges faster than SGD for transformers.

Transformers Learn Faster with Semantic Focus Toward understanding why adam converges faster than SGD for transformers

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.446017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.091192Z digest=sha256:2141a8420a22b23f0b5321ad388e70f48068e29b4447a9b2225738444d87e16e

Observation c9f67186-e199-4d0c-9ad6-15fa02f759f1 · outbound

This paper cites How does adaptive optimization impact local neural network geometry? Advances in Neural Information Processing Systems, 36: 0 8305--8384, 2023.

Transformers Learn Faster with Semantic Focus How does adaptive optimization impact local neural network geometry? Advances in Neural Information Processing Systems, 36: 0 8305--8384, 2023

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.431099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.095922Z digest=sha256:5cb2f01b544d6af84a9f03a7350cab0dfa3f6e8bea573ec99622e26f20015b3c

Observation dbc1b7cc-1d95-47c1-85fa-4b50bf0a234a · outbound

This paper cites Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be.

Transformers Learn Faster with Semantic Focus Noise is not the main factor behind the gap between sgd and adam on transformers, but sign descent might be

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.414833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.101370Z digest=sha256:461d071680967f3f752b7062a0810b0216a79af9a4434fac1842cba9cda93228

Observation 884fe00e-92a9-460e-848e-cb3bd47f188e · outbound

This paper cites Linear attention is (maybe) all you need (to understand transformer optimization).

Transformers Learn Faster with Semantic Focus Linear attention is (maybe) all you need (to understand transformer optimization)

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.399313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.105944Z digest=sha256:2fc87342f069a18cc439913bfd95e52c9225b986b3627fa1bdfc0fee63fc8123

Observation 0b632e26-88f9-40fd-b30c-2f0591d3f489 · outbound

This paper cites On the optimization and generalization of two-layer transformers with sign gradient descent.

Transformers Learn Faster with Semantic Focus On the optimization and generalization of two-layer transformers with sign gradient descent

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.383165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.110777Z digest=sha256:aa4eb6551cba8b13544ff36240903fc39f35f51c39a0875461f491961eae1c3d

Observation 7d1005fb-2019-4208-b0c1-d3163e3f2a97 · outbound

This paper cites Layer Normalization.

Transformers Learn Faster with Semantic Focus Layer Normalization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.114796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.114796Z digest=sha256:251daf109402f0dd44e42db32b6ea4b3453ee4467b94a9843d4f4f5fcf5fa421

Observation 46115e48-be86-45ce-b49a-e9ce4bdda4f9 · outbound

This paper cites Root mean square layer normalization.

Transformers Learn Faster with Semantic Focus Root mean square layer normalization

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.119039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.119039Z digest=sha256:05cced4dfc7bd0c181593b1eccfdd0cb79d63d2e8e2ad84584651226aa3c9dae

Observation ee4e4e51-cd62-4925-9254-1f342c52ca72 · outbound

This paper cites BERT : Pre-training of deep bidirectional transformers for language understanding.

Transformers Learn Faster with Semantic Focus BERT : Pre-training of deep bidirectional transformers for language understanding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.123158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.123158Z digest=sha256:c1761ddb09e0f31c1b3f1afaf38b5958a871196ad040ea96f201baf51a5f1eb9

Observation d1ad7295-411a-42bf-9db6-42b3ff685421 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Transformers Learn Faster with Semantic Focus Gaussian Error Linear Units (GELUs)

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.128361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.128361Z digest=sha256:78c57eba986bc3bccdb8efedfffd84fe8180aaedc5a3e24a287c52f3077f6705

Observation d521b3a9-96a1-4475-9a3a-d3942c18b279 · outbound

This paper cites Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs).

Transformers Learn Faster with Semantic Focus Fast and Accurate Deep Network Learning by Exponential Linear Units (ELUs)

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.133288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.133288Z digest=sha256:d95361d37d4d6d75f0bcf5c28cf5741f059c98e4708662686e1e0d6393f496ec

Observation 5fa87b52-75d6-4f21-b69c-f76cc01c4f8e · outbound

This paper cites Listops: A diagnostic dataset for latent tree learning.

Transformers Learn Faster with Semantic Focus Listops: A diagnostic dataset for latent tree learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.356077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.141387Z digest=sha256:bb7617985beef68bdd7cf8d24ac7d29ffd6ceb1bf8597bf5806a34447d39f794

Observation bfb46c1d-85a6-4889-b1b2-37873f04faca · outbound

This paper cites Mish: A Self Regularized Non-Monotonic Activation Function.

Transformers Learn Faster with Semantic Focus Mish: A Self Regularized Non-Monotonic Activation Function

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.146060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.146060Z digest=sha256:95db85f8fe78b4d8a24f793da53a0a90e830310e983f02e4b31a051cc7c97322

Observation 7a944d61-2998-44b7-a34b-676ba1b45bf1 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Transformers Learn Faster with Semantic Focus Adam: A Method for Stochastic Optimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.152690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.152690Z digest=sha256:18360414d7421a98a2a1dabdcf7a00e0a6e2e2bc7c538d0247fb3b5852c86e18

Observation 1c0085ca-f463-4103-952b-307c15eb115e · outbound

This paper cites The Lipschitz Constant of Self-Attention.

Transformers Learn Faster with Semantic Focus The Lipschitz Constant of Self-Attention

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.158224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.158224Z digest=sha256:f329f0ae130f6348a0765482ac90d08bca8e215b221c6f43fe7ff880a1a4eb18

Observation 12797a98-aa4b-437c-b189-96d7d2b884b1 · outbound

This paper cites Explicit Sparse Transformer: Concentrated Attention Through Explicit Selection.

Transformers Learn Faster with Semantic Focus Explicit Sparse Transformer: Concentrated Attention Through Explicit Selection

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:31.162921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:26:31.162921Z digest=sha256:bf426bfbcc8fe7bfb08b186bb022ed261a0194d8689ec4986b6a29dadeaf87d3

Observation 37334c11-9643-4161-800f-94b4c7a6bbe5 · outbound

This paper cites Visualizing the loss landscape of neural nets.

Transformers Learn Faster with Semantic Focus Visualizing the loss landscape of neural nets

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:33.275494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.168150Z digest=sha256:734343914e05c3d653125db17ddcacfd02822d1e5fb5abcd01f584f192af4cc6

Observation 10878189-944f-4b9b-8f6c-9907c722def7 · outbound

This paper cites Never train from scratch: Fair comparison of long-sequence models requires data-driven priors.

Transformers Learn Faster with Semantic Focus Never train from scratch: Fair comparison of long-sequence models requires data-driven priors

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:32.980495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.172682Z digest=sha256:f7e5f75f5ea4fc02190d84866469c160145579f4862faff5fb64159965aa6e5e

Observation c25273f1-6b34-4339-bc1d-998ffb03b3da · outbound

This paper cites Learning overparameterized neural networks via stochastic gradient descent on structured data.

Transformers Learn Faster with Semantic Focus Learning overparameterized neural networks via stochastic gradient descent on structured data

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:32.851131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.177991Z digest=sha256:591a314c46796e05a0cdd169166c58132e985f80e4bd5d113a6a3eb2167c5715

Observation 233ed32b-ea5f-4522-944c-d2bb9ef0050c · outbound

This paper cites A convergence theory for deep learning via over-parameterization.

Transformers Learn Faster with Semantic Focus A convergence theory for deep learning via over-parameterization

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:32.582899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.182659Z digest=sha256:5d33629c060a1a2fa9049df1e7ec8c898915be0558b7ab92d19bf7c96160949b

Observation ca681d40-ef8b-457c-962c-fffe1a113051 · outbound

This paper cites Gradient descent optimizes over-parameterized deep relu networks.

Transformers Learn Faster with Semantic Focus Gradient descent optimizes over-parameterized deep relu networks

Reference 67

Resolution
verified exact
doi, observed 2026-08-07T00:26:31.234860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.187525Z digest=sha256:dfa41f5afa58f07e910b3dbce57c0ad2be1d29b319808cd4f17a41646d9244df

Observation aa3a87b7-a9a4-47f5-a37e-2fd076f33f5b · outbound

This paper cites Convergence rates for the stochastic gradient descent method for non-convex objective functions.

Transformers Learn Faster with Semantic Focus Convergence rates for the stochastic gradient descent method for non-convex objective functions

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:32.280375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T00:26:31.193173Z digest=sha256:0f57c15a246f66dca28220777fcdfe83cdf71b7b1f8bee546bbd89215bc85a1f

Pith citing papers

No inbound Pith citation observations are available.