Pith. sign in

Paper Citation Record · LEDGER

Length Generalization of Causal Transformers without Position Encoding

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2404.12224.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.12224 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:13:39.605272Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b8f4df77-7f58-46a0-9c6f-c77f3c4f229f · inbound

Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization cites this paper.

Fourier Position Embedding: Enhancing Attention's Periodic Extension for Length Generalization Length Generalization of Causal Transformers without Position Encoding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:18:54.573270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:18:54.573270Z digest=sha256:7741afb923fa69db3fe3f28e692a8adb3e806aca4efa1a46335a8e34ffafe061

Observation be41265f-654f-4a6d-b09d-514a4dbd1335 · inbound

Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding cites this paper.

Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding Length Generalization of Causal Transformers without Position Encoding

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-10T22:49:45.128017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:49:45.128017Z digest=sha256:45ca7eb193110ae8543182351f49fc9cce8a84b2efb7d3d11bb1f9e67293f396

Observation c72e8e0f-57f3-4f34-b4ce-a6db9d23ebcd · inbound

Scalable-Softmax Is Superior for Attention cites this paper.

Scalable-Softmax Is Superior for Attention Length Generalization of Causal Transformers without Position Encoding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T20:34:01.257094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:34:01.257094Z digest=sha256:5db4b0b5d27cd38958326224bb6690ce624d7722a5ca632c62a11cb6591b195b

Observation 295eda4b-cbf4-4ef2-8810-4abfc8d4b5dc · inbound

Solving Empirical Bayes via Transformers cites this paper.

Solving Empirical Bayes via Transformers Length Generalization of Causal Transformers without Position Encoding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T20:21:58.381466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:21:58.381466Z digest=sha256:1b1440c84f678790d4abc6d021b91056cc21a489bf7ac4776291dca2a3d23e6d

Observation 859f64a4-a17e-4fc0-88dd-caab9bebb4fc · inbound

EfficientLLM: Efficiency in Large Language Models cites this paper.

EfficientLLM: Efficiency in Large Language Models Length Generalization of Causal Transformers without Position Encoding

Reference 294

Resolution
unresolved
no resolver link, observed 2026-08-15T20:13:39.605272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:13:39.605272Z digest=sha256:bc90ba93cb7c1edee5077043b32a657ba0745f75c74f53985974fd304532e206

Observation a7586f8e-18a6-4212-9f91-f085bd39dd63 · inbound

Understanding Transformer from the Perspective of Associative Memory cites this paper.

Understanding Transformer from the Perspective of Associative Memory Length Generalization of Causal Transformers without Position Encoding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.463146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.463146Z digest=sha256:d03d36c4a7cf8b4d91c727936a82749805980501998f12eefcfe88140d1e3af6

Observation a947700b-699d-4ab9-b363-beabfe0252b0 · inbound

Home-made Diffusion Model from Scratch to Hatch cites this paper.

Home-made Diffusion Model from Scratch to Hatch Length Generalization of Causal Transformers without Position Encoding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T04:34:37.047557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:34:37.047557Z digest=sha256:8427be12cad37c5ed815745caa032c083f7c90cb880d0bf5c86c51044cb16fe3

Observation 48e16509-8bb2-4695-9968-b1ebba7cf5a9 · inbound

Dual Triangle Attention: Effective Bidirectional Attention Without Positional Embeddings cites this paper.

Dual Triangle Attention: Effective Bidirectional Attention Without Positional Embeddings Length Generalization of Causal Transformers without Position Encoding

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:20:51.337939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:50:43.716813Z digest=sha256:b147798e214e99106620211dd32eeeec53edac8ed66be45bc2dbf2e43f0c2d0d

Observation e7be42e3-9ee3-4f0e-822c-25ace2c4df9d · inbound

ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory cites this paper.

ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory Length Generalization of Causal Transformers without Position Encoding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T17:09:59.034396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-25T23:52:52.985532Z digest=sha256:1fe22f360232bb7130dee8616ded79631670cb9ff924171c26754468ec26263f

Observation 4b4bbb94-280d-4f19-84a1-fbee2b9f7a9b · inbound

ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory cites this paper.

ATMA: Long-Context Language Modeling via Polar Attention and Gated-Delta Compression Memory Length Generalization of Causal Transformers without Position Encoding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:44:37.720777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T09:38:43.250420Z digest=sha256:4ab3ea8bc6c1ea8ea03efdbc1cbe796b3060f26b12e44bebee6271bb4296f6c5