Pith. sign in

Paper Citation Record · LEDGER

Retrofitting Linear Attention into Diffusion Language Models

As of 11 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2608.06628.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06628 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:14:52.259504Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5a2cd5a-19a3-4a67-9705-f5aaf2f93008 · outbound

This paper cites LLaDA2.1: Speeding up text diffusion via token editing.arXiv preprint arXiv:2602.08676,.

Retrofitting Linear Attention into Diffusion Language Models LLaDA2.1: Speeding up text diffusion via token editing.arXiv preprint arXiv:2602.08676,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.215311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.215311Z digest=sha256:698a6b737cb63014989dae82bfa8e43797634c5fe893ba5ae84bbd32a6e609cd

Observation a3bd89b5-400f-4c24-ae2a-c637e226a8f3 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Retrofitting Linear Attention into Diffusion Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.227519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.227519Z digest=sha256:a46e37caf43e9978a307675eafe063e2728e9f756e98f6e718927ce1ce05d91d

Observation f2c19505-7d58-4ab0-8f26-18f9c71be134 · outbound

This paper cites The diffusion duality.arXiv preprint arXiv:2506.10892,.

Retrofitting Linear Attention into Diffusion Language Models The diffusion duality.arXiv preprint arXiv:2506.10892,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.234749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.234749Z digest=sha256:2e656cde84cf9a34f3c060d310c51c6cede15e447c859aa4b4b4192ce228335a

Observation 23b64969-f29e-4ae8-bee2-6ee729f11043 · outbound

This paper cites Simple guidance mechanisms for discrete diffusion models.

Retrofitting Linear Attention into Diffusion Language Models Simple guidance mechanisms for discrete diffusion models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:14:53.431302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:14:52.237955Z digest=sha256:bdf6cae1990307d5c290952d31ffbf516aee7813e850b6298baef740e3c052fa

Observation 45fde0ac-3bb3-45ff-9b5a-00564e33aa02 · outbound

This paper cites Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference.

Retrofitting Linear Attention into Diffusion Language Models Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.241127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.241127Z digest=sha256:e52457dd6f63df67a333dfd8b96435e440efc86772b573ed1c7fd8ac015dc3ea

Observation 75a91339-753f-4b2c-ade5-355164e84f6f · outbound

This paper cites Discrete diffusion models exploit asymmetry to solve lookahead planning tasks.arXiv preprint arXiv:2602.19980,.

Retrofitting Linear Attention into Diffusion Language Models Discrete diffusion models exploit asymmetry to solve lookahead planning tasks.arXiv preprint arXiv:2602.19980,

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-08-10T04:14:52.860526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:14:52.244505Z digest=sha256:8d2889f4bf35c3c86a66a63828f54c1721c664f1d8d2df0d50581e0a875171a1

Observation b661e5b2-1b26-4a38-8fd4-2207cde2dd89 · outbound

This paper cites Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning.

Retrofitting Linear Attention into Diffusion Language Models Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.251220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.251220Z digest=sha256:a05fbeb4f3b111b0443aca4f73fe297231863cb1d65b5c5204d73a3179ad75aa

Observation dd92681b-0e94-4a2c-b337-b9e4a095f30a · outbound

This paper cites Dream 7B: Diffusion Large Language Models.

Retrofitting Linear Attention into Diffusion Language Models Dream 7B: Diffusion Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.255473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.255473Z digest=sha256:3a681c58f995ca59d93d9852639d6d3727f03077f3eacdb5f94918c2edb289a6

Observation f6e50143-e983-4ddc-aab8-923000149b22 · outbound

This paper cites LLaDA-MoE: A sparse MoE diffusion language model.arXiv preprint arXiv:2509.24389,.

Retrofitting Linear Attention into Diffusion Language Models LLaDA-MoE: A sparse MoE diffusion language model.arXiv preprint arXiv:2509.24389,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.259504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.259504Z digest=sha256:d1057c80730f9c70a3b07a22eec654b4374b822e5aad7784d1c68d525f6a122f

Observation b4fc0e3f-e488-4920-94dc-c5359c886ca1 · outbound

This paper cites Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions.

Retrofitting Linear Attention into Diffusion Language Models Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.223396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.223396Z digest=sha256:cdab485ab96808b89171224143da4592a702f8215257f3a2e42186db3ac23061

Observation 57eac887-b242-4466-82b9-7facf6251a04 · outbound

This paper cites Mercury: Ultra-Fast Language Models Based on Diffusion.

Retrofitting Linear Attention into Diffusion Language Models Mercury: Ultra-Fast Language Models Based on Diffusion

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.219401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.219401Z digest=sha256:2a49630d2d2bd65d490670b1ed93d0967287b645331b71882648e59184cab00a

Observation 4116384a-676a-4adb-beea-c56fef673046 · outbound

This paper cites an unresolved cited work.

Retrofitting Linear Attention into Diffusion Language Models Unresolved cited work

Reference 2024

Resolution
verified exact
arxiv_id, observed 2026-08-10T04:14:53.207417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:14:52.231187Z digest=sha256:8938feab6864ef1c8a905efe82b20a1895ad5d9df9441b478bd090f547e16c97

Observation 005303c5-fdda-42e1-911d-718910c9471c · outbound

This paper cites LLaDA2.0: Scaling Up Diffusion Language Models to 100B.

Retrofitting Linear Attention into Diffusion Language Models LLaDA2.0: Scaling Up Diffusion Language Models to 100B

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.210581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.210581Z digest=sha256:d4f499d0c18144482b218d67c7796ffc0a41aac7d3b771f81cc6f5ac7d760ad1

Observation 0b87d0c0-a694-4142-bb9f-d2dfbc6d6f56 · outbound

This paper cites Scaling behavior of discrete diffusion language models.arXiv preprint arXiv:2512.10858,.

Retrofitting Linear Attention into Diffusion Language Models Scaling behavior of discrete diffusion language models.arXiv preprint arXiv:2512.10858,

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.247619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.247619Z digest=sha256:18ca23c9292781978c62a74b86d18a2bedfbd3ad55235b06b43bef2249953617

Pith citing papers

No inbound Pith citation observations are available.