Pith. sign in

Paper Citation Record · LEDGER

Retrofitting Linear Attention into Diffusion Language Models

As of 11 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2608.06628.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06628 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:14:52.259504Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5a2cd5a-19a3-4a67-9705-f5aaf2f93008 · outbound

This paper cites LLaDA2.1: Speeding up text diffusion via token editing.arXiv preprint arXiv:2602.08676,.

Retrofitting Linear Attention into Diffusion Language Models LLaDA2.1: Speeding up text diffusion via token editing.arXiv preprint arXiv:2602.08676,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.215311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.215311Z digest=sha256:b44e40b0b7d70304a38b4b05172be4a014930332c5e0daa84f90fa73aa79ab52

Observation a3bd89b5-400f-4c24-ae2a-c637e226a8f3 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Retrofitting Linear Attention into Diffusion Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.227519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.227519Z digest=sha256:ed57a352709024c3b6f3944ee311e73dd7e2ea87fadb541029190fe09a948873

Observation f2c19505-7d58-4ab0-8f26-18f9c71be134 · outbound

This paper cites The diffusion duality.arXiv preprint arXiv:2506.10892,.

Retrofitting Linear Attention into Diffusion Language Models The diffusion duality.arXiv preprint arXiv:2506.10892,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.234749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.234749Z digest=sha256:29134aad461e207588ff509471f3a24e00dd94090e6663402d55bb119e9ce985

Observation 23b64969-f29e-4ae8-bee2-6ee729f11043 · outbound

This paper cites Simple guidance mechanisms for discrete diffusion models.

Retrofitting Linear Attention into Diffusion Language Models Simple guidance mechanisms for discrete diffusion models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:14:53.431302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:14:52.237955Z digest=sha256:5996cb76c28ddfdcaa769a1f4bd510dc8b83451960724824c1eec7027fc490fb

Observation 45fde0ac-3bb3-45ff-9b5a-00564e33aa02 · outbound

This paper cites Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference.

Retrofitting Linear Attention into Diffusion Language Models Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.241127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.241127Z digest=sha256:9d79eb6510b7f9ac752d0f979d7c48eb88925484b6ecfda0942c2ce5187ae674

Observation 75a91339-753f-4b2c-ade5-355164e84f6f · outbound

This paper cites Discrete diffusion models exploit asymmetry to solve lookahead planning tasks.arXiv preprint arXiv:2602.19980,.

Retrofitting Linear Attention into Diffusion Language Models Discrete diffusion models exploit asymmetry to solve lookahead planning tasks.arXiv preprint arXiv:2602.19980,

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-08-10T04:14:52.860526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:14:52.244505Z digest=sha256:12586af4f599cee4be29a0843e8f475752994334c91a6b46278b7fc20deeaf7e

Observation b661e5b2-1b26-4a38-8fd4-2207cde2dd89 · outbound

This paper cites Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning.

Retrofitting Linear Attention into Diffusion Language Models Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.251220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.251220Z digest=sha256:13bcb9f802ba8297dafcf784be9a0e057ef7570121efaac1319716b612ff79ae

Observation dd92681b-0e94-4a2c-b337-b9e4a095f30a · outbound

This paper cites Dream 7B: Diffusion Large Language Models.

Retrofitting Linear Attention into Diffusion Language Models Dream 7B: Diffusion Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.255473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.255473Z digest=sha256:04718691045be6e285e93864dda289c65edfb4337283ca84ff380db5e264d0c3

Observation f6e50143-e983-4ddc-aab8-923000149b22 · outbound

This paper cites LLaDA-MoE: A sparse MoE diffusion language model.arXiv preprint arXiv:2509.24389,.

Retrofitting Linear Attention into Diffusion Language Models LLaDA-MoE: A sparse MoE diffusion language model.arXiv preprint arXiv:2509.24389,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.259504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.259504Z digest=sha256:3399943aa2dd9a28f7aee3f47f58cb00006483b251d9a8d4905cbbbe5c9638cc

Observation b4fc0e3f-e488-4920-94dc-c5359c886ca1 · outbound

This paper cites Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions.

Retrofitting Linear Attention into Diffusion Language Models Train for the Worst, Plan for the Best: Understanding Token Ordering in Masked Diffusions

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.223396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.223396Z digest=sha256:071f4f0e50aacbe9254c01a85511799d18882369169a3af3da2c46e2593e51f5

Observation 57eac887-b242-4466-82b9-7facf6251a04 · outbound

This paper cites Mercury: Ultra-Fast Language Models Based on Diffusion.

Retrofitting Linear Attention into Diffusion Language Models Mercury: Ultra-Fast Language Models Based on Diffusion

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.219401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.219401Z digest=sha256:88fbbd612abd6563dcc79a02d05167a42acad0ce8df92dc19a3b5c4342b3d29c

Observation 4116384a-676a-4adb-beea-c56fef673046 · outbound

This paper cites an unresolved cited work.

Retrofitting Linear Attention into Diffusion Language Models Unresolved cited work

Reference 2024

Resolution
verified exact
arxiv_id, observed 2026-08-10T04:14:53.207417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:14:52.231187Z digest=sha256:962add63b3de22f800b8191281fa61f45e24932eb2605c2d8c4b096080ccee18

Observation 005303c5-fdda-42e1-911d-718910c9471c · outbound

This paper cites LLaDA2.0: Scaling Up Diffusion Language Models to 100B.

Retrofitting Linear Attention into Diffusion Language Models LLaDA2.0: Scaling Up Diffusion Language Models to 100B

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.210581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.210581Z digest=sha256:e61f3b59bc481eb5d4f3203b978c38376cf6af15468fbd6fac20750f1cfff7fb

Observation 0b87d0c0-a694-4142-bb9f-d2dfbc6d6f56 · outbound

This paper cites Scaling behavior of discrete diffusion language models.arXiv preprint arXiv:2512.10858,.

Retrofitting Linear Attention into Diffusion Language Models Scaling behavior of discrete diffusion language models.arXiv preprint arXiv:2512.10858,

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-10T04:14:52.247619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:14:52.247619Z digest=sha256:70a666178feb94f0b5cd01a83b96c0522bb4081248d7d05f02a82d82d35c1470

Pith citing papers

No inbound Pith citation observations are available.