Pith. sign in

Paper Citation Record · LEDGER

Attention Drift: What Autoregressive Speculative Decoding Models Learn

As of 5 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 2 inbound Pith citation observations for arXiv:2605.09992.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.09992 v1

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T02:14:31.291528Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T07:13:09.677634Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

22 of 22 outbound references displayed

  • verified exact3
  • verified fuzzy19
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0387b58c-05c3-4265-902c-2d54fae86713 · outbound

This paper cites Fast inference from transformers via speculative decoding.

Attention Drift: What Autoregressive Speculative Decoding Models Learn Fast inference from transformers via speculative decoding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.371364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:f9ea7d79024e5dda6fb3c526be45244527d422e79fdcf1bcdbb34f51a3edff96

Observation 04be06cf-0d12-4c53-b899-5e798f6d0449 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Attention Drift: What Autoregressive Speculative Decoding Models Learn Accelerating Large Language Model Decoding with Speculative Sampling

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T02:16:15.854513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:abfc41c5199d6b51ce2abe661423a5abd0b42ddf44f87b61e1e0bf5d13bf8bb4

Observation 8a614a83-1c78-4a4c-ac42-4753069b7f64 · outbound

This paper cites Duoattention: Efficient long-context LLM inference with retrieval and streaming heads.

Attention Drift: What Autoregressive Speculative Decoding Models Learn Duoattention: Efficient long-context LLM inference with retrieval and streaming heads

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.296799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:516454cf3ef6c062a80e8a26fbf53b5dc5becc73103104a2ae5f2888342ea980

Observation e47de48d-062c-49b2-b8cf-ce17fbb743bb · outbound

This paper cites Longspec: Long-context lossless speculative decoding with efficient drafting and verification.

Attention Drift: What Autoregressive Speculative Decoding Models Learn Longspec: Long-context lossless speculative decoding with efficient drafting and verification

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.323845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:15aaf951af7e5816b91548f5414c3333ac550411f9134828e867163e2a2827e0

Observation ca8dd930-7ca6-429c-b109-8e24fcf1f816 · outbound

This paper cites EAGLE-3: Scaling up inference acceleration of large language models via training-time test.

Attention Drift: What Autoregressive Speculative Decoding Models Learn EAGLE-3: Scaling up inference acceleration of large language models via training-time test

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.333157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:aee103fb41951e4aecc9be1a165162b78dc863165b868bdc6e34edace89cd743

Observation d259b6b1-273a-455d-ac82-82a933c51a06 · outbound

This paper cites Better & faster large language models via multi-token prediction.

Attention Drift: What Autoregressive Speculative Decoding Models Learn Better & faster large language models via multi-token prediction

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.352219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:85169f1c7cf81de774dca38545f1e6f9d1eaa64959aa2e659e321ebcb9ff1db6

Observation 51d116c3-4000-4dab-b313-d20aad3eb29b · outbound

This paper cites EAGLE: speculative sampling requires rethinking feature uncertainty.

Attention Drift: What Autoregressive Speculative Decoding Models Learn EAGLE: speculative sampling requires rethinking feature uncertainty

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.339075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:60ba1d187f84c36d584a5d1d3eedbb830906e6491e3a844e85febf5580ae7d63

Observation b0ac6791-5bd7-4930-b727-deb685133950 · outbound

This paper cites Eagle-2: Faster inference of language models with dynamic draft trees.

Attention Drift: What Autoregressive Speculative Decoding Models Learn Eagle-2: Faster inference of language models with dynamic draft trees

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.343701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:109aeb1fc5ec9a8fda338461ef676abf8016e8325b9ebcc3512fb98d1efc9a27

Observation cac23b1d-ddc5-43b3-8f45-f32c202f524b · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30.

Attention Drift: What Autoregressive Speculative Decoding Models Learn Attention is all you need.Advances in neural information processing systems, 30

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.302388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:af4a1ea923a23d5f77e9b3b01eb75ea49668c4861455d690039ab3d430892af1

Observation 132f235c-2a89-45ba-a87f-809672898e07 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Attention Drift: What Autoregressive Speculative Decoding Models Learn Efficient memory management for large language model serving with pagedattention

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.328690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:59f1c6a1551005471bdae4120da7d58f473822eadc15bd300ab0d7a3aeac5016

Observation 89ea6913-6648-4327-9853-1f4ef9799948 · outbound

This paper cites Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583.

Attention Drift: What Autoregressive Speculative Decoding Models Learn Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.361470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:1ea89c87accb1f556ced78d79e337de3ef4caced25cac81eb07406d450438cc9

Observation 720fd475-ba19-471b-bd0f-8dc0dbda3935 · outbound

This paper cites Efficient streaming language models with attention sinks.

Attention Drift: What Autoregressive Speculative Decoding Models Learn Efficient streaming language models with attention sinks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.356699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:5ffb3e085d7be94ad82420a50cd0612b5e4a2f2446d52e745f26b198bdede360

Observation 024fafc6-3e0e-4638-9983-41b1cf7eb0ae · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Attention Drift: What Autoregressive Speculative Decoding Models Learn Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.310433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:45511d72481800ddd5984f4cf918547940ba26fc96d8a40a15a8308ba44c92ec

Observation 1c9b20d8-3669-486f-a094-0ab4d07a05c8 · outbound

This paper cites Attention Residuals.

Attention Drift: What Autoregressive Speculative Decoding Models Learn Attention Residuals

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:39:04.582456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:da924e30092d644bdd7583a4994200a2079fa354556b20002c0f0d3b5e011a78

Observation 0c790067-0c52-427d-b849-3ba7b8d685e5 · outbound

This paper cites Magicdec: Breaking the latency-throughput tradeoff for long context generation with speculative decoding.

Attention Drift: What Autoregressive Speculative Decoding Models Learn Magicdec: Breaking the latency-throughput tradeoff for long context generation with speculative decoding

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.375683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:fa88d763b8e2532041adda1b38cae087789083c80acf096a9c0a982f40327047

Observation 0a9b10bc-5fe9-40dd-8e5e-fb2724a92de1 · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding.

Attention Drift: What Autoregressive Speculative Decoding Models Learn Longbench: A bilingual, multitask benchmark for long context understanding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.319566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:359ba4791b357a24a8f4f7e665af86d13bc7a71f94faf9ab93b0b41ee59dee4c

Observation de26b050-9c4e-451b-a912-62c944659849 · outbound

This paper cites Lee, Deming Chen, and Tri Dao.

Attention Drift: What Autoregressive Speculative Decoding Models Learn Lee, Deming Chen, and Tri Dao

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.366373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:0f7b2a84b951f4acbdfd9753379bb44bdce1100a94dae4a6b2e7a4297699d948

Observation 0aafff01-a860-40ae-bf4b-d938a95c0f41 · outbound

This paper cites Hydra: Sequentially-dependent draft heads for medusa decoding.

Attention Drift: What Autoregressive Speculative Decoding Models Learn Hydra: Sequentially-dependent draft heads for medusa decoding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.314741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:1a1caad9a2d19a6c9f8d8b925550f557c15071dd3627c01068524a7b43bb78db

Observation 54740bfb-1d04-4ef3-85ef-d25179d24f23 · outbound

This paper cites DFlash: Block Diffusion for Flash Speculative Decoding.

Attention Drift: What Autoregressive Speculative Decoding Models Learn DFlash: Block Diffusion for Flash Speculative Decoding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:05:01.756664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:2b1d11420ac426f46d5e4d014f3279b3a36fb946fc4c8338483d9db6d22273c7

Observation a3e90fda-dfdd-442b-971a-b4a880b43114 · outbound

This paper cites Gated at- tention for large language models: Non-linearity, sparsity, and attention-sink-free.

Attention Drift: What Autoregressive Speculative Decoding Models Learn Gated at- tention for large language models: Non-linearity, sparsity, and attention-sink-free

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.292130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:11041fef4c8e6b1112d49444e011dff1c8d97da9a993501466feb811286cd565

Observation ff74305a-cc0d-4a8d-9c7f-25cfc281dd91 · outbound

This paper cites On layer normalization in the transformer architecture.

Attention Drift: What Autoregressive Speculative Decoding Models Learn On layer normalization in the transformer architecture

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.306473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:e105f7372215d4d78b9b0f905f69090b16d629523f839abc5c2278fec08a9f44

Observation a3190540-889a-43fd-a487-a416e9835f3c · outbound

This paper cites On the role of attention masks and layernorm in transformers.

Attention Drift: What Autoregressive Speculative Decoding Models Learn On the role of attention masks and layernorm in transformers

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T23:36:57.348119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T02:14:31.291528Z digest=sha256:33b2f4d023faf4e3d7dcedc0f84516b91702dd64f4918cde7f25c111d2a5d433

Pith citing papers

Observation 3c641cb6-7648-4160-948f-83239dbf19fd · inbound

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation cites this paper.

DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation Attention Drift: What Autoregressive Speculative Decoding Models Learn

Reference 130

Resolution
unresolved
no resolver link, observed 2026-07-11T08:05:17.460513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T08:05:17.460513Z digest=sha256:d3e318c304caee9034c3a08f81a67440987ab719758376387ab12974b3fedeb7

Observation 7bbb746f-acd3-4787-8a93-307723dcbcf4 · inbound

Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context cites this paper.

Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context Attention Drift: What Autoregressive Speculative Decoding Models Learn

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T07:13:09.677634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:13:09.677634Z digest=sha256:a6236b1c91653bdecd4def2f89eda23a1abd9434b6a4a39bf5287456fb9aacfd