Pith. sign in

Paper Citation Record · LEDGER

FlashDecoding++: Faster Large Language Model Inference on GPUs

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2311.01282.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.01282 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:45:09.857737Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:39:46.798667Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9ef22293-e0aa-4262-89e0-326b6b8ee658 · inbound

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security cites this paper.

Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 270

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:57:26.799232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T00:57:26.303195Z digest=sha256:87432e4c067f3ac4eb4538ecc27f0a9195af9c19974144499ea771f6c595700b

Observation ac17d14e-2ccb-4c22-adc6-cccc746d29ca · inbound

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering cites this paper.

TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:09.857737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:09.857737Z digest=sha256:4079e1be21fe909928f61351bb0421dd8df5c4deeaa5791b8eab92b399711454

Observation 67557a58-4e61-4f7d-9e52-63de85969a1b · inbound

QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm cites this paper.

QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:55:50.245539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:55:50.245539Z digest=sha256:a36f9eba63cbea710d7e3b67c3c91e2ae48bf3b69e31cea3b19402ac200b7707

Observation a4678885-969e-4a5f-86b7-82f6600a50de · inbound

SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations cites this paper.

SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:28.045319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:28.045319Z digest=sha256:f5624f6b8984fd5e6ee2b0375416b06297349e858daa552f5fbc88012dcddaa2

Observation 08228ed5-dff6-4ef2-b535-85988aa99550 · inbound

Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure cites this paper.

Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 204

Resolution
unresolved
no resolver link, observed 2026-08-06T18:50:52.895202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:50:52.895202Z digest=sha256:16ea54453354ca770b402f522b2f5afbaec98f4fc9b4c809475c86b0d7ac1444

Observation 5b123bcb-f408-452d-add1-7d409451bda2 · inbound

Past-Future Scheduler for LLM Serving under SLA Guarantees cites this paper.

Past-Future Scheduler for LLM Serving under SLA Guarantees FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.797553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.797553Z digest=sha256:cf0e56ded0c067d8546e0fae89e982299fe8da2f209cd2f70bd04bc3fe495816

Observation cffde461-3275-4a7e-b093-b1caaca2907a · inbound

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference cites this paper.

Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T17:51:42.142783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T17:47:58.019030Z digest=sha256:abf5311edb4e49a1714f4a7218504055736897dc4254edc0898a26d47180400d

Observation 18b815b4-1dd2-4022-a073-4f091de5632a · inbound

FG-Attn: Leveraging Fine-Grained Sparse Attention in Video Diffusion Models cites this paper.

FG-Attn: Leveraging Fine-Grained Sparse Attention in Video Diffusion Models FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:02.030920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:02.030920Z digest=sha256:3d4c8b421a8da78f39136d3aeb94911d0b80613c9896ad6676bd38f70f6e5fed

Observation f319ba07-7273-4021-8f80-619847606956 · inbound

VFA: Relieving Vector Operations in Flash Attention with Global Maximum Pre-computation cites this paper.

VFA: Relieving Vector Operations in Flash Attention with Global Maximum Pre-computation FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:00.032774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:10:40.858525Z digest=sha256:2320144959236790ce73ace4005bf0ad1c9f01a985f11208cc2fec26c086a438

Observation 20f99678-05a1-4b93-a6aa-fdece8d46f93 · inbound

Prism: Symbolic Superoptimization of Tensor Programs cites this paper.

Prism: Symbolic Superoptimization of Tensor Programs FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:03:24.949228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T09:01:38.783602Z digest=sha256:3cb006db524b60f700ccb7aaff711262099b73bf57df4a9da05e6e61fa3d62bd

Observation ffda947e-5625-4717-b167-f7e1f580643f · inbound

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference cites this paper.

An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:05:54.233012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:56:28.828593Z digest=sha256:ba7c470fa6a1f8e0017f0f304bea3ec872a7119583338c523a731af4ddab776b

Observation 3dccb04e-b102-4f83-801d-e8eac0eec47f · inbound

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents cites this paper.

Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:36:59.405629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T01:07:14.691347Z digest=sha256:29958ccb1f64136d9e7e519e08cdbb147a77eb1857fe8f9c2046ff6206dbefbe

Observation 50cd2131-44bf-4af8-9404-d4bb2279ba9f · inbound

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill cites this paper.

ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:39:46.800083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T09:42:39.573568Z digest=sha256:797104b92c522136f462050015b015e6940711df57f6e7318a3d74df1f7650ef

Observation f9d0b40c-a746-425d-a2eb-54a225b0ff58 · inbound

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling cites this paper.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:22.543190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:22.543190Z digest=sha256:e23e74ce333e18507c3abe569e12397f47d4c95591d2c7def10311677d746d8d