Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2311.01282.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:45:09.857737Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T09:39:46.798667Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 9ef22293-e0aa-4262-89e0-326b6b8ee658 · inbound
Personal LLM Agents: Insights and Survey about the Capability, Efficiency and Security FlashDecoding++: Faster Large Language Model Inference on GPUs
Reference 270
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ac17d14e-2ccb-4c22-adc6-cccc746d29ca · inbound
TaDA: Training-free recipe for Decoding with Adaptive KV Cache Compression and Mean-centering FlashDecoding++: Faster Large Language Model Inference on GPUs
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67557a58-4e61-4f7d-9e52-63de85969a1b · inbound
QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm FlashDecoding++: Faster Large Language Model Inference on GPUs
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4678885-969e-4a5f-86b7-82f6600a50de · inbound
SD-Acc: Accelerating Stable Diffusion through Phase-aware Sampling and Hardware Co-Optimizations FlashDecoding++: Faster Large Language Model Inference on GPUs
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08228ed5-dff6-4ef2-b535-85988aa99550 · inbound
Compute Can't Handle the Truth: Why Communication Tax Prioritizes Memory and Interconnects in Modern AI Infrastructure FlashDecoding++: Faster Large Language Model Inference on GPUs
Reference 204
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b123bcb-f408-452d-add1-7d409451bda2 · inbound
Past-Future Scheduler for LLM Serving under SLA Guarantees FlashDecoding++: Faster Large Language Model Inference on GPUs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cffde461-3275-4a7e-b093-b1caaca2907a · inbound
Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference FlashDecoding++: Faster Large Language Model Inference on GPUs
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 18b815b4-1dd2-4022-a073-4f091de5632a · inbound
FG-Attn: Leveraging Fine-Grained Sparse Attention in Video Diffusion Models FlashDecoding++: Faster Large Language Model Inference on GPUs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f319ba07-7273-4021-8f80-619847606956 · inbound
VFA: Relieving Vector Operations in Flash Attention with Global Maximum Pre-computation FlashDecoding++: Faster Large Language Model Inference on GPUs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 20f99678-05a1-4b93-a6aa-fdece8d46f93 · inbound
Prism: Symbolic Superoptimization of Tensor Programs FlashDecoding++: Faster Large Language Model Inference on GPUs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ffda947e-5625-4717-b167-f7e1f580643f · inbound
An Efficient Hybrid Sparse Attention with CPU-GPU Parallelism for Long-Context Inference FlashDecoding++: Faster Large Language Model Inference on GPUs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3dccb04e-b102-4f83-801d-e8eac0eec47f · inbound
Vortex: Efficient and Programmable Sparse Attention Serving for AI Agents FlashDecoding++: Faster Large Language Model Inference on GPUs
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 50cd2131-44bf-4af8-9404-d4bb2279ba9f · inbound
ASAP: A Disaggregated and Asynchronous Inference System for MoE Prefill FlashDecoding++: Faster Large Language Model Inference on GPUs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f9d0b40c-a746-425d-a2eb-54a225b0ff58 · inbound
Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling FlashDecoding++: Faster Large Language Model Inference on GPUs
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.