Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-13T03:24:08.548348Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2607.09385.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-13T03:24:08.548348Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
24 of 24 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8f797521-39f6-4039-a6c6-7b6c7474d6f3 · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU AIOS: LLM Agent Operating System,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cabc7fb-7d3e-4737-bb29-1bfe8855bd16 · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c86430ee-29e0-452a-bc8e-e86670bc3a88 · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU A survey on privacy risks and protection in large language models,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0169cce9-deab-44a9-8f4c-7623e2f15737 · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU AMD XDNA NPU in Ryzen AI Processors,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 791d64a9-ac84-444e-b942-900736d56b01 · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51625f01-9fe2-453c-8512-c5041da42fcb · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU NITRO: LLM Inference on Intel Laptop NPUs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af5092ae-d8e8-4f46-86e4-7fe8a3f2fb01 · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e0449c7-df1e-4cf9-bb57-244104506728 · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec5e0bda-8a02-41f0-904c-6d37dd09f835 · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49393d3c-451d-4193-8683-07a989579fdb · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU PACE: An Optimal Piecewise Polynomial Approximation Unit for Flexible and Efficient Transformer Non- linearity Acceleration
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 298e2b47-4141-406d-9e07-8406b95e6788 · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Efficiency, Expressivity, and Extensibility in a Close- to-Metal NPU Programming Interface,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5179c24-4ad6-4893-8a0c-ee762a67cb6d · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04dc20c6-24a3-41c4-a0f9-208596bc28e4 · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Dato: A Task-Based Programming Model for Dataflow Accelerators
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e84660a6-0c51-43d8-9400-dd4e96c4dbab · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU The Llama 3 Herd of Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81227a4a-1e43-4901-bdfe-527f98b79a34 · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Attention is All you Need,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7938933-4197-4fa2-9022-0f06f642c0fb · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Online normalizer calculation for softmax
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cf0515c-fe3e-4696-b449-af178db110ce · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Qualcomm Hexagon DSP: An architecture optimized for mobile multimedia and communications,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf284211-aa75-478b-ba22-d8d79a5c40a4 · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU ASCEND-CC: Confidential Computing on Heterogeneous NPU for Emerging Generative AI Workloads,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89afd327-e73c-4ae8-a86b-eec2d7666a7d · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU SPARTA: Spatial Acceleration for Efficient and Scalable Horizontal Diffusion Weather Stencil Computation,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9be7f1d-72f1-4fb9-8a13-f64e5799deea · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU You Only Look Once: Unified, Real-Time Object Detection,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 939f641c-a518-4d11-a121-0072a340699f · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU 14.5 Envision: A 0.26-to-10TOPS/W subword- parallel dynamic-voltage-accuracy-frequency-scalable Convolutional Neural Network processor in 28nm FDSOI,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3e47a6b-c8dd-4345-ac33-9ceea2c88789 · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Basic Linear Algebra Subprograms for Fortran Usage,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 231500f8-a40c-4e67-a058-00385c9cff44 · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU AKG: automatic kernel generation for neural processing units using polyhedral transformations,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99b59607-afdf-41b0-a2ad-0073f3ef1f0b · outbound
STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU PyTorch: An Imperative Style, High-Performance Deep Learning Library,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.