Pith. sign in

Paper Citation Record · LEDGER

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU

As of 7 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2607.09385.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.09385 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T03:24:08.548348Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8f797521-39f6-4039-a6c6-7b6c7474d6f3 · outbound

This paper cites AIOS: LLM Agent Operating System,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU AIOS: LLM Agent Operating System,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:7feea25b0c5e0d1f2eaaccf439e3ddc21239f0319dfc61ed5a5bc09a2a25ef54

Observation 6cabc7fb-7d3e-4737-bb29-1bfe8855bd16 · outbound

This paper cites Intelligence per Watt: Measuring Intelligence Efficiency of Local AI.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Intelligence per Watt: Measuring Intelligence Efficiency of Local AI

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:0beeaeb8f1abfaa120cfccda5e2ca3f0653b862a38ce443d2a2825db98731dbc

Observation c86430ee-29e0-452a-bc8e-e86670bc3a88 · outbound

This paper cites A survey on privacy risks and protection in large language models,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU A survey on privacy risks and protection in large language models,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:334173863012fc6684534835943236c15648123909a2b53700855464586f44b0

Observation 0169cce9-deab-44a9-8f4c-7623e2f15737 · outbound

This paper cites AMD XDNA NPU in Ryzen AI Processors,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU AMD XDNA NPU in Ryzen AI Processors,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:6829ab1c44261255c8cf1bfee80f0ea99adddad0e91e271bbae7fb287b131b7d

Observation 791d64a9-ac84-444e-b942-900736d56b01 · outbound

This paper cites FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:3edccc47b2a59d035640a85e5bf037485d378cdc58221669973c9b162c55cb14

Observation 51625f01-9fe2-453c-8512-c5041da42fcb · outbound

This paper cites NITRO: LLM Inference on Intel Laptop NPUs.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU NITRO: LLM Inference on Intel Laptop NPUs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:56344907ecc5dd69adedc108662a83d1a0b12f3c620ad877a11b1a1237aaaa4a

Observation af5092ae-d8e8-4f46-86e4-7fe8a3f2fb01 · outbound

This paper cites LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:b933bf0cd39adacad4a6bc770e56d224ba15346ad179404f70bcd75a101ac64d

Observation 2e0449c7-df1e-4cf9-bb57-244104506728 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:0370045a750ab8e297108f3e3ec573ef55a12699e8c9ef923623af4fa24dd010

Observation ec5e0bda-8a02-41f0-904c-6d37dd09f835 · outbound

This paper cites ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU ITA: An Energy-Efficient Attention and Softmax Accelerator for Quantized Transformers,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:2be3ab910ba418ce99f3be176eebb77bec3e05480e5fd5f5cff619ede9cdd8d3

Observation 49393d3c-451d-4193-8683-07a989579fdb · outbound

This paper cites PACE: An Optimal Piecewise Polynomial Approximation Unit for Flexible and Efficient Transformer Non- linearity Acceleration.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU PACE: An Optimal Piecewise Polynomial Approximation Unit for Flexible and Efficient Transformer Non- linearity Acceleration

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:ef20afd58620eb6070439591d0bf05e4511b892c01a744ff6fd61d871281c9c0

Observation 298e2b47-4141-406d-9e07-8406b95e6788 · outbound

This paper cites Efficiency, Expressivity, and Extensibility in a Close- to-Metal NPU Programming Interface,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Efficiency, Expressivity, and Extensibility in a Close- to-Metal NPU Programming Interface,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:3bad0ee20a0bf002b395c9012d7e78646a4fbe64d286e7dba5f94c1c97bb6c7d

Observation c5179c24-4ad6-4893-8a0c-ee762a67cb6d · outbound

This paper cites Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:c383f13ede38d8d9755e4b72c53898e091336cbe5915dfaab2c73d7b6a118911

Observation 04dc20c6-24a3-41c4-a0f9-208596bc28e4 · outbound

This paper cites Dato: A Task-Based Programming Model for Dataflow Accelerators.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Dato: A Task-Based Programming Model for Dataflow Accelerators

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:0f5460c21456ba516e4c9c5e3d1a8ab0e6ac991794b4db6d67cf5453800f9aaf

Observation e84660a6-0c51-43d8-9400-dd4e96c4dbab · outbound

This paper cites The Llama 3 Herd of Models.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:aae7fcb72c98d39ad8af3e602ec62843913842a715041902c493e0a966c9bf21

Observation 81227a4a-1e43-4901-bdfe-527f98b79a34 · outbound

This paper cites Attention is All you Need,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Attention is All you Need,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:27da25557b93c39084a7335a8aca3b23653add4430beed6cbc36cd7c1230c811

Observation e7938933-4197-4fa2-9022-0f06f642c0fb · outbound

This paper cites Online normalizer calculation for softmax.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Online normalizer calculation for softmax

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:1f4f3121f88f18b420466c52cd37f040a48ee4d9f098298a035f15ca6c443815

Observation 7cf0515c-fe3e-4696-b449-af178db110ce · outbound

This paper cites Qualcomm Hexagon DSP: An architecture optimized for mobile multimedia and communications,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Qualcomm Hexagon DSP: An architecture optimized for mobile multimedia and communications,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:84e9db845933785472e8fdc4b4a37f0057214635717887b0bca25de5c533be3c

Observation bf284211-aa75-478b-ba22-d8d79a5c40a4 · outbound

This paper cites ASCEND-CC: Confidential Computing on Heterogeneous NPU for Emerging Generative AI Workloads,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU ASCEND-CC: Confidential Computing on Heterogeneous NPU for Emerging Generative AI Workloads,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:fb8904e126a05fcf0ab70c0ce891802dd54cb5066f7351f71592212491c8b7e4

Observation 89afd327-e73c-4ae8-a86b-eec2d7666a7d · outbound

This paper cites SPARTA: Spatial Acceleration for Efficient and Scalable Horizontal Diffusion Weather Stencil Computation,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU SPARTA: Spatial Acceleration for Efficient and Scalable Horizontal Diffusion Weather Stencil Computation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:f2e91391729f35cd94ade5379ca062ee075e191a15675a26ad52ad98233318a5

Observation e9be7f1d-72f1-4fb9-8a13-f64e5799deea · outbound

This paper cites You Only Look Once: Unified, Real-Time Object Detection,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU You Only Look Once: Unified, Real-Time Object Detection,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:082a6d420cd874e68e578c8ac80a50de12ed1f15ed7b5f32a3bfd04a4f952384

Observation 939f641c-a518-4d11-a121-0072a340699f · outbound

This paper cites 14.5 Envision: A 0.26-to-10TOPS/W subword- parallel dynamic-voltage-accuracy-frequency-scalable Convolutional Neural Network processor in 28nm FDSOI,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU 14.5 Envision: A 0.26-to-10TOPS/W subword- parallel dynamic-voltage-accuracy-frequency-scalable Convolutional Neural Network processor in 28nm FDSOI,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:7e930cccd766daa5ea81b0d4e701d9fed1e88f6d040752dc709069a53640dcc0

Observation c3e47a6b-c8dd-4345-ac33-9ceea2c88789 · outbound

This paper cites Basic Linear Algebra Subprograms for Fortran Usage,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU Basic Linear Algebra Subprograms for Fortran Usage,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:15dfc83d2685441e978fb2a3afb6e42cbc42bea0141de8a9ed80abdc9ed1d1f6

Observation 231500f8-a40c-4e67-a058-00385c9cff44 · outbound

This paper cites AKG: automatic kernel generation for neural processing units using polyhedral transformations,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU AKG: automatic kernel generation for neural processing units using polyhedral transformations,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:fc0cd323c3c21662dd79a7e48227c2e442406a1d031cc3d0556d2f9eee8ba33a

Observation 99b59607-afdf-41b0-a2ad-0073f3ef1f0b · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library,.

STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU PyTorch: An Imperative Style, High-Performance Deep Learning Library,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-13T03:24:08.548348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:24:08.548348Z digest=sha256:5ee62f247cb0a762ea772c36685fe8d696ba4ea6ead2f5756a7880fb9bcbc91f

Pith citing papers

No inbound Pith citation observations are available.