Pith. sign in

REVIEW 3 cited by

FastAttention: Extend FlashAttention2 to NPUs and Low-resource GPUs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.16663 v1 pith:N5XAX57C submitted 2024-10-22 cs.LG

classification cs.LG
keywords fastattentiongpusflashattentionnpusseriesstrategyspeeduptimes
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

FlashAttention series has been widely applied in the inference of large language models (LLMs). However, FlashAttention series only supports the high-level GPU architectures, e.g., Ampere and Hopper. At present, FlashAttention series is not easily transferrable to NPUs and low-resource GPUs. Moreover, FlashAttention series is inefficient for multi- NPUs or GPUs inference scenarios. In this work, we propose FastAttention which pioneers the adaptation of FlashAttention series for NPUs and low-resource GPUs to boost LLM inference efficiency. Specifically, we take Ascend NPUs and Volta-based GPUs as representatives for designing our FastAttention. We migrate FlashAttention series to Ascend NPUs by proposing a novel two-level tiling strategy for runtime speedup, tiling-mask strategy for memory saving and the tiling-AllReduce strategy for reducing communication overhead, respectively. Besides, we adapt FlashAttention for Volta-based GPUs by redesigning the operands layout in shared memory and introducing a simple yet effective CPU-GPU cooperative strategy for efficient memory utilization. On Ascend NPUs, our FastAttention can achieve a 10.7$\times$ speedup compared to the standard attention implementation. Llama-7B within FastAttention reaches up to 5.16$\times$ higher throughput than within the standard attention. On Volta architecture GPUs, FastAttention yields 1.43$\times$ speedup compared to its equivalents in \texttt{xformers}. Pangu-38B within FastAttention brings 1.46$\times$ end-to-end speedup using FasterTransformer. Coupled with the propose CPU-GPU cooperative strategy, FastAttention supports a maximal input length of 256K on 8 V100 GPUs. All the codes will be made available soon.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. STEEL: Sparsity-Aware Fused Attention for Energy-Efficient Long-Sequence Inference on AMD's XDNA NPU

    cs.DC 2026-07 conditional novelty 6.0 of 10

    STEEL maps fused FlashAttention onto XDNA NPUs with sparsity-aware pipeline placement, cutting energy ~9 imes vs CPU and ~1.75 imes vs GPU and beating prior XDNA attention by ~9.6× latency.

  2. Ascend to Science: Exploration of AI Chips for Scientific Computing

    cs.DC 2026-07 conditional novelty 5.0 of 10

    AI-oriented Ascend NPUs can run scientific workloads with FP32-like accuracy and competitive throughput when algorithms are reformulated and data movement is explicitly orchestrated.

  3. SMC-AI: Scaling Monte Carlo Simulation to Four Trillion Atoms with AI Accelerators

    physics.comp-ph 2026-04 unverdicted novelty 4.0 of 10

    SMC-AI scales Monte Carlo simulations to 4 trillion atoms on AI hardware clusters, achieving 32 times larger systems and 1.3 times higher throughput than prior records while decoupling ML models from the simulation core.

Pith tools