Pith. sign in

REVIEW 8 cited by

SageAttention2++: A More Efficient Implementation of SageAttention2

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.21136 v3 pith:7NBX5D2Z submitted 2025-05-27 cs.LG cs.AIcs.ARcs.CV

SageAttention2++: A More Efficient Implementation of SageAttention2

classification cs.LG cs.AIcs.ARcs.CV
keywords sageattention2attentionmatmulacceleratefasterinstructionacceleratesaccumulated
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

The efficiency of attention is critical because its time complexity grows quadratically with sequence length. SageAttention2 addresses this by utilizing quantization to accelerate matrix multiplications (Matmul) in attention. To further accelerate SageAttention2, we propose to utilize the faster instruction of FP8 Matmul accumulated in FP16. The instruction is 2x faster than the FP8 Matmul used in SageAttention2. Our experiments show that SageAttention2++ achieves a 3.9x speedup over FlashAttention while maintaining the same attention accuracy as SageAttention2. This means SageAttention2++ effectively accelerates various models, including those for language, image, and video generation, with negligible end-to-end metrics loss. The code will be available at https://github.com/thu-ml/SageAttention.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Speculative Decoding for Autoregressive Video Generation

    cs.CV 2026-04 conditional novelty 7.0

    A training-free speculative decoding method for block-based autoregressive video diffusion uses a quality router on worst-frame ImageReward scores to accept drafter proposals, achieving up to 2.09x speedup at 95.7% qu...

  2. MXAttention: Data-Free Optimal Scaling and Pre-Normalization Quantization for MXFP4 Attention

    cs.LG 2026-07 conditional novelty 6.5

    A closed-form MXFP4 scale boundary Qmax=7.25 plus pre-normalization softmax quantization recovers nearly all FP16 video quality lost by naive MXFP4 attention.

  3. CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention

    cs.CL 2026-07 conditional novelty 6.0

    CoSA couples an order-aware sparse-mask proxy with a skip-capable attention kernel, achieving higher long-context accuracy at lower compute budgets than existing sparse-attention baselines.

  4. Vidu S1: A Real-Time Interactive Video Generation Model

    cs.CV 2026-07 conditional novelty 6.0

    Vidu S1 generates voice-controlled interactive avatar video in real time at 540p/42 FPS with claimed infinite stable streams and top reported quality metrics.

  5. Vidu S1: A Real-Time Interactive Video Generation Model

    cs.CV 2026-07 conditional novelty 6.0

    Vidu S1 generates 540p interactive avatars at up to 42 FPS with voice control, using diffusion forcing, DMD distillation, and a serving stack.

  6. Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms

    eess.IV 2026-03 unverdicted novelty 6.0

    Video generation models can function as world simulators if efficiency gaps in spatiotemporal modeling are bridged via organized paradigms, architectures, and algorithms.

  7. P-Cast Precision in FP8 Attention: Sink-Induced Collapse and the Optimality of S=2^8

    cs.AR 2026-06 unverdicted novelty 5.0

    Forward KV iteration in FP8 attention produces P-collapse under attention sink; reverse iteration with S=256 removes it and is optimal among bit-exact scales.

  8. Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms

    eess.IV 2026-03 conditional novelty 5.0

    In twisted bilayer nodal d-wave superconductors, interlayer hopping creates nodes on the C2 axis and Bogoliubov flat bands when the single-layer Berry connection is parallel to that axis.