Pith. sign in

REVIEW 6 cited by

Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.04093 v4 pith:NLEIIDRU submitted 2024-08-07 cs.LG cs.CL

classification cs.LGcs.CL
keywords attentiontreedecodingacrossfasterlessnodesreduction
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Our formulation reveals that the reduction across the sequence axis can be efficiently computed in parallel through a tree reduction. Our algorithm, called Tree Attention, for parallelizing exact attention computation across multiple GPUs enables cross-device decoding to be performed asymptotically faster (up to 8x faster in our experiments) than state-of-the-art approaches such as Ring Attention, while also requiring significantly less communication volume and incurring 2x less peak memory. We demonstrate that Tree Attention speeds up decoding up to 4x on Llama 3.1-8B and can be applied to a variety of hardware and networking setups such as H100 DGX nodes, AMD MI300x nodes, and PCIe connected NVIDIA RTX 4090s. Our code is publicly available here: https://github.com/Zyphra/tree_attention

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. ZAYA1-8B Technical Report

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    ZAYA1-8B is a reasoning MoE model with 700M active parameters that matches larger models on math and coding benchmarks and reaches 91.9% on AIME'25 via Markovian RSA test-time compute.

  2. xGR: Efficient Generative Recommendation Serving at Scale

    cs.LG 2025-12 conditional novelty 6.0 of 10

    On real-world recommendation datasets, xGR sustains about 2.9–3.5× the throughput of vLLM/xLLM under a 200 ms P99 latency cap through GR-specific KV-cache, beam-search, and scheduling optimizations.

  3. Pixel-Resolved Long-Context Learning for Turbulence at Exascale: Resolving Small-scale Eddies Toward the Viscous Limit

    physics.flu-dyn 2025-07 conditional novelty 6.0 of 10

    A multiscale transformer with a new collective-based parallel attention method is claimed to be the first deep-learning model to reproduce small-scale turbulence statistics down to the viscous limit in 3D flow.

  4. ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution

    cs.LG 2026-07 conditional novelty 5.0 of 10

    ZUNA1.1, an open-source 380M EEG diffusion autoencoder, reconstructs variable-length, flexibly masked EEG at least as well as its predecessor and far better than spherical spline interpolation.

  5. ZONOS2 Technical Report

    cs.SD 2026-06 unverdicted novelty 4.0 of 10

    ZONOS2 8B is a scaled MoE TTS model with 900M active parameters trained on 6M hours of data that reports competitive SOTA results on naturalness, speaker similarity, WER, and a new ZTTS1-Eval benchmark while releasing...

  6. ZONOS2 Technical Report

    cs.SD 2026-06 unverdicted novelty 3.0 of 10

    ZONOS2 8B scales a prior TTS system to 8B parameters with MoE architecture and 6M hours of data, reporting competitive benchmark performance on naturalness and speaker similarity while releasing weights.

Pith tools