Pith. sign in

Rat: Bridging rnn efficiency and attention accuracy via chunk-based sequence modeling.arXiv preprint arXiv:2507.04416

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

fields

cs.CL 1 cs.LG 1

years

2026 2

representative citing papers

Dynamic Linear Attention

cs.CL · 2026-06-09 · unverdicted · novelty 5.0

DLA introduces adaptive state merging based on token information variation plus fixed-size memory management for linear attention, reporting better results than fixed-policy baselines on 16 datasets across three categories.

citing papers explorer

Showing 2 of 2 citing papers.

  • RAT+: Train Dense, Infer Sparse -- Recurrence Augmented Attention for Dilated Inference cs.LG · 2026-02-20 · conditional · none · ref 29 · 2 links

    A single dense-pretrained RAT+ model can be adapted with 1B tokens to run at dilation sizes D=2..128, matching dense accuracy at D=16 and losing only 1-3 points at D=64 on reasoning and long-context benchmarks, though retrieval tasks drop far more.

  • Dynamic Linear Attention cs.CL · 2026-06-09 · unverdicted · none · ref 7

    DLA introduces adaptive state merging based on token information variation plus fixed-size memory management for linear attention, reporting better results than fixed-policy baselines on 16 datasets across three categories.