Pith. sign in

REVIEW 3 cited by

Systems and Algorithms for Convolutional Multi-Hybrid Language Models at Scale

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.01868 v1 pith:TOUIXTXQ submitted 2025-02-25 cs.LG cs.AIcs.CLcs.DC

classification cs.LGcs.AIcs.CLcs.DC
keywords modelsmulti-hybridoperatorsalgorithmsarchitecturearchitecturesattentionconvolutional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We introduce convolutional multi-hybrid architectures, with a design grounded on two simple observations. First, operators in hybrid models can be tailored to token manipulation tasks such as in-context recall, multi-token recall, and compression, with input-dependent convolutions and attention offering complementary performance. Second, co-designing convolution operators and hardware-aware algorithms enables efficiency gains in regimes where previous alternative architectures struggle to surpass Transformers. At the 40 billion parameter scale, we train end-to-end 1.2 to 2.9 times faster than optimized Transformers, and 1.1 to 1.4 times faster than previous generation hybrids. On H100 GPUs and model width 4096, individual operators in the proposed multi-hybrid StripedHyena 2 architecture achieve two-fold throughput improvement over linear attention and state-space models. Multi-hybrids excel at sequence modeling over byte-tokenized data, as demonstrated by the Evo 2 line of models. We discuss the foundations that enable these results, including architecture design, overlap-add blocked kernels for tensor cores, and dedicated all-to-all and point-to-point context parallelism strategies.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evaluating DNA function understanding in genomic language models using evolutionarily implausible sequences

    q-bio.QM 2025-06 conditional novelty 7.0 of 10

    A new benchmark shows that genomic language models mostly fail to detect loss-of-function mutations in synthetic, evolutionarily implausible DNA, with accuracy tied to how likely the model finds the sequence.

  2. Native Multi-Dimensional Subquadratic Operators via Input Dependent Long Convolutions

    cs.LG 2026-07 conditional novelty 6.0 of 10

    HyenaND is an input-dependent, multi-dimensional convolution that runs in near-linear time and matches attention baselines on vision, medical, PDE, and genomics benchmarks.

  3. Exploring Diffusion Transformer Designs via Grafting

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Grafting uses activation distillation and lightweight fine-tuning to edit pretrained diffusion transformers into hybrid architectures with near-baseline quality at under 2% pretraining compute.

Pith tools