Pith. sign in

REVIEW 2 cited by

ReMamba: Equip Mamba with Effective Long-Sequence Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.15496 v4 pith:XEHXHXHA submitted 2024-08-28 cs.CL

ReMamba: Equip Mamba with Effective Long-Sequence Modeling

classification cs.CL
keywords mambaremambamodelscomprehendcontextsefficiencyinferencelong
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

While the Mamba architecture demonstrates superior inference efficiency and competitive performance on short-context natural language processing (NLP) tasks, empirical evidence suggests its capacity to comprehend long contexts is limited compared to transformer-based models. In this study, we investigate the long-context efficiency issues of the Mamba models and propose ReMamba, which enhances Mamba's ability to comprehend long contexts. ReMamba incorporates selective compression and adaptation techniques within a two-stage re-forward process, incurring minimal additional inference costs overhead. Experimental results on the LongBench and L-Eval benchmarks demonstrate ReMamba's efficacy, improving over the baselines by 3.2 and 1.6 points, respectively, and attaining performance almost on par with same-size transformer models.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Safety, Security, and Cognitive Risks in State-Space Models: A Systematic Threat Analysis with Spectral, Stateful, and Capacity Attacks

    cs.CR 2026-04 unverdicted novelty 6.0

    State-space models are vulnerable to three new attack types that corrupt state integrity, with experiments showing up to 156x output changes and 6x higher targeted corruption than random inputs.

  2. User-Centric Modeling of Transactional Sequences with Explainable State Space Models

    cs.LG 2026-07 conditional novelty 4.0

    Injecting a pretrained CoLES user embedding into Mamba as an initial hidden state or prefix token improves accuracy by up to 3.2 pp on three transaction benchmarks and speeds convergence 2–3x.