Pith. sign in

REVIEW 2 cited by

MambaPEFT: Exploring Parameter-Efficient Fine-Tuning for Mamba

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.03855 v3 pith:Z6IS6F65 submitted 2024-11-06 cs.CL cs.AIcs.CVcs.LG

classification cs.CLcs.AIcs.CVcs.LG
keywords mambapeftmethodsmodelstransformersbeendownstreameffectively
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

An ecosystem of Transformer-based models has been established by building large models with extensive data. Parameter-efficient fine-tuning (PEFT) is a crucial technology for deploying these models to downstream tasks with minimal cost while achieving effective performance. Recently, Mamba, a State Space Model (SSM)-based model, has attracted attention as a potential alternative to Transformers. While many large-scale Mamba-based models have been proposed, efficiently adapting pre-trained Mamba-based models to downstream tasks remains unexplored. In this paper, we conduct an exploratory analysis of PEFT methods for Mamba. We investigate the effectiveness of existing PEFT methods for Transformers when applied to Mamba. We also modify these methods to better align with the Mamba architecture. Additionally, we propose new Mamba-specific PEFT methods that leverage the distinctive structure of Mamba. Our experiments indicate that PEFT performs more effectively for Mamba than Transformers. Lastly, we demonstrate how to effectively combine multiple PEFT methods and provide a framework that outperforms previous works. To ensure reproducibility, we will release the code after publication.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Energy- and Memory-Efficient PEFT Methods for Personalized On-Device SLMs on Consumer GPUs

    cs.CL 2026-08 conditional novelty 5.0 of 10

    On consumer GPUs, LoRA+ gives the best energy-focused fine-tuning score in 19 of 24 small-model task configurations, while QLoRA wins the memory-focused score when peak VRAM is the binding constraint.

  2. JCAPT: A Joint Modeling Approach for CAPT

    cs.CL 2025-06 conditional novelty 4.0 of 10

    JCAPT, a Mamba-based joint APA and MDD model with phonological features and think tokens, improves mispronunciation detection and several scoring aspects on speechocean762 over JAM.

Pith tools