Pith. sign in

REVIEW 2 cited by

Model Compression Method for S4 with Diagonal State Space Layers using Balanced Truncation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.15993 v3 pith:O3UX3WIL submitted 2024-02-25 cs.LG

classification cs.LG
keywords modelmodelscompressionbalancedlayerstruncationaccuracymethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

To implement deep learning models on edge devices, model compression methods have been widely recognized as useful. However, it remains unclear which model compression methods are effective for Structured State Space Sequence (S4) models incorporating Diagonal State Space (DSS) layers, tailored for processing long-sequence data. In this paper, we propose to use the balanced truncation, a prevalent model reduction technique in control theory, applied specifically to DSS layers in pre-trained S4 model as a novel model compression method. Moreover, we propose using the reduced model parameters obtained by the balanced truncation as initial parameters of S4 models with DSS layers during the main training process. Numerical experiments demonstrate that our trained models combined with the balanced truncation surpass conventionally trained models with Skew-HiPPO initialization in accuracy, even with fewer parameters. Furthermore, our observations reveal a positive correlation: higher accuracy in the original model consistently leads to increased accuracy in models trained using our model compression method, suggesting that our approach effectively leverages the strengths of the original model.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. An Exact Instrument for State Usage in Selective State-Space Models, and the Input-Driven Migration It Reveals

    cs.LG 2026-07 accept novelty 7.0 of 10

    An exact per-mode Gram instrument shows trained Mamba models migrate state usage with input via Bt, and two-pass scheduled pruning matches unpruned quality at half state budget.

  2. A Survey of Mamba

    cs.LG 2024-08 unverdicted novelty 2.0 of 10

    The paper consolidates existing research on Mamba models, their architecture variants, adaptations to different data modalities, and applications across domains.

Pith tools