Pith. sign in

REVIEW 5 cited by

S7: Selective and Simplified State Space Layers for Sequence Modeling

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.03464 v1 pith:52MAS37P submitted 2024-10-04 cs.LG eess.SPmath.DS

classification cs.LGeess.SPmath.DS
keywords modelingsequenceinputstateacrossbenchmarkshandlereparameterization
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

A central challenge in sequence modeling is efficiently handling tasks with extended contexts. While recent state-space models (SSMs) have made significant progress in this area, they often lack input-dependent filtering or require substantial increases in model complexity to handle input variability. We address this gap by introducing S7, a simplified yet powerful SSM that can handle input dependence while incorporating stable reparameterization and specific design choices to dynamically adjust state transitions based on input content, maintaining efficiency and performance. We prove that this reparameterization ensures stability in long-sequence modeling by keeping state transitions well-behaved over time. Additionally, it controls the gradient norm, enabling efficient training and preventing issues like exploding or vanishing gradients. S7 significantly outperforms baselines across various sequence modeling tasks, including neuromorphic event-based datasets, Long Range Arena benchmarks, and various physical and biological time series. Overall, S7 offers a more straightforward approach to sequence modeling without relying on complex, domain-specific inductive biases, achieving significant improvements across key benchmarks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Low-latency Event-based Object Detection with Spatially-Sparse Linear Attention

    cs.CV 2026-03 conditional novelty 7.0 of 10

    SSLA-Det is the first fully asynchronous linear-attention detector for event cameras, reaching 0.375 mAP on Gen1 and 0.515 mAP on N-Caltech101 with over 20x lower per-event computation than prior async baselines.

  2. Three factor delay learning rules for spiking neural networks

    cs.NE 2026-01 conditional novelty 6.0 of 10

    A three-factor eligibility-trace rule lets LIF spiking networks learn synaptic and axonal delays online, matching offline backpropagation accuracy on SHD while cutting model size.

  3. Cortical-SSM: A Deep State Space Model for Motor Imagery Decoding from EEG Signals

    cs.CV 2025-10 conditional novelty 6.0 of 10

    Cortical-SSM, a dual state-space architecture with wavelet-based frequency features, reports state-of-the-art motor-imagery decoding accuracy on OpenBMI, Stieger2021, and a clinical ECoG-ALS dataset.

  4. Quantizing Small-Scale State-Space Models for Edge AI

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Quantization-aware training with a frozen state matrix lifts sequential MNIST accuracy from 40% under post-training quantization to 96%, and a heterogeneous precision scheme cuts memory by 6 times.

  5. ControlMambaIR: Conditional Controls with State-Space Model for Image Restoration

    cs.CV 2025-06 reject novelty 4.0 of 10

    A diffusion image restoration model with a Mamba condition network reports low LPIPS/FID on several benchmarks, but the PSNR losses and internal inconsistencies undermine the stated performance claims.

Pith tools