REVIEW 23 cited by
SiMBA: Simplified Mamba-Based Architecture for Vision and Multivariate Time series
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Transformers have widely adopted attention networks for sequence mixing and MLPs for channel mixing, playing a pivotal role in achieving breakthroughs across domains. However, recent literature highlights issues with attention networks, including low inductive bias and quadratic complexity concerning input sequence length. State Space Models (SSMs) like S4 and others (Hippo, Global Convolutions, liquid S4, LRU, Mega, and Mamba), have emerged to address the above issues to help handle longer sequence lengths. Mamba, while being the state-of-the-art SSM, has a stability issue when scaled to large networks for computer vision datasets. We propose SiMBA, a new architecture that introduces Einstein FFT (EinFFT) for channel modeling by specific eigenvalue computations and uses the Mamba block for sequence modeling. Extensive performance studies across image and time-series benchmarks demonstrate that SiMBA outperforms existing SSMs, bridging the performance gap with state-of-the-art transformers. Notably, SiMBA establishes itself as the new state-of-the-art SSM on ImageNet and transfer learning benchmarks such as Stanford Car and Flower as well as task learning benchmarks as well as seven time series benchmark datasets. The project page is available on this website ~\url{https://github.com/badripatro/Simba}.
Forward citations
Cited by 23 Pith papers
-
Let SSMs be ConvNets: State-space Modeling with Optimal Tensor Contractions
Treating state-space layers as tensor networks with CNN-style connectivity and optimized contraction orders yields hybrid SSM networks that outperform homogeneous SSMs on raw audio tasks and enable competitive streami...
-
Black-Mamba: Biologically-Inspired Leaky Accumulation for Conceptual Knowledge under Distribution Drift
Gating test-time memory writes on leaky accumulated surprisal preserves most of the adaptation benefit while roughly halving the number of updates.
-
DCVC-MB: Neural B-Frame Video Compression using State Space Models
DCVC-MB, a neural B-frame video codec using Mamba state-space fusion, reports BD-rate savings up to 8.98% over prior neural codecs and up to 30.45% over VTM-19.0-LDP.
-
Partial Ring Scan: Revisiting Scan Order in Vision State Space Models
Ring-based scanning with selective channel routing improves accuracy, speed, and rotation robustness of vision state-space models.
-
CLIMP: Contrastive Language-Image Mamba Pretraining
A fully Mamba-based (VMamba + Mamba LLM) CLIP model matches or beats transformer baselines on retrieval and OOD benchmarks, and natively supports high resolutions and dense captions.
-
Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
A dedicated accelerator for Vision Mamba using a Kogge-Stone systolic scan array and hybrid 8-bit quantization achieves 2.3x end-to-end speedup and 11.5x energy-efficiency gain over an edge GPU with less than 1% top-1...
-
Training-free Token Reduction for Vision Mamba
MTR uses Mamba's timescale parameter Δ as a token importance score to merge unimportant tokens, giving training-free inference speedups with small accuracy loss.
-
SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot
SparseSSM extends OBS-style second-order pruning to Mamba's discretized, time-shared state-transition matrix, pruning 50% of its weights in one pass without fine-tuning.
-
Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking
STTrack, a video tracker with a temporal state generator and mamba fusion modules, reports state-of-the-art success and accuracy numbers on five multimodal tracking benchmarks.
-
LV-CadeNet: A Long-View Feature Convolution-Attention Fusion Encoder-Decoder Network for EEG/MEG Spike Analysis
A conv-attention encoder-decoder with hand-crafted long-view morphology features reports state-of-the-art EEG and MEG spike classification, including a 13.58-point balanced-accuracy gain on a clinical MEG set.
-
FLDmamba: Integrating Fourier and Laplace Transform Decomposition with Mamba for Enhanced Time Series Prediction
FLDmamba combines a learnable Fourier filter on Mamba's step size with a damped-sinusoid output layer and reports superior long-term forecasting accuracy on standard benchmarks.
-
Exploring State-Space-Model based Language Model in Music Generation
SiMBA, a Mamba-based decoder, converges faster and matches or slightly trails a Transformer baseline in a single-codebook text-to-music system under limited compute.
-
Burst Image Super-Resolution via Multi-Cross Attention Encoding and Multi-Scan State-Space Decoding
A Transformer-Mamba hybrid with multi-cross attention and multi-scan state-space fusion reports state-of-the-art PSNR and SSIM on synthetic burst super-resolution benchmarks.
-
BiCrossMamba-ST: Speech Deepfake Detection with Bidirectional Mamba Spectro-Temporal Cross-Attention
BiCrossMamba-ST, a dual-branch bidirectional-Mamba spectro-temporal model with cross-attention, reports substantially lower error rates than AASIST and RawBMamba on ASVspoof 2021 LA and DF benchmarks.
-
Vision Mamba in Remote Sensing: A Comprehensive Survey of Techniques, Applications and Outlook
A structured review of roughly 120 Mamba-based remote sensing papers that proposes taxonomies for scan strategies and architectural integrations, and claims Mamba-based models often outperform CNN and Transformer base...
-
SSD-Poser: Avatar Pose Estimation with State Space Duality from Sparse Observations
SSD-Poser reconstructs full-body SMPL poses from three sparse HMD signals with a hybrid state-space/attention encoder and frequency-aware decoder, reporting state-of-the-art accuracy at real-time speed on AMASS.
-
Directing Mamba to Complex Textures: An Efficient Texture-Aware State Space Model for Image Restoration
TAMambaIR introduces a texture-aware state space model that prioritizes high-texture patches, achieving modest gains on restoration benchmarks with lower FLOPs than comparable Mamba models.
-
MV-GMN: State Space Model for Multi-View Action Recognition
MV-GMN, a state-space model with graph convolution, reports state-of-the-art accuracies on NTU RGB+D and PKU-MMD action recognition benchmarks.
-
Sequence Complementor: Complementing Transformers For Time Series Forecasting with Learnable Sequences
Adding three learnable 'complementor' sequences to each input channel of a PatchTST-style transformer lowers forecasting error by 1-5 percent on standard benchmarks, though the stated information-theoretic guarantee d...
-
Faster Vision Mamba is Rebuilt in Minutes via Merged Token Re-training
Merged-token retraining recovers Vision Mamba accuracy after token reduction, within minutes and with up to 1.5x faster inference.
-
MTS-UNMixers: Multivariate Time Series Forecasting via Channel-Time Dual Unmixing
MTS-UNMixers forecasts multivariate time series by decomposing data into shared time and channel components with Mamba networks, and reports improved benchmark accuracy.
-
$\text{S}^{3}$Mamba: Arbitrary-Scale Super-Resolution via Scaleable State Space Model
S3Mamba applies scale-modulated state space models to arbitrary-scale super-resolution, reporting marginal PSNR gains over prior INR-based methods.
-
Straightforward Bayesian A/B testing with Dirichlet posteriors
The submission is internally inconsistent: the abstract promises a Bayesian A/B testing method, but the full text is a different computer vision paper, leaving the claimed result unevaluable.
Discussion (0). Continue with ORCID to comment.