Pith. sign in

REVIEW 2 cited by

A Diagonal Structured State Space Model on Loihi 2 for Efficient Streaming Sequence Processing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.15022 v1 pith:3G4FTCN6 submitted 2024-09-23 cs.LG cs.AIcs.ETcs.NE

classification cs.LGcs.AIcs.ETcs.NE
keywords jetsonprocessingrecurrentefficientimplementationloihitimestoken-by-token
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep State-Space Models (SSM) demonstrate state-of-the art performance on long-range sequence modeling tasks. While the recurrent structure of SSMs can be efficiently implemented as a convolution or as a parallel scan during training, recurrent token-by-token processing cannot currently be implemented efficiently on GPUs. Here, we demonstrate efficient token-by-token inference of the SSM S4D on Intel's Loihi 2 state-of-the-art neuromorphic processor. We compare this first ever neuromorphic-hardware implementation of an SSM on sMNIST, psMNIST, and sCIFAR to a recurrent and a convolutional implementation of S4D on Jetson Orin Nano (Jetson). While we find Jetson to perform better in an offline sample-by-sample based batched processing mode, Loihi 2 outperforms during token-by-token based processing, where it consumes 1000 times less energy with a 75 times lower latency and a 75 times higher throughput compared to the recurrent implementation of S4D on Jetson. This opens up new avenues towards efficient real-time streaming applications of SSMs.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. QS4D: Quantization-aware training for efficient hardware deployment of structured state-space sequential models

    cs.LG 2025-07 conditional novelty 4.0 of 10

    Quantization-aware training allows S4D sequence models to run at much lower precision, cutting estimated hardware costs by up to two orders of magnitude while keeping accuracy.

  2. Quantizing Small-Scale State-Space Models for Edge AI

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Quantization-aware training with a frozen state matrix lifts sequential MNIST accuracy from 40% under post-training quantization to 96%, and a heterogeneous precision scheme cuts memory by 6 times.

Pith tools