REVIEW 3 major objections 5 minor 76 references
Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper claims a purpose-built systolic scan array can make Vision Mamba's selective-scan bottleneck 11.6x faster on edge hardware, yielding a 2.3x end-to-end speedup.
desk verdict A coherent edge-vision-Mamba accelerator with a genuinely new scan-array idea, but the headline speedups rest on an unvalidated simulator. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the Systolic Scan Array (SSA), a grid of Scan Processing Elements (SPEs), each with two multipliers and an adder, that implements the Kogge-Stone recurrence steps $P_n P_{n+1}$ and $P_{n+1} Q_n + Q_{n+1}$ while passing partial results only to neighbouring SPEs. This local, rhythmic dataflow eliminates explicit shared-memory writes of intermediate states and runs different state dimensions as independent parallel rows. The Long Input Support Unit (LISU) stitches partial states across sequence chunks, and the hybrid quantization scheme makes the required rescaling a shift rather than a multiplication; a 16-32 entry LUT-based SFU handles SiLU, exponential, and softplus.
What would settle it
Build the Mamba-X RTL described in Section 5 at the stated 65nm target, run the Tiny/Small/Base Vision Mamba models on an FPGA or test chip at 224 to 1024 resolutions, and measure end-to-end latency, energy, and ImageNet-1K top-1 accuracy; if selective-scan throughput is materially below 11.6x or accuracy loss exceeds about 1 percentage point, the central claim fails.
Extended reading notes
Core claim
Mamba-X's central claim is that a dedicated accelerator can make Vision Mamba practical on edge devices by targeting the selective SSM block. On a conventional edge GPU the paper finds the selective SSM consumes up to 60% of encoder latency for high-resolution inputs, because fused scan kernels limit parallelism along the state dimension and because small on-chip SRAM forces intermediate state vectors to spill off-chip. Mamba-X counters both problems with a systolic scan array that scans chunks of the sequence across state dimensions in parallel, a Long Input Support Unit that threads partial state between chunks, and an INT8 hybrid quantization scheme that rounds scaling factors to powers of two so rescaling becomes shift operations. The result, as the authors state it, is that scan throughput rises 11.6x, end-to-end latency falls 2.3x, energy efficiency rises 11.5x, and performance per area rises 601x while accuracy loss stays under 1%p.
Load-bearing premise
The load-bearing premise is in Section 5: Mamba-X is modeled as a cycle-level C++ simulator, and only area comes from synthesized RTL; if the simulator's cycle counts, power, or memory traffic are optimistic, the 11.6x, 11.5x, and 601x results would not transfer to real hardware.
Editorial extensions
If this is right
- If the reported figures hold, the selective scan stops being a serial bottleneck, so Vision Mamba can be deployed at higher resolutions on edge hardware without retraining the model.
- The hybrid INT8 quantization with power-of-two rescaling means the scan-path hardware needs only integer adders, multipliers, and shifters, reducing per-operation energy and on-chip buffer pressure.
- A $1.34\,\mathrm{mm}^2$ footprint scaled to 12nm (0.4% of the baseline GPU die) suggests the SSA and GEMM engine could be reused as a small IP block inside a larger edge SoC.
- The speedup is largest where the selective SSM dominates, so the reported gains should grow with input resolution from 224 to 1024, consistent with the paper's breakdown.
Reading between the lines
- The same systolic scan dataflow should extend to Mamba-style language models, whose selective recurrence is identical; that is an extrapolation the paper does not make.
- The power-of-two rounding of activation scaling factors is a robustness bet: a deployment distribution with different outliers than the 500 calibration images could push the accuracy loss above the reported margin, a test the paper does not perform.
- The 601x performance-per-area figure assumes the simulator's throughput and the RTL-synthesized area combine without DRAM scheduling, clocking, or integration overheads, so real-system numbers may be lower.
- The characterization suggests the accelerator's advantage narrows as model size grows and GEMMs dominate, pointing to a future co-design where the GEMM engine also receives attention.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes Mamba-X, an end-to-end accelerator for Vision Mamba inference on edge devices. The architecture combines a systolic scan array (SSA) for the selective scan operation, a long-input support unit (LISU) for inter-chunk dependencies, a LUT-based special function unit (SFU), and a hybrid channel/tensor-granularity INT8 quantization scheme with power-of-two scaling-factor approximation. The authors characterize Vision Mamba on an NVIDIA Jetson AGX Xavier, identify the selective SSM as the main bottleneck, and report that Mamba-X achieves an average 11.6x improvement in selective scan throughput, an average 11.5x end-to-end energy-efficiency improvement, a 601x increase in average performance/area, and under 1 percentage point top-1 accuracy loss on ImageNet-1K.
Significance. If the reported results hold, Mamba-X would be a useful contribution to edge inference for state-space vision models: the characterization of the GPU bottleneck is concrete, the SSA/LISU design is a plausible architectural response to the sequential scan dependency, and the quantization-aware SPE design is well motivated. The area estimate is grounded in RTL synthesis, which is a strength relative to purely analytical proposals. However, the central performance, energy, and area-efficiency numbers are produced by a cycle-level simulator that is not validated against RTL simulation or hardware, and the accuracy evaluation uses a calibration set drawn from the same ImageNet-1K test set on which accuracy is reported. These issues are load-bearing for the paper's headline claims.
major comments (3)
- [Section 5 (Performance) and Section 6.1] All speedup, energy-efficiency, and performance/area claims (11.6x, 11.5x, and 601x in Sections 1 and 6) rest on a cycle-level C++ simulator that is not validated against the synthesized RTL or any hardware prototype. Section 5 states that only area is estimated from RTL synthesis (Table 4), while energy is computed by multiplying RTL-synthesized power by the simulator's inference time. The manuscript does not specify how SSA or LISU cycle counts are derived, how DRAM refresh or bank conflicts are modeled, or whether the Jetson baseline is modeled at achieved throughput. Because any systematic timing optimism propagates directly into the energy-efficiency and performance/area results, please validate the simulator against RTL or measured hardware for at least one configuration, and provide a sensitivity analysis of the headline results to the main timing assumptions.
- [Section 4.4 and Section 6.3, Table 5] The accuracy evaluation is circular. Section 4.4 states that quantization scaling factors are calibrated using 500 randomly sampled images from the 50,000-image ImageNet-1K test set, and Section 6.3 reports final top-1/top-5 accuracy on that same dataset. Since the calibration images are part of the evaluation set, the reported 0.59-0.89 percentage-point accuracy losses in Table 5 are likely optimistic. Please calibrate on a disjoint set (for example, a held-out portion of the training split) and report accuracy excluding any calibration images, or quantify the degree to which the 500-image calibration set inflates the reported accuracies.
- [Tables 4-5 and Figures 17-18] The accuracy claim of 'under 1 percentage point degradation' is established only at 224x224 resolution, whereas the performance, energy, and area-efficiency claims are reported across 224, 512, 738, and 1024 resolutions. Quantization ranges, SFU breakpoint selection, and LISU inter-chunk behavior all interact with sequence length, so the accuracy result does not automatically transfer to the higher resolutions featured in the performance evaluation. Please measure accuracy at higher resolutions or explicitly qualify the accuracy claim to 224x224.
minor comments (5)
- [Table 1] Table 1 does not state which model size is used; the baseline accuracy of 76.04% matches the Tiny model in Table 5, so please label the table accordingly.
- [Section 6.2] The 601x performance/area improvement is reported without a precise definition; please state whether it is the ratio of normalized inverse latencies divided by area, and describe how the technology scaling from 65 nm to 12 nm is applied to the comparison with the 12 nm Jetson AGX Xavier.
- [Section 5 (Energy)] The energy model uses a fixed 4 pJ/bit for LPDDR4 off-chip transfers; please clarify whether the same value is applied to both the baseline GPU and Mamba-X, and whether DRAM refresh overhead is included in either model.
- [Equation (1)] The rounding notation in Equation (1) is not defined; please state explicitly that X_q is obtained by round-to-nearest with ties handled in a particular way, as the quantization results are sensitive to this choice.
- [Figure 8] The term 'oracular ideal GPU design' is informal; consider replacing it with a more precise description such as 'an infinite-capacity on-chip storage baseline'.
Circularity Check
Accuracy claims are partially circular because quantization scale factors are calibrated on a subset of the ImageNet-1K test set and then evaluated on that same test set; hardware speedup/energy claims are simulator-based but not circular.
-
fitted input called prediction
[Section 4.4 (Hybrid quantization) and Section 5 (Model accuracy); Table 5]
"We observe that using only 1% of the test dataset (500 randomly sampled images from the 50,000-image ImageNet-1K dataset) provides a robust estimation of global maximum and minimum values for configuring our scaling factors. ... When measuring model accuracy, we use the ImageNet-1K [8] dataset, which contains 50,000 images at a resolution of 224 x 224."
The quantization scaling factors s from Equation (1) are fitted to 500 images drawn from the ImageNet-1K test set, and Table 5 reports top-1/top-5 accuracy on the full ImageNet-1K test set, which includes those same 500 images. The reported less-than-1%p accuracy loss is therefore not an independent prediction on unseen data; the quantization parameters were tuned to the evaluation set, so the accuracy result is partially forced by the calibration. This is a fitted-input-called-prediction pattern: the fitted scale factors are reused in the accuracy evaluation on the same test instances.
full rationale
The central hardware claims (11.6x selective-scan throughput, 2.3x end-to-end speedup, 11.5x energy-efficiency, 601x performance/area) are produced by a cycle-level simulator and RTL-based area/power estimation; they do not reduce to fitted parameters or to self-citation. The main circularity is confined to the accuracy evaluation: the paper explicitly calibrates scale factors using 1% of the ImageNet-1K test set and then evaluates accuracy on that same test set. This is a real methodological leakage that partially invalidates the reported accuracy as an independent prediction, but it does not by construction force the hardware efficiency numbers. Because the central hardware derivation is independent and the accuracy circularity affects only one component of the paper's claims, a moderate score of 4 is appropriate rather than a higher score reserved for fully circular derivations.
Assumptions & free parameters
free parameters (6)
- Number of SSA instances =
8 (with 1, 2, 4 also shown)
- Chunk size =
16
- GEMM engine size =
64x64 PEs
- SFU LUT sizes and breakpoints =
16 entries for exponential, 32 for SiLU and softplus
- Activation and weight scaling factors for INT8 quantization =
Derived from 500 ImageNet calibration images
- SPE fixed-point intermediate precision =
2 extra fractional bits
assumptions (6)
- standard math Kogge-Stone parallel prefix algorithm is a correct inclusive-scan formulation for the selective SSM recurrence.
- domain assumption Zero-order hold discretization of the continuous-time SSM is the correct model for Vision Mamba.
- ad hoc to paper The cycle-level C++ simulator faithfully models Mamba-X and the edge GPU baseline.
- domain assumption Scaling 65nm synthesized area and energy to 12nm using the equations of Stillmaker and Baas is accurate for this design.
- ad hoc to paper The 500-image calibration set sampled from the ImageNet-1K test set is representative and does not inflate the reported accuracies.
- domain assumption The CUB-based fused selective SSM kernel is a representative state-of-the-art implementation on the edge GPU baseline.
invented entities (3)
-
Systolic Scan Array (SSA)
-
Scan Processing Element (SPE)
-
Long Input Support Unit (LISU)
Cite this review
Pith. "Pith review of Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices." pith.science (2026). https://pith.science/paper/PG3IHXLP
@misc{pith2026250802977,
author = {Pith},
title = {Pith review of: Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices},
year = {2026},
howpublished = {\url{https://pith.science/paper/PG3IHXLP}},
note = {Machine review of arXiv:2508.02977}
}
abstract
Transformers have proven effective in language modeling but are limited by high computational and memory demands that grow quadratically with input sequence length. State space models (SSMs) offer a promising alternative by reducing attention complexity from $O(L^2)$ to $O(L)$ while also lowering overall memory consumption. Vision Mamba adapts the SSM approach for computer vision tasks, achieving lower latency and memory consumption than traditional transformer models. However, deploying Vision Mamba on edge devices is challenging due to its sequential scan operations, which hinder GPU efficiency. We propose Mamba-X, an end-to-end Vision Mamba accelerator that includes a systolic scan array to maximize parallelism and minimize memory traffic, along with a hybrid, hardware-friendly quantization technique to reduce memory usage and improve hardware efficiency without sacrificing accuracy.
Figures
Figures from the paper (14 more)
Reference graph
Works this paper leans on
-
[1]
Arian Bakhtiarnia, Qi Zhang, and Alexandros Iosifidis. 2024. Efficient High-Resolution Deep Learning: A Survey.Comput. Surveys(2024)
work page 2024
-
[2]
Rajeev Balasubramonian, Andrew B Kahng, Naveen Muralimanohar, Ali Shafiee, and Vaishnav Srinivas. 2017. CACTI 7: New Tools for Interconnect Exploration in Innovative Off-Chip Memories.ACM Transactions on Architecture and Code Optimization (TACO)(2017)
work page 2017
-
[3]
Ali Behrouz and Farnoosh Hashemi. 2024. Graph Mamba: Towards Learning on Graphs with State Space Models. InProceedings of the ACM SIGKDD Conference on Knowledge Discovery and Data Mining
work page 2024
-
[4]
Nicolas Carion, Francisco Massa, Gabriel Synnaeve, Nicolas Usunier, Alexander Kirillov, and Sergey Zagoruyko. 2020. End-to-End Object Detection with Transformers. InProceedings of the European Confer- ence on Computer Vision (ECCV)
work page 2020
-
[5]
Chun-Fu (Richard) Chen, Quanfu Fan, and Rameswar Panda. 2021. CrossViT: Cross-Attention Multi-Scale Vision Transformer for Image Classification. InProceedings of the European Conference on Computer Vision (ECCV)
work page 2021
-
[6]
Yu-Hsin Chen, Joel Emer, and Vivienne Sze. 2016. Eyeriss: A Spatial Architecture for Energy-Efficient Dataflow for Convolutional Neural Networks. InProceedings of the International Symposium on Computer Architecture (ISCA)
work page 2016
-
[7]
Jyotikrishna Dass, Shang Wu, Huihong Shi, Chaojian Li, Zhifan Ye, Zhongfeng Wang, and Yingyan Lin. 2023. ViTALiTy: Unifying Low- rank and Sparse Approximation for Vision Transformer Acceleration with a Linear Taylor Attention. InProceedings of the International Symposium on High-Performance Computer Architecture (HPCA)
work page 2023
-
[8]
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei
Show all 76 references
-
[9]
Peiyan Dong, Mengshu Sun, Alec Lu, Yanyue Xie, Kenneth Liu, Zhenglun Kong, Xin Meng, Zhengang Li, Xue Lin, Zhenman Fang, and Yanzhi Wang. 2023. Heatvit: Hardware-efficient Adaptive Token Pruning for Vision Transformers. InProceedings of the International Symposium on High-Perf...
2023
-
[10]
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weis- senborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transform- ers for Image Reco...
2021
-
[11]
Yu Feng, Tianrui Ma, Yuhao Zhu, and Xuan Zhang. 2024. BlissCam: Boosting Eye Tracking Efficiency with Learned In-Sensor Sparse Sam- pling. InProceedings of the International Symposium on Computer Architecture (ISCA)
2024
-
[12]
Fu, Tri Dao, Khaled K
Daniel Y. Fu, Tri Dao, Khaled K. Saab, Armin W. Thomas, Atri Rudra, and Christopher Ré. 2023. Hungry Hungry Hippos: Towards Language Modeling with State Space Models. InarXiv preprint arXiv:2212.14052
2023 arXiv
-
[13]
Fung, Ivan Sham, George Yuan, and Tor M
Wilson W.L. Fung, Ivan Sham, George Yuan, and Tor M. Aamodt. 2007. Dynamic Warp Formation and Scheduling for Efficient GPU Control Flow. InProceedings of the International Symposium on Microarchitec- ture (MICRO)
2007
-
[14]
Albert Gu and Tri Dao. 2024. Mamba: Linear-Time Sequence Modeling with Selective State Spaces. InarXiv preprint arXiv:2312.00752
2024 arXiv
-
[15]
Albert Gu, Karan Goel, Ankit Gupta, and Christopher Ré. 2022. On the Parameterization and Initialization of Diagonal State Space Models. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS)
2022
-
[16]
Albert Gu, Karan Goel, and Christopher Ré. 2022. Efficiently Model- ing Long Sequences with Structured State Spaces. InarXiv preprint arXiv:2111.00396
2022 arXiv
-
[17]
Albert Gu, Isys Johnson, Karan Goel, Khaled Saab, Tri Dao, Atri Rudra, and Christopher Ré. 2021. Combining Recurrent, Convolu- tional, and Continuous-time Models with Linear State Space Layers. InProceedings of the Conference on Neural Information Processing Sys- tems (NeurIPS)
2021
-
[18]
Ramyad Hadidi, Jiashen Cao, Yilun Xie, Bahar Asgari, Tushar Krishna, and Hyesoon Kim. 2019. Characterizing the Deployment of Deep Neural Networks on Commercial Edge Devices. InProceedings of the International Symposium on Workload Characterization (IISWC)
2019
-
[19]
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. InProceedings of the Con- ference on Computer Vision and Pattern Recognition (CVPR)
2016
-
[20]
Mark Horowitz. 2014. 1.1 Computing’s Energy Problem (and What We Can Do About It). InProceedings of the International Solid State Circuits Conference (ISSCC)
2014
-
[21]
Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam
Andrew G. Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam
-
[22]
Norman P Jouppi, Doe Hyun Yoon, Matthew Ashcraft, Mark Gottscho, Thomas B Jablin, George Kurian, James Laudon, Sheng Li, Peter Ma, Xiaoyu Ma, Thomas Norrie, Nishant Patil, Sushma Prasad, Cliff Young, Zongwei Zhou, and David Patterson. 2021. Ten Lessons from Three Generations S...
2021
-
[23]
Norman P Jouppi, Cliff Young, Nishant Patil, David Patterson, Gau- rav Agrawal, Raminder Bajwa, Sarah Bates, Suresh Bhatia, Nan Bo- den, Al Borchers, Rick Boyle, Pierre-luc Cantin, Clifford Chao, Chris Clark, Jeremy Coriell, Mike Daley, Matt Dau, Jeffrey Dean, Ben Gelb, Tara V...
2017
-
[24]
Mahoney, and Kurt Keutzer
Sehoon Kim, Amir Gholami, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer. 2021. I-BERT: Integer-Only BERT Quantization. InPro- ceedings of the International Conference on Machine Learning (ICML)
2021
-
[25]
Nikita Kitaev, Lukasz Kaiser, and Anselm Levskaya. 2020. Reformer: The Efficient Transformer. InProceedings of the International Confer- ence on Learning Representations (ICLR)
2020
-
[26]
Kogge and Harold S
Peter M. Kogge and Harold S. Stone. 1973. A Parallel Algorithm for the Efficient Solution of a General Class of Recurrence Equations.IEEE Trans. Comput.100, 8 (1973), 786–793
1973
-
[27]
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton. 2012. Ima- geNet Classification with Deep Convolutional Neural Networks. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS)
2012
-
[28]
Kung. 1982. Why Systolic Architectures?Computer(1982)
1982
-
[29]
Hsiang Tsung Kung and Charles E Leiserson. 1979. Systolic Arrays (for VLSI). InSparse Matrix Proceedings
1979
-
[30]
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica
-
[31]
Junseo Lee, Seokwon Lee, Jungi Lee, Junyong Park, and Jaewoong Sim
-
[32]
Jungi Lee, Wonbeom Lee, and Jaewoong Sim. 2024. Tender: Accelerat- ing Large Language Models via Tensor Decomposition and Runtime Requantization. InProceedings of the International Symposium on Com- puter Architecture (ISCA)
2024
-
[33]
Seung Yul Lee, Hyunseung Lee, Jihoon Hong, SangLyul Cho, and Jae W. Lee. 2024. VGA: Hardware Accelerator for Scalable Long Sequence Model Inference. InProceedings of the International Symposium on Microarchitecture (MICRO)
2024
-
[34]
Jinhao Li, Shan Huang, Jiaming Xu, Jun Liu, Li Ding, Ningyi Xu, and Guohao Dai. 2024. MARCA: Mamba Accelerator with ReConfigurable Architecture. InProceedings of IEEE International Conference on Com- puter Aided Design (ICCAD)
2024
-
[35]
Lincan Li, Hanchen Wang, Wenjie Zhang, and Adelle Coster. 2024. STG-Mamba: Spatial-Temporal Graph Learning via Selective State Space Model. InarXiv preprint arXiv:2403.12418
2024 arXiv
-
[36]
Zhengang Li, Mengshu Sun, Alec Lu, Haoyu Ma, Geng Yuan, Yanyue Xie, Hao Tang, Yanyu Li, Miriam Leeser, Zhangyang Wang, Xue Lin, and Zhenman Fang. 2022. Auto-ViT-Acc: An FPGA-Aware Automatic Acceleration Framework for Vision Transformer with Mixed-Scheme Quantization. InProceed...
2022
-
[37]
Zhikai Li, Junrui Xiao, Lianwei Yang, and Qingyi Gu. 2023. Repq-vit: Scale Reparameterization for Post-Training Quantization of Vision Transformers. InProceedings of the International Conference on Com- puter Vision (ICCV)
2023
-
[38]
Dingkang Liang, Xin Zhou, Wei Xu, Xingkui Zhu, Zhikang Zou, Xiao- qing Ye, Xiao Tan, and Xiang Bai. 2024. PointMamba: A Simple State Space Model for Point Cloud Analysis. InProceedings of the Conference on Neural Information Processing Systems (NeurIPS)
2024
-
[39]
Feng Liang, Bichen Wu, Xiaoliang Dai, Kunpeng Li, Yinan Zhao, Hang Zhang, Peizhao Zhang, Peter Vajda, and Diana Marculescu. 2023. Open- Vocabulary Semantic Segmentation with Mask-Adapted CLIP. InPro- ceedings of the Conference on Computer Vision and Pattern Recognition (CVPR)
2023
-
[40]
Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. 2024. AWQ: Activation-Aware Weight Quantization for On- Device LLM Compression and Acceleration. InProceedings of Machine Learning and Systems (MLSys)
2024
-
[41]
Weikai Lin, Yu Feng, and Yuhao Zhu. 2025. MetaSapiens: Real-Time Neural Rendering with Efficiency-Aware Pruning and Accelerated Foveated Rendering. InProceedings of the International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS)
2025
-
[42]
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2023. Visual Instruction Tuning. InProceedings of the Conference on Neural Information Processing Systems (NeurIPS)
2023
-
[43]
Yue Liu, Yunjie Tian, Yuzhong Zhao, Hongtian Yu, Lingxi Xie, Yaowei Wang, Qixiang Ye, Jianbin Jiao, and Yunfan Liu. 2024. VMamba: Vi- sual State Space Model. InProceedings of the Conference on Neural Information Processing Systems (NeurIPS)
2024
-
[44]
Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. 2021. Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows. InProceedings of the International Conference on Computer Vision (ICCV)
2021
-
[45]
Zhenhua Liu, Yunhe Wang, Kai Han, Wei Zhang, Siwei Ma, and Wen Gao. 2021. Post-Training Quantization for Vision Transformer. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS)
2021
-
[46]
NVIDIA. 2018. Jetson AGX Xavier Series.https://www.nvidia.com/en- us/autonomous-machines/embedded-systems/jetson-agx-xavier/
2018
-
[47]
NVIDIA. 2022. NVIDIA CUB Library.https://nvidia.github.io/cccl/ cub/
2022
-
[48]
NVIDIA. 2025. NVIDIA Automatic Mixed Precision for Deep Learning. https://developer.nvidia.com/automatic-mixed-precision
2025
-
[49]
NVIDIA. 2025. NVIDIA Nsight Compute.https://developer.nvidia. com/nsight-compute
2025
-
[50]
Patro and Vijay S
Badri N. Patro and Vijay S. Agneeswaran. 2024. SiMBA: Simplified Mamba-Based Architecture for Vision and Multivariate Time Series. InarXiv preprint arXiv:2403.15360
2024 arXiv
-
[51]
Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean. 2023. Efficiently Scaling Transformer Inference. InProceedings of Machine Learning and Systems (MLSys)
2023
-
[52]
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. 2021. Learn- ing Transferable Visual Models From Natural Language Supervision. InProceeding...
2021
-
[53]
Enrico Reggiani, Renzo Andri, and Lukas Cavigelli. 2023. Flex-SFU: Accelerating DNN Activation Functions by Non-Uniform Piecewise Approximation. InDesign Automation Conference (DAC)
2023
-
[54]
Minsoo Rhu and Mattan Erez. 2013. The Dual-Path Execution Model for Efficient GPU Control Flow. InProceedings of the International Symposium on High-Performance Computer Architecture (HPCA)
2013
-
[55]
Sheng Shen, Zhen Dong, Jiayu Ye, Linjian Ma, Zhewei Yao, Amir Gho- lami, Michael W Mahoney, and Kurt Keutzer. 2020. Q-BERT: Hessian Based Ultra Low Precision Quantization of BERT. InProceedings of the AAAI Conference on Artificial Intelligence
2020
-
[56]
Kyomin Sohn. 2024. High-Bandwidth Memory and Processing-in- Memory in the Era of Generative AI. InProceedings of the International Solid State Circuits Conference (ISSCC)
2024
-
[57]
Aaron Stillmaker and Bevan Baas. 2017. Scaling Equations for the Accurate Prediction of CMOS Device Performance from 180nm to 7nm.Integration58 (2017), 74–81
2017
-
[58]
Mengshu Sun, Haoyu Ma, Guoliang Kang, Yifan Jiang, Tianlong Chen, Xiaolong Ma, Zhangyang Wang, and Yanzhi Wang. 2022. VAQF: Fully 13 Automatic Software-Hardware Co-Design Framework for Low-Bit Vision Transformer. InarXiv preprint arXiv:2201.06618
2022 arXiv
-
[59]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017. Attention is All You Need. InProceedings of the Conference on Neural Information Processing Systems (NeurIPS)
2017
-
[60]
Chloe Wang, Oleksii Tsepa, Jun Ma, and Bo Wang. 2024. Graph- Mamba: Towards Long-Range Graph Sequence Modeling with Selec- tive State Spaces. InarXiv preprint arXiv:2402.00789
2024 arXiv
-
[61]
Li, Madian Khabsa, Han Fang, and Hao Ma
Sinong Wang, Belinda Z. Li, Madian Khabsa, Han Fang, and Hao Ma
-
[62]
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han. 2023. SmoothQuant: Accurate and Efficient Post- Training Quantization for Large Language Models. InProceedings of the International Conference on Machine Learning (ICML)
2023
-
[63]
Mengwei Xu, Mengze Zhu, Yunxin Liu, Felix Xiaozhu Lin, and Xu- anzhe Liu. 2018. DeepCache: Principled Cache for Mobile Deep Vision. InProceedings of the Annual International Conference on Mobile Com- puting and Networking (MobiCom)
2018
-
[64]
Seungjae Yoo, Hangyeol Kim, and Joo-Young Kim. 2024. AdapTiV: Sign- Similarity Based Image-Adaptive Token Merging for Vision Trans- former Acceleration. InProceedings of the International Symposium on Microarchitecture (MICRO)
2024
-
[65]
Lee, and Minsoo Rhu
Dongho Yoon, Taehun Kim, Jae W. Lee, and Minsoo Rhu. 2024. A Quantitative Analysis of State Space Model-Based Large Language Model: Study of Hungry Hungry Hippos. InIEEE Computer Architec- ture Letters
2024
-
[66]
Haoran You, Zhanyi Sun, Huihong Shi, Zhongzhi Yu, Yang Zhao, Yon- gan Zhang, Chaojian Li, Baopu Li, and Yingyan Lin. 2023. ViTCoD: Vision Transformer Acceleration via Dedicated Algorithm and Accel- erator Co-Design. InProceedings of the International Symposium on High-Performa...
2023
-
[67]
Joonsang Yu, Junki Park, Seongmin Park, Minsoo Kim, Sihwa Lee, Dong Hyun Lee, and Jungwook Choi. 2022. NN-LUT: Neural Approx- imation of Non-linear Operations for Efficient Transformer Inference. InDesign Automation Conference (DAC)
2022
-
[68]
Zhihang Yuan, Chenhao Xue, Yiqi Chen, Qiang Wu, and Guangyu Sun. 2022. PTQ4ViT: Post-training Quantization for Vision Transform- ers with Twin Uniform Quantization. InProceedings of the European Conference on Computer Vision (ECCV)
2022
-
[69]
Yilong Zhao, Chien-Yu Lin, Kan Zhu, Zihao Ye, Lequn Chen, Size Zheng, Luis Ceze, Arvind Krishnamurthy, Tianqi Chen, and Baris Kasikci. 2024. Atom: Low-Bit Quantization for Efficient and Accurate LLM Serving. InProceedings of Machine Learning and Systems (MLSys)
2024
-
[70]
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang. 2024. Vision Mamba.https://github.com/ hustvl/Vim
2024
-
[71]
Lianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang, Wenyu Liu, and Xinggang Wang. 2024. Vision Mamba: Efficient Visual Represen- tation Learning with Bidirectional State Space Model. InProceedings of the International Conference on Machine Learning (ICML). 14
2024
-
[2009]
InPro- ceedings of the Conference on Computer Vision and Pattern Recognition (CVPR)
ImageNet: A Large-scale Hierarchical Image Database. InPro- ceedings of the Conference on Computer Vision and Pattern Recognition (CVPR)
-
[2017]
InarXiv preprint arXiv:1704.04861
MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications. InarXiv preprint arXiv:1704.04861
-
[2020]
InarXiv preprint arXiv:2006.04768
Linformer: Self-Attention with Linear Complexity. InarXiv preprint arXiv:2006.04768
2006 arXiv
-
[2023]
InProceedings of the ACM Symposium on Operating System Principles (SOSP)
Efficient Memory Management for Large Language Model Serving with Pagedattention. InProceedings of the ACM Symposium on Operating System Principles (SOSP)
-
[2024]
InProceedings of the International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS)
Gscore: Efficient Radiance Field Rendering via Architectural Support for 3D Gaussian Splatting. InProceedings of the International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS)
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.