CIMple delivers a 32 kb digital SRAM-based compute-in-memory accelerator for transformer self-attention that reaches 26.1 TOPS/W at 0.85 V in 28 nm with INT8 precision using dual-banked architecture and LUT-based split softmax.
Efficient Softmax approximation for deep neural networks with attention mechanism
3 Pith papers cite this work. Polarity classification is still indexing.
verdicts
UNVERDICTED 3representative citing papers
Presents quantization, checkpointing, softmax approximation, and logits masking to achieve substantial peak memory reductions in LoRA fine-tuning of 3B LLMs.
4-bit and 6-bit integer-only quantized Transformers implemented on Spartan-7 FPGA for AIoT time-series forecasting achieve 0.63% higher test loss than 8-bit baselines but up to 132x speedup and 48x lower energy.
citing papers explorer
-
CIMple: Standard-cell SRAM-based CIM with LUT-based split softmax for attention acceleration
CIMple delivers a 32 kb digital SRAM-based compute-in-memory accelerator for transformer self-attention that reaches 26.1 TOPS/W at 0.85 V in 28 nm with INT8 precision using dual-banked architecture and LUT-based split softmax.
-
Techniques for Peak Memory Reduction for LoRA Fine-tuning of LLMs on Edge Devices
Presents quantization, checkpointing, softmax approximation, and logits masking to achieve substantial peak memory reductions in LoRA fine-tuning of 3B LLMs.
-
Integer-only Quantized Transformers for Embedded FPGA-based Time-series Forecasting in AIoT
4-bit and 6-bit integer-only quantized Transformers implemented on Spartan-7 FPGA for AIoT time-series forecasting achieve 0.63% higher test loss than 8-bit baselines but up to 132x speedup and 48x lower energy.