PALUTE is a new PIM accelerator using in-DRAM LUTs on M3D DRAM that reports 1264 TPS at 0.16 W with 12.8x energy efficiency gains over CHIME for quantized edge LLM inference.
Mobile edge intelligence for large language models: A contemporary survey,
3 Pith papers cite this work, alongside 116 external citations. Polarity classification is still indexing.
representative citing papers
ST-SFLora reduces communication in split federated learning by selecting semantically important tokens via attention scores and jointly optimizing them with wireless bandwidth and power allocation.
VCON is a unified framework for smooth iterative DNN compression that uses parallel execution and an affine combination to progressively replace the original model with its compressed form during fine-tuning.
citing papers explorer
-
PALUTE: Processing-In-Memory Acceleration via Lookup Table for Edge LLM Inference
PALUTE is a new PIM accelerator using in-DRAM LUTs on M3D DRAM that reports 1264 TPS at 0.16 W with 12.8x energy efficiency gains over CHIME for quantized edge LLM inference.
-
Semantic-aware Token Selection and Resource Optimization for Communication-efficient Split Federated Fine-tuning in Edge Intelligence
ST-SFLora reduces communication in split federated learning by selecting semantically important tokens via attention scores and jointly optimizing them with wireless bandwidth and power allocation.
-
Vanishing Contributions: A Unified Framework for Smooth and Iterative Model Compression
VCON is a unified framework for smooth iterative DNN compression that uses parallel execution and an affine combination to progressively replace the original model with its compressed form during fine-tuning.