ShadowNPU presents shadowAttn, a co-designed sparse attention system that uses NPU pilot compute and techniques like graph bucketing and per-head sparsity to minimize CPU/GPU fallback during on-device LLM inference while maintaining accuracy.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
verdicts
UNVERDICTED 2representative citing papers
Presents quantization, checkpointing, softmax approximation, and logits masking to achieve substantial peak memory reductions in LoRA fine-tuning of 3B LLMs.
citing papers explorer
-
ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
ShadowNPU presents shadowAttn, a co-designed sparse attention system that uses NPU pilot compute and techniques like graph bucketing and per-head sparsity to minimize CPU/GPU fallback during on-device LLM inference while maintaining accuracy.
-
Techniques for Peak Memory Reduction for LoRA Fine-tuning of LLMs on Edge Devices
Presents quantization, checkpointing, softmax approximation, and logits masking to achieve substantial peak memory reductions in LoRA fine-tuning of 3B LLMs.