An NVCiM-assisted prompt tuning framework stores per-domain virtual tokens in non-volatile memory, retrieves them via a multi-scale search, and improves edge LLM accuracy and speed.
FL-NAS: Towards Fairness of NAS for Resource Constrained Devices via Large Language Models
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Neural Architecture Search (NAS) has become the de fecto tools in the industry in automating the design of deep neural networks for various applications, especially those driven by mobile and edge devices with limited computing resources. The emerging large language models (LLMs), due to their prowess, have also been incorporated into NAS recently and show some promising results. This paper conducts further exploration in this direction by considering three important design metrics simultaneously, i.e., model accuracy, fairness, and hardware deployment efficiency. We propose a novel LLM-based NAS framework, FL-NAS, in this paper, and show experimentally that FL-NAS can indeed find high-performing DNNs, beating state-of-the-art DNN models by orders-of-magnitude across almost all design considerations.
fields
cs.LG 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
NVCiM-PT: An NVCiM-assisted Prompt Tuning Framework for Edge LLMs
An NVCiM-assisted prompt tuning framework stores per-domain virtual tokens in non-volatile memory, retrieves them via a multi-scale search, and improves edge LLM accuracy and speed.