Small quantized language models (Yi, Phi, Llama3) achieve 5 to 12 tokens per second on a CPU-only Raspberry Pi 5 K3s cluster with under 50% CPU and RAM usage, supporting edge inference for 6G applications.
Mobile-llama: Instruction fine-tuning open-source llm for network anal- ysis in 5g networks
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.DC 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Generative AI on the Edge: Architecture and Performance Evaluation
Small quantized language models (Yi, Phi, Llama3) achieve 5 to 12 tokens per second on a CPU-only Raspberry Pi 5 K3s cluster with under 50% CPU and RAM usage, supporting edge inference for 6G applications.