Small quantized language models (Yi, Phi, Llama3) achieve 5 to 12 tokens per second on a CPU-only Raspberry Pi 5 K3s cluster with under 50% CPU and RAM usage, supporting edge inference for 6G applications.
5g nr; technical specification group radio access network; nr; release 18, 2023
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.DC 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Generative AI on the Edge: Architecture and Performance Evaluation
Small quantized language models (Yi, Phi, Llama3) achieve 5 to 12 tokens per second on a CPU-only Raspberry Pi 5 K3s cluster with under 50% CPU and RAM usage, supporting edge inference for 6G applications.