Small quantized language models (Yi, Phi, Llama3) achieve 5 to 12 tokens per second on a CPU-only Raspberry Pi 5 K3s cluster with under 50% CPU and RAM usage, supporting edge inference for 6G applications.
Learning to for- get: Continual prediction with lstm
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.DC 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Generative AI on the Edge: Architecture and Performance Evaluation
Small quantized language models (Yi, Phi, Llama3) achieve 5 to 12 tokens per second on a CPU-only Raspberry Pi 5 K3s cluster with under 50% CPU and RAM usage, supporting edge inference for 6G applications.