A 1B Llama Guard safety model, compressed to 4-bit weights and a 20-token output vocabulary, runs on a phone at 30+ tokens per second with English F1 slightly better than the full-precision 1B model.
Privileged bases in the transformer residual stream, 2023
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.DC 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations
A 1B Llama Guard safety model, compressed to 4-bit weights and a 20-token output vocabulary, runs on a phone at 30+ tokens per second with English F1 slightly better than the full-precision 1B model.