Pith. sign in

REVIEW 4 cited by

BitNet a4.8: 4-bit Activations for 1-bit LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.04965 v1 pith:W7PUI4TQ submitted 2024-11-07 cs.CL cs.LG

classification cs.CLcs.LG
keywords bitnetllmsactivationsinferencequantizationwhileenablingperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent research on the 1-bit Large Language Models (LLMs), such as BitNet b1.58, presents a promising direction for reducing the inference cost of LLMs while maintaining their performance. In this work, we introduce BitNet a4.8, enabling 4-bit activations for 1-bit LLMs. BitNet a4.8 employs a hybrid quantization and sparsification strategy to mitigate the quantization errors introduced by the outlier channels. Specifically, we utilize 4-bit activations for inputs to the attention and feed-forward network layers, while sparsifying intermediate states followed with 8-bit quantization. Extensive experiments demonstrate that BitNet a4.8 achieves performance comparable to BitNet b1.58 with equivalent training costs, while being faster in inference with enabling 4-bit (INT4/FP4) kernels. Additionally, BitNet a4.8 activates only 55% of parameters and supports 3-bit KV cache, further enhancing the efficiency of large-scale LLM deployment and inference.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SiLQ: Simple Large Language Model Quantization-Aware Training

    cs.LG 2025-07 conditional novelty 6.0 of 10

    SiLQ fine-tunes 8B-parameter LLMs with quantized weights, activations, and cache for a small fraction of extra training tokens, matching or beating leading post-training quantization methods.

  2. Scaling Law for Quantization-Aware Training

    cs.LG 2025-05 conditional novelty 6.0 of 10

    A unified QAT scaling law predicts 4-bit quantization error from model size, training tokens, and group size, showing activation outliers in the FC2 layer are the main W4A4 bottleneck.

  3. TAT-VPR: Ternary Adaptive Transformer for Dynamic and Efficient Visual Place Recognition

    cs.CV 2025-05 conditional novelty 5.0 of 10

    TAT-VPR is a ternary-quantized vision transformer for visual place recognition that can dynamically skip up to 40% of computation at runtime with under 1% Recall@1 loss on one benchmark.

  4. ReTern: Exploiting Natural Redundancy and Sign Transformations for Enhanced Fault Tolerance in Compute-in-Memory based Ternary LLMs

    cs.AR 2025-06 conditional novelty 4.0 of 10

    A training-free method that combines zero-weight bit-cell repair with per-column sign flips to make ternary LLMs on compute-in-memory accelerators substantially more tolerant to stuck-at faults.

Pith tools