Pith. sign in

hub

RPTQ: reorder-based post-training quantization for large language models

14 Pith papers cite this work. Polarity classification is still indexing.

14 Pith papers citing it

hub tools

citation-role summary

background 1

citation-polarity summary

roles

background 1

polarities

background 1

representative citing papers

BitNet Text Embeddings

cs.CL · 2026-06-24 · conditional · novelty 6.0

BITEMBED trains 1.58-bit ternary-weight LLM embedders with contrastive pre-training, supervised distillation, and multi-precision output training, matching FP16 teachers within ~0.6 MMTEB points at ~2x CPU speed.

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization

cs.LG · 2026-05-06 · unverdicted · novelty 6.0 · 2 refs

OSAQ suppresses weight outliers in LLMs via a closed-form additive transformation from the Hessian's stable null space, improving 2-bit quantization perplexity by over 40% versus vanilla GPTQ with no inference overhead.

An Empirical Study of OpenPangu Quantization on Ascend NPUs

cs.LG · 2026-06-19 · conditional · novelty 5.0 · 2 refs

On Ascend NPUs, OpenPangu-7B tolerates 4-bit weight quantization with minor accuracy loss, while OpenPangu-1B degrades sharply on math/code, and 2-bit/binary settings collapse to near-random output.

A Survey on Efficient Inference for Large Language Models

cs.CL · 2024-04-22 · accept · novelty 3.0

The paper surveys techniques to speed up and reduce the resource needs of LLM inference, organized by data-level, model-level, and system-level changes, with comparative experiments on representative methods.

citing papers explorer

Showing 14 of 14 citing papers.