Pith. sign in

Any-precision llm: Low-cost deployment of multiple, different-sized llms

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

fields

cs.LG 3 cs.AR 1

years

2026 4

verdicts

UNVERDICTED 4

representative citing papers

GRINQH: Graded Input-based Quantization Hierarchy for Efficient LLM Generation

cs.LG · 2026-06-22 · unverdicted · novelty 6.0

GRINQH introduces a graded input-based quantization hierarchy that dynamically assigns multi-precision weights using activation magnitudes as importance proxy, unifying quantization with sparsification to improve LLM decoding speed and quality trade-offs on Llama3 and Qwen3 models.

Multi-Bitwidth Quantization for LLMs Using Additive Codebooks

cs.LG · 2026-06-11 · unverdicted · novelty 5.0

Drop-by-Drop uses additive codebooks and Matryoshka-style training to produce one LLM model whose ordered codebook subsets give accurate reconstructions at successively higher bitwidths under a weighted MSE distortion.

Cassandra: Enabling Reasoning LLMs at Edge via Self-Speculative Decoding

cs.AR · 2026-05-26 · unverdicted · novelty 5.0

Cassandra is a self-speculative decoding system that builds a draft model via fine-grained data selection and optimized pruning/mantissa truncation, achieving up to 2.41x speedup over BF16 and 1.81x more tokens than Eagle-3 on Llama 3 8B without training.

citing papers explorer

Showing 4 of 4 citing papers.