Foldgpt: Simple and effective large lan- guage model compression scheme

Songwei Liu, Chao Zeng, Lianqiang Li, Chenqian Yan, Lean Fu, Xing Mei, Fangmin Chen · 2024 · arXiv 2407.00928

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

representative citing papers

S2O: Early Stopping for Sparse Attention via Online Permutation

cs.LG · 2026-02-26 · unverdicted · novelty 6.0

S2O uses online permutation and importance-based early stopping to increase effective sparsity in attention, delivering 7.51x attention and 3.81x end-to-end speedups on Llama-3.1-8B at 128K context with preserved accuracy.

citing papers explorer

Showing 1 of 1 citing paper.

S2O: Early Stopping for Sparse Attention via Online Permutation cs.LG · 2026-02-26 · unverdicted · none · ref 16
S2O uses online permutation and importance-based early stopping to increase effective sparsity in attention, delivering 7.51x attention and 3.81x end-to-end speedups on Llama-3.1-8B at 128K context with preserved accuracy.

Foldgpt: Simple and effective large lan- guage model compression scheme

fields

years

verdicts

representative citing papers

citing papers explorer