The Efficiency Frontier framework models LLM context management as a deployment-aware optimization problem balancing performance, token cost, and amortized preprocessing, with HotpotQA experiments showing 25% token reduction and over 50% cost savings for compression in high-performance regimes.
Longllmlingua: Accelerating and enhancing llms in long context sce- narios via prompt compression
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.CL 2years
2026 2representative citing papers
LLMLingua-2 at ~50% compression preserves semantics on LLaDA but substantially hurts GSM8K math accuracy, with failures driven more by omitted reasoning tokens than semantic drift.
citing papers explorer
-
The Efficiency Frontier: A Unified Framework for Cost-Performance Optimization in LLM Context Management
The Efficiency Frontier framework models LLM context management as a deployment-aware optimization problem balancing performance, token cost, and amortized preprocessing, with HotpotQA experiments showing 25% token reduction and over 50% cost savings for compression in high-performance regimes.
-
Prompt Compression in Diffusion Large Language Models: Evaluating LLMLingua-2 on LLaDA
LLMLingua-2 at ~50% compression preserves semantics on LLaDA but substantially hurts GSM8K math accuracy, with failures driven more by omitted reasoning tokens than semantic drift.