Pith. sign in

Flash-llm: Enabling cost-effective and highly-efficient large generative model inference with unstructured sparsity

5 Pith papers cite this work, alongside 1 external citations. Polarity classification is still indexing.

5 Pith papers citing it
1 external citations · external index

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 3 cs.DC 2

years

2026 3 2025 2

roles

background 1

polarities

background 1

representative citing papers

RAP: Runtime Adaptive Pruning for LLM Inference

cs.LG · 2025-05-22 · unverdicted · novelty 5.0

RAP is a reinforcement learning framework for runtime-adaptive pruning of LLMs that jointly optimizes model weights and KV-cache usage under varying memory budgets.

citing papers explorer

Showing 5 of 5 citing papers.