← back to paper
arxiv: 2607.22583 · 2 revisions
Multi-Objective Structured Pruning of LLMs for Latency and Model Size Optimization