Releases a public trace of coding-agent LLM sessions and characterizes workload features including long loops, context lengths, tool diversity, and cache hit rates for serving optimization.
Abdi, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.LG 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
STS repurposes draft-model attention scores from speculative decoding to build token-and-head-wise sparsity masks, delivering 2.67x speedup at ~90% sparsity on NarrativeQA with negligible accuracy loss.
citing papers explorer
-
TraceLab: Characterizing Coding Agent Workloads for LLM Serving
Releases a public trace of coding-agent LLM sessions and characterizes workload features including long loops, context lengths, tool diversity, and cache hit rates for serving optimization.
-
STS: Efficient Sparse Attention with Speculative Token Sparsity
STS repurposes draft-model attention scores from speculative decoding to build token-and-head-wise sparsity masks, delivering 2.67x speedup at ~90% sparsity on NarrativeQA with negligible accuracy loss.