Releases a public trace of coding-agent LLM sessions and characterizes workload features including long loops, context lengths, tool diversity, and cache hit rates for serving optimization.
Zhang, Zhilin Yang, Xinyu Zhou, Mingxing Zhang, and Jiezhong Qiu
2 Pith papers cite this work. Polarity classification is still indexing.
fields
cs.LG 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
Flux Attention uses a context-aware Layer Router to dynamically assign full or sparse attention to each LLM layer, achieving up to 2.8x prefill and 2.0x decode speedups with competitive performance on long-context and reasoning tasks.
citing papers explorer
-
TraceLab: Characterizing Coding Agent Workloads for LLM Serving
Releases a public trace of coding-agent LLM sessions and characterizes workload features including long loops, context lengths, tool diversity, and cache hit rates for serving optimization.
-
Flux Attention: Context-Aware Hybrid Attention for Efficient LLMs Inference
Flux Attention uses a context-aware Layer Router to dynamically assign full or sparse attention to each LLM layer, achieving up to 2.8x prefill and 2.0x decode speedups with competitive performance on long-context and reasoning tasks.