Stream2LLM introduces adaptive scheduling and preemption for append-mode and update-mode context streaming in disaggregated LLM deployments, delivering up to 11x TTFT improvements on real-world workloads while preserving throughput.
Data Organization
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.DB 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Stream2LLM: Overlap Context Streaming and Prefill for Reduced Time-to-First-Token (TTFT)
Stream2LLM introduces adaptive scheduling and preemption for append-mode and update-mode context streaming in disaggregated LLM deployments, delivering up to 11x TTFT improvements on real-world workloads while preserving throughput.