A water-filling token allocation policy optimally distributes output tokens across sequential LLM agents to maximize reliability under latency and cost constraints.
We introduced performance models quantifying trade-offs among reliability, latency, and cost, and derived closed-form optimal token allocation policies for sequential workflows
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Toward Reliable Design of LLM-Enabled Agentic Workflows: Optimizing Latency-Reliability-Cost Tradeoffs
A water-filling token allocation policy optimally distributes output tokens across sequential LLM agents to maximize reliability under latency and cost constraints.