A water-filling token allocation policy optimally distributes output tokens across sequential LLM agents to maximize reliability under latency and cost constraints.
Consider a sequential workflow com- prising n= 5 LLM agents with heterogeneous reliability parameters βj ∈ {0.001,0.002,0.0005,0.003,0.0015}
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Toward Reliable Design of LLM-Enabled Agentic Workflows: Optimizing Latency-Reliability-Cost Tradeoffs
A water-filling token allocation policy optimally distributes output tokens across sequential LLM agents to maximize reliability under latency and cost constraints.