A multi-agent test harness with a probabilistic state controller generates diverse conversations for an LLM assistant, achieving a 3.3% break rate versus 5.8% for human testers at 10-12x speed.
User simulation for reinforcement learning of dialogue management policies: Initial progress report
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
method 1
citation-polarity summary
fields
cs.AI 1years
2025 1verdicts
CONDITIONAL 1roles
method 1polarities
use method 1representative citing papers
citing papers explorer
-
Configurable multi-agent framework for scalable and realistic testing of llm-based agents
A multi-agent test harness with a probabilistic state controller generates diverse conversations for an LLM assistant, achieving a 3.3% break rate versus 5.8% for human testers at 10-12x speed.