Turing-RL uses an LLM-based Turing reward in RL to train user simulators that produce responses indistinguishable from real users, outperforming matching baselines on chat and Reddit domains per LLM and human evaluations.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
Learning User Simulators with Turing Rewards
Turing-RL uses an LLM-based Turing reward in RL to train user simulators that produce responses indistinguishable from real users, outperforming matching baselines on chat and Reddit domains per LLM and human evaluations.