{"id":"9da13248-febd-4e9a-b941-cf51e0a4383c","arxiv_id":"2606.19336","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Turing-RL uses an LLM-based Turing reward in RL to train user simulators that produce responses indistinguishable from real users, outperforming matching baselines on chat and Reddit domains per LLM and human evaluations.","lead":"The paper proposes Turing-RL, a reinforcement learning method that trains user simulator models using a reward from an LLM judge scoring how indistinguishable generated responses are from real user responses given history. This could improve training of AI assistants and evaluation of personalization systems by focusing on realism rather than exact matching to ground truth.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No evidence that indistinguishability gains transfer to downstream utility (agent training/evaluation)","rationale":"The identified concern is identical to the reader's weakest_assumption. The abstract-only limitation already correctly flags the missing transfer evidence; the full-text reference does not alter this gap in the reported results.","tokens_in":1639,"tokens_out":276,"duration_ms":24384,"concrete_test":"Train a downstream dialogue agent using Turing-RL simulators versus the paper's baselines in the conversational-chat domain; measure agent task-completion rate or human-rated interaction quality on a held-out test set. If the Turing-RL condition shows no statistically significant improvement, the effectiveness claim for simulator learning does not hold.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The claim that Turing-RL is effective for learning user simulators requires that higher scores on LLM-judge and human indistinguishability metrics produce simulators that are more useful for the stated applications (training agents, personalization evaluation, etc.). The abstract reports outperformance versus response-matching baselines on LLM and human metrics but contains no experiments measuring transfer—e.g., no results on agent success rates, personalization accuracy, or other utility proxies when the simulator is used in the loop. This leaves the central effectiveness claim dependent on an untested correlation between surface human-likeness and functional utility.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes Turing-RL, a reinforcement learning approach to train LLM-based user simulators. Instead of maximizing likelihood or similarity to ground-truth responses, it uses a discriminative Turing reward from a separate LLM judge that scores how indistinguishable a generated response is from a real user's response given conversation history. Experiments in conversational chat and Reddit discussion domains report that Turing-RL outperforms response-matching baselines on both LLM-judge and human indistinguishability metrics, leading to the claim that optimizing for indistinguishability is effective for learning user simulators applicable to agent training, personalization evaluation, and social science research.","tokens_in":1768,"tokens_out":407,"duration_ms":14477,"significance":"If the indistinguishability improvements transfer to downstream utility, the method offers a principled alternative to direct response matching for user simulation. The Turing-reward formulation is a clear conceptual contribution and the consistent outperformance on proxy metrics across two domains is a positive signal. However, the significance is limited by the absence of any experiments linking the reported metric gains to the applications listed in the abstract.","major_comments":[{"comment":"Abstract: The central claim that Turing-RL is effective for the stated applications (training agent assistants, personalization evaluation, etc.) rests on an untested correlation between higher LLM-judge/human indistinguishability scores and functional utility. No experiments measure downstream performance such as agent success rates, personalization accuracy, or other utility proxies when the learned simulators are placed in the loop.","section":"Abstract"},{"comment":"Abstract: No information is supplied on the judge LLM (model identity, prompting procedure, training data, or fine-tuning), statistical significance of the reported gains, or controls for potential confounds such as judge-simulator data overlap. These omissions make it impossible to assess whether the Turing reward is reliable or whether the outperformance is robust.","section":"Abstract"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback. We address the two major comments point by point below and outline planned revisions to the manuscript.","responses":[{"response":"We agree that the manuscript contains no downstream experiments linking indistinguishability gains to the motivating applications. The abstract and introduction present those applications as potential use cases for improved user simulators rather than as claims of demonstrated utility. The core contribution is the demonstration that optimizing for indistinguishability yields better simulators on the reported metrics. We will revise the abstract and add a dedicated limitations paragraph to clarify the scope and note the absence of downstream validation.","revision_made":"partial","referee_comment":"[Abstract] Abstract: The central claim that Turing-RL is effective for the stated applications (training agent assistants, personalization evaluation, etc.) rests on an untested correlation between higher LLM-judge/human indistinguishability scores and functional utility. No experiments measure downstream performance such as agent success rates, personalization accuracy, or other utility proxies when the learned simulators are placed in the loop."},{"response":"We acknowledge these omissions in the current draft. In the revised version we will add an appendix section detailing the judge LLM (including model name and version), the exact prompting procedure, any fine-tuning or training data used for the judge, statistical significance tests (e.g., paired t-tests or bootstrap confidence intervals) for all reported gains, and an explicit discussion of data-overlap controls between the judge and the simulator training sets.","revision_made":"yes","referee_comment":"[Abstract] Abstract: No information is supplied on the judge LLM (model identity, prompting procedure, training data, or fine-tuning), statistical significance of the reported gains, or controls for potential confounds such as judge-simulator data overlap. These omissions make it impossible to assess whether the Turing reward is reliable or whether the outperformance is robust."}],"tokens_in":1342,"tokens_out":410,"duration_ms":21141,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core proposal is to train user simulators via RL with a reward from an LLM judge that scores how well a generated turn could have come from the real user given the history. This replaces the usual approach of maximizing log probability or similarity to a single reference response. They run it on conversational chat and Reddit discussion data and report consistent wins on both LLM-based and human judgments of human-likeness.\n\nThat is the actual new piece: treating indistinguishability itself as the training signal inside the RL loop rather than response matching. The experiments appear to be set up cleanly enough to show the metric improvements hold across the two settings.\n\nThe main limitation is that the results stop at those indistinguishability metrics. The abstract lists applications like training agent assistants and evaluating personalization systems, yet there are no experiments that close the loop and measure whether the Turing-RL simulators produce better agent performance, more accurate personalization estimates, or any other functional outcome. Without that transfer data it is hard to know whether the judge is capturing something useful or just stylistic surface features.\n\nDetails on the judge model, training procedure, and statistical tests are also thin in the provided abstract, which makes it difficult to assess how robust the reported gains really are or whether the judge shares data or architecture with the simulator in ways that could create indirect leakage.\n\nThis is the kind of paper that would interest people building or evaluating conversational agents who are already thinking about alternative reward signals. A reader looking for a new training trick could extract the idea and try it, but anyone needing evidence that the method moves the needle on actual utility will come away unsatisfied. The work is coherent on its own terms and worth sending out for referee comments so the authors can address the missing transfer tests and evaluation details.","headline":"Turing-RL gets better indistinguishability scores than matching baselines on two domains, but the paper shows no evidence these gains improve the downstream uses it claims to target.","tokens_in":2262,"tokens_out":433,"would_cite":false,"duration_ms":17652,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Training user simulators to produce responses indistinguishable from real users by an LLM judge outperforms training them to match exact ground-truth replies.","keywords":["user simulators","Turing rewards","reinforcement learning","LLM judges","indistinguishability","conversational agents","Reddit discussions","social simulation"],"falsifier":"A downstream experiment in which agents trained on Turing-RL simulators show no measurable advantage over agents trained on matching-based simulators when both interact with real human users.","tokens_in":2556,"feed_emoji":"🤖","tokens_out":653,"duration_ms":16462,"temperature":0.7,"pith_summary":"The paper proposes training user simulator models through reinforcement learning where an LLM judge scores responses for how well they could have come from the real user given the conversation history. This Turing-test style reward replaces the usual goal of matching a single correct reply. The method is evaluated in conversational chat and Reddit forum domains, where it beats standard approaches on both automated and human metrics. A sympathetic reader would care because improved simulators could make agent training, system evaluation, and social research more realistic without needing perfect response copies for every turn.","feed_headline":"User simulators learn by fooling LLM judges","feed_subtitle":"Rewarding responses an LLM cannot distinguish from real users beats exact matching in chat and forum domains.","key_machinery":"The discriminative Turing reward, an LLM-judge score measuring how well a generated response blends with the user's history as if produced by the real user.","core_discovery":"Turing-RL trains an LLM user simulator via reinforcement learning with a discriminative Turing reward that an LLM judge assigns based on how indistinguishable the generated response is from what the real user might have said. Rather than maximizing log probability or similarity to a ground-truth reply, the simulator optimizes for passing the judge's human-likeness test. Across conversational chat and Reddit discussion domains the resulting models score higher on both LLM and human evaluation metrics than baselines that rely on response matching.","pith_inferences":["If the LLM judge signal holds, the approach could lower the cost of collecting large human response datasets for simulator training.","The same reward structure might apply to simulating other interactive behaviors where exact matching is difficult or unnecessary.","Substituting human judges for the LLM judge could be tested in high-stakes evaluation settings to check transfer of the signal.","Measuring real-human performance of agents trained on these simulators would directly test whether the claimed utility gains materialize."],"forward_implications":["User simulators no longer require exact ground-truth responses at every turn to be effective.","The indistinguishability objective improves performance consistently in both open chat and structured forum settings.","LLM-based and human judgments both favor simulators trained with the Turing reward over matching baselines.","Better simulators can directly support agent assistant training and personalization system evaluation."],"fun_headline_variants":["Turing rewards boost user simulator realism","LLM judge rewards train realistic user sims","Indistinguishable responses via Turing-RL training","Turing-RL outperforms matching for user simulators"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"An LLM judge supplies a reliable signal of human-likeness that corresponds to the simulator's actual usefulness in downstream tasks rather than just stylistic similarity.","fun_headline_variants_meta":{"raw":{"variants":["Turing rewards boost user simulator realism","LLM judge rewards train realistic user sims","Indistinguishable responses via Turing-RL training","Turing-RL outperforms matching for user simulators"]},"model":"grok-4.3","cost_usd":0.003735,"raw_usage":{"total_tokens":1911,"prompt_tokens":619,"num_sources_used":0,"completion_tokens":55,"cost_in_usd_ticks":37349500,"prompt_tokens_details":{"text_tokens":619,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1237,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":619,"tokens_out":55,"duration_ms":8056,"temperature":1.0,"reasoning_tokens":1237,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T20:50:42.657230+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A downstream experiment in which agents trained on Turing-RL simulators show no measurable advantage over agents trained on matching-based simulators when both interact with real human users.","supporting_citations":[],"review_version":1}