{"id":"09fb24ab-6407-4095-aa33-3abc1b4c13e5","arxiv_id":"2506.01881","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A dialogue simulation framework with hidden user states shows that moderate uncertainty about a user's profile can sometimes improve LLM assistant outcomes, but the evidence is limited to a single model configuration.","lead":"This paper introduces a framework that simulates conversations between a user LLM with private internal states and an agent LLM that only sees the visible dialogue, to study when an AI assistant should act versus ask clarifying questions. The authors release a 4,800-dialogue dataset and report that, in one configuration, hiding part of the user profile leads to better satisfaction than full transparency.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 40–60% 'moderate uncertainty advantage' rests on GPT-4o-mini's self-reported inner thoughts and a GPT-4o judge reading those same thoughts; no external or statistical anchor separates the effect from a simulation artifact.","rationale":"The reader's weakest assumption identifies the central vulnerability: STORM's outcome variables are not independently measured but narrated by the same GPT-4o-mini model that plays the UserLLM, and Clarify is judged on those narrations. This is not merely an external-validity concern; it is a measurement-confound risk because the treatment (profile uncertainty) alters the UserLLM's prompt and thereby the content of the inner-thought text that both satisfaction and Clarify read. The paper is internally coherent in its framework and tooling, and the dataset/dashboard are real deliverables, but Table 2 contains no variance estimates, so even internal reproducibility is unverified. The conditional verdict is therefore appropriate: the framework and dataset can stand as contributions, but the 'moderate uncertainty can outperform transparency' headline should not be treated as design guidance until it survives validation on observable dialogue, whether through human raters or repeated-seed statistical testing. My read does not move the verdict; it reinforces the same condition.","tokens_in":20727,"tokens_out":5144,"duration_ms":55317,"concrete_test":"Run a human validation on the exact Claude contrast: take 200 dialogues randomly sampled from the p=0% and p=60% no-profile conditions, strip all generated inner thoughts, satisfaction tags, and profile metadata, and have naive human raters judge, turn by turn, whether the user's intent became clearer and how satisfied the user appears, based only on the observable dialogue. If the human-rated clarity improvement and satisfaction advantage for 60% over 0% do not reproduce the reported 18% and 0.92-vs-0.88 patterns, the central claim is a simulator self-report artifact rather than a real dialogue phenomenon.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim — that 40–60% profile uncertainty can outperform complete transparency — depends on the Claude row in Table 2 where w/o-profile satisfaction (0.92) exceeds w/ profile (0.88) at 60% uncertainty, plus an 18% internal-clarity improvement. Both quantities are measured inside the simulator: satisfaction is 'derived from user inner thoughts' (Section 3.2), and the Clarify judge (GPT-4o) is fed those same inner thoughts in the turn-pair analysis prompt (Appendix M.1). The treatment variable, uncertainty level p, directly changes the UserLLM's prompt and therefore the text it writes inside [INNER_THOUGHTS] and [SATISFACTION] tags. The apparent improvement could be the simulated user telling the judge 'I got clearer' under the 60% condition, rather than a real change in latent intent formation. The paper's self-enhancement rebuttal (Section 3.2), noting that GPT-4o-mini does not score highest as an agent, does not address this confound: it tests whether the UserLLM favors itself, not whether its self-narrated hidden states are causally reliable. In addition, Table 2 reports only point estimates with no error bars or repeated-seed variance, so the 0.92-vs-0.88 difference and the 18% clarity gain could be ordinary sampling noise. The framework and dataset are useful contributions, but the headline empirical conclusion is not yet anchored to observable behavior.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces STORM, a simulation framework for studying what the authors call the Intent-Action Alignment Problem: deciding when a user's utterance is clear enough for an agent to act. STORM pairs a UserLLM, which has access to a full profile and to per-turn hidden states (inner thoughts, satisfaction, intent clarity, emotion), with an AgentLLM, which observes only the dialogue history. The user profile can be masked to varying degrees (p = 0%, 40%, 60%, 80% uncertainty), and metrics are introduced for satisfaction, clarification effectiveness (Clarify), and a composite Satisfaction-Seeking Actions (SSA) score. Experiments across four models and 4,800 simulated dialogues are used to claim that moderate uncertainty (40–60%) can outperform complete transparency in certain scenarios, with the Claude 3.7 Sonnet at 60% uncertainty without a profile reported as the main counterintuitive result. Additional contributions include a dialogue corpus, a visualization dashboard, and a formal notation for asymmetric information in dialogue.","tokens_in":21096,"tokens_out":4943,"duration_ms":51418,"significance":"If the empirical claims held up, the paper would make a useful contribution: a formal, extensible framework for studying asymmetric information in dialogue, a public corpus and dashboard that lower the barrier for follow-up work, and a thought-provoking privacy implication (calibrated information asymmetry as a design feature rather than a defect). The formalization of hidden states and the explicit masking of profile attributes are genuinely useful building blocks. On the other hand, the headline empirical claim is currently anchored only to simulator-internal self-reports, and the quantitative support consists of point estimates without error bars or significance tests. The paper's strengths are the framework, the dataset, and the visualization tool; its weakness is that the central behavioral conclusion is not yet empirically demonstrated.","major_comments":[{"comment":"The headline claim that moderate uncertainty (40–60%) can outperform complete transparency rests entirely on simulator-internal measurements. Satisfaction is extracted from UserLLM-generated [SATISFACTION] tags, and the Clarify score is judged by GPT-4o from a prompt that includes the UserLLM's [INNER_THOUGHTS] for the next turn (Appendix M.1). Because the uncertainty level p is part of the UserLLM prompt, the treatment directly changes the text inside those tags; the Claude 60% result (0.92 vs 0.88, plus the reported 18% internal-clarity improvement) could therefore be an artifact of the simulated user telling the judge that it became clearer under the 60% condition, rather than evidence of a real change in intent formation. The Section 3.2 rebuttal addresses only whether GPT-4o-mini favors itself as an agent; it does not establish that LLM-generated inner thoughts are causally reliable proxies for human intent formation. An external anchor (human evaluation, observable task outcome, or at least a judge blinded to the uncertainty condition) is needed before this result can be transferred to human-AI collaboration.","section":"§3.1–3.2, Table 2, Appendix M.1"},{"comment":"All values in Table 2 are point estimates with no error bars, confidence intervals, or significance tests, and the number of dialogues per condition is not stated. The critical differences are small (e.g., average satisfaction 0.92 vs 0.88; high-satisfaction rate 86.7% vs 80.7%), so without repeated-seed variance or a bootstrap/permutation analysis the 60%-uncertainty advantage could be ordinary sampling noise. Please report per-condition sample sizes and variability across independent simulation runs or seeds.","section":"Table 2"},{"comment":"The SSA metric is not parameter-free and its construction is data-dependent. The normalization factor λ = 7.75 is defined as the maximum observed Clarify score in the same 4,800-dialogue dataset, so every SSA value in Table 2 is normalized by an in-sample constant; the Appendix F ranking (Llama > Gemini > GPT > Claude) can be changed by rescaling this constant. Moreover, the Clarify definition C = w1Δt(h) + w2Δt(s) + w3gt in Section 2.2 never reports the values of w1, w2, w3, and the SSA weights wα = 0.7, wβ = 0.3 are described only as illustrative. Please report all weights, justify the choice of λ (e.g., through a held-out calibration set), and provide sensitivity analyses over these free parameters.","section":"§3.2, Appendix B.1"},{"comment":"The manuscript repeatedly refers to The UserLLM's hidden state ht as \"ground-truth\" internal state, but ht is generated by the same family of LLMs being evaluated and is available only in simulation. This is not a fatal flaw—simulation is a legitimate first step—but the terminology overstates the epistemic status of the data and invites the circularity concern raised in Section 3.1's own discussion of the difficulty of obtaining true internal-state data. The hidden states should be described as simulated or LLM-generated latent states, and the claims about \"internal cognitive improvements\" (e.g., the 18% clarity improvement) should be qualified accordingly.","section":"§1, §2.2"}],"minor_comments":[{"comment":"The hidden-state tuple is written ht = ⟨st, ct, it, et⟩, but the surrounding text names only satisfaction, intent clarity, and emotional state; please define the fourth component or correct the tuple.","section":"§2.2"},{"comment":"The two displayed forms of the SSA metric are algebraically different (SSA = λ(wαSavg + wβCclarify) in Section 3.2 versus SSA = wα(Savg·λ) + wβCclarify in Appendix B.1); please present one canonical definition and derive any equivalent form.","section":"§3.2 vs Appendix B.1"},{"comment":"Clarify and SSA are reported only in the \"w/o Profile\" columns; please state explicitly whether these metrics were computed only in the no-profile condition or whether profile-aware values exist but were omitted.","section":"Table 2"},{"comment":"The Executive Summary reports response-appropriateness gains of 15–28%, intent-alignment improvements of 45–65%, and satisfaction improvements of 4–23%, but I could not locate the experiments or tables supporting these numbers in the main text or appendices; please add the supporting analysis or remove the claims.","section":"Appendix E"},{"comment":"There are formatting errors in the author block and references (e.g., the project website link appears garbled, and the Allen et al. reference spells the author name as \"Horvtz\"); these should be cleaned before publication.","section":"References and front matter"}],"recommendation":"major_revision","confidential_remarks":"The central claim is original and worth pursuing, but the empirical evidence is not yet strong enough for a journal-level acceptance: the main result is measured inside the simulator, the SSA normalization is in-sample, and no inferential statistics are reported. I see no integrity concerns; the public dataset and dashboard are positive assets. The manuscript also has some annex material (Appendix E, Appendix F) that makes larger claims than the main text supports, which should be reconciled. For this journal, I would ask the authors to either add external validation or substantially soften the transfer-to-human-AI conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the paper is worth reading for its framework and dataset, but don't buy the headline empirical claim yet.\n\nWhat's actually new: STORM gives a clean formalization of asymmetric information in dialogue—user has hidden state, agent sees only behavior—and ships a 4,800-dialogue corpus with per-turn hidden-state annotations, plus a visualization dashboard. The Clarify metric, which scores an agent by its effect on the user's internal clarity, is a nice idea even if the implementation is simulation-dependent. Framing calibrated information asymmetry as a design feature (rather than a defect) is a genuinely interesting hypothesis that connects to privacy and bias. Credit for releasing code, data, and a useful tool.\n\nWhere it's soft: the central result—that 40–60% uncertainty can beat transparency—rests on a single Claude row in Table 2 (0.92 vs 0.88 satisfaction) plus an 18% clarity gain, both measured entirely inside the simulator. Satisfaction is extracted from the UserLLM's self-narrated inner thoughts, and the GPT-4o judge reads those same thoughts when scoring Clarify. The treatment variable (uncertainty p) directly modifies the user prompt, so the simulated user may simply be telling the judge it got clearer. The self-enhancement rebuttal doesn't touch this confound. Table 2 has no error bars or significance tests, and the SSA normalization uses the in-sample maximum Clarify score, making that metric data-dependent. The Clarify weights are also unreported. So the empirical conclusion is not anchored to observable behavior.\n\nWhat survives: the framework itself is coherent, the dataset is a plausible resource, and the paper is transparent about many of its design choices. The authors don't overclaim—they say \"certain scenarios\"—but the abstract's \"moderate uncertainty can outperform complete transparency\" is stronger than the evidence supports.\n\nWho this is for: people working on LLM user simulation, dialogue evaluation methodology, and personalization/privacy trade-offs. It deserves a serious peer review, but the right outcome is major revision: add repeated-seed variance, validate the inner-thought measurements against something external, report all weights, and reframe the 40–60% finding as hypothesis-generating. I'd cite it for the dataset and framework, not for the empirical result.","headline":"A useful framework and dataset from which the headline empirical claim about moderate uncertainty is not yet supported.","tokens_in":21603,"tokens_out":2384,"would_cite":true,"duration_ms":24668,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"STORM, a dialogue-simulation framework, claims that hiding 40-60% of a user profile can make an AI assistant more helpful than revealing everything, because incomplete information curbs presumptive reasoning and fosters clarifying…","keywords":["intent-action alignment","asymmetric information","dialogue systems","intent clarity","uncertainty calibration","LLM role-play simulation","inner thoughts","clarification-vs-action tradeoff"],"falsifier":"Run the masked-profile dialogue protocol with human users instead of a simulated UserLLM, eliciting goal clarity directly after each turn (self-reported clarity, or third-party judges who see only the transcript). If a fully informed assistant does not lose ground to a 60%-masked one on users' own clarity ratings — or if the 18% internal-clarity improvement reported for Claude disappears or reverses — then the central empirical claim fails.","tokens_in":20544,"feed_emoji":"💬","tokens_out":8857,"duration_ms":78182,"temperature":0.7,"pith_summary":"This paper tries to establish that the failure of dialogue systems is often not a misunderstanding of words but a misalignment of timing: the system cannot tell when a user's request, though semantically parseable, is ready for action. It builds STORM, a simulation framework in which a simulated user with full access to its own private goals and emotions talks to an assistant that sees only the observable conversation, and it claims that hiding part of the user's profile can improve the interaction. Concretely, the paper reports that a moderate uncertainty level — 40-60% of the profile masked — can outperform complete transparency, because complete profiles invite presumptive reasoning while moderate uncertainty elicits more open, clarifying questions. A sympathetic reader would care because this reframes information asymmetry and privacy as design levers rather than defects, and offers metrics that track internal cognitive progress instead of only surface satisfaction.","feed_headline":"Moderate uncertainty beats full transparency in AI dialogue","feed_subtitle":"Simulated dialogues show 40-60% profile masking beats total access: agents ask open questions instead of presuming.","key_machinery":"\nSTORM (State Trajectory oriented Representation Model), formalized as a five-domain tuple $\\{\\mathcal{T}, \\mathcal{U}, \\mathcal{E}, \\mathcal{R}, \\mathcal{H}\\}$ spanning tasks, user profiles, expressions, responses, and hidden states. The load-bearing object is the hidden-state vector $h_t = \\langle s_t, c_t, i_t, e_t \\rangle$, which the paper treats as encoding satisfaction, intent clarity, emotion, and the user's inner thoughts; it is private to the UserLLM and invisible to the AgentLLM. The uncertainty parameter $p \\in \\{0\\%, 40\\%, 60\\%, 80\\%\\}$ decides what fraction of the user profile the agent cannot see, and it is the dial whose tuning produces the paper's central results. The Clarify metric — a third-party judge (GPT-4o) scoring turn by turn whether an agent response improved the user's internal intent clarity — together with the intent-evolution measure $\\Delta_t(h) = h_t.\\text{clarity} - h_{t-1}.\\text{clarity}$, converts the unobservable process of \"becoming clearer\" into a quantifiable trajectory that the paper compares across models and uncertainty levels.","core_discovery":"The paper's central claim is that the Intent-Action Alignment Problem — knowing when an utterance is not just understood but truly ready for system action — can be studied as information-asymmetry dynamics, and that calibrated asymmetry is a design lever rather than a defect. In STORM, a UserLLM with full access to its own hidden state (goals, emotions, satisfaction, and recorded \"inner thoughts\") converses with an AgentLLM that sees only the dialogue history, producing 4,800 annotated dialogues over 600 profiles and four assistant models. The central empirical finding is that profile access boosts satisfaction by 15-40%, but that a moderate masking level (40-60% of the profile hidden) can beat full transparency: Claude at 60% uncertainty scored 0.92 satisfaction without a profile versus 0.88 with one, and its responses improved users' internal clarity by 18% relative to the 0% baseline. The paper explains this by observing that complete profiles lead to stereotypical, presumptive answers, whereas moderate uncertainty pushes assistants toward open, assumption-free questions; it generalizes this into task-dependent advice (simple tasks prefer low uncertainty, exploratory medical and housing tasks prefer high uncertainty) and into the claim that limiting information acts as an implicit bias mitigator.","pith_inferences":["If the 40-60% sweet spot survives replication with human users, privacy and performance stop being a trade-off: an assistant could be deliberately deprived of access to demographic data and behave better because of it, a consequence the paper gestures at but does not implement.","The transfer bottleneck is the inner-thought ground truth; a natural next experiment is a think-aloud human study comparing self-reported goal clarity under a masked versus a fully informed assistant, which the paper does not run.","The clarity trajectories STORM records could directly train a wait-versus-act stopping rule for production dialogue systems, connecting the framework to predictive wait-or-answer policies that the paper cites but does not integrate."],"forward_implications":["Calibrated masking becomes a design parameter: system builders can choose how much user data to expose to an assistant, and 40-60% masking outperforms full transparency in several configurations.","Task complexity predicts the right uncertainty level: simple tech support works best with low uncertainty, while medical and housing decisions favor higher masking levels because users stay internally uncertain longer.","Satisfaction alone misleads: successful clarification correlates with internal cognitive improvement more than with expressed satisfaction, so satisfaction-only evaluation misranks dialogue systems.","Model-specific deployment makes sense: Llama clarifies goals best (Clarify 7.58-7.75), Gemini is robust to missing profiles, Claude maximizes satisfaction, and GPT-4o-mini is consistent but flat.","Strategic information limitation acts as an implicit bias mitigator: with full profiles, agents stereotype (for instance, assuming elderly users need simplified help), while at optimal uncertainty they assess individuals."],"supporting_citations":[{"why":"Grounds the premise that user intent matures gradually rather than appearing fully formed, which motivates modeling intents as dynamic hidden states.","marker":"Subramonyam et al. (2023)"},{"why":"CAMEL-AI is the two-agent generation pipeline that STORM is built upon, supplying the user-assistant conversation scaffolding.","marker":"(Li et al., 2023)"},{"why":"Establishes that human evaluation of conversations is an open problem, motivating the new synthetic evaluation methodology.","marker":"Smith et al. (2022)"},{"why":"Supplies the Big Five personality dimensions that seed STORM's user profiles and drive variation in expression style.","marker":"Barrick & Mount (1991)"},{"why":"Documents the limitations of static, goal-oriented evaluation paradigms that STORM aims to replace with developmental intent tracking.","marker":"Zhou et al. (2023)"},{"why":"Defines the wait-or-answer decision in dialogue systems, the class of problem that the Intent-Action Alignment framing extends.","marker":"Lin et al. (2021)"}],"fun_headline_variants":["When AI knows less, it converses better","Moderate masking beats full transparency in dialogue","AI dialogue: partial access can outperform full access","The 40-60% uncertainty sweet spot for AI assistants","Hiding some info improves AI conversation performance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the LLM-generated \"inner thoughts\" and satisfaction scores are faithful measurements of how a real user's intent actually becomes clearer; if they are not, the 18% clarity gain and the 40-60% advantage reported here may not transfer to human-AI conversation.","fun_headline_variants_meta":{"raw":{"variants":["When AI knows less, it converses better","Moderate masking beats full transparency in dialogue","AI dialogue: partial access can outperform full access","The 40-60% uncertainty sweet spot for AI assistants","Hiding some info improves AI conversation performance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000257,"raw_usage":{"total_tokens":1608,"prompt_tokens":1005,"completion_tokens":603,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":621,"completion_tokens_details":{"reasoning_tokens":531}},"tokens_in":621,"tokens_out":603,"duration_ms":7386,"temperature":1.0,"reasoning_tokens":531,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:31:55.663401+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the masked-profile dialogue protocol with human users instead of a simulated UserLLM, eliciting goal clarity directly after each turn (self-reported clarity, or third-party judges who see only the transcript). If a fully informed assistant does not lose ground to a 60%-masked one on users' own clarity ratings — or if the 18% internal-clarity improvement reported for Claude disappears or reverses — then the central empirical claim fails.","supporting_citations":[],"review_version":1}