{"id":"60301ac9-0846-4289-92ed-ff50dd380151","arxiv_id":"2607.08326","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"LLMs collapse advice into a single supportive persona; Inverse-Process Distillation restores human-like persona diversity, yet raters still prefer the collapsed default.","lead":"Frontier LLMs give almost all personal advice in one warm, supportive voice, while top human advisors shift among five postures by situation. A training method can restore that variety, but experienced raters still prefer the warm default—especially when challenge is warranted.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Diagnosis and IPD repair treat top-rated community persona distributions as the policy target, yet the paper’s own preference study shows experienced raters reject that target (especially for challenge).","rationale":"The reader correctly isolates the human-upvote reference as the weakest assumption that underwrites both diagnosis and repair metrics. The descriptive collapse (models flat near 100% Healer across 14 contexts) is robust and does not require the reference to be optimal; the load-bearing step is treating match to that reference as successful repair. The preference study already supplies the decisive evidence that the reference is rejected by the very population whose judgment is later invoked, so no stronger internal inconsistency is needed. The LLM-judge pipeline and Stoic–Doomer confusion are secondary and largely owned by the authors. Keeping the CONDITIONAL verdict is therefore appropriate; the concern does not overturn the empirical findings but correctly limits how far “repaired toward human distribution” can be read as better advice.","tokens_in":29260,"tokens_out":562,"duration_ms":22873,"concrete_test":"On the 165 posts from the preference study, collect independent ideal (H, E) labels from a new panel of experienced advisors (or re-use the 199) for each situation alone, without seeing any response. Recompute JS(Instruct ∥ rater-ideal) versus JS(IPD ∥ rater-ideal). If IPD no longer reduces divergence by ~80% (or increases it), the repair claim fails relative to a preference-aligned target.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that frontier models exhibit persona collapse and that Inverse-Process Distillation repairs it (JS cut ~80%) is defined relative to the persona distribution of top-rated human comments (§3.2, §4, §5, Eq. 2). The paper repeatedly uses this distribution for Neff, JS, SFT targets, and confusion matrices while stating it is only an empirical benchmark, not ground truth. The preregistered human study (§6) then shows 199 experienced advice-givers prefer the unrepaired Healer default over every repaired model, with the largest penalty precisely on Stoic/Doomer items. Thus the quantity being optimized (match to upvote-derived personas) is not the quantity the same population endorses as better advice. If upvotes reward rhetoric, local norms, or performative harshness rather than wise stance selection, both the diagnosis of collapse-as-failure and the claim of successful repair are mis-aimed; the preference result is then not a surprising verifier problem but evidence that the reference itself is the wrong objective.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper formalizes advice-giving as situation-conditioned persona selection on a two-axis space (hedonic valence H and agentic depth E) with five landmark personas, and defines persona collapse as compression of diverse situations into a single default posture. On 1,281 Reddit advice posts across 14 contexts, top-rated human responses vary systematically (Neff ≈ 3.8), while GPT-5.1, Claude Opus 4.5, and Gemini 3 Pro put ≥89% of responses in the Healer region regardless of context. On an 8,262-item multi-source corpus, three open Instruct models show the same pattern; post-training trajectory analysis on OLMo3 shows capacity without selection in the Base model and capacity loss with only weak selection gains after SFT/DPO/RLVR. Inference-time plan-first prompting worsens collapse; LoRA SFT variants (Direct, Persona, Inverse-Process Distillation) restore Neff near the human reference and cut JS divergence by roughly half to 80% depending on model, with residual Stoic–Doomer confusion. A preregistered blinded study (N=199 experienced advice-givers) finds raters prefer the unrepaired Instruct default over all repaired models on tone fit, understanding, and accountability, most strongly when the human gold persona is confrontational, with modest within-session drift toward the repaired models.","tokens_in":29610,"tokens_out":1634,"duration_ms":15309,"significance":"If the descriptive claims hold, the paper identifies a policy-level failure mode for a high-stakes, high-volume LLM use case that existing persona and advice evaluations largely miss: not inability to be warm, but inability to allocate stance by situation. Strengths include a multi-context diagnostic corpus, three frontier and three open models, bootstrap CIs, confusion matrices, an OLMo post-training trajectory, Inverse-Process Distillation as an abductive process-supervision method for open-ended advice, and a preregistered mixed-effects human study. The preference result is itself a contribution: it frames advice as a verifier problem where immediate likeability and crowd-upvote persona distributions can diverge from each other and from longer-term agency support. The work is a useful stress test for alignment when rewards are delayed and subjective.","major_comments":[{"comment":"The central diagnosis and repair claims are defined relative to the persona distribution of top-rated community comments (§3.2 Eqs. 1–2; §4; §5; Table 13). The paper states this is an empirical benchmark not ground truth, yet Neff, JS, SFT targets, and κ/macro-recall all optimize or score against it. The preregistered study (§6, Figs. 6–8) then shows experienced advice-givers prefer the unrepaired Healer default, with the largest penalty on Stoic/Doomer items. This is not merely a surprising verifier problem: it is evidence that the quantity being repaired may not be the quantity the same population endorses as better advice. The manuscript needs a clearer separation of (i) descriptive claim that models are less context-varying than upvote-selected humans, (ii) normative claim that matching that distribution is desirable, and (iii) the preference result as a test of (ii). Without that, ‘","section":null},{"comment":"Persona labels for humans and models, and the (H,E) inputs to SFT-Persona and IPD teacher traces, all come from the same LLM-as-judge pipeline (gpt-5.4-nano; Appendix B, Table 5: ~60% axis-averaged exact agreement). Diagnosis, training targets, and evaluation are therefore coupled to one labeling scheme. The human preference study mitigates this for preference claims but not for Neff/JS/κ claims. At minimum the paper should report inter-judge reliability against human raters on a non-trivial subset, sensitivity of collapse/repair conclusions to judge model or region boundaries, and whether Stoic–Doomer confusion (§5.5, Fig. 5) is partly a judge-axis artifact (shared H=-1, differing E).","section":null},{"comment":"SFT variants restore distributional diversity but under-produce Healer (~20–24% vs human 42.3% on OLMo3; Table 13) and systematically map human-Stoic items to Doomer (0.57–0.68 on OLMo3; similar on Llama/Qwen). The abstract’s ‘cuts divergence by ~80%’ is therefore a marginal-distribution success with limited item-level selection success (κ ~0.20–0.22). The paper should qualify ‘repair’ accordingly and discuss whether IPD mainly teaches challenging tone rather than constructive agency support—the distinction the framework itself treats as load-bearing (§3.1).","section":null},{"comment":"The human study is single-turn and short-form; SFT replies are much shorter (~60 vs ~130 words). Raters may be penalizing curtness, missing conversational groundwork, or length/style confounds rather than persona fit alone (§6.3, Box 1; Limitations §8). Planned analyses should control for length or present length-matched ablations; otherwise the preference gap cannot be cleanly attributed to persona selection versus execution quality.","section":null}],"minor_comments":[{"comment":"Figure 3 caption refers to ‘Figure 3A’ in the text but the figure is not panel-labeled that way in the manuscript body; align caption and callouts.","section":null},{"comment":"Appendix numbering jumps (Appendix G Qualitative Repair Example appears after F; H follows); renumber for navigation.","section":null},{"comment":"Table 6 reports only Healer vs non-Healer for frontier models; a full five-way breakdown (as for humans) would strengthen the collapse claim.","section":null},{"comment":"Clarify whether the ~80% JS reduction in the abstract is relative to Instruct baseline or to the theoretical maximum, and report the exact baseline/repaired JS pair used for that headline number.","section":null},{"comment":"CareerVillage is named CareerNet in Appendix D Table 7; use one name consistently.","section":null},{"comment":"Box 1 examples are highly illustrative; state how they were selected (worst-case, median, or representative of the Stoic/Doomer gap) to avoid selection concerns.","section":null}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong fit for a computational social science / AI alignment venue that values empirical policy-level failures over pure method novelty. The main risk is overclaiming ‘repair of good advice’ when the optimized target is upvote-derived personas that the paper’s own raters reject. If the authors reframe cleanly as distributional diagnosis + preference misalignment, the contribution is solid; if they insist on repair-as-welfare without longitudinal outcomes, I would lean reject on overclaim. No integrity concerns; preregistration and multi-model reporting are genuine strengths."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: frontier models really do flatten advice into one warm default across contexts, humans don’t, and when you force the model back toward the human mix, experienced raters still prefer the warm default—especially when challenge is warranted. That last result is the paper’s sharpest contribution, not a side note.\n\nWhat’s new is the framing. Treating advice as situation-conditioned persona selection on a two-axis space (hedonic tone × agency support), defining collapse as a policy-level failure, and measuring it with Neff and JS against a human reference is cleaner than the usual per-response warmth/helpfulness scores. The diagnostic evidence is strong: 1,281 posts, 14 contexts, three frontier models stuck at ~90%+ Healer while human non-Healer mass swings from ~10% to 90%. The OLMo post-training trajectory is useful—capacity without selection in the base, then capacity stripped and selection only weakly improved. Plan-first prompting making collapse worse is a nice negative result. IPD is a reasonable answer-conditioned process scaffold; the ~80% JS cut is real on their metrics, and the Stoic–Doomer confusion is honestly reported.\n\nSoft spots, in proportion. The load-bearing reference is top-rated community comments. The authors say it isn’t ground truth, and they’re right—upvotes can reward rhetoric or local norms. The stress-test note that “repair toward upvotes” is then rejected by raters is fair as a caution, but it doesn’t break the paper: they present that rejection as a verifier problem for advice, not as proof that the repaired policy is better welfare. Still, diagnosis and repair both live inside the same LLM-judge H/E pipeline, so circularity is real even if the preference study is independent. Single-turn setup also makes confrontational SFT replies look blunter than they would with rapport. No code/data in the manuscript is a practical annoyance for a methods-heavy claim.\n\nMath and stats look fine for this genre—bootstrap CIs, confusion matrices, preregistered mixed models, N=199 after exclusions. Citations sit in the right neighborhood (sycophancy, persona work, therapist responsiveness) without obvious padding.\n\nWho it’s for: alignment, HCI, and anyone building advice or coaching systems. Worth a serious referee. I’d engage with it; the preference finding alone is something people will argue about productively.","headline":"Solid empirical diagnosis of advice persona collapse; the real payload is the preference–repair tension, which the authors mostly own rather than paper over.","tokens_in":30198,"tokens_out":587,"would_cite":true,"duration_ms":10993,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Frontier LLMs collapse almost all personal advice into one warm persona, while humans shift stance by situation—and people still prefer the collapsed default.","keywords":["persona collapse","LLM advice","situation-conditioned persona","Inverse-Process Distillation","hedonic valence","agency support","human preference","alignment"],"falsifier":"If longitudinal or outcome-linked studies showed that matching the human persona mixture (including more challenge) does not improve decision quality, agency, or safety relative to the unconditional Healer default—or if re-labeling the same posts with expert clinicians erased the human–model gap—the diagnosis of collapse and the case for repairing toward that reference would fail.","tokens_in":30157,"feed_emoji":"🤖","tokens_out":746,"duration_ms":7521,"temperature":0.7,"pith_summary":"People now take hard personal questions to language models: relationships, moral dilemmas, money, crisis. Good human advisors do not answer every case with the same posture; they comfort in crisis, challenge denial, and stay procedural on logistics. This paper argues that post-training locks modern assistants into a single supportive default, and that this is a policy failure, not just a style quirk. The authors map advice onto two axes—immediate affective tone and depth of agency support—yielding five advisory personas. Across 1,281 real advice posts, top-rated human replies systematically shift among those personas by context, while three frontier models put over 90% of replies in the supportive “Healer” mode no matter the context. Asking the model to plan its stance first makes the collapse worse. A training method that reconstructs the situational reading behind each human reply restores much of the human persona distribution, cutting divergence by roughly 80%. Yet in a blinded study, 199 experienced advice-givers still prefer the collapsed default—especially when the situation calls for challenge—though that preference softens a little across repeated exposures.","feed_headline":"AI advice collapses into one warm persona","feed_subtitle":"Humans shift stance by situation; repaired models match them—yet raters still prefer the default","key_machinery":"Persona collapse: failure of situation-conditioned selection among five advisory modes (Healer, Stoic Challenger, Technician, Enabler, Doomer) in a space of hedonic valence and agentic depth. The main repair is Inverse-Process Distillation—reconstructing a per-item situational reading that could have produced each human reply, then training on that scaffold rather than answers alone.","core_discovery":"The paper establishes that frontier LLMs exhibit persona collapse in advice: they compress heterogeneous situations into an almost unconditional supportive persona, while top-rated human advice varies systematically across five stances defined by hedonic tone and agency support. Simple planning prompts deepen the collapse; Inverse-Process Distillation restores distributional diversity toward the human reference, but experienced raters still prefer the default Assistant, most strongly when confrontation is called for.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["LLMs collapse advice into one warm persona","Models force a single supportive stance on all advice","AI advice flattens diverse situations to one persona","Human advisors shift stance; LLMs stay locked warm","Persona collapse: frontier models ignore context shifts"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The load-bearing premise is that top-rated community comments are a fair empirical benchmark for how an advisor should shift stance by situation, even though the paper treats them as a descriptive reference rather than true ground truth.","fun_headline_variants_meta":{"raw":{"variants":["LLMs collapse advice into one warm persona","Models force a single supportive stance on all advice","AI advice flattens diverse situations to one persona","Human advisors shift stance; LLMs stay locked warm","Persona collapse: frontier models ignore context shifts"]},"model":"grok-4.5","effort":"low","cost_usd":0.00345,"raw_usage":{"total_tokens":1143,"prompt_tokens":804,"num_sources_used":0,"completion_tokens":54,"cost_in_usd_ticks":34500000,"prompt_tokens_details":{"text_tokens":804,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":285,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":804,"tokens_out":54,"duration_ms":3206,"temperature":1.0,"reasoning_tokens":285,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-10T09:20:43.297947+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If longitudinal or outcome-linked studies showed that matching the human persona mixture (including more challenge) does not improve decision quality, agency, or safety relative to the unconditional Healer default—or if re-labeling the same posts with expert clinicians erased the human–model gap—the diagnosis of collapse and the case for repairing toward that reference would fail.","supporting_citations":[],"review_version":1}