{"id":"1dc3960d-34e7-44fd-9192-7e3fd2e68faa","arxiv_id":"2608.00478","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A single-participant Chinese dataset and training framework produce Ekova, a five-persona agent for self-discovery dialogue, outperforming prompt-based baselines on style accuracy by 9.5 points.","lead":"This paper introduces Personality Support (PS), a new goal for dialogue agents: helping users understand themselves, not just feel better. It builds a Chinese dataset, trains five style-specific personas, and wraps them into a persistent agent called Ekova.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"SA and PCL are measured by the same Kimi K2.5 judge that OrthoTune optimizes in Eq. 2, so the headline 9.5pp SA gap may reflect judge overfitting rather than genuine style boundaries.","rationale":"The reader's weakest assumption was the single-participant design, and they mentioned 'the same LLM judge used in training' in passing, but they did not make it the primary load-bearing concern. I agree that the single-participant design limits generalizability, but the more fundamental problem is internal: the SA and PCL metrics are scored by Kimi K2.5, the very model whose style judgments OrthoTune's training loss explicitly optimizes (Eq. 2). This is a direct training/evaluation leak that can inflate the headline numbers even on the author's own data. It is not merely a matter of transferring to other users; it threatens the validity of the central empirical claim as reported. The paper's human SC results show a much smaller advantage for OrthoTune over prompted GPT-5.5, which is consistent with the concern that p_φ SA is being gamed. The LPS-Helpfulness correlation also relies on the same judge for PCL, so the calibration evidence is entangled. Therefore, the manuscript should be accepted only under the condition that the core comparison is re-evaluated with an independent judge and ideally human annotators, and preferably with multi-participant data. This does not reject the conceptual PS paradigm or the dataset construction, but it means the headline empirical claims are not yet trustworthy. Hence, keeping the verdict CONDITIONAL rather than moving to ACCEPT or REJECT is the appropriate adjustment.","tokens_in":13035,"tokens_out":5040,"duration_ms":59057,"concrete_test":"Re-run the Table 3 comparison with an evaluation judge not used anywhere in training: score SA and PCL with GPT-5.5 or Claude-Sonnet-4.6 (using the same style descriptions and scoresheet) on the same DSD test set, and have human annotators rate Style Consistency on a random 300-response subset. Compute the OrthoTune minus best-prompt-baseline deltas under the independent judge and under p_φ. If the SA delta shrinks below ~3 points or reverses, the headline 9.5pp SA claim and the associated minimal-unit-boundary conclusion are not supported. If the independent judge reproduces the gap, the concern is resolved.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The most load-bearing weakness is not only the single-participant scope; it is that the headline automatic metrics are evaluated by the same model that the training objective optimizes. In Eq. 2, OrthoTune adds ε·log p_φ(s_i | a_{1:|A|}) to the SFT loss, with p_φ = frozen Kimi K2.5. In §6, Style Accuracy (SA) is defined as the fraction of responses whose top p_φ confidence score matches the target style, and PCL is rated by the same Kimi K2.5. Thus the 9.5 pp SA advantage over the strongest prompt-only baseline and the 63.0 pp diagonal-vs-off-diagonal gap in the cross-style matrix (Table 4) are measured on a criterion that OrthoTune was explicitly trained to maximize. The paper's justification that Kimi K2.5 is 'absent from the baseline pool' addresses self-evaluation among baselines, but not training/evaluation leakage: p_φ is inside the training loop and inside the metric. Consequently, the claim that 'style prompting alone cannot instantiate minimal-unit boundaries' rests on a potentially gamed score. This is independent of the single-participant generalization limitation: it threatens validity even on the author's own test set. The human SC ratings (4.36 vs 4.18) show a much smaller advantage, suggesting the p_φ SA gap may overstate the true style-boundary difference. The LPS correlation with Helpfulness (0.907) also uses PCL from the same judge, so the calibration claim is entangled as well.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a new paradigm, Personality Support (PS), distinct from Emotional Support (ES), targeting cognitive clarity and self-articulation. It presents DSD, a Chinese single-user longitudinal dataset of 8,590 samples across five 'minimal units' (Coach, Warm, Tsukkomi, Real, Gonzo); DeepSupport, a multi-persona system trained with OrthoTune, which uses style-specific LoRA adapters and a style-consistency regularizer driven by a frozen Kimi K2.5 judge p_phi (Eq. 2); and Ekova, a persistent agent with cross-session memory and adaptive routing. The headline result is that OrthoTune-trained models outperform the strongest prompt-based baseline by 16.3% average relative gain, with a 9.5 percentage-point SA gap and a 16.6% relative SDI improvement. The paper claims that style prompting alone cannot instantiate the minimal-unit boundaries, and that parameter-level specialization plus inference-time guards are required.","tokens_in":13465,"tokens_out":2324,"duration_ms":27625,"significance":"If the empirical claims held, this would be a substantive contribution: it proposes a new task family (PS), a novel multi-style fine-tuning framework with a style-consistency regularizer, and a persistent-agent architecture that routes across support personas. The paper is explicit about its limitations (single-participant design, Chinese-native scope, non-clinical positioning), and it ships code, which is a strength. The key difficulty is that the central automatic metrics are entangled with the training signal, and the evaluation is built on a single user's data. Thus the significance is conditional on whether the style-boundary results survive an independent evaluation protocol.","major_comments":[{"comment":"The central claim that OrthoTune achieves a 9.5pp SA gain and a 63.0pp diagonal-vs-off-diagonal cross-style gap is compromised by training/evaluation leakage. The style-consistency regularizer in Eq. (2) explicitly maximizes log p_phi(s_i | a_{1:|A|}) with p_phi = frozen Kimi K2.5, and the SA metric is defined as the fraction of responses whose top p_phi confidence score matches the target style. PCL is also rated by the same Kimi K2.5. The paper's justification that Kimi K2.5 is absent from the baseline pool addresses self-evaluation among baselines, but not the fact that the metric is optimized during training. This threatens validity even on the author's own test set. The human SC ratings (4.36 vs 4.18) show a much smaller advantage, suggesting the automatic SA gap may overstate the true style-boundary difference. Please re-evaluate with a judge not used in training, or report human S","section":"§5, Eq. (2) and §6 Metrics"},{"comment":"The single-participant design is a load-bearing limitation for the generalizability of the 'minimal unit' hypothesis and the 16.3% average relative gain. All training and test data come from one user's longitudinal stream; the five style definitions, the style judge, and the DSD dataset are all derived from that individual. The conclusion acknowledges this ('The single-participant design limits generalizability'), but the manuscript's broader claims about PS as a general paradigm and about the necessity of parameter-level specialization would require at least a second user's data or an explicitly framed proof-of-concept status. Please either extend the evaluation to multiple participants or substantially temper the general claims.","section":"§4 and §7"},{"comment":"Warm is the largest subset (5,200 samples) and is augmented with '30% template-based synthetic data for cold-start and low-disclosure scenarios.' It is unclear whether this synthetic data is included in the 90/10 train/test split. If synthetic templates appear in the test set, the Warm SA/SDI/PCL numbers (Table 3) and the ablation results (w/o Synthetic Data, Table 5) could be inflated. Please clarify the split and, if synthetic data is in the test set, report results on the human-only portion.","section":"§4 Warm subset and Table 1"}],"minor_comments":[{"comment":"Typo: 'Conversation' is misspelled as 'Coversation' in two places in the figure. Also, the figure is dense and the flow from 'Users' to the style subsets is hard to follow; consider simplifying or adding numbered steps.","section":"Figure 3"},{"comment":"The text states 'OrthoTune outperforms all baselines on every metric at p < 0.01' but does not specify the statistical test, the number of samples, or how multiple comparisons were handled. Please provide test details (e.g., paired bootstrap or Wilcoxon) and report effect sizes.","section":"§6 p-value claim"},{"comment":"Krippendorff's alpha is reported as 'above 0.7' globally, but Table 3 and the method description state that 3 raters were used while Appendix B mentions 9 annotators. Please reconcile the rater count and report per-metric alpha values.","section":"§6 Human evaluation"},{"comment":"Several 2026 references (e.g., GPT-5.5, Claude Sonnet 4.6, Kimi K2.5) are cited without version identifiers or full author lists, making reproducibility harder. If these are preprints or model cards, please include stable identifiers or access dates.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper has a promising high-level idea and a transparently reported system, but the evaluation protocol is the key blocker. The training/evaluation leakage via the shared p_phi judge is, in my view, a correct and serious concern that the authors need to address with an independent evaluator or by de-emphasizing the automatic SA/PCL numbers. The single-participant limitation is acknowledged but should be moved from a concluding caveat to a central framing constraint, otherwise the paper overclaims generality. I do not see this as a reject: the framework, dataset, and interface are contributions that could be published after a rigorous re-evaluation, but the current headline numbers should not be taken at face value."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this paper is worth a look for the paradigm shift it proposes — Personality Support as a distinct objective from Emotional Support — but the empirical headline numbers are compromised by a training/evaluation circularity. The PS framing, the five minimal units (Warm, Tsukkomi, Real, Gonzo, Coach), and the OrthoTune recipe of per-style LoRA adapters plus a frozen style judge are genuinely new relative to the ES literature. The DSD dataset, while single-participant, is collected through real longitudinal interaction, and the paper is unusually candid about its limits: it explicitly flags the single-participant design in the conclusion, lists persona-specific failure modes, and positions Ekova as a cognitive-clarity tool, not a clinical one. That honesty earns credit.\n\nThe soft spot is load-bearing. The style judge p_phi is a frozen Kimi K2.5 used in Eq. 2 as a regularizer during training, and the same judge is used to compute Style Accuracy and PCL in evaluation. So the 9.5 pp SA advantage over the strongest prompt-based baseline, and the 63-point diagonal/off-diagonal gap, are measured on a criterion the model was explicitly optimized to satisfy. The paper's argument that Kimi K2.5 is absent from the baseline pool addresses self-evaluation among baselines, not training/evaluation leakage. The human SC ratings show a much smaller advantage (4.36 vs 4.18), which is consistent with the judge being gamed. The 0.907 LPS-helpfulness correlation is also entangled because PCL comes from the same judge. This threatens validity within the author's own test set, independently of the single-participant generalizability issue.\n\nThat said, the single-participant limitation is acknowledged and not the primary problem. The primary problem is the shared judge. If the paper separates trainer and evaluator, or releases the data and code for independent verification, the underlying claim — that style prompting alone cannot instantiate minimal-unit boundaries — could survive. As is, the empirical support is weak but the conceptual contribution is strong.\n\nWho is this for? Researchers working on dialogue style control, emotional support, or LLM-as-judge methodology. It deserves a serious referee, but the referee should push for a training/evaluation judge split or a fully independent evaluation.\n\nI'd send it to review with a request for major revision; the ideas are worth engaging with even if the current evidence doesn't carry them.","headline":"Genuinely new PS framing and five-style taxonomy, but the headline gains are compromised because the same Kimi judge trains the model and grades it.","tokens_in":13965,"tokens_out":2749,"would_cite":true,"duration_ms":31851,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes a distinct 'personality support' paradigm—helping users articulate how they see themselves—and reports that a five-style agent trained with parameter-level adapters beats style-prompted baselines by 16.3% on average.","keywords":["personality support","emotional support","self-discovery dialogue","multi-persona dialogue","LoRA adapters","style consistency","longitudinal dataset","Chinese dialogue"],"falsifier":"Give Ekova to a new cohort of users and compare per-user PCL/SDI gains and style accuracy against the best style-prompted baseline; the central claim fails if the 9.5-point style-accuracy gap shrinks, the LPS-helpfulness correlation drops well below 0.9, or Tsukkomi and Gonzo are rated as unhelpful by most participants.","tokens_in":12910,"feed_emoji":"💬","tokens_out":7949,"duration_ms":82873,"temperature":0.7,"pith_summary":"Emotional support systems optimize for one thing: making users feel better in the moment. This paper argues that a second, complementary goal has been neglected—helping users see themselves more clearly—and names that goal Personality Support (PS). To test the idea, the author collected 8,590 Chinese dialogue samples from her own longitudinal interaction, decomposed PS into five minimal conversational styles (Coach, Warm, Tsukkomi, Real, Gonzo), trained one adapter per style under a style-consistency regularizer, and wrapped the result in a persistent agent called Ekova with cross-session memory. The paper reports that the trained personas outperform the strongest prompt-based baseline by 16.3% average relative gain, with a 9.5-point style-accuracy lead over style-prompted GPT-5.5 and a composite objective (LPS) that correlates with human helpfulness at 0.907. The central claim is that function-level support styles require parameter-level specialization plus inference-time guards, rather than prompt descriptions alone.","feed_headline":"Five support personas beat prompt-only models by 16.3 percent","feed_subtitle":"A new 'personality support' agent claims cognitive clarity needs trained style boundaries, not just prompting.","key_machinery":"The central mechanism is the minimal-unit decomposition of Personality Support, enforced by OrthoTune, a training framework that gives each style its own LoRA adapter on a shared Qwen2.5-32B base and adds a style-consistency regularizer: a frozen LLM judge outside the baseline pool scores each response's conformity to the target style, and that score enters the training loss. At inference, each persona has a dedicated guard—Coach retrieves exemplars from a style bank, Warm retrieves from a conversation memory buffer, Tsukkomi filters out non-ironic register, and Gonzo rejects outputs without cross-domain analogy markers—while Real runs without post-processing. Ekova then unifies the five per","core_discovery":"The paper's central claim is that Personality Support is a measurable and trainable objective distinct from affect regulation. It formalizes PS as LPS(U,A) = λ·PCL(U,A) + (1−λ)·SDI(U,A), where PCL measures convergence of the user's problem statement to an actionable formulation and SDI measures movement from generic to personal disclosure. The Minimal Unit Hypothesis then asserts that five conversational styles—Warm, Tsukkomi, Real, Gonzo, Coach—each make a non-replicable contribution to LPS, such that removing any one strictly reduces the system's ability to advance PCL or SDI. The reported evidence for this is a 63.0-point gap between diagonal and off-diagonal style-accuracy, with OrthoTun","pith_inferences":["The paper's single-user grounding means the five-style decomposition is a hypothesis about conversational cognition, not a proven universal typology; a multi-participant replication with the same data collection protocol would test whether Coach, Warm, Tsukkomi, Real, and Gonzo remain the minimal covering set for other people's self-discovery processes.","Because LPS is a weighted sum with λ tuned to human helpfulness, a natural extension is per-user λ adaptation: users who respond better to irony or analogy would get different blends, which the paper's architecture already supports through manual persona selection.","If cross-session memory compounds self-understanding, the agent's value should increase with usage; a longitudinal study could measure whether PCL/SDI gains accumulate over weeks rather than plateauing after the first session.","The claimed 16.3% average relative gain is against prompt-based baselines on a single-user test set; the margin could shrink or grow with other base models, other languages, or other users, so the headline number should be read as evidence for the mechanism, not a fixed performance guarantee."],"forward_implications":["Multi-style support requires parameter-level specialization and inference guards; prompt descriptions alone don't instantiate minimal-unit boundaries.","PS has its own evaluation metrics—PCL and SDI—so progress in this paradigm should be measured by self-articulation, not just affect.","Ekova's cross-session memory means a persistent PS agent can carry disclosed context across sessions, compounding self-understanding; this motivates longitudinal evaluation.","The five units and their guard designs provide a reusable recipe for building adaptive, user-customizable support agents."],"supporting_citations":[{"why":"Defines emotional support conversation and provides ESConv, the paradigm and baseline OrthoTune must beat.","marker":"Liu et al., 2021"},{"why":"Supplies the generic-to-personal-to-core-pattern disclosure progression used to define SDI and justify the single-participant longitudinal design.","marker":"Pennebaker, 1997"},{"why":"SoulChat provides a Chinese emotional-support baseline and dataset that the PS framing distinguishes from.","marker":"Chen et al., 2023"},{"why":"LoRA is the adaptation method behind the per-style adapters in OrthoTune.","marker":"Hu et al., 2022"},{"why":"The frozen style judge p_phi provides the style-consistency training signal and the PCL evaluation, outside the baseline pool.","marker":"Kimi Team et al., 2025"},{"why":"Qwen2.5-32B-Instruct is the shared base LLM for all five personas.","marker":"Yang et al., 2025"},{"why":"CPsyCoun is a Chinese counseling baseline and dataset comparison point.","marker":"Zhang et al., 2024"},{"why":"Companion-chatbot user experiences are cited to argue that companionship adds no cognitive function beyond the five minimal units.","marker":"Ta et al., 2020"}],"fun_headline_variants":["Five personas beat prompts by 16.3% for self-discovery","Personality support AI: 16.3% gain over prompt-only","Ekova: 5-style agent lifts self-discovery dialogue 16.3%","Self-discovery AI outdoes prompts with 5 trained personas"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The load-bearing premise is that the five style definitions, the training data, and the evaluations are all grounded in one person's interaction stream, so the style boundaries and the 16.3% gain are inferred from a single user's data.","fun_headline_variants_meta":{"raw":{"variants":["Five personas beat prompts by 16.3% for self-discovery","Personality support AI: 16.3% gain over prompt-only","Ekova: 5-style agent lifts self-discovery dialogue 16.3%","Self-discovery AI outdoes prompts with 5 trained personas"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000578,"raw_usage":{"total_tokens":2573,"prompt_tokens":768,"completion_tokens":1805,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":1728}},"tokens_in":512,"tokens_out":1805,"duration_ms":17174,"temperature":1.0,"reasoning_tokens":1728,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T00:50:32.022201+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give Ekova to a new cohort of users and compare per-user PCL/SDI gains and style accuracy against the best style-prompted baseline; the central claim fails if the 9.5-point style-accuracy gap shrinks, the LPS-helpfulness correlation drops well below 0.9, or Tsukkomi and Gonzo are rated as unhelpful by most participants.","supporting_citations":[{"cited_title":"2021 , publisher =","cited_arxiv_id":null,"evidence_quote":"Defines emotional support conversation and provides ESConv, the paradigm and baseline OrthoTune must beat."},{"cited_title":"2023 , address =","cited_arxiv_id":null,"evidence_quote":"SoulChat provides a Chinese emotional-support baseline and dataset that the PS framing distinguishes from."},{"cited_title":"2024 , address =","cited_arxiv_id":null,"evidence_quote":"CPsyCoun is a Chinese counseling baseline and dataset comparison point."}],"review_version":1}