{"id":"01e566b4-70c7-4c1e-9f2c-971c8dc1b094","arxiv_id":"2607.27851","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Proposes a 'capability-sustaining' longitudinal paradigm for emotional dialogue, backed by an audit showing current systems focus on relief and never measure long-term capability.","lead":"This paper argues that AI support chatbots should be judged on whether repeated use keeps users able to regulate emotions, cope, make their own decisions, and stay connected to people—not just whether each chat feels better. A small audit of past work shows most systems aim for short-term relief, and the paper offers a framework and evaluation plan for a capability-focused research agenda.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The motivating gap (95% relief, 0 longitudinal) rests on an arXiv-only, title/abstract-level, AI-coded sample; if human full-text coding reclassifies even one system paper, the field-level claim weakens.","rationale":"The reader identified the audit as the load-bearing empirical premise: arXiv-only, title/abstract-level, AI-coded. I agree. The paper is transparent about this, but transparency does not make the sample representative, and the headline claims ('95% relief', 'None evaluates capability or longitudinal outcomes') are stated prominently. The proposed check—human full-text recoding of the 60 system papers—directly tests whether the zero counts are a measurement artifact. If the check confirms the counts, the audit is credible within its scope; if not, the motivation is exaggerated. Since the CSED paradigm proposal could still stand as a normative framework even if the empirical gap is slightly overstated, the reader's CONDITIONAL verdict is appropriate; I do not change it.","tokens_in":17984,"tokens_out":2950,"duration_ms":30823,"concrete_test":"Recode the 60 system-building papers with two independent human annotators reading full texts (not just abstracts), using the same D1/D3/D4 codebook and preregistered definitions for 'capability outcome' and 'longitudinal'. Report chance-adjusted agreement and recompute the 0/91 and 57/60 counts. If any paper is reclassified as having a capability outcome or longitudinal evaluation, the audit's headline gap is not robust; if all remain, the claim survives within the sampled scope.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical premise is the audit's claim that system-building research is relief-oriented and lacks capability/longitudinal evaluation: 57/60 relief (95%), 0/91 capability outcomes, 0/91 longitudinal horizon. This is inferred from 91 arXiv records coded at title/abstract level by LLMs (Appendix B). The sample frame is narrow (arXiv only, title/abstract only, seven query phrases), and the codebook is applied by AI coders with paradigm-level kappa ≈ 0.716. The paper itself limits the claim to the sample, but the paradigm's motivation—that the field has a strategy gap—requires the sample to represent the field. If the 0/91 count is a coding artifact (e.g., capability or longitudinal terms appear in full texts but not in abstracts) or if relevant work appears in non-arXiv venues, the motivating gap is substantially weakened. The authors acknowledge these limitations, but the counts remain load-bearing because the entire 'gap' narrative and the need for CSED depend on them.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a third paradigm for emotional dialogue, Capability-Sustaining Emotional Dialogue (CSED), whose goal is to provide effective support while sustaining users' capacities for emotion regulation, active coping, self-endorsed decision making, and social connectedness across the full interaction lifecycle, including repeated use, non-use, transition, and termination. The motivation rests on two audits: a PRISMA-ScR-guided scoping audit of 91 year-stratified arXiv records (60 system-building papers) coded at title/abstract level, and a function-level coding of 300 ESConv supporter turns. The headline findings are that 95% of system-building papers pursue relief-oriented goals, none evaluates capability or longitudinal outcomes, and only one considers dependency/autonomy/termination risk; in ESConv, 43.0% of sampled turns contain capability-relevant functions, with generic suggestion-giving contributing 22.0%. The paper then gives an illustrative process model connecting latent capability to four evaluation timescales and lifecycle constraints, formalizes benchmark nesting and short-horizon non-identifiability, and derives a research and governance agenda. The authors are transparent about the scoped and preliminary nature of the evidence and release a protocol for extending the audit to deployed model behavior.","tokens_in":18315,"tokens_out":5639,"duration_ms":60304,"significance":"If the motivating gap is accepted, CSED is a potentially valuable synthesis: it connects design commitments to a process model, multiple evaluation horizons, and lifecycle governance in a way that much current emotional-dialogue work does not. The paper's strengths include exact, internally consistent arithmetic; a clear claim-traceability table (Table 7); explicit reliability checks on the AI-only coding (Appendix C.2); structural propositions that locate common benchmarks as nested cases; and a candid Limitations section. The combination of a formalizable paradigm with a reproducible audit protocol is useful for a field moving toward deployed, sustained-use systems. The central risk is not the formalism but the empirical audit that motivates it: the field-level 'gap' is inferred from a narrow arXiv/title-abstract sample coded only by LLMs, and one of the ESConv headline numbers depends on a grouping that is ambiguous as written.","major_comments":[{"comment":"The load-bearing claim that 'none evaluates capability or longitudinal outcomes' and that '95% of system papers pursue relief' is computed from 91 arXiv records coded at title/abstract level by LLMs. The Limitations section admits this, but the Abstract and the 'What the Audit Establishes' subsection state the result as a field-level fact without the sample scope. Because the entire CSED motivation rests on the existence of this gap, this is a substantive issue. Please either add a validation subsample (for example, human full-text coding of all 60 system papers for capability/longitudinal/risk terms, or a multi-database check) or consistently rephrase all occurrences as 'in the arXiv title/abstract sample' and soften the 'strategy gap' language accordingly.","section":"Abstract; 'What the Audit Establishes'; Appendix B.1"},{"comment":"The 43.0% 'capability-relevant' figure counts F4 (problem solving, 66/300 turns) as capability-relevant, but the same paragraph describes this 22.0% as 'generic suggestion-giving.' If generic suggestions are not capability-sustaining, excluding F4 lowers the figure to 63/300 = 21.0%, which materially changes the claim. The paper should define whether problem-solving qualifies as capability-relevant and, if so, why 'generic suggestion-giving' is still a subset of it; alternatively, separate 'generic suggestion' from 'capability-relevant problem solving' and recompute all headline percentages.","section":"'What the Audit Establishes'; Table 5; Figure 2d"},{"comment":"The reliability of the ESConv function coding is moderate: fine-grained Cohen's kappa ranges 0.547–0.751 (mean 0.631), paradigm-level kappa 0.646–0.845 (mean 0.716), and 26 of 57 disagreements cross the relief/capability/process boundary. Since the entire corpus claim depends on the F3–F7 grouping and on label quality, relying solely on LLM coders without human construct validation is a genuine limitation. A small human-coded subset (for example, 40–60 turns with adjudication) should be reported before the 43.0% figure is used as a motivational anchor.","section":"Appendix C.2"}],"minor_comments":[{"comment":"In the reference list, 'Zao-Sanders, M.; Hill, K.; New; Freitas, J. D.; ...' appears to have a missing author name after 'New.' Please verify.","section":"References"},{"comment":"The claim-traceability table is useful, but the 'source artifact' entries are symbolic labels. In a reproducibility-focused paper, consider providing exact file names, paths, and checksums in the artifact manifest, at least for the frozen arXiv records and the 300-turn coding file.","section":"Appendix D.1 / Table 7"},{"comment":"The four-part illustration is information-dense; in the printed version, the small text under 'Support strategy intensity' and the lifecycle stages may be hard to read. A vector version with larger fonts or a separate table would improve clarity.","section":"Figure 3"},{"comment":"The Lagrangian signs are consistent with the inequality directions, but it would help to state explicitly that λ_i are multipliers for the inequalities Dep≤δ, Aut≥α0, and Soc≥σ0, since the sign convention differs from the usual 'all constraints ≤0' form and may confuse readers.","section":"Equation (14)"}],"recommendation":"major_revision","confidential_remarks":"The empirical audit is the weak link. The paper is honest about limitations, but the abstract and main-text 'gap' statements outrun the evidence. The ESConv F4 grouping issue is easily fixable and changes a headline number, so it should be addressed before acceptance. If the authors add a modest full-text/human validation subsample or re-scope the language, the paper would be a solid conceptual contribution well within the journal's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper proposes making capability-sustaining support across repeated use the organizing goal for emotional dialogue research, and it backs the proposal with a small audit and a clearly labeled formal sketch. The conceptual contribution is real: pulling together attachment, overreliance, termination loss, and autonomy into one framework called CSED, with six design commitments, four evaluation timescales, and lifecycle constraints. That is a coherent agenda paper, not a breakthrough result.\n\nWhat it does well: the audit is transparent and reproducible in principle. Exact counts, coding rules, reliability coefficients, and a released protocol are all there. The arithmetic checks out, and the limitations are stated honestly in the appendix. The formalization is also handled with restraint — it is illustrative, and Remark 1 shows existing benchmark objectives are nested cases, which is a fair and useful way to position the framework. The 43% capability-relevant function share in ESConv is a genuine data point that people will want to build on.\n\nSoft spots: the load-bearing claim is the audit's 95% relief / 0% longitudinal / 1-in-60 risk finding. That rests on 91 arXiv records coded from titles and abstracts by LLMs, with paradigm-level kappa around 0.72. A single system-building paper with capability evaluation buried in the full text but not the abstract would dent the 0/91, and relevant work in non-arXiv venues is invisible to the sample. The authors explicitly scope their claims to the sample, but the abstract and intro generalize to 'the field.' That gap between the evidence and the motivating narrative is the main weakness.\n\nAlso, the formal model depends on Assumption 1 — measurable noisy proxies for latent capability — which is cited but not validated here. The authors know this, and they frame the model as a research agenda, but it means the evaluation stack is a proposal, not a working instrument. The ESConv function mapping is a judgment call; fine-grained reliability is moderate, so the specific 4% reappraisal and 0.3% boundary numbers should be handled with care.\n\nNone of this is fatal. The paper is internally consistent, careful with its limitations, and the case for longitudinal evaluation is sound. The issues are about the strength of the empirical foundation relative to the scope of the framing, not about the framework's value.\n\nWho should read it: anyone building emotional support systems, evaluating mental-health chatbots, or writing governance for AI companions. It deserves a serious referee. I would send it out, with a request that the authors either temper the field-level claims or add human coding of a subset and a full-text check. Either way, I'd engage with it.","headline":"A useful roadmap for longitudinal emotional dialogue, with an honest audit whose sample limits the strength of the field-level gap claim.","tokens_in":18782,"tokens_out":1905,"would_cite":true,"duration_ms":19915,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that emotional-dialogue research has concentrated on immediate relief and neglected whether support sustains users' capacities across repeated use, and it proposes a new longitudinal paradigm to close that gap.","keywords":["emotional support dialogue","empathetic dialogue","capability-sustaining emotional dialogue","longitudinal evaluation","user capability","dependency risk","AI companion safety","termination and handoff"],"falsifier":"Run the audit on a full-text, multi-database sample with human coders: if a non-trivial share of system-building papers reports capability outcomes or longitudinal evaluation, the field-level gap claim fails. Separately, a preregistered longitudinal trial comparing a relief-only policy with a CSED-constrained policy would falsify the central benefit if capability gains do not differ while relief stays comparable.","tokens_in":17876,"feed_emoji":"💬","tokens_out":4885,"duration_ms":42115,"temperature":0.7,"pith_summary":"The paper argues that emotional-dialogue research has optimized two things: feeling understood (empathetic dialogue) and feeling better in the moment (emotional support conversation). For users who return to a system over weeks or months, the paper claims a third goal matters more: whether the system sustains the user's own capacities to regulate emotion, cope actively, make self-endorsed decisions, and stay socially connected. A targeted audit of 91 sampled papers finds 95% of system-building work pursues relief, none measures capability outcomes, and none evaluates longitudinally; a turn-level analysis of a standard support corpus finds capability-relevant functions are uneven, with generic suggestions far outpacing reappraisal and self-efficacy support. On this evidence, the paper proposes CSED—capability-sustaining emotional dialogue—as a longitudinal paradigm with six design commitments, four evaluation timescales, and explicit lifecycle constraints, alongside a formal process model that traces transient emotion and latent capability across repeated sessions.","feed_headline":"95% of emotional-dialogue systems chase relief, not capability","feed_subtitle":"A longitudinal paradigm argues support must preserve users' coping, choice, and social ties across repeated use and shutdown.","key_machinery":"The load-bearing object is the latent capability state c_k = (c_reg, c_cop, c_aut, c_soc) carried across sessions, evolving through an unknown transition kernel that takes dialogue as one input among stressors and offline life. Because capability is latent, the framework rests on Assumption 1—noisy, validated proxies m_k = M(s_k) + eta_k—which connects the formal model to evaluation. From that state, the paper derives four evaluation timescales (response, conversation, longitudinal, termination) and three lifecycle constraints (dependency exposure, autonomy preservation, social-connectedness drift), and shows that standard empathetic and support benchmarks are parameter-restricted cases of t","core_discovery":"The central discovery is the documented gap: the field's objectives and evaluation horizons stop at the session, so system behavior that helps now but erodes capability over time is invisible. CSED makes the full longitudinal interaction—repeated sessions, non-use, re-engagement, transition, termination—the unit of inquiry, and defines success as effective support plus sustained user capability in regulation, coping, autonomy, and social connectedness. The paper formalizes this with a latent capability vector c_k, a transition kernel Phi(s_k, d_k, epsilon_k), noisy proxies m_k = M(s_k) + eta_k, and a four-horizon objective subject to constraints on dependency, autonomy, and connectedness dri","pith_inferences":["The audit's proportions (95% relief, 0 longitudinal) are about a small, specific sample—91 title-and-abstract preprints coded by AI. If the same coding is applied to full text and multiple databases, the numbers may shift; the paper's own released protocol invites exactly that test.","CSED predicts a 'comfort trap': systems that score high on immediate preference and attachment can show flat or declining capability and rising dependency. Measuring the share of regulation episodes routed to the system over time would give a direct test.","The autonomy constraint implies that user preference is not a sufficient success signal; a system can be preferred and disempowering. That reframes pairwise human preference as only one term in a constrained objective.","The formal framework is implementable only once Assumption 1's instruments exist; the next bottleneck is validated in-situ measurement of emotion-regulation, coping, autonomy, and connectedness, not better generation models."],"forward_implications":["Evaluation must add longitudinal and termination horizons; otherwise a policy can help in the session while degrading capability over time, and that harm stays invisible.","Common benchmarks—response-level empathy and session-level emotional change—become special cases of the CSED objective, so the paradigm does not discard existing work but recontextualizes it.","Policy design should compare relief-only, resilience-activation, and state-conditioned policies under shared safety constraints, expecting preference-optimized systems to violate dependency and autonomy thresholds more often.","Data collection should record repeated measures, non-use, action initiator, revision of system guidance, and contact with human support, so capability transfer beyond the dialogue is observable.","Termination and model change become designed parts of support: forewarning, closure, memory dignity, and transfer readiness are measurable and tunable."],"fun_headline_variants":["95% of emotional-dialogue systems ignore capability: audit finds none evaluate it","Emotional support AI: from relief to sustained capability – new paradigm","Capability-sustaining dialogue: the missing metric in emotional AI","Longitudinal emotional dialogue: keeping users resilient, not just soothed","Audit: 95% of emotional-dialogue systems target relief, zero target capability"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The audit—91 title-and-abstract preprints sampled by year and coded by AI—is assumed to represent the state of system-building emotional-dialogue research; if that sample is unrepresentative, the motivating gap weakens, and the formal model additionally assumes latent capability can be measured through noisy proxies, which is cited but not validated here.","fun_headline_variants_meta":{"raw":{"variants":["95% of emotional-dialogue systems ignore capability: audit finds none evaluate it","Emotional support AI: from relief to sustained capability – new paradigm","Capability-sustaining dialogue: the missing metric in emotional AI","Longitudinal emotional dialogue: keeping users resilient, not just soothed","Audit: 95% of emotional-dialogue systems target relief, zero target capability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00157,"raw_usage":{"total_tokens":6132,"prompt_tokens":796,"completion_tokens":5336,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":540,"completion_tokens_details":{"reasoning_tokens":5237}},"tokens_in":540,"tokens_out":5336,"duration_ms":33844,"temperature":1.0,"reasoning_tokens":5237,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T00:15:22.749637+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the audit on a full-text, multi-database sample with human coders: if a non-trivial share of system-building papers reports capability outcomes or longitudinal evaluation, the field-level gap claim fails. Separately, a preregistered longitudinal trial comparing a relief-only policy with a CSED-constrained policy would falsify the central benefit if capability gains do not differ while relief stays comparable.","supporting_citations":[],"review_version":1}