{"id":"305a0b90-5949-43f7-83e7-92db0851a8a8","arxiv_id":"2602.12450","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"LLM-based scoring of 1,610 therapy sessions finds therapist empathy and exploration are followed by more client disclosure, while prior-session rapport is associated with less self-directed negative emotion—but the claimed rapport moderation is never modeled.","lead":"The authors used large language models to score transcripts of roughly 1,600 therapy sessions for empathy, rapport, disclosure, and emotion, then modeled how therapist actions relate to the client's next response. It matters because it points toward automatic, large-scale measurement of therapy process that could support training and AI mental-health tools.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline claim that rapport is a moderator is not estimable from the reported SEM: Table 7 contains only main effects, no Rapport×Empathy or Rapport×Exploration interactions, and §6.4 defers moderation tests to future work.","rationale":"The reader's weakest assumption—low LLM–human agreement on internalizing emotions—is a genuine validity threat, but it is not the most load-bearing gap for the paper's signature claim. Even granting perfect measurement, the reported SEM cannot identify moderation because interaction terms are absent: Table 7 lists only main effects, and §6.4 explicitly defers 'moderation more directly' to future work. The Abstract and Conclusion nevertheless assert that rapport 'moderates' and 'conditions' the effects of empathy and exploration. A main effect of Rapport on self-directed negative emotion (β=−0.02) is not evidence of moderation. This is an internal inconsistency between the fitted model and the headline conclusion, not a matter of interpretive convention. The reader did note 'No interaction terms are shown' in the rationale, so agreement is partial, but their formal weakest_assumption centered on ICCs rather than on the absent moderation test. I keep the REJECT verdict: the missing interaction analysis is central, not cosmetic, and the measurement concerns also remain unresolved.","tokens_in":25900,"tokens_out":6611,"duration_ms":62114,"concrete_test":"Re-fit the SEM in §5.1 with latent or observed interaction terms Rapport×Empathy and Rapport×Exploration predicting all four client outcomes (e.g., product-indicator or latent moderated structural equations in lavaan), and include the lagged prior client-state covariate promised in the Abstract. Report interaction coefficients with confidence intervals and a nested model comparison with vs. without interactions. If the interactions are nonsignificant or inconsistent with the claimed moderating direction, the central claim should be revised to a direct-effects claim only.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—'empathy and exploration acting directly and rapport as a contextual moderator'—requires an interaction between session-level rapport and the two therapist-behavior variables in the SEM. The model reported in §5.1/Table 7 includes only main-effect predictors: log session ID, Rapport, Exploration, and Empathy. No Rapport×Empathy or Rapport×Exploration term is defined, estimated, or discussed; Figure 3 shows only direct arrows from Rapport to outcomes. A main effect of Rapport (e.g., β=−0.02, p=.05, on self-directed negative emotion) cannot tell us that rapport 'dampens empathy's association' or 'conditions the impact of exploration,' as the Conclusion claims. This is not a subtle inferential gap: the paper's own Future Work section (§6.4) says 'future work should test sequential mediation and moderation more directly,' conceding that the moderation model was not actually fit. The ICC issues with internalizing emotions (Table 5) are a separate validity threat; the moderation gap would remain even if every measure were perfectly reliable. Thus the distinguishing portion of the headline claim is unsupported by the reported analysis.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper develops an LLM-based measurement pipeline for psychotherapy process constructs—therapist empathy (Emotional Reaction, Interpretation, Exploration, Reflection), self-disclosure, nine emotions, and session-level rapport—and validates the scores against human annotations. It then applies the pipeline to roughly 243k utterances from 1,610 Alexander Street sessions and fits a multivariate SEM with utterance-level therapist behaviors and session-level rapport as predictors of next-turn client disclosure and emotional expression. The reported main effects suggest empathy and exploration predict disclosure and emotional expression, and prior-session rapport is negatively associated with self-directed negative emotion. The abstract and conclusion additionally claim that rapport moderates the effects of therapist empathy and exploration on client affect.","tokens_in":26220,"tokens_out":7320,"duration_ms":67568,"significance":"If the measurement-validation results hold, the framework is a useful contribution: the authors provide detailed prompt rubrics, a human-validation protocol with ICC reporting, and a large transcript corpus. This could support scalable psychotherapy process research. However, as reported, the SEM does not estimate the moderation interactions that the headline claim requires, and the low LLM–human reliability of the internalizing emotion indicators threatens the most distinctive emotion findings. The paper would need substantial reanalysis—not just copyediting—to support its central claims.","major_comments":[{"comment":"The abstract and §7 claim that rapport is a contextual moderator that 'dampens empathy’s association with negative emotions and conditions the impact of exploration.' However, Table 7 contains only main effects (log session ID, Rapport, Exploration, Empathy); no Rapport×Empathy or Rapport×Exploration interaction is defined or estimated, and Figure 3 shows only direct arrows. §6.4 itself states that 'future work should test sequential mediation and moderation more directly,' conceding that moderation was not actually tested. This is not a wording issue: the distinguishing part of the headline claim is not identified by the reported analysis.","section":"§5.1, Table 7, §7, Abstract"},{"comment":"The self-directed negative emotion factor is built from fear, anxiety, sadness, and depression plus reverse enjoyment. In Table 5, the LLM–human ICC 95% confidence intervals for these four internalizing emotions all include zero (Fear 0.454 [-0.13,0.71]; Anxiety 0.460 [-0.19,0.74]; Sadness 0.526 [-0.16,0.77]; Depression 0.500 [-0.06,0.72]). The key emotion findings—empathy’s larger association with self-directed than outward-directed negative emotion (β=0.12 vs 0.03) and rapport’s negative association (β=-0.02)—depend on this composite. No sensitivity analysis excluding or downweighting these low-ICC indicators is reported. The acknowledgment in §4.4.2 does not address the impact on the SEM estimates.","section":"§4.4.2, Table 5, §5"},{"comment":"The abstract states that the SEM estimates moment-to-moment relationships 'controlling for prior client state and context,' but the predictor set in Table 7 includes only log session ID, Rapport, Exploration, and Empathy. There is no lagged client outcome (e.g., previous-turn disclosure or previous-turn emotion) in the model. Without such a control, the reported associations may be confounded by client-level trends and autocorrelation, and the causal language in §7 ('exert immediate effects') is not justified.","section":"Abstract, §5.1, Table 7"},{"comment":"The manuscript gives two contradictory PCA descriptions for the emotion items. The first says 'Across nine client emotion indicators (N=122,939), PCA produced a three-component structure explaining 71.5%' with surprise as the third component. The following paragraph says 'PCA yielded a two-component solution that explained about 59% of the variance' and that surprise 'did not relate meaningfully to either component (uniqueness U2=.97).' Since Table 7 and Figure 3 treat surprise as a separate outcome, the three-component solution appears to be the one used, but the text must be reconciled. This inconsistency undermines the reproducibility of the factor construction.","section":"§5, PCA"}],"minor_comments":[{"comment":"The caption says 'rapport was associated with reduced negative emotions and a slight decrease in disclosure.' Table 7 shows the disclosure coefficient as 0.00 (p=.71), which is not a decrease. Please correct the caption.","section":"Figure 3 caption"},{"comment":"Describing an ICC of 0.45 as 'fair' is misleading when the 95% confidence interval includes zero. Report the CI explicitly in the text and avoid implying acceptable reliability.","section":"§4.4.1, Table 5"},{"comment":"Typo in the user prompt: 'herapist’s Response' should be 'therapist’s Response.'","section":"Appendix A.2"},{"comment":"The paper does not include a data or code availability statement. Given the detailed prompts and the reproducibility value of the pipeline, such a statement would strengthen the manuscript.","section":"General"},{"comment":"The abstract reports a mean Pearson r=.66, but the validation section uses mixed metrics (e.g., F1 for Reflection, ICC and r for others). Please clarify how this average is computed and over which constructs.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The gap between the abstract/conclusion moderation claim and the estimated main-effects model is severe; if the authors cannot supply interaction analyses and sensitivity analyses in revision, I would recommend rejection. The paper would be much stronger if reframed primarily as a measurement-validation contribution, with the SEM results reported as exploratory main effects."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the measurement work is real and worth building on, but the paper's central claim—that rapport moderates how empathy and exploration affect clients—is not actually estimated anywhere in the reported SEM, and the emotion PCA section contradicts itself. Don't publish this as-is; do read it if you work on automated therapy-process measurement.\n\nWhat's genuinely new: combining LLM-based construct scoring with SEM on roughly 1,600 sessions of psychotherapy transcripts, with human validation for each construct (Table 5), detailed prompts in the appendix, and a specific new empirical finding—empathy's association with self-directed negative emotion (β=0.12) is much larger than with outward-directed negative emotion (β=0.03), and prior-session rapport predicts less self-directed negative emotion. The measurement stack is legible and reproducible in spirit, and the authors disclose the internalizing-emotion reliability problems honestly in §4.4.2 even though they proceed to use those scores anyway.\n\nThe soft spots are real. First, the abstract and conclusion claim rapport acts as a moderator, but Table 7 contains only main effects. There is no Rapport×Empathy or Rapport×Exploration term anywhere, and §6.4 says moderation should be tested in future work. That is not a minor misstatement; it is the distinguishing claim of the paper. Second, the reported SEM does not include the lagged prior client state that the abstract says it controls for. Third, the internalizing emotion measures—fear, sadness, anxiety, depression—have LLM–human ICCs around 0.45–0.53 with 95% CIs crossing zero, and they are the core of the self-directed negative emotion factor; without sensitivity analyses, the empathy asymmetry and the rapport-reduction result are fragile. Fourth, the PCA section reports a three-component solution and then, in the very next paragraph, a two-component solution; Table 6 shows three columns. That kind of leftover inconsistency needs cleaning.\n\nNone of this is fatal to the underlying project. The main-effect associations are cheap to re-estimate with interaction terms, lagged outcomes, and reliability corrections. This paper is for psychotherapy process researchers and HCI/mental-health people who care about LLM-based measurement, and it would benefit from a serious referee who asks for those analyses. I would not cite the current version, but I would want to see the revision after major revision.","headline":"Real measurement contribution, but the moderation headline is not in the model table and the emotion PCA section contradicts itself; worth a serious referee after major revision.","tokens_in":26657,"tokens_out":4943,"would_cite":false,"duration_ms":46829,"reading_group":"yes","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Using LLM ratings of 1,610 therapy sessions, this paper claims that moment-to-moment therapist empathy and exploration directly increase client disclosure and shape emotional expression, while rapport works as a contextual moderator rather","keywords":["psychotherapy process","LLM-based assessment","structural equation modeling","therapist empathy","therapeutic rapport","self-disclosure","negative emotion expression","turn-level analysis"],"falsifier":"Re-estimate the SEM using only emotion items with adequate LLM–human agreement (anger, contempt, disgust, enjoyment, surprise) or replace the LLM emotion scores with human-coded labels on a subset of sessions; if the empathy coefficient for self-directed negative emotion collapses toward zero, the asymmetry claim fails. Also test the interaction term rapport×empathy in the model; if it is not significant, the moderator claim is unsupported.","tokens_in":25813,"feed_emoji":"💬","tokens_out":3461,"duration_ms":33445,"temperature":0.7,"pith_summary":"The paper tries to establish that core psychotherapy processes can be measured automatically and modeled moment-to-moment from transcripts. It claims that therapist empathy and exploration in one turn directly increase client self-disclosure and shift emotional expression in the next turn, while rapport built in prior sessions does not directly amplify disclosure or emotions but changes how those behaviors matter. If right, this would make therapy process research scalable and give trainers, supervisors, and AI-based tools a concrete, turn-level signal to work from.","feed_headline":"Empathy, exploration move clients; rapport moderates","feed_subtitle":"Turn-level LLM ratings of 1,600+ therapy sessions show what actually shifts disclosure and emotion.","key_machinery":"The carrying mechanism is a two-step pipeline: (1) construct-specific LLM prompts that rate each therapist turn, client turn, and session segment on theory-grounded scales (EPITOME empathy, WAI bond items, intimacy-based disclosure, PANAS-like emotions), with human validation (mean Pearson r=.66); (2) a structural equation model with utterance-level variables (therapist behavior in prior turn → client outcome in current turn), session-level rapport from prior session, and controls for prior client state and session index. PCA collapses nine emotions into self-directed negative, outward-directed negative, and surprise factors.","core_discovery":"Across 1,610 sessions and about 243,000 utterances, the authors built LLM-based scores for therapist empathy (emotional reaction, interpretation, exploration, reflection), rapport (observer-rated bond), and client self-disclosure and nine emotions, validated them against human raters, and fitted a structural equation model. They found empathy (β=0.12) and exploration (β=0.10) predict self-directed negative emotions, empathy weakly predicts outward-directed emotions (β=0.03), both predict disclosure (β=0.09), and prior-session rapport predicts slightly lower self-directed negative emotion (β=-0.02) rather than more disclosure. The paper interprets this as a dual-route model: empathy and explo","pith_inferences":["The low LLM–human agreement for fear, anxiety, sadness, and depression (ICCs 0.45–0.53 with confidence intervals including zero) means the self-directed negative emotion factor—and therefore the empathy-asymmetry result—may rest on noisy measurements; a focused re-analysis with human-coded emotion labels would test this.","The paper claims rapport moderates associations, but the reported SEM shows only direct effects; an explicit empathy-by-rapport interaction term would be the natural falsifier.","Because rapport is observer-rated (WAI-O) rather than client-rated, the null direct effect may reflect who is doing the rating; client-rated alliance could behave differently.","The Alexander Street corpus is a single, professionally transcribed, Western collection; generalizing to other languages, modalities, or cultural display norms remains untested."],"forward_implications":["If correct, turn-level process modeling becomes feasible for thousands of sessions without human coding, enabling large-scale tests of therapy mechanisms.","Therapist training systems could give moment-by-moment feedback—'this empathic reflection is likely to increase disclosure'—grounded in the estimated paths.","The dual-route model predicts that empathy's effect on internal distress expression is larger than its effect on outward anger, a testable signature of supported emotional processing.","Designers of chatbot therapists could embed these contingencies to make simulated clients respond to trainee empathy and exploration in realistic ways.","Observer-rated rapport not directly increasing disclosure would push process research to distinguish client-perceived alliance from observer-rated bond."],"fun_headline_variants":["LLM ratings of 1,600 therapy sessions show empathy, exploration shift client affect","Empathy, exploration act directly; rapport moderates—1,600-session LLM study","Therapist empathy and exploration drive client disclosure and emotion—rapport moderates","Turn-level LLM ratings: empathy, exploration predict disclosure; rapport moderates"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The claim that empathy mainly increases self-directed rather than outward-directed negative emotion depends on LLM scores for fear, anxiety, sadness, and depression being meaningful measures, but their agreement with human raters is low enough that the 95% confidence interval includes zero.","fun_headline_variants_meta":{"raw":{"variants":["LLM ratings of 1,600 therapy sessions show empathy, exploration shift client affect","Empathy, exploration act directly; rapport moderates—1,600-session LLM study","Therapist empathy and exploration drive client disclosure and emotion—rapport moderates","Turn-level LLM ratings: empathy, exploration predict disclosure; rapport moderates"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001087,"raw_usage":{"total_tokens":4415,"prompt_tokens":813,"completion_tokens":3602,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":557,"completion_tokens_details":{"reasoning_tokens":3522}},"tokens_in":557,"tokens_out":3602,"duration_ms":24035,"temperature":1.0,"reasoning_tokens":3522,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T23:47:43.563773+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-estimate the SEM using only emotion items with adequate LLM–human agreement (anger, contempt, disgust, enjoyment, surprise) or replace the LLM emotion scores with human-coded labels on a subset of sessions; if the empathy coefficient for self-directed negative emotion collapses toward zero, the asymmetry claim fails. Also test the interaction term rapport×empathy in the model; if it is not significant, the moderator claim is unsupported.","supporting_citations":[],"review_version":1}