{"id":"dc5ed767-2b27-4a11-9b91-a6e4265b3be6","arxiv_id":"2509.04465","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"LLM emotion labeling of dispute dialogues explains up to ~40% of variance in subjective outcomes (vs ~5% in prior negotiation work) and reveals anger escalation and compassion de-escalation patterns.","lead":"This paper tests whether large language models can read emotions in online buyer-seller dispute chats, and whether those emotion readings predict how satisfied people feel about the outcome. GPT-4o emotion labels, fed into simple regressions, explain up to roughly 40% of the variance in subjective outcomes, far more than earlier models.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LLM emotion labels may capture general negativity rather than distinct emotions, confounding the high R² and emotion-specific findings.","rationale":"The reader's weakest assumption flags exactly this construct-validity concern. I agree. The paper's validation with self-reported frustration shows strong correlations, but these are consistent with a general negativity artifact. A simple regression controlling for a global tone factor would discriminate between the two interpretations. If the increment is non-significant, the emotion labels are not measuring narrow emotions as claimed. This would not necessarily invalidate the practical observation that negative tone predicts dissatisfaction, but it would undermine the specific emotion findings and the 'explain' language. Given that the reader already set CONDITIONAL, my read does not change the verdict; the proposed test is necessary to move toward ACCEPT. Therefore, verdict_should_be = UNCHANGED.","tokens_in":14175,"tokens_out":5085,"duration_ms":50498,"concrete_test":"Re-run the SVI regressions (Section V-B) with a general tone covariate—e.g., the average intensity of anger, fear, and sadness, or the first principal component of all seven emotion channels—as a baseline. Then test whether the specific emotion channels (especially Anger and Compassion) add significant incremental R² beyond this baseline. If they do not, the high R² and the emotion-specific interpretations reflect global negativity rather than distinct emotional expressions.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central quantitative claim—that LLM-based emotion intensities explain over 40% of variance in subjective dispute outcomes—assumes these labels measure distinct emotional states, not global negativity or adversarial tone. The validation for this assumption is thin: a 2-item self-reported frustration scale (Table II) and human annotations on 100 utterances (Section VI-B, Fig. 5). The correlation pattern in Table II (anger +0.54, fear +0.36, joy –0.36, compassion –0.18 with frustration) is exactly what one would expect from a single negativity/sentiment factor. No analysis tests whether specific emotion channels add incremental predictive value beyond a general tone variable in explaining SVI. If they do not, the reported R² is driven by overall negativity, not by the specific emotions the paper interprets (anger spirals, compassion pathways), and the comparison to 5% in negotiation research is misleading. This is a construct-validity threat to the core contribution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes 2,025 buyer–seller dispute dialogues from the KODIS corpus, using GPT-4o (and other LLMs) to produce soft emotion-intensity labels for each utterance. Compared with T5-Twitter, the LLM labels correlate better with self-reported frustration and with third-party human annotations on a small sample. The authors then use multiple linear regression to predict the four Subjective Value Inventory (SVI) subscales from dialogue-level mean emotion intensities, reporting R2 values up to roughly 0.4, and they examine turn-level trajectories of anger and compassion for disputes ending in impasse versus resolution. The central claims are that automatically recognized emotions explain substantially more variance in dispute outcomes than in prior negotiation research and that the trajectories reveal an anger-escalation spiral and a compassion pathway to resolution.","tokens_in":14376,"tokens_out":4517,"duration_ms":49464,"significance":"If the central estimate survives proper evaluation, the result is substantial: prior negotiation studies reported roughly 5% variance explained by recognized emotion, whereas this paper reports up to ~40% in disputes, and it applies in a domain—dispute resolution—that is understudied in the affective-computing literature. The paper also demonstrates a methodological advantage of context-aware LLM annotation over a fine-tuned T5 baseline, with machine-checkable comparisons across multiple LLMs and human annotations. These strengths make the contribution potentially valuable for emotion-aware agent design and for social-science theories of conflict escalation. However, the quantitative claims currently rest on in-sample R2 values, a prompt-selection procedure that uses the evaluation subset, and a construct-validity argument that has not ruled out a single negativity/positivity factor.","major_comments":[{"comment":"The headline R2 values are computed in-sample, and the prompt configuration is selected on the same evaluation subset. The text says 'We perform these tests on a 20% subset (N = 406) of the corpus' and Table III then reports mean R2 for the ablated configurations, with the best GPT4o configuration used in Fig. 3. Since each model has 7 emotion-intensity predictors (per side) and there is no held-out split, cross-validation, or significance testing, the reported R2—and especially the improvement over T5—can reflect overfitting and selection. Because the abstract and Discussion rely on the 'over 40% of the variance' claim, please report out-of-sample R2 (e.g., 5-fold CV, or train on the remaining 80% and test on the 20%), separate prompt selection from evaluation, and provide confidence intervals or at least standard errors.","section":"§V-B, Table III, Fig. 3"},{"comment":"The validation evidence does not establish that the LLM emotion labels isolate distinct emotional constructs rather than a single valence/negativity factor. The correlation pattern in Table II—anger +0.54, fear +0.36, joy −0.36, compassion −0.18 with self-reported frustration—is exactly what a unidimensional negativity factor would predict. The SVI regressions in Fig. 3 may therefore be driven by overall adversarial tone rather than by specific emotions. To support the specific anger-spiral and compassion-pathway claims, please test incremental validity: do anger and compassion add significant variance to SVI after controlling for an aggregate negative-tone index (e.g., mean intensity of negative labels or a lexical sentiment score)? Additionally, the human-annotation benchmark uses only 100 utterances with no inter-annotator agreement or per-emotion confidence intervals in Fig. 5; repor","section":"§V-A, Table II; §VI-B, Fig. 5"},{"comment":"The escalation-spiral and compassion-pathway conclusions are based on raw mean trajectories with no inferential test. The text states that sellers 'reciprocate the buyer's anger' in impasse dialogues and 'stick to the script' in resolved dialogues, but no statistical comparison of the trajectories (e.g., an outcome × turn interaction in a mixed-effects model, or per-timepoint tests with multiple-comparison correction) is reported. Given that the labels are themselves generated by GPT-4o, the trajectory differences could reflect annotation artifacts. Please add formal comparisons of the trajectory shapes and report uncertainty around the means.","section":"§V-C, Fig. 4a–4b"}],"minor_comments":[{"comment":"Clarify whether the 19% impasse rate is dyad-level or participant-level; the Discussion says '19% of participants failed to reach an agreement,' which is a different statement.","section":"§III-C"},{"comment":"The checkmark columns in Table III are hard to parse. Define what each column corresponds to (IC learning, compassion label, dialogue history, neutral label) and describe the in-context examples: who hand-annotated them, how many, and whether they are from the same corpus.","section":"§IV-B, Table III"},{"comment":"Report the number of annotators per utterance, the total number of annotations (the text says N=336, presumably annotators), and the inter-annotator agreement. Also clarify that 'average human' is the mean over annotations, and add error bars to the correlation coefficients.","section":"§VI-B, Fig. 5"},{"comment":"The abstract says the paper investigates 'subjective and objective outcomes,' but the regression analysis only predicts subjective SVI scores; objective resolution (impasse vs. resolution) appears only in descriptive trajectory plots. Please align the wording with what is actually estimated.","section":"Abstract and §VII"},{"comment":"The dialogue snippet contains 'traud'—if this is meant to be 'fraud,' correct it in the figure; if intentional (as a non-word), mark it as such.","section":"Fig. 1"},{"comment":"Typo: 'University of Southern Calirofnia' should be 'University of Southern California.'","section":"Author affiliations"},{"comment":"The sentence 'we do not find much evidence of interactions between those demographic factors' is not supported by any reported analysis. Either include the analysis or remove the claim.","section":"§VIII"}],"recommendation":"major_revision","confidential_remarks":"The paper is a good fit for this venue and the corpus is valuable. The main risk is overinterpretation of in-sample R2 and of model-generated emotion labels as distinct constructs. The requested fixes (out-of-sample evaluation, incremental-validity controls, and trajectory inference) are feasible within the scope of the manuscript; they do not require new data collection. One additional editorial concern: the KODIS corpus is described as previously published (ref. [51]), and the present paper should make clear which analyses are new relative to that corpus paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real news here is that LLM-based emotion intensity labels, applied to a large buyer-seller dispute corpus, explain a striking share of variance in how people feel about the outcome—up to 40% on some SVI subscales. That is a much bigger signal than the ~5% seen in prior negotiation work, and the paper is the first to bring this kind of analysis to text disputes. The trajectory plots (anger spirals, early compassion) are also genuinely suggestive and align with existing conflict theory. Credit where due: the KODIS corpus is large, the comparison against T5 is a fair baseline, and the human annotation benchmark, though small, is a reasonable external check.\n\nThe soft spots are real but not disqualifying. First, the headline R² is computed in-sample on the 20% subset used to select the best prompt configuration. That is a recipe for optimistic estimates, and the paper reports no cross-validation, no adjusted R², and no confidence intervals for the regressions. The \"over 40%\" figure is also a maximum subscale value, not a typical result; the Discussion's \"thirty to forty percent\" is closer to the truth but still lacks uncertainty bounds. Second, the construct validity concern is legitimate. The correlation pattern in Table II (anger +0.54, fear +0.36, joy −0.36, compassion −0.18 with self-reported frustration) is exactly what a general negativity factor would produce. The paper never tests whether specific emotion channels add incremental predictive power beyond a single valence/tone variable. If they do not, the emotion-specific interpretations—anger reciprocation, compassion pathway—could be artifacts of overall adversarial tone. That is not a fatal flaw, because the descriptive trajectories are still interesting, but it does undercut the more precise claims.\n\nMinor issues: the 100-utterance human annotation set is thin, inter-annotator agreement is not reported, and the authors do not release prompts, in-context examples, or the annotation code. All of that should be required for reproducibility.\n\nBottom line: this is a useful, well-motivated paper for affective computing, conflict resolution, and computational social science. The central quantitative claim is not yet credible as stated, but the domain, data, and analysis framework are worth engaging with. I would send it to peer review and ask for out-of-sample evaluation, a valence-only baseline, significance testing, and full prompt transparency before accepting. A serious referee could turn this into a solid contribution.","headline":"LLM emotion labels explain a lot of variance in dispute outcomes—but the headline R² is in-sample, prompt-tuned on the test subset, and not clearly distinct from a general negativity factor.","tokens_in":14857,"tokens_out":2275,"would_cite":true,"duration_ms":26933,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Emotion intensities generated by GPT-4o from dispute text explain more than 40% of the variance in how participants feel about the process and their relationship, far exceeding prior negotiation-based emotion recognition.","keywords":["dispute resolution","emotion recognition","large language models","affective computing","subjective value","anger escalation","compassion","negotiation"],"falsifier":"Give the same regression pipeline a single generic negativity or adversarial-tone score per dialogue, such as a lexical hostility measure or a human rating of how negative the conversation was, alongside or instead of the seven emotion labels. If the generic score alone recovers most of the R-squared, or if adding it collapses the emotion coefficients, the specific emotion-reading claim is not supported. Alternatively, collect human soft-label emotion annotations on a large random sample of dialogues and run the same regressions: if human labels explain far less than GPT-4o's labels, the high","tokens_in":14071,"feed_emoji":"⚖️","tokens_out":8009,"duration_ms":76382,"temperature":0.7,"pith_summary":"Disputes evoke stronger emotions than negotiations, and anger can spiral into retaliation rather than compromise. The paper asks whether automatically recognized emotional expressions from text can predict how people experience a dispute's outcome. Using 2,025 buyer-seller disputes from the KODIS corpus, it shows that prompting GPT-4o to annotate each utterance with soft intensities over anger, joy, fear, sadness, surprise, compassion, and neutral yields labels that, averaged over a dialogue, explain in some cases more than 40% of the variance in the Subjective Value Inventory's four subscales, with buyers' feelings about process and relationship approaching half the variance. This is far above the roughly 5% explained by recognized anger in prior negotiation research, and the same pattern holds across several LLMs that also match human annotators better than the T5 baseline. The labels also trace two dynamics: seller reciprocation of buyer anger tracks impasse, while early seller compassion and buyer reciprocation track resolution.","feed_headline":"LLM emotion reads explain 40% of dispute-outcome feelings","feed_subtitle":"AI-annotated anger and compassion predict how satisfied disputants feel, far beyond prior negotiation results.","key_machinery":"The machinery is LLM-based utterance-level emotion intensity annotation followed by linear regression. The paper prompts GPT-4o with dialogue history, a seven-label scheme replacing 'love' with 'compassion' and adding 'neutral', soft labels summing to one, and a small set of in-context examples to score each turn; these intensity vectors are averaged over a dialogue and used as predictors of each SVI subscale. The R-squared of those regressions is the paper's measure of explanatory power. For dynamics, the same annotations are averaged by role and turn to trace anger and compassion trajectories in resolved versus impasse dialogues.","core_discovery":"The central claim is that automatically recognized emotional expression is a measurable and substantial correlate of dispute outcomes. Concretely, the paper reports that when GPT-4o is prompted to produce per-utterance soft emotion-intensity vectors, including dialogue context, a label set of anger, joy, fear, sadness, surprise, compassion, and neutral, and a few in-context examples, the dialogue-level averages predict a disputant's subjective outcome: linear regressions on these labels explain, in some cases, over 40% of the variance in SVI subscales, with buyers' feelings about process and relationship approaching half the variance. The paper further claims that this explanatory power is n","pith_inferences":["A control regression that adds a generic negativity or adversarial-tone score alongside the seven emotion labels would show whether the specific emotion labels carry distinct information; the paper does not include such a control.","The causal reading of the emotion-outcome link would require an intervention experiment, such as randomly assigning an AI mediator to encourage compassion in one arm; the correlational design alone cannot rule out that emotions merely mirror an underlying dispute quality.","The claim may transfer to other dispute settings such as workplace, family, or legal mediation, but the corpus is a single online role-play scenario in English, so cross-setting generalization is untested.","An early-warning system that reads only the first few turns is a concrete design implication: the reported trajectories separate impasse from resolution almost immediately after the first round."],"forward_implications":["Emotion labels alone, without dialogue content, can flag high-risk disputes early: dialogues that end in impasse show anger reciprocation within the first few turns.","Mediation agents could use such labels to detect escalation in real time and intervene to defuse anger or encourage compassion.","Label design matters: adding a neutral category and swapping love for compassion improved predictive performance, so downstream tasks should shape emotion-taxonomy choices in affective computing.","LLM-based emotion recognition outperforms the fine-tuned T5 baseline both in matching human annotations and in explaining subjective outcomes across four LLMs, making it a ready replacement in conflict research pipelines.","Disputes should be treated as a distinct genre from negotiations in emotion research: impasse rates and emotion-outcome correlations differ sharply."],"supporting_citations":[{"why":"Supplies the KODIS corpus of 2,025 buyer-seller dispute dialogues and associated self-report measures.","marker":"[51]"},{"why":"Supplies the T5-Twitter emotion-recognition baseline and the prior negotiation finding of about 5% explained variance that this paper extends and outperforms.","marker":"[25]"},{"why":"Supplies the Subjective Value Inventory, the four-subscale outcome measure used as the dependent variable.","marker":"[54]"},{"why":"Supplies the CaSiNo negotiation corpus and the prior collection framework, including the 3% impasse-rate comparison point.","marker":"[31]"},{"why":"Supplies the conflict-escalation theory predicting that anger reciprocation produces an escalatory spiral, which the turn-by-turn analysis tests and replicates.","marker":"[20]"},{"why":"Supplies prior evidence that anger and compassion influence negotiation performance, grounding the compassion pathway result.","marker":"[56]"},{"why":"Supplies the frustration subscale from the tactic questionnaire used to validate the emotion labels against self-report.","marker":"[52]"},{"why":"Supplies the comparison of LLMs on negotiation dialogue inferences, used to justify selecting GPT-4o as the primary annotation model.","marker":"[55]"}],"fun_headline_variants":["AI emotion reads explain 40% of dispute satisfaction","GPT-4o emotion labels forecast dispute outcomes","Emotion-aware AI predicts how disputes feel","LLM emotion analysis explains dispute satisfaction","AI reads emotions in disputes, predicts outcomes"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The claim stands or falls on whether GPT-4o's soft emotion labels measure the emotions actually expressed, rather than an overall adversarial tone or generic negativity that happens to correlate with bad outcomes.","fun_headline_variants_meta":{"raw":{"variants":["AI emotion reads explain 40% of dispute satisfaction","GPT-4o emotion labels forecast dispute outcomes","Emotion-aware AI predicts how disputes feel","LLM emotion analysis explains dispute satisfaction","AI reads emotions in disputes, predicts outcomes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000267,"raw_usage":{"total_tokens":1400,"prompt_tokens":639,"completion_tokens":761,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":383,"completion_tokens_details":{"reasoning_tokens":707}},"tokens_in":383,"tokens_out":761,"duration_ms":7723,"temperature":1.0,"reasoning_tokens":707,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T14:26:14.854549+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give the same regression pipeline a single generic negativity or adversarial-tone score per dialogue, such as a lexical hostility measure or a human rating of how negative the conversation was, alongside or instead of the seven emotion labels. If the generic score alone recovers most of the R-squared, or if adding it collapses the emotion coefficients, the specific emotion-reading claim is not supported. Alternatively, collect human soft-label emotion annotations on a large random sample of dialogues and run the same regressions: if human labels explain far less than GPT-4o's labels, the high","supporting_citations":[{"cited_title":"Kodis: A multicultural dispute resolution dialogue corpus,","cited_arxiv_id":null,"evidence_quote":"Supplies the KODIS corpus of 2,025 buyer-seller dispute dialogues and associated self-report measures."},{"cited_title":"To- wards emotion-aware agents for improved user satisfaction and partner perception in negotiation dialogues,","cited_arxiv_id":null,"evidence_quote":"Supplies the T5-Twitter emotion-recognition baseline and the prior negotiation finding of about 5% explained variance that this paper extends and outperforms."},{"cited_title":"What do people value when they negotiate? mapping the domain of subjective value in negotiation","cited_arxiv_id":null,"evidence_quote":"Supplies the Subjective Value Inventory, the four-subscale outcome measure used as the dependent variable."},{"cited_title":"Casino: A corpus of campsite negotiation dialogues for automatic ne- gotiation systems,","cited_arxiv_id":null,"evidence_quote":"Supplies the CaSiNo negotiation corpus and the prior collection framework, including the 3% impasse-rate comparison point."},{"cited_title":"Conflict escalation in organizations,","cited_arxiv_id":null,"evidence_quote":"Supplies the conflict-escalation theory predicting that anger reciprocation produces an escalatory spiral, which the turn-by-turn analysis tests and replicates."},{"cited_title":"The influence of anger and compassion on negotiation performance,","cited_arxiv_id":null,"evidence_quote":"Supplies prior evidence that anger and compassion influence negotiation performance, grounding the compassion pathway result."},{"cited_title":"Dignity, face, and honor cultures: A study of negotiation strategy and outcomes in three cultures,","cited_arxiv_id":null,"evidence_quote":"Supplies the frustration subscale from the tactic questionnaire used to validate the emotion labels against self-report."},{"cited_title":"Are llms effective negotiators? systematic evaluation of the multifaceted capabilities of llms in negotiation dialogues,","cited_arxiv_id":null,"evidence_quote":"Supplies the comparison of LLMs on negotiation dialogue inferences, used to justify selecting GPT-4o as the primary annotation model."}],"review_version":1}