{"id":"3924949d-c3a9-4a10-abf2-3e6c5a139389","arxiv_id":"2504.12723","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A new multicultural corpus of dispute resolution dialogues, with initial evidence that expressed anger accompanies escalation and impasse.","lead":"KODIS is a new dataset of thousands of text conversations in which people from over 75 countries act out a customer service dispute. The paper's early analysis finds that more anger in these chats is linked to breakdowns in the conversation, and that emotional patterns vary by culture.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The initial analysis's emotion measurements rely on GPT-4o without cross-cultural validation, so the escalation and cultural-emotion claims in Section 3 are not yet established.","rationale":"The reader's weakest assumption identifies the absence of cross-culturally validated GPT-4o emotion annotations as the key risk, and the manuscript itself acknowledges this limitation. My stress-test confirms that the escalation and cultural-emotion findings in Section 3 stand or fall on the validity of these annotations, with no human ground truth available. I also note that Figure 4 lacks statistical tests and Table 2 contains a counterexample to the stated generalization claim, but these are secondary to the measurement validity problem. The proposed human-annotation study would directly test whether GPT-4o's emotion scores track human perception across the five focal cultures. Because the reader already issued a CONDITIONAL verdict conditioned on this gap, I see no reason to change the verdict; the concern is real but does not undermine the corpus's descriptive value, which is the paper's primary resource contribution.","tokens_in":14179,"tokens_out":5643,"duration_ms":58292,"concrete_test":"Select a stratified random sample of 200 human-human dialogues (40 each from US, UK, Canada, Mexico, South Africa). Recruit native or fluent annotators from each country to rate every utterance for perceived anger intensity on a 0–1 scale. Compute per-country and per-dialogue correlations (e.g., Spearman) between GPT-4o anger scores and human ratings. Then reproduce Figure 4 (anger by exchange number, impasse vs. agreement) and Figure 5 (mean anger by country, within vs. cross-culture) using only human anger labels, and test the escalation trajectory with a mixed-effects model (exchange number × outcome). If GPT-human agreement varies substantially by country, or if the cultural differences and escalation pattern disappear with human labels, the Section 3 conclusions are unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central empirical contribution is that anger expressions escalate in disputes and vary by culture, but every such claim in Section 3 depends on GPT-4o's per-utterance emotion scores. The only reported validation (Appendix A.1) is a correlation with self-reported frustration, collected once after the dispute and aggregated over all participants; it does not establish per-turn or per-culture validity. Section 5 explicitly concedes that there are no independent human annotations and notes that cross-cultural emotion recognition by large pretrained models is 'particularly fraught' (citing Havaldar et al., 2023). The escalation pattern in Figure 4 is presented without confidence intervals or significance tests, so even with valid labels the claimed trajectory is not statistically supported. The cultural transfer result in Table 2 is also internally inconsistent: for South Africa, the US-trained model yields higher R² (0.157) than the within-culture model (0.047), contradicting the claim that generalization degrades for dissimilar cultures. Because the abstract's central claim is that the initial analysis supports the escalation and cultural-difference theories, the absence of cross-culturally validated emotion measurement is the load-bearing weakness: if GPT-4o's anger scores are culturally biased, both Section 3 findings could be artifacts.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents KODIS, a dyadic dispute resolution dialogue corpus collected from over 4,000 participants across 75+ countries, using an online role-play scenario about a buyer-seller dispute. The corpus includes pre-dispute dispositional and preference measures, dialogue data (human-human and human-AI), and post-dispute outcomes such as subjective value, objective utility, and perspective-taking. The authors describe the collection framework, the task design, the culture clustering (Dignity/Face/Honor), and then report an initial analysis in which GPT-4o emotion scores on dialogue turns are used to study escalation, to predict subjective outcomes, and to examine cultural transfer. The abstract claims that this analysis supports theories of anger-driven escalatory spirals and reveals cultural differences in emotional expression. The paper also includes a limitations section acknowledging the absence of independent human annotation and the cross-cultural risks of LLM-based emotion measurement.","tokens_in":14373,"tokens_out":3166,"duration_ms":36374,"significance":"If the corpus is released as described, it is a genuinely useful community resource: it extends the study of negotiation dialogues to dispute resolution, includes both within- and cross-country dyads, integrates a rich set of self-report and outcome measures, and provides a detailed, reproducible collection framework. The comparison with the CaSiNo corpus and the replication of a higher impasse rate in disputes than in deal-making are valuable 'face validity' checks. However, the initial analysis that motivates the corpus is exploratory and currently rests on emotion labels from a single LLM model without independent human validation, and some of the reported results are internally inconsistent. The resource contribution is solid; the empirical claims about escalation and cultural variation need to be either strengthened with proper validation and statistical testing or scaled back.","major_comments":[{"comment":"The central escalation claim is not statistically supported as presented. Figure 4 plots mean anger scores by role, dialogue turn, and outcome (impasse vs. agreement) but provides no confidence intervals, error bars, or significance tests. The reader cannot rule out that the apparent divergence between impasse and agreement dialogues reflects sampling noise. Please add appropriate inferential statistics, such as a mixed-effects model with random intercepts for dyad (and possibly for role), or per-turn comparisons with multiple-comparison correction, and report effect sizes.","section":"Section 3.2, Figure 4"},{"comment":"The claim that a US-trained regression explains subjective outcomes 'better for the countries similar and worse for dissimilar ones' is contradicted by the numbers in Table 2. For South Africa, the cross-culture R² (0.157) is higher than the within-culture R² (0.047), and for Mexico the cross-culture value (0.170) is also higher than the within-culture value (0.137). This pattern actually suggests the transferred model often outperforms the within-culture baseline for non-US countries. The text needs to be corrected, or the analysis re-run with clearly specified baselines (e.g., a no-information baseline or a within-culture model trained on the same sample size) and with appropriate comparison metrics.","section":"Section 3.3, Table 2"},{"comment":"The only validation of the GPT-4o emotion labels is a correlation between utterance-level emotion scores and a single self-reported frustration measure collected after the dispute and aggregated across all participants. This does not establish per-turn validity, nor does it provide any evidence of cross-cultural measurement invariance. Section 5 explicitly concedes that there are no independent human annotations and that LLM-based emotion recognition in a cross-cultural setting is 'particularly fraught.' Because both the escalation findings and the cultural-difference findings in Section 3 rely entirely on these labels, the authors should either collect human annotations on a stratified sample (by culture and role) and report agreement, or substantially weaken the abstract's and Section 3's causal and comparative claims.","section":"Section 3.1, Appendix A.1, Section 5"},{"comment":"The cultural-emotion comparison in Figure 5 is presented as raw mean emotion scores without any statistical tests for differences across countries or between within-culture and cross-culture dyads. As a result, the claim that emotional expression varies by culture in a meaningful way is not established. Please provide inferential tests (e.g., an ANOVA or mixed-effects model with country and dyad type as factors) and report effect sizes and confidence intervals, or explicitly label this as a purely descriptive visualization.","section":"Section 3.3, Figure 5"}],"minor_comments":[{"comment":"The Dignity, Face, and Honor scale is described as an 18-item instrument, but no reliability statistics (e.g., Cronbach's alpha) are reported for the current sample. At least the scale reliability for each of the three subscales should be provided.","section":"Section 2.5.2"},{"comment":"The k-means clustering uses k=5 determined by the elbow method, but no stability analysis or silhouette-type validation is reported. The resulting clusters are interpreted as Dignity/Face/Honor clusters, but this mapping is an inference; consider reporting the cluster centroids on the three subscales to support the interpretation.","section":"Section 2.5.3, Figure 3"},{"comment":"The GPT-4o prompt instructs that all emotion scores should sum to one, but the paper does not state whether this constraint was satisfied in the model output and how it was handled in the regression analyses. Please clarify the post-processing steps.","section":"Section 2.4.2, Appendix A.3"},{"comment":"The text says a random subset of 406 dialogues was used for the regression, but the corpus statistics in Table 1 suggest far more human-human dyads; please clarify whether 406 is the number of dialogues, dyads, or participants, and how this subset was selected.","section":"Section 3.3"},{"comment":"The reference list is extensive and generally appropriate, but a few entries are missing DOIs or page ranges (e.g., the GPT-4 technical report); consider completing bibliographic details for consistency with journal style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The corpus itself is a strong contribution and the authors are unusually transparent about the limitations of their LLM-based analysis. The main issue is that the abstract and Section 3 make claims that go beyond the reported evidence: the escalation and cultural-transfer findings rely on unvalidated emotion labels and lack appropriate significance testing, and Table 2 actually undermines one of the stated conclusions. I think this is fixable within a revision by (a) adding human validation for a stratified sample, (b) re-analyzing with proper statistical tests, and (c) recalibrating the claims to match the evidence. I would not recommend rejection, but the paper should not be accepted in its current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The corpus is the contribution here. KODIS fills a real gap: dispute resolution is distinct from deal-making, and until now there was no large, multicultural dyadic corpus built for it. The expert-designed scenario, the preference elicitation, the Dignity/Face/Honor measures, and the human-AI extension are all sensible and well-documented. This will be a useful resource for NLP and social science researchers for years. That alone justifies sending the paper out for review.\n\nThe initial analysis, though, is weaker than the abstract implies. The abstract says the analysis “supports” escalation and cultural-difference theories, but every emotion measurement comes from GPT-4o with no independent human validation. The paper is admirably candid about this in Section 5, but the framing in Section 3 and the abstract still treats the findings as evidence. The escalation figure has no significance tests or intervals; it is a visual pattern. The cross-cultural transfer result in Table 2 is internally inconsistent: for South Africa, the US-trained model explains more variance (R² = 0.157) than the within-culture model (0.047), which contradicts the stated conclusion that transfer degrades for dissimilar cultures. That should have been flagged and explained, not left for the reader to trip over.\n\nThere is also a participant count discrepancy: the abstract says 4,061, but summing Table 1 (within-country dyads × 2, mixed dyads × 2, plus LLM dialogues) gives roughly 4,743. Minor, but it should be reconciled. The comparison of KODIS impasse rates to CaSiNo is suggestive, but the tasks, populations, and time periods differ, so calling it a “replication” is generous.\n\nWho is this for? Anyone working on negotiation dialogue, conflict escalation, or culturally aware NLP. The corpus is the prize; the analysis is a demo. The authors should be pushed to reframe it as exactly that: a pilot analysis that motivates the resource, not a test of psychological theory. If they add significance tests, subset-validate the emotion labels (at least on one culture pair), and fix the consistency issues, this becomes a solid contribution.\n\nMy recommendation: send it to peer review. It deserves referee time. The corpus will be cited, and the limitations, while real, are not fatal to the resource. But the editors should insist that the abstract and conclusions match the evidence.","headline":"KODIS is a genuinely useful new corpus for dispute-resolution research; the initial analysis, however, overclaims what the GPT-4o emotion labels can support.","tokens_in":14912,"tokens_out":3064,"would_cite":true,"duration_ms":32209,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new dialogue corpus spanning over 75 countries aims to show that expressed anger drives disputes into escalation and impasse.","keywords":["dispute resolution","dialogue corpus","cross-cultural emotion","conflict escalation","anger expression","Dignity Face Honor","customer service dispute","LLM emotion annotation"],"falsifier":"Collect independent human annotations of perceived anger and compassion for a stratified sample of KODIS turns from each of the five most common countries, then re-run the escalation and cross-cultural transfer analyses using those labels; if the reciprocal-anger pattern and the country transfer gaps disappear or reverse, the paper's central empirical claims would not survive.","tokens_in":13974,"feed_emoji":"⚖️","tokens_out":5276,"duration_ms":56867,"temperature":0.7,"pith_summary":"The paper introduces KODIS, a large dyadic dispute-resolution corpus built around a buyer-seller customer-service conflict, with thousands of dialogues from participants in over 75 countries. It argues that this resource fills a gap left by existing deal-making negotiation corpora, because disputes are backward-looking, emotionally intense conflicts rather than forward-looking opportunities. The initial analysis claims to support conflict-spiral theory: dialogues ending in impasse show reciprocal anger escalating across turns, while dialogues ending in agreement show one side refusing to reciprocate. The corpus also reveals cultural differences in emotional expression and is released openly for the NLP and social-science communities.","feed_headline":"75-country dispute corpus links anger spirals to impasses","feed_subtitle":"The paper opens a rare multicultural dataset for studying how emotions escalate or resolve real disputes.","key_machinery":"The load-bearing mechanism is a designed dispute task: a buyer and seller argue over a misdelivered basketball jersey, each with a different version of the facts, with four issues (refund, reviews, and apology) and monetary bonuses tied to self-elicited preference points. Surrounding this task, the corpus collects Dignity, Face, and Honor cultural-norm measures, risk propensity, pre-dispute preference vectors from which integrative potential is computed, post-dispute tactics, and the Subjective Value Inventory. The escalation analysis works by comparing GPT-4o-annotated anger trajectories across dialogue turns for impasse versus agreement dialogues, and the cultural analysis works by clustering countries on the cultural measures and testing how well an emotion-to-outcome regression trained on US dialogues predicts outcomes in other countries.","core_discovery":"The central claim is that KODIS provides the first large-scale multicultural dialogue corpus specifically for dispute resolution, distinct from deal-making corpora, and that the corpus's initial results confirm dispute theory's prediction that anger expressions provoke escalation rather than concession. In dialogues that end in impasse, the buyer's early anger is reciprocated by the seller and then grows further; in dialogues that end in agreement, sellers avoid reciprocating and buyer anger declines. The paper also claims that emotional expressions alone explain nearly half the variance in participants' subjective feelings about the process and their partner, and that models trained on United States dialogues transfer well to culturally similar countries but worse to culturally distant ones, evidence of cultural differences in emotional expression.","pith_inferences":["A direct test of the paper's empirical core would replace GPT-4o emotion labels with human annotations drawn from each of the five most common countries; if the reciprocal-anger pattern and the country transfer gaps disappear under those labels, the reported cultural differences would reflect the annotator model's biases rather than disputants' behavior.","Because one-minute-unmatched participants were paired with an AI partner without being told, the corpus also enables studies of how the mere suspicion of an AI counterpart changes emotional expression and dispute tactics; the paper leaves these comparisons largely unanalyzed.","The per-dyad integrative-potential measure means the corpus can support a direct test of whether cultural distance predicts failure to realize joint gains, an outcome analysis that goes beyond the paper's emotion-focused study.","The country-transfer results could be turned into a practical benchmark: the performance gap when an emotion-to-outcome model trained in one culture is applied to another is a quantitative measure of cultural bias in emotion classifiers."],"forward_implications":["If the escalating-anger finding is right, automated dispute-intervention systems could target the moment one side reciprocates the other's anger as the earliest detectable signal that an impasse is forming.","If emotional expression predicts subjective outcomes as strongly as reported, dialogue systems can estimate user dissatisfaction from affective tone alone, without parsing the substantive content of the dispute.","If the cross-cultural transfer results are right, emotion models trained on one culture will systematically misread disputes elsewhere, so culture-specific calibration will be needed for any AI that mediates or evaluates conflict.","If the deal-making versus dispute-resolution contrast holds, NLP benchmarks for negotiation should treat concession-making and conflict-escalation as separate capabilities rather than assuming one dialogue skill covers both."],"supporting_citations":[{"why":"Supplies the CaSiNo deal-making corpus and online matching framework that KODIS adapts and compares against for impasse rates.","marker":"Chawla et al., 2021"},{"why":"Provides the Dignity, Face, and Honor model of culture and negotiation tactics that motivates the task design and process measures.","marker":"Aslani et al., 2016"},{"why":"Is the conflict-escalation theory whose reciprocal-anger spiral the Section 3 analysis claims to replicate.","marker":"Pruitt, 2007"},{"why":"Supplies the contrasting result that anger elicits concessions in deal-making, sharpening the dispute-specific prediction.","marker":"Van Kleef et al., 2004"},{"why":"Is the evidence cited that GPT models give the most accurate emotion inferences on negotiation dialogues.","marker":"Kwon et al., 2024"},{"why":"Is the prior emotion-satisfaction analysis method on CaSiNo that KODIS's GPT emotion regression follows.","marker":"Chawla et al., 2023a"},{"why":"Provides the Dignity, Face, and Honor scale used to cluster countries and define cultural differences.","marker":"Leung and Cohen, 2011"},{"why":"Grounds the conceptual distinction between forward-looking deal-making and backward-looking disputes that motivates the corpus.","marker":"Brett, 2007"}],"fun_headline_variants":["Anger spirals in cross-cultural disputes: new 75-country corpus","Dispute corpus: anger escalates unless sellers stay cool","75-country dispute data shows anger fuels impasses","Anger spiral theory confirmed by multicultural dispute corpus","How anger escalates or resolves: 75-country dispute dialogue corpus"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The initial empirical findings assume GPT-4o's emotion scores measure the emotions a human listener would perceive across every culture studied, but the paper validates them only against participants' self-reported frustration and concedes it has no independent human annotations.","fun_headline_variants_meta":{"raw":{"variants":["Anger spirals in cross-cultural disputes: new 75-country corpus","Dispute corpus: anger escalates unless sellers stay cool","75-country dispute data shows anger fuels impasses","Anger spiral theory confirmed by multicultural dispute corpus","How anger escalates or resolves: 75-country dispute dialogue corpus"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000354,"raw_usage":{"total_tokens":1826,"prompt_tokens":747,"completion_tokens":1079,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":363,"completion_tokens_details":{"reasoning_tokens":997}},"tokens_in":363,"tokens_out":1079,"duration_ms":9254,"temperature":1.0,"reasoning_tokens":997,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:23:42.802993+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect independent human annotations of perceived anger and compassion for a stratified sample of KODIS turns from each of the five most common countries, then re-run the escalation and cross-cultural transfer analyses using those labels; if the reciprocal-anger pattern and the country transfer gaps disappear or reverse, the paper's central empirical claims would not survive.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the CaSiNo deal-making corpus and online matching framework that KODIS adapts and compares against for impasse rates."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Dignity, Face, and Honor model of culture and negotiation tactics that motivates the task design and process measures."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the conflict-escalation theory whose reciprocal-anger spiral the Section 3 analysis claims to replicate."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the contrasting result that anger elicits concessions in deal-making, sharpening the dispute-specific prediction."},{"cited_title":"Lucas, and Jonathan Gratch","cited_arxiv_id":null,"evidence_quote":"Is the evidence cited that GPT models give the most accurate emotion inferences on negotiation dialogues."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Dignity, Face, and Honor scale used to cluster countries and define cultural differences."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Grounds the conceptual distinction between forward-looking deal-making and backward-looking disputes that motivates the corpus."}],"review_version":1}