{"id":"be8ac911-4afd-48ce-b214-64af55dd84d5","arxiv_id":"2412.15993","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new 300-argument German corpus with crowdsourced discrete emotion labels shows LLMs overpredict negative emotions and only partially beat direct binary emotionality prompts.","lead":"Researchers added discrete emotion labels, such as anger or interest, to 300 German arguments and asked three language models to label the same texts. The models tended to overpredict negative emotions like fear and anger, and the new labeled corpus is offered as a resource for argument and emotion research.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Strict gold labels conflate annotator disagreement with absence of emotion, undermining the discrete-vs-binary comparison and the negative-bias finding.","rationale":"The reader identifies majority-aggregated gold as the weakest assumption; I agree. The no-majority-to-NO EMOTION rule is not a neutral tie-break: with three annotators and a ten-class label set, the absence of a majority is the expected outcome for most emotions, so the strict gold standard is dominated by the tie-break rather than by annotation signal. This affects the paper's headline in two ways: the category-inferred-vs-binary emotionality comparison and the negative-bias finding. The bias finding is qualitatively supported by the GPT fear examples, but the quantitative precision/recall numbers in Table 9 are not trustworthy under the strict rule. The corpus and relaxed evaluation are useful, and the issue is fixable by reanalysis, so I retain the reader's CONDITIONAL verdict rather than moving to reject.","tokens_in":19953,"tokens_out":7939,"duration_ms":74933,"concrete_test":"Recompute Tables 6 and 7/9 in strict mode with 'no majority' excluded from the per-class denominator, and with NO EMOTION as gold only when at least two annotators explicitly choose NO EMOTION; then compare the binary-vs-discrete F1 ordering and fear/anger precision. If any reported conclusion shifts materially, the central claims are artifacts of the tie-break rule.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strict evaluation in §4.3 assigns NO EMOTION whenever three human labels have no majority. Table 5 reports that for most discrete classes—including FEAR, PRIDE, and GUILT—no argument had all three annotators agree, and several classes have no two-annotator agreement. Consequently, many instances that are genuinely emotional but disputed are converted into gold NO EMOTION. The central claim that category-based prompting improves binary emotionality (Table 6) and the negative-bias pattern of high recall/low precision for anger and fear (Tables 7 and 9) are both computed against this recoded gold. A model that overpredicts fear is scored as false positive precisely where annotators split, which can manufacture the reported low precision. The relaxed evaluation partially addresses subjectivity, but the main headline and bias analysis rely on strict mode. The no-majority handling must be treated as uncertainty, not absence of emotion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Emo-DeFaBel, a crowdsourced corpus of 300 German argumentative texts from DeFaBel, each annotated by three workers with emotion categories, binary emotionality, stance, familiarity, and convincingness. It then evaluates three LLMs (Falcon-7b-instruct, Llama-3.1-8B-instruct, GPT-4o-mini) under binary, closed-domain, and open-domain emotion prompts in zero-shot, one-shot, and chain-of-thought variants. The main reported findings are that inferring binary emotionality from discrete emotion labels improves over direct binary prompting, and that LLMs exhibit a high-recall/low-precision bias toward anger and fear. The paper also reports correlations between emotion categories and perceived convincingness.","tokens_in":20075,"tokens_out":7413,"duration_ms":67778,"significance":"The corpus is a useful and timely resource: it is, to my knowledge, the first argument corpus with discrete emotion category annotations, the annotation protocol is documented in detail, and the data and code are publicly available. The qualitative analysis of GPT's fear predictions is also valuable, since it illustrates that LLMs rely on lexical cues while human annotations depend on stance and perceived personal relevance. However, the headline quantitative claims currently rest on an evaluation design that treats annotator disagreement as absence of emotion and on aggregations whose evaluation mode is not consistently reported. These issues must be resolved before the central claims can be accepted.","major_comments":[{"comment":"The strict evaluation recodes every argument without a three-way majority as NO EMOTION. Table 5 shows that this affects most emotional categories: FEAR, PRIDE, and GUILT have essentially no majority agreement (FEAR has 100% single-annotator and 0% two- or three-annotator agreement), and INTEREST reaches only 5% three-way agreement. Consequently, the strict gold standard encodes 'disputed' as 'non-emotional'. Since Table 6 (binary emotionality) and Tables 7 and 9 (discrete emotions) are the basis of the two central claims, the high-recall/low-precision pattern for anger and fear and the apparent advantage of category-based prompting may be artifacts of this recoding. Please report the binary and per-class results in relaxed mode as well, or treat no-majority instances as uncertain/unlabeled, and state explicitly which evaluation mode underlies each table.","section":"§4.3, Table 5"},{"comment":"Under the strict rule defined in §4.3, FEAR has zero gold-positive instances, yet Table 9 reports FEAR recall values between .62 and .75 for Falcon and GPT. Similarly, PRIDE has no majority agreement and GUILT is never annotated, so their strict gold counts are zero or near-zero. If Table 9 is computed in relaxed mode, the caption and surrounding text must say so; if it is strict, the recall computation is unexplained. This ambiguity directly affects the 'bias toward negative emotions' conclusion and must be clarified.","section":"Table 5 vs. Table 9"},{"comment":"The abstract claim that 'emotion categories enhance the prediction of emotionality' is not supported for all models. In Table 6, Falcon achieves the same F1 (.67) in binary, closed-domain, and open-domain settings, and Llama improves over binary only when compared with the poorly performing one-shot binary prompt (F1 .06 vs .67); its zero-shot binary F1 is already .61. Moreover, the apparent closed/open gains for GPT and Llama come from raising recall to 1.0 while precision drops to about .50, meaning the models essentially predict 'emotional' for nearly all arguments. The conclusion should be qualified to 'some models' and should address the precision-recall tradeoff explicitly.","section":"Abstract, §5.2.2, Table 6"},{"comment":"The claim that high recall with low precision for anger and fear holds 'across all prompt settings and models' is contradicted by Table 9: GPT has low anger recall (.08–.23 across settings), and Llama's fear recall is not consistently high (.19–.75). The bias statement should be restricted to the models and emotions for which it actually holds, such as Falcon fear, Llama anger, and GPT fear, rather than being presented as a universal finding.","section":"Abstract, §5.2.4, Table 9"}],"minor_comments":[{"comment":"The column headers '=1', '≤2', '≤3' should be clarified as 'exactly one annotator', 'at least two annotators', and 'all three annotators' to avoid ambiguity.","section":"Table 5"},{"comment":"Both tables should state in their captions whether the results are from strict or relaxed evaluation; currently the reader must infer this from the text.","section":"Tables 6 and 9"},{"comment":"The sentence 'we distribute one count of a false negative prediction across the set of gold labels' needs a formal definition; it is unclear whether the distribution is uniform, whether it applies only to relaxed mode, and how it interacts with macro-averaging.","section":"§4.3"},{"comment":"The cost '0.20C' appears to have a corrupted Euro symbol; it should read '€0.20'.","section":"§5.1"},{"comment":"The mapping of free-text labels such as 'Verwirrung' (confusion) to SURPRISE and 'Unsicherheit' (uncertainty) to FEAR is plausible but should be justified or at least flagged as a potentially consequential annotation decision.","section":"Appendix B"}],"recommendation":"major_revision","confidential_remarks":"The corpus release and annotation design are the strongest contribution, and the qualitative analysis of fear predictions is insightful. The LLM evaluation needs to be reworked around a defensible gold standard that does not conflate disagreement with absence of emotion, and the abstract should be made consistent with the per-model results. I would not recommend rejection, but the headline claims need substantial revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's my read of arXiv:2412.15993. The main contribution is Emo-DeFaBel, 300 German arguments with crowdsourced discrete emotion labels, and I think that part is real. It's the first argument corpus I know of with such annotations, and the annotation design is careful: native-speaker screening, attention checks, and a free-text emotion field that they map back to the label set. The authors also release code and data, which is good.\n\nThe more interesting scientific result is the LLM bias pattern: across models and prompts, anger and fear get high recall but low precision in the relaxed evaluation. That's a useful observation for anyone using LLMs to annotate emotions. The qualitative analysis of why GPT over-predicts fear (it keys on words like 'cancer' and 'accident' while human readers don't feel immediate threat) is a nice touch.\n\nNow the soft spots. The abstract says 'emotion categories enhance the prediction of emotionality.' Looking at Table 6, that's true for Llama and GPT but not for Falcon, and the improvement comes mostly from a precision-recall tradeoff: you get recall up to 1.0 but precision down at .50. That is not an unambiguous enhancement. The claim needs qualification.\n\nThere's a numerical inconsistency in Section 5.2.1: '326 annotations of statement-argument pairs' vs. the 300 arguments they set out to annotate. That needs clearing up.\n\nThe convincingness-emotion correlations (Figure 2) are presented without any significance testing. With rare emotions like pride (4 annotations), those correlations are likely noise. A sentence of caution would do.\n\nYou might have seen a stress-test note arguing that the no-majority-to-NO_EMOTION recoding in the strict evaluation undermines the bias finding. I don't think that lands. The per-class bias analysis in Table 9 is computed in the relaxed mode, where any human label counts as correct. The strict mode only affects the aggregate strict numbers, which are already low. The no-majority handling is a reasonable choice for a subjective task, though it does mean the strict numbers should be interpreted with that in mind.\n\nOverall, this is a solid resource paper, not a breakthrough. The corpus is small and German-only, and the annotation agreement is low, which limits what you can conclude from it. But the authors are upfront about the subjectivity, and they provide the data for others to reanalyze. I'd send it to a serious referee. The main fixes are: qualify the abstract claim, fix the count inconsistency, and add caveats to the correlation claims.\n\nLet me know if you want to discuss.","headline":"A genuine first corpus and a useful LLM bias finding, but the headline claim is overgeneralized and there are some reporting issues.","tokens_in":20599,"tokens_out":8682,"would_cite":true,"duration_ms":71578,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Discrete emotion categories improve prediction of emotionality in arguments and expose a negative-emotion bias in language models.","keywords":["emotion annotation","emotion categories","argument mining","large language models","prompting","German corpus","convincingness","negative bias"],"falsifier":"Re-annotate a random subset of, say, 100 arguments from Emo-DeFaBel with a larger panel of ten annotators per argument and recompute the majority-vote gold labels; if the emotionality flags or dominant emotions change for a substantial share of arguments, the reported precision, recall, and bias numbers for fear and anger do not rest on a stable ground truth.","tokens_in":19735,"feed_emoji":"😡","tokens_out":6323,"duration_ms":52973,"temperature":0.7,"pith_summary":"The paper argues that knowing which emotion an argument evokes is not a luxury but a necessity: discrete emotion categories improve the prediction of whether an argument is emotional at all. It introduces Emo-DeFaBel, a corpus of 300 German arguments, each annotated by three crowd workers for one of ten emotion categories, and uses it to test three prompting strategies on three instruction-tuned language models. The central empirical claim is that deriving a binary emotionality label from a discrete emotion label performs better than asking language models for the binary label directly. The paper also finds that the models systematically over-predict fear and anger, showing high recall and low precision for those categories. A secondary result links convincingness to emotion categories: arguments evoking joy or pride are rated more convincing, while anger is associated with lower convincingness.","feed_headline":"Discrete emotion labels beat direct emotionality checks","feed_subtitle":"Naming the emotion first beats a yes/no question, exposing a fear-and-anger bias in LLMs.","key_machinery":"The load-bearing object is Emo-DeFaBel, a corpus of 300 German persuasive arguments drawn from the DeFaBel corpus, each annotated by three crowd workers with one of ten emotion categories (joy, anger, fear, sadness, disgust, surprise, pride, interest, shame, guilt) or no emotion. The evaluation machinery is a two-mode scoring scheme: a strict mode in which the majority vote of the three annotators is the gold label, with no emotion assigned when no majority exists, and a relaxed mode in which any of the three annotator labels counts as correct. This machinery allows the authors to compare three output spaces (binary, closed-domain, open-domain) and three prompting techniques (zero-shot, one-shot, chain-of-thought), and it is what turns the discrete categories into a testable claim about binary emotionality.","core_discovery":"The authors claim to have built the first argumentative corpus labeled with discrete emotion categories and to have shown that these categories are not just a fine-grained extra but a better route to the binary emotionality judgments that argument mining already uses. In their experiments on a German argument corpus, asking GPT-4o-mini, Llama-3.1-8B-Instruct, and Falcon-7b-instruct to name the dominant emotion and then converting that name to an emotionality flag gives equal or better binary performance than asking for the flag directly, while also providing information a binary label cannot carry. The human annotations reveal a correlation between discrete emotion and perceived convincingness, with joy and pride rated higher and anger lower. All three models show a pronounced negative-emotion bias, especially high recall for fear and anger, which the paper explains as the models fixating on lexical threat cues rather than the reader's felt experience.","pith_inferences":["If the discrete-to-binary benefit holds beyond this German corpus, other subjective annotation tasks that currently use binary flags could be redesigned around a small closed set of categories, with the binary label derived afterward.","The negative-emotion bias is likely driven by lexical cues such as cancer, accident, or explosion; a stress test would be to prompt models with reader-stance information or to decorrelate these cue words and see whether fear and anger precision rises.","Given the low inter-annotator agreement, majority voting is a questionable gold standard; a distributional or multi-label treatment of emotion would likely change the ranking of the models.","The correlation between pride and joy with convincingness suggests a testable causal claim: if argument quality is judged under manipulated emotion primes, the same argument may be rated differently depending on which emotion it evokes."],"forward_implications":["Asking an LLM to select a discrete emotion label first is a viable route to binary emotionality detection in arguments, often outperforming a direct binary prompt.","Discrete emotion annotations of arguments are worth collecting, because categories such as anger, joy, and pride carry signals about convincingness that a binary label cannot express.","LLM-based emotion labeling should not be used without correction for its negative-emotion bias: fear and anger are over-predicted, while shame, guilt, and pride are almost never predicted.","Prompting technique matters little for this task; zero-shot, one-shot, and chain-of-thought produce similar overall performance.","The strict-versus-relaxed evaluation gap implies that LLM labels are often one of several plausible human labels, so single-label evaluation underestimates their utility."],"supporting_citations":[{"why":"Supplies the DeFaBel corpus from which the 300 German arguments are drawn.","marker":"Velutharambath et al. (2024)"},{"why":"Establishes binary emotionality and convincingness in arguments, the prior line of work this paper extends.","marker":"Habernal and Gurevych (2017)"},{"why":"Shows GPT-4 can reproduce human emotion labels in parliamentary speeches, motivating the use of LLMs for annotation.","marker":"Tarkka et al. (2024)"},{"why":"Introduces in-context one-shot prompting used in the one-shot condition.","marker":"Brown et al. (2020)"},{"why":"Introduces chain-of-thought prompting used in the chain-of-thought condition.","marker":"Wei et al. (2022)"},{"why":"Provides the finding that prompt choice is not dominant in low-data regimes, used to interpret the small differences across prompting settings.","marker":"Le Scao and Rush (2021)"},{"why":"TruthfulQA is the source of the statements behind the DeFaBel argument topics.","marker":"Lin et al. (2022)"},{"why":"Provides the Potato tool used for the crowdsourced annotation interface.","marker":"Pei et al. (2022)"}],"fun_headline_variants":["Discrete emotions beat yes/no in argument tasks","LLMs overpredict anger and fear in arguments","Emotion categories improve binary emotionality prediction","Crowdsourced emotion labels for arguments outperform binary","First German corpus with discrete argument emotions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The results are only as solid as the gold standard, which is the majority vote of three crowd annotators with no emotion used whenever no majority exists, even though annotators rarely agreed on emotion categories.","fun_headline_variants_meta":{"raw":{"variants":["Discrete emotions beat yes/no in argument tasks","LLMs overpredict anger and fear in arguments","Emotion categories improve binary emotionality prediction","Crowdsourced emotion labels for arguments outperform binary","First German corpus with discrete argument emotions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000224,"raw_usage":{"total_tokens":1463,"prompt_tokens":951,"completion_tokens":512,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":567,"completion_tokens_details":{"reasoning_tokens":443}},"tokens_in":567,"tokens_out":512,"duration_ms":6872,"temperature":1.0,"reasoning_tokens":443,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T10:53:30.816212+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a random subset of, say, 100 arguments from Emo-DeFaBel with a larger panel of ten annotators per argument and recompute the majority-vote gold labels; if the emotionality flags or dominant emotions change for a substantial share of arguments, the reported precision, recall, and bias numbers for fear and anger do not rest on a stable ground truth.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the DeFaBel corpus from which the 300 German arguments are drawn."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows GPT-4 can reproduce human emotion labels in parliamentary speeches, motivating the use of LLMs for annotation."}],"review_version":1}