REVIEW 4 major objections 5 minor 1 cited by
Fearful Falcons and Angry Llamas: Emotion Category Annotations of Arguments by Humans and LLMs
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Discrete emotion categories improve prediction of emotionality in arguments and expose a negative-emotion bias in language models.
desk verdict A genuine first corpus and a useful LLM bias finding, but the headline claim is overgeneralized and there are some reporting issues. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is Emo-DeFaBel, a corpus of 300 German persuasive arguments drawn from the DeFaBel corpus, each annotated by three crowd workers with one of ten emotion categories (joy, anger, fear, sadness, disgust, surprise, pride, interest, shame, guilt) or no emotion. The evaluation machinery is a two-mode scoring scheme: a strict mode in which the majority vote of the three annotators is the gold label, with no emotion assigned when no majority exists, and a relaxed mode in which any of the three annotator labels counts as correct. This machinery allows the authors to compare three output spaces (binary, closed-domain, open-domain) and three prompting techniques (zero-shot, one-shot, chain-of-thought), and it is what turns the discrete categories into a testable claim about binary emotionality.
What would settle it
Re-annotate a random subset of, say, 100 arguments from Emo-DeFaBel with a larger panel of ten annotators per argument and recompute the majority-vote gold labels; if the emotionality flags or dominant emotions change for a substantial share of arguments, the reported precision, recall, and bias numbers for fear and anger do not rest on a stable ground truth.
Extended reading notes
Core claim
The authors claim to have built the first argumentative corpus labeled with discrete emotion categories and to have shown that these categories are not just a fine-grained extra but a better route to the binary emotionality judgments that argument mining already uses. In their experiments on a German argument corpus, asking GPT-4o-mini, Llama-3.1-8B-Instruct, and Falcon-7b-instruct to name the dominant emotion and then converting that name to an emotionality flag gives equal or better binary performance than asking for the flag directly, while also providing information a binary label cannot carry. The human annotations reveal a correlation between discrete emotion and perceived convincingness, with joy and pride rated higher and anger lower. All three models show a pronounced negative-emotion bias, especially high recall for fear and anger, which the paper explains as the models fixating on lexical threat cues rather than the reader's felt experience.
Load-bearing premise
The results are only as solid as the gold standard, which is the majority vote of three crowd annotators with no emotion used whenever no majority exists, even though annotators rarely agreed on emotion categories.
Editorial extensions
If this is right
- Asking an LLM to select a discrete emotion label first is a viable route to binary emotionality detection in arguments, often outperforming a direct binary prompt.
- Discrete emotion annotations of arguments are worth collecting, because categories such as anger, joy, and pride carry signals about convincingness that a binary label cannot express.
- LLM-based emotion labeling should not be used without correction for its negative-emotion bias: fear and anger are over-predicted, while shame, guilt, and pride are almost never predicted.
- Prompting technique matters little for this task; zero-shot, one-shot, and chain-of-thought produce similar overall performance.
- The strict-versus-relaxed evaluation gap implies that LLM labels are often one of several plausible human labels, so single-label evaluation underestimates their utility.
Reading between the lines
- If the discrete-to-binary benefit holds beyond this German corpus, other subjective annotation tasks that currently use binary flags could be redesigned around a small closed set of categories, with the binary label derived afterward.
- The negative-emotion bias is likely driven by lexical cues such as cancer, accident, or explosion; a stress test would be to prompt models with reader-stance information or to decorrelate these cue words and see whether fear and anger precision rises.
- Given the low inter-annotator agreement, majority voting is a questionable gold standard; a distributional or multi-label treatment of emotion would likely change the ranking of the models.
- The correlation between pride and joy with convincingness suggests a testable causal claim: if argument quality is judged under manipulated emotion primes, the same argument may be rated differently depending on which emotion it evokes.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces Emo-DeFaBel, a crowdsourced corpus of 300 German argumentative texts from DeFaBel, each annotated by three workers with emotion categories, binary emotionality, stance, familiarity, and convincingness. It then evaluates three LLMs (Falcon-7b-instruct, Llama-3.1-8B-instruct, GPT-4o-mini) under binary, closed-domain, and open-domain emotion prompts in zero-shot, one-shot, and chain-of-thought variants. The main reported findings are that inferring binary emotionality from discrete emotion labels improves over direct binary prompting, and that LLMs exhibit a high-recall/low-precision bias toward anger and fear. The paper also reports correlations between emotion categories and perceived convincingness.
Significance. The corpus is a useful and timely resource: it is, to my knowledge, the first argument corpus with discrete emotion category annotations, the annotation protocol is documented in detail, and the data and code are publicly available. The qualitative analysis of GPT's fear predictions is also valuable, since it illustrates that LLMs rely on lexical cues while human annotations depend on stance and perceived personal relevance. However, the headline quantitative claims currently rest on an evaluation design that treats annotator disagreement as absence of emotion and on aggregations whose evaluation mode is not consistently reported. These issues must be resolved before the central claims can be accepted.
major comments (4)
- [§4.3, Table 5] The strict evaluation recodes every argument without a three-way majority as NO EMOTION. Table 5 shows that this affects most emotional categories: FEAR, PRIDE, and GUILT have essentially no majority agreement (FEAR has 100% single-annotator and 0% two- or three-annotator agreement), and INTEREST reaches only 5% three-way agreement. Consequently, the strict gold standard encodes 'disputed' as 'non-emotional'. Since Table 6 (binary emotionality) and Tables 7 and 9 (discrete emotions) are the basis of the two central claims, the high-recall/low-precision pattern for anger and fear and the apparent advantage of category-based prompting may be artifacts of this recoding. Please report the binary and per-class results in relaxed mode as well, or treat no-majority instances as uncertain/unlabeled, and state explicitly which evaluation mode underlies each table.
- [Table 5 vs. Table 9] Under the strict rule defined in §4.3, FEAR has zero gold-positive instances, yet Table 9 reports FEAR recall values between .62 and .75 for Falcon and GPT. Similarly, PRIDE has no majority agreement and GUILT is never annotated, so their strict gold counts are zero or near-zero. If Table 9 is computed in relaxed mode, the caption and surrounding text must say so; if it is strict, the recall computation is unexplained. This ambiguity directly affects the 'bias toward negative emotions' conclusion and must be clarified.
- [Abstract, §5.2.2, Table 6] The abstract claim that 'emotion categories enhance the prediction of emotionality' is not supported for all models. In Table 6, Falcon achieves the same F1 (.67) in binary, closed-domain, and open-domain settings, and Llama improves over binary only when compared with the poorly performing one-shot binary prompt (F1 .06 vs .67); its zero-shot binary F1 is already .61. Moreover, the apparent closed/open gains for GPT and Llama come from raising recall to 1.0 while precision drops to about .50, meaning the models essentially predict 'emotional' for nearly all arguments. The conclusion should be qualified to 'some models' and should address the precision-recall tradeoff explicitly.
- [Abstract, §5.2.4, Table 9] The claim that high recall with low precision for anger and fear holds 'across all prompt settings and models' is contradicted by Table 9: GPT has low anger recall (.08–.23 across settings), and Llama's fear recall is not consistently high (.19–.75). The bias statement should be restricted to the models and emotions for which it actually holds, such as Falcon fear, Llama anger, and GPT fear, rather than being presented as a universal finding.
minor comments (5)
- [Table 5] The column headers '=1', '≤2', '≤3' should be clarified as 'exactly one annotator', 'at least two annotators', and 'all three annotators' to avoid ambiguity.
- [Tables 6 and 9] Both tables should state in their captions whether the results are from strict or relaxed evaluation; currently the reader must infer this from the text.
- [§4.3] The sentence 'we distribute one count of a false negative prediction across the set of gold labels' needs a formal definition; it is unclear whether the distribution is uniform, whether it applies only to relaxed mode, and how it interacts with macro-averaging.
- [§5.1] The cost '0.20C' appears to have a corrupted Euro symbol; it should read '€0.20'.
- [Appendix B] The mapping of free-text labels such as 'Verwirrung' (confusion) to SURPRISE and 'Unsicherheit' (uncertainty) to FEAR is plausible but should be justified or at least flagged as a potentially consequential annotation decision.
Circularity Check
No significant circularity: empirical annotation study evaluated against external human labels.
full rationale
This is an empirical annotation and evaluation paper. The central claims (discrete emotion labels are useful; LLMs overpredict negative emotions) are assessed against crowdsourced human annotations collected independently of the LLM outputs. The derivation of binary emotionality from discrete emotion predictions (Section 4.1 and Table 6) is an evaluation design choice: it compares alternative prompting routes against the same human gold standard. Although the strict no-majority rule recodes annotator disagreement as NO EMOTION (Section 4.3), that is a gold-standard construction choice that could raise measurement concerns, not circular reasoning: the paper does not fit any parameter to the target result, does not define its input in terms of its output, and does not rely on load-bearing self-citation. The cited prior work is background and resource selection (DeFaBel), not an argument that presupposes the conclusions.
Assumptions & free parameters
assumptions (3)
- domain assumption The ten closed-set emotion labels (JOY, ANGER, FEAR, SADNESS, DISGUST, SURPRISE, PRIDE, INTEREST, SHAME, GUILT) plus NO EMOTION are an appropriate and sufficient set for emotions evoked by German arguments.
- domain assumption Majority vote of three annotators, with NO EMOTION for ties, is a usable gold standard for strict evaluation.
- domain assumption LLM outputs can be reliably mapped to the label set by JSON parsing or string search.
Cite this review
Pith. "Pith review of Fearful Falcons and Angry Llamas: Emotion Category Annotations of Arguments by Humans and LLMs." pith.science (2026). https://pith.science/paper/SF3FBS32
@misc{pith2026241215993,
author = {Pith},
title = {Pith review of: Fearful Falcons and Angry Llamas: Emotion Category Annotations of Arguments by Humans and LLMs},
year = {2026},
howpublished = {\url{https://pith.science/paper/SF3FBS32}},
note = {Machine review of arXiv:2412.15993}
}
read the original abstract
Arguments evoke emotions, influencing the effect of the argument itself. Not only the emotional intensity but also the category influence the argument's effects, for instance, the willingness to adapt stances. While binary emotionality has been studied in arguments, there is no work on discrete emotion categories (e.g., "Anger") in such data. To fill this gap, we crowdsource subjective annotations of emotion categories in a German argument corpus and evaluate automatic LLM-based labeling methods. Specifically, we compare three prompting strategies (zero-shot, one-shot, chain-of-thought) on three large instruction-tuned language models (Falcon-7b-instruct, Llama-3.1-8B-instruct, GPT-4o-mini). We further vary the definition of the output space to be binary (is there emotionality in the argument?), closed-domain (which emotion from a given label set is in the argument?), or open-domain (which emotion is in the argument?). We find that emotion categories enhance the prediction of emotionality in arguments, emphasizing the need for discrete emotion annotations in arguments. Across all prompt settings and models, automatic predictions show a high recall but low precision for predicting anger and fear, indicating a strong bias toward negative emotions.
Figures
Figures from the paper (3 more)
Forward citations
Cited by 1 Pith paper
-
Investigating Subjective Factors of Argument Strength: Storytelling, Emotions, and Hedging
Storytelling and hedging help subjective persuasion in online debate but hurt objective argument quality, while emotions show mostly domain-independent effects.
Reference graph
Works this paper leans on
-
[1]
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. 2023. https://arxiv.org/abs/2311.16867 The falcon series of open language models . Preprint, arXiv:2...
arXiv 2023
-
[2]
Christopher Bagdon, Prathamesh Karmalkar, Harsha Gurulingappa, and Roman Klinger. 2024. https://doi.org/10.18653/v1/2024.naacl-long.439 `` you are an expert annotator '' : Automatic best -- worst-scaling annotations for emotion intensity modeling . In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Lin...
-
[4]
Mohamed Benlamine, Ramla Ghali, Serena Villata, Claude Frasson, Fabien Gandon, and Elena Cabrio. 2017 b . https://doi.org/10.1007/978-3-319-58071-5_50 Persuasive argumentation and emotions: An empirical evaluation with users . Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics)
-
[5]
Benlamine, Maher Chaouachi, Serena Villata, Elena Cabrio, Claude Frasson, and Fabien L
Mohamed S. Benlamine, Maher Chaouachi, Serena Villata, Elena Cabrio, Claude Frasson, and Fabien L. Gandon. 2015. https://api.semanticscholar.org/CorpusID:11320420 Emotions in argumentation: an empirical evaluation . In International Joint Conference on Artificial Intelligence
work page 2015
-
[6]
Gerd Bohner, Kimberly Crow, Hans-Peter Erb, and Norbert Schwarz. 1992. https://doi.org/10.1002/ejsp.2420220602 Affect and persuasion: Mood effects on the processing of message content and context cues and on subsequent behavior . European Journal of Social Psychology, 22:511--530
-
[7]
Boster, Shannon Cruz, Brian Manata, Briana N
Franklin J. Boster, Shannon Cruz, Brian Manata, Briana N. DeAngelis, and Jie Zhuang. 2016. https://doi.org/10.1080/15534510.2016.1142892 A meta-analytic review of the effect of guilt on compliance . Social Influence, 11(1):54--67
arXiv 2016
-
[8]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...
2020
-
[9]
Felix Casel, Amelie Heindl, and Roman Klinger. 2021. https://aclanthology.org/2021.konvens-1.5 Emotion recognition under consideration of the emotion component process model . In Proceedings of the 17th Conference on Natural Language Processing (KONVENS 2021), pages 49--61, D \"u sseldorf, Germany. KONVENS 2021 Organizers
work page 2021
Show all 48 references
-
[10]
Yongchao Chen, Jacob Arkin, Yilun Hao, Yang Zhang, Nicholas Roy, and Chuchu Fan. 2024. https://doi.org/10.18653/v1/2024.emnlp-main.226 PR ompt optimization in multi-step tasks ( PROMST ): Integrating human feedback and heuristic-based sampling . In Proceedings of the 2024 Conf...
2024 doi
-
[11]
Long Cheng, Qihao Shao, Christine Zhao, Sheng Bi, and Gina-Anne Levow. 2024. https://aclanthology.org/2024.wassa-1.49 TEII : Think, explain, interact and iterate with large language models to solve cross-lingual emotion detection . In Proceedings of the 14th Workshop on Comput...
2024
-
[12]
Svetlana Churina, Preetika Verma, and Suchismita Tripathy. 2024. https://aclanthology.org/2024.wassa-1.38 WASSA 2024 shared task: Enhancing emotional intelligence with prompts . In Proceedings of the 14th Workshop on Computational Approaches to Subjectivity, Sentiment, & Socia...
2024
-
[13]
Chunhui Du, Jidong Tian, Haoran Liao, Jindou Chen, Hao He, and Yaohui Jin. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.150 Task-level thinking steps help large language models for challenging classification task . In Proceedings of the 2023 Conference on Empirical Method...
2023 doi
-
[14]
Roxanne El Baff, Henning Wachsmuth, Khalid Al Khatib, and Benno Stein. 2020. https://doi.org/10.18653/v1/2020.acl-main.287 A nalyzing the P ersuasive E ffect of S tyle in N ews E ditorial A rgumentation . In Proceedings of the 58th Annual Meeting of the Association for Computa...
2020 doi
-
[15]
Natalia Evgrafova, Veronique Hoste, and Els Lefever. 2024. https://aclanthology.org/2024.politicalnlp-1.5 Analysing pathos in user-generated argumentative text . In Proceedings of the Second Workshop on Natural Language Processing for Political Sciences @ LREC-COLING 2024, pag...
2024
-
[16]
Marcio Fonseca and Shay Cohen. 2024. https://doi.org/10.18653/v1/2024.findings-acl.478 Can large language models follow concept annotation guidelines? a case study on scientific and financial domains . In Findings of the Association for Computational Linguistics ACL 2024, page...
2024 doi
-
[17]
Fabrizio Gilardi, Meysam Alizadeh, and Maël Kubli. 2023. https://doi.org/10.1073/pnas.2305016120 Chatgpt outperforms crowd workers for text-annotation tasks . Proceedings of the National Academy of Sciences, 120(30)
2023 doi
-
[18]
Shiota, and Samantha L
Vladas Griskevicius, Michelle N. Shiota, and Samantha L. Neufeld. 2010. https://doi.org/10.1037/a0018421 Influence of different positive emotions on persuasion processing: A functional evolutionary approach. Emotion, 10(2):190--206
2010 doi
-
[19]
Ivan Habernal and Iryna Gurevych. 2016. https://doi.org/10.18653/v1/P16-1150 Which argument is more convincing? analyzing and predicting convincingness of web arguments using bidirectional LSTM . In Proceedings of the 54th Annual Meeting of the Association for Computational Li...
2016 doi
-
[20]
Ivan Habernal and Iryna Gurevych. 2017. https://doi.org/10.1162/COLI_a_00276 Argumentation mining in user-generated web discourse . Computational Linguistics, 43(1):125--179
2017 doi
-
[21]
Svetlana Kiritchenko and Saif Mohammad. 2018. https://doi.org/10.18653/v1/S18-2005 Examining gender and race bias in two hundred sentiment analysis systems . In Proceedings of the Seventh Joint Conference on Lexical and Computational Semantics, pages 43--53, New Orleans, Louis...
2018 doi
-
[22]
Roman Klinger, Orph \'e e De Clercq, Saif Mohammad, and Alexandra Balahur. 2018. https://doi.org/10.18653/v1/W18-6206 IEST : WASSA -2018 implicit emotions shared task . In Proceedings of the 9th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media A...
2018 doi
-
[23]
Yurie Koga, Shunsuke Kando, and Yusuke Miyao. 2024. https://aclanthology.org/2024.inlg-main.12 Forecasting implicit emotions elicited in conversations . In Proceedings of the 17th International Natural Language Generation Conference, pages 145--152, Tokyo, Japan. Association f...
2024
-
[24]
Takeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2024. Large language models are zero-shot reasoners. In Proceedings of the 36th International Conference on Neural Information Processing Systems, NIPS '22, Red Hook, NY, USA. Curran Associates Inc
2024
-
[25]
Teven Le Scao and Alexander Rush. 2021. https://doi.org/10.18653/v1/2021.naacl-main.208 How many data points is a prompt worth? In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pa...
2021 doi
-
[26]
Sophia Yat Mei Lee and Helena Yan Ping Lau. 2020. https://aclanthology.org/2020.lrec-1.203 An event-comment social media corpus for implicit emotion analysis . In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 1633--1642, Marseille, France. Euro...
2020
-
[27]
Howard Leventhal and Grevilda Trembly. 1968. https://doi.org/10.1111/j.1467-6494.1968.tb01466.x Negative emotions and persuasion . Journal of Personality, 36(1):154--168
1968
-
[28]
Moxin Li, Wenjie Wang, Fuli Feng, Yixin Cao, Jizhi Zhang, and Tat-Seng Chua. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.95 Robust prompt optimization for large language models against distribution shifts . In Proceedings of the 2023 Conference on Empirical Methods in Na...
2023 doi
-
[29]
Stephanie Lin, Jacob Hilton, and Owain Evans. 2022. https://doi.org/10.18653/v1/2022.acl-long.229 T ruthful QA : Measuring how models mimic human falsehoods . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pa...
2022 doi
-
[30]
Llama Team, AI @ Meta . 2024. https://arxiv.org/abs/2407.21783 The llama 3 herd of models . Preprint, arXiv:2407.21783
2024 arXiv
-
[31]
Stephanie Lukin, Pranav Anand, Marilyn Walker, and Steve Whittaker. 2017. https://aclanthology.org/E17-1070 Argument strength is in the eye of the beholder: Audience effects in persuasion . In Proceedings of the 15th Conference of the E uropean Chapter of the Association for C...
2017
-
[32]
Usman Malik, Simon Bernard, Alexandre Pauchet, Clément Chatelain, Romain Picot-Clémente, and Jérôme Cortinovis. 2024. https://doi.org/10.1109/ACCESS.2024.3354705 Pseudo-labeling with large language models for multi-label emotion classification of french tweets . IEEE Access, 1...
2024
-
[33]
Saif Mohammad. 2011. https://aclanthology.org/W11-1514 From once upon a time to happily ever after: Tracking emotions in novels and fairy tales . In Proceedings of the 5th ACL - HLT Workshop on Language Technology for Cultural Heritage, Social Sciences, and Humanities , pages ...
2011
-
[34]
Saif Mohammad, Xiaodan Zhu, and Joel Martin. 2014. https://doi.org/10.3115/v1/W14-2607 Semantic role labeling of emotions in tweets . In Proceedings of the 5th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis, pages 32--41, Baltimore, M...
2014 doi
-
[35]
Andrew Nedilko. 2023. https://doi.org/10.18653/v1/2023.wassa-1.61 Generative pretrained transformers for emotion detection in a code-switching setting . In Proceedings of the 13th Workshop on Computational Approaches to Subjectivity, Sentiment, & Social Media Analysis , pages ...
2023 doi
-
[36]
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...
2024 arXiv
-
[37]
Jiaxin Pei, Aparna Ananthasubramaniam, Xingyao Wang, Naitian Zhou, Apostolos Dedeloudis, Jackson Sargent, and David Jurgens. 2022. https://doi.org/10.18653/v1/2022.emnlp-demos.33 POTATO : The portable text annotation tool . In Proceedings of the 2022 Conference on Empirical Me...
2022 doi
-
[38]
Richard Petty, David Schumann, Steven Richman, and Alan Strathman. 1993. https://doi.org/10.1037/0022-3514.64.1.5 Positive mood and persuasion: Different roles for affect under high and low-elaboration conditions . Journal of Personality and Social Psychology, 64:5--20
1993 doi
-
[39]
M Pfau, A Szabo, J Anderson, Josh Morrill, J Zubric, and H-H H-Wan. 2006. https://doi.org/10.1111/j.1468-2958.2001.tb00781.x The role and impact of affect in the process of resistance to persuasion . Human Communication Research, 27:216 -- 252
2006
-
[40]
Laria Reynolds and Kyle McDonell. 2021. https://arxiv.org/abs/2102.07350 Prompt programming for large language models: Beyond the few-shot paradigm . Preprint, arXiv:2102.07350
2021 arXiv
-
[41]
Otto Tarkka, Jaakko Koljonen, Markus Korhonen, Juuso Laine, Kristian Martiskainen, Kimmo Elo, and Veronika Laippala. 2024. https://aclanthology.org/2024.parlaclarin-1.11 Automated emotion annotation of F innish parliamentary speeches using GPT -4 . In Proceedings of the IV Wor...
2024
-
[42]
Enrica Troiano, Laura Oberl \"a nder, and Roman Klinger. 2023. https://doi.org/10.1162/coli_a_00461 Dimensional modeling of emotions in text with appraisal theories: Corpus creation, annotation reliability, and prediction . Computational Linguistics, 49(1):1--72
2023 doi
-
[43]
Aswathy Velutharambath, Amelie W \"u hrl, and Roman Klinger. 2024. https://aclanthology.org/2024.lrec-main.243 Can factual statements be deceptive? the D e F a B el corpus of belief-based deception . In Proceedings of the 2024 Joint International Conference on Computational Li...
2024
-
[44]
Henning Wachsmuth, Nona Naderi, Yufang Hou, Yonatan Bilu, Vinodkumar Prabhakaran, Tim Alberdingk Thijm, Graeme Hirst, and Benno Stein. 2017. https://aclanthology.org/E17-1017 Computational argumentation quality assessment in natural language . In Proceedings of the 15th Confer...
2017
-
[45]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, brian ichter, Fei Xia, Ed Chi, Quoc V Le, and Denny Zhou. 2022. https://proceedings.neurips.cc/paper_files/paper/2022/file/9d5609613524ecf4f15af0f7b31abca4-Paper-Conference.pdf Chain-of-thought prompting elicits reasoning...
2022
-
[46]
Worth and Diane M
Leila T. Worth and Diane M. Mackie. 1987. https://api.semanticscholar.org/CorpusID:145366874 Cognitive mediation of positive affect in persuasion . Social Cognition, 5:76--94
1987
-
[47]
Qinyuan Ye, Mohamed Ahmed, Reid Pryzant, and Fereshte Khani. 2024. https://doi.org/10.18653/v1/2024.findings-acl.21 Prompt engineering a prompt engineer . In Findings of the Association for Computational Linguistics ACL 2024, pages 355--385, Bangkok, Thailand and virtual meeti...
2024 doi
-
[48]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...
-
[49]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.