Pith. sign in

REVIEW 4 major objections 4 minor 60 references

Investigating Subjective Factors of Argument Strength: Storytelling, Emotions, and Hedging

T0 review · 4 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read Storytelling and hedging boost subjective persuasion but suppress objective argument quality in regression tests on two datasets.

desk verdict A genuinely new joint regression with a plausible but unsecured cross-domain annotation pipeline; worth refereeing with validation requested. read the letter →

arxiv 2507.17409 v1 pith:JPPEJTRT submitted 2025-07-23 cs.CL

classification cs.CL
keywords argumentstrengthpersuasionqualitystorytellingemotionshedgingregressionanalysisautomatedannotation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper investigates whether three subjective rhetorical devices—personal storytelling, emotional language, and hedging (uncertainty markers like 'probably' or 'I think')—strengthen or weaken arguments, and whether their impact depends on how argument strength is defined. Using regression on two standard datasets, IBM ARGQ for objective argument quality and Cornell CMV for subjective persuasion, the authors find that storytelling and hedging significantly help persuasion but significantly hurt objective quality. Emotion effects are mostly stable across the two datasets: guilt/shame and disgust reduce argument strength, while fear and sadness increase it, and the authors interpret this as evidence that the rhetoric use of emotions matters more than the domain. The paper also contributes automated annotation methods for the three features, since no existing dataset labels all of them.

What carries the argument

The central object is a pair of parallel regressions on two datasets that operationalize objective and subjective argument strength: OLS linear regression on IBM ARGQ's continuous quality score and logistic regression on Cornell CMV's binary persuasion label. The independent variables come from automated annotation layers: a 10-fold RoBERTa ensemble for storytelling trained on mixed domains, masked emotion classifiers per emotion trained on CROWD-ENVENT event descriptions, and a lexicon-based hedge detector with syntactic disambiguation rules. The design's power comes from the contrast between the datasets, which differ in collection, length, and annotation procedure, so any feature whose coefficient sign is consistent across both is attributed to the feature itself, while a sign flip is attributed to the difference between objective and subjective strength.

What would settle it

A manual annotation of a random sample of 200–300 arguments from each dataset by human judges for storytelling, the ten emotions, and hedging, followed by re-running the regressions with the human labels, would settle whether the automated-label coefficients reflect real effects or measurement error; if the signs or significances change, the paper's central contrast is an artifact of the annotation pipeline.

Watch

Extended reading notes

Core claim

The central claim is that the value of a subjective rhetorical feature is not intrinsic; it flips sign with the facet of argument strength being measured. On IBM ARGQ, where short impersonal arguments are scored by averaged crowd judgments of adequacy, storytelling has a significant negative coefficient ($\beta = -0.182$) and hedge count is significantly negative ($\beta = -0.011$). On Cornell CMV, where a delta means one reader changed their mind, storytelling raises the odds of persuasion by a factor of 1.148 and hedge count by 1.030. Emotions behave differently: disgust and guilt/shame significantly lower argument strength in both datasets, while fear and sadness significantly raise it, which the authors trace to the difference between emotional attacks aimed at a participant and emotional appeals to universal concerns. The regression models themselves explain only a few percent of variance, so the contribution is about the direction and significance of the effects, not their size.

Load-bearing premise

The automated annotations of storytelling, emotions, and hedges are accurate enough on the target corpora to serve as independent variables, but they are only validated on held-out training-domain data, not on IBM ARGQ or Cornell CMV text.

Editorial extensions

If this is right

  • Argument-quality models built for objective settings such as essays or debates should treat personal stories and hedges as likely weaknesses, not strengths.
  • Persuasion-oriented systems, in contrast, can treat hedges and personal narratives as positive signals of credibility-building.
  • Emotion-aware argument models can expect the same directional effects across domains: disgust and guilt/shame lower strength, fear and sadness raise it.
  • Because explained variance is low (about 4% for quality and 1% for persuasion), these features complement rather than replace contextual and demographic predictors.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A direct test of the paper's interpretation would be to re-run the analysis with a third dataset that contains both objective and subjective strength labels for the same arguments; the framework predicts that feature coefficients align with the label definition, not with the topic or platform.
  • The emotion results suggest that classifying emotional attacks versus emotional appeals, rather than raw emotion categories, could be a stronger predictor of argument strength than any single emotion label.
  • The automated annotation transfer assumption is fragile: because classifiers are validated only on training-domain held-out data, a manual validation sample from the target corpora would reveal whether the sign flips are genuine or measurement artifacts.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. This paper studies the joint impact of three subjective features—storytelling, emotions, and hedging—on two argument-strength corpora. The authors automatically annotate IBM ARGQ (5.3k short debate arguments, objective quality scores) and Cornell CMV (11.5k Reddit comments, binary persuasion labels) using a RoBERTa ensemble for storytelling and ten emotion classifiers trained on CROWD-ENVENT, plus lexicon-based hedge detection. They then run univariate and stepwise regressions with interactions. They report that storytelling and hedging have negative coefficients on IBM ARGQ but positive effects on Cornell CMV persuasiveness, while discrete emotions have largely consistent effects across corpora, with disgust and guilt/shame negative and fear/sadness positive. The paper also contributes the annotated datasets and code.

Significance. If the empirical results hold, the main contrast—subjective features helping persuasion but hurting objective quality—would be a valuable contribution to computational argumentation, and the released annotated datasets would support further research. The paper has genuine strengths: it compares multiple annotation strategies (e.g., masked versus unmasked emotion training, mixed-domain ensembles), it reproduces prior storytelling results, and it is unusually explicit about the limitations of its cross-domain annotation. The data and code are shared. However, the significance of the empirical findings is currently limited by the absence of target-domain validation of the automatically predicted features and by inferential choices discussed below.

major comments (4)
  1. [§4.1, §4.2, Table 2, Limitations] The independent variables for storytelling and emotions are machine predictions with no target-domain validation. For IBM ARGQ, the storytelling classifier predicts only 0.8% positives (45 of 5.3k instances in Table 2), and all reported held-out F1 scores in Table 5 come from training-domain test sets (Falk and Lapesa 2022; CROWD-ENVENT), not from IBM ARGQ or Cornell CMV. The Limitations section explicitly justifies the emotion annotations via good results on the heldout training data. Measurement error in binary predictors is non-classical and can bias regression coefficients in either direction, especially with rare positives. Since the headline storytelling/hedging contrast is identified by these predicted labels, the authors should supply target-domain validation (for example, manual evaluation of a random sample, particularly the 45 IBM ARGQ storytelling positives), run a sensitivity analysis under misclassification assumptions, or substantially soften the contrast claims.
  2. [§3, Discussion] The central contrast between objective argument quality and individualized persuasion is confounded by domain differences. IBM ARGQ and Cornell CMV differ simultaneously in text length, register, collection procedure, annotation process, and topic, as the authors acknowledge in §3 when they state that the number of differences disallows a comparison without confounding factors. The opposite signs for storytelling and hedging may therefore reflect domain, length, or register effects rather than the objective/subjective distinction. The paper should frame the result as a comparison of two datasets rather than of two argument-strength constructs, or add within-domain controls such as length, topic, or matched sub-samples.
  3. [§5, Table 9] The stepwise regression procedure selects features and interactions by AIC and significance tests on the same data, and the resulting p-values are reported without correction for this selection. In addition, Table 3 tests multiple emotion categories plus storytelling and hedging variants against two corpora, and some Cornell CMV effects are only marginal (for example storytelling p=0.015, pride p=0.042, relief p=0.007) with no multiple-comparison correction. With final explained variance of 3.96% for IBM ARGQ and 1.36% for Cornell CMV, the Limitations statement that high significance nonetheless shows importance overstates the evidence, because post-selection p-values are not valid and the effect sizes are very small. The authors should hold out a validation set for model selection, correct for multiple testing, and report effect sizes with confidence intervals.
  4. [§5 Discussion, Conclusion] The abstract claim that the influence of emotions depends on their rhetoric utilization rather than the domain is not supported by the regression design. The models measure emotion categories such as anger, fear, and disgust, but not rhetoric utilization such as emotional attack versus appeal; the attack/appeal distinction is introduced post hoc from a few examples in Table 4 and a qualitative reading. To support this claim, the authors would need to annotate or automatically measure attack/appeal usage and include it in the regression, or explicitly label the claim as an untested hypothesis.
minor comments (4)
  1. [§5] The sentence listing the features as 1 storytelling plus 9 emotions excluding boredom and surprise plus 6 hedging variants does not match Table 3, which includes boredom and reports only one hedge count; please harmonize the feature count and the table.
  2. [§5] The expression beta-story = 0.138, beta-hedge) = 0.030 contains a misplaced parenthesis; please clarify whether the Cornell CMV estimates are log-odds or odds ratios.
  3. [Table 5] The header standard deviance should be standard deviation; the current wording is ambiguous and nonstandard.
  4. [§4.1] The storytelling training data includes a subset of r/ChangeMyView, so the Cornell CMV predictions are not fully cross-domain; please state this explicitly when discussing cross-domain robustness.

Circularity Check

0 steps flagged · score 2.0 of 10

Empirical regression with automated annotations; no equation-level circularity; the only same-author citation is not load-bearing for the central contrast.

full rationale

The paper's derivation chain is an empirical regression: Section 4 produces automated annotations via classifiers trained on external or prior corpora, and Section 5 regresses those annotations on the target variables. No independent variable is defined in terms of the dependent variable, and no fitted parameter is renamed as a prediction. The central claim—storytelling and hedging have opposite-signed effects on objective argument quality versus subjective persuasion—is a regression outcome, not an identity or a restatement of the annotation models' outputs. The one place where same-author prior work appears is the storytelling annotation setup, which follows Falk and Lapesa (2022) for training data and ensemble design; however, the paper re-trains the classifiers and reports its own held-out F1 = 0.82, and the prior work is not used to forbid alternatives or to supply the regression result. Thus the self-citation is not load-bearing. The Limitations paragraph's reliance on held-out training-domain performance for the emotion classifiers is an explicit cross-domain validity limitation, not a circularity: it concerns measurement error, not a reduction of the conclusion to the input. No circular step can be exhibited by quoting an equation or construction, so no specific circular step is reported.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claims rest on the validity of predicted annotations and on the mapping of IBM ARGQ to objective quality and Cornell CMV to subjective persuasion. No new physical or conceptual entities are postulated; the emotional attack versus appeal distinction is an interpretive hypothesis, not a tested variable.

assumptions (4)
  • ad hoc to paper The automatically predicted feature annotations (storytelling, emotions, hedging) are valid proxies for the true features in both target corpora.
    The paper validates classifiers only on held-out training domains and assumes transfer to IBM ARGQ and Cornell CMV; see the Limitations paragraph.
  • domain assumption IBM ARGQ crowd-averaged scores operationalize objective argument quality, and Cornell CMV deltas operationalize individualized subjective persuasion.
    Section 3 assigns these conceptualizations while acknowledging that the datasets differ in many confounded ways.
  • standard math OLS and logistic regression are appropriate functional forms, with independent errors and no serious multicollinearity.
    Section 5 uses these models but reports no residual or multicollinearity diagnostics.
  • domain assumption Hedge lexicons and disambiguation rules transfer to both corpora, including Reddit abbreviation usage.
    Section 4.3 says the lexicons are adapted from similar semi-formal domains, but no target-domain evaluation is provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Investigating Subjective Factors of Argument Strength: Storytelling, Emotions, and Hedging." pith.science (2026). https://pith.science/paper/JPPEJTRT

@misc{pith2026250717409,
  author       = {Pith},
  title        = {Pith review of: Investigating Subjective Factors of Argument Strength: Storytelling, Emotions, and Hedging},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JPPEJTRT}},
  note         = {Machine review of arXiv:2507.17409}
}
abstract

In assessing argument strength, the notions of what makes a good argument are manifold. With the broader trend towards treating subjectivity as an asset and not a problem in NLP, new dimensions of argument quality are studied. Although studies on individual subjective features like personal stories exist, there is a lack of large-scale analyses of the relation between these features and argument strength. To address this gap, we conduct regression analysis to quantify the impact of subjective factors $-$ emotions, storytelling, and hedging $-$ on two standard datasets annotated for objective argument quality and subjective persuasion. As such, our contribution is twofold: at the level of contributed resources, as there are no datasets annotated with all studied dimensions, this work compares and evaluates automated annotation methods for each subjective feature. At the level of novel insights, our regression analysis uncovers different patterns of impact of subjective features on the two facets of argument strength encoded in the datasets. Our results show that storytelling and hedging have contrasting effects on objective and subjective argument quality, while the influence of emotions depends on their rhetoric utilization rather than the domain.

Figures

Figures reproduced from arXiv: 2507.17409 by the authors.

Figure 2
Figure 2. Interaction between fear (x-axis) and sadness (standard deviations shown through hue and dashing) on IBM ARGQ with confidence intervals. of argument strength. To that end, we first de￾termined the feasibility of large-scale automated annotation of our subjective features, to then sys￾tematically reveal correlations through a regression analysis. We could reveal a significant effect of almost all observed features on… view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

60 extracted references · 47 canonical work pages

  1. [1]

    Khalid Al-Khatib, Henning Wachsmuth, Matthias Hagen, and Benno Stein. 2017. https://doi.org/10.18653/v1/D17-1141 Patterns of argumentation strategies across topics . In Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, pages 1351--1357, Copenhagen, Denmark. Association for Computational Linguistics

  2. [2]

    Liliana Ardissono, Guido Boella, and Leonardo Lesmo. 1999. Politeness and speech acts. In Proc. Workshop on Attitude, Personality and Emotions in User-Adapted Interaction, pages 41--55. Citeseer

  3. [3]

    Aristotle. 2007. On Rhetoric: A Theory of Civic Discourse. (Kennedy, G.A., translator), Oxford University Press

  4. [4]

    Mohamed Sahbi Benlamine, Maher Chaouachi, Serena Villata, Elena Cabrio, Claude Frasson, and Fabien Gandon. 2015. https://hal.inria.fr/hal-01152966 Emotions in argumentation: an empirical evaluation . In International Joint Conference on Artificial Intelligence, IJCAI 2015, pages 156--163

  5. [5]

    Mohamed Sahbi Benlamine, Serena Villata, Ramla Ghali, Claude Frasson, Fabien Gandon, and Elena Cabrio. 2017. https://doi.org/10.1007/978-3-319-58071-5_50 Persuasive argumentation and emotions: An empirical evaluation with users . In Human-Computer Interaction. User Interface Design, Development and Multimodality, pages 659--671, Cham. Springer Internation...

  6. [6]

    Laura W. Black. 2008. https://doi.org/10.1111/j.1468-2885.2007.00315.x Deliberation, Storytelling, and Dialogic Moments . Communication Theory, 18(1):93--116

  7. [7]

    Laura W. Black. 2013. https://doi.org/10.16997/jdd.153 Framing Democracy and Conflict Through Storytelling in Deliberative Groups . Journal of Public Deliberation, 9(1):art. 4

  8. [8]

    G D Bryant and G R Norman. 1979. The communication of uncertainty. In Proceedings of the Eighteenth Annual Conference on Research in Medical Education

Show all 60 references
  1. [9]

    Moitreya Chatterjee, Sunghyun Park, Han Suk Shim, Kenji Sagae, and Louis-Philippe Morency. 2014. https://doi.org/10.3115/v1/W14-5908 Verbal behaviors and persuasiveness in online multimedia content . In Proceedings of the Second Workshop on Natural Language Processing for Soci...

  2. [10]

    Roxanne El Baff, Henning Wachsmuth, Khalid Al Khatib, and Benno Stein. 2020. https://doi.org/10.18653/v1/2020.acl-main.287 A nalyzing the P ersuasive E ffect of S tyle in N ews E ditorial A rgumentation . In Proceedings of the 58th Annual Meeting of the Association for Computa...

  3. [11]

    Katharina Esau. 2018. https://doi.org/doi:10.1515/auk-2018-0003 Capturing citizens' values: On the role of narratives and emotions in digital participation . Analyse & Kritik, 40(1):55--72

  4. [12]

    Neele Falk and Gabriella Lapesa. 2022. https://doi.org/10.18653/v1/2022.acl-long.379 Reports of personal experiences and stories in argumentation: datasets and analysis . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long...

  5. [13]

    Neele Falk and Gabriella Lapesa. 2023. https://doi.org/10.18653/v1/2023.acl-long.132 S tory ARG : a corpus of narratives and personal experiences in argumentative texts . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long...

  6. [14]

    Michael Fromm, Max Berrendorf, Johanna Reiml, Isabelle Mayerhofer, Siddharth Bhargava, Evgeniy Faerman, and Thomas Seidl. 2022. https://doi.org/10.48550/ARXIV.2205.09803 Towards a holistic view on argument quality prediction . arXiv preprint

  7. [15]

    Marlène Gerber, André Bächtiger, Susumu Shikano, Simon Reber, and Samuel Rohr. 2018. https://doi.org/10.1017/S0007123416000144 Deliberative abilities and influence in a transnational deliberative poll (europolis) . British Journal of Political Science, 48(4):1093--1118

  8. [16]

    Martin Gleize, Eyal Shnarch, Leshem Choshen, Lena Dankin, Guy Moshkowich, Ranit Aharonov, and Noam Slonim. 2019. https://doi.org/10.18653/v1/P19-1093 Are you convinced? choosing the more convincing evidence with a S iamese network . In Proceedings of the 57th Annual Meeting of...

  9. [17]

    Lynn Greschner and Roman Klinger. 2024. https://arxiv.org/abs/2412.15993 Fearful falcons and angry llamas: Emotion category annotations of arguments by humans and llms . Preprint, arXiv:2412.15993

  10. [18]

    Shai Gretz, Roni Friedman, Edo Cohen - Karlik, Assaf Toledo, Dan Lahav, Ranit Aharonov, and Noam Slonim. 2020. https://doi.org/10.1609/AAAI.V34I05.6285 A large-scale dataset for argument quality ranking: Construction and analysis . In The Thirty-Fourth AAAI Conference on Artif...

  11. [19]

    Kathrin Grosse, Maria P Gonzalez, Carlos I Chesnevar, and Ana G Maguitman. 2015. Integrating argumentation and sentiment analysis for mining opinions from twitter. AI Communications, 28(3):387--401

  12. [20]

    Ivan Habernal and Iryna Gurevych. 2016. https://doi.org/10.18653/v1/P16-1150 Which argument is more convincing? analyzing and predicting convincingness of web arguments using bidirectional LSTM . In Proceedings of the 54th Annual Meeting of the Association for Computational Li...

  13. [21]

    Ivan Habernal and Iryna Gurevych. 2017. https://doi.org/10.1162/COLI_a_00276 Argumentation mining in user-generated web discourse . Computational Linguistics, 43(1):125--179

  14. [22]

    Paul Hoggett and Simon Thompson. 2012. https://books.google.de/books?id=53P5Fg4QL8YC Politics and the Emotions: The Affective Turn in Contemporary Political Studies . Bloomsbury Publishing

  15. [23]

    Ken Hyland. 1998. https://doi.org/10.1075/pbns.54 Hedging in scientific research articles . John Benjamins

  16. [24]

    Jumayel Islam, Lu Xiao, and Robert E. Mercer. 2020. https://aclanthology.org/2020.lrec-1.380 A lexicon-based approach for detecting hedges in informal text . In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 3109--3113, Marseille, France. Europe...

  17. [25]

    Jordan, Diane L

    Michelle E. Jordan, Diane L. Schallert, Yangjoo Park, SoonAh Lee, Yueh hui Vanessa Chiang, An-Chih Janne Cheng, Kwangok Song, Hsiang-Ning Rebecca Chu, Taehee Kim, and Haekyung Lee. 2012. https://doi.org/10.1080/0163853X.2012.722851 Expressing uncertainty in computer-mediated d...

  18. [26]

    Roman Kazakov, Kseniia Petukhova, and Ekaterina Kochmar. 2024. https://doi.org/10.18653/v1/2024.semeval-1.164 P et K az at S em E val-2024 task 3: Advancing emotion classification with an LLM for emotion-cause pair extraction in conversations . In Proceedings of the 18th Inter...

  19. [27]

    Johannes Kiesel, Milad Alshomary, Nicolas Handke, Xiaoni Cai, Henning Wachsmuth, and Benno Stein. 2022. https://aclanthology.org/2022.acl-long.306 Identifying the human values behind arguments . In Proceedings of the 60th Annual Meeting of the Association for Computational Lin...

  20. [28]

    Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. 2017. https://proceedings.neurips.cc/paper/2017/hash/9ef2ed4b7fd2c810847ffa5fa85bce38-Abstract.html Simple and scalable predictive uncertainty estimation using deep ensembles . In Advances in Neural Information ...

  21. [29]

    John Lawrence and Chris Reed. 2019. https://doi.org/10.1162/coli_a_00364 Argument mining: A survey . Computational Linguistics, 45(4):765--818

  22. [30]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. https://doi.org/10.48550/ARXIV.1907.11692 Roberta: A robustly optimized bert pretraining approach . arXiv preprint

  23. [31]

    Stephanie Lukin, Pranav Anand, Marilyn Walker, and Steve Whittaker. 2017. https://aclanthology.org/E17-1070 Argument strength is in the eye of the beholder: Audience effects in persuasion . In Proceedings of the 15th Conference of the E uropean Chapter of the Association for C...

  24. [32]

    Kelvin Luu, Chenhao Tan, and Noah A. Smith. 2019. https://doi.org/10.1162/tacl_a_00281 Measuring online debaters ' persuasive skill from text over time . Transactions of the Association for Computational Linguistics, 7:537--550

  25. [33]

    John Lyons. 1977. https://doi.org/10.1017/CBO9780511620614.010 Modality , volume 2, page 787–849. Cambridge University Press

  26. [34]

    Rousiley C. M. Maia, Danila Cal, Janine Bargas, and Neylson J. B. Crepalde. 2020. https://doi.org/10.1017/S1755773919000328 Which types of reason-giving and storytelling are good for deliberation? assessing the discussion dynamics in legislative and citizen forums . European P...

  27. [35]

    Rousiley C. M. Maia and Gabriella Hauber. 2020. The emotional dimensions of reason-giving in deliberative forums. Policy Sciences, 53:33--59

  28. [36]

    Aaron Maladry, Pranaydeep Singh, and Els Lefever. 2024. https://doi.org/10.18653/v1/2024.wassa-1.43 Findings of the WASSA 2024 EXALT shared task on explainability for cross-lingual emotion in tweets . In Proceedings of the 14th Workshop on Computational Approaches to Subjectiv...

  29. [37]

    Pedro Mart \'i n. 2003. The pragmatic rhetorical strategy of hedging in academic writing. Vigo International Journal of Applied Linguistics (VIAL), 0

  30. [38]

    Ammar Mohammed and Rania Kora. 2023. https://doi.org/10.1016/j.jksuci.2023.01.014 A comprehensive review on ensemble deep learning: Opportunities and challenges . Journal of King Saud University - Computer and Information Sciences, 35(2):757--774

  31. [39]

    Lily Ng, Anne Lauscher, Joel Tetreault, and Courtney Napoles. 2020. https://aclanthology.org/2020.argmining-1.13 Creating a domain-diverse corpus for theory-based argument quality assessment . In Proceedings of the 7th Workshop on Argument Mining, pages 117--126, Online. Assoc...

  32. [40]

    Nathan Ong, Diane Litman, and Alexandra Brusilovsky. 2014. https://doi.org/10.3115/v1/W14-2104 Ontology-based argument mining and automatic essay scoring . In Proceedings of the First Workshop on Argumentation Mining, pages 24--28, Baltimore, Maryland. Association for Computat...

  33. [41]

    Joonsuk Park and Claire Cardie. 2018. https://aclanthology.org/L18-1257 A corpus of e R ulemaking user comments for measuring evaluability of arguments . In Proceedings of the Eleventh International Conference on Language Resources and Evaluation ( LREC 2018) , Miyazaki, Japan...

  34. [42]

    Isaac Persing, Alan Davis, and Vincent Ng. 2010. https://aclanthology.org/D10-1023 Modeling organization in student essays . In Proceedings of the 2010 Conference on Empirical Methods in Natural Language Processing, pages 229--239, Cambridge, MA. Association for Computational ...

  35. [43]

    Isaac Persing and Vincent Ng. 2013. https://aclanthology.org/P13-1026 Modeling thesis clarity in student essays . In Proceedings of the 51st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 260--269, Sofia, Bulgaria. Association fo...

  36. [44]

    Prince, J

    E.F. Prince, J. Frader, and C. Bosk. 1982. On hedging in physician discourse. Linguistics and the Professions, Alex Publishing Corporation, pages 83--97

  37. [45]

    Litman, Richard Correnti, Lindsay Clare Matsumura, Elaine Wang, and Zahid Kisa

    Zahra Rahimi, Diane J. Litman, Richard Correnti, Lindsay Clare Matsumura, Elaine Wang, and Zahid Kisa. 2014. Automatic scoring of an analytical response-to-text assessment. In Intelligent Tutoring Systems, pages 601--610, Cham. Springer International Publishing

  38. [46]

    Liliana Mamani Sanchez and Carl Vogel. 2015. https://aclanthology.org/W15-0302 A hedging annotation scheme focused on epistemic phrases for informal language . In Proceedings of the Workshop on Models for Modality Annotation, London, UK. Association for Computational Linguistics

  39. [47]

    Skipper Seabold and Josef Perktold. 2010. statsmodels: Econometric and statistical modeling with python. In 9th Python in Science Conference

  40. [48]

    Edwin Simpson and Iryna Gurevych. 2018. https://doi.org/10.1162/tacl_a_00026 Finding convincing arguments using scalable B ayesian preference learning . Transactions of the Association for Computational Linguistics, 6:357--371

  41. [49]

    Manfred Stede. 2020. Automatic argumentation mining and the role of stance and sentiment. Journal of Argumentation in Context, 9(1):19--41

  42. [50]

    Reid Swanson, Brian Ecker, and Marilyn Walker. 2015. https://doi.org/10.18653/v1/W15-4631 Argument mining: Extracting arguments from online dialogue . In Proceedings of the 16th Annual Meeting of the Special Interest Group on Discourse and Dialogue, pages 217--226, Prague, Cze...

  43. [51]

    Chenhao Tan, Vlad Niculae, Cristian Danescu - Niculescu - Mizil, and Lillian Lee. 2016. https://arxiv.org/abs/1602.01103 Winning arguments: Interaction dynamics and persuasion strategies in good-faith online discussions . CoRR, abs/1602.01103

  44. [52]

    Assaf Toledo, Shai Gretz, Edo Cohen-Karlik, Roni Friedman, Elad Venezian, Dan Lahav, Michal Jacovi, Ranit Aharonov, and Noam Slonim. 2019. https://doi.org/10.18653/v1/D19-1564 Automatic argument quality assessment - new datasets and methods . In Proceedings of the 2019 Confere...

  45. [53]

    Enrica Troiano, Laura Oberländer, and Roman Klinger. 2023. https://doi.org/10.1162/coli_a_00461 Dimensional Modeling of Emotions in Text with Appraisal Theories: Corpus Creation, Annotation Reliability, and Prediction . Computational Linguistics, 49(1):1--72

  46. [54]

    Enrica Troiano, Sebastian Pad \'o , and Roman Klinger. 2019. https://doi.org/10.18653/v1/P19-1391 Crowdsourcing and validating event-focused emotion corpora for G erman and E nglish . In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, p...

  47. [55]

    Morgan Ulinski and Julia Hirschberg. 2019. https://doi.org/10.18653/v1/W19-4001 Crowdsourced hedge term disambiguation . In Proceedings of the 13th Linguistic Annotation Workshop, pages 1--5, Florence, Italy. Association for Computational Linguistics

  48. [56]

    Vasilieva

    I. Vasilieva. 2004. Gender-specific use of boosting and hedging adverbs in english computer-related texts -- a corpus-based study. In International Conference on Language, Politeness and Gender, pages 2--5

  49. [57]

    Henning Wachsmuth, Nona Naderi, Yufang Hou, Yonatan Bilu, Vinodkumar Prabhakaran, Tim Alberdingk Thijm, Graeme Hirst, and Benno Stein. 2017. https://aclanthology.org/E17-1017 Computational argumentation quality assessment in natural language . In Proceedings of the 15th Confer...

  50. [58]

    Zhongyu Wei, Yang Liu, and Yi Li. 2016. https://doi.org/10.18653/v1/P16-2032 Is this post persuasive? ranking argumentative comments in online forum . In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 195...

  51. [59]

    Margaryta Zaitseva. 2023. https://doi.org/10.12958/2227-2631-2023-1-47-152-162 Some observations on altering hedging phenomenon in courtroom discourse . LINGUISTICS, 1(47):152--162

  52. [60]

    Timon Ziegenbein, Shahbaz Syed, Felix Lange, Martin Potthast, and Henning Wachsmuth. 2023. https://doi.org/10.18653/v1/2023.acl-long.238 Modeling appropriate language in argumentation . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics ...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.