Pith. sign in

REVIEW 3 major objections 3 minor 39 references

Argument Invention from First Principles

T0 review · 3 major / 3 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read A small taxonomy of recurring argument themes can cover most debate motions and be matched automatically.

desk verdict A genuinely useful taxonomy and dataset for argument invention, with a solid internal evaluation but an external validation that does not support the abstract's strongest claim. read the letter →

arxiv 1908.08336 v1 pith:SDA4XTZ5 submitted 2019-08-22 cs.CL

classification cs.CL
keywords argumentinventionfirstprinciplesdebatemotionstaxonomyofargumentsmotion-CoPAmatchingparliamentarynaturallanguageprocessinggeneration
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that competitive debaters' 'first principles' can be made explicit: a small taxonomy of recurring argument themes, each carrying two opposing commonplace claims, can cover most debate motions and can be matched to a new motion automatically. The authors define 37 such Classes of Principled Arguments (CoPAs) over 689 motions and report that 87% of motions belong to at least one CoPA, with crowd annotators agreeing on the matching. They also report that claims from matched CoPAs are often at least implicit in professional debate speeches, which they take as evidence that the taxonomy reflects real debating practice rather than being an artificial construct. If the claim holds, a debater or system facing an unfamiliar topic can be given ready-made arguments and counterarguments almost immediately, which is the practical promise of argument invention.

What carries the argument

The central object is the Class of Principled Arguments (CoPA), defined as a pair $c=(A,M)$, where $A$ is a set of two concise claims of opposing stance toward the class's theme and $M$ is the set of motions for which those claims are plausible. A motion is a pair (action, topic), such as (ban, smoking), and a CoPA 'matches' a motion when its claims can plausibly be made in deliberating that motion. This structure carries the argument because matching a new motion to a CoPA immediately yields two ready-made argumentative claims, one for each side, which can be instantiated by replacing the special [TOPIC] token inside a claim with the motion's topic. The same (motion, CoPA) pairs also define the supervised learning task used to predict membership for new motions.

What would settle it

Show annotators a set of speeches paired with (a) the CoPA claims the paper matched to each motion, (b) claims from CoPAs judged non-matching for that motion, and (c) generic claims unattached to any CoPA; if the implicit-mention rates for (b) and (c) are close to those for (a), the taxonomy's agreement with professional debate would be explained by generic wording rather than by thematic coverage.

Watch

Extended reading notes

Core claim

The paper's central claim is that a Class of Principled Arguments — a pair of concise opposing claims about a recurring theme, together with the set of motions to which those claims plausibly apply — is a workable unit for automatic argument invention. On its own terms, the taxonomy is a 'first attempt' rather than a finished theory, but it claims the basic properties are already in place: 87% of 689 motions matched at least one CoPA; the average motion matched about 1.95 CoPAs; annotators agreed on the matching with reasonably high kappa; a classifier ensemble reached 86% precision for the highest-scoring CoPA at a threshold that yields a prediction for half the motions; and in recorded professional speeches, 66% of aligned (speech, claim) pairs were judged positive, mostly as implicit mentions. The paper therefore treats motion-to-CoPA matching as the actionable core: once a motion is matched, its two claims supply the thesis and antithesis around which deliberation can be built.

Load-bearing premise

The claim that the taxonomy 'coincides with what professional debaters actually argue' rests on the assumption that a CoPA claim marked as at least implicit in a speech is evidence of real use, even though the speech's motion was already matched to that CoPA by the authors' own annotation and the claims were deliberately written in generic language.

Editorial extensions

If this is right

  • A debater facing an unfamiliar motion can be handed at least one relevant CoPA almost immediately: the ensemble matcher returns a correct highest-scoring CoPA for 86% of motions at a threshold covering half of motions.
  • Argument invention becomes a two-stage operation: automatically match the motion to CoPAs, then instantiate the CoPA's two opposing claims (filling in the [TOPIC] token) to obtain pro and con arguments.
  • The paper's dataset of 689 motions and 37 CoPAs gives argument-generation systems a reusable collection of claims known to be plausible across many topics.
  • Because most professional-speech matches are implicit rather than explicit, the paper suggests that first-principles arguments operate below the surface of skilled debate, making them a reliable foundation for open-domain argument assistance.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the motion-to-CoPA matchers could be coupled with a claim-generation model to make a practical debate-preparation tool, a step the paper describes as future work.
  • Editorial inference: the 'coincides with professional debaters' claim would be strengthened by control conditions comparing matched CoPA claims with non-matching CoPA claims and with generic claims in the same speeches; the paper does not report such a comparison.
  • Editorial inference: the same CoPA structure could extend outside parliamentary debate, for example to essay writing or policy analysis, where a writer would be offered the two sides of a recurring clash relevant to their topic.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 3 minor

Summary. The paper introduces a taxonomy of 37 'Classes of Principled Arguments' (CoPAs), each consisting of two opposing generic claims, and annotates 689 debate motions for membership in these classes. It reports inter-annotator agreement (kappa 0.60-0.78) for the motion-CoPA annotations, evaluates several classifiers for automatic motion-CoPA matching under leave-one-motion-out, and conducts a speech-based study to argue that the CoPA claims are actually used by professional debaters. The authors claim the taxonomy is coherent, covers most motions, coincides with professional debating practice, and facilitates automatic argument invention. The main technical contributions are the formal definition of CoPAs, the annotated dataset, and a comparative evaluation of matching methods.

Significance. If the central claims are established, this is a valuable resource for argument invention, connecting classical rhetorical topoi with modern computational argumentation. The paper's strengths include a clear formalization of the problem, a carefully constructed dataset with crowd-sourced validation, and a multi-method evaluation with a naive baseline. The explicit acknowledgment that the taxonomy is a 'first attempt' and the discussion of limitations (e.g., the difficulty of phrasing claims without context) are commendable. However, the external validation against professional speeches is currently under-controlled, and the dataset itself is not available in the preprint, which limits independent verification. The core resource and matching experiments appear sound and publishable after addressing these issues.

major comments (3)
  1. [Sections 4.3 and 6.2] The external validation of the claim that the taxonomy 'coincides with what professional debaters actually argue' is not established because the annotation protocol lacks a control condition. For each speech, annotators were shown only claims from CoPAs that the authors' own annotation had already matched to the speech's motion, and a claim was counted as positive if it was merely implicit. As the paper itself notes in Section 6.2, the high positive rate 'is probably due to the rather generic phrasing of the claims.' Because generic claims could be judged implicit in almost any speech, the study cannot distinguish 'debaters use these arguments' from 'these are plausible general statements.' A proper test would require annotating non-matching CoPAs or generic control claims, or measuring specificity by comparing matched versus unmatched CoPAs for the same speech.
  2. [Section 6.2, opposing-stance result] The 5% positive rate for opposing-stance claims is cited as evidence of annotation quality, but it does not address the specificity concern. It shows only that annotators do not label every claim positive. Without a control set of non-matching claims, the result is compatible with a process that labels claims positive based on general topical relevance rather than on the specific CoPA, so the opposing-stance statistic does not rescue the external validation.
  3. [Data availability and reproducibility] The manuscript repeatedly refers to supplementary material (Sections 4.1, 4.2, 6.2, and Appendix B) but the full dataset is not included in the arXiv preprint. Since the paper's contribution is a dataset and taxonomy, the absence of the data (or a clear, working link) makes the kappa statistics and the leave-one-out classifier evaluation impossible to verify independently. Please provide the complete annotation set and all code needed to reproduce the experiments in the revision.
minor comments (3)
  1. [Figure 2] The caption states that distance between vertices is 'indicative of intersection size,' but no scale or visual legend is provided; please add a legend or a brief explanation of how distance maps to intersection size.
  2. [Section 3] The definition of a motion as (action, topic) is clear for policy actions, but for analysis actions such as 'brings more harm than good' the mapping is less intuitive; please clarify how these actions are interpreted in the framework.
  3. [Section 5, BA-k] The parameter k=5 is set for BA-k, but no sensitivity analysis is reported; please state whether the results are robust to this choice or add a short analysis.

Circularity Check

1 steps flagged · score 4.0 of 10

Speech-validation claim is partially circular: claims were built to be generic/as-is and then counted as 'implicit' in speeches, so the external check partly re-expresses the construction property.

  1. self definitional [Section 4.3 and Section 6.2 (CoPA claims in recorded speeches)]
    "For each motion we extracted the CoPAs to which it belongs according to our annotation ... They were asked whether each claim was (i) explicitly made by the speaker, was (ii) implicit in the speech or was (iii) not mentioned at all. ... This is probably due to the rather generic phrasing of the claims, which in the first place were constructed to be applicable 'as-is' in multiple contexts."

    The abstract's claim that the taxonomy 'coincides with what professional debaters actually argue' rests on this speech study. The study only presents claims from CoPAs that the authors' own annotation already matched to the motion, and a positive label includes 'implicit'. The authors explain the high implicit rate by the fact that the claims were constructed to be applicable as-is across contexts. Thus the positive labels partly re-express the construction property of generic applicability rather than an independent measurement of actual debating practice. There is no control condition with non-matching CoPAs or with deliberately generic distractor claims; the 5% opposing-stance result shows that annotators attend to stance, but it does not establish specificity to the matched CoPA.

full rationale

The central taxonomy construction and the automatic motion-CoPA matching pipeline are not circular: the CoPAs are manually authored, the motion-CoPA labels are human annotations, and the matching classifiers are evaluated in a leave-one-motion-out framework on held-out motions. The speech-based external validation, however, has a partial circular element. For each speech, the claims shown to annotators come only from CoPAs that the authors' own annotation had already matched to the motion, and a claim is counted as positive if it is merely 'implicit'. The paper itself concedes that the high implicit rate 'is probably due to the rather generic phrasing of the claims, which in the first place were constructed to be applicable as-is in multiple contexts.' This means the validation cannot cleanly distinguish 'debaters use these arguments' from 'these are broad statements that can be read into almost any speech'. The limitation is real but does not infect the whole paper: the leave-one-out classifier results, the inter-annotator agreement on motion-CoPA membership, and the manual taxonomy itself stand on independent evidence. No load-bearing self-citation chain was found; the cited prior work is used as data or method, not as a substitute for proof. Overall, the circularity score is 4 because the abstract's strongest external claim is partially supported by a design that is partly self-confirming, while the rest of the contribution retains independent content.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central contribution is a manually-authored taxonomy and dataset, so the load-bearing assumptions are about data representativeness and annotation ground truth rather than fitted mathematical constants. No scalar constants are fitted to data; the classifier hyperparameters (e.g., BA-k support k=5, KNN threshold 0.5) are design choices for the evaluation, not for the taxonomy itself.

assumptions (3)
  • domain assumption The 689 motions are a representative sample of the universe of debatable motions.
    All coverage and generalization claims (87% of motions match a CoPA; an average of 1.95 CoPAs per motion) are computed on this self-selected set; Section 4.2 describes collecting 589 additional motions but no sampling procedure or external yardstick.
  • domain assumption The two annotators' consensus on CoPA-motion relevance is an adequate ground truth.
    The dataset labels come from two annotators, with crowd-sourced validation on a sample (kappa 0.60 to 0.78). This supports consistency but not whether another team would produce the same taxonomy.
  • domain assumption Word2vec embeddings and Wikipedia link-based similarity capture topic relatedness for the matching task.
    The KNN and LR methods rely on these semantic relatedness measures without task-specific calibration, so their quality constrains the reported matching precision.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Argument Invention from First Principles." pith.science (2026). https://pith.science/paper/SDA4XTZ5

@misc{pith2026190808336,
  author       = {Pith},
  title        = {Pith review of: Argument Invention from First Principles},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SDA4XTZ5}},
  note         = {Machine review of arXiv:1908.08336}
}
read the original abstract

Competitive debaters often find themselves facing a challenging task -- how to debate a topic they know very little about, with only minutes to prepare, and without access to books or the Internet? What they often do is rely on "first principles", commonplace arguments which are relevant to many topics, and which they have refined in past debates. In this work we aim to explicitly define a taxonomy of such principled recurring arguments, and, given a controversial topic, to automatically identify which of these arguments are relevant to the topic. As far as we know, this is the first time that this approach to argument invention is formalized and made explicit in the context of NLP. The main goal of this work is to show that it is possible to define such a taxonomy. While the taxonomy suggested here should be thought of as a "first attempt" it is nonetheless coherent, covers well the relevant topics and coincides with what professional debaters actually argue in their speeches, and facilitates automatic argument invention for new topics.

Figures

Figures reproduced from arXiv: 1908.08336 by the authors.

Figure 1
Figure 1. Distribution of the number of motions per [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Graph of CoPAs, where edges indicate non-empty intersection and distance between vertices is indicative [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Fraction of overlapping motions among classes; the value of entry (i,j) is the fraction of motions in class i which also appear in class j. Green indicates high values, red low ones. the ensemble method for a threshold yielding a prediction for half the motions attains a precision of 75% ( [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (2 more)
Figure 5
Figure 5. Figure 5: P@1 vs. coverage of the various motion￾CoPA matching methods [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Precision results when ignoring the three gen [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

39 extracted references · 34 canonical work pages

  1. [1]

    URL: " 'urlintro :=

    ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRINGS urlintro eprinturl eprintpr...

  2. [2]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...

  3. [3]

    Khalid Al-Khatib, Henning Wachsmuth, Matthias Hagen, Jonas K \"o hler, and Benno Stein. 2016. Cross-domain mining of argumentative text through distant supervision. In Proceedings of the 2016 conference of the north american chapter of the association for computational linguistics: human language technologies, pages 1395--1404

  4. [4]

    Bob Altemeyer. 1981. Right-wing authoritarianism. University of Manitoba press

  5. [5]

    Aristotle and George Alexander Kennedy. 1991. On rhetoric: A theory of civic discourse. Oxford University Press New York

  6. [6]

    Yonatan Bilu and Noam Slonim. 2016. Claim synthesis via predicate recycling. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), volume 2, pages 525--530

  7. [7]

    Boydstun, Justin H

    Dallas Card, Amber E. Boydstun, Justin H. Gross, Philip Resnik, and Noah A. Smith. 2015. The media frames corpus: Annotations of frames across issues. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Short Papers), pages 438–--444

  8. [8]

    Steffen Eger, Johannes Daxenberger, and Iryna Gurevych. 2017. Neural end-to-end learning for computational argumentation mining. In Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), volume 1, pages 11--22

Show all 39 references
  1. [9]

    Liat Ein Dor, Alon Halfon, Yoav Kantor, Ran Levy, Yosi Mass, Ruty Rinott, Eyal Shnarch, and Noam Slonim. 2018. Semantic relatedness of wikipedia concepts-benchmark data and a working solution. In Proceedings of the Eleventh International Conference on Language Resources and Ev...

  2. [10]

    Jim AC Everett. 2013. The 12 item social and economic conservatism scale (secs). PloS one, 8(12):e82131

  3. [11]

    Cheryl Glenn, Melissa A Goldthwaite, and Robert Connors. 2008. The St. Martin's guide to teaching writing. Bedford/St. Martin's

  4. [12]

    Ivan Habernal and Iryna Gurevych. 2016. Which argument is more convincing? analyzing and predicting convincingness of web arguments using bidirectional lstm. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), vol...

  5. [13]

    John T Jost, Sally Blount, Jeffrey Pfeffer, and Gy \"o rgy Hunyady. 2003. Fair market ideology: Its cognitive-motivational underpinnings. Research in organizational behavior, 25:53--91

  6. [14]

    Christian Kock. 2009. Choice is not true or false: The domain of rhetorical argumentation. Argumentation, 23(1):61--80

  7. [15]

    Janice M Lauer. 2004. Invention in rhetoric and composition. Parlor Press LLC

  8. [16]

    Ran Levy, Yonatan Bilu, Daniel Hershcovich, Ehud Aharoni, and Noam Slonim. 2014. Context dependent claim detection. In Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers, pages 1489--1500

  9. [17]

    Ran Levy, Ben Bogin, Shai Gretz, Ranit Aharonov, and Noam Slonim. 2018. Towards an argumentative content search engine using weak supervision. In Proceedings of the 27th International Conference on Computational Linguistics, pages 2066--2081

  10. [18]

    Marco Lippi and Paolo Torroni. 2015. Argument mining: A machine learning perspective. In International Workshop on Theory and Applications of Formal Argumentation, pages 163--176. Springer

  11. [19]

    Marco Lippi and Paolo Torroni. 2016. Argumentation mining: State of the art and emerging trends. ACM Transactions on Internet Technology (TOIT), 16(2):10

  12. [20]

    Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. http://arxiv.org/abs/1301.3781 Efficient estimation of word representations in vector space . CoRR, abs/1301.3781

  13. [21]

    Shachar Mirkin, Guy Moshkowich, Matan Orbach, Lili Kotlerman, Yoav Kantor, Tamar Lavee, Michal Jacovi, Yonatan Bilu, Ranit Aharonov, and Noam Slonim. 2018. Listening comprehension over argumentative content. In Proceedings of the 2018 Conference on Empirical Methods in Natural...

  14. [22]

    Raquel Mochales Palau and Marie-Francine Moens. 2009. Argumentation mining: the detection, classification and structure of arguments in text. In Proceedings of the 12th international conference on artificial intelligence and law, pages 98--107. ACM

  15. [23]

    Rebecca J Passonneau and Bob Carpenter. 2014. The benefits of a model of annotation. Transactions of the Association for Computational Linguistics, 2:311--326

  16. [24]

    Chaim Perelman. 1971. The new rhetoric. In Pragmatics of natural languages, pages 145--149. Springer

  17. [25]

    Ella Rabinovich, Benjamin Sznajder, Artem Spector, Ilya Shnayderman, Ranit Aharonov, David Konopnicki, and Noam Slonim. 2018. http://aclweb.org/anthology/D18-1522 Learning concept abstractness using weak supervision . In Proceedings of the 2018 Conference on Empirical Methods ...

  18. [26]

    Niklas Rach, Saskia Langhammer, Wolfgang Minker, and Stefan Ultes. 2018. Utilizing argument mining techniques for argumentative dialogue systems. In Proceedings of the 9th International Workshop On Spoken Dialogue Systems (IWSDS)

  19. [27]

    Chris Reed and Glenn Rowe. 2004. Araucaria: Software for argument analysis, diagramming and representation. International Journal on Artificial Intelligence Tools, 13(04):961--979

  20. [28]

    Ruty Rinott, Lena Dankin, Carlos Alzate Perez, Mitesh M Khapra, Ehud Aharoni, and Noam Slonim. 2015. Show me your evidence-an automatic method for context dependent evidence detection. In Proceedings of the 2015 conference on empirical methods in natural language processing, p...

  21. [29]

    Holli A Semetko and Patti M Valkenburg. 2000. Framing european politics: A content analysis of press and television news. Journal of communication, 50(2):93--109

  22. [30]

    Eyal Shnarch, Carlos Alzate, Lena Dankin, Martin Gleize, Yufang Hou, Leshem Choshen, Ranit Aharonov, and Noam Slonim. 2018. Will it blend? blending weak and strong labeled data in a neural network for argumentation mining. In Proceedings of the 56th Annual Meeting of the Assoc...

  23. [31]

    Jim Sidanius and Felicia Pratto. 2001. Social dominance: An intergroup theory of social hierarchy and oppression. Cambridge University Press

  24. [32]

    Tim Sonnreich. 2012. Monash association of debaters guide to debating (tips, tactics and first principles)

  25. [33]

    Christian Stab and Iryna Gurevych. 2014. Annotating argument components and relations in persuasive essays. In Proceedings of COLING 2014, the 25th International Conference on Computational Linguistics: Technical Papers, pages 1501--1510

  26. [34]

    de Vreese

    Claes H. de Vreese. 2005. News framing: Theory and typology. Information Design Journal, 13:51--62

  27. [35]

    Douglas Walton. 2013. Argumentation schemes for presumptive reasoning. Routledge

  28. [36]

    Douglas Walton and Thomas F Gordon. 2012. The carneades model of argument invention. Pragmatics & Cognition, 20(1):1--31

  29. [37]

    Douglas Walton and Thomas F Gordon. 2017. Argument invention with the carneades argumentation system. SCRIPTed, 14:168

  30. [38]

    Douglas Walton and Thomas F Gordon. 2018. How computational tools can help rhetoric and informal logic with argument invention. Argumentation, pages 1--27

  31. [39]

    Douglas Walton, Christopher Reed, and Fabrizio Macagno. 2008. Argumentation schemes. Cambridge University Press

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.