{"id":"2b924cab-71a9-453a-91f9-112a6ba46ad4","arxiv_id":"1908.08336","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The authors formalize a taxonomy of 37 recurring debate arguments and show that relevance to new topics can be predicted automatically using a dataset of 689 motions.","lead":"This paper defines a taxonomy of 37 recurring debate themes, called Classes of Principled Arguments, and a dataset of 689 controversial motions annotated for which themes apply. It shows that standard classifiers can predict the relevant themes for a new topic, supporting a new approach to automatic argument invention.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 6.2's external validation lacks a control condition: matching CoPA claims may be labelled 'implicit' merely because they are generic, so the claim that the taxonomy 'coincides with what professional debaters actually argue' is not established.","rationale":"I agree with the reader's weakest assumption and make it more precise by tying it to the missing control condition. The paper's internal contributions—the 37-class taxonomy, the 689-motion annotation with reported kappas, and the leave-one-out classifier evaluation—are credible and clearly documented. The construction guidelines and the generalization to unseen motions are genuine strengths. The problem is specifically the abstract's 'coincides with professional debaters' sentence: the speech study presents only matched CoPAs, counts 'implicit' as positive, and the authors themselves attribute the high implicit rate to generic phrasing. Without a control condition comparing non-matched CoPA claims or generic filler claims, the positive rates in Section 6.2 cannot establish that professional debaters specifically deploy these themes. The proposed control experiment is feasible and would settle the question. This does not change the reader's verdict: the paper already deserves CONDITIONAL, because the issue is addressable with additional analysis and does not undermine the taxonomy's internal coherence or the classifier results.","tokens_in":14102,"tokens_out":3445,"duration_ms":35776,"concrete_test":"Run a controlled annotation on a sample of the 184 speeches: for each speech, alongside the matched CoPA claims, present (a) claims from randomly selected non-matching CoPAs and (b) generic control claims not from any CoPA but matched for length and style (e.g., 'the current system has serious problems and must be reconsidered'). Use the same explicit/implicit/not-mentioned protocol and majority-vote scoring as Section 4.3. The Section 6.2 claim survives only if the matched-condition positive rate (39% excluding general CoPAs, or 66% overall) is significantly above both control conditions; if the control conditions produce comparable positive rates, the external validation is not informative.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central external claim in the abstract—that the taxonomy 'coincides with what professional debaters actually argue'—rests on the speech-annotation study in Sections 4.3 and 6.2, which has no control condition. For each speech, annotators are shown only claims from CoPAs that the authors' own annotation already matched to the speech's motion, and a claim is counted as positive if it is merely 'implicit' in the speech; only 5% of the (speech, claim) pairs are explicit mentions. Section 6.2 itself concedes that the high implicit rate 'is probably due to the rather generic phrasing of the claims, which in the first place were constructed to be applicable as-is in multiple contexts.' If the positive labels are driven by generic applicability rather than by specific debating practice, the study cannot distinguish 'debaters use these arguments' from 'these are plausible general statements that can be read into almost any speech.' The opposing-stance 5% result shows only that annotators are not labelling everything positive; it does not measure specificity to the matched CoPA. Thus the strongest externally falsifiable component of the paper is not established by the current design. The dataset being omitted from the preprint is a separate, secondary issue; even with the data released, the missing control remains.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a taxonomy of 37 'Classes of Principled Arguments' (CoPAs), each consisting of two opposing generic claims, and annotates 689 debate motions for membership in these classes. It reports inter-annotator agreement (kappa 0.60-0.78) for the motion-CoPA annotations, evaluates several classifiers for automatic motion-CoPA matching under leave-one-motion-out, and conducts a speech-based study to argue that the CoPA claims are actually used by professional debaters. The authors claim the taxonomy is coherent, covers most motions, coincides with professional debating practice, and facilitates automatic argument invention. The main technical contributions are the formal definition of CoPAs, the annotated dataset, and a comparative evaluation of matching methods.","tokens_in":14271,"tokens_out":2896,"duration_ms":30768,"significance":"If the central claims are established, this is a valuable resource for argument invention, connecting classical rhetorical topoi with modern computational argumentation. The paper's strengths include a clear formalization of the problem, a carefully constructed dataset with crowd-sourced validation, and a multi-method evaluation with a naive baseline. The explicit acknowledgment that the taxonomy is a 'first attempt' and the discussion of limitations (e.g., the difficulty of phrasing claims without context) are commendable. However, the external validation against professional speeches is currently under-controlled, and the dataset itself is not available in the preprint, which limits independent verification. The core resource and matching experiments appear sound and publishable after addressing these issues.","major_comments":[{"comment":"The external validation of the claim that the taxonomy 'coincides with what professional debaters actually argue' is not established because the annotation protocol lacks a control condition. For each speech, annotators were shown only claims from CoPAs that the authors' own annotation had already matched to the speech's motion, and a claim was counted as positive if it was merely implicit. As the paper itself notes in Section 6.2, the high positive rate 'is probably due to the rather generic phrasing of the claims.' Because generic claims could be judged implicit in almost any speech, the study cannot distinguish 'debaters use these arguments' from 'these are plausible general statements.' A proper test would require annotating non-matching CoPAs or generic control claims, or measuring specificity by comparing matched versus unmatched CoPAs for the same speech.","section":"Sections 4.3 and 6.2"},{"comment":"The 5% positive rate for opposing-stance claims is cited as evidence of annotation quality, but it does not address the specificity concern. It shows only that annotators do not label every claim positive. Without a control set of non-matching claims, the result is compatible with a process that labels claims positive based on general topical relevance rather than on the specific CoPA, so the opposing-stance statistic does not rescue the external validation.","section":"Section 6.2, opposing-stance result"},{"comment":"The manuscript repeatedly refers to supplementary material (Sections 4.1, 4.2, 6.2, and Appendix B) but the full dataset is not included in the arXiv preprint. Since the paper's contribution is a dataset and taxonomy, the absence of the data (or a clear, working link) makes the kappa statistics and the leave-one-out classifier evaluation impossible to verify independently. Please provide the complete annotation set and all code needed to reproduce the experiments in the revision.","section":"Data availability and reproducibility"}],"minor_comments":[{"comment":"The caption states that distance between vertices is 'indicative of intersection size,' but no scale or visual legend is provided; please add a legend or a brief explanation of how distance maps to intersection size.","section":"Figure 2"},{"comment":"The definition of a motion as (action, topic) is clear for policy actions, but for analysis actions such as 'brings more harm than good' the mapping is less intuitive; please clarify how these actions are interpreted in the framework.","section":"Section 3"},{"comment":"The parameter k=5 is set for BA-k, but no sensitivity analysis is reported; please state whether the results are robust to this choice or add a short analysis.","section":"Section 5, BA-k"}],"recommendation":"major_revision","confidential_remarks":"The paper's internal resource construction is solid and the matching experiments are informative. The main risk is the overstatement of the external validation: the speech study as designed cannot support the abstract's claim about professional debaters without a control condition. Ensuring data availability is also essential for a resource paper. These are fixable with additional experiments and a revised interpretation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, the genuinely new thing: the paper defines 37 Classes of Principled Arguments (CoPAs), annotates 689 debate motions with CoPA memberships, and evaluates several baseline classifiers for automatic matching. The taxonomy is documented in an appendix, the internal annotation agreement is respectable (kappa 0.60–0.78), and the leave-one-out classifier evaluation is appropriate. The related work is well placed against topic-specific argument databases and argumentation schemes. This is a reusable resource for computational argumentation, and the idea of formalizing first-principles debating as a small set of recurring themes is clearly a useful step.\n\nThe soft spot is exactly where the stress test lands. The abstract claims the taxonomy 'coincides with what professional debaters actually argue in their speeches,' but the speech-annotation study in Section 4.3/6.2 does not establish that. Annotators are shown claims only from CoPAs that the authors themselves already matched to the speech's motion, and 'implicit' counts as positive; only 5% of pairs are explicit. Section 6.2 itself concedes that the high implicit rate is probably due to the deliberately generic phrasing of the claims. With no control condition—no non-matching CoPAs, no generic claims—the study cannot distinguish 'debaters use these arguments' from 'these are plausible general statements that can be read into almost any speech.' That is a load-bearing flaw for the headline claim, though not for the resource itself.\n\nThe classifier results are credible: the ensemble reaches 86% precision for its top prediction at half coverage, and dropping the three general CoPAs still gives 75%. The dataset is not actually included in the preprint; the authors point to a supplementary website, but for a resource paper the data should be available to reviewers and readers. That is a secondary but real issue.\n\nCitation pattern looks fine—no obvious self-citation inflation, and the related work on argumentation schemes and framing is apt. The paper shows clear thinking; the external validation problem is one of design, not of incoherence.\n\nWho is this for? Anyone working on argument generation, debate systems, or computational rhetoric. It deserves a serious referee: the resource is valuable, the internal evidence holds up, and the external validation can be fixed with a control condition and a softened abstract claim. I would send it to peer review, but I would not accept it as-is.","headline":"A genuinely useful taxonomy and dataset for argument invention, with a solid internal evaluation but an external validation that does not support the abstract's strongest claim.","tokens_in":14899,"tokens_out":1631,"would_cite":true,"duration_ms":18747,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A small taxonomy of recurring argument themes can cover most debate motions and be matched automatically.","keywords":["argument invention","first principles","debate motions","taxonomy of arguments","motion-CoPA matching","parliamentary debate","natural language processing","argument generation"],"falsifier":"Show annotators a set of speeches paired with (a) the CoPA claims the paper matched to each motion, (b) claims from CoPAs judged non-matching for that motion, and (c) generic claims unattached to any CoPA; if the implicit-mention rates for (b) and (c) are close to those for (a), the taxonomy's agreement with professional debate would be explained by generic wording rather than by thematic coverage.","tokens_in":13839,"feed_emoji":"⚖️","tokens_out":8073,"duration_ms":69784,"temperature":0.7,"pith_summary":"This paper tries to establish that competitive debaters' 'first principles' can be made explicit: a small taxonomy of recurring argument themes, each carrying two opposing commonplace claims, can cover most debate motions and can be matched to a new motion automatically. The authors define 37 such Classes of Principled Arguments (CoPAs) over 689 motions and report that 87% of motions belong to at least one CoPA, with crowd annotators agreeing on the matching. They also report that claims from matched CoPAs are often at least implicit in professional debate speeches, which they take as evidence that the taxonomy reflects real debating practice rather than being an artificial construct. If the claim holds, a debater or system facing an unfamiliar topic can be given ready-made arguments and counterarguments almost immediately, which is the practical promise of argument invention.","feed_headline":"37 recurring arguments cover 87 percent of debate motions","feed_subtitle":"If true, debaters get ready-made pro and con claims for unfamiliar topics automatically.","key_machinery":"The central object is the Class of Principled Arguments (CoPA), defined as a pair $c=(A,M)$, where $A$ is a set of two concise claims of opposing stance toward the class's theme and $M$ is the set of motions for which those claims are plausible. A motion is a pair (action, topic), such as (ban, smoking), and a CoPA 'matches' a motion when its claims can plausibly be made in deliberating that motion. This structure carries the argument because matching a new motion to a CoPA immediately yields two ready-made argumentative claims, one for each side, which can be instantiated by replacing the special [TOPIC] token inside a claim with the motion's topic. The same (motion, CoPA) pairs also define the supervised learning task used to predict membership for new motions.","core_discovery":"The paper's central claim is that a Class of Principled Arguments — a pair of concise opposing claims about a recurring theme, together with the set of motions to which those claims plausibly apply — is a workable unit for automatic argument invention. On its own terms, the taxonomy is a 'first attempt' rather than a finished theory, but it claims the basic properties are already in place: 87% of 689 motions matched at least one CoPA; the average motion matched about 1.95 CoPAs; annotators agreed on the matching with reasonably high kappa; a classifier ensemble reached 86% precision for the highest-scoring CoPA at a threshold that yields a prediction for half the motions; and in recorded professional speeches, 66% of aligned (speech, claim) pairs were judged positive, mostly as implicit mentions. The paper therefore treats motion-to-CoPA matching as the actionable core: once a motion is matched, its two claims supply the thesis and antithesis around which deliberation can be built.","pith_inferences":["Editorial inference: the motion-to-CoPA matchers could be coupled with a claim-generation model to make a practical debate-preparation tool, a step the paper describes as future work.","Editorial inference: the 'coincides with professional debaters' claim would be strengthened by control conditions comparing matched CoPA claims with non-matching CoPA claims and with generic claims in the same speeches; the paper does not report such a comparison.","Editorial inference: the same CoPA structure could extend outside parliamentary debate, for example to essay writing or policy analysis, where a writer would be offered the two sides of a recurring clash relevant to their topic."],"forward_implications":["A debater facing an unfamiliar motion can be handed at least one relevant CoPA almost immediately: the ensemble matcher returns a correct highest-scoring CoPA for 86% of motions at a threshold covering half of motions.","Argument invention becomes a two-stage operation: automatically match the motion to CoPAs, then instantiate the CoPA's two opposing claims (filling in the [TOPIC] token) to obtain pro and con arguments.","The paper's dataset of 689 motions and 37 CoPAs gives argument-generation systems a reusable collection of claims known to be plausible across many topics.","Because most professional-speech matches are implicit rather than explicit, the paper suggests that first-principles arguments operate below the surface of skilled debate, making them a reliable foundation for open-domain argument assistance."],"supporting_citations":[{"why":"It defines the argument-invention task and the Carneades approach that this paper extends from topic-specific argumentation to reusable principled arguments.","marker":"Walton and Gordon (2012, 2017)"},{"why":"It supplies the recorded professional debate speeches used as external evidence that CoPA claims are at least implicit in real debating.","marker":"Mirkin et al. (2018)"},{"why":"It provides the Wikipedia concept similarity measure used by the nearest-neighbour motion matcher.","marker":"Ein Dor et al. (2018)"},{"why":"It supplies the concept-related sentence retrieval technique that underpins the Naive Bayes and recurrent neural network matchers.","marker":"Rabinovich et al. (2018)"},{"why":"It represents the earlier claim-synthesis-by-predicate-recycling approach against which the CoPA framework is positioned as a richer basis for de novo argument generation.","marker":"Bilu and Slonim (2016)"}],"fun_headline_variants":["37 first-principles argument pairs map to 87% of debate motions","Taxonomy of 37 recurring arguments covers 87% of debate topics","37 argument pairs supply both sides for 87% of debate topics","First-principles taxonomy matches 87% of debate motions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that the taxonomy 'coincides with what professional debaters actually argue' rests on the assumption that a CoPA claim marked as at least implicit in a speech is evidence of real use, even though the speech's motion was already matched to that CoPA by the authors' own annotation and the claims were deliberately written in generic language.","fun_headline_variants_meta":{"raw":{"variants":["37 first-principles argument pairs map to 87% of debate motions","Taxonomy of 37 recurring arguments covers 87% of debate topics","37 argument pairs supply both sides for 87% of debate topics","First-principles taxonomy matches 87% of debate motions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00096,"raw_usage":{"total_tokens":4070,"prompt_tokens":907,"completion_tokens":3163,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":523,"completion_tokens_details":{"reasoning_tokens":3088}},"tokens_in":523,"tokens_out":3163,"duration_ms":25659,"temperature":1.0,"reasoning_tokens":3088,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:42:33.258010+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Show annotators a set of speeches paired with (a) the CoPA claims the paper matched to each motion, (b) claims from CoPAs judged non-matching for that motion, and (c) generic claims unattached to any CoPA; if the implicit-mention rates for (b) and (c) are close to those for (a), the taxonomy's agreement with professional debate would be explained by generic wording rather than by thematic coverage.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It defines the argument-invention task and the Carneades approach that this paper extends from topic-specific argumentation to reusable principled arguments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the recorded professional debate speeches used as external evidence that CoPA claims are at least implicit in real debating."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It provides the Wikipedia concept similarity measure used by the nearest-neighbour motion matcher."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the concept-related sentence retrieval technique that underpins the Naive Bayes and recurrent neural network matchers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It represents the earlier claim-synthesis-by-predicate-recycling approach against which the CoPA framework is positioned as a richer basis for de novo argument generation."}],"review_version":1}