{"id":"950131ed-aad1-438d-aabe-6a645aff57a5","arxiv_id":"1909.00393","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A new dataset of 55 general-purpose claim-rebuttal pairs, annotated across 50 debate motions and 200 speeches, shows that generic rebuttals are often plausible and that claim-frequency priors outperform text-based baselines.","lead":"This paper introduces GPR-KB-55, a dataset of 55 general-purpose claims and rebuttals that apply across many debate topics, along with annotations showing where such claims appear in 200 speeches. It matters because it provides a reusable resource for building systems that automatically rebut long argumentative texts, and it shows that even a simple frequency baseline beats modern NLP models on this task.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 87% rebuttal-plausibility result lacks a control condition and a stable denominator; without comparing to generic rebuttals and reporting the variable number of raters per pair, it cannot support the central claim.","rationale":"The paper's contribution is a dataset and a manual knowledge base, and that contribution is real. The central overclaim is in Section 4.4 and the abstract: a 55-pair knowledge base with 103 plausibility judgements is taken to show that general-purpose rebuttals are viable in the vast majority of contexts. The reader's conditional verdict is appropriate because the weakness is in the evaluation, not in the data construction. I agree with the reader's weakest assumption and would keep the verdict CONDITIONAL or, equivalently, UNCHANGED: the authors should add a control annotation, report rater denominators and confidence intervals, and ideally re-annotate mention labels with experts before the viability claim can be accepted. No internal contradiction or misconduct is present; the inference is simply under-supported as reported.","tokens_in":11775,"tokens_out":5688,"duration_ms":56385,"concrete_test":"Re-run the Section 4.4 annotation on the same 103 speech-rebuttal pairs with three blind conditions: (a) the actual GPR rebuttal, (b) a length-matched generic filler sentence, and (c) a rebuttal drawn from a different GPR claim. Use the same crowd pool and majority rule, but record the number of annotators who reach the rebuttal stage in each condition. If the plausible rate for condition (b) or (c) is statistically indistinguishable from 87% (e.g., within the 95% Wilson interval), the central claim is not supported. Also compute the 95% confidence interval for the 87% and report per-condition rater counts; if fewer than 10 raters rate a large share of pairs, the majority statistic is unstable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that 'most rebuttals are appropriate in the vast majority of contexts' rests on one number: 87% of 103 speech-rebuttal pairs rated plausible by majority vote (Section 4.4). That number is load-bearing and under-supported for three reasons. First, there is no control condition: annotators are not shown generic rebuttals, non-rebuttal texts, or rebuttals matched to the wrong claim, so the high 'plausible' rate may reflect acquiescence or a low bar rather than the specific GPR-KB content. Second, the 103 pairs are sampled from speeches already labeled as mentioning the claim in Section 4.2, and the rebuttal is shown only to annotators who again marked the claim as mentioned; the paper does not report how many annotators rated each rebuttal, so the 87% is a majority over a variable, self-selected subset rather than a well-defined statistic over pairs. Third, the mention labels used for sampling have pairwise kappa of only 0.37 (Section 4.2); if many sampled 'mentions' are false positives, the rebuttals are being judged against claims the speaker never made, making 'plausible in context' vacuous. These issues compound: low label reliability inflates the sample, the filtering step removes the disputed cases, and the absence of a control leaves no way to know whether the remaining 87% says anything about general-purpose rebuttal quality. The paper's own Section 4.4 notes that unanimous failures were topic-related, but this is reported without quantitative follow-up. The dataset itself is still valuable, but the stated viability conclusion is not yet established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a new task in natural language understanding: producing a rebuttal in response to a long argumentative text, and proposes a General-Purpose Rebuttal Knowledge Base (GPR-KB) containing 55 manually authored general claims with matching rebuttals. The GPR-KB is evaluated through four annotation experiments over 200 speeches on 50 debate motions: cross-topic relevancy of claims, mention detection in speeches, sentence-level localization, and plausibility of rebuttals in context. The authors also provide baseline results for detecting whether a GP-claim is mentioned in a speech, showing that a frequency-based prior is a strong baseline. The central claims are that GP-claims are relevant across many topics, are commonly mentioned in spoken content, and that the pre-written rebuttals are judged plausible in 87% of 103 evaluated speech-rebuttal pairs.","tokens_in":12106,"tokens_out":5000,"duration_ms":46401,"significance":"If the findings are robust, the paper makes a useful contribution: it defines a novel task, releases a multi-layer dataset (GPR-KB-55), and demonstrates that a compact set of general claims can often be matched to real argumentative speeches. Strengths include the expert-debater authorship of the knowledge base without access to the target motions, the use of multiple annotation layers with known-answer control questions, and the honest reporting of a strong frequency-prior baseline. The paper also provides a clear comparison with the iDebate18 resource. However, the headline rebuttal-plausibility result currently rests on a single 87% majority-vote statistic that lacks a control condition and a stable denominator, and the mention-detection labels that drive coverage and sampling have only fair pairwise agreement. These issues are load-bearing for the paper's central empirical claims and require additional analysis or experiments.","major_comments":[{"comment":"The central claim that 'most rebuttals are appropriate in the vast majority of contexts' rests on the 87% plausibility rate over 103 speech-rebuttal pairs, but the experiment has no control condition: annotators are shown the GPR-KB rebuttal only after marking the claim as mentioned, and are never offered generic rebuttals, rebuttals matched to a different claim, or non-rebuttal texts. Without such a control, the high 'plausible' rate may reflect acquiescence or a low decision bar rather than the specific content of the GPR-KB. Additionally, because only annotators who marked the claim as mentioned proceed to rate the rebuttal, the number of raters contributing to each majority label is variable and self-selected; the paper does not report the distribution of the number of ratings per pair. Please add a control condition, report the per-pair rater counts, and provide confidence intervals for the 87% estimate.","section":"Section 4.4, Results"},{"comment":"The mention annotations used to sample the 103 rebuttal pairs and to compute coverage and prior statistics have pairwise Cohen's kappa of only 0.37, conventionally 'fair' agreement. The paper treats the majority label as ground truth throughout, but provides no reliability analysis of these majority labels beyond the 7% error rate on known-answer questions. If a substantial fraction of the sampled 'mentions' are false positives, the rebuttal-plausibility judgments in Section 4.4 are partially evaluated against claims the speaker never made. Please report agreement between individual annotators and the majority label (as is done for the rebuttal task), and perform a sensitivity analysis of the coverage, average-mention, and prior statistics under stricter mention thresholds (e.g., requiring 7 or more of 10 annotators, or unanimous agreement).","section":"Section 4.2"},{"comment":"The manual analysis of unanimously inappropriate rebuttals is reported only qualitatively: 'this stemmed from the rebuttal being inappropriate for the topic, rather than a specific speech' is stated without counts, a coding scheme, or a quantitative follow-up. Given that the paper decides to stop collecting further annotation based in part on this analysis, please provide the number of unanimously negative cases, examples, and a systematic check of whether excluding or down-weighting topic-inappropriate cases changes the 87% estimate.","section":"Section 4.4, Analysis"}],"minor_comments":[{"comment":"The phrase 'would consist a plausible rebuttal response' should be 'would constitute a plausible rebuttal response' or 'would be a plausible rebuttal response.'","section":"Section 4.4, first paragraph"},{"comment":"The text uses 'F 1-score' with an awkward space; use 'F1 score' consistently throughout.","section":"Section 5, last paragraph"},{"comment":"The similarity thresholds for sentence-claim pairing differ (0.5 for GP-claims and 0.7 for iDebate claims) without explanation; a brief justification would help readers interpret the two annotation pools.","section":"Section 4.3"},{"comment":"The column header 'Annotated pairs' is redundant with the table title; consider renaming it 'Total pairs' for clarity.","section":"Table 3"},{"comment":"The sentence 'She was not given access to any of the iDebate18 motions' uses a gendered pronoun where 'the debater' or 'the author' would be more neutral, though this is a minor style point.","section":"Section 3.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a resource paper with a potentially valuable dataset, and the revision path is clear: the rebuttal-plausibility experiment needs a control condition and a stable denominator, and the mention-label robustness needs quantification. The editor may also want to verify that the dataset release link remains accessible and that the annotation guidelines in the appendix are complete. If the authors address the Section 4.4 concerns, the paper would be a solid acceptance candidate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know: this is mainly a dataset paper, and the dataset is the contribution. The 87% rebuttal-plausibility number gets cited as proof that the general-rebuttal approach works, but that specific claim is under-supported. Don't let that obscure what's here.\n\nWhat's new: a curated set of 55 general claims and rebuttals, authored without access to the target motions, plus annotations linking them to 50 motions and 200 speeches. That's a genuinely useful resource, and the paper is transparent about the annotation process—multiple annotators, known-answer checks, agreement scores, and a public release. The baseline result showing a claim-prior beating text-based methods is a useful warning sign for the field, and the comparison to iDebate18 claims is informative.\n\nSection 4.4 is the weak spot. The claim that most rebuttals are appropriate in most contexts rests on 103 pairs, a majority-vote decision, no control condition, and no confidence intervals. The two-step procedure means only annotators who said the claim was mentioned go on to rate the rebuttal, and the paper doesn't say how many ratings each rebuttal actually got after that filter. Add the 0.37 kappa on mention detection, and the 87% is a real observation but not a firm estimate. The paper's own discussion of kappa with biased labels is fair, but it doesn't fix the missing control.\n\nI don't think this is a fatal flaw. The dataset is still solid, and the authors are honest about the main failure mode—topic-inappropriate rebuttals. But the abstract's stronger conclusion should be read as a hypothesis to test, not a result.\n\nWho should read it: anyone building rebuttal systems or working on argument mining. I'd send it to peer review; the dataset alone justifies that. Just don't let the stronger conclusion in the abstract survive without a better experiment.","headline":"A genuinely useful dataset paper; the 87% rebuttal-plausibility result is thinner than the abstract suggests, but the dataset itself justifies a serious look.","tokens_in":12686,"tokens_out":2397,"would_cite":true,"duration_ms":23017,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 55-entry knowledge base of general claims and their counterarguments can rebut speeches across many debate topics, with 87% of pre-written rebuttals judged plausible.","keywords":["general-purpose rebuttal","rebuttal generation","argument mining","debate dataset","listening comprehension","claim detection","crowdsourced annotation","knowledge base"],"falsifier":"Re-annotate the 3,246 speech-claim pairs under an adjudicated protocol with expert arbiters; if the share of speeches containing at least one GP-claim falls well below the paper's 100 percent coverage, the claim that the 55 general claims appear across all speeches would be refuted. Likewise, ask annotators to judge the same 103 speech-rebuttal pairs with rebuttals randomly swapped across claims: if the swapped rebuttals are judged plausible as often as the true ones, the 87 percent result would reflect generic acceptability rather than genuine counter-argument quality.","tokens_in":11603,"feed_emoji":"⚖️","tokens_out":9422,"duration_ms":80519,"temperature":0.7,"pith_summary":"This paper proposes a new task: automatically producing a critical rebuttal to a long argumentative speech, rather than to a short utterance. It argues that a small set of general-purpose claims, statements that recur across many debate topics, can be paired with pre-written counterarguments and reused to rebut speeches on topics never seen in advance. To support this, the authors built the GPR-KB-55 dataset: 55 manually authored claim-rebuttal pairs with annotations linking them to 50 debate motions and 200 recorded speeches. The annotations show that 46% of claim-motion pairs are relevant, that at least one general claim is mentioned in every speech, and that 87% of the pre-written rebuttals were judged plausible in context. If correct, this gives automatic rebuttal systems a topic-independent starting point, and the dataset is released for research.","feed_headline":"55 generic rebuttals can counter speeches on most topics","feed_subtitle":"A hand-built 55-claim knowledge base covers all 200 speeches, and 87% of its rebuttals were judged plausible.","key_machinery":"The carrying object is the General-Purpose Rebuttal Knowledge Base (GPR-KB): 55 manually authored claim-rebuttal pairs in which each claim is a short sentence with slots such as '[ACTION] [TOPIC]' that get instantiated for a specific motion, and each rebuttal is a context-free counterargument. It carries the argument by being small enough to curate by hand yet general enough that annotators label 46% of claim-motion pairs as relevant, and it is paired with a three-stage annotation pipeline that measures relevance, stance, mention in speech, sentence-level location, and rebuttal plausibility.","core_discovery":"The central discovery is that general-purpose argumentative claims are both frequent in real spoken debates and rebuttable without topic-specific knowledge. In the authors' terms, Section 4.2 shows GP-claims 'are often used in spoken content discussing various topics': 41% of relevant claim-speech pairs were labeled as mentioned (34% implicitly), and the 55 claims cover all 200 speeches, compared with 86.5% coverage for the topic-specific claims in the prior dataset. Section 4.4 shows that 'most rebuttals are appropriate in the vast majority of contexts': annotators judged 87% of 103 speech-rebuttal pairs plausible. A surprising baseline result is that simply predicting the most frequently mentioned claims, without reading the speech, competes with or beats text-based detectors, which the authors take as evidence that claim-frequency priors must be controlled for in evaluation.","pith_inferences":["The 87% plausibility rate may overestimate quality because annotators first decided the claim was mentioned and then judged its rebuttal, an anchoring setup; a control condition presenting the same rebuttals for claims they were not written for would test whether the rate reflects generic acceptability.","The low mention-detection agreement suggests 'implicit mention' is partly in the ear of the listener; an automatic system trained on these labels may be learning to reproduce what audiences read into a speech rather than what the speaker literally said, which may still be the right target for rebuttal.","Because only 5% of the 55 claims were relevant to all 50 motions, a deployed system would likely need to extend the GPR-KB or instantiate claims dynamically to keep coverage high on niche topics."],"forward_implications":["A concise, hand-curated set of 55 general claims can supply at least one candidate rebuttal for every speech in the dataset, removing the need for topic-specific argument lists.","Pre-written, context-free rebuttals can be judged plausible responses in 87% of the cases tested, so automatic rebuttal systems can be built without generating fresh content per speech.","The prior baseline, predicting the claims most frequently mentioned in training speeches, is strong enough that any text-based claim detector must be compared against it, and outperforming it requires more than semantic similarity.","The GPR-KB-55 dataset roughly triples the average number of claim mentions found per speech relative to the topic-specific annotations (6.7 vs. 1.8), giving listening-comprehension research a larger test bed.","Relevance and stance of general claims to a motion can be annotated with moderate agreement, making it feasible to filter claims for a given topic before rebuttal selection."],"supporting_citations":[{"why":"Supplies the 200 speeches, 50 motions, and the listening-comprehension task that the annotated GPR data is built on.","marker":"Mirkin et al. (2018)"},{"why":"Provides the CoPA argumentation model used in the implemented rebuttal system, grounding the claim that the KB can drive automatic rebuttal.","marker":"Bilu et al. (2019)"},{"why":"Companion listening-comprehension method that mines topical claims from a news corpus; the paper positions GPR-KB as an alternative when topic-specific mining is unavailable.","marker":"Lavee et al. (2019)"},{"why":"BERT is the neural baseline for claim detection in speeches, needed for the comparison showing the prior baseline is strong.","marker":"Devlin et al. (2018)"},{"why":"Cited to justify why Cohen's kappa is ill-suited for the skewed rebuttal-plausibility labels, supporting the 87% result's measurement choice.","marker":"Jeni et al. (2013)"},{"why":"Defines the kappa inter-annotator agreement metric used throughout the annotation experiments.","marker":"Cohen (1960)"},{"why":"Word2vec embeddings are used to select candidate sentences for sentence-level claim annotation and as a detection baseline.","marker":"Mikolov et al. (2013)"}],"fun_headline_variants":["55 generic rebuttals cover all 200 speeches, 87% plausible","Generic rebuttals: 87% plausible across 200 speeches","Frequency beats text: generic rebuttal detection surprise","41% of debated claims are implicitly rebuttable","55 universal rebuttals defeat topic-specific claims"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that crowdsourced majority votes for whether a speaker mentioned a general claim are accurate enough to measure coverage, even though annotators rarely agree with each other (a standard agreement score of 0.37), and the rebuttal-plausibility result was collected without a control condition.","fun_headline_variants_meta":{"raw":{"variants":["55 generic rebuttals cover all 200 speeches, 87% plausible","Generic rebuttals: 87% plausible across 200 speeches","Frequency beats text: generic rebuttal detection surprise","41% of debated claims are implicitly rebuttable","55 universal rebuttals defeat topic-specific claims"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000528,"raw_usage":{"total_tokens":2508,"prompt_tokens":868,"completion_tokens":1640,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":484,"completion_tokens_details":{"reasoning_tokens":1559}},"tokens_in":484,"tokens_out":1640,"duration_ms":11144,"temperature":1.0,"reasoning_tokens":1559,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:53:42.403012+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate the 3,246 speech-claim pairs under an adjudicated protocol with expert arbiters; if the share of speeches containing at least one GP-claim falls well below the paper's 100 percent coverage, the claim that the 55 general claims appear across all speeches would be refuted. Likewise, ask annotators to judge the same 103 speech-rebuttal pairs with rebuttals randomly swapped across claims: if the swapped rebuttals are judged plausible as often as the true ones, the 87 percent result would reflect generic acceptability rather than genuine counter-argument quality.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the 200 speeches, 50 motions, and the listening-comprehension task that the annotated GPR data is built on."},{"cited_title":"Towards Effective Rebuttal: Listening Comprehension using Corpus-Wide Claim Mining","cited_arxiv_id":"1907.11889","evidence_quote":"Companion listening-comprehension method that mines topical claims from a news corpus; the paper positions GPR-KB as an alternative when topic-specific mining is unavailable."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Cited to justify why Cohen's kappa is ill-suited for the skewed rebuttal-plausibility labels, supporting the 87% result's measurement choice."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the kappa inter-annotator agreement metric used throughout the annotation experiments."}],"review_version":1}