{"id":"e7a5f5fb-ac74-45d4-a778-be89f52a5234","arxiv_id":"2502.09644","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Arguments are encoded as vectors of stances on topic-specific concepts, and these vectors are used to predict subtle agreement, disagreement, and orthogonality between arguers.","lead":"This paper presents Perspectivized Stance Vectors: for each debate topic, a list of relevant concepts is built, and each argument is scored as for, against, or neutral on every concept, giving a vector. The hope is that such fine-grained vectors expose where opponents genuinely agree, which could serve as a starting point for resolving conflicts.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fine-grained evaluation rests on concept-level annotations with Krippendorff's α=0.03, so the central 'fine-grained' claim is not empirically supported.","rationale":"The reader's weakest assumption focused on completeness and granularity of the top-k ConceptNet signature. I agree that is a real concern, but the single most load-bearing issue is different: the fine-grained gold labels themselves are unreliable, with concept-level inter-annotator agreement at α = 0.03 and most topics annotated by only one person. Without reliable labels, the quantitative evidence for the paper's distinctive fine-grained claim is essentially unvalidated. This does not overturn the conditional verdict, because the paper also offers global acceptability annotations with α = 0.64, a large unannotated same-stance sanity check, and interpretable case studies; these support the framework as a proposal. However, the conditional acceptance must require either improved annotation reliability or a scaled-back claim about fine-grained resolution-point identification. The proposed concrete test would directly determine whether the low α is an artifact of one difficult topic or a fundamental property of the annotation task.","tokens_in":24931,"tokens_out":3311,"duration_ms":35968,"concrete_test":"Re-annotate the 25 opposite-stance argument pairs from a second topic with three trained annotators using the same concept-level guidelines, and compute Krippendorff's α per class and overall. If α remains below ~0.2, the fine-grained gold standard is not reliable, and Table 4 (middle) should be reported as exploratory only. As a complementary check, recompute the P0 AUC in Table 4 (middle) restricted to concept-instances where at least two of three annotators agree on the label; if the restricted AUC does not improve substantially over the full-set AUC, the model is not tracking even the most consensual judgments.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that PSVs enable fine-grained, perspective-structured (dis)agreement analysis that identifies actionable resolution points. The direct evidence for this is Section 4.3.2, where perspectivized acceptability predictions are compared to human gold labels for 'argument pairs – concept-level'. Table 5 (Appendix B.1) reports Krippendorff's α = 0.03 for this task, essentially chance agreement. Moreover, only one of the five topics was annotated by all three annotators; the remaining topics were labeled by a single annotator, so the gold set is not only noisy but largely unverified. The AUC values in Table 4 (middle), including the best P0 results (0.62/0.75/0.76), are computed against labels that are not reliably reproducible. If annotators cannot agree on whether two arguments agree, disagree, or are neutral on a given concept, then high AUC against a single annotator's labels does not show that PSV-derived scores track shared perspective structure. The paper acknowledges this in §B.1 but still interprets the case study as identifying common ground and conflicts. The granularity problem in Table 2 (53.3% appropriate granularity for unfiltered concepts) compounds this: even if labels were reliable, a large fraction of the signature dimensions may be too coarse or too fine to support the claimed resolution-point identification.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces Perspectivized Stance Vectors (PSVs) as a representation of arguments on a debate topic, built by selecting issue-specific concept sets from ConceptNet, predicting stances (for/against/neutral) toward those concepts with LLM-based classifiers, and aggregating the resulting vectors into acceptability scores. The authors evaluate signature quality, stance prediction, and global and perspectivized acceptability on a small manually annotated subset of PAKT, and illustrate the framework with stakeholder-level case studies. The main thesis is that PSVs can reveal fine-grained shared and opposing perspectives that support deliberation and conflict resolution.","tokens_in":25223,"tokens_out":3855,"duration_ms":31431,"significance":"If the representation and its evaluations are reliable, the paper offers a genuinely interpretable, unsupervised alternative to binary stance classification, with a clear formalization of stance vectors and aggregation operators, and several automatic sanity checks (same-stance separation, agreement/disagreement correlation, and a large-scale unannotated evaluation). The authors provide code and data, and the case studies demonstrate a useful proof-of-concept for stakeholder-level analysis. The main risk is that the fine-grained evaluation—the evidence for the paper's central claim—is anchored in annotations with near-chance inter-annotator agreement and in test-set selection of aggregation functions, so the quantitative support for the central claim is currently weak.","major_comments":[{"comment":"The central fine-grained claim—that PSVs identify which perspectives two arguments agree or disagree on—is supported only by the concept-level argument-pair acceptability evaluation, yet the gold labels for this task have Krippendorff's α = 0.03 (Table 5), which is effectively chance agreement. Moreover, only one of the five topics (Animal Hunting) was labeled by all three annotators; the remaining four topics were each labeled by a single annotator, so the 1,250 concept-level labels in the table are mostly unverified. Under these conditions, the P0 AUC values of 0.62/0.75/0.76 reported in the middle block of Table 4 cannot be read as evidence that the method tracks shared perspective structure. The paper acknowledges the low α in §B.1 but proceeds to interpret the case study as common-ground/conflict identification; this needs to be addressed by, e.g., re-annotating with adjudication, reporting per-annotator results, or explicitly demoting these numbers to exploratory.","section":"§4.3.2, Table 5, Appendix B.1"},{"comment":"All aggregation methods (S, S0, SD, P, P0, PD) are evaluated on the same manually annotated test sets, and the best-performing method (P0) is then presented as the paper's method without any held-out validation. Because the method family is large and the annotation set is small (125 argument pairs for global and 1,250 for perspectivized acceptability), the reported AUCs (e.g., 0.62/0.75/0.76) are likely optimistically biased by test-set selection. The authors should either pre-register the aggregation choice, use nested cross-validation, or report selection-corrected performance.","section":"§4.3, Table 4"},{"comment":"The signature analysis shows that unfiltered ConceptNet-based signatures have only 53.3% precision for 'appropriate granularity' (Table 2). Since a large fraction of PSV dimensions may be too general or too specific, the 'actionable points' identified in the case study (e.g., Section 5) may reflect artifacts of the signature-induction process rather than true perspective structure. The paper reports filtering variants but does not analyze how granularity failures affect the acceptability scores or the case-study conclusions. At a minimum, the authors should report the acceptability results for the appropriate-granularity subset and for the filtered signatures, and discuss the sensitivity of the qualitative findings to the signature choice.","section":"§3.1.1, Table 2"},{"comment":"The stance-prediction module that feeds the PSVs has macro F1 of only 50.2% (GPT4o zero-shot, Table 3), and the confusion matrix (Figure 6) shows that neutral stances are frequently misclassified as negative. Because stance values are the core input to all aggregation functions, these errors propagate directly into the acceptability scores; the paper does not report how much of the disagreement/agreement signal is due to stance-classification error as opposed to aggregation quality. An error propagation or sensitivity analysis (e.g., comparing predicted stances against oracle stances) would make the pipeline's strengths and limitations much clearer.","section":"§3.1.2, Table 3, Figure 6"}],"minor_comments":[{"comment":"The pairwise GPT4o prompt contains 'Argument 2: {argument_1}', which appears to be a copy-paste error; it should read 'Argument 2: {argument_2}'.","section":"Appendix A.5"},{"comment":"The note that Krippendorff's α is computed only on one topic should be made more prominent in the main text, since the reported α values are per-topic rather than pooled across all 5 topics; this affects how readers interpret the quality of the full annotation set.","section":"Table 5"},{"comment":"The grouping of 'agreement' and 'partial agreement' into one class is reasonable, but the paper should report the class frequencies before and after grouping; with only two original 'agreement' pairs, the stability of the resulting AUC values is questionable.","section":"§4.3.1"},{"comment":"The paper states 'Data and code are available at [GitHub]' but also says the annotated data 'will be published upon acceptance'; please clarify which items are currently released and include the annotation guidelines and labels in the public repository.","section":"Data availability"},{"comment":"The histograms in Figures 2 and 9 do not have explicit legends; please label which color corresponds to same-stance and which to different-stance pairs, and specify the axis units.","section":"Figures 2 and 9"}],"recommendation":"major_revision","confidential_remarks":"The paper's strengths are its clean formalization of perspective-based stance vectors and the interpretable, scalable pipeline; the main weakness is the empirical support for the central fine-grained claim, which currently rests on near-chance inter-annotator agreement and test-set selection of aggregation methods. These issues are fixable with re-annotation or a more cautious framing, but they are substantial enough that the paper should not be accepted as is. No concerns about academic integrity."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I agree with the conditional verdict, and the stress-test identifies the core problem. The idea of representing arguments as vectors of stances toward issue-specific concepts is worth taking seriously. The signature induction via frequency contrasts and the aggregation family (S, S0, SD, P, P0, PD) are clean formalizations, and the orthogonality distinction is a real step beyond binary stance. The same-stance sanity check on 326k pairs is a nice idea, and the case study at least shows the kind of interpretable output the framework is meant to produce.\n\nBut the central empirical claim falls apart. The concept-level pair annotation—the direct gold for perspectivized (dis)agreement—has Krippendorff's α = 0.03. That is effectively random. Only one of five topics was annotated by all three annotators; the rest is a single annotator's judgment. AUCs of 0.62–0.76 against those labels do not show that PSVs track shared perspective structure. On top of that, the best aggregation method was selected on the test set, so the numbers that exist are optimistic. The paper acknowledges the low IAA in §B.1 but then proceeds to interpret the case study as identifying common ground and conflicts. That is an overclaim.\n\nThere is also the granularity issue: Table 2 shows only 53.3% of unfiltered concepts have appropriate granularity. So even if labels were reliable, a large chunk of the 100 dimensions is too coarse or too specific to support the resolution-point claim. The ablation without PSVs, where GPT4o is prompted directly, lands at 0.62/0.66/0.69 compared to the best PSV results of 0.62/0.75/0.76, which does not exactly strengthen the case.\n\nThis is not a fundamentally wrong idea. It is an evaluation that is too weak for the claim. The framework could be publishable if the authors either redo the annotation with a more constrained task and better guidelines, report results on a held-out development set rather than selecting the method on the test set, or scale the claims back to illustrative case studies. As it stands, I would not cite the empirical results, but I would point to the formalization as related work.\n\nThe paper deserves a serious referee—it is not a desk reject—but it needs major revision. I would send it back with a clear request to fix the evaluation or soften the claims. For a reading group, it is worth a session on annotation reliability and test-set selection; it is a good negative example. My own verdict would be major revision, not acceptance in the current form.","headline":"The PSV framework is a genuine formal step beyond binary stance, but the fine-grained evaluation rests on annotations with Krippendorff's α=0.03, so the central claim is not empirically supported.","tokens_in":25759,"tokens_out":2577,"would_cite":false,"duration_ms":22473,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that encoding arguments as vectors of stances toward issue-specific concepts can expose the shared and opposing grounds between arguers, moving computational argumentation from debate to deliberation.","keywords":["computational argumentation","perspectivized stance vectors","deliberation","fine-grained disagreement","stance classification","ConceptNet","acceptability scores","argument mining"],"falsifier":"On a fresh debate topic not used in the paper, induce a PSV signature with the paper's top-k method, then have annotators list the concepts on which opposing arguments actually agree or disagree. If most of the annotator-identified conflict concepts fall outside the signature or are rated neutral by the stance predictor, the signature has missed the decisive perspectives and the central claim fails.","tokens_in":24747,"feed_emoji":"⚖️","tokens_out":7069,"duration_ms":56524,"temperature":0.7,"pith_summary":"Debates present a conflict as a binary choice: for or against. This paper asks what happens if an argument is instead represented as a vector of stances toward the specific concepts that the debate turns on — for example, 'hunting for food' versus 'trophy hunting'. The authors' claim is that this finer representation, called a Perspectivized Stance Vector, reveals exactly which perspectives two arguers share, which they oppose, and which are simply irrelevant to one of them, and that this information can identify actionable points for resolving a conflict. That is the step that could move computational argumentation from winning debates to supporting deliberation.","feed_headline":"Stance vectors expose the shared ground in any debate","feed_subtitle":"Locating the exact perspectives where arguers agree and clash makes conflicts actionable for compromise.","key_machinery":"The load-bearing object is the Perspectivized Stance Vector (PSV), an n-dimensional vector in which each dimension is one concept from a topic-specific signature and each entry is a stance value from the set {against, neutral, in favor}. The signature is induced without supervision: arguments in a debate corpus are aligned to ConceptNet concept graphs, and the top-k concepts by differential frequency across PRO and CON arguments are selected. Stance values are then predicted per concept (best by zero-shot GPT4o prompting), and pairs of vectors are aggregated by functions such as P0, which computes agreement, orthogonality, and disagreement contributions per concept before averaging to a global acceptability score.","core_discovery":"The paper's central claim is that arguments on a contested issue can be usefully represented as Perspectivized Stance Vectors (PSVs): a fixed, topic-specific list of concepts — the signature — together with a predicted stance label (against, neutral, in favor) for each concept. Comparing two PSVs dimension by dimension yields agreement, orthogonality, and disagreement scores per perspective, so an argument pair is no longer reduced to a single 'agree or clash' verdict. The authors show that this representation uncovers partial agreement between arguments that take opposite global stances, uncovers disagreement among arguments that share a global stance, and — aggregated across stakeholder groups — surfaces shared ground such as hunters and environmentalists both condemning poaching. Their evaluations report best performances of 50.2% macro F1 for stance prediction with GPT4o and 0.62-0.76 ROC-AUC for perspectivized acceptability with the P0 aggregation, which they argue is enough to point toward concrete compromise proposals.","pith_inferences":["If the signature were extended beyond ConceptNet with, say, LLM-generated concepts, the method might capture perspectives outside the static knowledge graph, improving on the paper's observed 53.3% appropriate-granularity ceiling for unfiltered concepts.","The neutral (orthogonal) dimension is the natural negotiation space: perspectives where one arguer has no stance may be the least costly concessions, an implication the paper identifies but does not pursue.","A learned weighting over the fixed signature could replace the paper's uniform averaging and plausibly raise global acceptability prediction, at the cost of the training data the authors note are currently unavailable.","Because the case study shows that same-stance arguers disagree on perspectives such as 'control' or 'pleasure,' stakeholder maps built from PSVs could serve as a diagnostic for split coalitions within advocacy groups, not just across them."],"forward_implications":["PSV-based analysis can detect partial agreement between argument pairs of opposite global stance, turning a binary opposition into a list of specific contested and shared perspectives.","The same framework reveals disagreement within a global stance, showing that arguers who are both PRO or both CON may hold that position for different reasons.","Because signatures are induced unsupervised from ConceptNet, the approach transfers to new debate topics without training a topic-specific classifier.","The interpretable per-concept scores beat direct pairwise LLM prompting on perspectivized acceptability (0.62-0.76 vs 0.62-0.69 AUC) while scaling linearly in the number of arguments rather than quadratically.","Aggregating PSVs by stakeholder group yields maps of shared and conflicting perspectives that can guide moderators toward compromise entry points."],"supporting_citations":[{"why":"Supplies the PAKT corpus of stance-annotated arguments on binary issues that the experiments and case study run on.","marker":"(Plenz et al., 2024)"},{"why":"Provides ConceptNet, the commonsense knowledge graph whose nodes become the perspective concepts in a PSV signature.","marker":"(Speer et al., 2017)"},{"why":"Contributes the method for aligning arguments to ConceptNet concept graphs that underlies signature induction.","marker":"(Plenz et al., 2023b)"},{"why":"Represents the LLM-based consensus-generation approach that the paper positions against, claiming interpretability and explicit conflict resolution as advantages.","marker":"(Bakker et al., 2022)"},{"why":"Introduces syntopical graphs for viewpoint detection, the prior work that PSVs extend by also detecting orthogonality between arguments.","marker":"(Barrow et al., 2021)"}],"fun_headline_variants":["PSVs reveal hidden agreement in conflicting arguments","Per-issue stance vectors map both clash and common ground","Track stance per concept to find compromise points in debates","Beyond agree/disagree: stance vectors show partial convergence","Shared perspectives emerge from fine-grained stance vectors"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that the set of top-k concepts selected by differential frequency across PRO and CON arguments is a complete and appropriately granular set of perspectives for a debate topic, so that a fixed 100-dimensional PSV captures the dimensions of (dis)agreement that matter.","fun_headline_variants_meta":{"raw":{"variants":["PSVs reveal hidden agreement in conflicting arguments","Per-issue stance vectors map both clash and common ground","Track stance per concept to find compromise points in debates","Beyond agree/disagree: stance vectors show partial convergence","Shared perspectives emerge from fine-grained stance vectors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000234,"raw_usage":{"total_tokens":1509,"prompt_tokens":968,"completion_tokens":541,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":584,"completion_tokens_details":{"reasoning_tokens":468}},"tokens_in":584,"tokens_out":541,"duration_ms":5316,"temperature":1.0,"reasoning_tokens":468,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T15:25:57.159409+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"On a fresh debate topic not used in the paper, induce a PSV signature with the paper's top-k method, then have annotators list the concepts on which opposing arguments actually agree or disagree. If most of the annotator-identified conflict concepts fall outside the signature or are rated neutral by the stance predictor, the signature has missed the decisive perspectives and the central claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides ConceptNet, the commonsense knowledge graph whose nodes become the perspective concepts in a PSV signature."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Represents the LLM-based consensus-generation approach that the paper positions against, claiming interpretability and explicit conflict resolution as advantages."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces syntopical graphs for viewpoint detection, the prior work that PSVs extend by also detecting orthogonality between arguments."}],"review_version":1}