{"id":"76a33d82-a078-46a2-ab0e-b592ca4facd8","arxiv_id":"1908.03181","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Sentiment polarity alone does not reliably indicate stance in tweets, as shown on a new dataset and the SemEval benchmark.","lead":"A study of tweets on four controversial topics shows that positive or negative language is a poor guide to whether a post supports or opposes the topic. The finding warns against using simple sentiment scores to measure public opinion on social media.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No significant objection to the central claim; the missing CD-dataset reliability metrics limit the validation, but SemEval independently supports the conclusion.","rationale":"The reader's weakest assumption is label reliability, especially in the new CD dataset. I agree that the CD dataset lacks reported inter-annotator agreement and is not released, which makes its specific numbers unverifiable. However, this is not load-bearing for the paper's main conclusion, because the SemEval benchmark independently demonstrates the same low alignment: Favor tweets are mostly negative, and only 8.9% of all SemEval tweets are Favor-Positive. The paper's qualitative examples also show genuine sentiment-stance mismatches. Thus the central claim is supported even if the CD dataset were set aside entirely. The paper's contribution is partly the new context-dependent dataset, so the missing release and reliability metrics justify keeping the CONDITIONAL verdict rather than upgrading to ACCEPT. The concern does not warrant rejection because the SemEval evidence is public and well established, and the paper's conclusion is modest.","tokens_in":6875,"tokens_out":13399,"duration_ms":139815,"concrete_test":"Re-annotate a random 200-tweet-per-topic sample of the CD dataset with fresh annotators, compute Krippendorff's alpha for sentiment and stance separately, and recompute the Figure 4b alignment proportions on the high-agreement subset. If alpha is below 0.6 or the proportions shift by more than 10 percentage points, the CD-based validation is unreliable; otherwise the concern does not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"No significant objection identified. The central claim is that sentiment polarity alone cannot stand in for stance. The most decisive evidence is the existing SemEval benchmark: within Favor tweets only 35.0% are positive (56.1% are negative), and overall only 8.9% of tweets are Favor-Positive versus 33.9% Against-Negative, leaving 57.2% mixed (Figures 2a, 4a). These figures do not depend on the new CD dataset. The CD dataset is a real limitation: five-annotator majority vote with no reported inter-annotator agreement and no public release, so its specific numbers (e.g., 65.1% mixed, 4.0% Favor-Positive) cannot be independently checked. If CD annotation noise differed by class it could distort the CD-based percentages. But since SemEval alone yields the same qualitative conclusion and the Table 2 examples show genuine cases, the central claim survives. The only remaining load-bearing assumption is that SemEval's gold labels are correct; that is a standard benchmark, but the paper provides no sensitivity analysis to label noise.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper examines whether sentiment polarity can serve as a proxy for stance in social-media posts. It analyzes the SemEval-2016 Task 6 stance benchmark (4,063 tweets, five topics) and a new 'context-dependent' dataset of 6,324 reply tweets across four controversial topics, each annotated by five workers via majority vote. The authors report sentiment-by-stance distributions, the overall proportion of tweets in which stance and sentiment agree (Favor-Positive and Against-Negative), and TF-IDF top-word Jaccard similarities. They conclude that sentiment and stance are not highly aligned and that sentiment polarity alone should not be used to infer stance.","tokens_in":7090,"tokens_out":5321,"duration_ms":54703,"significance":"The paper addresses a practically important question: public-opinion studies frequently use sentiment as a proxy for stance, and this work provides quantitative and qualitative evidence that the two constructs differ. The use of the established SemEval benchmark as an external validation is a clear strength, and the four examples in Table 2 convincingly illustrate the phenomenon. The contribution would be more complete with reliability statistics for the new dataset and statistical tests for the aggregate comparisons; as it stands, the strength of the quantitative claims is limited by the absence of such measures. No code or data release is provided, so the new dataset cannot be independently checked.","major_comments":[{"comment":"The CD dataset is central to the paper's empirical contribution, but no inter-annotator agreement is reported for the five-annotator majority-vote labels. Because the entire alignment analysis depends on the reliability of these sentiment and stance gold labels, the paper should report Fleiss' kappa (or equivalent) per topic and per label type, and should report the distribution of agreement. Without this, the CD-based percentages (e.g., 65.1% mixed and 4.0% Favor-Positive in Fig. 4b) cannot be distinguished from annotation noise. The dataset also does not appear to be released; please provide a public link or a clear availability statement.","section":"Section 3 (Data collection)"},{"comment":"The central claim that sentiment and stance are 'not highly aligned' is supported only by raw proportions and informal comparisons of Jaccard curves; no significance test, confidence interval, or association measure (e.g., chi-square, Cramér's V, mutual information) is reported. The apparent mismatch is large, but the reader cannot assess whether the observed pattern exceeds what would be expected by chance or by the marginal distributions of the labels. Please add statistical tests or resampling baselines for the contingency tables in Fig. 2 and the Jaccard curves in Fig. 3, and consider a sensitivity analysis that varies the SemEval gold labels.","section":"Sections 4.1 and Appendix A"},{"comment":"RQ2 asks 'When does positive/negative sentiment indicate support/against stance?', but the analysis only establishes that sentiment is not a reliable indicator overall; it does not identify conditions under which sentiment and stance do align beyond illustrative examples in Table 2. The Discussion's statement that negative sentiment 'could help' with against stances is not operationalized. Either provide a concrete characterization (e.g., by topic, target, or textual pattern) or explicitly narrow the paper's contribution to answering RQ1.","section":"Sections 1 and 5 (RQ2)"}],"minor_comments":[{"comment":"The surname is misspelled as 'Jacquard'; it should be 'Jaccard' throughout.","section":"Appendix A"},{"comment":"The phrase 'The words choice gap exists' is awkward; consider 'There is a vocabulary gap between in-favor stance and positive sentiment' or similar.","section":"Section 5"},{"comment":"The captions do not state whether the displayed percentages are normalized per row or per column; please make this explicit so the distributions can be interpreted correctly.","section":"Figures 1, 2, and 4"},{"comment":"Since the CD labels are claimed to follow SemEval annotation guidelines, including the actual annotation instruction sheet in an appendix or supplementary material would help readers judge label quality.","section":"Section 3"},{"comment":"The paper would benefit from a short limitations paragraph explicitly acknowledging the reliance on gold labels and the absence of significance testing, rather than leaving these points implicit.","section":"Overall"}],"recommendation":"major_revision","confidential_remarks":"To the editor: The paper is within the scope of SocInfo and the core message is plausible. My main concerns are the missing inter-annotator agreement for the new dataset and the purely descriptive quantitative analysis. If the authors can supply reliability statistics, add basic statistical tests, and make the CD dataset available (or state clearly why it cannot be released), I would be willing to support acceptance after a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper is a small, careful empirical study that confirms something the field already largely accepted: sentiment polarity alone does not tell you a Twitter user's stance. The authors themselves cite Mohammad et al. (2016, 2017) to that effect, so the conceptual result is not new. What is new is an annotated context-dependent dataset of 6,324 reply tweets across four controversial topics, with the parent tweet shown to annotators so they could judge stance in context. That dataset is potentially valuable, but the paper does not release it, and it reports no inter-annotator agreement for the five-worker majority labels. That is the main soft spot.\n\nThe strengths: the descriptive analysis is clear and the examples in Table 2 are genuinely illustrative (e.g., \"Life is our first and most basic human right\" labeled positive sentiment and Against stance on abortion). Using the SemEval 2016 stance dataset as a second, independent validation is good practice, and the SemEval numbers alone support the conclusion: within Favor tweets, 56% are negative, and only 8.9% of all tweets are Favor-Positive. So even if the CD dataset vanished, the central claim holds.\n\nThe weaknesses are proportionate. There are no significance tests, confidence intervals, or sensitivity analysis to label noise. The \"mixed\" category in Figure 4 is coarse—it lumps everything that is not Against-Negative or Favor-Positive, including neutral sentiment and None stance, into one bucket, which inflates the apparent mismatch. The appendix Jaccard analysis is underdeveloped and has a typo (\"Jacquard\"). But these are fixable and do not undermine the qualitative conclusion.\n\nFor a reader, this is a useful cautionary citation and a reasonable dataset paper if the authors ever release the data. As a research contribution it is confirmatory rather than groundbreaking, and a desk editor could reasonably note that the novelty is modest. However, the work is coherent, honest, and relevant to anyone using sentiment as a proxy for public opinion on social media. I would send it out to a serious referee, primarily to push for IAA, a data-release plan, and sharper metrics, not because the paper is flawed beyond repair.","headline":"A modest, honest confirmation that sentiment polarity does not equal stance; the new dataset is the real asset but it is not released and the analysis lacks reliability metrics.","tokens_in":7524,"tokens_out":1707,"would_cite":false,"duration_ms":21098,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Sentiment alone cannot tell support from opposition in tweets","keywords":["stance detection","sentiment analysis","public opinion","event analysis","social media","Twitter","stance-sentiment alignment"],"falsifier":"Compute inter-annotator agreement on the CD dataset and recompute the sentiment-stance distributions using only tweets all five annotators labelled identically; if the against-negative and favour-positive proportions move substantially toward alignment, the paper's mismatch is inflated by label noise.","tokens_in":6727,"feed_emoji":"💬","tokens_out":5292,"duration_ms":53544,"temperature":0.7,"pith_summary":"Sentiment polarity—whether a tweet reads positive, negative, or neutral—is often treated as a shortcut for detecting which side of a controversial issue the writer supports. This paper tests that shortcut directly, building a new context-dependent dataset of 6,324 reply tweets on four topics with separate gold labels for sentiment and stance, and re-analysing the established SemEval stance benchmark. The result is that sentiment and stance are far from interchangeable: negative tweets make up the bulk of both favour and against stances, while tweets that are both positive and in favour are rare. The authors conclude that public-opinion readings built on sentiment polarity alone will systematically misrepresent true support.","feed_headline":"Sentiment alone cannot tell support from opposition in tweets","feed_subtitle":"New four-topic dataset and SemEval reanalysis show favor-positive tweets at under 9 percent.","key_machinery":"The load-bearing objects are two sentiment-and-stance annotated tweet collections: the SemEval-2016 Task 6 benchmark (4,063 tweets, five topics) and a new context-dependent (CD) dataset of 6,324 reply tweets on four topics, annotated by five workers per tweet with majority-vote gold labels and the parent tweet shown for context. The analysis itself rests on conditional distributions of sentiment given stance (and vice versa), which expose the misalignment, and on Jaccard similarity between the top TF-IDF-weighted vocabularies of sentiment labels and stance labels, which quantifies how different the two expression systems are. The example rows in Table 2—positive-sentiment tweets that oppose, negative-sentiment tweets that support—carry the qualitative force of the argument.","core_discovery":"The paper's central claim is that stance and sentiment are different dimensions of expression, and that polarity cannot be used on its own to infer stance toward a topic. In both the SemEval dataset and the new CD dataset, negative sentiment is the dominant polarity for supporting and opposing stances alike: over 54% of favour tweets in both datasets carry negative sentiment. Only about 33.9% of SemEval tweets and 30.9% of CD tweets are the 'matching' combination of against plus negative, while favour plus positive tweets make up just 8.9% and 4.0% respectively. A Jaccard analysis of the most frequent words per label shows less than 20% overlap between the vocabulary of favour stance and the vocabulary of positive sentiment, while against stance shares more words with negative sentiment. The authors therefore state that simple sentiment polarity cannot substitute for stance detection.","pith_inferences":["Because the new dataset consists of reply tweets, the mismatch the paper finds may be especially pronounced in conversational contexts, where rebuttals, irony, and defensive agreement routinely separate the literal polarity of words from the position being defended; extending the same analysis to standalone posts could show a weaker or stronger gap.","A direct, testable extension would be to train a stance classifier on sentiment features only and compare its accuracy with the majority-class baseline; the distributions here imply it would barely beat the baseline, which would quantify the ceiling of polarity-based stance inference.","The fact that favour-positive agreement is so low (4–9%) suggests that support in public discourse is often expressed as criticism of the other side rather than praise of the target; if that is general, positive sentiment is a poor instrument for detecting approval in any contested domain."],"forward_implications":["Sentiment polarity should not be used as a stand-alone signal for stance detection or public-opinion measurement.","Studies that equated negative sentiment with opposition and positive sentiment with support should be revisited, since the favour-positive match is below 9% in both datasets.","Stance classifiers that combine sentiment with other features have the right shape; the paper supports using sentiment as a complement, not a substitute.","Negative sentiment is a partially informative cue for against-stances, but it is diluted by the large share of negative favour tweets, so any threshold-based polarity rule will overcount opposition.","Future annotation efforts should label stance and sentiment as separate dimensions, as the CD dataset does, rather than deriving one from the other."],"supporting_citations":[{"why":"Supplies the SemEval stance benchmark: the five-topic tweet dataset, gold stance and sentiment labels, and the annotation guidelines reused for the new CD dataset.","marker":"[15]"},{"why":"Prior study of stance and sentiment in tweets that concluded sentiment helps only in combination with other features; the paper extends and sharpens this conclusion.","marker":"[16]"},{"why":"Earlier analysis of the interaction between stance and sentiment that the paper draws on to frame the alignment question.","marker":"[20]"},{"why":"A joint sentiment-target-stance model showing sentiment alone is inadequate for stance classification, supporting the paper's critique.","marker":"[6]"},{"why":"A SemEval system that used sentiment features alongside others, cited as evidence that sentiment is insufficient on its own.","marker":"[7]"},{"why":"Defines stance as a standpoint toward a proposition, framing the task the paper studies.","marker":"[4]"}],"fun_headline_variants":["Stance ≠ sentiment: over half of supportive tweets carry negative tone","Only 9% of support tweets are positive—stance ≠ sentiment","Supportive tweets skew negative—polarity doesn't reveal stance","Negative sentiment dominates both sides—stance is not polarity"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The correctness of the gold labels—the majority-vote annotations in the CD dataset and the SemEval labels—is the load-bearing premise; if annotation noise differs across sentiment and stance classes, the observed mismatch could be a labeling artifact rather than a property of how people express stance.","fun_headline_variants_meta":{"raw":{"variants":["Stance ≠ sentiment: over half of supportive tweets carry negative tone","Only 9% of support tweets are positive—stance ≠ sentiment","Supportive tweets skew negative—polarity doesn't reveal stance","Negative sentiment dominates both sides—stance is not polarity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000995,"raw_usage":{"total_tokens":4176,"prompt_tokens":866,"completion_tokens":3310,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":482,"completion_tokens_details":{"reasoning_tokens":3239}},"tokens_in":482,"tokens_out":3310,"duration_ms":25350,"temperature":1.0,"reasoning_tokens":3239,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T14:20:48.930459+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute inter-annotator agreement on the CD dataset and recompute the sentiment-stance distributions using only tweets all five annotators labelled identically; if the against-negative and favour-positive proportions move substantially toward alignment, the paper's mismatch is inflated by label noise.","supporting_citations":[{"cited_title":"In: SemEval@ NAACL-HLT","cited_arxiv_id":null,"evidence_quote":"Supplies the SemEval stance benchmark: the five-topic tweet dataset, gold stance and sentiment labels, and the annotation guidelines reused for the new CD dataset."},{"cited_title":"ACM Transactions on Internet Technology (TOIT) 17(3), 26 (2017)","cited_arxiv_id":null,"evidence_quote":"Prior study of stance and sentiment in tweets that concluded sentiment helps only in combination with other features; the paper extends and sharpens this conclusion."},{"cited_title":"In: Proceedings of the Fifth Joint Conference on Lexical and Computational Semantics","cited_arxiv_id":null,"evidence_quote":"Earlier analysis of the interaction between stance and sentiment that the paper draws on to frame the alignment question."},{"cited_title":"In: COLING","cited_arxiv_id":null,"evidence_quote":"A joint sentiment-target-stance model showing sentiment alone is inadequate for stance classification, supporting the paper's critique."},{"cited_title":"In: Proceedings of the 10th International Work- shop on Semantic Evaluation (SemEval-2016)","cited_arxiv_id":null,"evidence_quote":"A SemEval system that used sentiment features alongside others, cited as evidence that sentiment is insufficient on its own."},{"cited_title":"Discourse processes 11(1), 1–34 (1988)","cited_arxiv_id":null,"evidence_quote":"Defines stance as a standpoint toward a proposition, framing the task the paper studies."}],"review_version":1}