{"id":"cb600730-b030-4629-b233-8dd51b7568ba","arxiv_id":"2502.05255","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"COVID-19 mentions in climate discussions are strongly associated with higher incivility and disagreement on Twitter and Reddit, an effect linked to anti-internationalist populist attitudes.","lead":"This study analyzed millions of tweets and Reddit comments about climate change and found that posts also mentioning COVID-19 were more hostile and contentious, with spikes tied to pandemic events. It shows that political anger around public health can spill into climate discussions along pre-existing populist divides, which matters for how science is communicated and trusted.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing assumption is that Perspective API scores, binarized at 0.5, validly measure incivility in climate/COVID tweets; without a human validation or threshold sensitivity check, the COVID-toxicity association could be partly a classifier artifact.","rationale":"The reader's weakest assumption—that Perspective API and DistilBERT validly operationalize incivility and contentiousness—identifies the same general area of risk, and I agree that measurement validity is the most load-bearing issue. My stress test sharpens this to the specific binarization of Perspective scores and the absence of any human validation on the actual study tweets. Account fixed effects and fractional logistic regression address confounding and bounded outcomes, but they do not address differential measurement error in the outcome itself. If the classifier systematically scores COVID-adjacent tweets as toxic because of correlated linguistic surface features rather than substantive incivility, every downstream estimate—including the 14% Cox hazard ratio—would be biased. The paper does provide illustrative examples and cites the official Perspective definition, which is useful but not a substitute for a stratified precision/recall assessment. The BEAST-based selection of the post-escalation analysis window is a secondary concern because it can amplify whatever gap the classifier measures, but it is not the core issue. Given the promised code and data release, the proposed validation is feasible and would settle the concern. The reader's conditional verdict already reflects the need for additional verification, so my analysis does not move the verdict; it reinforces the condition and specifies a concrete check.","tokens_in":20167,"tokens_out":7030,"duration_ms":79370,"concrete_test":"Draw a stratified sample of approximately 2,000 general climate tweets from the study period, stratified by COVID-keyword presence and by Perspective score bins centered on 0.5 (for example, 0.3–0.7). Have at least three annotators label each post as uncivil or not using the paper's stated definition. Compare the COVID versus non-COVID gap in human-rated incivility with the Perspective-scored gap, and recompute Fig. 1.B and the Cox model using (a) human labels where available, (b) a threshold sweep at 0.3/0.5/0.7, and (c) the continuous Perspective score. If the COVID coefficient attenuates to near zero under human labels or is highly sensitive to the threshold, the measurement-artifact concern lands; if the gap persists across all specifications, the central association is supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The 14% Cox hazard ratio and the LPM coefficients in Fig. 1.B/C are computed from Perspective API toxicity scores binarized at 0.5 (Sections 4.2.2, 4.3.2, and S5.2). Perspective is trained on Wikipedia talk pages and NYT comments, not on climate/COVID tweets, and no validation on a sample from the study corpus is reported. The COVID-19 keyword set (Table S3.1) includes terms such as 'mask', 'vaccin', 'pandemic', and 'Fauci'; if tweets containing those strings are stylistically different (for example, more all-caps, profanity, or anti-elite framing) in ways Perspective treats as toxic regardless of whether human readers would call the post uncivil, then the COVID-19 coefficients and the spillover conclusion are at least partly measurement artifacts. This is load-bearing because every outcome—post toxicity, conversation-level toxicity onset, and the international-organization temporal analysis—passes through the same classifier and threshold. The example tweets are illustrative, not a systematic precision/recall check. A secondary amplifier is that the main regression window (August 14, 2020–August 26, 2021) is selected from BEAST change points estimated on the same toxicity outcome, which can inflate the apparent gap.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"Using approximately 38.6 million climate-related tweets, 311,000 full Twitter conversations, and 2.1 million Reddit comments from February 2019 through August 2021, the paper studies whether the COVID-19 pandemic changed the tone of public engagement with climate science. The authors operationalize incivility with Perspective API toxicity scores (binarized at 0.5), contentiousness with a fine-tuned DistilBERT disagreement classifier, and contrarian climate claims with CARDS. They report that climate posts mentioning COVID-19 or Anthony Fauci are more likely to be toxic across Twitter and Reddit; that COVID-19-related climate conversations show higher disagreement; and that these patterns track pandemic events, with evidence that anti-internationalist populist keywords link climate and vaccine skepticism. The main analyses are supported by account fixed effects, fractional logistic regression, and a Cox proportional hazards model of toxicity onset.","tokens_in":20365,"tokens_out":7780,"duration_ms":72866,"significance":"The paper provides a large-scale, cross-platform descriptive account of a timely question: whether a public health crisis can spill into the tone of climate discussions. The central association—COVID-19 references co-occurring with higher toxicity—is consistent across three social media systems and several model specifications, which is a real empirical strength. If the measurement and sampling-window concerns are addressed, the findings would be a valuable contribution to science communication and affective polarization literatures. The stated commitment to release code and data is also a positive feature. However, the reliance on a single black-box toxicity classifier without any validation on the study corpus is a serious validity threat, and the outcome-driven selection of the analysis window further limits the strength of the causal claims.","major_comments":[{"comment":"The validity of the key dependent variable rests entirely on Perspective API toxicity scores binarized at 0.5. The model was trained on Wikipedia talk pages and New York Times comments, not on climate/COVID tweets, and no precision/recall evaluation is reported on a sample from the study corpus. Because the COVID-19 keyword set (Table S3.1) includes terms such as 'mask,' 'vaccin,' and 'Fauci,' which may co-occur with stylistic features (e.g., all-caps, profanity, anti-elite framing) that Perspective flags as toxic regardless of human judgment, the association between COVID-19 content and toxicity could be driven by systematic misclassification. I request a validation sub-study: human annotation of a random sample of posts stratified by platform and COVID-19 keyword presence, reporting agreement statistics, along with a threshold sensitivity analysis (e.g., continuous toxicity score with cutoffs 0.3, 0.5, 0.7) to demonstrate that the main coefficients in Fig. 1.B and Fig. 1.C are not artifacts of the 0.5 cutoff.","section":"Sections 4.2.2, 4.3.1, 4.3.2"},{"comment":"The main regression and Cox analyses are restricted to the window August 14, 2020–August 26, 2021, which is selected from BEAST change points estimated on the same toxicity time series (Section 4.4). This outcome-driven sample selection can exaggerate the apparent association and yield over-optimistic confidence intervals, because the analysts have essentially chosen the period where the divergence is largest. The authors should report results for alternative analysis windows (e.g., the full pandemic period from March 2020 onward, or a window defined a priori from event dates) and show that the COVID-19 coefficients in Figs. 1.B and 1.C are stable across these choices.","section":"Sections 4.3.1 and S5.2"},{"comment":"The claim that the increase in the toxicity gap for international-organization keywords is 'attributed to the COVID-19 shock' requires the parallel trends assumption and the absence of confounding events beyond the pandemic. No placebo tests are provided (e.g., using a set of unrelated international topics that were not central to COVID-19 politics), and the Q1–Q2 2020 spike could also coincide with the U.S. presidential primary season or other political shocks. Please provide falsification exercises or soften the causal language to 'is associated with the pandemic period.'","section":"Section 2.3, Figure 3"}],"minor_comments":[{"comment":"The text states that the binary classifier returns f1 scores of 0.806 and 0.716 for disagreement and lack of disagreement, respectively, but Table S4.1 reports the three-class disagreement f1 as 0.716. Clarify which value corresponds to which class and reconcile the text with the table.","section":"Section 4.2.3 and Table S4.1"},{"comment":"The phrase 'the number of replies until one exceeds 0.5 toxicity probability is significantly lower' is informal; since the outcome is the hazard of toxicity onset, it would be clearer to say 'the hazard of toxicity onset is significantly higher' or 'replies become toxic sooner.'","section":"Section 2.1"},{"comment":"The example tweets are illustrative but the selection criteria are not described. State whether they were hand-picked, randomly sampled, or selected to match specific toxicity ranges, so readers can gauge their representativeness.","section":"Tables 1 and 3"},{"comment":"The opening sentence 'Affective polarization and political sorting drive public antagonism' states a causal mechanism as fact rather than as a hypothesis the paper tests. Consider rephrasing to avoid overclaiming.","section":"Abstract"},{"comment":"The inset panels for climate science tweets and Reddit comments are very small and difficult to read; consider presenting them as separate panels or with larger font sizes.","section":"Figure 1.A"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope well and the empirical scale is impressive. The main barrier to acceptance is the lack of measurement validation for the central toxicity construct; this is fixable with a targeted annotation study and sensitivity analyses. Please also ensure the data and code are actually available at the stated DOI at the time of resubmission, since the current wording says 'will be made publicly available.'"},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a real contribution. The authors bring together large multi-platform data—general climate Twitter, climate-science Twitter, and Reddit—and show that COVID-19 references in climate posts are associated with higher toxicity and disagreement, with temporal spikes lining up with pandemic events and with pre-existing anti-internationalist populist cues. The 14% higher hazard of toxicity onset in replies to COVID-19-rooted climate threads is a concrete finding. The account fixed effects and fractional logit robustness checks are proper, and the Cox resampling scheme sensibly handles conversation dependence. Credit where due: the outcome classifiers (Perspective, DistilBERT, CARDS) are external benchmarks, not fitted to this data, and the paper reports held-out performance for the disagreement model.\n\nThe stress-test note is partly right. Perspective is trained on Wikipedia talk pages and NYT comments; the paper does not report a human validation or threshold sensitivity check on climate/COVID tweets. Since every outcome variable passes through that classifier binarized at 0.5, the magnitude of the COVID-toxicity association is uncertain—some of it could come from stylistic features (caps, profanity, anti-elite framing) that Perspective flags as toxic even when human readers would not call the post uncivil. That said, the example tweets in Tables 1 and 3 are genuinely vitriolic, the association is consistent across three datasets and several specifications, and the disagreement results rest on a separately fine-tuned model with reported f1 scores. So this is a significant caveat, not a fatal flaw. A validation on a random sample of a few hundred posts from their own corpus, or a threshold sweep, would materially harden the paper.\n\nSecond soft spot: the main analysis window (Aug 14, 2020–Aug 26, 2021) is chosen from BEAST change points estimated on the same toxicity outcome. That can inflate the apparent gap between COVID and non-COVID posts. It does not undermine the cross-sectional LPMs or the Cox results estimated within the window, but the causal 'spillover' framing in the abstract runs ahead of the observational design.\n\nWho is this for: people working on affective polarization, science communication, or contentious politics online. It deserves a serious referee. If the authors release the promised code/data and add a classifier validation, I would be happy to cite it.","headline":"A solid, well-run observational study of incivility spillover from COVID-19 to climate discourse; the real caveat is the unvalidated toxicity classifier, not the design.","tokens_in":20950,"tokens_out":3651,"would_cite":true,"duration_ms":34423,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The COVID-19 pandemic measurably spilled hostility into public climate debates on social media along pre-existing populist divides.","keywords":["affective polarization","climate change","COVID-19","incivility","contentiousness","social media","science communication","populism"],"falsifier":"Take a random sample of climate posts that mention COVID-19 keywords and a matched sample that does not, have human annotators rate incivility blind to the research hypothesis, and check whether the COVID-19 posts are actually ruder; if the human-rated difference is near zero, the central association is an artifact of the toxicity classifier.","tokens_in":19917,"feed_emoji":"🗯️","tokens_out":9469,"duration_ms":81071,"temperature":0.7,"pith_summary":"The paper tries to show that the hostile, polarized way the public talked about COVID-19 spilled over into how the same publics talked about climate change. Analyzing 134 weeks of English-language posts—38.6 million climate tweets, 2.1 million Reddit comments, and full reply trees of hundreds of thousands of Twitter conversations—it finds that climate posts and replies mentioning COVID-19 were consistently more toxic and more contentious, and that replies to conversations starting with a COVID-19 tweet had roughly 14% higher odds of turning toxic at any given moment. The authors argue this spillover was not random: it activated along pre-pandemic political cleavages, especially anti-internationalist populist beliefs that tied climate-policy opposition to vaccine hesitancy. If true, this means affective polarization does not stay inside one issue but becomes entrenched across science-policy domains, changing how publics engage with climate science.","feed_headline":"COVID-19 talk made climate replies 14% likelier to turn toxic","feed_subtitle":"Pandemic-era antagonism spilled into climate debates on Twitter and Reddit along existing populist divides.","key_machinery":"The argument is carried by a set of machine-classifier measures and survival/regression models applied to conversation trees. Toxicity, standing in for incivility, is the Perspective API's probability that a post is \"rude, disrespectful, or unreasonable,\" binarized at 0.5; disagreement, standing in for contentiousness, comes from a DistilBERT model fine-tuned on the DEBAGREEMENT dataset of comment-reply pairs; obstructionist climate claims are scored with the Augmented CARDS model. These outcomes are linked to keyword indicators for COVID-19, Fauci, and international organizations through linear probability models with day and account fixed effects, fractional logistic regressions, a Cox proportional hazards model of time-to-toxicity in conversation threads, and BEAST time-series change-point estimation. The Cox model is the piece that gives the headline 14% figure, because it measures whether conversations that begin with COVID-19 content reach a toxic reply sooner.","core_discovery":"On its own terms, the paper claims that COVID-19 functioned as a cross-domain conduit for affective polarization in public engagement with science. Across Twitter and Reddit, climate-related posts that referenced COVID-19 (and especially Anthony Fauci) had higher toxicity probabilities than otherwise similar posts, and climate conversations related to COVID-19 showed elevated disagreement between message pairs. In general climate conversations, a root tweet mentioning COVID-19 shortened the time until a reply became toxic, with about 14% higher odds of toxicity onset at any given time; in conversations explicitly engaging climate-science publications, COVID-19 content appeared both in the initial tweet and in disagreeing responses, but rarely escalated to personal attacks. The paper further claims that the increase in incivility toward international organizations in climate discussions during the pandemic—measured against a pre-pandemic baseline—shows that the spillover traveled through pre-existing anti-internationalist populist sentiments rather than through COVID-19 content alone.","pith_inferences":["An implication the authors leave implicit is that the same spillover mechanism should appear in other science-policy clashes—for example, AI governance or gene editing—whenever elite cues align those issues with the same populist cleavages; that is a testable prediction, not a claim established here.","Because the data end in August 2021 and rely on keyword matching, the 14% hazard estimate is likely a lower bound for later pandemic waves; one could check with post-2021 data whether the effect persisted, grew, or decayed.","A direct extension would repeat the analysis for a later health emergency, such as mpox, using human-annotated incivility labels instead of classifier scores, to see whether cross-domain spillover is a general crisis phenomenon.","The anti-internationalist pathway suggests that decoupling vaccine and climate messages from international-elite frames might dampen spillover, but whether such decoupling works is not tested in the paper."],"forward_implications":["If the spillover is real, crisis-driven antagonism in one science-policy domain should be expected to raise the temperature of debate in other domains that share political cleavage lines.","Public-health shocks like COVID-19 can leave a lasting mark on climate discourse: the paper shows toxicity toward international organizations stayed elevated from Q1 2020 until Q3 2021, well after the initial crisis phase.","Science communicators and platform moderators should treat incivility signals as cross-domain: a post that looks like ordinary climate skepticism may carry pandemic-era hostility.","Because the effect appeared on both Twitter and Reddit and survived account fixed effects, it is not just a selection artifact of which users choose to post about COVID-19; the same individuals became more toxic when discussing it.","Conversations that explicitly engage climate-science publications are not immune: they showed more contentious disagreement involving COVID-19, though less personal toxicity, than general climate talk."],"supporting_citations":[{"why":"Defines and supplies the Perspective API toxicity score used as the paper's measure of incivility.","marker":"Jigsaw 2017"},{"why":"Provides the personal-attack annotation dataset that grounds the toxicity scoring approach.","marker":"Wulczyn, Thain and Dixon 2017"},{"why":"Supplies the DEBAGREEMENT comment-reply pairs used to fine-tune the disagreement classifier.","marker":"Pougué-Biyong et al. 2021"},{"why":"Provides the DistilBERT model that the disagreement classifier is built from.","marker":"Sanh et al. 2019"},{"why":"Supplies the Augmented CARDS model for detecting obstructionist climate claims, used as a control and in the disagreement analysis.","marker":"Rojas et al. 2024"},{"why":"Provides the BEAST time-series method that links toxicity and disagreement spikes to pandemic event timing.","marker":"Zhao et al. 2019"},{"why":"Supplies fractional logistic regression, the robustness specification confirming the main probability-model results.","marker":"Papke and Wooldridge 2008"},{"why":"Provides the partisan-sorting account of affective polarization that motivates the cross-domain spillover mechanism.","marker":"Mason 2015"},{"why":"Documents elite-level polarization around COVID-19, the contextual condition the paper says made science-policy links salient.","marker":"Green et al. 2020"},{"why":"Supplies the Reddit climate-change comment dataset used as a second platform.","marker":"SocialGrep 2022"}],"fun_headline_variants":["Pandemic rage leaked into climate debates on social media","COVID-19 incivility spilled into climate chats on Twitter and Reddit","How COVID-19 made climate talks more toxic online","Climate replies turned nastier when they mentioned COVID-19","Polarization from COVID-19 contaminated climate discourse"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the machine-scored measures of toxicity and disagreement genuinely capture rudeness and contentiousness—if the scoring system flags posts merely for mentioning COVID-19 or Fauci, the claimed spillover could be an artifact.","fun_headline_variants_meta":{"raw":{"variants":["Pandemic rage leaked into climate debates on social media","COVID-19 incivility spilled into climate chats on Twitter and Reddit","How COVID-19 made climate talks more toxic online","Climate replies turned nastier when they mentioned COVID-19","Polarization from COVID-19 contaminated climate discourse"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000679,"raw_usage":{"total_tokens":3067,"prompt_tokens":908,"completion_tokens":2159,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":524,"completion_tokens_details":{"reasoning_tokens":2077}},"tokens_in":524,"tokens_out":2159,"duration_ms":15111,"temperature":1.0,"reasoning_tokens":2077,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T20:07:04.252872+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of climate posts that mention COVID-19 keywords and a matched sample that does not, have human annotators rate incivility blind to the research hypothesis, and check whether the COVID-19 posts are actually ruder; if the human-rated difference is near zero, the central association is an artifact of the toxicity classifier.","supporting_citations":[],"review_version":1}