{"id":"41dc125e-baaf-4cef-8e48-a710b632deff","arxiv_id":"2412.05741","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"Toxic and insulting comments are more likely near the end of YouTube political conversations, but the study does not establish that toxicity causes the silence.","lead":"This study analyzes 32 million YouTube comments on US news channels around the 2020 election and fits hidden Markov models to conversations. It reports that in a latent 'terminal' state, toxic and insulting posts are more likely, which the authors interpret as evidence that toxicity silences minority opinions.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'silence' state is an artifact of appending an end-of-conversation symbol; without an independent measure of user disengagement, the HMM cannot establish that toxicity silences users.","rationale":"The reader's weakest assumption is correct and is the most load-bearing point in the paper. The HMM is fitted to sequences that all end in an artificially appended 0, so the identification of a terminal state is essentially guaranteed by construction. The paper's central claim, however, requires that this terminal state corresponds to self-censorship—a theoretically meaningful silence by users who refrain from participating because of toxicity. Nothing in the data measures silence: no user-level abstention, no deleted comments, no moderation logs, no natural stopping reason. The 10-day window and YouTube's first-level reply structure alone can end a conversation. Therefore the reported RRX3 > 1 in Section 2.4 is compatible with a simpler explanation: the last observed comment in a thread is somewhat more likely to be toxic, for reasons unrelated to silencing. This is an external-validity gap, not an internal inconsistency in the HMM machinery. The paper itself acknowledges in Section 2.4 that 'we may lack direct information about the specific states a conversation undergoes,' which underscores the concern. I do not see a way to recover the silencing claim from the current operationalization; the natural fix is to validate the terminal state against user-level churn, which is exactly the test proposed. The large descriptive dataset and the robustness checks across channels and topic clusters are real strengths, but they do not bear on the validity of the 'silence' label. Since the central claim in the title and abstract overstates what the analysis can identify, the rejection stands.","tokens_in":21500,"tokens_out":7453,"duration_ms":74582,"concrete_test":"Link the inferred states to user-level attrition. For each conversation, identify all participants and determine, from the full dataset, whether each participant posted any comment on any of the six channels within, say, 30 days after that conversation's last comment. Compute the attrition rate per conversation (fraction of participants with no subsequent comment) and regress it on the HMM's posterior probability that the conversation is in state Z1 at its final observed comment, controlling for channel, video, conversation length, and toxicity history. If Z1 really represents silencing, conversations whose final state is Z1 should show markedly higher attrition. If the posterior probability of Z1 does not predict actual user disengagement, the 'silence' label is unsupported and the central claim cannot be sustained.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that toxic behavior silences users depends on interpreting the HMM's terminal latent state as 'self-censorship'. In SI S4 the authors state they 'appended a zero (X1=0) to the end of every sequence/conversation to indicate its conclusion.' Since every observed sequence therefore ends in the artificial symbol 0, any fitted HMM must place P(X1)>0 in some state; the paper reports P(X1|Z1)>0 and P(X1|Z2)=0 (Section 2.4) and labels Z1 the terminal state. The load-bearing step is the semantic leap from 'the observed comment stream ended' to 'users were silenced by toxicity'. The model contains no observation of user absence, deleted comments, or suppressed replies. The stream could have stopped because of the 10-day cutoff, YouTube's first-level reply limitation, thread exhaustion, moderator action, or simple loss of interest. The relative risk RRX3=P(X3|Z1)/P(X3|Z2) only shows that toxic/insulting comments co-occur with the constructed end marker; it does not show that toxicity caused the end, nor that the end is a silence phenomenon rather than a property of where the platform truncates threads. Without an external validation of Z1 as a genuine silence state, the central claim collapses into the trivial statement that conversations have an end, plus a weak correlation between end positions and toxicity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper studies toxic and insulting behavior in YouTube conversations under six US news channels during the 2020 US election period. The authors classify comments with Google's Perspective API, define conversations as top-level comments plus replies received within 10 days, append an artificial end-of-conversation symbol X1=0 to every sequence, and fit a two-state hidden Markov model with observations X1 (end), X2 (non-toxic/non-insulting), and X3 (toxic/insulting). They report that one state, Z1, emits the end marker with positive probability while the other does not, and that the relative risk of toxic/insulting comments in Z1 versus Z2 exceeds 1 for most channels. The authors interpret Z1 as a latent terminal state consistent with toxicity-driven silence and conclude that toxic behavior silences users. They also repeat the analysis with topic-based video clusters and find similar relative-risk patterns.","tokens_in":21816,"tokens_out":8987,"duration_ms":84405,"significance":"The paper brings a very large dataset (32.5M comments) to an important question about self-censorship in online political deliberation. The descriptive finding that toxic and insulting comments are relatively more probable near the end of conversations is potentially interesting, and the topic-cluster robustness check is a plus. However, the interpretive leap from 'the observed comment stream ended' to 'users were silenced by toxicity' is not supported by the design. The end marker is an artifact of the data-collection procedure and platform structure, not an observation of user disengagement. Without an independent measure of silencing (e.g., a user who stops replying after receiving a toxic comment, or a conversation that ends earlier than expected given its activity), the central claim remains a restatement of the correlation between conversation position and toxicity. As a descriptive study of toxicity at the end of threads, the paper might be salvageable, but the current framing overstates the evidence.","major_comments":[{"comment":"The operationalization of 'silence' as the appended end-of-conversation symbol X1=0 has no construct validity for self-censorship. The authors state in SI S4 that they 'appended a zero (X1=0) to the end of every sequence/conversation to indicate its conclusion.' Since every conversation is artificially terminated by the data-collection window (10 days) and by YouTube's first-level-reply structure, the terminal state Z1, which is identified by its emission of X1, cannot be interpreted as a state of user silence without external validation. The abstract's claim that the authors 'observe patterns of self-censorship' is therefore not supported by the measurements.","section":"SI S4; Section 2.4; Abstract"},{"comment":"The relative risks RRX2 and RRX3 are computed from the same fitted emission probabilities that define the latent states, so the finding RRX3>1 is a property of the fitted model rather than an independent test of a silencing mechanism. The paper does not report uncertainty estimates or confidence intervals for these ratios across the repeated HMM fits described in SI Table 4; adding such intervals would be necessary to assess whether the >1 pattern is statistically stable.","section":"Section 2.4, Eqs. (1)-(2)"},{"comment":"The discussion goes beyond the observational evidence by describing 'a circular causality between toxicity and disengagement.' The data cannot distinguish between toxicity causing users to stop commenting, users ceasing to comment for unrelated reasons and the conversation then turning toxic, or platform and moderator effects truncating threads. A concrete alternative explanation is that YouTube only permits first-level replies (as the authors note in Section 2.1) and the 10-day cutoff, so the 'end' marker is independent of user silence.","section":"Section 3 (Discussion)"},{"comment":"The claim that state Z1 is 'characterized by reduced user activity' (abstract) is not measured. The model has no observation of user activity beyond the presence or absence of comments; the appended X1 symbol is not a measure of activity level. This characterization should be removed or replaced with a user-level analysis (e.g., reply rates or return probabilities) to support the silence interpretation.","section":"Abstract; Section 2.4"}],"minor_comments":[{"comment":"There is a typographical error: 'P (X1|Z2 = 0)' should be 'P (X1|Z2) = 0' in the sentence about the four-state model; the same notation appears again in the paragraph describing Figure 4.","section":"Section 2.4"},{"comment":"The sentence 'It is possible to use these mentions to reconstruct sub-threads, but doing so is beyond the scope of this research. Using these mentions, it is therefore possible to reconstruct sub-threads; however, this is beyond the scope of this work.' is repeated with slight variation; the duplication should be removed.","section":"Section 2.1"},{"comment":"The relative risk values in Figure 4 are described only in prose and without error bars; please provide uncertainty or bootstrap intervals, or state that they are omitted.","section":"Section 2.4 / Figure 4"},{"comment":"The bias scores use a decimal comma (e.g., '-2,40' for ABC News); the manuscript should use decimal points consistently.","section":"Table 3"},{"comment":"Reference [66] (a protein-protein interactions overview) is used as the source for the HMM description; a standard HMM reference (e.g., Rabiner 1989) would be more appropriate.","section":"SI S4"},{"comment":"The paper says hSBM 'automatically identifies the number of topics and hierarchical levels,' but then it says 'we work with the clusters at the fourth hierarchical level.' Please clarify whether the fourth level is chosen by the model or by the authors.","section":"Section 2.5"}],"recommendation":"reject","confidential_remarks":"The central finding appears to be an artifact of appending an end-of-conversation marker. The paper would need either an independent measure of user disengagement or a substantial reframing to a descriptive claim about toxicity at thread endings. Given the title and abstract, I do not see how the current manuscript can be accepted. If the authors can provide user-level evidence of silencing, a resubmission may be worth considering."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The real contribution here is narrower than the title claims: the authors show, across six US news channels and nine topic clusters, that a fitted HMM state associated with the end of a conversation emits toxic and insulting comments at elevated rates. That is a genuine, reproducible descriptive pattern, and it holds under two different ways of grouping videos (channel and topic). The dataset is large, the methods section is transparent about the sampling and fitting procedure, and the robustness check with 3–5 hidden states shows they did not just pick the two-state model to make the result appear. Credit where it is due: this is a competently executed observational study of a pattern that, as far as I know, has not been reported before.\n\nThe soft spot is the interpretive leap. The 'silence' state is constructed by appending an artificial end-of-conversation symbol to every sequence. That guarantees an HMM with a terminal state; the only question is what that state emits. Finding that it emits more toxicity is interesting, but it does not tell you why the conversation ended. The stream could have stopped because of the 10-day cutoff, YouTube's thread structure, moderator action, or simple loss of interest. The authors label the state 'toxicity-driven silence' and the abstract says they 'observe patterns of self-censorship.' That is not supported by the data, and the stress-test note is right: without an independent measure of user disengagement, the causal claim collapses into a correlation between end positions and toxicity. The paper itself acknowledges some limitations, but not the load-bearing one.\n\nOther issues are secondary. The relative risks in Figures 4 and 6 are shown without error bars, which matters because the sampling across realizations presumably produced variance. The sample is restricted to top-level comments and first-level replies, and reply-less comments are excluded by construction—a selection that could plausibly correlate with both toxicity and conversation length. No code or data are released, which makes the robustness claims hard to verify.\n\nWho is this for? Researchers working on toxicity dynamics or platform moderation who want a well-described observational baseline on end-of-conversation behavior. It is not evidence for the spiral of silence, and the authors should not sell it as such. A serious referee could push them to reframe the paper as a descriptive finding and add the missing uncertainty quantification. I would not desk-reject it, but I would send it back with the causal language substantially softened.\n\nRecommendation: accept for peer review, with the expectation of major revision. The descriptive core is worth publishing; the current framing is not.","headline":"Solid descriptive finding on end-of-conversation toxicity, but the causal 'silencing' claim outruns the design.","tokens_in":22331,"tokens_out":1158,"would_cite":false,"duration_ms":14230,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper finds that YouTube political conversations tend to end in a terminal state where toxic and insulting comments are more likely, and interprets this as evidence that hostile behavior silences participants.","keywords":["social media","toxic content","hidden Markov model","spiral of silence","political deliberation","self-censorship","YouTube comments"],"falsifier":"Compare matched threads where the final comment is toxic versus non-toxic: if the terminal state reflects self-censorship, users in toxic-ending threads should show a measurable drop in subsequent commenting on the same channel or video relative to users in non-toxic-ending threads. In addition, place the appended zero at a random position within a subset of sequences; if the terminal-state toxicity disappears, it is an artifact of the appended marker rather than a property of real conversations.","tokens_in":21290,"feed_emoji":"💬","tokens_out":9776,"duration_ms":79482,"temperature":0.7,"pith_summary":"Online political threads under six major US news channels tend to end not with a reasoned conclusion but with a burst of toxic and insulting comments, and the paper argues this pattern is evidence of self-censorship. Using around 32.5 million YouTube comments from the 2020 US presidential election period, the authors fit a two-state hidden Markov model to each conversation, treating comments as toxic, non-toxic, or an appended end marker. They find a terminal latent state that emits both the end of the thread and a higher probability of toxic and insulting posts relative to the earlier state, for almost every channel and across topic clusters. If right, this means public comment sections understate the diversity of political opinion because toxicity pushes participants out while amplifying hostile voices.","feed_headline":"Toxic comments cluster at the end of political threads","feed_subtitle":"A hidden Markov model links the final stage of YouTube comment threads to a silence state.","key_machinery":"A two-state hidden Markov model fitted to comment sequences, with three observed symbols: $X_1 = 0$ (an end-of-conversation marker appended to every thread), $X_2 = 1$ (a non-toxic and non-insulting comment), and $X_3 = 2$ (a toxic or insulting comment, scored by the Perspective API toxicity classifier). The model estimates transition probabilities between two latent conversational states and emission probabilities from each state to each symbol. The terminal state $Z_1$ is identified by $P(X_1|Z_1) > 0$ and $P(X_1|Z_2) = 0$; its character is quantified by the relative risks $RR_{X_2} = P(X_2|Z_1)/P(X_2|Z_2)$ and $RR_{X_3} = P(X_3|Z_1)/P(X_3|Z_2)$.","core_discovery":"The paper's central discovery is a latent silence state at the end of online political conversations. After converting each YouTube thread into a sequence of toxic, non-toxic, and end-of-conversation symbols, a two-state hidden Markov model learns that conversations almost never end in the ordinary-posting state Z2; instead they terminate in state Z1, where non-toxic and non-insulting posts are relatively rare and toxic and insulting posts are relatively common, with the relative risk exceeding 1 for nearly all channels and CNN as the exception. Because the end marker was appended by the authors, this finding asserts that the way a thread closes is systematically toxic, and the authors interpret that as toxicity-driven self-censorship: users exposed to hostile replies stop contributing, leaving the final word to toxic content. The pattern is robust across channel-level and topic-level groupings, with the strongest effects for Fox News and for topics such as police brutality, the Black Lives Matter movement, and COVID-19 vaccination.","pith_inferences":["An extension would be to test whether individual users who receive toxic replies subsequently reduce their commenting activity in the same channel; that would move the terminal-state pattern from an aggregate signature to a per-user behavioral claim.","The appended zero conflates 'the conversation ended' with 'the participants were silenced'; a control that appends the end marker at a random position in the sequence would show whether the terminal state's toxicity is a real conversational pattern or a modeling artifact.","The same two-state hidden Markov model could be applied to Reddit or Twitter threads to see whether terminal-state toxicity is a general property of online deliberation or specific to YouTube's reply structure."],"forward_implications":["Platforms that want to sustain civil discussion should treat a rise in toxic replies as a warning that a thread is approaching a terminal state dominated by hostile content.","If the terminal state reflects self-censorship, then visible comment sections understate the diversity of political opinion, because moderate voices drop out while toxic ones have the last word.","The U-shaped pattern across channel bias ratings suggests polarization is self-reinforcing: channels further from the political center show higher terminal-state toxicity, which may push moderate users away.","Because the effect appears across topic clusters rather than only channel groupings, the silencing dynamic is likely tied to the content being discussed, not just to the outlet hosting it."],"supporting_citations":[{"why":"Supplies the comparison baseline for toxicity prevalence across platforms and the threshold used to classify a comment as toxic.","marker":"[33]"},{"why":"Introduces the personal-attack detection approach that underlies the Perspective API toxicity scores.","marker":"[34]"},{"why":"Provides the modern Perspective API classifier used to assign toxicity and insult scores to each comment.","marker":"[35]"},{"why":"Introduces the hidden Markov model framework used to infer latent conversational states.","marker":"[40]"},{"why":"Details the learning problem and parameter estimation that the paper applies to fit its two-state model.","marker":"[42]"},{"why":"Provides the rationale for preferring the two-state model over alternatives with more states.","marker":"[57]"},{"why":"Supplies the topic-modeling method used to replicate the channel-level finding across video topics.","marker":"[59]"}],"fun_headline_variants":["Toxicity silences online political dissent","Political threads end toxic, silencing minority views","Model links online toxicity to self-censorship","When toxicity rises, online speakers go quiet"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that appending an end-of-conversation symbol to every comment sequence and fitting a two-state hidden Markov model produces a hidden state that genuinely represents self-censorship, rather than simply the structural fact that every thread has an endpoint.","fun_headline_variants_meta":{"raw":{"variants":["Toxicity silences online political dissent","Political threads end toxic, silencing minority views","Model links online toxicity to self-censorship","When toxicity rises, online speakers go quiet"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000256,"raw_usage":{"total_tokens":1546,"prompt_tokens":890,"completion_tokens":656,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":600}},"tokens_in":506,"tokens_out":656,"duration_ms":7445,"temperature":1.0,"reasoning_tokens":600,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:23:41.996839+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare matched threads where the final comment is toxic versus non-toxic: if the terminal state reflects self-censorship, users in toxic-ending threads should show a measurable drop in subsequent commenting on the same channel or video relative to users in non-toxic-ending threads. In addition, place the appended zero at a random position within a subset of sequences; if the terminal-state toxicity disappears, it is an artifact of the appended marker rather than a property of real conversations.","supporting_citations":[{"cited_title":"Standford University","cited_arxiv_id":null,"evidence_quote":"Introduces the hidden Markov model framework used to infer latent conversational states."},{"cited_title":"Department of Computer Science, Brown University","cited_arxiv_id":null,"evidence_quote":"Details the learning problem and parameter estimation that the paper applies to fit its two-state model."},{"cited_title":"Journal of Agricultural, Bio- logical and Environmental Statistics 22, 270–293 (2017) https://doi.org/10.1007/s13253-017-0283-8","cited_arxiv_id":null,"evidence_quote":"Provides the rationale for preferring the two-state model over alternatives with more states."}],"review_version":1}