{"id":"3138954e-b134-420f-b433-b8610872308a","arxiv_id":"2606.21996","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Cross-lingual analysis of 1.76M Singapore comments finds culturally specific hate targets but shared binding moral grammar and threat frames across languages.","lead":"This paper audits online hate across English, Chinese, and Malay communities in Singapore by annotating a large social media corpus with large language models. It reports that targets of hate differ by language while threat frames and moral foundations are more consistent.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Validation covers only binary hate detection; structural labels (frames, morals) lack human cross-lingual checks","rationale":"The reader's weakest_assumption correctly identifies LLM reliability as critical, but the load-bearing gap is narrower and more technical: the validation protocol stops at binary hate and does not extend to the structural variables whose cross-lingual stability is the paper's main result. Replicating findings with a second model mitigates some risk but cannot replace missing human ground truth on those variables. This moves the verdict from UNVERDICTED to CONDITIONAL pending the proposed check.","tokens_in":1909,"tokens_out":334,"duration_ms":29852,"concrete_test":"Create a new stratified human gold set of 300 comments (100 per language) labeled for moral foundations and threat frames by bilingual annotators; compute per-language κ between Phi-4 and humans—if any language or category falls below 0.70, recompute the V statistics on the human labels.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires accurate LLM extraction of targets (V=0.25), threat frames, and moral foundations (V=0.08/0.07) to demonstrate monotonic decline in cross-lingual divergence. The reported benchmark (Phi-4 accuracy 0.95, κ=0.91) applies only to hate detection against a human gold set; no equivalent human adjudication is described for the finer-grained moral grammar or frame classifications that drive the layered-contingency thesis. Model-specific artifacts in non-English languages could therefore produce the observed convergence even if human judgments diverge.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.3","summary":"The manuscript analyzes a 2025 Singapore corpus (31M items, 1.76M comments on 11 identity groups) across English, Chinese, and Malay to test cross-lingual structure in online hate. After benchmarking eight LLMs against a human gold set and selecting Phi-4 (accuracy 0.95, κ=0.91), it reports that target selection is culturally specific (language × target V=0.25) while threat frames and binding moral foundations (sanctity/loyalty dominant, V=0.08/0.07) converge, supporting a 'layered cultural contingency' thesis; hate is contempt-driven and anti-immigrant, with selective amplification and no event linkage. Absolute prevalence is treated as ill-defined due to model disagreement (κ≈0.42).","tokens_in":2033,"tokens_out":545,"duration_ms":20116,"significance":"If the structural annotations are reliable, the monotonic decline in divergence from targets to morals provides a testable, empirically grounded distinction between culturally variable and shared components of hate speech, with direct implications for multilingual moderation. The human gold set for binary detection, replication under a second model, and focus on relative structure rather than absolute rates are positive features.","major_comments":[{"comment":"The reported validation (accuracy 0.95, κ=0.91 on independent manual check) applies exclusively to binary hate detection against the human gold set. No parallel human cross-lingual adjudication is described for the target, threat-frame, or moral-foundation labels whose divergence statistics (V=0.25 → 0.08/0.07) carry the central layered-contingency claim.","section":"Methods (LLM benchmarking and annotation pipeline)"},{"comment":"Post-selection of Phi-4 as the primary annotator, combined with the acknowledged low inter-model agreement on prevalence (κ≈0.42), leaves open the possibility that model-specific artifacts in non-English languages drive the observed convergence in frames and morals even if human judgments diverge.","section":"Results (model replication and structural comparisons)"}],"minor_comments":[{"comment":"Clarify whether the second-model replication applied the identical prompt templates and label schemas used for Phi-4 or introduced any adjustments.","section":"Methods"},{"comment":"The abstract states that 'volume does not track 2025 Singapore key events'; provide the exact event list and statistical test used for this null result.","section":"Results"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful review and for noting the strengths of the human gold set for binary detection, the model replication, and the focus on relative structure. We respond to each major comment below.","responses":[{"response":"The referee correctly identifies that human validation was performed only for binary hate detection. Target, threat-frame, and moral-foundation annotations were produced by the selected LLM and subjected to full replication under the second model; the monotonic decline in divergence (V=0.25 to 0.08/0.07) is reproduced in both. We will add an explicit limitations paragraph acknowledging the lack of human cross-lingual adjudication for the structural labels and will clarify how cross-model consistency functions as the primary robustness check for those annotations.","revision_made":"partial","referee_comment":"[Methods (LLM benchmarking and annotation pipeline)] The reported validation (accuracy 0.95, κ=0.91 on independent manual check) applies exclusively to binary hate detection against the human gold set. No parallel human cross-lingual adjudication is described for the target, threat-frame, or moral-foundation labels whose divergence statistics (V=0.25 → 0.08/0.07) carry the central layered-contingency claim."},{"response":"All structural comparisons were replicated under the second model, and the convergence patterns in threat frames and moral foundations remain unchanged. Because the low inter-model κ on prevalence is already acknowledged, the manuscript centers on relative structure; divergent model artifacts would be expected to produce inconsistent structural signals across models, which is not observed. We will expand the methods and results sections to report the replication statistics specifically for the non-binary labels and to state that this replication directly tests against language-specific model biases.","revision_made":"yes","referee_comment":"[Results (model replication and structural comparisons)] Post-selection of Phi-4 as the primary annotator, combined with the acknowledged low inter-model agreement on prevalence (κ≈0.42), leaves open the possibility that model-specific artifacts in non-English languages drive the observed convergence in frames and morals even if human judgments diverge."}],"tokens_in":1591,"tokens_out":463,"duration_ms":32559,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main result is that cross-lingual divergence in online hate drops steadily from targets (V=0.25) to moral foundations and emotions (V=0.08 and 0.07) in Singapore's English-Chinese-Malay setting. Targets are more culturally specific, but the binding morals (sanctity, loyalty) and contempt-driven frames are shared, with hate often anti-immigrant rather than anti-system.\n\nThe work does a few things right. The corpus is large and timely, drawn from 2025 Facebook, Reddit, and YouTube. Benchmarking eight LLMs against a human gold set for hate detection, then replicating with a second model, gives some grounding. Reporting relative structures instead of absolute prevalence is sensible given the kappa ceiling around 0.42.\n\nThe soft spot is the validation gap the stress test flags. Human adjudication covers only binary hate detection; there is no equivalent check described for the threat frames or moral foundation labels that carry the layered-contingency claim. If Phi-4 introduces language-specific artifacts in Chinese or Malay on those finer categories, the observed convergence could be an artifact rather than a real pattern. The abstract does not report cross-lingual human agreement on the structural annotations.\n\nThis is for computational social scientists and moderation researchers working on multilingual platforms. The Singapore case is a clean natural experiment and the data effort is real. It deserves peer review so referees can examine the annotation pipeline for the moral and frame layers, even if the binary detection step looks reasonable.","headline":"Singapore multilingual hate shows targets varying by language while moral frames converge, but the structural labels rest on unvalidated LLM output beyond binary detection.","tokens_in":2482,"tokens_out":382,"would_cite":false,"duration_ms":28040,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Online hate in Singapore shows language-specific targets but shared moral and emotional structures across communities.","keywords":["online hate","cross-lingual","Singapore","moral foundations","threat frames","social media","content moderation","multiculturalism"],"falsifier":"Human raters from each language group rating the same comments differently from the LLM on moral foundations or emotions at levels inconsistent with the high agreement scores reported.","tokens_in":2803,"feed_emoji":"","tokens_out":624,"duration_ms":32241,"temperature":0.7,"pith_summary":"The paper studies online hate discussions in English, Chinese, and Malay within Singapore's multicultural setting using millions of social media comments. It claims that the specific out-groups targeted vary by language, but the threat frames, moral foundations, and emotions expressed in that hate are much more uniform. This indicates a layered structure to hate where cultural differences are stronger at the level of targets than at the level of underlying grammar. The analysis relies on LLM annotations validated against human judgments to compare these layers. Understanding this could help in designing moderation that addresses common patterns rather than isolated language differences.","feed_headline":"Hate targets vary by language but moral structures stay consistent","feed_subtitle":"Singapore data shows divergence drops from targets to frames and morals in English, Chinese and Malay hate speech.","key_machinery":"Layered cultural contingency, which describes how divergence falls monotonically from what is hated to how and why it is hated.","core_discovery":"The paper's core claim is that cross-lingual divergence in online hate decreases as analysis moves from targets to structures: language-by-target association is V=0.25, but drops to V=0.08 for moral foundations and V=0.07 for emotion, with binding morals of sanctity and loyalty prominent at 55-75% and hate being contempt-driven with anti-immigration focus.","pith_inferences":["Moderation policies might benefit from targeting shared structural elements like moral language instead of specific topics.","Patterns observed here could be tested in other multilingual societies with parallel language communities.","Further human validation per language could test the reliability of the LLM-based cross-lingual comparisons.","Volume of hate not tracking events suggests it's driven by ongoing social dynamics rather than temporary triggers."],"forward_implications":["Out-group targets for hate are specific to each language public.","Threat frames and binding morals like sanctity and loyalty are shared across languages.","Hateful comments overall receive less amplification than neutral ones, but anti-immigrant hate is amplified more.","Hate expresses out-group grievances rather than anti-system ones.","Absolute hate rates are hard to define due to low inter-model agreement, so relative structures are emphasized."],"fun_headline_variants":["Hate targets differ by language but morals align across tongues","Language shapes targets yet moral frames stay consistent cross-lingually","Targets vary culturally while binding morals converge in multilingual hate","Divergence drops from targets to shared moral structures in online hate"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The LLM chosen for annotation accurately reflects human perceptions of hate, frames, and morals in all three languages.","fun_headline_variants_meta":{"raw":{"variants":["Hate targets differ by language but morals align across tongues","Language shapes targets yet moral frames stay consistent cross-lingually","Targets vary culturally while binding morals converge in multilingual hate","Divergence drops from targets to shared moral structures in online hate"]},"model":"grok-4.3","cost_usd":0.003836,"raw_usage":{"total_tokens":2052,"prompt_tokens":820,"num_sources_used":0,"completion_tokens":67,"cost_in_usd_ticks":38362000,"prompt_tokens_details":{"text_tokens":820,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1165,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":820,"tokens_out":67,"duration_ms":11604,"temperature":1.0,"reasoning_tokens":1165,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T11:00:47.916738+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Human raters from each language group rating the same comments differently from the LLM on moral foundations or emotions at levels inconsistent with the high agreement scores reported.","supporting_citations":[],"review_version":1}