{"id":"81bfb8b3-9a11-47b9-8df1-a87dc8be93cc","arxiv_id":"2507.15600","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Conflicting narrative roles and selected actors in German Twitter debates reveal discursive fault lines that may explain issue alignment between left and right camps.","lead":"This paper studies German Twitter debates on Ukraine, Covid, and climate change to show that political camps tell conflicting stories about the same events, for example disagreeing about NATO's role or invoking Bill Gates. It proposes narratives as a lens for studying polarization and issue alignment, but the evidence is largely qualitative and based on the author's own prior pipeline.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The sign-disagreement conflict networks in §3.2.2 rest on LLM valence labels validated on only 100 tweets at 86% agreement, with admitted irony/sarcasm failures; systematic camp-correlated label noise could manufacture the paper's central conflicting-narrative evidence.","rationale":"The paper's central claim is that conflicting narratives can be systematically extracted from sign-disagreement edges in left/right actantial networks. That claim depends on the LLM-based relation labeling being accurate and unbiased across camps. The paper's own Appendix A.3 reports only a thin validation: 100 tweets, 86% agreement, no confidence intervals, and a stated failure on irony and sarcasm. This is exactly the place where the argument is least secure, because the conflict-network evidence is the quantitative backbone of the systematic comparison, while the close-read examples are qualitative illustrations. The concern is not that the LLM is generally useless; the reader already noted this as the weakest assumption, and I agree. The specific mechanism that makes 86% agreement potentially insufficient is that a conflict edge is defined by a sign difference between two independently labeled camps. Even uncorrelated noise at this rate creates a nontrivial fraction of spurious sign disagreements, and correlated noise—plausible because irony and sarcasm are camp-dependent rhetorical devices—could systematically create conflicts where none exist. The paper itself flags this limitation in the Appendix, so the concern is not manufactured. However, the close-read examples in §4 and Appendix A.4 do provide genuine illustrations of the phenomena, and the qualitative narrative findings are plausible. The correct conclusion is conditional: the method could work, but the central systematic evidence is not yet robust enough to be accepted without additional validation. This does not change the reader's verdict; it reinforces it.","tokens_in":26589,"tokens_out":3419,"duration_ms":41211,"concrete_test":"Build a camp-stratified sample of about 500 tweets that occur on or near sign-disagreement edges and the same number from agreement edges; have two independent annotators, blind to camp, label each (ARG0, ARG1) relation as supportive, conflictive, or neutral and flag irony. Then recompute all conflict edges under three alternative labelings: (a) human labels only, (b) a second open-weight LLM (e.g., Llama-3.1-70B) operating on the original German tweets with the same prompt, and (c) the context-free dictionary baseline from §A.3. If the conflict network retains fewer than 70% of its edges under (a) and (b), or if a bootstrap 95% confidence interval for LLM-vs-human agreement is wider than ±8%, then the sign-disagreement evidence in §4 should be treated as provisional rather than demonstrative.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central evidence for conflicting narratives is the set of edges whose sign (supportive/conflictive/neutral) differs between the left and right actantial networks (§3.2.2, Figs. 3, 5, 7). Those signs are produced by the Phi-4 LLM prompt in §A.2 after German-to-English translation and AMR parsing. The only validation (§A.3) is 86% agreement with human coding on 100 tweets, with no confidence intervals, no per-class breakdown, and an explicit admission that irony and sarcasm are systematically misclassified. A 14% raw label error matters more than it might appear because conflict edges require a sign disagreement between two independently labeled camps. Under an independent-error model with per-edge error q, roughly 2q(1−q) of edge pairs will show spurious sign disagreement, so a 14% error rate could already generate a substantial share of apparent conflicts. More seriously, the errors are unlikely to be independent: ambiguity from irony, reported speech, and implied stance is not uniformly distributed across ideological camps, and the prompt's rule that reported or direct speech is 'neutral' may interact with camp-specific rhetorical style. If right-leaning tweets use sarcasm or indirect attribution more often, the LLM's known failure mode could systematically flip signs in exactly the right-leaning network, producing conflict edges that reflect classifier bias rather than narrative divergence. The Appendix's own example of 'elect Armin Laschet now, because only he can build us an ark' is precisely such a case: the LLM labels a sarcastic supportive-sounding sentence as supportive, which would hide a real conflict or create a false alignment. Because no sensitivity analysis is reported and no data or code are released, the sign-conflict networks cannot currently be distinguished from label-noise artifacts.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes conflicting narratives in German Twitter discourse around the war in Ukraine, Covid, and climate change. Using a pipeline that combines AMR parsing, translation, and LLM-based labeling of actantial relations, it constructs actantial networks for left- and right-leaning camps and identifies edges whose supportive/conflictive/neutral sign disagrees between the camps. These sign-disagreement networks are interpreted as fault lines in polarized debates, revealing, for example, opposite characterizations of NATO's role in Ukraine and the emergence of Bill Gates in the right-leaning Covid narrative. The paper also presents qualitative evidence for cross-issue narrative alignment, such as recurring antagonism toward the media and the Green Party on the right and a solidarity meta-narrative on the left. The central methodological novelty is the LLM-assisted labeling of actantial links, validated in Appendix A.3 against 100 hand-coded examples.","tokens_in":26925,"tokens_out":3322,"duration_ms":41042,"significance":"If the central claims hold, the paper offers a systematic and interpretable way to surface narrative divergence in large-scale social media corpora, bridging computational text analysis and close reading. The strengths include a fully documented pipeline, the inclusion of many direct tweet quotations that ground the interpretation, a validation comparison against a dictionary-based baseline (46% versus 86% agreement), and an explicit discussion of limitations. The bookkeeping of sign-disagreement networks is a concrete operationalization of 'conflicting narratives' that could be reused by other researchers. However, the evidential weight of the sign-disagreement networks and the issue-alignment section depends on the reliability of the LLM labels and on the systematicity of the qualitative analysis, both of which need strengthening before the claims can be fully accepted.","major_comments":[{"comment":"The validation of the LLM-based relation labeling is too thin to support the central sign-disagreement evidence. The paper reports 86% agreement on 100 tweets, but gives no confidence intervals, no per-class breakdown (supportive/conflictive/neutral), and no breakdown by ideological camp. Because conflict edges in §3.2.2 require a sign disagreement between two independently labeled networks, a 14% raw error rate, if non-independent across camps, could generate a large share of spurious conflicts. The admitted irony/sarcasm failure is particularly worrying: the Appendix's own example of 'elect Armin Laschet now, because only he can build us an ark' is a mislabeled tweet of exactly the kind that could feed a conflict network. I request a stratified validation on left and right subcorpora, a per-class error matrix, and a sensitivity analysis showing how many edges in Figures 3, 5, and 7 remain conflicted when label noise is simulated or when a score-magnitude threshold (e.g., |σ| > 0.1) is imposed.","section":"§A.3 and §3.2.2"},{"comment":"The issue-alignment claim rests on a curated set of example tweets selected after the fact, with no systematic quantification. The paper introduces the notion that recurring actors and analogies bind issues together, but it does not measure, for instance, overlap in the sets of antagonized actors across the three issue-specific conflict networks, or the frequency of cross-issue keyphrases like 'climate lockdown' in the entire corpus. Without such a systematic component, the 'first evidence' language in the abstract and Section 4.4 overstates what is currently a qualitative observation. Please either add a quantitative analysis of narrative alignment (e.g., network overlap, co-occurrence of actors across issue corpora) or clearly rephrase the claim as a hypothesis-generating observation.","section":"§4.4"},{"comment":"The formal definition of a conflict edge uses sign(α_l(i,j)) ≠ sign(α_r(i,j)), but α is never explicitly defined; the text elsewhere defines a score σ ∈ [−1, 1]. This inconsistency matters because a threshold on the score magnitude is needed to avoid treating numerically negligible sign differences as narrative conflicts. A score of +0.01 in one camp and −0.01 in the other would count as a conflict edge, which is arguably noise. Please define α, clarify its relationship to σ, report the score distributions for the edges included in the conflict networks, and test robustness to the threshold used.","section":"§3.2.2"},{"comment":"The claim that certain actants (e.g., Bill Gates, vaccination side effects) are absent from one side's narrative is inferred from their absence in the filtered networks, but those networks are restricted to the top-100 central nodes in each camp and are further thresholded by retweet counts. Absence in the displayed network may reflect the centrality cutoff or the weight threshold rather than a genuine narrative omission. Please report the rank or unthresholded weight of the supposedly absent actors in the relevant networks, or verify by direct corpus search that these actors are effectively not discussed in the corresponding camp's tweets for the issue at hand.","section":"§4.2.2 and §3.2"}],"minor_comments":[{"comment":"There are several typographical issues, including 'warin Ukraine' in the abstract and 'recieved' in Section 3.1.2; a careful proofreading pass is needed.","section":"Abstract and §1"},{"comment":"The prompt definition lists 'approves of' twice in the supportive relations definition; the duplicate should be removed.","section":"§A.2"},{"comment":"The notation alternates between σ and α for the edge score, and the relationship between 'score' and the qualitative sign used in the figures is not always explicit; unifying the notation would improve clarity.","section":"§3.2.1 and §4.1.1"},{"comment":"The caption states that the color bar is 'only plotted here for subsequent actantial networks,' which is confusing because Figure 2 itself uses the color scale; please rephrase the caption to describe the color encoding directly.","section":"Figure 2 caption"},{"comment":"The discussion would benefit from a short paragraph on the representativeness of the selected three issues relative to the full issue list in Table A1; for example, the paper could state whether the conflict-network findings are likely to generalize to less-polarized topics.","section":"§5 Discussion"}],"recommendation":"major_revision","confidential_remarks":"The manuscript builds heavily on the author's own prior work (Pournaki et al. 2025; Pournaki and Willaert 2024) for the dataset, the ideological camp assignment, and the narrative extraction method. This is not a problem per se, but the current paper inherits any limitations of those methods without re-validating them. The referees should be aware that the validation of the LLM labeling is conducted on a sample drawn by the authors, and the conflict-network evidence could be sensitive to the prompt design; an independent replication of the labeling on a small sample by the journal's reviewers might be valuable. The paper does not mention code/data availability, which is a limitation given the computational nature of the work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Punchline: this is a credible and readable application of the author's narrative-extraction method to German Twitter data, and it offers a useful typology of conflicting narratives. The weaknesses are real but not disqualifying.\n\nWhat is new: the two conflict types—opposite roles for same actants, different actants for same event—are clearly demonstrated with rich examples. The LLM-based relation labeling is a sensible extension of the earlier dictionary-based approach, and the validation, though small, shows a large improvement over the dictionary baseline. The close reading is careful and the quotes are well chosen.\n\nSoft spots: the validation is thin (100 tweets, 86% agreement, no CI, admitted irony/sarcasm failures). Because the sign-disagreement networks are built on those labels, camp-correlated label noise could produce spurious conflicts. The paper partially mitigates this by grounding every claimed conflict in direct quotes, so the qualitative findings would survive even if some network edges are mislabeled. The issue-alignment section is explicitly exploratory and relies on curated examples; it should be framed as hypothesis-generation, not evidence. No data or code is released, and the pipeline depends on multiple unpublished thresholds. The heavy reliance on the author's own prior work is not a problem in itself, but it makes independent checking harder.\n\nThe paper is honest about its limitations, and the Discussion is appropriately cautious. The central argument—that camps construct opposing realities through narrative role assignment—holds up as a descriptive claim, even if the computational scaffolding is less solid than it appears.\n\nWho this is for: computational social scientists studying polarization and narrative; qualitative researchers looking for a systematic way to identify fault lines. It deserves a serious referee, but the referees should push for more validation, sensitivity analysis, and data release.\n\nRecommendation: send to peer review with the expectation of major revisions; the conceptual core is sound.","headline":"A useful qualitative typology of conflicting narratives, undermined only by thin LLM validation and a curated issue-alignment section; deserves peer review despite the soft spots.","tokens_in":27482,"tokens_out":3309,"would_cite":true,"duration_ms":33675,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Conflicting narratives in German tweets reveal the discursive fault lines behind political polarization.","keywords":["polarization","narratives","actantial networks","Twitter/X","issue alignment","abstract meaning representation","LLM text annotation","German Twittersphere"],"falsifier":"Take a fresh sample of tweets that produced sign-opposite edges in the conflict networks, have annotators blind to the author’s camp label the valence by hand, and check whether the conflicts survive; a rate near chance, or a collapse of sign-opposite edges when ironic or sarcastic tweets are removed, would show the conflicts are artifacts of label noise rather than genuine narrative divergence.","tokens_in":26389,"feed_emoji":"🗣️","tokens_out":11214,"duration_ms":102095,"temperature":0.7,"pith_summary":"This paper tries to establish that the discursive side of political polarization can be read from conflicting narratives in social-media text. Working with about twenty million German tweets on trending topics from 2021 to 2023, it divides users into left- and right-leaning camps from retweet networks and then compares the actantial roles each camp assigns to the same public actors. The evidence shows two recurring conflict patterns: the same actor is cast as helper in one camp and culprit in the other (NATO and the USA in the Ukraine war), and different actors are emplotted for the same event (Bill Gates, vaccination side effects, and the phrase “climate lockdown” appear mainly on the right). The paper also reports first signs of narrative alignment, where recurring antagonists such as the media and the Green Party, and repeated appeals to freedom versus solidarity, bind separate issues into one ideological story. A reader should care because this gives a concrete textual handle on the fault lines behind structural polarization.","feed_headline":"Rival camps on Twitter tell opposite stories of NATO, Covid, climate","feed_subtitle":"A new analysis maps polarization as conflicting roles: NATO as helper or culprit, Putin as savior or war criminal.","key_machinery":"The paper’s central object is the actantial network: nodes are actors (NATO, Russia, vaccine, Bill Gates, the collective “we”), and a directed edge from actor i to actor j records how often tweets express a relationship from i to j, with a weight equal to the retweet count and a score between −1 and +1 indicating conflictive, neutral, or supportive valence. To build these networks, the corpus is translated from German to English, parsed into Abstract Meaning Representation graphs, and the ARG0–ARG1 pairs are extracted as candidate relationships; an open-weight LLM (Phi-4) then labels each relationship using the original tweet as context, replacing an earlier context-free verb dictionary. Conflicting narratives are defined operationally as edges whose score has opposite sign in the left and right networks; keeping only the most central actors yields conflict networks that pinpoint the disputed links. The same networks, restricted to the node “we,” serve as identity narratives, and the comparison of recurring antagonists and cross-issue keyphrases provides the mechanism for studying narrative alignment.","core_discovery":"The central discovery is that conflicting narratives can be extracted and displayed as sign-opposite relationships in actantial networks, and that these networks surface the interpretive fault lines on which polarization is built. For the war in Ukraine, the left-leaning network attributes to NATO and the USA a supportive role in defending Ukraine and ending the war, while the right-leaning network attributes to them a conflictive role as instigators who profit from the war; the same event of Putin’s arrest warrant for deporting children is narrated as justice in one camp and as saving children from a war zone in the other. For Covid, the camps disagree on the vaccine’s effectiveness and side effects, and the right emplots Bill Gates and fear-mongering politicians where the left emplots solidarity and scientific guidance. For climate change, the left treats floods and extreme weather as consequences of emissions and defends the Last Generation protests, while the right relativizes CO2, condemns climate terrorists, and frames climate policy as a threat to freedom. Across all three issues the paper reports a recurring pattern: the right leans on a meta-narrative of individual freedom under attack by elites, the media, and the Green Party, while the left leans on solidarity and shared responsibility, and cross-issue phrases like “climate lockdown” suggest these narratives actively align opinions.","pith_inferences":["Beyond the paper: because the validation explicitly finds irony and sarcasm mislabeled, a sarcasm-aware labeler could change which edges appear as conflicts; the present conflict networks should be re-examined with such a model before the sign-opposite links are taken at face value.","Beyond the paper: if “climate lockdown” is a genuine alignment device, then tracking the spread of such cross-issue keyphrases over time would predict which issues become ideologically bundled before structural alignment appears in retweet networks.","Beyond the paper: the same actantial-network comparison could be applied to mainstream media coverage, testing whether the Twitter camps’ narratives are mirrored, amplified, or challenged by legacy outlets."],"forward_implications":["If the central claim is right, polarization is visible not only in who retweets whom but in the stories told: the same actors (NATO, vaccines, floods) can simultaneously play savior and culprit in the two camps.","Sign-opposite links in actantial networks give a systematic way to locate the specific points of tension in a polarized debate, turning “discursive fault lines” into a searchable object.","Narrative alignment through shared antagonists and analogies such as “climate lockdown” offers a discursive mechanism that could explain why polarization spreads across otherwise unrelated topics.","The extraction method is reusable for any event covered on social media when translation and AMR parsing are available, making cross-country and cross-platform comparisons possible.","Analyzing actantial networks over time could detect narrative shifts as links change sign or new actors enter the plot, connecting narratives to the dynamics of polarization."],"supporting_citations":[{"why":"Supplies the Twitter dataset of German trending topics from 2021 to 2023 and the structural retweet-based polarization and issue-alignment results this paper builds on.","marker":"(Pournaki et al., 2025)"},{"why":"Provides the narrative-signal extraction method, turning AMR graphs into actantial networks, which the present paper extends with LLM labeling and conflict extraction.","marker":"(Pournaki and Willaert, 2024)"},{"why":"Defines Abstract Meaning Representation, the semantic parsing framework used to turn tweets into structured graphs from which actants are extracted.","marker":"(Banarescu et al., 2013)"},{"why":"Supplies the mbart multilingual translation model used to render German tweets into English for AMR parsing.","marker":"(Tang et al., 2020)"},{"why":"Provides Phi-4, the open-weight LLM that labels actantial relationships as supportive, conflictive, or neutral with tweet context.","marker":"(Abdin et al., 2024)"},{"why":"Provides VerbAtlas, the verb-family dictionary baseline whose context-free labeling is compared against the LLM approach in validation.","marker":"(Di Fabio et al., 2019)"},{"why":"Establishes retweet networks as a proxy for opinion clusters, the premise for assigning tweets to left- and right-leaning camps.","marker":"(Conover et al., 2011)"},{"why":"Supplies the stochastic blockmodel method used to detect opinion clusters within each retweet network.","marker":"(Peixoto, 2019)"},{"why":"Provides BERTopic, the topic-modeling pipeline that assigns tweets to issues such as Covid, Ukraine, and climate change.","marker":"(Grootendorst, 2022)"}],"fun_headline_variants":["Twitter camps cast NATO as hero or villain","Conflicting narratives reveal polarization's fault lines","Same events, opposite stories: how Twitter polarizes","Narrative networks expose ideological splits","Freedom vs solidarity: meta-stories drive polarization"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the LLM’s supportive/conflictive/neutral labels reflect the narrative valence German speakers actually intended, after automatic translation and AMR parsing; the paper’s own validation reaches 86% agreement with human coding on 100 tweets and explicitly concedes that irony and sarcasm are mislabeled.","fun_headline_variants_meta":{"raw":{"variants":["Twitter camps cast NATO as hero or villain","Conflicting narratives reveal polarization's fault lines","Same events, opposite stories: how Twitter polarizes","Narrative networks expose ideological splits","Freedom vs solidarity: meta-stories drive polarization"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000323,"raw_usage":{"total_tokens":1853,"prompt_tokens":1024,"completion_tokens":829,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":640,"completion_tokens_details":{"reasoning_tokens":762}},"tokens_in":640,"tokens_out":829,"duration_ms":10223,"temperature":1.0,"reasoning_tokens":762,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:27:18.727781+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a fresh sample of tweets that produced sign-opposite edges in the conflict networks, have annotators blind to the author’s camp label the valence by hand, and check whether the conflicts survive; a rate near chance, or a collapse of sign-opposite edges when ironic or sarcastic tweets are removed, would show the conflicts are artifacts of label noise rather than genuine narrative divergence.","supporting_citations":[],"review_version":1}