Pith. sign in

REVIEW 4 major objections 5 minor 13 references

Israel-Hamas war through Telegram, Reddit and Twitter

T0 review · 4 major / 5 minor · reviewed 2026-08-10 · deepseek-v4-flash

Pith's one-line read The paper claims that online discussion of the Gaza war is polarized into military and humanitarian narratives, with anger and fear dominant.

desk verdict The paper's own data description is internally contradictory (125K vs 70K vs 105K messages; 2025 vs 2015-2024), so the central claim about manipulation is not supported and the dataset cannot be used as-is. read the letter →

arxiv 2502.00060 v1 pith:NRL7JQES submitted 2025-01-30 cs.SI cs.AIcs.LG

classification cs.SIcs.AIcs.LG
keywords Israel-PalestineconflictTelegramsocialmediadiscourseBERTopicsentimentanalysistopicmodelingpolarizationpropaganda
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that online discussion of the Israel–Hamas war, especially on Telegram, is dominated by two opposed narrative poles—military operations and civilian suffering—and by the emotions of anger and fear. The authors argue that this polarized emotional structure is the signature of how political factions and outside actors shape public opinion during the conflict, and they treat sentiment-topic patterns as possible evidence of manipulation and propaganda. They support the claim with volume analysis, entity extraction, LDA and BERTopic topic modeling, and sentiment and emotion classification across 125,054 Telegram messages, 2,001 tweets, and about 2.1 million Reddit comments. A sympathetic reader would care because the finding would turn social-media emotion into a measurable indicator of coordinated narrative-shaping in real-time crises.

What carries the argument

The load-bearing machinery is a pipeline of four techniques applied to text corpora: volume analysis over time and channels; entity extraction from frequent hashtags and words; topic modeling with LDA and BERTopic, a transformer-embedding method that clusters semantically similar messages into topics; and sentiment and emotion classification using pretrained Hugging Face models, including a fine-grained emotion model that labels anger, fear, hope, and similar states. LDA and BERTopic supply the topic clusters; the sentiment and emotion classifiers supply the emotional tone; and the paper treats the correlation between the two as evidence about how narratives are being shaped.

What would settle it

Rebuild the corpus with random or exhaustive collection and verified timestamps, ideally all Telegram messages in these channels plus a time-matched Twitter sample from the same months, and rerun the identical topic and emotion pipeline; if anger and fear dominance and the military-versus-humanitarian split shrink or vanish, the polarization result is an artifact of which channels and posts were selected.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that conflict discourse across Telegram, Twitter, and Reddit is emotionally polarized and topically clustered: anger and fear dominate, especially in channels reporting combat directly, and the discussion splits into humanitarian themes (Gaza, children, hospitals, displacement) and military or operational themes (strikes, brigades, occupation, resistance). The authors interpret this configuration as polarized narratives being the hallmark of how political factions and outsiders mold public opinion, with the sentiment-topic alignment taken to reveal trends that may show manipulation and attempts of propaganda. In short, the paper claims that the emotional tone of what people say online during the war is not a random reflection of events but a structured product of narrative competition.

Load-bearing premise

The whole argument rests on the assumption that the hand-picked Telegram channels, the 2,001-tweet set, and the Kaggle Reddit dump collectively stand in for conflict discourse, and that the observed emotion-topic patterns reflect narrative manipulation rather than ordinary, organic reactions to a war.

Editorial extensions

If this is right

  • If the reading is right, Telegram is not just a messaging app in this conflict but a primary arena where military and humanitarian narratives compete for emotional engagement.
  • Topic clusters of about 19 percent general Palestinian issues and 17.9 percent geographical and humanitarian aspects imply that a large share of discourse centers on civilian impact, not strategy.
  • The dominance of anger and fear across conflict-reporting channels implies that emotion, not information, carries the engagement in these spaces.
  • Sentiment-topic alignment, if confirmed, gives platforms and researchers a measurable signal for detecting possible propaganda campaigns during crises.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's own cumulative volume plots show message numbers rising sharply after late 2020 and especially after 7 October 2023; a natural extension would compare the emotion mix before and after that spike to separate event-driven anger from sustained narrative-driven anger.
  • With only 2,001 tweets, the Twitter leg is too thin for platform-level conclusions; the same pipeline on a larger, time-matched Twitter sample would test whether the Telegram polarization pattern generalizes.
  • If outside actors truly mold opinion, their fingerprint should be coordinated hashtag bursts across otherwise unconnected channels; the entity table offers candidate tags such as #FreePalestine and #GazaUnderAttack for a burst-coordination test the paper does not run.
  • The paper states the Telegram collection period inconsistently, as starting on 23 October 2025 in the abstract and as spanning 2015 to 2024 in the methodology; resolving that dating would determine whether the reported volume and emotion patterns describe one continuous corpus or two different collections.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper analyzes online discourse about the 2023-2025 Israel-Hamas war using a Telegram corpus, a 2,001-tweet Twitter dataset, and a Reddit dataset of about 2.1 million comments. The stated methods are volume analysis, entity extraction, LDA and BERTopic topic modeling, and sentiment/emotion analysis. The authors report that message volume increased sharply after October 7, 2023, that topics cluster around military operations and humanitarian impact, and that anger and fear dominate the emotional tone. The conclusion goes further, claiming that polarized narratives indicate manipulation and propaganda by political factions and outsiders. The manuscript also deposits a dataset on Zenodo. However, the paper contains severe internal inconsistencies in the reported dataset size and date range, and the causal claim about manipulation is not supported by the descriptive analysis.

Significance. If the descriptive findings were reliable and the dataset were properly documented, the paper could offer a useful case study of multilingual, multi-platform discourse during a geopolitical crisis. The authors have made an effort to share data (Zenodo DOI 10.5281/zenodo.14710657), and the combination of LDA, BERTopic, and emotion analysis is a reasonable toolkit for exploratory work. The main obstacle is that the descriptive results rest on an ill-specified corpus: the paper reports mutually incompatible message counts (125K, 70,313, 70,321, 105K) and an impossible collection start date (23 October 2025). Moreover, the central interpretive claim—that observed sentiment-topic correlations reveal attempts at manipulation and propaganda—has no causal or baseline support and is partly predetermined by choosing already-affiliated Telegram channels. The paper therefore cannot, in its current form, support its advertised conclusions.

major comments (4)
  1. [Abstract; Section 5; Table 1; Section 5.3.2; Section 7] The Telegram corpus is described in irreconcilable ways. The Abstract states '125K messages ... spanning from 23 October 2025 until today' (a future date relative to the paper's submission); Section 5 states a 9-year span with the oldest message on 2015-10-23 and the youngest on 2024-10-24, totaling 70,313 messages; Table 1 lists 125,054 messages as of 2025-01-20; Section 5.3.2 says 51,403 of 70,321 messages were analyzed; and Section 7 refers to 105,000 messages. These numbers cannot all be correct, and the abstract's date range is temporally impossible. Because the volume analysis (Fig 1), topic percentages (Section 5.3.1), and sentiment-topic associations are all computed on this corpus, the reader cannot determine the actual sample size or time period, which undermines every downstream quantitative claim and the reproducibility of the study.
  2. [Abstract; Section 7] The claim that polarized narratives are a 'hallmark of how political factions and outsiders mold public opinion' and that sentiment-topic trends 'may show manipulation and attempts of propaganda' is not supported by the analysis. The study reports descriptive statistics: message frequencies, entity counts, topic proportions, and sentiment/emotion bar charts. There is no measurement of coordinated or inauthentic behavior, no comparison baseline for organic discourse, no identification of actor intent, and no statistical test linking topic-sentiment associations to manipulation. The observed dominance of anger and fear and the presence of military and humanitarian topics are equally consistent with organic audience responses to a violent conflict. Moreover, since the Telegram channels were hand-picked from already-affiliated sources (Section 1.1, Table 1), the 'polarization' conclusion is partly an artifact of sample selection rather than an empirical discovery.
  3. [Abstract; Section 3; Section 5.3; Section 6.3] The paper claims to apply the same topic and sentiment analysis to Telegram, Twitter, and Reddit, but no topic modeling results are presented for the Reddit dataset. Section 5.3 reports LDA and BERTopic results for Telegram and BERTopic for Twitter; Section 6.3 provides only a single Reddit sentiment bar plot. The cross-platform comparative analysis advertised in the Abstract and in the Contributions list is therefore not actually carried out for Reddit, which is a major gap relative to the paper's stated scope.
  4. [Section 5.2] The entity extraction step is not reproducible. Section 5.2 states that hashtags and words were categorized into 'predefined entities' 'based on frequency and manually defined rule,' but no rule, category definitions, inter-annotator agreement, or validation procedure is described. Table 3 mixes hashtags (e.g., '#FreePalestine') with unigrams (e.g., 'gaza') and assigns frequencies without explaining how these are grouped into entities. This makes the subsequent topic interpretation and any claims based on entity prevalence difficult to verify.
minor comments (5)
  1. [Throughout] The manuscript contains numerous typos and mechanical errors, e.g., 'Thesw datasets' (Contributions), 'T able 1' (Table 1 caption), 'weo' in Section 5.3.3, 'enttiets' (Section 5.2), 'discusssed' (Section 5.3.1), and inconsistent capitalization of 'BERTopic'/'BertTopic'. These should be corrected in a thorough revision.
  2. [References] Several references appear as raw LaTeX or are broken, including 'citegayo2011limits,' 'citepreoctiuc2015studying,' 'citeSocialCapitalMarkets' in Section 3, and 'In [?]articleal2019multi' in Section 2. The reference list also has inconsistent formatting; all citations must be properly resolved.
  3. [Section 5.3.1] The LDA section states there are 8 topics but later refers to 'topics 6-10' and 'Topic 9' and 'Topic 10' (Section 5.3.1, paragraphs after Fig 9). The number of topics and the topic numbering need to be made consistent.
  4. [Section 5.1] The text says the main message volume 'was initiated at the end of 2020' and attributes this to events including the COVID-19 pandemic, while the Conclusion states the volume increased dramatically after October 7, 2023. These statements should be reconciled with the actual time series and the stated collection period.
  5. [Appendix] The captions in Appendix 8.1 are generic ('Overall Caption for All Topics') and do not describe the subfigures individually; the subfigures in Fig 20 are labeled with placeholder captions such as 'Caption for Topic 7.' These need to be replaced with informative captions.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: descriptive NLP analysis does not predict fitted quantities, so no derivation reduces to its inputs.

full rationale

The paper is a descriptive, observational social-media study. It applies off-the-shelf tools (BERTopic, LDA, HuggingFace sentiment/emotion classifiers) to three externally sourced or independently collected text corpora and reports topic proportions, sentiment distributions, and their correlations. No parameter is fitted to a subset of the data and then used to predict a closely related quantity; the topic and sentiment outputs are computed directly from the text and are not derived from the conclusions about polarization or manipulation. The claim that the findings 'hint at polarized narratives' and 'may show manipulation' is an interpretive gloss on descriptive trends, not a formal derivation from fitted inputs. The selection of already-affiliated Telegram channels affects external validity and representativeness, but that is a sampling and inference concern, not definitional circularity. The internal inconsistencies in dataset size and date range (abstract's 125K messages from '23 October 2025' versus Section 5's 70,313 messages spanning 2015-10-23 to 2024-10-24, versus Section 5.3.2's 51,403 analyzed messages) are serious correctness/reproducibility defects, but they do not make any stated equation or prediction equivalent to its inputs by construction. No self-citation is load-bearing; the cited third-party Telegram/Reddit/Twitter datasets and pre-trained models are independent sources. Therefore the circularity score is 0.

Assumptions & free parameters 2 free parameters · 4 assumptions · 0 invented entities

The analysis introduces no mathematical free parameters or invented entities, but it relies on researcher-chosen model settings and several domain assumptions about data representativeness, model transferability, and manual annotation. These assumptions are not independently validated.

free parameters (2)
  • LDA topic count = 8
    Authors chose to present 8 LDA topics for Telegram (Section 5.3.1) with no model selection criterion; topics are interpreted subjectively.
  • BERTopic hyperparameters = defaults (not specified)
    BERTopic models are trained with library defaults (MiniLM embeddings, CountVectorizer) with no reported parameters or seeds, making the exact topic set unreproducible.
assumptions (4)
  • domain assumption English-language sentiment and emotion models transfer to conflict-related Telegram, Reddit, and Twitter text.
    The paper applies off-the-shelf Hugging Face models (Section 6) without validating them on the target corpora, which include multilingual and dialectal content.
  • domain assumption The 17 selected Telegram channels provide a representative view of conflict discourse.
    Channels are listed in Table 1 with no inclusion criteria; they range from 1 to 49,178 messages and include overtly affiliated actors such as AlQassamBrigades and The Jerusalem Post.
  • ad hoc to paper Manual entity categorization rules are consistent and objective.
    Section 5.2 says entities are defined 'based on frequency and manually defined rule' without specifying the rule.
  • domain assumption Different platforms and time periods can be directly compared.
    Telegram data spans 2015 to 2024 (or 2025 per the abstract), the Twitter set is 2001 tweets from a GitHub user, and Reddit data comes from Kaggle; no alignment of time windows is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Israel-Hamas war through Telegram, Reddit and Twitter." pith.science (2026). https://pith.science/paper/NRL7JQES

@misc{pith2026250200060,
  author       = {Pith},
  title        = {Pith review of: Israel-Hamas war through Telegram, Reddit and Twitter},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/NRL7JQES}},
  note         = {Machine review of arXiv:2502.00060}
}
read the original abstract

The Israeli-Palestinian conflict started on 7 October 2023, have resulted thus far to over 48,000 people killed including more than 17,000 children with a majority from Gaza, more than 30,000 people injured, over 10,000 missing, and over 1 million people displaced, fleeing conflict zones. The infrastructure damage includes the 87\% of housing units, 80\% of public buildings and 60\% of cropland 17 out of 36 hospitals, 68\% of road networks and 87\% of school buildings damaged. This conflict has as well launched an online discussion across various social media platforms. Telegram was no exception due to its encrypted communication and highly involved audience. The current study will cover an analysis of the related discussion in relation to different participants of the conflict and sentiment represented in those discussion. To this end, we prepared a dataset of 125K messages shared on channels in Telegram spanning from 23 October 2025 until today. Additionally, we apply the same analysis in two publicly available datasets from Twitter containing 2001 tweets and from Reddit containing 2M opinions. We apply a volume analysis across the three datasets, entity extraction and then proceed to BERT topic analysis in order to extract common themes or topics. Next, we apply sentiment analysis to analyze the emotional tone of the discussions. Our findings hint at polarized narratives as the hallmark of how political factions and outsiders mold public opinion. We also analyze the sentiment-topic prevalence relationship, detailing the trends that may show manipulation and attempts of propaganda by the involved parties. This will give a better understanding of the online discourse on the Israel-Palestine conflict and contribute to the knowledge on the dynamics of social media communication during geopolitical crises.

Figures

Figures reproduced from arXiv: 2502.00060 by the authors.

Figure 1
Figure 1. Comparison of message volumes across platforms Additionally we show the barplot of the distribution of messages per channel in figure 2 [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Barplot showing the distribution of messages per channel on Telegram February 4, 2025 7/23 [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Barplot showing the distribution of messages per channel on Reddit [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (17 more)
Figure 4
Figure 4. Figure 4: Messages per subreddit [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Barplot showing the distribution of tweets per dataset file February 4, 2025 8/23 [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: A 2x2 grid of subfigures showing various analyses. 5.2 Entity extraction The first step of our analysis was to extract the main thematic entities the online discussion were centered on. To achieve this, we process the messages from our dataset and extract the most comm…
Figure 7
Figure 7. Figure 7: Word clouds for each topic in the analysis on Twitter dataset 5.3 Topic Analysis 5.3.1 LDA Towards the analysis of the topics discusssed in the corpus we apply state of the art methods for topic analysis including Latent Dirichlet allocation (LDA) and BertTopic analysi…
Figure 8
Figure 8. Figure 8: First 4 topics(1-4)on Telegram from LDA topic analysis As shown in figure 8, in Topic 1 (19% of tokens) appears to focus on general discourse about Palestinian issues, with key terms including: ”palestinian”, ”resistance”, ”people” as primary terms Media-related terms …
Figure 9
Figure 9. Figure 9: The next 4 topics(4-8) on Telegram 5.3.2 Bert Topic Model However conventional methods like LDA and NMF, possess some drawbacks due to their manual process of specifying topics in advance and other lacking capabilities. By contrast, modern transformers have revolutioni…
Figure 10
Figure 10. Figure 10: A network of connections between the top ten most important topics on Telegram February 4, 2025 14/23 [PITH_FULL_IMAGE:figures/full_fig_p014_10.png]
Figure 11
Figure 11. Figure 11: Network of 309 topics on Telegram [PITH_FULL_IMAGE:figures/full_fig_p015_11.png]
Figure 12
Figure 12. Figure 12: Distribution of entities by applying Bert topic analysis, on the number of messages February 4, 2025 15/23 [PITH_FULL_IMAGE:figures/full_fig_p015_12.png]
Figure 13
Figure 13. Figure 13: Network of topics for Twitter dataset, contaning 309 topics February 4, 2025 16/23 [PITH_FULL_IMAGE:figures/full_fig_p016_13.png]
Figure 14
Figure 14. Figure 14: The network of topics for Twitter [PITH_FULL_IMAGE:figures/full_fig_p017_14.png]
Figure 15
Figure 15. Figure 15: The heatmap network of Twitter topics February 4, 2025 17/23 [PITH_FULL_IMAGE:figures/full_fig_p017_15.png]
Figure 16
Figure 16. Figure 16: Granular Sentiment across Various Channels 6.2 Telegram The final step is to perform granular sentiment analysis in order to reveal specific emotions like hope, support, anger, and despair. Towards this goal we incorporate a model that can classify texts into more det…
Figure 17
Figure 17. Figure 17: Granular Sentiment across Various Channels We apply a pre-trained emotion analysis model from Hugging Face’s transformers library. The dominant emotion for each non empty, English, message is identified to finally generate specific emotion analysis as shown in bar plo…
Figure 18
Figure 18. Figure 18: Sentiment for Reddit 7 Conclusion The discussion of the war between Israel and Hamas in the current study provides the first comprehensive insight, while additional comparison was also drawn from Twitter and Reddit. An analysis of 105,000 messages on Telegram, 2,001 t…
Figure 19
Figure 19. Figure 19: Overall Caption for All Topics February 4, 2025 22/23 [PITH_FULL_IMAGE:figures/full_fig_p022_19.png]
Figure 20
Figure 20. Figure 20: Overall Caption for All Topics February 4, 2025 23/23 [PITH_FULL_IMAGE:figures/full_fig_p023_20.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

13 extracted references · 12 canonical work pages

  1. [1]

    The Pushshift Telegram Dataset

    Baumgartner J, Zannettou S, Squire M, Blackburn J. The Pushshift Telegram Dataset. Proceedings of the International AAAI Conference on Web and Social Media. 2020;14(1):840–847. doi:10.1609/icwsm.v14i1.7348

  2. [2]

    Narratives online: Shared stories in social media

    Page R. Narratives online: Shared stories in social media. Cambridge University Press; 2018

  3. [3]

    middleeasteye. Israel-Palestine war: How an Israeli Telegram channel is used to in- cite violence against Palestinians; 2020.https://www.middleeasteye.net/news/ israel-palestine-war-telegram-incite-violence-psychological-warfare-palestinians

  4. [4]

    How Telegram disruption impacts jihadist platform migration

    Amarasingam A, Maher S, Winter C. How Telegram disruption impacts jihadist platform migration. Centre for Research and Evidence on Security Threats. 2021; p. 27

  5. [5]

    Analysing the Impact of Telegram (Social Media) on the Political Public Sphere: A Study on the Political Public Sphere and Deliberative Democracy

    Jongbloed A. Analysing the Impact of Telegram (Social Media) on the Political Public Sphere: A Study on the Political Public Sphere and Deliberative Democracy. unknown; 2024

  6. [6]

    IS and the Jihadist information highway–Projecting influence and religious identity via Telegram

    Prucha N. IS and the Jihadist information highway–Projecting influence and religious identity via Telegram. Perspectives on Terrorism. 2016;10(6):48–58

  7. [7]

    Quantifying Extreme Opinions on Reddit Amidst the 2023 Israeli-Palestinian Conflict

    Guerra A, Lepre M, Karakus O. Quantifying Extreme Opinions on Reddit Amidst the 2023 Israeli-Palestinian Conflict. arXiv preprint arXiv:241210913. 2024

  8. [8]

    Tweeting# Palestine: Twitter and the mediation of Palestine

    Siapera E. Tweeting# Palestine: Twitter and the mediation of Palestine. International Journal of Cultural Studies. 2014;17(6):539–555

Show all 13 references
  1. [9]

    # GazaUnderAttack: Twitter, Palestine and diffused war

    Siapera E, Hunt G, Lynn T. # GazaUnderAttack: Twitter, Palestine and diffused war. Information, Communication & Society. 2015;18(11):1297–1319

  2. [10]

    The self and other: portraying Israeli and Palestinian identities on Twitter

    Deegan J, Hogan J, Feeney S, O’Rourke BK. The self and other: portraying Israeli and Palestinian identities on Twitter. Irish communication review. 2018;16(1):8

  3. [11]

    Sentiment Analysis of Israel-Palestine Conflict Comments Using Sentiment Intensity Analyzer and TextBlob

    Sharkar ME, Hosen MJ, Abdullah M, Islam S, Rana S, Sultana N. Sentiment Analysis of Israel-Palestine Conflict Comments Using Sentiment Intensity Analyzer and TextBlob. In: 2024 15th International Conference on Computing Communication and Networking Technologies (ICCCNT). IEEE;...

  4. [12]

    GitHub-rizqikapratamaa; 2024

    GitHub. GitHub-rizqikapratamaa; 2024. https://github.com/rizqikapratamaa/ Sentiment-Analysis-on-Twitter-Tweets-about-the-Israel-Palestine-Conflict

  5. [13]

    Daily Public Opinion on Israel-Palestine War; 2024

    Asaniczka. Daily Public Opinion on Israel-Palestine War; 2024. Available from: https://www.kaggle.com/dsv/9832871. February 4, 2025 21/23 8 Appendix 8.1 Additional volume plots (a) Cumulative distribution function of messages in ’GAZA NOW IN ENGLISH’ channel (b) Cumulative dis...

Pith tools

Reviewed August 10, 2026 · model on record in the stance chip above.