{"id":"13a52776-4c17-4699-9f68-280055d743b6","arxiv_id":"2504.14037","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Arabic online conspiracy discourse clusters into six narrative categories (gender/feminist, geopolitical, government cover-ups, apocalyptic, Judeo-Masonic, geoengineering) according to NER and Top2Vec analysis of 1,641 curated texts.","lead":"The paper uses topic modeling and named entity recognition to map conspiracy narratives in Arabic Facebook posts and blogs. It identifies six recurring categories, including geopolitical, apocalyptic, and Judeo-Masonic themes, drawn from a hand-curated corpus of 1,641 texts.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Dataset representativeness is the load-bearing risk: the six categories may be artifacts of a conspiracy-preselected convenience sample from one blog and unnamed Facebook pages.","rationale":"The reader identified the representativeness of ArCons as the weakest assumption, and I agree. The central claim is explicitly general: 'Arabic online conspiracism revolves around six major conspiracy theories.' Yet the dataset is a convenience sample, assembled because documents already related to conspiracy theories, from one blog and unnamed Facebook pages. The authors do acknowledge in Section 3.3 that the data does not fully depict Arabic conspiracy theorists' web interests, which is an honest limitation but also directly contradicts the strength of the Section 4.2 claim. The topic-modeling step cannot repair this, because the input distribution determines the topics: a corpus preselected for conspiracist content will naturally yield conspiratorial themes, and the post-hoc reduction from 11 to 9 to 6 categories introduces additional subjective freedom. No internal contradiction in the pipeline was found, and the methods are standard for an exploratory study, so rejection is not warranted. Conditional acceptance remains appropriate, provided the authors either narrow the claim to 'within the curated corpus' or validate against a broader, independently sampled corpus.","tokens_in":18213,"tokens_out":2568,"duration_ms":24568,"concrete_test":"Build an independent validation corpus without conspiracy preselection: collect a stratified random sample of Arabic political/social posts from multiple platforms (Twitter/X, Telegram, YouTube, Facebook public pages) over 2010-2023 using neutral query terms (e.g., politics, economy, religion, science) or platform-wide random sampling, then apply the identical preprocessing and Top2Vec pipeline. Compare emergent topics to the six categories. If the same six categories recur at comparable prevalence, the claim survives; if new major categories appear or the six fail to reproduce, the original finding is a sampling artifact. A cheaper first step: re-run the analysis on the 392 blog documents alone versus the 1,249 Facebook documents alone; if the two subsamples yield different topic structures, the unified six-category result is not stable across sources.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that Arabic online conspiracism 'revolves around six major conspiracy theories' (Section 4.2)—requires that ArCons is a representative sample of that discourse. Section 3.1 shows otherwise: 1,641 documents were curated 'centered around conspiracy theories' from one blog (24%) and unspecified Facebook pages/groups, with no description of page selection, search queries, or inclusion protocol. The curation step explicitly excluded non-conspiratorial documents, so the corpus cannot establish which narratives dominate the broader Arabic online landscape; it can only describe themes within the curators' chosen sample. The paper itself concedes in Section 3.3 that 'our data does not fully depict Arabic conspiracy theorists web interests.' Because the six categories are derived from topic labels assigned post hoc by the authors (11 topics reduced to 9 then grouped into 6, with no inter-annotator or external validation), the categories may reflect the source material's editorial slant or the annotators' expectations rather than empirical structure. This is a selection-bias problem: the sampling frame constrains the possible topics, so the general claim is not supported by the evidence as presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes a curated corpus of 1,641 Arabic texts from Facebook and one blog (the ArCons dataset) to identify conspiratorial narratives. Using the Marifa NER model and Top2Vec topic modeling, with manual interpretation of topics, the authors claim that Arabic online conspiracism revolves around six major categories: gender/feminist, geopolitical, government cover-ups, apocalyptic, Judeo-Masonic, and geoengineering conspiracies. They also describe temporal shifts in frequent terms (2010-2016 vs 2017-2023) and list prominent named entities. The paper situates these findings in the cultural and historical context of the Arab region and discusses theoretical, practical, and policy implications.","tokens_in":18450,"tokens_out":3177,"duration_ms":30480,"significance":"If the central descriptive claim were adequately supported, this paper would fill a genuine gap: most computational conspiracy-theory research targets English-language or Western contexts, while Arabic online content remains understudied. The paper contributes a new Arabic dataset, applies Arabic-specific NLP tools, and proposes a taxonomy that aligns with established qualitative scholarship on Middle Eastern conspiracism. The combination of NER and topic modeling is reasonable, and the cultural contextualization is a strength. However, the empirical evidence is currently weakened by sampling and validation limitations. The paper would be a useful exploratory starting point if the claims are reframed to the corpus and the category induction is made more transparent and reproducible.","major_comments":[{"comment":"The central claim that 'Arabic online conspiracism revolves around six major conspiracy theories' (Section 4.2) is not supported by the sampling design. The corpus is a curated convenience sample: 1,641 documents selected from one blog (24%) and unspecified Facebook pages, explicitly chosen because they center on conspiracy theories. No search queries, page selection criteria, or inclusion protocol are documented, and Section 3.3 itself concedes that 'our data does not fully depict Arabic conspiracy theorists web interests.' Because the sampling frame constrains the topics that can appear, the six categories cannot be generalized to the broader Arabic online discourse; they are at best descriptive of this particular collected corpus. The authors should either build a more systematically sampled corpus (e.g., random posts from defined public pages with documented keyword discovery) or explicitly restrict all conclusions to the ArCons sample.","section":"3.1–3.2 and 4.2"},{"comment":"The reduction from 11 Top2Vec topics to 9 and then to 6 narrative categories is performed manually with no quantitative validation. No topic-coherence metrics (e.g., NPMI, topic coherence), no inter-annotator agreement, and no external benchmark are reported. The mapping of topics to labels such as 'gender/feminist' or 'geoengineering' appears reasonable but may reflect analyst expectations rather than structure intrinsic to the data. To make the central taxonomy credible, the authors should provide a coding protocol, have at least two annotators independently assign topics to categories and report agreement (e.g., Cohen's kappa), and show representative top documents or document-topic distributions for each category.","section":"3.5 and 4.1"},{"comment":"The temporal analysis is descriptive only and does not substantiate the claim of a thematic evolution. The split at 2016/2017 is arbitrary, and the comparison of most frequent terms is presented without statistical testing or topic-proportion analysis over time. The observation that the second period is more 'introspective' and 'contemplative' is based on raw term lists and could be an artifact of corpus composition or preprocessing. The authors should compute topic proportions per year or per period and test for significance, or soften the temporal claim to an explicitly qualitative observation.","section":"3.3"}],"minor_comments":[{"comment":"The title of Section 3.5 contains a typo: 'conspirasionist' should be 'conspiracist.'","section":"3.5"},{"comment":"The text refers to 'Table 6' when discussing Topic 07, but the paper only presents a single topic table (Table 4); the reference should be corrected.","section":"4.1.6"},{"comment":"The list of NER categories says 'five distinct categories' but then enumerates seven (locations, nationalities, persons, jobs, organizations, events, and products). Please correct the count or the list.","section":"3.4"},{"comment":"No evaluation or error analysis is provided for the Marifa NER model on this specific corpus, which includes dialectal Arabic. Reporting a sample of manually checked entity extractions would help assess reliability.","section":"3.4"},{"comment":"No data availability statement or link to the ArCons dataset is provided; for reproducibility, the authors should state whether and how the dataset can be accessed.","section":"3.1"},{"comment":"The paper does not discuss ethical considerations or platform terms of service for collecting and publishing Facebook posts; a brief ethical statement would be appropriate.","section":"2.1 and 3.1"},{"comment":"The Arabic text in Table 2 appears garbled or reversed in several places (e.g., 'ىربك' and 'ﺭالودلﺍ'), making the table difficult to read; the Arabic strings should be proofread.","section":"Table 2"},{"comment":"The keyword field contains 'Digital Environement,' which should be 'Digital Environment.'","section":"Introduction and Keywords"}],"recommendation":"major_revision","confidential_remarks":"The paper's exploratory contribution is plausible and addresses a real gap, but the load-bearing issues are (i) the inability to generalize from the curated sample to 'Arabic online conspiracism' and (ii) the lack of any validation of the manual topic-to-category mapping. These are fixable within a revision if the authors reframe the claims to the corpus and add transparent coding/validation steps. The NER component is largely illustrative, so the review has focused on the topic-modeling pipeline and the sampling frame."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a competent exploratory study of Arabic conspiracy content, but the title and the central claim outrun the data. The six-category taxonomy is a reasonable reading of what the authors curated, not a proven map of Arabic online conspiracism.\n\nWhat's genuinely useful: it applies a modern topic-modeling pipeline (Top2Vec with NER) to Arabic-language conspiracy content, which is genuinely underserved in the literature. The literature review is broad, and the discussion ties each category to regional history and existing conspiracy scholarship. The temporal word-frequency comparison is simple but honest. The authors also state their own caveat in Section 3.3 that the data does not fully depict Arabic conspiracy theorists' web interests, and they list limitations.\n\nSoft spots, in order of severity. First, dataset representativeness. The 1,641 documents were selected 'centered around conspiracy theories' from one blog and unnamed Facebook pages. There is no description of page selection, search queries, or inclusion rules. That means the six categories and the temporal shift could be artifacts of source choice. The paper's own limitation statement basically concedes this. Second, the category construction is entirely manual. Eleven topics are merged into nine, then six, with no inter-annotator reliability, no topic coherence metrics, and no external validation. The labels may be plausible but they are not independently checked. Third, there is no released data or code, so the results cannot be reproduced. Fourth, a minor point: the choice of Top2Vec over LDA/NMF is asserted without quantitative comparison, and the temporal split at 2016 is arbitrary.\n\nNone of these sink the paper; they just limit what it can claim. The right fix is to reframe the conclusion as 'six categories within our curated corpus,' add a validation section (even a small human-labeled set with agreement scores), and release the data and preprocessing scripts. As it stands, it is a valuable starting point for cross-cultural conspiracy research, but not a general map of Arabic online conspiracism.\n\nI would give it a serious referee. The topic is important, the application is novel, and the flaws are addressable. I would recommend major revision rather than desk rejection. For a reading group, it is a good example of how sampling choices interact with topic-modeling claims.","headline":"Useful exploratory map of conspiratorial themes in a curated Arabic corpus, but the headline generalization about 'Arabic online conspiracism' overreaches what the sampling supports.","tokens_in":18936,"tokens_out":2488,"would_cite":true,"duration_ms":21205,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Arabic online conspiratorial discourse organizes into six narrative categories.","keywords":["Conspiracy theories","Misinformation","Arabic Natural Language Processing","Topic Modeling","Named Entity Recognition","Social networks","Digital environment","Arabic online content"],"falsifier":"Collect an independent, pre-registered sample of Arabic social-media posts using broad conspiratorial and neutral keywords across multiple countries and platforms, apply the same preprocessing, NER, and Top2Vec pipeline, and check whether the same six categories recur; if gender and feminist or geoengineering themes do not appear in the independent sample, the ArCons-based taxonomy is an artifact of corpus selection.","tokens_in":18056,"feed_emoji":"🕵️","tokens_out":5959,"duration_ms":53070,"temperature":0.7,"pith_summary":"This paper claims that Arabic-language conspiracy theorizing on blogs and Facebook is structured rather than amorphous, clustering into six recurring narrative categories: gender and feminist plots, geopolitical schemes, government cover-ups, apocalypticism, Judeo-Masonic plots, and geoengineering theories such as HAARP and earthquake manipulation. The claim matters because most computational research on online conspiracy theories has focused on English-language, Western settings, leaving Arabic-speaking digital environments largely unmapped. Using a hand-curated corpus of 1,641 Arabic texts spanning 2010 to 2023, the paper identifies the categories with topic modeling and named-entity extraction and argues that they are shaped by regional history, culture, and contemporary events. If the taxonomy holds, it gives media monitors, fact-checkers, and platform moderators a concrete checklist of conspiratorial narratives to look for in Arabic content.","feed_headline":"Arabic online conspiracy talk splits into six narratives","feed_subtitle":"A 1,641-text Arabic corpus maps the six recurring themes, from feminist plots to HAARP-induced earthquakes.","key_machinery":"The mechanism carrying the argument is a two-stage computational pipeline applied to a corpus the paper calls ArCons: 1,641 Arabic documents collected from one blog and Facebook pages between 2010 and 2023. First, a pre-trained Arabic named-entity recognizer (Marifa NER) extracts persons, locations, organizations, events, jobs, and products, sketching the actors and incidents in the discourse. Second, Top2Vec, a neural topic-modeling algorithm that embeds documents and words in a shared vector space and clusters them into topics, produces eleven topics that the authors consolidate into six narrative categories. NER supplies the actors and events; Top2Vec supplies the thematic structure; together they convert raw Arabic text into the paper's taxonomy of conspiratorial narratives.","core_discovery":"The central claim is that Arabic online conspiracism revolves around six major conspiracy theories: gender and feminist conspiracies, geopolitical conspiracies, government cover-ups, apocalypticism, Judeo-Masonic conspiracies, and geoengineering conspiracies. The paper derives this taxonomy from eleven Top2Vec topics that it consolidates into nine, then six, and corroborates it with named entities extracted from the same corpus, including figures like Hitler, Rothschild, and Putin, organizations like NATO, the Freemasons, and NASA, and events like the Arab Spring, the 2023 Turkey-Syria earthquake, and the HAARP project. It also claims a temporal shift: from 2010 to 2016 the discourse centered on geopolitics and international actors, whereas from 2017 to 2023 it broadened toward society, gender roles, religion, pseudo-science, and cosmic themes. These patterns, the paper argues, are culturally rooted in the region's colonial history, U.S. involvement, and political instability, while also responding to immediate triggers such as disasters and global media events.","pith_inferences":["The six categories are plausibly transferable as labels for training a supervised Arabic conspiracy-detection model, a step the paper does not itself take.","Because the corpus skews toward Algerian and Maghrebi sources and a single blog, the taxonomy likely under-represents Gulf and Levant narrative variants; a balanced multi-country corpus would test this.","The prominence of gender and feminist conspiracy themes may reflect platform visibility and algorithmic amplification of Western gender-politics controversies after the Arab Spring, rather than the most widely believed conspiracy theory among offline Arab publics.","Applying the same pipeline to newer crises, such as the post-2023 Gaza war or the Sudan conflict, would provide a live test of whether the six categories are exhaustive or whether event-specific categories emerge."],"forward_implications":["Moderators and fact-checkers can treat the six categories as an initial codebook for tagging Arabic conspiracy content.","The temporal shift implies that current Arabic conspiracist discourse leans less on regional geopolitics and more on society, gender, pseudo-science, and cosmic or religious themes.","The named-entity watchlists identify concrete anchors, such as Rothschild, NATO, the Freemasons, HAARP, and the Turkey earthquake, around which Arabic conspiracy narratives crystallize.","Event-driven spikes, particularly following the 2023 Turkey-Syria earthquake, suggest that Arabic conspiracy monitoring should activate around disasters and science-related news, not only political events.","The combined NER and topic-modeling pipeline can be reused on new Arabic corpora to detect whether the six categories persist or new ones emerge."],"supporting_citations":[{"why":"Top2Vec is the topic-modeling algorithm whose output produces the eleven topics later consolidated into the paper's six categories.","marker":"[39]"},{"why":"The review documents the English-language and Western bias in online conspiracy research, which is the gap the paper positions itself against.","marker":"[12]"},{"why":"Gray's analysis supplies the cultural and historical interpretation of Middle Eastern conspiracism that the paper uses to explain the geopolitical categories.","marker":"[19]"},{"why":"The MENA survey provides evidence of widespread conspiracy belief and its association with anti-Western and anti-Jewish attitudes.","marker":"[21]"},{"why":"This study of Egyptian tweets shows conspiratorial thinking functioning as a coping mechanism in Arabic online discourse, supporting the paper's framing.","marker":"[30]"},{"why":"Salvati and colleagues define gender-ideology conspiracy beliefs, grounding the paper's first category.","marker":"[34]"},{"why":"The Al Jazeera report documents conspiracy theories that followed the 2023 Turkey-Syria earthquake, evidence for the geoengineering category.","marker":"[75]"}],"fun_headline_variants":["Six conspiratorial themes dominate Arabic online discourse","Arabic social media's six conspiracy categories revealed","Top2Vec maps six conspiracy narratives in Arabic online chatter","From feminist plots to HAARP: six Arabic conspiracy themes","Arabic conspiracy taxonomy: six narratives from 1,641 texts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the 1,641 documents the authors selected because they already related to conspiracy theories, drawn from one blog and unspecified Facebook pages, are representative enough of Arabic online conspiracism that the six categories and the temporal shift reflect the broader landscape rather than the quirks of that particular selection.","fun_headline_variants_meta":{"raw":{"variants":["Six conspiratorial themes dominate Arabic online discourse","Arabic social media's six conspiracy categories revealed","Top2Vec maps six conspiracy narratives in Arabic online chatter","From feminist plots to HAARP: six Arabic conspiracy themes","Arabic conspiracy taxonomy: six narratives from 1,641 texts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00092,"raw_usage":{"total_tokens":3929,"prompt_tokens":908,"completion_tokens":3021,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":524,"completion_tokens_details":{"reasoning_tokens":2944}},"tokens_in":524,"tokens_out":3021,"duration_ms":17699,"temperature":1.0,"reasoning_tokens":2944,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:57:05.142810+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect an independent, pre-registered sample of Arabic social-media posts using broad conspiratorial and neutral keywords across multiple countries and platforms, apply the same preprocessing, NER, and Top2Vec pipeline, and check whether the same six categories recur; if gender and feminist or geoengineering themes do not appear in the independent sample, the ArCons-based taxonomy is an artifact of corpus selection.","supporting_citations":[{"cited_title":"Top2vec: Distributed representations of topics","cited_arxiv_id":null,"evidence_quote":"Top2Vec is the topic-modeling algorithm whose output produces the eleven topics later consolidated into the paper's six categories."},{"cited_title":"Schäfer, and Jing Zeng","cited_arxiv_id":null,"evidence_quote":"The review documents the English-language and Western bias in online conspiracy research, which is the gap the paper positions itself against."},{"cited_title":"Explaining conspiracy theories in modern arab middle eastern political discourse: Some problems and limitations of the literature","cited_arxiv_id":null,"evidence_quote":"Gray's analysis supplies the cultural and historical interpretation of Middle Eastern conspiracism that the paper uses to explain the geopolitical categories."},{"cited_title":"Conspiracy and misperception belief in the middle east and north africa","cited_arxiv_id":null,"evidence_quote":"The MENA survey provides evidence of widespread conspiracy belief and its association with anti-Western and anti-Jewish attitudes."},{"cited_title":"Essam, Mostafa M","cited_arxiv_id":null,"evidence_quote":"This study of Egyptian tweets shows conspiratorial thinking functioning as a coping mechanism in Arabic online discourse, supporting the paper's framing."},{"cited_title":"What is hiding behind the rainbow plot? the gender ideology and LGBTQ + lobby conspiracies","cited_arxiv_id":null,"evidence_quote":"Salvati and colleagues define gender-ideology conspiracy beliefs, grounding the paper's first category."},{"cited_title":"A mysterious american weapon or a russian submarine? | aljazeera.net","cited_arxiv_id":null,"evidence_quote":"The Al Jazeera report documents conspiracy theories that followed the 2023 Turkey-Syria earthquake, evidence for the geoengineering category."}],"review_version":1}