{"id":"76f62ff0-7872-4df8-bf67-291218d1134c","arxiv_id":"2502.08319","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"MultiProSE adds manual sentiment and emotion labels to the 8,000-paragraph Arabic ArPro propaganda corpus and reports BERT and GPT-4o-mini baselines.","lead":"MultiProSE adds manual sentiment and emotion labels to the existing 8,000-paragraph Arabic ArPro propaganda corpus, creating a new multi-label benchmark. A general reader should look at this paper to see whether Arabic propaganda research now has a resource for studying how sentiment, emotion, and propagandistic content interact.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Annotation guidelines may impose a deterministic sentiment-emotion mapping, making the claimed interaction analysis circular.","rationale":"I agree with the reader's identification: the weakest load-bearing assumption is that sentiment and emotion are annotated independently. The paper's own Section 3.3 undermines this by stating that the known correlation was 'incorporated' into the guidelines, which would guarantee a near-deterministic mapping. Unless the actual annotations contradict the stated mapping, the dataset cannot support the claimed contribution of analyzing interactions between opinion dimensions. This concern is more fundamental than the placeholder URL or the split inconsistencies, which are administrative and correctable. The proposed contingency-table check would settle whether the mapping is actually present in the released data. If the mapping is enforced, the paper should either re-annotate the texts or substantially weaken the interaction claims; conditional acceptance with a required verification is therefore appropriate.","tokens_in":11259,"tokens_out":6194,"duration_ms":60002,"concrete_test":"Once the data are released, compute the 3×5 contingency table of sentiment labels (positive/negative/neutral) against emotion labels (happiness/sadness/anger/fear/none) across all 8,000 texts. If the positive-sentiment row is almost entirely happiness (e.g., >95%) and the negative-sentiment row is almost entirely anger+sadness+fear (e.g., >95%), while cross-combinations like positive+sadness or negative+happiness are absent or negligible, then the annotation guidelines enforced the mapping and the interaction analysis in Figure 3 is circular. If a meaningful fraction (say >5%) of texts exhibit cross-combinations, the independence claim is supported. Supplementary check: inspect the released annotation guidelines (Appendix A) for language that instructs annotators to choose emotion labels based on sentiment polarity.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central value proposition is that MultiProSE enables studying how opinion dimensions interact (Abstract, Section 1, Section 4.3). This requires sentiment and emotion to be annotated as independent constructs. However, Section 3.3 states: 'a text annotated as positive will have happiness as the emotion label, while a text with negative sentiment will be annotated with anger, sadness, or fear as emotion labels. Therefore, in the annotation guidelines, these details have been considered and incorporated.' If the guidelines instruct annotators to apply this mapping, then the sentiment and emotion labels are deterministically coupled by the schema, not discovered from the text. Consequently, the correlations in Figure 3 and any downstream claim about the interaction between propaganda, sentiment, and emotion are at least partly artifacts of annotation design. The sentence claiming this 'ensure[s] that sentiment is annotated independently' is internally contradictory: incorporating a fixed mapping cannot produce independence. This is the most load-bearing issue because it attacks the dataset's scientific purpose, not merely its metadata or availability.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MultiProSE, an Arabic corpus of 8,000 news paragraphs that extends the existing ArPro propaganda dataset with manually annotated sentiment and emotion labels. It describes the annotation protocol, a gold-data quality-control mechanism, inter-annotator agreement results, label distributions, and baseline experiments with AraBERT, XLM-RoBERTa, and GPT-4o-mini. The stated contributions are a new Arabic benchmark for propaganda detection, sentiment analysis, and emotion recognition, together with an analysis of how these opinion dimensions interact in news text.","tokens_in":11419,"tokens_out":6250,"duration_ms":65844,"significance":"If the dataset is actually released and the annotation caveats are resolved, MultiProSE would be a useful Arabic multi-task benchmark. The design includes three paid native-speaker annotators with doctoral degrees, a pre-exam and gold-data phase, and reported inter-annotator agreement in the substantial range (Light's kappa 0.7074-0.8128; Fleiss' kappa 0.7093-0.7650), which are genuine strengths. The baseline results from multiple-seed runs also provide useful reference points. However, the manuscript currently overstates the dataset size, presents inconsistent train/test splits, uses a placeholder repository link, and, most importantly, describes an annotation guideline that deterministically couples sentiment and emotion labels while claiming these dimensions are annotated independently. These issues must be corrected before the benchmark and interaction-analysis claims can be accepted.","major_comments":[{"comment":"Section 3.3 states that a positive text will have happiness as its emotion label, that a negative text will be labeled anger, sadness, or fear, and that these details were 'incorporated' into the guidelines; the same section claims this 'ensure[s] that sentiment is annotated independently' of emotion. These statements are contradictory. If annotators followed the guidelines, sentiment and emotion are coupled by the annotation schema, so the co-occurrence patterns in Figure 3 and any sentiment-emotion interaction analysis are partly products of the design rather than discoveries about the text. Please re-annotate with independent dimensions, or substantially re-frame the claims and report the analyses conditional on the enforced mapping.","section":"3.3 / Figure 3"},{"comment":"The abstract's statement that 8,000 articles make MultiProSE 'the largest propaganda dataset to date' is contradicted by Table 1, which lists TSHP-17 with 22,580 articles and QProp with 51,294 articles. If the intended claim is 'largest Arabic manually annotated propaganda dataset,' it should be stated precisely and supported; otherwise the global claim should be removed.","section":"Abstract / Table 1"},{"comment":"The dataset split is reported inconsistently: Section 3.1 gives 6,002/672/1,326 for train/validation/test; Section 3.7 gives 6,680/1,320 train/test with no validation set; Section 4.1 says 75%/8.5%/16.5%. These numbers cannot all describe the released data, and the discrepancy directly affects the comparability of the baseline results. Specify the exact released split and use it consistently throughout.","section":"3.1 / 3.7 / 4.1"},{"comment":"The availability statement in the Abstract is not currently verifiable: footnote 1 reads 'https://github.com/xxx/xxx', which is a placeholder. For a dataset paper the repository URL, the dataset, and the annotation guidelines must be accessible; please provide the actual link and confirm the license.","section":"Footnote 1"},{"comment":"The term 'multi-label' is used for the dataset, but each text receives a single propaganda label, a single sentiment label, and a single emotion label. This is multi-task, not multi-label in the standard sense. If the intended meaning is that each text carries several label dimensions, please state this explicitly; if some tasks are multi-label (e.g., multiple propaganda techniques per text), then the annotation and evaluation sections need to reflect that.","section":"Title / 3.7"}],"minor_comments":[{"comment":"The majority-voting description is unclear: if three annotators disagree, the text says a sixth annotator may be added, but with three annotators majority voting already produces a decision; please specify the conflict-resolution rule and when and how additional annotators are consulted.","section":"3.2"},{"comment":"The text refers to both 'Light's Kappa' and 'Lights' index'; please standardize the name of the measure to Light's kappa throughout.","section":"3.6"},{"comment":"The model ArabicBERT is mentioned in the results discussion but is not introduced in Section 4.1, which describes only AraBERT and XLM-RoBERTa; please define it or remove the reference.","section":"4.2"},{"comment":"The citation to MELD [38] is misleading here because MELD does not enforce a deterministic sentiment-emotion mapping; please make the relationship precise or use a different justification.","section":"3.3"},{"comment":"The sentence about 12 batches and 24,000 collected annotations follows from 8,000 texts times three annotations, but the relation to the reported rounds and batches should be stated explicitly to avoid confusion.","section":"3.5"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a resource/benchmark venue, and the annotation effort is credible. The main editorial concern beyond the technical revisions is that the 'first Arabic' and 'multi-label' claims need careful calibration relative to ArPro, ArAIEval, and the actual annotation schema; I would not reject on novelty grounds, but the final version should make the incremental contribution precise and ensure the repository and data are genuinely public."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First, what you should know: this is a real annotation effort filling a real gap—Arabic propaganda detection with sentiment and emotion labels—but the paper's central claim about studying how opinion dimensions interact is undercut by its own annotation guidelines, and the resource is not yet verifiable because the release link is a placeholder and the dataset size claim is contradicted by the paper's own table.\n\nThe good: three PhD-level native annotators, a quality-control pipeline with gold data and trust scores, and inter-annotator agreement in the substantial range (Light's Kappa 0.81 for sentiment, 0.71 for emotion; Fleiss' Kappa 0.77 and 0.71). That is honest work. The extension of ArPro with sentiment and emotion labels is new, and the baselines (AraBERT, XLM-R, GPT-4o-mini) are standard and plausible. For someone who simply wants a combined Arabic propaganda-sentiment-emotion benchmark, this is a usable starting point.\n\nThe soft spots, in order. Most serious: Section 3.3 encodes a deterministic mapping—positive sentiment gets 'happiness,' negative sentiment gets 'anger/sadness/fear'—and calls this 'independent' annotation. That is contradictory. If annotators follow the mapping, the sentiment-emotion correlations in Figure 3 are schema artifacts, not discovered relationships. Propaganda-sentiment and propaganda-emotion correlations are less affected, since propaganda comes from ArPro independently, but the paper's stated purpose of analyzing how opinion dimensions interact is directly compromised.\n\nSecond, the 'largest propaganda dataset' claim is false; Table 1 lists TSHP-17 with 22,580 articles and QProp with 51,294. Minor, but embarrassing and easy to fix.\n\nThird, the split statistics are inconsistent: Section 3.1 gives 6002/672/1326, Section 3.7 says 6680/1320 with no validation, and Section 4.1 describes a 75/8.5/16.5 split. These need reconciling.\n\nFourth, the GitHub link is a literal placeholder. A dataset paper without accessible data and code cannot be evaluated as a resource.\n\nWho this is for: Arabic NLP researchers who want a combined propaganda/sentiment/emotion benchmark and can treat the labels with the schema's constraints in mind. The annotation effort deserves a serious referee, and the paper is repairable, but it needs major revisions, an honest discussion of the annotation mapping, and a real release before it can serve as a benchmark for interaction studies.","headline":"A genuine Arabic propaganda-sentiment-emotion annotation effort, but the guidelines impose the sentiment-emotion correlation it claims to discover, and the placeholder release makes it unverifiable.","tokens_in":11967,"tokens_out":3672,"would_cite":false,"duration_ms":33058,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MultiProSE is the first Arabic dataset to label the same 8,000 news articles for propaganda, sentiment, and emotion, and the paper reports baseline results for all three tasks.","keywords":["Arabic NLP","propaganda detection","sentiment analysis","emotion recognition","multi-label dataset","news corpus","Modern Standard Arabic","annotation quality"],"falsifier":"Inspect the released sentiment-by-emotion contingency table: if almost every positive text is labeled happiness and almost every negative text is labeled anger, sadness, or fear, the two dimensions are not independent and the reported propaganda-emotion links largely reflect the annotation rule. A direct check is to re-annotate a random sample with fresh annotators who receive no sentiment-emotion mapping and compare the resulting agreement and correlations.","tokens_in":11090,"feed_emoji":"📰","tokens_out":15299,"duration_ms":132135,"temperature":0.7,"pith_summary":"MultiProSE is the authors' answer to a gap: Arabic, despite being the fourth most-used internet language, had no corpus combining propaganda detection with sentiment and emotion labels on the same news texts. The dataset contains 8,000 Modern Standard Arabic news paragraphs, each annotated for propaganda (inherited from ArPro), sentiment (positive/negative/neutral), and emotion (happiness/sadness/anger/fear/none). The authors position it as the first Arabic dataset of this kind and report substantial inter-annotator agreement, with baseline results from two BERT-style models and GPT-4o-mini. If the resource holds up, it gives Arabic NLP a public benchmark for three tasks and a way to study how opinion dimensions interact in news.","feed_headline":"8,000 Arabic news paragraphs gain three annotation layers","feed_subtitle":"MultiProSE adds sentiment and emotion to every propaganda-labeled news text","key_machinery":"The central object is MultiProSE itself, an 8,000-text corpus in which every Modern Standard Arabic news paragraph carries three labels: propaganda (true/false, inherited from ArPro), sentiment (positive/negative/neutral), and emotion (happiness/sadness/anger/fear/none). The mechanism that makes it a usable benchmark is the annotation protocol: three paid native-speaker annotators with doctoral training, a gold-data quality-control phase with a 70% trust threshold, a qualifying exam, majority voting, and a consolidation phase that can add up to six annotators. The guidelines encode a mapping from sentiment to emotion, pairing positive with happiness and negative with anger, sadness, or fear, so the protocol both produces the labels and shapes how the dimensions relate to each other.","core_discovery":"The paper's central claim is that a manually annotated Arabic news corpus can carry propaganda, sentiment, and emotion labels on every text and still be annotated reliably. MultiProSE extends ArPro's 8,000 paragraphs by adding sentiment and emotion labels under a schema adapted from a six-basic-emotions model, dropping disgust and adding a 'none' category. The authors report averaged pairwise kappa values of 0.7074-0.8128 and multi-annotator kappa values of 0.7093-0.7650, which they read as substantial agreement. Baseline results show GPT-4o-mini reaching Micro-F1 scores of 0.842 for sentiment and 0.750 for emotion, while AraBERT and GPT-4o-mini both reach 0.769 for propaganda. The authors also report that propaganda is more frequent in negative and positive sentiment texts than in neutral ones, with anger and happiness the most common emotions among propagandistic paragraphs.","pith_inferences":["Editorial extension: the paper does not test its motivating claim that sentiment and emotion features improve propaganda detection; the natural next experiment is to feed sentiment and emotion predictions into the propaganda classifier and measure the gain.","Editorial extension: since the propaganda labels come from ArPro, MultiProSE's novelty is the added sentiment and emotion layers, and the 'largest propaganda dataset' claim should be read in that light rather than as a new collection of propaganda annotations.","Editorial extension: the sentiment-emotion mapping written into the guidelines means the interaction distributions in Figure 3 should be re-examined with fresh annotators who are not given that mapping, to see whether the correlations persist."],"forward_implications":["MultiProSE provides a public Arabic benchmark with fixed train/test splits for propaganda detection, sentiment analysis, and emotion recognition, allowing direct comparison of future models.","The baselines set reference numbers: GPT-4o-mini reaches Micro-F1 0.842 for sentiment and 0.750 for emotion, and AraBERT ties GPT-4o-mini at 0.769 for propaganda.","The corpus enables analysis of how opinion dimensions co-occur, with the reported data showing propaganda concentrated in negative and positive sentiment and most often paired with anger or happiness.","Because the dataset, guidelines, and code are released, researchers can extend the work to span-level annotation or Arabic sentiment and emotion lexicons, as the paper suggests for future work."],"supporting_citations":[{"why":"ArPro supplies the 8,000 Arabic news paragraphs and the propaganda labels that MultiProSE extends; all propaganda results depend on this source.","marker":"[26]"},{"why":"This dataset is the basis for the paper's rule that sentiment labels correlate with emotion labels, a rule written into the annotation guidelines.","marker":"[38]"},{"why":"This study is the cited evidence that persuasion techniques correlate with emotional salience, motivating the added sentiment and emotion layers.","marker":"[5]"},{"why":"This task's guidelines and 70% accuracy threshold are adapted for MultiProSE's quality control.","marker":"[35]"},{"why":"This reference supplies the kappa interpretation that lets the authors describe the agreement as substantial.","marker":"[40]"},{"why":"AraBERT is an Arabic pretrained baseline that ties for best propaganda-detection Micro-F1.","marker":"[47]"},{"why":"XLM-RoBERTa is the multilingual pretrained baseline whose lower scores support the paper's point that Arabic-specific models do better.","marker":"[48]"},{"why":"GPT-4o-mini is the LLM baseline that reaches the best sentiment and emotion Micro-F1 scores, establishing the benchmark numbers.","marker":"[45]"},{"why":"This work motivates dropping disgust and adding a 'none' label to the emotion scheme.","marker":"[37]"}],"fun_headline_variants":["First Arabic dataset labels every news text for propaganda, sentiment, and emotion","8,000 Arabic articles now carry propaganda, sentiment, and emotion labels","MultiProSE: first Arabic multi-label dataset for propaganda, sentiment, emotion","Arabic news corpus tagged for propaganda, sentiment, and emotion","Largest Arabic propaganda dataset now includes sentiment and emotion"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that sentiment and emotion can be annotated as independent dimensions even though the guidelines tell annotators to pair positive sentiment with happiness and negative sentiment with anger, sadness, or fear; if annotators follow that rule, the dataset's sentiment-emotion correlations are built in rather than discovered.","fun_headline_variants_meta":{"raw":{"variants":["First Arabic dataset labels every news text for propaganda, sentiment, and emotion","8,000 Arabic articles now carry propaganda, sentiment, and emotion labels","MultiProSE: first Arabic multi-label dataset for propaganda, sentiment, emotion","Arabic news corpus tagged for propaganda, sentiment, and emotion","Largest Arabic propaganda dataset now includes sentiment and emotion"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001131,"raw_usage":{"total_tokens":4694,"prompt_tokens":932,"completion_tokens":3762,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":3672}},"tokens_in":548,"tokens_out":3762,"duration_ms":27786,"temperature":1.0,"reasoning_tokens":3672,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T05:34:32.397120+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the released sentiment-by-emotion contingency table: if almost every positive text is labeled happiness and almost every negative text is labeled anger, sadness, or fear, the two dimensions are not independent and the reported propaganda-emotion links largely reflect the annotation rule. A direct check is to re-annotate a random sample with fresh annotators who receive no sentiment-emotion mapping and compare the resulting agreement and correlations.","supporting_citations":[],"review_version":1}