{"id":"77ddf367-c1fa-4877-a588-67cee638d19e","arxiv_id":"2505.11959","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"EmoHopeSpeech is a proposed Arabic-English emotion and hope speech corpus whose claimed reliability is contradicted by its own tables.","lead":"The paper introduces EmoHopeSpeech, a bilingual Arabic-English corpus annotated for emotions and hope speech. Its headline reliability and model scores conflict with the numbers in the body, so the resource needs verification before use.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's reliability figures (kappa 0.75–0.85, F1 0.69) directly contradict the body's reported kappa 0.31–0.56 and F1 ≤0.47; the central dataset-quality claim is unsupported without resolving this internal inconsistency.","rationale":"The reader's verdict of REJECT rests on the internal inconsistency between the abstract's validation metrics and the body's reported metrics. My independent reading confirms this is the most load-bearing concern: a dataset paper stands or falls on the credibility of its annotation quality and size claims. The abstract explicitly states kappa 0.75–0.85 and F1 0.69 as validation, while Table 12 reports kappa 0.31–0.56 and Table 11 reports F1 ≤0.47. Moreover, Section 3.1 and Table 9 agree that the English dataset has 4,036 annotated rows, not the 10,036 claimed in the abstract. These contradictions mean the paper's central claim—that EmoHopeSpeech is a large, reliable bilingual resource—is not supported by the submitted evidence. The concern is internal inconsistency, not disagreement with field consensus, so it is a correctness risk rather than a novelty dispute. The released dataset and code enable a concrete check, so the issue is resolvable rather than permanent. My recommendation is unchanged from the reader's: reject as submitted, but a corrected version with consistent numbers could be salvageable. I find no reason to move the verdict in the other direction.","tokens_in":13231,"tokens_out":3624,"duration_ms":31700,"concrete_test":"Download the released dataset from Zenodo (10.5281/zenodo.14669301) and the code from github.com/raﬁulbiswas/hopespeech. Then: (1) count the rows in the English CSV and verify whether they sum to 10,036 or to the 4,036 implied by Table 9; (2) recompute Fleiss' kappa on the 200-instance annotation sample (Section 6) using the released annotations; (3) rerun the LR, Naive Bayes, AraBERT, and BERT baselines from Section 5.2 on the released splits and compare F1 to the abstract's 0.69 and Table 11's ≤0.47. If the released data and code reproduce the body's tables (kappa 0.31–0.56, F1 ≤0.47, English n=4,036), the abstract's validation claims are refuted by the paper's own evidence. If they reproduce the abstract's figures, then the body's tables and Section 3.1 are the erroneous parts and must be corrected.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that EmoHopeSpeech is a large, reliable bilingual emotion and hope speech dataset. The abstract explicitly uses Fleiss' kappa of 0.75–0.85 and baseline F1 of 0.69 as evidence that the annotations are 'worthy.' The body, however, reports Table 12 kappas of only 0.31–0.56 (fair to moderate agreement) and Table 11 F1 scores no higher than 0.47. Additionally, Section 3.1 states that only 4,036 of the 10,036 filtered English rows were actually annotated, and Table 9's English category counts sum to exactly 4,036 — contradicting the abstract's '10,036 entries for English.' These are not cosmetic differences: the abstract's numbers are the stated validation of annotation quality, while the body's numbers show only modest agreement and near-majority-class baselines. Because the paper's contribution is a dataset, the reliability of the labels is the load-bearing premise. If the body's numbers are correct, the abstract overstates both the dataset's size and its annotation reliability; if the abstract is correct, the body's evaluation tables and Section 3.1 are wrong. Either way, the submitted manuscript does not coherently support the central claim of a validated, large bilingual resource.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents EmoHopeSpeech, a bilingual (Arabic/English) corpus annotated for emotion intensity, emotion complexity, emotion cause, and hope speech categories with subcategories. The authors describe data collection from existing emotion datasets, filtering, annotation by native speakers, inter-annotator agreement analysis, chi-square association tests between emotion features and hope speech, and baseline classification experiments with logistic regression, naive Bayes, and transformer models (AraBERT/BERT). The central claims are that the dataset is large and that its annotations are reliable, supported by the abstract's reported Fleiss' Kappa of 0.75–0.85 and F1-score of 0.69.","tokens_in":13538,"tokens_out":6816,"duration_ms":62807,"significance":"If the reported reliability and size figures were accurate, EmoHopeSpeech would be a valuable bilingual resource for affective computing, particularly for Arabic, and the public release via Zenodo and GitHub is a positive feature. The inclusion of multiple annotation layers (intensity, complexity, cause, hope subcategories) addresses a genuine gap. However, as submitted, the manuscript contains direct contradictions between the abstract and the body's evaluation tables, internal inconsistencies in the dataset statistics, and a circular validation argument. These issues affect the central claim of a validated large-scale resource, so the paper's contribution as currently stated is not established. The dataset may still be useful to the community, but the claims need to be substantially corrected and the validation strengthened.","major_comments":[{"comment":"The abstract states that Fleiss' Kappa revealed 0.75–0.85 agreement and that baseline F1-score 0.69 validates the annotations, but Table 12 reports kappas of 0.31–0.56 and Table 11 reports F1 scores no higher than 0.47. The paper's own conclusion (Section 7) characterizes the kappas as 'moderate.' Since annotation reliability is the central claim of a dataset paper, this direct contradiction must be resolved: either the body's tables are wrong or the abstract's figures are unsupported. The manuscript cannot be accepted with both sets of numbers standing.","section":"Abstract; §6, Table 12; §5.2, Table 11"},{"comment":"Section 3.1 states that of 10,036 English rows remaining after filtering, only 4,036 were annotated and included in the study, yet the abstract reports '10,036 entries for English' as part of the dataset. Table 9's English counts sum to 4,036, confirming that the released/analyzed English resource is 4,036 rows, not 10,036. The abstract overstates the dataset size by a factor of 2.5, which is a load-bearing misrepresentation of the resource.","section":"§3.1; Abstract"},{"comment":"The sentence 'This evaluation not only validated the robustness of the annotation process' uses model performance on the same annotated pool as evidence of label quality. Because the models are trained and evaluated on annotations produced by the same process, high or low F1 does not independently validate the annotations; it only measures learnability under the annotators' labels. An external gold subset, a comparison with existing benchmark labels, or at minimum a clear train/test split and discussion of the ceiling imposed by the measured kappa is needed to support the reliability claim.","section":"§5.2"},{"comment":"The dataset statistics are internally inconsistent: Table 7 English emotion counts sum to 4,031 while Table 9 English hope-speech category counts sum to 4,036; Table 8's English complexity, intensity, and cause counts sum to 3,946, 3,980, and 3,987 respectively. These discrepancies are unexplained and undermine the statistical analyses in Section 5.1, which rely on the same counts. The authors should recompute and reconcile all descriptive statistics against the released data.","section":"§4, Tables 7–9"}],"minor_comments":[{"comment":"The abstract printed at the top of the paper reports '23,456 entries for Arabic,' while the full-text abstract and Section 3.1 report 27,456; this should be corrected to a single number.","section":"Abstract (first page)"},{"comment":"In the English hope-speech prediction block, the model is labeled 'AraBERT,' but the text says BERT-based-uncased was used for English; this appears to be a copy-paste error.","section":"Table 11"},{"comment":"The Fleiss' Kappa calculation is described only as being on 'a sample of 200 instances'; the selection procedure for this sample and whether it is representative of the full corpus should be stated.","section":"§6"},{"comment":"The heading 'Corelation' is misspelled; it should be 'Correlation.'","section":"§5.1"},{"comment":"The in-text citation 'Divakaran, Girish, and Shashirekha' is missing a year; the reference list gives 2024, so the citation should include it.","section":"§2"},{"comment":"The dataset release paragraph is grammatically incomplete ('We have made dataset publicly available to the 10.5281/zenodo.14669301'); it should provide the full URL and say 'available at' instead of 'to the.'","section":"§9"}],"recommendation":"major_revision","confidential_remarks":"The most serious problem is not a methodological disagreement but a factual one: the abstract and body report incompatible numbers for the paper's headline reliability statistics. I would require the authors to reconcile these before any further consideration. If the body numbers are correct, the abstract must be rewritten to state the actual kappa range and F1 scores, and the circular validation language in §5.2 should be removed or replaced with an external check."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I read EmoHopeSpeech because the topic is timely and the dataset could be a real resource for Arabic emotion and hope speech work. The annotation scheme is the genuinely new part: combining emotion intensity, complexity, cause, and hope speech (with subcategories) in a single bilingual corpus, and doing it for Arabic where hope speech resources are scarce. The guidelines are detailed with examples in both languages, and the authors offer the data and code openly on Zenodo and GitHub. That is a real contribution, and the cross-linguistic angle is a good idea.\n\nThe problem is that the paper's reliability claims are internally inconsistent, and this is a dataset paper where label quality is the load-bearing premise. The abstract reports Fleiss' kappa of 0.75–0.85 and F1 of 0.69, but the body's Table 12 reports kappas of 0.31–0.56, and Table 11 reports F1 scores no higher than 0.47. The abstract also says the English set has 10,036 entries, while Section 3.1 clearly states only 4,036 rows were annotated, and Table 9's counts sum to exactly 4,036. The Arabic count itself is inconsistent between the submitted abstract (23,456) and the full text (27,456). These are not small rounding errors; they are the numbers used to justify \"worthy\" annotations. The body's own discussion is honest about the moderate agreement, so this looks like a badly written abstract rather than fabricated results, but the paper as submitted cannot be trusted until the numbers are reconciled.\n\nThere is a secondary weakness: Section 5.2 treats baseline model performance on the authors' own annotations as validation of annotation quality. That is circular. A model can learn patterns from noisy labels, and F1 around 0.45–0.47 is close to what one would expect from the majority class; it does not license the claim that the annotations are robust.\n\nWho is this for? Researchers working on Arabic emotion or hope speech might find the dataset useful after correction, but the current version is not a reliable substrate for that work. The paper deserves a serious referee, not a desk reject, because the underlying resource is new and the methodological flaws are fixable. My recommendation: send it to review with a request for major revision, specifically requiring the abstract to match the body, the English dataset size to be reported honestly, and the validation claims to be toned down to what the data actually show. As is, I would not cite it.","headline":"Potentially useful bilingual dataset, but the submitted abstract's reliability numbers contradict the body's own tables, and the paper cannot currently support its central claim.","tokens_in":14007,"tokens_out":2740,"would_cite":false,"duration_ms":27595,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A bilingual Arabic-English dataset links emotion labels to hope speech.","keywords":["emotion recognition","hope speech detection","Arabic NLP","bilingual dataset","annotation reliability","Fleiss kappa","cross-linguistic analysis","social media text"],"falsifier":"Recompute Fleiss' kappa on the full released Arabic and English files using the paper's own category definitions. If the full-corpus values match the body's 0.31-0.56 range rather than the abstract's 0.75-0.85, the reliability claim is contradicted; likewise, re-running the best classifier on an independently re-annotated sample should come near the abstract's F1 of 0.69 if that validation claim is correct.","tokens_in":13049,"feed_emoji":"💬","tokens_out":9013,"duration_ms":79589,"temperature":0.7,"pith_summary":"EmoHopeSpeech is a new resource of 27,456 Arabic and 10,036 English social-media posts, each carrying a basic emotion label plus deeper annotations: emotion intensity, complexity, cause, and hope-speech category with subcategories. The authors aim to provide the first large English-Arabic dataset that supports joint analysis of emotion and hope speech, addressing what they see as a gap for Arabic in particular. If the resource works as claimed, it would let researchers compare how positivity and emotion are expressed across languages, and give model builders a training ground for detecting uplifting rather than only toxic content. The paper validates the resource through annotator agreement and baseline classifiers, and its abstract reports high validation scores.","feed_headline":"A bilingual dataset links emotion labels to hope speech","feed_subtitle":"27,456 Arabic and 10,036 English posts annotated for emotion intensity, cause, and hope category.","key_machinery":"The load-bearing instrument is the multi-axis annotation scheme, which records emotion and hope speech as layered structures rather than a single sentiment score. It is what lets the paper compute relationships between emotion characteristics and hope categories, and it defines the gold standard on which the baseline models are trained. The evaluation machinery pairs that scheme with Fleiss' kappa computed on a shared sample of 200 instances, plus logistic regression, multinomial naive Bayes, and fine-tuned AraBERT and BERT classifiers for hope-speech prediction.","core_discovery":"The central claim is that the paper has built the first large-scale bilingual dataset combining emotion and hope speech annotations in English and Arabic. The contribution is the annotation scheme itself: each text receives an emotion label, an intensity level, a complexity level, a cause category, a hope-speech category (hope, not hope, counter, neutral, hate/negativity), and a hope subcategory (inspirational, solidarity, resilience, spiritual). The paper argues that the annotations are reliable enough to support machine-learning detection, and that statistically significant chi-square associations between emotion features and hope categories--for example emotion cause versus hope category ($\\chi^2=810.3$ for Arabic and $245.6$ for English)--show the two phenomena are genuinely coupled in the data.","pith_inferences":["Editorial inference: the body's agreement values (Fleiss kappa 0.31-0.56) are below the abstract's (0.75-0.85), so a cautious user should treat the main hope-speech category as the most dependable label and the emotion-cause and subcategory labels as exploratory.","Because emotion cause is the strongest correlate of hope category, a natural follow-up is a cause-aware model that predicts hope speech using emotion-cause labels, testing whether it beats the paper's plain baselines.","The large distributional differences between the two languages (e.g., 80% neutral labels in English versus 13% in Arabic) may reflect the source corpora as much as cultural expression; a rigorous cross-cultural comparison would require matched sampling across genres."],"forward_implications":["Cross-linguistic comparisons become concrete: the paper reports that Arabic skews toward medium-complexity and internal-reflection emotions, while English skews toward simple complexity and low intensity.","Joint emotion-hope modeling becomes feasible because emotion cause is strongly associated with hope category in both languages, giving models a predictive signal beyond the hope label itself.","The four hope subcategories turn hope detection from a binary task into a finer-grained classification, which is useful for applications that want to distinguish encouragement, solidarity, resilience, and spiritual support.","The body's baseline results (best F1 of 0.47 for Arabic and 0.45 for English) give later work concrete numbers to beat rather than a claim that the task is solved."],"supporting_citations":[{"why":"Supplies the prior multilingual hope speech dataset (English, Tamil, Malayalam) that motivates and contrasts with the new Arabic-English resource.","marker":"Chakravarthi 2020"},{"why":"Provides the two-level hope speech detection model and taxonomy that the annotation scheme adapts.","marker":"Balouchzahi, Sidorov, and Gelbukh 2023"},{"why":"Source of the Arabic Emotional Tone tweets used in the corpus.","marker":"Al-Khatib and El-Beltagy 2018"},{"why":"Source of the Arabic poetry emotion data used in the corpus.","marker":"Shahriar, Al Roken, and Zualkernan 2023"},{"why":"Source of the Arabic multi-label hate speech data that was re-annotated for this dataset.","marker":"Zaghouani, Mubarak, and Biswas 2024"},{"why":"Source of the English emotion-labeled text data.","marker":"Anjali 2024"},{"why":"Provides BERT, the English baseline model used to evaluate hope-speech prediction.","marker":"Devlin et al. 2018"},{"why":"Provides AraBERT, the Arabic baseline model used to evaluate hope-speech prediction.","marker":"Antoun, Baly, and Hajj 2020"}],"fun_headline_variants":["Bilingual corpus pairs emotions with hope categories","First annotated set for emotion and hope in Arabic and English","23,456 Arabic and 10,036 English posts for emotion and hope","EmoHopeSpeech dataset links emotion intensity to hope categories"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the validation figures in the abstract--Fleiss kappa 0.75-0.85 and F1 0.69--are the correct ones, while the body reports kappa 0.31-0.56 and F1 no higher than 0.47; if the abstract's numbers are wrong, the claim that the annotations are trustworthy collapses.","fun_headline_variants_meta":{"raw":{"variants":["Bilingual corpus pairs emotions with hope categories","First annotated set for emotion and hope in Arabic and English","23,456 Arabic and 10,036 English posts for emotion and hope","EmoHopeSpeech dataset links emotion intensity to hope categories"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001748,"raw_usage":{"total_tokens":6849,"prompt_tokens":838,"completion_tokens":6011,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":454,"completion_tokens_details":{"reasoning_tokens":5943}},"tokens_in":454,"tokens_out":6011,"duration_ms":39110,"temperature":1.0,"reasoning_tokens":5943,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:43:45.084551+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute Fleiss' kappa on the full released Arabic and English files using the paper's own category definitions. If the full-corpus values match the body's 0.31-0.56 range rather than the abstract's 0.75-0.85, the reliability claim is contradicted; likewise, re-running the best classifier on an independently re-annotated sample should come near the abstract's F1 of 0.69 if that validation claim is correct.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the prior multilingual hope speech dataset (English, Tamil, Malayalam) that motivates and contrasts with the new Arabic-English resource."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the two-level hope speech detection model and taxonomy that the annotation scheme adapts."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the Arabic Emotional Tone tweets used in the corpus."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the Arabic poetry emotion data used in the corpus."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the Arabic multi-label hate speech data that was re-annotated for this dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the English emotion-labeled text data."}],"review_version":1}