{"id":"724c9f9e-266f-4aa4-97e0-442aea598d75","arxiv_id":"2508.15810","paper_version":1,"verdict":"UNVERDICTED","confidence":"UNKNOWN","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"The submission cannot be reviewed as a coherent paper: its abstract and full text are two different papers, so the abstract's claims have no supporting body.","lead":"This arXiv record's abstract describes an Arabic hate and hope detection system, but the supplied full text is a different paper about a scientific journal cover benchmark. Because the abstract and body describe two unrelated studies, the claimed results cannot be checked against any supporting methods or data.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Submission package mismatch: the abstract's MAHED results have no supporting methods or experiments in the supplied full text, so the central empirical claim is unverifiable.","rationale":"The reader's weakest_assumption identifies the same load-bearing concern: the abstract and the supplied full text are different papers, so the abstract's empirical claims have no supporting derivation in the submitted document. I agree with that assessment. The strongest claim from the abstract is plausible as a standalone statement, but the condition required for evaluating it—that the submitted full text describes the MAHED 2025 experiments—is contradicted by the document itself. The full text's own header and content identify it as arXiv:2508.15802, a benchmark paper about scientific journal covers. The abstract never appears in that body, and the body never mentions Arabic, hate speech, hope, memes, or MAHED 2025. Per the review rule, I treat this mismatch as in-scope evidence rather than dismissing it as pipeline noise. The mismatch does not, by itself, prove the reported F1 scores are wrong, but it makes them impossible to check from the submitted record. I am not raising a substantive objection to the scientific content of the MAC paper, because that is not the paper under review. I am also not accusing the authors of misconduct; the mismatch could originate from the submission pipeline. However, the consequence is the same: the central claim of the abstract is unverifiable from the provided materials. Since the reader's verdict is already UNVERDICTED, my read does not change it. A single concrete check—fetching the canonical arXiv record for 2508.15810 and checking whether it contains the MAHED experiments—would settle whether this concern actually lands.","tokens_in":18167,"tokens_out":2591,"duration_ms":26636,"concrete_test":"Retrieve the canonical arXiv record for 2508.15810 (via arXiv API or arXiv HTML/PDF) and inspect its full text. Check whether it contains the MAHED 2025 dataset description, task definitions, fine-tuning details, and result tables. If the canonical body matches the abstract, verify the macro-F1 numbers against the MAHED 2025 evaluation script and official leaderboard. If the canonical body is still the MAC paper, the submitted record contains no supporting experiments and the paper remains unverdictable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that fine-tuned GPT-4o-mini (Arabic text) and Gemini Flash 2.5 (memes) achieve macro-F1 scores of 72.1%, 57.8%, and 79.6% and secure first place overall on the MAHED 2025 challenge. For that claim to be assessable, the submitted document must contain the corresponding experimental setup: dataset splits, fine-tuning hyperparameters, evaluation protocol, and result tables. The supplied full text does not provide any of this. Its title page identifies it as arXiv:2508.15802, 'MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding' by Jiang et al., and the body (§1–§6) is entirely about scientific journal-cover understanding. The abstract's reported numbers are therefore assertions without derivation in this record. This is not an internal inconsistency in the MAC paper; it is a mismatch between the abstract and the artifact submitted as its full text. Because the reported scores, the 'first place overall' claim, and the claim that fine-tuned commercial LLMs outperform alternatives all depend on experiments absent from the record, the weakest assumption—that the full text corresponds to the abstract—fails. The correct status is unverdictable, not accepted or rejected, unless the canonical document resolves the mismatch.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The abstract of this submission claims that fine-tuned GPT-4o-mini (for Arabic textual speech) and Gemini Flash 2.5 (for Arabic memes) achieve macro-F1 scores of 72.1%, 57.8%, and 79.6% on tasks 1, 2, and 3 of the ArabicNLP MAHED 2025 challenge, and that the proposed system secured first place overall. The supplied full text, however, is a different paper: 'MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding' (arXiv:2508.15802), which proposes a benchmark of scientific journal covers and evaluates MLLMs on image-to-text and text-to-image matching. This full text contains no Arabic datasets, no mention of MAHED, no hate-speech/hope/emotion tasks, no fine-tuning configurations, and no results tables reporting macro-F1 scores. The central empirical claim of the abstract is therefore unsupported by any experimental description in the submitted artifact.","tokens_in":18271,"tokens_out":3493,"duration_ms":38398,"significance":"If the claimed results were substantiated, they would be of practical value for Arabic content moderation and would provide evidence that fine-tuned commercial LLMs can perform strongly on an external, community-organized benchmark. A positive feature is that the evaluation target (MAHED 2025) is external to the authors, so the claimed outcome is not defined circularly in terms of the method itself. However, the significance cannot currently be assessed because the submitted manuscript body is a coherent but unrelated benchmark paper. No supporting experimental details, dataset splits, hyperparameters, evaluation protocol, or code are present in the record. Thus the claimed contribution is, at this stage, an abstract-level assertion without a verifiable basis.","major_comments":[{"comment":"The abstract reports the paper's only occurrence of the three macro-F1 scores (72.1%, 57.8%, 79.6%) and the 'first place overall' claim. The supplied full text is arXiv:2508.15802, 'MAC: A Live Benchmark for Multimodal Large Language Models in Scientific Understanding.' This body does not mention Arabic, MAHED, GPT-4o-mini, Gemini Flash 2.5, hope, hate speech, offensive language, or emotions. No table or equation in the full text reports macro-F1 or any result for tasks 1-3. The central claim of the paper is therefore entirely unverified by the submitted artifact.","section":"Abstract vs. full-text manuscript"},{"comment":"The experimental sections evaluate models on MAC-2025 scientific-cover matching using accuracy, ECE, NLL, and RMS. There are no training splits, fine-tuning hyperparameters, prompt templates, evaluation scripts, or held-out test details for the Arabic tasks described in the abstract. Even if the full text were read as the intended paper, it would not support the premise that the reported MAHED scores come from the official evaluation protocol on a held-out test set without information leakage. This is a load-bearing gap: the abstract's numbers cannot be checked or reproduced.","section":"Full text, §4 and Tables 1-8"},{"comment":"The submission's title page identifies the full text as a COLM 2025 paper by Jiang et al. with arXiv:2508.15802, while the abstract describes a different MAHED 2025 system paper. This is not a minor presentation issue; it means the document as submitted is internally inconsistent. The authors (or the editorial office) need to supply the canonical manuscript matching the abstract before the claimed results can receive substantive review.","section":"Manuscript identity and internal consistency"}],"minor_comments":[{"comment":"The challenge name is spelled inconsistently: 'MAHED 2025' appears in one sentence and 'Mahed 2025' in another. Please standardize.","section":"Abstract"},{"comment":"Several figure captions and text passages in the supplied full text contain placeholder glyphs (e.g., '����'), indicating an encoding or rendering problem. If this manuscript is to be considered, these should be repaired.","section":"Full text, figures and captions"},{"comment":"Once the correct manuscript is supplied, the authors should add a reproducibility appendix listing the training/validation splits, fine-tuning budgets, model versions, and the exact evaluation protocol used for the MAHED 2025 test set.","section":"General"}],"recommendation":"uncertain","confidential_remarks":"This appears to be a submission-package mismatch: the abstract and the full text describe two different arXiv papers. I cannot render a content-based verdict on the claimed MAHED results without the correct manuscript. I recommend that the editor verify whether the wrong PDF was uploaded and, if so, obtain the canonical version. If the submitted full text is genuinely what the authors intend to publish, then the central claim is unsupported and the paper should not proceed in its current form."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing you should know up front: this submission cannot be read as a single paper. The abstract and metadata match an Arabic hate/hope/emotion detection system for the MAHED 2025 shared task, with claimed macro-F1 scores of 72.1/57.8/79.6 and a first-place finish. The supplied full text, however, is the MAC benchmark paper (arXiv 2508.15802) about multimodal scientific understanding, with different authors and no mention of Arabic, hate speech, or MAHED. That is not a minor packaging glitch; it means the central empirical claims are assertions with no accompanying methods, tables, or evaluation details in the record.\n\nWhat is genuinely here: nothing from the full text bears on the abstract. The abstract itself describes a reasonable applied recipe—evaluate base LLMs, fine-tune GPT-4o-mini for text and Gemini Flash 2.5 for memes, and compare against pretrained embeddings on an external benchmark. If the missing body actually backs up those numbers, this is a modest but legitimate contribution to Arabic content moderation, the kind of system description that shared-task venues publish routinely. But I cannot give it credit for experiments I cannot see, and the document I was sent is the wrong artifact.\n\nSoft spots, in proportion: the mismatch is load-bearing rather than cosmetic. Even taking the abstract at face value, it reports no training splits, hyperparameters, checkpoints, or evaluation protocol, so the scores are not independently checkable. The abstract also cites no prior work, which is defensible for a short shared-task report but does not help situate the result. The full text's citation list is irrelevant to this submission. None of this is a scientific objection to the underlying work the authors presumably intended to describe—it is a statement that the submitted record does not contain that work.\n\nWho is this for? If a corrected version appears, the intended audience is practitioners building Arabic moderation and emotion-detection systems. As-is, the paper is useful to nobody because it self-contradicts on its own terms.\n\nMy recommendation to the editor: do not send this to peer review. Send it back to the authors to upload the correct full text, and then evaluate the corrected package. If the corrected version matches the abstract, it deserves a normal shared-task review. The current record does not.","headline":"The submission package is broken: the abstract describes an Arabic shared-task paper, but the full text is a different paper on scientific-journal covers, so the reported MAHED results have no derivable support in this record.","tokens_in":18935,"tokens_out":1574,"would_cite":false,"duration_ms":18142,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Although its full text is a different benchmark paper, the abstract claims fine-tuned GPT-4o-mini and Gemini Flash 2.5 win the MAHED 2025 Arabic hope/hate/emotion tasks with macro F1 up to 79.6%.","keywords":["Arabic NLP","hate speech detection","hope detection","emotion detection","multimodal memes","LLM fine-tuning","MAHED 2025","content moderation"],"falsifier":"Run the three MAHED 2025 tasks on the official test split with the same fine-tuned model versions and the same evaluation script; the claim fails if the macro F1 scores do not reproduce the reported 72.1%, 57.8%, and 79.6% within normal variance. A simpler check available now: the submission's own full text is a different paper, so the manuscript itself contains no task-specific result tables to verify.","tokens_in":17876,"feed_emoji":"🏆","tokens_out":4389,"duration_ms":47567,"temperature":0.7,"pith_summary":"The paper sets out to show that fine-tuned large language models beat both base LLMs and pre-trained embedding models on three Arabic-language tasks—detecting hope, hate/offensive speech, and emotion—in both plain text and memes, using the MAHED 2025 challenge data. Its headline claim is that GPT-4o-mini fine-tuned on Arabic text and Gemini Flash 2.5 fine-tuned on Arabic memes achieve macro F1 scores of up to 72.1%, 57.8%, and 79.6% on the three tasks, securing first place overall. A reader should care because Arabic content moderation is under-resourced, and a winning recipe built from off-the-shelf commercial models would be immediately usable. However, the full text supplied under this title is a different paper, describing a benchmark called MAC for multimodal scientific understanding; none of the experiments behind the abstract's numbers appear in it. The abstract's claim therefore cannot be checked from the body as submitted.","feed_headline":"Fine-tuned LLMs top Arabic hope-hate-emotion challenge","feed_subtitle":"The winning recipe: GPT-4o-mini for text, Gemini Flash 2.5 for memes, with macro F1 up to 79.6 percent.","key_machinery":"The central mechanism is modality-matched fine-tuning: taking a general-purpose pretrained LLM and adapting it on task-specific Arabic text or Arabic memes rather than using it zero-shot. The abstract credits this mechanism for the reported macro-F1 gains and the first-place finish, claiming that the fine-tuned commercial models extract the signal that base LLMs and pre-trained embedding models miss. The full text as supplied does not describe this mechanism or its experiments.","core_discovery":"The discovery the authors report is that task-specific fine-tuning of commercial LLMs outperforms both base LLMs and pre-trained embedding classifiers on Arabic hope, hate/offensive, and emotion detection, and that the best model differs by modality: GPT-4o-mini fine-tuned on Arabic textual speech leads for text tasks, while Gemini Flash 2.5 fine-tuned on Arabic memes leads for multimodal meme tasks. They report macro F1 scores of 72.1%, 57.8%, and 79.6% for tasks 1, 2, and 3, and first place in the MAHED 2025 challenge. The manuscript body provided here, however, is the text of a different paper (MAC-2025, a live benchmark for multimodal scientific understanding), so the experiments, datase","pith_inferences":["Editorial inference: If the reported scores hold, the practical bottleneck likely shifts to annotation quality and class balance, since macro F1 is sensitive to label distribution; a testable extension is per-class error analysis to identify which hate and emotion categories remain confused.","Editorial inference: Because the winning models are commercial APIs, the exact scores may not be reproducible later as model versions change; a useful extension is to record API versions and checkpoints at evaluation time.","Editorial inference: The same fine-tuning recipe would plausibly transfer to dialectal or code-switched Arabic social media, and could be tested on Egyptian or Maghrebi datasets to measure generalization beyond the MAHED split.","Editorial inference: The mismatch between the abstract and the supplied full text suggests the method description may live in a separate report; if the MAC text is a misattached file, the challenge-system paper could still be released with training details and per-task breakdowns."],"forward_implications":["Fine-tuned commercial LLM APIs could serve as ready Arabic content moderation systems, avoiding the need to train models from scratch.","Text and meme detection would be best treated as separate pipelines, since the winning models differ by modality.","The reported 79.6% macro F1 top score would mark the current ceiling for automatic Arabic hope/hate/emotion detection with commercial API models on the MAHED data.","On shared challenges like MAHED, the competitive edge would shift from architecture design to fine-tuning configuration, data selection, and evaluation strategy."],"supporting_citations":[],"fun_headline_variants":["Fine-tuned LLMs win Arabic hope-hate-emotion","Per-modality tuning tops Arabic text and meme detection","GPT-4o-mini and Gemini Flash 2.5 lead Arabic content moderation","Arabic emotion detection: fine-tuning beats base LLMs","Fine-tuned LLMs hit 79.6% F1 on Arabic multimodal tasks"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The claim collapses if the reported macro-F1 numbers were not produced by the described fine-tuned GPT-4o-mini and Gemini Flash 2.5 models on the official MAHED 2025 held-out test set, because the submitted full text contains no experiments to support them.","fun_headline_variants_meta":{"raw":{"variants":["Fine-tuned LLMs win Arabic hope-hate-emotion","Per-modality tuning tops Arabic text and meme detection","GPT-4o-mini and Gemini Flash 2.5 lead Arabic content moderation","Arabic emotion detection: fine-tuning beats base LLMs","Fine-tuned LLMs hit 79.6% F1 on Arabic multimodal tasks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000548,"raw_usage":{"total_tokens":2492,"prompt_tokens":819,"completion_tokens":1673,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":1595}},"tokens_in":563,"tokens_out":1673,"duration_ms":16161,"temperature":1.0,"reasoning_tokens":1595,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:59:41.258027+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the three MAHED 2025 tasks on the official test split with the same fine-tuned model versions and the same evaluation script; the claim fails if the macro F1 scores do not reproduce the reported 72.1%, 57.8%, and 79.6% within normal variance. A simpler check available now: the submission's own full text is a different paper, so the manuscript itself contains no task-specific result tables to verify.","supporting_citations":[],"review_version":1}