{"id":"c6c1a858-ac99-4887-a6ac-57b93bcf7286","arxiv_id":"2508.17378","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A 115-response UI test of ALLaM 34B claims strong Arabic abilities, but the dialect score in the main table conflicts with the paper's own heat map and the underlying data is missing.","lead":"A short evaluation of ALLaM 34B through the closed HUMAIN Chat interface reports high Arabic performance, but its own dialect heat map contradicts the headline score and no raw data is released. The paper is a useful cautionary example of why small, artifact-free UI evaluations are hard to trust.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported dialect results are internally inconsistent: Table 1's Dialect mean (4.21) cannot be reconciled with Figure 3's per-dialect averages (~3.3), undermining the central claim of cultural grounding.","rationale":"I read the paper as a UI-level empirical evaluation whose central claim is the robustness and cultural grounding of ALLaM-34B. For that claim to hold, the reported scores must at least be internally coherent. The clearest load-bearing failure is the conflict between Table 1 and Figure 3: the dialect category mean cannot be 4.21 if the five dialect conditions in Fig. 3 average about 3.3. This is not a disagreement with consensus; it is an internal arithmetic contradiction in the evidence the paper itself presents. It also aligns with the authors' own qualitative observation that the model replies in MSA or English for dialect prompts, which contradicts 'strong dialect fidelity.' The reader's stated weakest assumption—that HUMAIN Chat really runs ALLaM-34B—is also important and untestable, but the internal inconsistency is more direct and does not depend on external facts. I did not find a reason to reject the LLM-as-judge methodology per se; the decisive issue is the unreproducible and inconsistent score reporting. A test that simply asks for the raw score matrix and recomputes one category mean would settle whether the headline claim has a factual basis. For these reasons I agree with the reader's rejection; no verdict adjustment is needed beyond the reader's REJECT.","tokens_in":6324,"tokens_out":3837,"duration_ms":35881,"concrete_test":"Release the full score matrix (23 prompts × 5 runs × 3 judges × 5 metrics) or, minimally, per-prompt category means. Recompute the Dialect category mean from the raw dialect-prompt scores. If the recomputed mean is ≈3.3 (consistent with Fig. 3) rather than 4.21, Table 1 and the abstract's '4.21/5' dialect claim are wrong and the central conclusion is unsupported. Independently check the three adversarial rows in Table 1: if every run and judge produced identical 4.20 scores, verify whether the scoring was accidentally constant rather than genuinely robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—'robust and culturally grounded Arabic LLM ... practical readiness'—rests on the reported category scores, especially dialect fidelity. Those scores are internally inconsistent. Table 1 reports Dialect mean 4.21 [4.09, 4.34]. Figure 3 gives per-dialect overall means: Najdi 3.8, Hijazi 3.7, Egyptian 3.7, Levantine 2.7, Moroccan 2.7; their average is 3.32, nearly a full point lower. The dialect-fidelity metric in the same figure averages (3.9+3.8+3.7+2.9+2.6)/5 = 3.38, not 4.21. Section 3.2's qualitative findings corroborate the lower figure: for Najdi/Hijazi/Egyptian prompts the model frequently answers in English or MSA rather than the requested dialect. Absent raw per-prompt scores, the only way to reconcile Table 1 with Figure 3 is to assume the 'Dialect' category in Table 1 includes different or weighted prompts than the five dialects in Figure 3—but that is not stated, and the text says 'Dialect prompts score 4.21 on average.' This internal contradiction directly undermines the abstract's dialect-fidelity claim and the headline assertion of cultural grounding.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a UI-level evaluation of the ALLaM 34B model as deployed in the closed HUMAIN Chat web service. The authors constructed 23 prompts spanning MSA, five dialects, code-switching, knowledge, reasoning, generation, and safety/adversarial categories, collected five responses per prompt (115 total), and scored each response with three frontier LLM judges on accuracy, fluency, instruction following, safety, and dialect fidelity. They report category-level means with 95% confidence intervals, claim near-perfect performance on code-switching and generation, strong performance on MSA and reasoning, and a dialect score of 4.21/5. The paper concludes that ALLaM 34B is a robust, culturally grounded Arabic LLM ready for real-world deployment.","tokens_in":6611,"tokens_out":4748,"duration_ms":45511,"significance":"If the measurements were trustworthy, a UI-level evaluation of a closed Arabic-centric commercial model would be a useful contribution, particularly the dialectal breakdown and the multi-metric LLM-judge pipeline. The paper also includes qualitative examples and human-judge validation, which are appropriate steps. However, the significance is heavily contingent on the internal consistency of the reported scores and on the verifiability of the claim that HUMAIN Chat is actually powered by ALLaM 34B. As the manuscript stands, those two issues are not resolved, and the central claims are therefore not supported by the evidence presented.","major_comments":[{"comment":"The reported Dialect category mean of 4.21 (95% CI [4.09, 4.34]) in Table 1 is directly contradicted by the per-dialect averages shown in Figure 3: Najdi 3.8, Hijazi 3.7, Egyptian 3.7, Levantine 2.7, Moroccan 2.7, which average 3.32. The same figure gives dialect-fidelity scores of 3.9, 3.8, 3.7, 2.9, and 2.6, averaging 3.38, also far below 4.21. Section 3.2's own narrative states that Levantine drops to 2.73 and Moroccan is weaker at 3.3. This internal inconsistency means the abstract's headline claim of improved dialect fidelity (4.21/5) is not supported by the paper's own data. The authors must provide the per-prompt raw scores and either reconcile the aggregation or correct the error.","section":"§3.1, Table 1 vs. §3.2, Figure 3"},{"comment":"The evaluation assumes that HUMAIN Chat is running ALLaM 34B, but the service is closed, has no public API, and the paper provides no evidence that the responses were produced by that model rather than by a different model, a wrapper, or a post-processing pipeline. Since the paper's title and abstract attribute every measurement to ALLaM 34B, this assumption is load-bearing for every reported score. Without a verification protocol (e.g., testing known distinguishing behaviors or comparing with public ALLaM checkpoints) or a clear reframing of the claims as being about the HUMAIN Chat service, the conclusions about ALLaM 34B specifically are unverifiable.","section":"§2.2, sampling protocol and abstract"},{"comment":"The three adversarial categories (Prompt Injection, Jailbreak, Data Exfiltration) each report a mean of exactly 4.20 with a zero-width confidence interval. Given that each category is based on distinct prompts with five runs each and three independent LLM judges scoring multiple metrics, an exactly zero variance across all runs and judges is implausible and suggests either a data-processing artifact or a reporting error. The paper should disclose the underlying score distributions or explain how the zero variance arose; as presented, these numbers undermine confidence in the measurement pipeline.","section":"§3.1, Table 1, adversarial categories"}],"minor_comments":[{"comment":"The title in the header reads 'UI-L EVEL EVALUATION OF ALL AM 34B' with inconsistent spacing; the model name is also written variously as 'ALLaM 34B' and 'ALLaM-34B'. Please standardize the formatting.","section":"Title and abstract"},{"comment":"The sentence beginning 'Dialectal fidelity:p Uneven performance' contains a stray colon and letter 'p'; it should read 'Dialectal fidelity: Uneven performance'.","section":"§3.2"},{"comment":"The word 'acroos' is a typo for 'across'.","section":"Figure 1 caption"},{"comment":"The Arabic prompt and response excerpts are garbled due to encoding issues, making the qualitative evidence unreadable. Please provide properly typeset Arabic text or transliterations.","section":"§3.2 qualitative examples"},{"comment":"The human evaluation validation is described in one sentence with no information on the number of raters, the number of responses reviewed, or the agreement measure used. Adding these details would strengthen the reliability discussion.","section":"§2.4 human evaluation"},{"comment":"The phrase 'improved dialect fidelity' implies a comparison to a previous baseline, but no prior evaluation or baseline is presented in the paper. Please either state the baseline or remove the word 'improved'.","section":"Abstract"}],"recommendation":"reject","confidential_remarks":"The paper's central claims rest on internally inconsistent data (Table 1 vs. Figure 3) and on an unverifiable attribution of the closed service's outputs to ALLaM 34B. The zero-variance adversarial scores further suggest a reporting problem. Even though the dialect inconsistency could be corrected by recomputation, the model identity issue is fundamental and cannot be resolved without access to the service internals. The authors would need to substantially reframe the paper as an evaluation of HUMAIN Chat rather than of ALLaM 34B, and provide verifiable raw data, before it could meet the standards for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nShort version: this is a routine UI-level LLM evaluation, not a research contribution, but the dialect material is genuinely informative. The headline claim—'robust and culturally grounded ... practical readiness'—is not supported by the paper's own numbers.\n\nWhat's new: a closed-chat evaluation of ALLaM 34B with 115 outputs scored by three LLM judges. That is a datapoint, not a method or benchmark. No artifacts, no raw scores, no baseline, so reproducibility is limited. Credit where due: the qualitative examples and the heat map do real work. The observation that the model often answers MSA or English when asked for dialects is the most useful finding, and the discussion of uneven dialect support is honest.\n\nThe soft spots are serious. Table 1 reports Dialect mean 4.21 [4.09, 4.34], but Figure 3's per-dialect overall scores (Najdi 3.8, Hijazi 3.7, Egyptian 3.7, Levantine 2.7, Moroccan 2.7) average about 3.3, and the dialect-fidelity column averages about 3.4. The text even says Najdi, Hijazi, and Egyptian have 'perfect dialect fidelity,' while the figure gives 3.9, 3.8, and 3.7. These cannot all be right. Without raw per-prompt scores, a referee cannot tell whether the table or the figure is wrong, and the abstract leans on the higher number.\n\nThe adversarial categories are also suspicious: prompt injection, jailbreak, and data exfiltration all report exactly 4.20 with zero variance across three judges and five runs. That is implausible unless the scores were rounded or aggregated in a way not described.\n\nThe claimed human validation is a sentence with no protocol, sample size, or agreement numbers, so it carries no weight. The closed-UI setting means we cannot verify that HUMAIN Chat is actually serving ALLaM 34B; that is an inherent limitation, not a sin, but it should soften any deployment claim.\n\nCitation pattern is fine; the self-citation to the author's culturally aligned benchmark is relevant, and the ALLaM paper is cited. No citation red flags.\n\nNet: the paper is a workmanlike technical report with one good qualitative insight and a load-bearing numerical inconsistency. I would not cite it. If the author corrects the tables/figures, posts the prompt pack and raw scores, and rewrites the abstract to match the actual dialect results, it could be a modest useful datapoint. As is, I'd desk-reject rather than spend referee time on unreconciled numbers.\n\nBest","headline":"Useful dialect data points buried under an overclaimed abstract and internally inconsistent tables; worth a skim, not a cite.","tokens_in":7121,"tokens_out":4335,"would_cite":false,"duration_ms":40519,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Through 115 judged chat replies, ALLaM-34B shows near-perfect code-switching and generation, strong MSA and safety, and uneven dialect fidelity, which is taken as evidence that the model is culturally grounded and deployment-ready.","keywords":["ALLaM-34B","UI-level evaluation","Arabic LLM","dialectal Arabic","LLM-as-a-judge","code-switching","safety evaluation","HUMAIN Chat"],"falsifier":"Run the same 23 prompts against an open-weight ALLaM-34B deployment with the same three judges; if the category means and dialect heat map differ materially, the UI-level scores describe the service rather than the model. A cheaper check is to probe HUMAIN Chat with prompts engineered to expose another model's known fingerprint and see whether the responses match ALLaM-34B's documented behavior.","tokens_in":6123,"feed_emoji":"🗣️","tokens_out":6575,"duration_ms":61074,"temperature":0.7,"pith_summary":"ALLaM-34B, an Arabic-centric large language model served behind a closed chat interface, is evaluated here through 115 UI-level responses to 23 prompts spanning formal Arabic, five dialects, code-switching, reasoning, knowledge, generation, and adversarial safety. Three frontier LLM judges scored every response on accuracy, fluency, instruction following, safety, and dialect fidelity, and category means came out at 4.92/5 for code-switching and generation, 4.74 for MSA, 4.64 for reasoning, 4.54 for safety, and 4.21 for dialect. The paper reads these scores as evidence that the model is technically strong and culturally grounded enough for real-world Arabic deployment, while flagging Levantine and Moroccan dialect weakness and a tendency to fall back to formal MSA. A sympathetic reader would care because most Arabic LLM benchmarks are translated from English and miss the cultural and dialectal dimensions this protocol tries to measure.","feed_headline":"Arabic LLM hits 4.92/5 on code-switching and generation","feed_subtitle":"UI-level test shows code-switching and generation near 5/5, with dialect gaps in Levantine and Moroccan.","key_machinery":"The load-bearing mechanism is the four-stage evaluation pipeline: a balanced 23-prompt pack, five repeated submissions through the chat UI to capture decoding variability, independent scoring by three frontier LLM judges on five Likert-scale metrics, and aggregation into category means with 95% confidence intervals. The judges are GPT-5, Gemini 2.5 Pro, and Claude Sonnet-4, each rating accuracy, fluency, instruction following, safety, and dialect fidelity, with the overall score defined as the mean of the applicable metrics. A human evaluation subset is added to validate the automated ratings, and dialect results are visualized as metric-by-dialect heat maps. This pipeline is what converts raw chat interactions into the paper's comparative claim about Arabic capability.","core_discovery":"The central claim is that ALLaM-34B, as accessed through HUMAIN Chat, delivers consistently high-quality Arabic generation and code-switching (4.92/5 both), strong modern-standard-Arabic handling (4.74), solid reasoning (4.64), stable safety behavior (4.54), and moderate dialect fidelity (4.21), with confidence intervals narrow enough to claim reliability. Across the five tested dialects, Najdi, Hijazi, and Egyptian reach roughly 3.7–3.8 overall, while Levantine drops to 2.73 and Moroccan to about 3.3, driven by accuracy loss and a recurring fallback into MSA or English retrieval-style output. The paper also claims the model consistently refuses prompt-injection, jailbreak, and data-exfiltration attempts, with all three adversarial categories scoring 4.20 at zero variance.","pith_inferences":["Because HUMAIN Chat is closed and no model-identity probe is possible, the scores could describe a wrapper, post-processed outputs, or a different model; verifying against a weight-accessible ALLaM-34B deployment is a direct test.","The paper's own conclusion concedes the closed interface, the small 23-prompt pack, and LLM judges as limitations, which reinforces that the deployment-readiness claim rests on unverified model identity and judge alignment.","LLM judges may inflate fluency and generation scores because they reward polished prose; a native-speaker preference test would be needed to confirm the claim of cultural groundedness beyond surface fluency.","The zero-variance 4.20 adversarial scores likely reflect a judge ceiling or prompt simplicity, so a more varied adversarial suite could widen the safety gap between categories."],"forward_implications":["Arabic-English code-switching and generative writing are likely ready for user-facing products, since the near-ceiling scores and tight intervals indicate consistent behavior across runs.","Dialectal coverage is the main actionable gap: Najdi, Hijazi, and Egyptian are usable, while Levantine and Moroccan need more corpus work, dialect-specific adapters, and benchmarks that reward authentic dialect rather than formal MSA.","Safety behavior on the tested adversarial prompts is consistent, so the deployed service can probably handle routine injection and jailbreak attempts, though harder attacks remain untested.","Repeated sampling through a closed UI with multiple independent judges is a transferable protocol for evaluating any model-only service that exposes no API.","The observed drift into MSA or English on dialect prompts means that user-facing dialect features would need wrappers or constrained decoding to force the requested register."],"supporting_citations":[{"why":"Establishes the ALLaM model family, its Arabic vocabulary expansion, and its pretraining mixture, which the evaluation assumes are running behind the chat service.","marker":"[4]"},{"why":"Supplies the LLM-as-a-judge methodology used to score all responses on the five-point metrics.","marker":"[8]"},{"why":"Serves as one of the three judge models, GPT-5, that produced the accuracy, fluency, and safety scores.","marker":"[5]"},{"why":"Serves as the Gemini 2.5 Pro judge in the scoring pipeline.","marker":"[6]"},{"why":"Serves as the Claude Sonnet-4 judge in the scoring pipeline.","marker":"[7]"},{"why":"Motivates culturally aligned Arabic evaluation as the gap that the prompt pack is designed to fill.","marker":"[3]"}],"fun_headline_variants":["ALLaM-34B scores 4.92/5 on Arabic code-switching and generation","Arabic model: generation 4.92/5, Levantine dialect 2.73/5","Code-switching and generation at 4.92/5 for ALLaM-34B","Levantine Arabic largest gap in ALLaM-34B test","ALLaM-34B Arabic code-switching scores 4.92/5"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole evaluation depends on the unverified assumption that the text served by HUMAIN Chat is generated directly by ALLaM-34B, with no wrapper or post-processing.","fun_headline_variants_meta":{"raw":{"variants":["ALLaM-34B scores 4.92/5 on Arabic code-switching and generation","Arabic model: generation 4.92/5, Levantine dialect 2.73/5","Code-switching and generation at 4.92/5 for ALLaM-34B","Levantine Arabic largest gap in ALLaM-34B test","ALLaM-34B Arabic code-switching scores 4.92/5"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000715,"raw_usage":{"total_tokens":3260,"prompt_tokens":1033,"completion_tokens":2227,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":649,"completion_tokens_details":{"reasoning_tokens":2112}},"tokens_in":649,"tokens_out":2227,"duration_ms":17563,"temperature":1.0,"reasoning_tokens":2112,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:04:18.240852+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 23 prompts against an open-weight ALLaM-34B deployment with the same three judges; if the category means and dialect heat map differ materially, the UI-level scores describe the service rather than the model. A cheaper check is to probe HUMAIN Chat with prompts engineered to expose another model's known fingerprint and see whether the responses match ALLaM-34B's documented behavior.","supporting_citations":[{"cited_title":"Gpt-5 is here","cited_arxiv_id":null,"evidence_quote":"Serves as one of the three judge models, GPT-5, that produced the accuracy, fluency, and safety scores."},{"cited_title":"Claude sonnet 4","cited_arxiv_id":null,"evidence_quote":"Serves as the Claude Sonnet-4 judge in the scoring pipeline."},{"cited_title":"Towards inclusive arabic llms: A culturally aligned benchmark in arabic large language model evaluation","cited_arxiv_id":null,"evidence_quote":"Motivates culturally aligned Arabic evaluation as the gap that the prompt pack is designed to fill."}],"review_version":1}