{"id":"a25441cf-c2ab-4429-8881-6fbafa0ebb5a","arxiv_id":"2505.02656","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A new manually diacritized benchmark of Arabic Wikipedia proper nouns with English glosses, plus a GPT-4o evaluation that reaches 73% exact-match accuracy.","lead":"Researchers built a new dataset of 3,000 Arabic proper nouns from Wikipedia with full vowel markings and English translations, then tested how well the GPT-4o model can recover the correct markings. The model gets about 73% right, showing the task is hard and the dataset is a useful benchmark for Arabic language tools.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The gold-standard claim hinges on labels that were seeded by GPT-4o and checked by a single annotator; if the IAA study inherits the same model-anchoring, the reported 73% benchmark is not an independent measure of task difficulty.","rationale":"The reader's weakest assumption matches the concern I would raise: the gold labels and the IAA were produced in a pipeline that starts from GPT-4o proposals, so the benchmark's headline accuracy is not an independent measurement. The paper is otherwise careful: the dataset is released, guidelines are detailed, and the few-shot examples are drawn from CP-SAMA, not from the test set, which avoids direct train/test leakage. My concern is specifically about label independence, not about the authors' diligence. I did not find a separate load-bearing issue in the evaluation: the use of a single model is a scope limitation the authors acknowledge, and the exact-match metric is standard for a resource paper. The proposed blind reannotation test is feasible because the dataset is public and the annotation protocol is well specified.","tokens_in":15922,"tokens_out":3997,"duration_ms":44769,"concrete_test":"Sample 400 entries stratified by entity class and Freeman-similarity bin. Have two new native Arabic linguists independently diacritize each entry from scratch with no GPT-4o suggestion, using the published guidelines; do not reveal the gold label until after both annotators commit. Compute blind-annotator agreement with the released gold and with each other. If blind-vs-gold agreement is close to the reported 92.4% (and blind human-human agreement is comparable), anchoring is not material and the benchmark stands. If blind-vs-gold agreement falls materially below 92.4% while blind human-human agreement remains high, the gold is anchored in the original GPT-4o proposals; in that case recompute the Table 8 few-shot accuracy against the blind labels, and use the drop as the corrected estimate of GPT-4o performance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"CP-WIKI-D3K's value as a gold standard depends on the gold diacritizations being correct in a way that is not biased toward the system being benchmarked. In Section 5.2, the primary annotator was given GPT-4o's diacritized proposal (after automatic post-processing) for every entry; Section 5.3 shows she changed only 909 of 3,362 proposals (27%), leaving 73% of gold labels identical to the model's output. The IAA study in Section 5.4 used 'the same annotation process' for the second annotator, so both annotators saw the same GPT-4o starting point. The 92.4% agreement therefore measures consistency under a shared anchor, not independent accuracy; a systematic skew toward GPT-4o's preferred vowelization would inflate both the agreement and the headline 73% exact-match score measured against that gold. This is not merely a theoretical worry: the authors' own Limitations section states that 'multiple correct variants may exist depending on regional, historical, or phonetic conventions,' which is exactly the regime where anchoring is hardest to detect, since deviations from the proposal can be rationalized as acceptable variation. Table 7's disagreements (Kasra vs. Sukun, Sukun vs. Damma, etc.) show the annotation decisions are not forced by the Arabic script. The load-bearing condition is therefore not the raw number of annotations but the independence of the gold labels from the model they are used to evaluate.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CP-WIKI-D3K, a dataset of 3,000 unique Arabic Wikipedia proper nouns (3,362 Arabic--English gloss pairs) with manual lemma-level diacritizations, English glosses, and detailed annotation guidelines. The authors report 92.4% inter-annotator agreement on 500 samples and benchmark GPT-4o under several prompting conditions, achieving 73.0% exact-match accuracy with few-shot Arabic+Gloss prompting. They analyze the interplay of frequency, phonological similarity, and accuracy, and provide an error taxonomy. The stated contribution is a publicly available gold-standard dataset for Arabic proper noun diacritization paired with transliteration, along with a baseline benchmark.","tokens_in":16227,"tokens_out":6763,"duration_ms":72296,"significance":"If the dataset is genuinely gold-standard, this is a valuable resource: it is the first publicly available benchmark of its size for Arabic proper noun diacritization with English glosses, and it ships annotation guidelines, automated well-formedness checks, and a detailed, honest error analysis. The paper also makes a clear methodological point about the difficulty of the task and the value of English-gloss information. However, the strength of the contribution depends on the gold labels being independent of the benchmarked model, which is currently not demonstrated.","major_comments":[{"comment":"The gold-standard claim rests on annotations that are not independent of the model being benchmarked. Section 5.2 states that the primary annotator received a GPT-4o proposal for every entry, and Section 5.3 reports that only 909 of 3,362 proposals (27%) were changed. Section 5.4 then describes the IAA study as using \"the same annotation process,\" meaning the second annotator also saw GPT-4o proposals. The 92.4% agreement therefore measures consistency under a shared anchor, not independent correctness, and the 73.0% exact-match accuracy in Table 8 is computed against gold labels of which 73% are verbatim GPT-4o outputs. This is a form of circularity: the benchmark score is inflated whenever the annotator accepts a plausible but non-unique GPT-4o vowelization, and the paper's own Limitations section concedes that multiple correct variants exist. I request that the authors: (a) run an additional IAA study on a random sample where the second annotator diacritizes from scratch without the GPT-4o proposal; (b) report the overlap between the final gold and the initial GPT-4o proposals, with an analysis of the 27% changes; or (c) if such validation is not feasible, reframe the resource as a human-verified corpus and temper the \"gold-standard\" claim accordingly.","section":"§5.2–§5.4"},{"comment":"The evaluation metric is not aligned with the paper's own acknowledgement of multiple valid diacritizations. The Limitations section states that \"multiple correct variants may exist depending on regional, historical, or phonetic conventions,\" and Section 6.4 says of the non-matching outputs that they \"are plausible and acceptable alternatives in most cases.\" Nevertheless, Table 8 reports exact-match accuracy as the headline result, and the Levenshtein distance used alongside it still treats the gold form as the only correct output. Under this metric, a valid regional variant is counted as an error, so the 73.0% figure does not cleanly measure task difficulty or model quality. The authors should either (a) have a human judge assess the acceptability of a sample of non-matching outputs and report an adjusted accuracy, (b) mark entries with multiple valid forms and allow multiple gold references, or (c) explicitly restrict the claim to \"exact match against the corpus convention\" rather than \"correct diacritization.\"","section":"§6.2, §6.4"}],"minor_comments":[{"comment":"The class percentages sum to over 100% (77.1+25.5+2.0 = 104.6 for CP-WIKI and 85.2+35.0+2.0 = 122.2 for CP-WIKI-D3K), which suggests the classes are not mutually exclusive; please state this explicitly and consider reporting a multi-label breakdown. The shift in the Name category from 25.5% to 35.0% is also larger than \"broadly similar\" suggests; please test or discuss whether the random sample is representative.","section":"Table 5"},{"comment":"The claim that \"99.45% of all entries had no diacritics\" is stated without specifying the dataset or the detection method; please clarify whether this refers to CP-WIKI or CP-WIKI-D3K and how diacritics were detected.","section":"Introduction"},{"comment":"The bullet list of postprocessing operations contains a typo (\"consoant\") and the final bullet, \"Mapping Non-Arabic Arabic-script letters,\" would benefit from concrete examples of the mapped letters; please add them for reproducibility.","section":"§5.2"},{"comment":"Reporting only a raw agreement percentage without a chance-corrected measure (e.g., Cohen's kappa) or a confidence interval makes the 92.4% figure harder to interpret; please add such statistics, and state how the 38 disagreements were resolved in the final gold.","section":"§5.4"},{"comment":"The few-shot examples are \"manually manipulated\" from CP-SAMA; please document exactly how many examples were altered and what the manipulations were, so that other researchers can replicate the prompt setup.","section":"§6.1"},{"comment":"The reported correlation of -0.95 between accuracy and edit distance is expected because the same reference is used for both; consider removing this correlation or rephrasing it as a sanity check rather than an empirical finding.","section":"§6.3"}],"recommendation":"major_revision","confidential_remarks":"The main concern is the circularity between the gold construction and the benchmarked model: the IAA study as designed cannot rule out anchoring, and the headline 73% accuracy is partly self-referential. I believe the authors can address this with a supplementary from-scratch annotation study or by carefully re-scoping the dataset's claims. The paper is transparent and well-written, and the resource is likely to be useful regardless of the eventual framing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things before you read it. The paper fills a real gap: it releases the first public benchmark for Arabic proper-noun diacritization paired with English glosses, with annotation guidelines and data. That is genuinely useful. But the gold standard is less independent than the term suggests. The primary annotator was given GPT-4o's diacritized proposal for every entry and changed only 27% of them; the second annotator in the IAA study used the same process, so the 92.4% agreement partly measures consistency under a shared anchor. The 73% GPT-4o accuracy is measured against labels it helped generate. This is a real soft spot, but the authors are transparent about it, and the resource is still usable if read as a human-vetted, model-assisted dataset rather than an unbiased gold standard.\n\nWhat the paper does well: the annotation guidelines address genuinely hard issues (consonant clusters, final ya, determiner/plural removal); the automated well-formedness checks are a nice touch; the error analysis is honest, including the observation that many mismatches are plausible alternative vowelizations, which means exact-match is conservative. The frequency/Freeman-similarity analysis adds useful context, and the comparison with prior work is fair.\n\nSoft spots beyond anchoring: only 500 of 3,362 entries were independently re-annotated, and disagreements were not adjudicated; the dataset is location-heavy (85%); and \"gold-standard\" overclaims for a task where the authors themselves acknowledge multiple valid diacritizations. I'd prefer \"manually curated reference.\"\n\nWho is this for? Arabic NLP researchers working on diacritization, transliteration, or name normalization. A reading group could productively discuss annotation bias. The paper deserves peer review. A serious referee should ask for a sample re-annotated without GPT-4o proposals and for variance across runs. If those hold, the resource is solid.","headline":"A genuinely useful Arabic NLP resource—a manually diacritized proper-noun dataset with English glosses—but the gold labels were seeded by the very model being benchmarked, so the headline accuracy and IAA should be read with that in mind.","tokens_in":692,"tokens_out":1682,"would_cite":true,"duration_ms":39500,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that Arabic Wikipedia proper-noun diacritization is a hard, undertooled task, and backs the claim with a manually diacritized 3,362-pair gold dataset on which GPT-4o attains 73.0% exact-match accuracy.","keywords":["Arabic diacritization","proper nouns","transliteration","Arabic Wikipedia","benchmark dataset","GPT-4o","lemmatization","named entities"],"falsifier":"Have a fresh annotator independently diacritize a random 500-entry sample from the raw Arabic-plus-gloss inputs with no model suggestion, then measure agreement with the published gold labels; if agreement falls materially below the reported 92.4%, the gold standard is partly anchored to the initial GPT-4o proposals.","tokens_in":15749,"feed_emoji":"🔤","tokens_out":3933,"duration_ms":46474,"temperature":0.7,"pith_summary":"The paper tries to establish that diacritizing Arabic Wikipedia proper nouns can be turned into a measurable benchmark task by pairing each undiacritized Arabic name with its English Wikipedia gloss, and that even a strong large language model falls well short of mastery. It introduces a gold-standard dataset of 3,000 unique Arabic proper nouns, expanded to 3,362 Arabic-gloss pairs, with manual lemma-level diacritizations and English equivalents. Benchmarking GPT-4o on recovering the full diacritization, the paper reports 73.0% exact-match accuracy in the best setting, arguing that the task is genuinely difficult and that better resources and models are needed. If the claim is right, the field gains the first public benchmark of this scale for the intersection of Arabic diacritization, transliteration, and proper-noun lemmatization.","feed_headline":"Arabic name diacritization stalls at 73% even with English hints","feed_subtitle":"New 3,362-pair gold dataset shows GPT-4o still misses many proper-noun vowels; glosses help but don't solve it.","key_machinery":"The load-bearing object is the annotated dataset CP-WIKI-D3K, which pairs undiacritized Arabic proper nouns with English glosses and gold lemma-level diacritizations. The task is framed as a mapping from (Arabic input, English gloss) to a fully diacritized Arabic lemma, guided by an explicit annotation scheme: retain input spelling except for required Hamza corrections, allow consonant clusters in foreign names, and remove definite articles and plural suffixes. The evaluation machinery is exact-match accuracy plus Levenshtein edit distance, applied after a post-processing pipeline that enforces well-formedness rules such as inserting Fatha before Alif, normalizing Shadda-vowel order, and removing final diacritics.","core_discovery":"The central claim is that CP-WIKI-D3K, a randomly sampled subset of 3,000 unique Arabic-script proper nouns from Arabic Wikipedia paired with 3,362 English glosses, is a reliable gold-standard resource for the joint task of diacritizing and lemmatizing proper nouns. Each entry was annotated with a fully diacritized Arabic lemma following a maximal-diacritization scheme adapted to foreign names, with clitics such as the definite article removed and plural demonym endings stripped. The paper further claims that GPT-4o, when given the undiacritized Arabic and the English gloss plus a few examples, reaches only 73.0% exact-match accuracy, demonstrating that the task resists current models. Error analysis shows most failures are plausible alternative diacritizations rather than consonant errors, with systematic overuse of Fatha and Shadda and underuse of Kasra and Sukun.","pith_inferences":["A multi-reference or human-acceptability evaluation would likely raise reported performance well above 73%, since the paper itself shows many mismatches are plausible variants.","The same Arabic-English pair structure could support fine-tuned smaller models, which may surpass prompt-only GPT-4o at lower cost, though the paper does not test this.","The systematic overuse of Fatha and Shadda points to a concrete modeling fix: training or prompting that explicitly penalizes gemination and favors Kasra/Sukun in foreign names.","The dataset could double as a testbed for transliteration evaluation, since the gloss provides a Romanization target that constrains the Arabic vowelization."],"forward_implications":["Adding the English gloss substantially helps: few-shot accuracy rises from 55.9% with Arabic only to 73.0% with Arabic plus gloss.","Frequency is a strong predictor: accuracy climbs from 64.8% in the lowest-frequency quartile to 79.7% in the highest.","Phonetically dissimilar but frequent names, such as 'Egypt', are still predicted accurately, so pure transliteration similarity is not the only useful signal.","Many of the remaining errors are defensible alternative diacritizations, implying that a single-reference exact-match score underestimates the model's linguistic plausibility."],"supporting_citations":[{"why":"Supplies the CP-WIKI proper-noun list from which the 3,000 entries were sampled.","marker":"Khairallah et al. (2024)"},{"why":"Provides the maximal Arabic diacritization guidelines that the annotation scheme adapts.","marker":"Elgamal et al. (2024)"},{"why":"Is the GPT-4o model that the paper benchmarks on the new dataset.","marker":"OpenAI et al. (2024)"},{"why":"Supplies the Freeman similarity score used to measure Arabic-English phonological similarity in the analysis.","marker":"Freeman et al. (2006)"},{"why":"Supplies the edit-distance metric used alongside exact-match accuracy.","marker":"Levenshtein (1966)"},{"why":"Supplies the Arabic frequency list used to bin entries by frequency and analyze accuracy.","marker":"Khalifa et al. (2021)"},{"why":"Earlier work at the intersection of diacritization and transliteration that this dataset extends.","marker":"Mubarak et al. (2009)"},{"why":"Prior automatic diacritization of transliterated words with much smaller training and test sets, providing the comparison this resource improves on.","marker":"Darwish et al. (2017)"}],"fun_headline_variants":["Arabic name diacritization: GPT-4o only 73% on new gold dataset","New benchmark: 3K Arabic proper nouns, GPT-4o hits 73%","GPT-4o fails 27% of Arabic proper noun diacritizations","Arabic Wikipedia names: gold dataset reveals GPT-4o limits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The gold labels are assumed to be correct and unbiased, but they rest on a single primary annotator who edited GPT-4o-generated proposals, with only 500 of 3,362 entries independently re-annotated.","fun_headline_variants_meta":{"raw":{"variants":["Arabic name diacritization: GPT-4o only 73% on new gold dataset","New benchmark: 3K Arabic proper nouns, GPT-4o hits 73%","GPT-4o fails 27% of Arabic proper noun diacritizations","Arabic Wikipedia names: gold dataset reveals GPT-4o limits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000442,"raw_usage":{"total_tokens":2210,"prompt_tokens":886,"completion_tokens":1324,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":502,"completion_tokens_details":{"reasoning_tokens":1236}},"tokens_in":502,"tokens_out":1324,"duration_ms":12083,"temperature":1.0,"reasoning_tokens":1236,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:44:53.421669+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a fresh annotator independently diacritize a random 500-entry sample from the raw Arabic-plus-gloss inputs with no model suggestion, then measure agreement with the published gold labels; if agreement falls materially below the reported 92.4%, the gold standard is partly anchored to the initial GPT-4o proposals.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the CP-WIKI proper-noun list from which the 3,000 entries were sampled."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Freeman similarity score used to measure Arabic-English phonological similarity in the analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the Arabic frequency list used to bin entries by frequency and analyze accuracy."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Earlier work at the intersection of diacritization and transliteration that this dataset extends."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Prior automatic diacritization of transliterated words with much smaller training and test sets, providing the comparison this resource improves on."}],"review_version":1}