{"id":"b7700189-1efa-4c01-9dca-ca2cd217783b","arxiv_id":"2508.05722","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A manually aligned English-Arabic parallel corpus of 51,671 healthcare sentences is presented.","lead":"The paper introduces PEACH, a new sentence-aligned parallel English-Arabic corpus of healthcare texts with more than 51,000 sentence pairs. It aims to support machine translation, linguistics, and patient-information research in a domain where such resources are scarce.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Gold-standard alignment quality is asserted but unsupported; without inter-annotator agreement, a sampling quality test, or a documented alignment protocol, the central quality claim cannot be evaluated, and the corpus's downstream utility remains unverified.","rationale":"The reader's weakest-assumption analysis identified the absence of alignment-quality evidence as the most load-bearing issue, and I agree. The strongest claim is about the corpus's usefulness as a gold standard, not merely its existence or size. That usefulness is directly conditioned on alignment accuracy, which is asserted but not demonstrated. The available text only gives aggregate numbers and the phrase 'manually aligned.' Without an inter-annotator agreement study or a documented protocol, the label 'gold-standard' is an unsupported quality claim. The full text is corrupted, so we cannot rule out that such details exist, but on the evidence present the central quality claim is unverified. I do not see internal inconsistencies in the reported counts: the average English sentence length (590,517/51,671 ≈ 11.4) and Arabic (567,707/51,671 ≈ 11.0) both fall within the stated 9.52–11.83 range, which slightly increases confidence in the descriptive statistics. However, the alignment-quality issue remains the decisive unknown. Therefore, the reader's UNVERDICTED verdict is appropriate, and my review does not change it.","tokens_in":2082,"tokens_out":2557,"duration_ms":26586,"concrete_test":"Obtain a random sample of 500 sentence pairs from the released corpus, have two independent bilingual Arabic-English annotators re-align the same source documents, and compute inter-annotator agreement (e.g., Cohen's kappa on exact sentence-pair matching). If kappa is below 0.9, the 'gold-standard' characterization is not supported and the corpus quality claim should be revised. If kappa is high and the sample is representative, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that PEACH is a 'gold-standard,' manually aligned parallel English-Arabic healthcare corpus. The only evidence for 'gold-standard' in the available text (the abstract) is the statement that the corpus is manually aligned. No inter-annotator agreement measure, no random-sample quality check, no alignment protocol, and no exclusion criteria are reported. All listed downstream applications—bilingual lexicon induction, domain-specific machine translation, readability assessment—depend on the correctness of sentence-level alignments. If the alignment error rate is non-negligible, the corpus's value as a reference standard drops sharply, regardless of its size. The abstract also states the corpus is publicly accessible, but no URL or repository identifier appears in the provided text, so even basic availability cannot be confirmed from the preprint. The full text is corrupted, preventing a check for later sections that may contain such details. None of this demonstrates the claim is false; it shows that a core quality assertion currently rests on a single unoperationalized word, 'manually aligned,' without verifiable evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces PEACH, a sentence-aligned parallel English-Arabic healthcare corpus built from patient information leaflets and educational materials. The abstract reports 51,671 parallel sentences, approximately 590,517 English and 567,707 Arabic tokens, and mean sentence lengths between 9.52 and 11.83 words. It claims the corpus is manually aligned and therefore a gold-standard resource, publicly accessible, and intended for applications such as bilingual lexicon induction, domain-specific machine translation, healthcare MT evaluation, readability assessment, and translator education. The submitted full text is largely unreadable because it is corrupted and ends with an unrelated arXiv identifier, so the technical sections describing construction and evaluation cannot be reviewed.","tokens_in":2329,"tokens_out":3660,"duration_ms":35332,"significance":"If the size and alignment-quality claims hold, PEACH would be a valuable resource for English-Arabic healthcare NLP, filling a domain-specific gap in parallel corpora. The reported counts are concrete and presumably checkable, but the corpus's impact depends on alignment accuracy. The paper names plausible downstream applications and makes an availability claim; with a documented protocol and a quality evaluation, it could support reproducible research in low-resource and healthcare translation. The current significance is conditional on evidence that is not yet presented.","major_comments":[{"comment":"The submitted full text is corrupted: it consists of garbled characters and ends with the unrelated identifier 'arXiv:2508.05736v1 [quant-ph] 7 Aug 2025', not the paper's own identifier (2508.05722). The technical sections describing corpus construction, alignment protocol, annotation, and evaluation cannot be reviewed. This is load-bearing: the central claims about manual alignment and gold-standard quality are unverifiable in this submission.","section":"Full text"},{"comment":"The 'gold-standard' claim is asserted solely from the phrase 'manually aligned'. No inter-annotator agreement, sample-based quality inspection, alignment protocol, or exclusion criteria are reported; if such details appear later, they are inaccessible in the corrupted full text. All listed downstream applications are sensitive to alignment noise, so a documented protocol and a quantitative quality check (e.g., random-sample accuracy or IAA) are necessary, or the 'gold-standard' label should be moderated.","section":"Abstract"},{"comment":"The corpus is said to be 'publicly accessible', but no URL, repository identifier, license, or version information is provided. For a dataset paper, this information is essential for reproducibility and reuse.","section":"Abstract"}],"minor_comments":[{"comment":"The reported average sentence lengths 'between 9.52 and 11.83 words' should specify whether these are per-language corpus means or per-document averages, and include standard deviations.","section":"Abstract"},{"comment":"The token counts (590,517 English; 567,707 Arabic) should state the tokenization method; for Arabic, whitespace tokenization versus morphological segmentation can change counts substantially.","section":"Abstract"},{"comment":"The extraneous arXiv identifier at the end of the full text must be removed; if it is a conversion artifact, the authors should verify that the submitted PDF/text matches the intended paper.","section":"Full text"}],"recommendation":"major_revision","confidential_remarks":"The corruption of the full text and the mismatched arXiv ID make it impossible to review the technical content. I assume this is a submission or parsing error rather than a deliberate issue, but the editor may wish to confirm with the authors that the uploaded file is the correct, complete manuscript before sending it back. The quality-evidence gap in the abstract is independent of that issue and needs to be addressed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is a 51k-sentence parallel English-Arabic corpus focused on healthcare, which fills a real gap for low-resource MT and health literacy work. The domain choice (patient leaflets and educational materials) is sensible, and the reported counts are straightforward. If the corpus is actually public, it could be a solid resource for lexicon induction, domain adaptation, and readability research.\n\nThat said, the central quality claim collapses under scrutiny. The abstract calls PEACH a \"gold-standard corpus\" solely because it is \"manually aligned.\" Manual alignment is a necessary condition, but not sufficient, for gold-standard status. There is no inter-annotator agreement, no random-sample quality check, no documented alignment protocol, and no exclusion criteria. Without that, the word \"gold-standard\" is doing all the work. The abstract also says the corpus is publicly accessible but gives no URL or repository ID, so even basic availability is unverified. On top of that, the arXiv full text is corrupted mojibake, which isn't the authors' fault but means I could not check whether the full paper contains the missing methodology. As it stands, the abstract alone does not support the quality claim.\n\nOn the positive side, the paper does not overreach in its list of downstream applications; those are realistic. The sentence-length statistics (9.52–11.83 words) suggest a homogeneous text type, which is useful information. And this is a resource paper, not a theory paper, so the absence of a novel method is fine.\n\nI agree with the stress-test note that the gold-standard label is the load-bearing issue. I disagree with framing this as a fatal flaw; a corpus can still be useful even if the alignment quality is not yet proven. But the omitted evidence matters, and the authors should be pushed to provide it.\n\nThis paper is for researchers working on English-Arabic MT or health communication. It deserves a serious referee if the full text includes a proper data statement, alignment protocol, and availability link. As-is, the abstract alone is too thin for a confident verdict. I would send it to peer review because domain-specific parallel corpora for low-resource language pairs are under-supplied, and the authors are doing needed work. I would require them to substantiate the gold-standard claim and to fix the corrupted upload.","headline":"A potentially useful English-Arabic healthcare corpus, but the gold-standard claim is asserted, not demonstrated, and the full text is unreadable.","tokens_in":2753,"tokens_out":2566,"would_cite":false,"duration_ms":27530,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A new manually aligned English–Arabic healthcare corpus, PEACH, offers 51,671 parallel sentences for translation research.","keywords":["parallel corpus","English-Arabic","healthcare","sentence alignment","machine translation","patient information leaflets","gold standard","NLP"],"falsifier":"Have two independent bilingual annotators re-align a random sample of roughly 500 sentence pairs from PEACH and measure their agreement rate; if agreement falls well short of standard gold-corpus thresholds, or if a substantial fraction of pairs are found to be misaligned, the gold-standard claim is not supported.","tokens_in":1996,"feed_emoji":"🩺","tokens_out":1926,"duration_ms":21372,"temperature":0.7,"pith_summary":"This paper introduces PEACH, a sentence-aligned parallel corpus of English–Arabic healthcare texts drawn from patient information leaflets and educational materials. The corpus contains 51,671 parallel sentences, roughly 590,517 English tokens and 567,707 Arabic tokens, with average sentence lengths between 9.52 and 11.83 words. The author's central claim is that because alignment was done manually, PEACH qualifies as a gold-standard resource, one that can support contrastive linguistics, translation studies, and natural language processing. A sympathetic reader would care because high-quality, domain-specific parallel data is scarce for English–Arabic, and healthcare is a high-stakes domain where machine translation errors can have real consequences.","feed_headline":"New English-Arabic healthcare corpus pairs 51,671 sentences","feed_subtitle":"Manually aligned patient leaflets and education text support translation research and domain-specific machine translation.","key_machinery":"The central mechanism is manual sentence alignment: pairs of English and Arabic sentences from healthcare documents are aligned by hand, which the paper treats as the defining guarantee of corpus quality. This alignment is what elevates PEACH from a raw parallel text collection to a gold-standard resource for tasks that depend on reliable translation correspondences.","core_discovery":"PEACH is a publicly available, sentence-aligned parallel corpus built from healthcare texts, specifically patient information leaflets and educational materials, in English and Arabic. The paper reports its size as 51,671 parallel sentences, about 590,517 English word tokens and 567,707 Arabic word tokens, with average sentence lengths ranging from 9.52 to 11.83 words. The corpus is described as manually aligned, a property the author takes to make it a gold-standard resource. The paper argues that this resource enables downstream uses including bilingual lexicon extraction, domain adaptation of large language models for machine translation, evaluation of user perceptions of machine translat","pith_inferences":["The paper's 'gold-standard' designation rests on the manual alignment procedure alone; without reported inter-annotator agreement or quality sampling, the label is an assertion to be tested rather than a demonstrated property.","Because PEACH is relatively small (around 51k sentences) and domain-specific, its main value may lie in fine-tuning or evaluation rather than in training large general-purpose models from scratch.","A natural next step would be to publish alignment guidelines and an inter-annotator agreement score, which would strengthen the corpus's credibility as a gold standard for future benchmarking.","The healthcare domain makes this corpus especially relevant for studying the consequences of translation errors, since mistranslations in patient instructions can carry clinical risk; this suggests downstream work should pair PEACH with user-centered safety evaluations."],"forward_implications":["PEACH can be used to derive bilingual healthcare lexicons, particularly for terminology in patient leaflets and educational materials.","The corpus supports domain adaptation of machine translation systems, potentially improving English–Arabic translation quality in medical settings.","Researchers can use PEACH to evaluate how end users perceive machine-translated healthcare content, including safety and comprehension.","The corpus enables readability and lay-friendliness assessments of patient information leaflets by comparing English and Arabic versions.","As an openly accessible, manually aligned resource, PEACH can serve as a training and evaluation set in translation studies curricula."],"supporting_citations":[],"fun_headline_variants":["51,671 healthcare sentences aligned in English-Arabic corpus","PEACH: gold-standard English-Arabic health corpus with 51k pairs","Manually aligned English-Arabic healthcare corpus hits 51,671 sentences","New bilingual health corpus: English-Arabic, 51,671 sentence pairs","PEACH corpus: 51,671 aligned English-Arabic healthcare sentences"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The corpus's value as a gold standard rests on the assumption that every one of the 51,671 sentence pairs was aligned accurately by hand, yet the paper provides no inter-annotator agreement, quality sampling, or detailed alignment protocol to verify that consistency.","fun_headline_variants_meta":{"raw":{"variants":["51,671 healthcare sentences aligned in English-Arabic corpus","PEACH: gold-standard English-Arabic health corpus with 51k pairs","Manually aligned English-Arabic healthcare corpus hits 51,671 sentences","New bilingual health corpus: English-Arabic, 51,671 sentence pairs","PEACH corpus: 51,671 aligned English-Arabic healthcare sentences"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000204,"raw_usage":{"total_tokens":1175,"prompt_tokens":644,"completion_tokens":531,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":388,"completion_tokens_details":{"reasoning_tokens":441}},"tokens_in":388,"tokens_out":531,"duration_ms":4535,"temperature":1.0,"reasoning_tokens":441,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:17:38.966489+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have two independent bilingual annotators re-align a random sample of roughly 500 sentence pairs from PEACH and measure their agreement rate; if agreement falls well short of standard gold-corpus thresholds, or if a substantial fraction of pairs are found to be misaligned, the gold-standard claim is not supported.","supporting_citations":[],"review_version":1}