{"id":"e8e2bb9b-2344-4d38-bb8e-27b3a5d55f75","arxiv_id":"2504.18582","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":7,"one_line_summary":"Fine-tuning Wav2Vec 2.0 on a custom Kurdish corpus is reported to cut speaker diarization error by 7.2 percentage points and raise cluster purity by about 13 percentage points, though the paper contains conflicting numbers.","lead":"This paper fine-tunes a pre-trained Wav2Vec speech model on a small Kurdish audio corpus to identify who speaks when in recordings. It reports lower diarization error and purer speaker clusters versus a non-fine-tuned baseline, but the paper's own results are inconsistent.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Two mutually exclusive result sets in the same paper leave the 7.2-pp DER improvement unsupported; the experimental protocol must be reconciled before the claim can be assessed.","rationale":"I read the paper as a straightforward empirical claim: applying transfer learning through Wav2Vec 2.0 fine-tuning improves Kurdish speaker diarization. The evidence for that claim is Table 1 and Table 3, but the same paper contains a second, mutually inconsistent set of results in Section 4.2. This is the most load-bearing problem because it concerns the exact number being claimed. Reference errors and missing artifacts compound the problem, but they would not by themselves invalidate a clean experiment. If the Section 4.2 numbers are real, then Table 1 is wrong; if Table 1 is real, then Section 4.2 is wrong. There is also an unstated assumption that the baseline was generated by the same diarization pipeline, which the paper never specifies. I would not accept the claim as reproducible without reconciling these two result sets and describing the baseline protocol. This strengthens, rather than changes, the reader's REJECT verdict.","tokens_in":15723,"tokens_out":5620,"duration_ms":50320,"concrete_test":"Obtain, or require the authors to release, the raw per-run metric logs for the five independent runs and the exact segmentation, clustering, and scoring commands used for both models. Recompute the means for DER, cluster purity, and SNR from those logs on the same test split described in Section 4.1. If the recomputed means match Table 1 (22.8/15.6, 76.4/89.1, 12.5/18.7), the contradiction is a Section 4.2 reporting error and the claim can be assessed on Table 1 alone; if they match Section 4.2 (15.2/8.0, 78.5/85.3, 10.4/14.7), Table 1 and Table 3 are invalid. Without the logs, no unique result set exists to verify.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim depends entirely on Table 1: DER 22.8% to 15.6%, cluster purity 76.4% to 89.1%, and SNR 12.5 dB to 18.7 dB. Section 4.2 reports a different evaluation of the same comparison: DER fell by 50.3% from 15.2% to 8.0%, with cluster purity 78.5% to 85.3% and SNR 10.4 dB to 14.7 dB. Both result sets have the same 7.2-percentage-point absolute DER gap, so the abstract's 7.2% reduction could come from either version; the paper therefore does not pin down a single empirical effect. The comparison also presupposes that the pre-trained baseline was scored through the same segmentation, clustering, and evaluation pipeline as the fine-tuned model, but no such pipeline is described anywhere; Section 4.2 only recounts the fine-tuning procedure. Since an unstated change in the baseline pipeline could itself produce a 7.2-point gap, the central claim is not uniquely supported as written.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes fine-tuning Wav2Vec 2.0 on a custom Kurdish audio corpus for speaker diarization, reporting improvements in Diarization Error Rate (DER), cluster purity, and Signal-to-Noise Ratio (SNR) relative to a pre-trained Wav2Vec baseline. It describes the dataset, preprocessing and augmentation steps, a dual cross-entropy/CTC fine-tuning objective, hyperparameter choices, and results for full, few-shot, one-shot, and zero-shot settings. The central claim is that fine-tuning reduces DER by 7.2 percentage points and improves cluster purity by about 13 percentage points, with implications for low-resource speech technology.","tokens_in":15948,"tokens_out":6137,"duration_ms":55689,"significance":"If the headline results were reproducible, the paper would address a genuine gap: speaker diarization for Kurdish is understudied, and a dedicated Kurdish corpus plus a transfer-learning recipe would be a useful community resource. The paper also makes a plausible case that data augmentation and fine-tuning can help in low-resource settings. However, the current manuscript does not support the central claim: the evaluation is internally contradictory, the baseline diarization pipeline is unspecified, and no code or data are released. These issues prevent external verification and make the reported effect size uninterpretable as written.","major_comments":[{"comment":"The paper reports two irreconcilable sets of results for the same comparison. Table 1 gives baseline DER 22.8% and fine-tuned DER 15.6%, cluster purity 76.4% and 89.1%, and SNR 12.5 dB and 18.7 dB. Section 4.2 and Figure 8 report baseline DER 15.2% and fine-tuned DER 8.0%, cluster purity 78.5% and 85.3%, and SNR 10.4 dB and 14.7 dB. Both sets share the same 7.2-percentage-point DER gap, so the abstract's '7.2%' does not uniquely refer to either experiment. The authors must identify which numbers are the actual evaluation and provide a single consistent set of results; as written, the central empirical claim is unsupported.","section":"§4.2, §4.3, Table 1"},{"comment":"The baseline to which the fine-tuned model is compared is never defined as a diarization system. A Wav2Vec encoder alone does not produce speaker diarization; one must specify voice activity detection, speaker segmentation, embedding extraction, clustering, handling of the number of speakers, and the scoring tool. None of these steps is described for either the baseline or the fine-tuned model. Without this information, the reported DER drop cannot be attributed to fine-tuning rather than to a difference in the evaluation pipeline.","section":"§3.3, §4.2"},{"comment":"The claimed statistical significance is not substantiated. Section 4.1 states that 'the exact p-value of less than 0.05 was obtained for all analysed parameters' and that the DER decrease was 'confirmed statistically different in five independent runs, SD = ±0.5%', but no test statistic, degrees of freedom, confidence intervals, or baseline variance are reported. Table 1 provides only a single standard-deviation column without clarifying whether it applies to both conditions or how it was computed; this is insufficient support for an inferential claim.","section":"§4.1, Table 1"},{"comment":"The metric definitions are too incomplete to verify the reported improvements. DER is defined with a simple sum over total speech time, but standard diarization scoring requires a collar, forgiveness, and overlap handling; cluster purity is described only as 'the proportion of correctly grouped segments to the total number of segments', which presupposes the clustering that is never specified. SNR is an audio-quality metric, not a diarization metric, so the SNR gain in Table 1 reflects preprocessing rather than fine-tuning and should not be presented as part of the model's diarization improvement.","section":"§3.3.2, Table 1"},{"comment":"The evaluation rests entirely on a custom, unreleased Kurdish corpus whose composition is not described in sufficient detail. Section 3.1.1 gives only folder counts and a qualitative list of limitations; there is no information on total duration, exact number of speakers per file beyond folder labels, dialect distribution, speaker demographics, or annotation protocol, and the dataset is not made available. Because every result is obtained on this private test set, the findings cannot be checked or compared with any external benchmark.","section":"§3.1.1, §4.1"}],"minor_comments":[{"comment":"The objectives list refers to 'Improve Word2Vec for Kurdish' and the text later alternates between Word2Vec and Wav2Vec; the paper is about Wav2Vec 2.0, so the terminology should be corrected throughout.","section":"§1"},{"comment":"The relative DER decrease from 15.2% to 8.0% is 47.4%, not 50.3% as stated; please recalculate and report the correct value.","section":"§4.2"},{"comment":"Table 3 contains unexplained columns (LP, IDA, RI) and an RI value of 0.0 for the baseline; these should be defined or removed.","section":"Table 3"},{"comment":"There are duplicated subsection headings labeled 2.2.3, and reference [26] does not appear to support the claim about low-resource diarization; the references should be verified.","section":"§2.2.3"},{"comment":"Several passages appear unfinished or out of place, such as the sentence beginning 'A logo is a provisional description'; the manuscript needs careful editing.","section":"§3.1.1"}],"recommendation":"reject","confidential_remarks":"The topic is suitable for the journal and the corpus-building effort is potentially valuable, but the current manuscript is not a coherent experimental report. The two result sets are mutually exclusive, the baseline pipeline is missing, and the dataset is unavailable, so the headline comparison is unverifiable. I would encourage the authors to resubmit after a complete rewrite that reports one consistent evaluation and releases the corpus or at least a reproducible subset."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is the target language: no one has built a Kurdish speaker diarization benchmark, and the paper addresses a real gap. The authors assembled a purpose-built corpus (269 files, up to four speakers), document their augmentation and hyperparameter choices, and honestly list dataset limitations such as dialect coverage and overlapping speech. That is useful scaffolding for the community.\n\nThe problem is the experimental reporting. The paper contains two mutually exclusive evaluations of the same comparison. The prose around Figure 8 says baseline DER 15.2%, fine-tuned 8.0%; the table labeled Table 1 in Section 4.3 says 22.8% and 15.6%. Both have the same 7.2-point absolute gap, which is suspicious. Cluster purity and SNR also contradict each other: 76.4%/89.1% and 12.5/18.7 dB versus 78.5%/85.3% and 10.4/14.7 dB. So the paper does not pin down a single empirical effect. The abstract's 7.2% reduction could be read from either version, but the supporting numbers disagree.\n\nThe baseline pipeline is also undefined. We never learn what segmentation, embedding, clustering, or scoring was used for the pre-trained Wav2Vec baseline. If the baseline ran through a weaker pipeline, the improvement is meaningless. The claim of p<0.05 over five runs has no test statistic, only standard deviations. The dataset and code are not released, so there is no way to check.\n\nThe citation pattern is sloppy: [19] is a medical imaging paper, [26] is about dollar store vaccine distribution. There are duplicated section numbers and a stray line about a \"logo\" that looks like an editing artifact. These are not fatal by themselves, but they reinforce the impression that the manuscript was not carefully checked.\n\nTo be fair, the core idea is plausible: fine-tuning a self-supervised multilingual model on Kurdish should help. The authors deserve credit for identifying the gap and building a first corpus. But the central claim is load-bearing, and the paper's own data does not uniquely support it. This needs a major revision: reconcile the numbers, specify the baseline pipeline, and release the dataset or at least a detailed experimental protocol.\n\nWho is this for? Researchers working on Kurdish speech technology and low-resource diarization. They would want to know about the corpus and the approach. But as a citable result, it is not trustworthy in current form. A serious editor should send it to peer review rather than desk reject, because the reported gap is large and the language is genuinely underserved; a reviewer can demand the fixes. I would not accept it as is, and I would not cite it until the numbers are reconciled.","headline":"First Wav2Vec 2.0 fine-tuning for Kurdish diarization with a purpose-built corpus, but two contradictory result sets make the headline 7.2% DER improvement unsupported as written.","tokens_in":16522,"tokens_out":2603,"would_cite":false,"duration_ms":25381,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fine-tuning Wav2Vec 2.0 on a small Kurdish corpus cuts diarization error by 7.2 points.","keywords":["Speaker Diarization","Kurdish Speech Processing","Wav2Vec 2.0","Self-Supervised Learning","Transfer Learning","Low-Resource Languages","Diarization Error Rate","Cluster Purity"],"falsifier":"Run the pre-trained and fine-tuned Wav2Vec models through one fixed diarization pipeline—same segmentation, same clustering, same scoring tool—on a held-out Kurdish test set. If the pre-trained baseline scores near the paper's other reported baseline value of 15.2% DER rather than 22.8%, the claimed 7.2-point improvement evaporates; a score near 22.8% would support the claim.","tokens_in":15530,"feed_emoji":"🎙️","tokens_out":6258,"duration_ms":57567,"temperature":0.7,"pith_summary":"Speaker diarization—splitting an audio stream by who is speaking—works well for high-resource languages but has been barely studied for Kurdish, which lacks labeled corpora and spans several dialects. This paper tries to close that gap by taking a pre-trained Wav2Vec 2.0 model and fine-tuning it on a small, purpose-built Kurdish audio corpus. On the paper's reported five-run evaluation, fine-tuning lowers the Diarization Error Rate from 22.8% to 15.6% (a 7.2-point drop) and raises cluster purity from 76.4% to 89.1% (a 12.7-point gain). If those numbers hold, the same transfer-learning recipe could make speaker separation practical for other under-resourced languages.","feed_headline":"Fine-tuning Wav2Vec cuts Kurdish diarization errors by 7.2 points","feed_subtitle":"A small labeled Kurdish corpus is enough to sharply improve speaker separation for a low-resource language.","key_machinery":"The load-bearing object is Wav2Vec 2.0, a self-supervised speech encoder that learns representations from raw audio using a convolutional network plus a transformer context network. The paper fine-tunes it with a dual loss: cross-entropy for speaker classification and connectionist temporal classification (CTC) for aligning speaker transitions. Data preprocessing—noise reduction, normalization, segmentation, and augmentation with synthetic noise, pitch shifts, and speed changes—feeds the model cleaner Kurdish speech.","core_discovery":"The paper's claim is that a self-supervised multilingual speech model, Wav2Vec 2.0, can be adapted to Kurdish speaker diarization with only a small curated dataset. The argument is that the pre-trained model already encodes general acoustic and phonetic structure, and fine-tuning on Kurdish audio with two losses—cross-entropy for assigning segments to speakers and CTC loss for aligning speaker boundaries in time—shifts those representations toward Kurdish. The reported evidence shows DER falling from 22.8% to 15.6%, cluster purity rising from 76.4% to 89.1%, and SNR rising from 12.5 dB to 18.7 dB, with standard deviations of ±0.5%, ±0.7%, and ±0.3 dB across five runs. The paper takes these gains as showing that transfer learning plus data augmentation can overcome the absence of large labeled Kurdish diarization datasets.","pith_inferences":["The paper does not isolate how much of the gain comes from fine-tuning versus preprocessing; an ablation that keeps the audio pipeline fixed while varying only the model weights would separate those contributions.","The same recipe could be tested on other under-resourced languages with dialect variation and code-switching, such as comparing Sorani versus Kurmanji audio, to see whether transfer learning generalizes beyond Kurdish.","A useful extension would measure how DER and cluster purity scale with the size of the labeled Kurdish corpus; the few-shot/one-shot/zero-shot table suggests the curve may be steep, which would tell practitioners how much annotation is worth paying for."],"forward_implications":["If the reported gains replicate, Kurdish media transcription, meeting analysis, and call-center speaker separation can be built without waiting for large annotated Kurdish corpora.","The combination of Wav2Vec 2.0 pretraining, a small curated dataset, and dual-loss fine-tuning becomes a template for other low-resource languages whose phonetics differ from the pretraining data.","Full fine-tuning outperforms few-shot, one-shot, and zero-shot variants, but all tuned variants beat the untuned baseline, so even a handful of labeled Kurdish examples appears to help.","The simultaneous drop in DER and rise in cluster purity implies the fine-tuned model produces both fewer wrong speaker labels and more coherent speaker clusters, not just a trade-off between the two."],"supporting_citations":[{"why":"Supplies the Wav2Vec 2.0 architecture that the paper fine-tunes.","marker":"[5]"},{"why":"Introduces the original Wav2Vec pretraining objective that the approach builds on.","marker":"[6]"},{"why":"Shows cross-lingual pretraining plus fine-tuning works for low-resource speech, the direct precedent for transferring to Kurdish.","marker":"[34]"},{"why":"Defines the transfer-learning mechanism the paper relies on to adapt multilingual representations.","marker":"[35]"},{"why":"Provides the CTC loss used to align speaker transitions during fine-tuning.","marker":"[45]"},{"why":"Supplies dropout regularization used to avoid overfitting on the small Kurdish corpus.","marker":"[47]"},{"why":"Defines the diarization error rate metric used for the headline comparison.","marker":"[15]"}],"fun_headline_variants":["Wav2Vec fine-tuning slashes Kurdish diarization error by 7.2%","Kurdish speaker separation boosted 13% via Wav2Vec transfer","Low-resource diarization improved: Wav2Vec fine-tune on Kurdish","Wav2Vec adapts to Kurdish, cutting diarization errors 7.2 points","Fine-tuned Wav2Vec lifts Kurdish diarization accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim depends on the pre-trained baseline being scored through the same diarization pipeline and test set as the fine-tuned model; if the baseline pipeline was weaker, the 7.2-point drop is not caused by fine-tuning.","fun_headline_variants_meta":{"raw":{"variants":["Wav2Vec fine-tuning slashes Kurdish diarization error by 7.2%","Kurdish speaker separation boosted 13% via Wav2Vec transfer","Low-resource diarization improved: Wav2Vec fine-tune on Kurdish","Wav2Vec adapts to Kurdish, cutting diarization errors 7.2 points","Fine-tuned Wav2Vec lifts Kurdish diarization accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00091,"raw_usage":{"total_tokens":3902,"prompt_tokens":929,"completion_tokens":2973,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":2864}},"tokens_in":545,"tokens_out":2973,"duration_ms":18996,"temperature":1.0,"reasoning_tokens":2864,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:58:51.413208+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the pre-trained and fine-tuned Wav2Vec models through one fixed diarization pipeline—same segmentation, same clustering, same scoring tool—on a held-out Kurdish test set. If the pre-trained baseline scores near the paper's other reported baseline value of 15.2% DER rather than 22.8%, the claimed 7.2-point improvement evaporates; a score near 22.8% would support the claim.","supporting_citations":[{"cited_title":"data augmentation","cited_arxiv_id":null,"evidence_quote":"Supplies the Wav2Vec 2.0 architecture that the paper fine-tunes."},{"cited_title":"The approach starts by providing a comprehensive depiction of the dataset, including its organization and the preprocessing procedures executed to make it suitable for training","cited_arxiv_id":null,"evidence_quote":"Introduces the original Wav2Vec pretraining objective that the approach builds on."},{"cited_title":"FocusNet: imbalanced large and small organ segmentation with an end -to-end deep neural network for head and neck CT images,","cited_arxiv_id":null,"evidence_quote":"Shows cross-lingual pretraining plus fine-tuning works for low-resource speech, the direct precedent for transferring to Kurdish."},{"cited_title":"End -to-end neural speaker diarization with self-attention,","cited_arxiv_id":null,"evidence_quote":"Defines the transfer-learning mechanism the paper relies on to adapt multilingual representations."},{"cited_title":"Effectiveness of self -supervised pre-training for asr,","cited_arxiv_id":null,"evidence_quote":"Provides the CTC loss used to align speaker transitions during fine-tuning."},{"cited_title":"EEND-SS: Joint end-to-end neural speaker diarization and speech separation for flexible number of speakers,","cited_arxiv_id":null,"evidence_quote":"Supplies dropout regularization used to avoid overfitting on the small Kurdish corpus."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the diarization error rate metric used for the headline comparison."}],"review_version":1}