{"id":"2085f973-4c9f-4a74-bf13-937d25f26341","arxiv_id":"2508.02354","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":3,"one_line_summary":"A new Danish speech dataset with 96 participants yields 67% COPD detection accuracy with openSMILE features and logistic regression, suggesting speech as a cross-linguistic screening signal.","lead":"This paper presents a Danish speech dataset from 96 participants, half with COPD and half healthy controls, and reports a best accuracy of 67% for COPD detection using openSMILE features and logistic regression. A generalist might read it because speech-based screening could offer a remote, non-invasive way to flag COPD, but the modest accuracy and small sample size limit immediate clinical impact.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The supplied full text is an unrelated paper (Text2Lip), so the reported 67% accuracy cannot be checked for a held-out protocol or confound control; the central screening claim rests on an unverifiable evaluation.","rationale":"The reader's verdict is UNVERDICTED because the manuscript text does not correspond to the claimed COPD paper. My review agrees with that assessment, but I would sharpen the load-bearing concern: the abstract's 67% accuracy is only interpretable if the missing full text establishes a clean evaluation protocol. The reader's weakest assumption about confounds and evaluation artifacts is correct, and it is the right substantive worry. However, the more immediate problem is evidentiary: the full text provided is an unrelated paper, so even the basic experimental setup cannot be inspected. I do not think this rises to rejection based on the abstract alone, because a small feasibility study with a modest accuracy is plausible and the claim itself is modest. But the current submission, as received, gives no way to verify the number, so the appropriate verdict remains UNVERDICTED. No change to the reader's verdict is needed; my concern reinforces it rather than moving it.","tokens_in":11176,"tokens_out":2105,"duration_ms":24778,"concrete_test":"Obtain the actual full text for arXiv:2508.02354 (e.g., the arXiv PDF or the authors' repository) and locate the evaluation protocol behind the 67% result. Check specifically: (1) Is the accuracy computed on a held-out test set or via cross-validation with participant-level splits rather than utterance-level splits? (2) Are age, sex, smoking history, and recording conditions balanced or statistically adjusted between COPD and control groups? (3) How many feature sets and models were compared, and is 67% the best of many selected post hoc? If the actual paper reports a participant-disjoint, confound-aware evaluation, the central claim is supported; if not, the 67% cannot be interpreted as evidence for COPD screening.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that a logistic regression on openSMILE features achieved 67% accuracy for COPD detection in 96 Danish speakers, and that this supports speech-based COPD screening. For that claim to be meaningful, the 67% must reflect disease-related acoustic signal rather than participant-level confounds or evaluation artifacts. The abstract alone does not provide the decisive details: no held-out test set, no cross-validation scheme, no statement about participant-level splits, no class-balance or demographic-matching information, and no indication of how many feature/model combinations were tried before reporting the 'best' accuracy. More importantly, the manuscript body supplied for arXiv:2508.02354 is a different paper entirely (Text2Lip, a talking-face generation paper), so the COPD study's methods, dataset description, and experimental protocol are not available for review at all. This is an explicit missing-support condition: the evidence needed to validate the central claim is absent from the submitted text. The concern is not that the authors are necessarily wrong, but that the only evidence in scope is an abstract that cannot support the screening conclusion without the missing evaluation details.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission presents an abstract for a study on COPD detection from Danish speech, reporting a dataset of 96 participants (48 with COPD, 48 healthy) performing reading, coughing, and sustained-vowel tasks, and a best classification accuracy of 67% using openSMILE features with logistic regression, which is claimed to support speech-based screening for COPD. The manuscript body, however, is the unrelated Text2Lip paper (arXiv:2508.02362v1) on text-driven talking-face generation; no methods, dataset documentation, or evaluation details for the COPD study appear in the submitted text. As a result, the only checkable evidence is the abstract, which lacks the experimental protocol needed to validate the headline accuracy.","tokens_in":11369,"tokens_out":5521,"duration_ms":60505,"significance":"If the COPD result were fully documented, the dataset would be a potentially useful contribution: Danish is an under-represented language in respiratory speech-biomarker research, and even a modest baseline accuracy combined with a public dataset would give the community a point of comparison. The abstract is also appropriately restrained in claiming 'potential' rather than clinical deployment. However, the submitted manuscript does not contain that documentation, and the single abstract-level accuracy figure cannot support the screening conclusion on its own; no confidence interval, cross-validation scheme, or confounder analysis is provided. Because the full text is a different paper, the significance cannot be assessed beyond the abstract, and the central claim is not verifiable from the submitted materials.","major_comments":[{"comment":"The manuscript body after the abstract is the Text2Lip paper, arXiv:2508.02362v1, with a different title, author list, and subject area; it contains no description of the Danish COPD dataset, the three speech tasks, the openSMILE/x-vector pipelines, or the evaluation protocol. This is a load-bearing missing-support condition: the central claim of the abstract cannot be checked against any portion of the submitted manuscript.","section":"Full text (all sections)"},{"comment":"The phrase 'best accuracy of 67%' is reported without specifying the evaluation protocol: no held-out test set, no cross-validation scheme, no statement of whether splits are at the participant or utterance level, and no confidence interval. With n=96 participants, a single point estimate is insufficient to establish that the classifier exceeds chance or that it generalizes, and utterance-level splits would create within-speaker leakage.","section":"Abstract"},{"comment":"No demographic or clinical matching is described: the abstract does not state whether the COPD and control groups are balanced for age, sex, smoking history, or recording conditions, nor whether any such variables were adjusted for in the model. If the logistic regression separates groups on these attributes rather than on disease-related acoustic signal, the 67% accuracy would not support the screening conclusion; a concrete check is to report the accuracy of a classifier trained on demographic and recording variables alone and demonstrate the incremental value of speech features.","section":"Abstract"},{"comment":"The use of the word 'best' indicates that multiple feature sets and model families were compared, but the selection procedure and the number of configurations tried are not reported. Without nested validation or reporting all model and feature results, the chosen 67% figure is potentially optimistically biased, and the conclusion that the findings support a screening tool is not warranted by the abstract alone.","section":"Abstract"}],"minor_comments":[{"comment":"The phrase 'different levels of COPD' is undefined; specify diagnostic criteria and severity staging such as GOLD grades.","section":"Abstract"},{"comment":"No ethics approval, informed-consent, or dataset-availability information is provided, which is expected for a clinical speech dataset paper.","section":"Abstract"},{"comment":"The reference list consists entirely of talking-face and computer-vision literature, so the submitted bibliography cannot contextualize the COPD claims made in the abstract.","section":"Full text"}],"recommendation":"reject","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The submission has a serious problem: the full text under arXiv:2508.02354 is not the COPD paper. It is Text2Lip, a talking-face generation paper from a different set of authors. There is no COPD methodology, no dataset description, no experimental protocol to check. I can only go on the abstract, and that is not enough to referee.\n\nWhat the abstract does describe is a plausible small feasibility study: 96 Danish participants, half with COPD, three speech tasks, openSMILE features versus x-vector embeddings, and a best accuracy of 67% from logistic regression. A new Danish dataset is a real contribution to the speech-biomarker literature, which is dominated by English and a few other languages. The authors are also appropriately modest in their conclusion—they say the findings \"support the potential\" of speech analysis, not that it is ready for screening. That is the right tone for a study this size.\n\nThe soft spots are equally clear, even from the abstract alone. Sixty-seven percent accuracy on 96 people is weak, and the abstract gives no confidence intervals, no cross-validation scheme, no statement about participant-level splits, and no mention of controlling for age, sex, or recording conditions. The reader's stress test is right: the result could easily reflect confounds or evaluation artifacts, not disease state. The mismatch between the abstract and the full text is the bigger issue. I cannot tell whether the actual COPD manuscript addresses any of these concerns, because it is not here.\n\nMy recommendation is to desk reject this submission as-is. A peer review system cannot engage with a manuscript whose supplied text is a different paper. If the authors correct the upload and the real COPD paper is similar to the abstract, it would be worth a light review—maybe a borderline workshop or dataset paper—but it would need far more evaluation detail to support the screening claim. As submitted, there is nothing to verify.","headline":"The abstract describes a modest Danish COPD speech dataset and a 67% accuracy result, but the submitted full text is an unrelated talking-face paper, so the CODP work cannot be reviewed as submitted.","tokens_in":11896,"tokens_out":1418,"would_cite":false,"duration_ms":17189,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Danish speech recordings tell COPD from healthy at 67% accuracy","keywords":["COPD","speech analysis","Danish","openSMILE","x-vectors","logistic regression","screening","acoustic biomarkers"],"falsifier":"Classify the same recordings with the COPD and healthy groups matched one-to-one on age and sex, or add age and sex as explicit features and see whether accuracy collapses toward chance. If accuracy falls to chance after matching, the screening claim is not supported.","tokens_in":11003,"feed_emoji":"🫁","tokens_out":3409,"duration_ms":33638,"temperature":0.7,"pith_summary":"This paper builds a new Danish speech corpus from 96 participants, half with diagnosed COPD and half healthy, each recorded reading, coughing, and sustaining vowels. It then tests simple classifiers: hand-crafted openSMILE acoustic features and learnt x-vector embeddings. The best model, logistic regression on openSMILE features, reaches 67% accuracy at separating the two groups. The authors take this as support for speech-based COPD screening that is non-invasive, remote, and scalable, while acknowledging that validity across languages still needs testing.","feed_headline":"Danish speech recordings tell COPD from healthy at 67% accuracy","feed_subtitle":"A new 96-speaker Danish corpus of reading, cough, and vowel tasks supports speech as a cheap remote screening signal.","key_machinery":"The argument is carried by two feature families: openSMILE, a standard extraction pipeline for hand-crafted acoustic descriptors such as pitch, energy, and cepstral coefficients, and x-vectors, dense embeddings trained to represent speaker characteristics. A logistic regression baseline on the openSMILE features is the best performer. The three speech tasks, reading, coughing, and sustained vowels, are meant to expose different aspects of respiratory function, so their acoustic features can jointly distinguish the disease group.","core_discovery":"The paper's central claim is that acoustic characteristics of Danish speech carry enough COPD-related signal for a lightweight classifier to separate patients from controls at 67% accuracy. The result is obtained on balanced groups of 48 COPD patients and 48 healthy controls across three speech tasks, with hand-crafted openSMILE features outperforming the learnt x-vector embeddings in this setup. This is offered as evidence that speech-based analysis can serve as a scalable screening tool for COPD, with the caveat that 67% accuracy is a baseline for future work rather than a clinical threshold.","pith_inferences":["The paper does not report age and sex matching between the COPD and healthy groups, so part of the 67% could reflect these confounds; a matched re-analysis would be a natural next experiment.","A dedicated cough-only classifier is a testable extension: cough acoustics directly reflect airflow limitation and may outperform the combined model.","If the same recordings are re-classified after a day or a week, stability of decisions would show whether the signal is disease-related or vocal-day variation.","The claim of scalability assumes the model survives recording-device variation; a cross-device test (phone vs. studio microphone) would stress that assumption."],"forward_implications":["If 67% accuracy reproduces on new Danish speakers, a lightweight screening tier for COPD is feasible in Danish without specialized equipment.","The comparison of feature families suggests hand-crafted acoustic descriptors capture more COPD-relevant signal than generic speaker embeddings in this small corpus.","The three-task protocol can be reused as a template for building comparable COPD speech datasets in other languages.","The achieved accuracy, while above chance, is not high enough for diagnosis, so the realistic use case is referral or triage before clinical testing."],"supporting_citations":[],"fun_headline_variants":["Danish speech signals COPD at 67% accuracy","COPD spotted in Danish speech, 67% baseline","67% accuracy: speech analysis on Danish COPD data","Danish speech tasks hint at COPD, 67% result","COPD detection via Danish speech: 67% in 96 speakers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result depends on the 67% accuracy being driven by COPD-related acoustic differences rather than by age, sex, recording order, or other differences between the COPD and healthy groups.","fun_headline_variants_meta":{"raw":{"variants":["Danish speech signals COPD at 67% accuracy","COPD spotted in Danish speech, 67% baseline","67% accuracy: speech analysis on Danish COPD data","Danish speech tasks hint at COPD, 67% result","COPD detection via Danish speech: 67% in 96 speakers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00014,"raw_usage":{"total_tokens":1094,"prompt_tokens":810,"completion_tokens":284,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":426,"completion_tokens_details":{"reasoning_tokens":203}},"tokens_in":426,"tokens_out":284,"duration_ms":3402,"temperature":1.0,"reasoning_tokens":203,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:38:52.456400+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Classify the same recordings with the COPD and healthy groups matched one-to-one on age and sex, or add age and sex as explicit features and see whether accuracy collapses toward chance. If accuracy falls to chance after matching, the screening claim is not supported.","supporting_citations":[],"review_version":1}