{"id":"a3fa5da5-1f28-4dfd-b552-94bc354f0ef8","arxiv_id":"2508.04404","paper_version":2,"verdict":"UNVERDICTED","confidence":"UNKNOWN","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":3,"one_line_summary":"Perfusion MRI descriptors from brain regions are reported to separate distal ischemic stroke from seizure mimics with an AUROC of 0.90 in a 162-patient cohort.","lead":"The abstract reports that a machine-learning model reading MRI perfusion images can separate true distal strokes from seizure-induced stroke mimics. If the numbers hold, emergency teams could catch small strokes that CT misses and avoid treating seizure patients as stroke patients.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The supplied full text is a different paper (FlexQ, arXiv:2508.04405v2), so the stroke MRI abstract's AUROC claim has no supporting methods or results in the document; the central claim is unverifiable as submitted.","rationale":"The reader's verdict of UNVERDICTED is exactly right, but the strongest reason is more basic than the weakest_assumption field suggests. The reader focuses on reference-standard label correctness, which is an important scientific concern if the paper's methods were present. Here, however, the manuscript body is an entirely different paper (FlexQ), so the abstract's claims cannot be checked at all. The label-leakage issue is downstream: it presupposes a full text that defines the cohort and the labeling protocol. Because no such methods appear in the supplied document, the central claim has no derivable support. This is not a verdict that the AUROC result is wrong; it is a verdict that the submitted artifact cannot be audited. The concrete test is therefore document verification first, then label-standard and reproducibility checks once the correct manuscript is obtained. The current UNVERDICTED status should stand, with the recommendation that the authors confirm the correct full-text file is attached to the arXiv record.","tokens_in":25967,"tokens_out":2753,"duration_ms":30575,"concrete_test":"Download the source/PDF directly from the arXiv record for 2508.04404 (not the supplied body) and compare it with the abstract. If the file is the FlexQ text, the verdict remains UNVERDICTED until the correct MRI manuscript is provided. If the correct full text is obtained, first verify the reference standard: identify who labeled the 162 cases, whether DSC perfusion features were available to the labelers, and whether labels were assigned blinded to PMDs; then reproduce the logistic-regression AUROC/AUPRC from the stated features and cohort using the linked code.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The single load-bearing problem is that the document supplied as the full text is not this paper. The title, authors, abstract, sections, tables, and references all belong to FlexQ, an INT6 LLM quantization framework (arXiv:2508.04405v2); there is no mention of DSC-MRI, perfusion map descriptors, distal stroke, seizures, the 162-patient cohort, or the logistic regression. The abstract's central claim (AUROC 0.90, AUPRC 0.74, specificity 92%, sensitivity 73%) therefore has zero supporting methods or results inside the manuscript. This is not a claim that the MRI result is false; it is an unverifiability failure. The abstract itself is also incomplete: the GitHub URL is truncated ('.../PMD_analysis{github.com/...' ). Under the rule that every passage of the manuscript is in-scope evidence, the absent methods, absent reference-standard definition, absent feature definitions, and absent validation protocol are all decisive: there is no way to reconstruct or audit the analysis. The reader's concern about label leakage is a plausible secondary risk, but it cannot even be assessed until the correct full text is available.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission carries the title \"Discriminating Distal Ischemic Stroke from Seizure-Induced Stroke Mimics Using Dynamic Susceptibility Contrast MRI\" and an abstract reporting that perfusion-map descriptors (PMDs) from DSC-MRI distinguish distal acute ischemic stroke from seizure-induced mimics in a retrospective cohort of 162 patients (129 AIS, 33 seizures). The abstract states that a logistic regression model achieved AUROC 0.90, AUPRC 0.74, specificity 92%, and sensitivity 73%. However, the accompanying full text is not this paper. It is the FlexQ manuscript (arXiv:2508.04405v2), an INT6 LLM quantization framework, and contains no mention of DSC-MRI, PMDs, stroke, seizures, the 162-patient cohort, feature definitions, statistical tests, or validation methodology. As submitted, therefore, the abstract's central claims have no supporting methods or results in the manuscript.","tokens_in":26214,"tokens_out":2481,"duration_ms":30499,"significance":"The clinical question is genuine and important: distal ischemic strokes are often radiologically inconspicuous on CT, and seizure-induced stroke mimics are a common source of diagnostic uncertainty. If the reported result were supported, region-wise DSC-MRI perfusion descriptors with hemispheric asymmetry would be a plausible, interpretable contribution to an emergency-setting diagnostic pathway. The abstract also makes a concrete, falsifiable performance claim and indicates an intention to release code, which are positive features. Nevertheless, the submission as it stands provides no verifiable evidence for these claims, and the central results cannot be checked by a reader. The scientific significance of the abstract is therefore real but entirely unrealized in the submitted document.","major_comments":[{"comment":"The provided full text is a different paper: 'FlexQ: Efficient Post-training INT6 Quantization for LLM Serving via Algorithm-System Co-Design' (arXiv:2508.04405v2). Its title, authors, abstract, sections, tables, figures, and references concern 6-bit LLM quantization and contain no material on DSC-MRI, perfusion map descriptors, ischemic stroke, seizures, the 162-patient cohort, or logistic regression. The abstract's AUROC 0.90, AUPRC 0.74, specificity 92%, and sensitivity 73% therefore have no derivable support in this manuscript. This is not a local omission but an absence of the entire methods and results section; the central claim is unverifiable as submitted.","section":"Full text (all sections)"},{"comment":"Even taking the abstract at face value, the reported AUROC 0.90, AUPRC 0.74, specificity 92%, and sensitivity 73% are presented without confidence intervals, without a description of the cross-validation or split protocol, and without the classification threshold that yields the stated sensitivity/specificity pair. Given a class imbalance of 129:33, an AUROC alone cannot establish practical utility, and the absence of any validation details makes overfitting or threshold optimization on the same cohort impossible to rule out. The correct full text must supply these details.","section":"Abstract (performance metrics)"},{"comment":"The abstract states that statistical analyses identified brain regions with significant group differences in PMDs and that a logistic regression model was then trained on PMDs. It does not state whether the same 162 patients were used both to select the significant regions/PMD types and to evaluate the classifier. If selection was performed on the full cohort and the model was evaluated on the same cohort, the reported AUROC would be optimistically biased; nested or outer cross-validation is required. In addition, the reference standard for the 129 AIS and 33 seizure labels is not defined anywhere in the submission. Since the paper's premise is precisely that this distinction is clinically and radiologically difficult, label definition is load-bearing; if expert labels partly incorporated perfusion information, the discrimination could be circular. These issues must be addressed explicit","section":"Abstract (feature selection and reference standard)"}],"minor_comments":[{"comment":"The GitHub URL is truncated and malformed: 'https://github.com/Marijn311/PMD_extraction_and_analysis{github.com/Marijn311/PMD_extraction_and_analysis'. The intended link should be provided in full and checked.","section":"Abstract (GitHub link)"},{"comment":"The abstract does not define 'PMD' or specify which perfusion descriptors are included, nor does it define 'distal' (e.g., vessel territory, occlusion level, or infarct size). Definitions are needed for reproducibility regardless of the final text.","section":"Abstract (terminology)"},{"comment":"There is no statement of ethics approval, imaging acquisition parameters, or inclusion/exclusion criteria for the retrospective cohort. These should appear in any revised submission.","section":"General"}],"recommendation":"reject","confidential_remarks":"The mismatch between the abstract and the full text is fundamental: the manuscript body is an unrelated LLM-quantization paper. This appears to be a submission error, but as submitted there is no way to peer-review the reported MRI study. If a correct full text exists, it should be resubmitted as a new manuscript with the methods and validation details described in the major comments. I saw no evidence specific to the scientific content to judge the underlying clinical claim beyond the abstract."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I looked at arXiv:2508.04404. The abstract is about a clinically important problem: using DSC-MRI perfusion descriptors to tell distal occlusions from seizure mimics in stroke triage. That's a real gap, and an AUROC of 0.90 with 92% specificity on 162 patients would be a meaningful step if it holds.\n\nBut here's the problem: the full text I have is not this paper. It's the FlexQ paper on INT6 LLM quantization (arXiv:2508.04405v2), by different authors, with no mention of MRI or stroke. So there are no methods, no feature definitions, no statistical tests, no validation details, nothing behind the abstract's numbers. You literally cannot check the result.\n\nThe abstract itself is thin even on its own. It doesn't give confidence intervals for the AUROC, doesn't define the reference standard (how were the 129 AIS and 33 seizures diagnosed?), and doesn't say whether the 'significant' regions were selected on the same cohort that trained the logistic regression. If the authors picked regions on the full dataset and then fit the model on that same dataset, the performance estimate would be optimistically biased. That concern is secondary right now because the main document doesn't contain the analysis at all, but it's worth flagging for the eventual full version.\n\nWhat's good here: the clinical framing is sensible, and the proposed features (region-wise PMDs plus hemispheric asymmetry) are interpretable and plausible. The idea isn't obviously wrong. There's just no way to evaluate it from what was submitted.\n\nBottom line: this is a desk-reject as submitted. The mistake is likely a bad upload—the abstract and the full text are different works. The right response is to tell the authors to put up the correct paper. Once the real manuscript is available, it should go to peer review; the question is worth a serious referee. I would not cite or use this version for anything.","headline":"The stroke MRI abstract is clinically plausible, but the attached full text is a different paper, leaving the AUROC 0.90 claim without supporting methods—desk-reject as submitted, ask for the correct manuscript.","tokens_in":26751,"tokens_out":3616,"would_cite":false,"duration_ms":37334,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Perfusion MRI patterns separate distal strokes from seizure-induced mimics with 0.90 AUROC in a 162-patient cohort.","keywords":["distal acute ischemic stroke","stroke mimics","seizure","dynamic susceptibility contrast MRI","perfusion map descriptors","hemispheric asymmetry","logistic regression","magnetic resonance perfusion imaging"],"falsifier":"Re-run the analysis on a prospectively collected cohort where seizure versus distal-stroke labels are adjudicated by clinicians blinded to DSC-MRI perfusion maps and PMDs, and check whether the logistic regression's AUROC remains near 0.90. A large drop would indicate the reported discrimination came from label leakage or overfitting. A cheaper check: report cross-validated performance stratified by label source (for example discharge diagnosis versus imaging-confirmed) and show the separation persists when ambiguous perfusion reads are excluded.","tokens_in":25833,"feed_emoji":"🧠","tokens_out":6954,"duration_ms":68474,"temperature":0.7,"pith_summary":"Distinguishing a true distal acute ischemic stroke from a seizure that mimics a stroke is hard, especially when the blocked vessel is too small to show on CT. The paper argues that magnetic resonance perfusion imaging carries enough signal to help make that call: region-wise perfusion descriptors, concentrated in temporal and occipital lobes and complemented by hemispheric asymmetry, separated 129 distal strokes from 33 seizure mimics in a retrospective cohort. A logistic regression trained on these descriptors achieved an area under the ROC curve of 0.90, with 92% specificity and 73% sensitivity. If the result holds, perfusion MRI could act as an interpretable secondary test that reduces both missed distal occlusions and unnecessary treatment of mimics.","feed_headline":"Perfusion MRI flags distal strokes among seizure mimics, AUROC 0.90","feed_subtitle":"Temporal and occipital perfusion descriptors sorted 129 strokes from 33 seizure mimics; CT often misses these.","key_machinery":"The central object is the region-wise perfusion map descriptor (PMD): a set of quantitative features computed from dynamic susceptibility contrast (DSC) perfusion maps, organized by brain region. The pipeline extracts these descriptors, tests which regions show significant group differences, confirms them with hemispheric asymmetry analysis, and feeds the PMDs into a logistic regression classifier. The descriptors turn raw perfusion imaging into interpretable, region-localized signal that carries the stroke-versus-seizure discrimination.","core_discovery":"The paper claims that region-wise perfusion map descriptors (PMDs) extracted from dynamic susceptibility contrast (DSC) magnetic resonance perfusion images can discriminate distal acute ischemic stroke (AIS) from seizure-induced stroke mimics. Statistical analyses of the 162-patient retrospective cohort identified significant group differences mainly in temporal and occipital lobe regions, and hemispheric asymmetry analyses highlighted the same regions as discriminative. A logistic regression model trained on the PMDs achieved an AUROC of 0.90, an AUPRC of 0.74, a specificity of 92%, and a sensitivity of 73%. The authors read these results as evidence that MRP-based PMDs are interpretable fe","pith_inferences":["The cohort is imbalanced (129 strokes versus 33 mimics), so the AUPRC of 0.74 is the more conservative estimate of real-world performance once base rates shift; triage settings should plan around that number rather than the AUROC.","The strongest test of the claim would be a prospective or externally validated study in which the reference standard is adjudicated by a panel blinded to the perfusion features the model consumes; the abstract does not define the reference standard or report how labels were assigned.","If the temporal and occipital PMD signal is real, it may reflect seizure-induced hyperperfusion or post-ictal changes; comparing against other mimic types (migraine, conversion disorder, hypoglycemia) would reveal whether the descriptors mark stroke-versus-mimic generally or seizure specifically.","A natural next experiment is to measure whether adding these PMDs to clinical variables or to a neuroradiologist's reading improves diagnostic accuracy beyond either alone; the abstract does not include such a comparison."],"forward_implications":["If the discrimination generalizes, DSC-MRI perfusion descriptors could serve as a diagnostic aid for distal occlusions that CT-based protocols routinely miss.","The temporal and occipital regions highlighted by the analysis give radiologists concrete locations to scrutinize when a distal stroke is suspected in a seizure mimic.","Hemispheric asymmetry in these perfusion descriptors is a candidate feature to fold into future diagnostic scoring systems.","At the reported operating point, 73% sensitivity at 92% specificity means the trade-off can be tuned: raising sensitivity would shift the model toward fewer missed strokes at the cost of more mimics receiving acute treatment.","Because the model is a logistic regression on interpretable descriptors, clinicians can inspect which regions and perfusion values drive each prediction, unlike a black-box image classifier."],"supporting_citations":[],"fun_headline_variants":["Perfusion MRI sorts distal strokes from seizure mimics with 0.90 AUROC","MRI perfusion descriptors separate distal stroke from seizure mimics","Temporal-occipital perfusion patterns flag distal stroke vs seizure mimics","DSC-MRI perfusion maps hit AUROC 0.90 in stroke-seizure mimic task","MR perfusion PMDs distinguish distal stroke from seizure mimics, AUROC 0.90"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that the reference-standard labels — which of the 162 retrospective patients truly had distal acute ischemic stroke and which had seizures — are correct and were assigned without using the DSC perfusion features the model was built on.","fun_headline_variants_meta":{"raw":{"variants":["Perfusion MRI sorts distal strokes from seizure mimics with 0.90 AUROC","MRI perfusion descriptors separate distal stroke from seizure mimics","Temporal-occipital perfusion patterns flag distal stroke vs seizure mimics","DSC-MRI perfusion maps hit AUROC 0.90 in stroke-seizure mimic task","MR perfusion PMDs distinguish distal stroke from seizure mimics, AUROC 0.90"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000292,"raw_usage":{"total_tokens":1575,"prompt_tokens":811,"completion_tokens":764,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":663}},"tokens_in":555,"tokens_out":764,"duration_ms":8125,"temperature":1.0,"reasoning_tokens":663,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T00:00:24.998422+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the analysis on a prospectively collected cohort where seizure versus distal-stroke labels are adjudicated by clinicians blinded to DSC-MRI perfusion maps and PMDs, and check whether the logistic regression's AUROC remains near 0.90. A large drop would indicate the reported discrimination came from label leakage or overfitting. A cheaper check: report cross-validated performance stratified by label source (for example discharge diagnosis versus imaging-confirmed) and show the separation persists when ambiguous perfusion reads are excluded.","supporting_citations":[],"review_version":1}