{"id":"5cd94375-df44-4a5d-a59c-4151234d3b96","arxiv_id":"2606.04722","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"StrokeTimer combines self-supervised disentanglement and energy-guided contrastive learning to classify NCCT images into three stroke onset time windows, reporting macro AUC 0.69 and F1 0.57 on multi-center data.","lead":"StrokeTimer is a machine learning system that analyzes routine non-contrast CT brain scans to classify ischemic stroke onset into one of three time windows: under 4.5 hours, 4.5-6 hours, or over 6 hours. If accurate, it could help doctors decide on time-sensitive treatments like clot removal when the exact start time of the stroke is unknown.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Lack of center-stratified validation leaves cross-scanner robustness unproven","rationale":"The identified concern matches the reader's weakest assumption exactly. Because the original review was abstract-only and the full text supplies no counter-evidence to this specific risk, the UNVERDICTED / LOW verdict is unaffected.","tokens_in":1815,"tokens_out":263,"duration_ms":18004,"concrete_test":"Re-run the full pipeline with registry-stratified splits (train on MR CLEAN Registry, test on MR CLEAN LATE and vice versa); if macro AUC falls below 0.60 on the held-out registry, the cross-center robustness claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that self-supervised disentanglement plus energy-guided contrastive learning extracts subtle, consistent ischemic changes across centers/scanners despite heterogeneity and imbalance. The reported macro AUC 0.69 / F1 0.57 (50% gain, p<0.005) on pooled MR CLEAN Registry + LATE data would not support this if the model instead exploits center-specific acquisition artifacts. No description in the provided material indicates leave-one-registry-out or center-stratified splits; pooled evaluation alone cannot rule out this alternative explanation.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces StrokeTimer, a framework that combines self-supervised disentanglement learning with energy-guided contrastive learning to estimate ischemic stroke onset time from non-contrast CT (NCCT) scans. Onset times are binned into three clinically relevant intervals (<4.5 h, 4.5–6 h, >6 h). On a pooled multi-center dataset from the MR CLEAN Registry and MR CLEAN LATE cohorts, the method reports a macro AUC of 0.69 and macro F1-score of 0.57, representing an approximately 50% improvement over the strongest baseline (p < 0.005). Model explanations are said to highlight radiological biomarkers such as gray-white matter blurring. Code is released at https://github.com/BrainVas/StrokeTimer.","tokens_in":1934,"tokens_out":511,"duration_ms":13681,"significance":"If the reported gains are shown to arise from the proposed disentanglement and contrastive components rather than center-specific acquisition artifacts, the work would be significant for supporting time-critical reperfusion decisions in acute stroke, where NCCT is the most widely available modality and onset-time uncertainty is common. The explicit release of code strengthens reproducibility and allows independent verification of the empirical claims.","major_comments":[{"comment":"Experimental evaluation (likely §4): The manuscript evaluates on pooled data from MR CLEAN Registry and MR CLEAN LATE without describing center-stratified splits, leave-one-registry-out validation, or scanner-specific hold-out sets. Because the central claim is robustness to “center-scanner-related heterogeneity,” the absence of such partitioning leaves open the possibility that performance exploits cohort-specific acquisition signatures rather than the subtle ischemic changes asserted in the abstract. A concrete test (e.g., per-center AUC tables or cross-registry evaluation) is required to substantiate the robustness claim.","section":"Experimental evaluation"}],"minor_comments":[{"comment":"Abstract and §3: The loss formulations for the energy-guided contrastive term and the disentanglement objective are referenced but not written out with explicit equations; adding the mathematical definitions would improve clarity without altering the central narrative.","section":"Methods"},{"comment":"Table 1 or results section: Baseline descriptions should include the exact architectures and training protocols used for the “strongest baseline” so that the 50% relative improvement can be directly reproduced.","section":"Results"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive feedback on the experimental evaluation. The concern regarding potential exploitation of center-specific acquisition signatures is well-taken and directly relevant to our robustness claims. We address this point below and will incorporate the requested analyses in the revision.","responses":[{"response":"We agree that the current pooled evaluation does not sufficiently isolate the contribution of the proposed methods from potential center-specific effects. MR CLEAN Registry and MR CLEAN LATE represent distinct national cohorts with differing inclusion criteria and acquisition protocols, yet we did not report explicit cross-registry or center-stratified results. In the revised manuscript we will add leave-one-registry-out experiments (training on one cohort and testing on the other) together with per-center AUC and macro-F1 tables. These additions will allow direct assessment of whether performance gains persist across acquisition heterogeneity.","revision_made":"yes","referee_comment":"Experimental evaluation (likely §4): The manuscript evaluates on pooled data from MR CLEAN Registry and MR CLEAN LATE without describing center-stratified splits, leave-one-registry-out validation, or scanner-specific hold-out sets. Because the central claim is robustness to “center-scanner-related heterogeneity,” the absence of such partitioning leaves open the possibility that performance exploits cohort-specific acquisition signatures rather than the subtle ischemic changes asserted in the abstract. A concrete test (e.g., per-center AUC tables or cross-registry evaluation) is required to substantiate the robustness claim."}],"tokens_in":1480,"tokens_out":313,"duration_ms":12672,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"StrokeTimer combines self-supervised disentanglement with energy-guided contrastive learning for three-class onset-time prediction on NCCT and reports a macro AUC of 0.69 plus F1 of 0.57 on pooled MR CLEAN Registry and LATE data, beating baselines by nearly 50 percent.\n\nThe work does a few things right. It targets a genuine clinical gap where onset time is often unknown and standard imaging changes are subtle. The authors handle long-tailed classes and scanner heterogeneity explicitly, release code, and show saliency maps that align with known radiological signs like gray-white blurring. That combination applied to this exact task is new relative to the cited literature.\n\nThe soft spot is the validation. No center-stratified or leave-one-registry-out splits are described, so the pooled numbers could reflect site-specific acquisition patterns rather than consistent ischemic features. The abstract also skips architecture details, exact loss formulations, and data-split descriptions, which makes it difficult to credit the lift to the claimed components. These gaps are real but fixable with additional experiments.\n\nThe paper is aimed at medical-image researchers working on stroke or time-sensitive diagnostics. A reader who needs concrete ideas for contrastive learning on imbalanced NCCT data will find usable pieces, even if the current evidence does not yet prove cross-scanner reliability.\n\nIt deserves peer review so referees can examine the full methods and request the missing stratified results.","headline":"StrokeTimer combines self-supervised disentanglement and energy-guided contrastive learning for three-class NCCT onset estimation and reports solid gains on pooled multi-center data, but the robustness claim rests on unstratified evaluation.","tokens_in":2477,"tokens_out":371,"would_cite":false,"duration_ms":14428,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"StrokeTimer estimates ischemic stroke onset time from non-contrast CT using self-supervised disentanglement and energy-guided contrastive learning.","keywords":["ischemic stroke","onset time estimation","non-contrast CT","self-supervised learning","contrastive learning","disentanglement","medical imaging","stroke classification"],"falsifier":"An independent multi-center NCCT test set on which the model drops to near-chance macro AUC of approximately 0.33 would falsify the claim that the learned representations are robust.","tokens_in":2744,"feed_emoji":"🧠","tokens_out":671,"duration_ms":20634,"temperature":0.7,"pith_summary":"The paper aims to establish that subtle early ischemic changes on non-contrast CT can be extracted to categorize onset time into three windows even when data is imbalanced and comes from multiple centers with different scanners. StrokeTimer combines self-supervised disentanglement learning with energy-guided contrastive learning to build representations that handle these issues. This matters because treatment eligibility for reperfusion depends on knowing how much time has passed since onset, yet that time is frequently unknown in practice. On a large multi-center dataset the method reaches macro AUC of 0.69 and F1 of 0.57, well above near-chance baselines. Model explanations point to known signs such as gray-white matter blurring and hypodense areas.","feed_headline":"Stroke onset time estimated from CT at 0.69 AUC","feed_subtitle":"Disentanglement plus energy-guided contrastive learning overcomes imbalance and scanner variation for three time windows.","key_machinery":"Self-supervised disentanglement learning combined with energy-guided contrastive learning, which isolates ischemic features from acquisition variability and class imbalance.","core_discovery":"StrokeTimer integrates self-supervised disentanglement learning with energy-guided contrastive learning to capture subtle ischemic patterns while addressing long-tailed data distributions under acquisition variability. On a large multi-center NCCT dataset from MR CLEAN Registry and MR CLEAN LATE it achieves a macro AUC of 0.69 and a macro F1-score of 0.57 for classifying onset into <4.5 h, 4.5-6 h, and >6 h, improving the strongest baseline by nearly 50 percent.","pith_inferences":["The same disentanglement strategy could be tested on other subtle time-dependent imaging features such as tumor growth or cardiac ischemia.","If the consistency assumption fails on new scanners, adding explicit domain-adaptation terms might be required.","Replacing the three discrete windows with a regression head on the same backbone would test whether continuous onset prediction is also feasible."],"forward_implications":["Onset time can be placed into three clinically actionable windows from routine NCCT scans with usable accuracy.","Interpretability maps align with established radiological biomarkers such as gray-white matter blurring.","The framework maintains performance across scanner and center differences where prior methods fail.","Automatic estimation becomes feasible in settings where patient history is unreliable."],"fun_headline_variants":["StrokeTimer yields 0.69 AUC for NCCT stroke onset classification","Disentanglement and contrastive learning for CT stroke timing at 0.69 AUC","Multi-center validation of StrokeTimer at 0.69 AUC on NCCT onset","Three time windows for stroke onset estimated from CT at 0.69 AUC"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Subtle early ischemic changes on NCCT are consistent enough across centers and scanners to be learned despite pronounced class imbalance and heterogeneity.","fun_headline_variants_meta":{"raw":{"variants":["StrokeTimer yields 0.69 AUC for NCCT stroke onset classification","Disentanglement and contrastive learning for CT stroke timing at 0.69 AUC","Multi-center validation of StrokeTimer at 0.69 AUC on NCCT onset","Three time windows for stroke onset estimated from CT at 0.69 AUC"]},"model":"grok-4.3","cost_usd":0.007145,"raw_usage":{"total_tokens":3346,"prompt_tokens":761,"num_sources_used":0,"completion_tokens":83,"cost_in_usd_ticks":71449500,"prompt_tokens_details":{"text_tokens":761,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2502,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":761,"tokens_out":83,"duration_ms":15522,"temperature":1.0,"reasoning_tokens":2502,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T07:06:41.743424+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"An independent multi-center NCCT test set on which the model drops to near-chance macro AUC of approximately 0.33 would falsify the claim that the learned representations are robust.","supporting_citations":[],"review_version":1}