{"id":"fd9dd0e3-0c54-4cd6-95d8-585ed612f022","arxiv_id":"2508.14286","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"OccluNet, a spatio-temporal attention-based detector, is claimed to outperform frame-based YOLOv11 baselines for occlusion detection in DSA sequences.","lead":"The paper introduces OccluNet, a deep learning model that detects blocked vessels in DSA video from acute stroke patients, reporting 89.02% precision and 74.87% recall. The submitted full text, however, appears to be an unrelated paper on tactile graphics, so the model's methods could not be reviewed.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"OccluNet's reported metrics cannot be checked because the submitted manuscript contains no OccluNet methods or evaluation; the central claim is unsupported by the provided text.","rationale":"The reader's verdict is UNVERDICTED, and I agree that the abstract's performance claim cannot be verified from this submission. My concern is slightly broader than the reader's weakest_assumption. The reader identified label accuracy and leakage as the key uncheckable assumptions. I agree those are the main technical risks, but the load-bearing issue is that there is no OccluNet manuscript at all in the provided text: no methods, no results, no evaluation details. This means every element of the central claim is unsupported, not just the weakest assumption. I therefore do not propose a different verdict: UNVERDICTED remains the correct assessment. The concrete test of patient-wise cross-validation would directly address the most likely source of inflated performance if the real paper is retrieved. I do not see grounds for REJECT because the abstract may be accurate in its original context; the problem is evidentiary, not an identified flaw in the model. The recommendation is UNCHANGED, with the reader's low confidence and unverdictable status upheld.","tokens_in":14046,"tokens_out":1974,"duration_ms":24175,"concrete_test":"Obtain the actual OccluNet paper and source code (the GitHub URL in the abstract is the intended source), and rerun the evaluation with patient-wise grouped cross-validation: assign all DSA frames from each MR CLEAN Registry patient to either training or test, never both. Then recompute precision and recall at the same confidence threshold and compare to 89.02%/74.87%. If the per-patient grouping changes the metrics by more than a few points, leakage is the likely cause; if the code or full methods are unavailable, the claim remains unverifiable.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the abstract's performance assertion: precision 89.02%, recall 74.87%, and significant improvement over YOLOv11 baselines on MR CLEAN Registry DSA sequences. For this claim to hold, the evaluation must be well-defined: a clear patient-level train/test split, a consistent frame/sequence unit of analysis, a fixed confidence threshold, and ground-truth labels that are accurate and independent of the model. However, the submitted full text is an unrelated manuscript on tactile graphics (arXiv 2508.14289). It contains no OccluNet architecture, no training protocol, no dataset split, no loss function, no inference procedure, no statistical test, and no evaluation code. Therefore the abstract's numbers are unverifiable assertions. This is not an internal inconsistency in OccluNet itself, but an evidentiary gap: the only document claiming these results does not describe the experiment that produced them. The reader's weakest_assumption about label accuracy and leakage is relevant, but the more fundamental problem is that no methods are present to even test that assumption. Without the full manuscript, source code, or a reproducible evaluation protocol, the claim cannot be distinguished from a miscalculated metric, a leakage-inflated result, or a correct but unverifiable outcome.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript's abstract claims a novel deep learning system, OccluNet, for automated vascular occlusion detection in digital subtraction angiography (DSA) sequences, reporting precision of 89.02% and recall of 74.87% on the MR CLEAN Registry, with two spatio-temporal attention variants significantly outperforming YOLOv11 baselines. However, the full text submitted is a completely different paper on tactile graphics for blind and low-vision users (“They Aren’t Built For Me”), with no mention of OccluNet, DSA, YOLOX, MR CLEAN, or any occlusion detection experiments. Thus, the central claim of the abstract is entirely unsupported by the body of the manuscript.","tokens_in":14292,"tokens_out":2325,"duration_ms":26392,"significance":"If the claimed result were substantiated, OccluNet would be a practically valuable contribution to endovascular thrombectomy workflows, offering a spatio-temporal detector that leverages temporal consistency across angiography sequences. The comparison with strong baselines such as YOLOv11 and the two attention variants would also be informative. However, the submitted manuscript contains none of the architecture, training protocol, data split, evaluation methodology, or statistical comparisons that would support this claim. Consequently, the significance cannot be assessed from the current submission, because the purported evidence is absent.","major_comments":[{"comment":"The submitted full text is an entirely unrelated paper on tactile graphics, titled “They Aren’t Built For Me”, with its own abstract about a Cleveland-McGill replication with blind participants. It contains no OccluNet architecture, no DSA dataset, no YOLOX/YOLOv11 baselines, no training or inference details, and no occlusion detection evaluation. The abstract’s claims about precision/recall and significant improvement over baselines therefore have no supporting methods or results anywhere in the manuscript. This is a load-bearing mismatch that invalidates the central claim; it cannot be remedied by minor edits.","section":"Full text (all sections after abstract)"},{"comment":"The abstract announces source code availability and two spatio-temporal attention variants (pure temporal and divided space-time), but none of these artifacts appear in the full text. No dataset split, patient-level or frame-level unit of analysis, confidence threshold, ground-truth labeling procedure, or statistical test is described. Even the self-reported limitations in Section 7.2 of the full text concern the tactile graphics study, not OccluNet. The manuscript therefore provides no basis to verify the abstract’s quantitative claims or to assess potential leakage, class imbalance, or test-set overfitting.","section":"Abstract vs. full text"}],"minor_comments":[{"comment":"The arXiv identifier in the full text (2508.14289) does not match the submitted identifier (2508.14286), and the title/author list of the full text differs from that implied by the abstract. These inconsistencies should be resolved if the submission is ever corrected.","section":"Metadata"},{"comment":"The full text is a camera-ready paper for IEEE VIS 2025 with its own acknowledgments and references; it is clearly a distinct document. If the intended submission were actually OccluNet, the manuscript body, figures, and references would need to be replaced entirely.","section":"Full-text formatting"}],"recommendation":"reject","confidential_remarks":"This appears to be a submission error at a fundamental level: the abstract and the full text describe two different papers. The OccluNet claims are completely unverifiable because the supporting paper is absent. Even with a complete rewrite of the full text, the current submission would not meet the standards for review. I recommend rejection without further review, although the editorial office may wish to check whether a different version was intended."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about arXiv:2508.14286. First, the abstract describes OccluNet, a YOLOX-plus-temporal-attention model for occlusion detection in DSA sequences, with precision and recall of 89.02% and 74.87% on the MR CLEAN Registry. Second, the full text attached to this submission is an entirely different paper about tactile graphics for blind users. That is not a minor formatting issue; it means we have no methods, no training protocol, no dataset split, no statistical tests, and no evaluation details for OccluNet. The only evidence for the central claim is the abstract itself.\n\nWhat is potentially new is the architecture combination and the application to a clinically meaningful problem: automated occlusion detection could reduce time-to-recanalization during endovascular thrombectomy. The specific performance numbers and the provided GitHub link suggest real work behind the scenes. Credit where due: the clinical motivation is sound, the numbers are concrete, and sharing code is good practice.\n\nBut the soft spot is the entire submission. The stress-test note is right—this is an evidentiary gap, not an internal inconsistency. We cannot check whether the performance advantage over YOLOv11 is genuine or an artifact of patient-level leakage, class imbalance, or test-set overfitting because none of the experiment is described. We cannot even confirm the model is new; maybe this combination has already been published elsewhere. The reader's worry about label accuracy is valid but secondary; the first problem is that no methods exist in this artifact to evaluate.\n\nIt is possible the authors accidentally uploaded the wrong file—that would fit with the tactile graphics paper being a legitimate, separate study. But as submitted, this is not a coherent manuscript for OccluNet. The tactile graphics content should not be reviewed as if it were the OccluNet paper, and the abstract alone cannot carry the claim.\n\nMy recommendation: do not send this to peer review in its current state. Return it to the authors and ask them to resubmit with the correct full text. If the real OccluNet paper appears with the architecture details, data split, and error bars, then it deserves serious referee time because the clinical target is important and the approach is plausible. This version does not.","headline":"This submission is a structural mismatch: the abstract describes OccluNet, but the full text is an unrelated paper on tactile graphics, so the central claim is unverifiable from this artifact.","tokens_in":14817,"tokens_out":2271,"would_cite":false,"duration_ms":24548,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Spatio-temporal model detects stroke occlusions in DSA with 89% precision","keywords":["occlusion detection","digital subtraction angiography","spatio-temporal deep learning","transformer attention","YOLOX","endovascular thrombectomy","acute ischemic stroke","MR CLEAN Registry"],"falsifier":"Train OccluNet and a YOLOv11 frame-wise baseline on the same MR CLEAN Registry data, but with a strict patient-level train/test split; if the precision-recall gap shrinks to less than a few points, the temporal-attention advantage is an artifact of information leakage. Also, if a YOLOv11 trained on the identical frames with identical augmentation and test-time settings matches or beats OccluNet, the claim that sequence modeling helps would be refuted.","tokens_in":13967,"feed_emoji":"🧠","tokens_out":3921,"duration_ms":38141,"temperature":0.7,"pith_summary":"The paper proposes OccluNet, a deep-learning model that detects vascular occlusions in digital subtraction angiography (DSA) sequences by combining the YOLOX object detector with transformer-based temporal attention. The authors claim that using the temporal structure of the image sequence improves detection over analyzing single frames, and they report precision of 89.02% and recall of 74.87% on the MR CLEAN Registry. They compare against YOLOv11 baselines trained on individual frames or minimum intensity projections, and find OccluNet significantly better. The practical goal is to support endovascular thrombectomy (EVT) decision-making, where rapid, accurate identification of blockages is critical. A reader should care because an automated detector that works on real clinical angiography could reduce time-to-treatment in acute ischemic stroke.","feed_headline":"Spatio-temporal AI spots stroke blockages at 89% precision","feed_subtitle":"OccluNet watches the whole DSA sequence, not single frames, and beats the frame-by-frame YOLOv11 baseline on clinical angiograms.","key_machinery":"OccluNet's core is a spatio-temporal attention pipeline: YOLOX produces frame-level detection features, and a transformer encoder with temporal attention aggregates them across the DSA sequence, so the model can reward detections that are temporally consistent and suppress flicker. The divided space-time variant factorizes attention into spatial and temporal steps; the pure temporal variant applies attention only across time. The machinery's job is to turn a set of per-frame detections into a sequence-level decision, which is what the paper claims outperforms single-frame detection.","core_discovery":"On the paper's terms, the central discovery is that adding a transformer-based temporal attention module to a single-stage detector (YOLOX) lets the model exploit consistency across DSA frames, improving occlusion detection beyond what the same detector sees in a single frame or a minimum-intensity projection. Two spatio-temporal variants—pure temporal attention and divided space-time attention—achieve similar performance, with the best configuration reaching 89.02% precision and 74.87% recall on MR CLEAN Registry data. The authors interpret this as evidence that temporal modeling, rather than the specific attention factorization, is what drives the gain, and that OccluNet is a viable automa","pith_inferences":["The paper likely understates the risk of patient-level leakage: if frames from the same patient appear in both train and test, the temporal attention could memorize patient-specific anatomy rather than generalize. A patient-stratified cross-validation would be the natural check.","The model's ability to use temporal consistency may extend beyond occlusions to other time-varying contrast phenomena in angiography, such as collateral flow grading.","A direct comparison could be made against a 3D-CNN or video transformer baseline to see how much of the gain is from attention specifically versus any temporal encoder."],"forward_implications":["If OccluNet's performance holds, EVT teams could get an automated second reader that flags occlusions from the full sequence rather than single frames.","The finding that both attention variants perform similarly suggests a simpler pure-temporal model may suffice, reducing compute cost in practice.","The reported precision-recall balance implies OccluNet could be used as a screening tool, with false negatives at roughly 25% warranting human review.","The success on MR CLEAN Registry data supports testing on prospective, multi-center DSA datasets to confirm clinical usefulness."],"supporting_citations":[],"fun_headline_variants":["Temporal attention boosts stroke occlusion AI to 89% precision","Watching DSA sequence, not frames, spots strokes at 89% precision","Temporal attention model finds stroke blockages with 89% precision","OccluNet: sequence-level AI beats single-frame detection for strokes","AI that watches DSA video finds blockages better than frame-by-frame"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"That the MR CLEAN Registry ground-truth labels are correct and that the reported 89%/75% numbers come from a fair split, not from patient leakage, class imbalance, or tuning on the test set.","fun_headline_variants_meta":{"raw":{"variants":["Temporal attention boosts stroke occlusion AI to 89% precision","Watching DSA sequence, not frames, spots strokes at 89% precision","Temporal attention model finds stroke blockages with 89% precision","OccluNet: sequence-level AI beats single-frame detection for strokes","AI that watches DSA video finds blockages better than frame-by-frame"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000396,"raw_usage":{"total_tokens":1902,"prompt_tokens":728,"completion_tokens":1174,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":472,"completion_tokens_details":{"reasoning_tokens":1078}},"tokens_in":472,"tokens_out":1174,"duration_ms":8977,"temperature":1.0,"reasoning_tokens":1078,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T18:38:17.533115+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train OccluNet and a YOLOv11 frame-wise baseline on the same MR CLEAN Registry data, but with a strict patient-level train/test split; if the precision-recall gap shrinks to less than a few points, the temporal-attention advantage is an artifact of information leakage. Also, if a YOLOv11 trained on the identical frames with identical augmentation and test-time settings matches or beats OccluNet, the claim that sequence modeling helps would be refuted.","supporting_citations":[],"review_version":1}