{"id":"8e905e17-1059-4ddf-b322-d67d67baca0b","arxiv_id":"2605.24012","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A two-stage deep learning framework automates TMPFC calculation from angiography and matches expert manual measurements with r=0.98 in a 655-patient multi-center cohort.","lead":"This paper develops a deep learning system called DL-TMPFC to automatically calculate the TIMI Myocardial Perfusion Frame Count from routine coronary angiograms. It aims to provide an objective, fast way to detect coronary microvascular dysfunction without manual counting or extra invasive tests.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Territory-aware segmentation and first/last frame detection accuracy under real-world angiographic variations is the unverified linchpin for all reported agreement metrics","rationale":"The reader's weakest assumption correctly isolates the single component whose failure would falsify every downstream number; the full-text availability does not remove the need for explicit segmentation validation metrics or external testing, so the UNVERDICTED stance remains appropriate.","tokens_in":1890,"tokens_out":354,"duration_ms":18604,"concrete_test":"Run the territory-aware segmentation network on a fresh 100-sequence multi-center hold-out set with independent expert territory masks and frame annotations; if mean Dice per territory falls below 0.80 or mean absolute frame-selection error exceeds 2 frames, recompute the DL-TMPFC vs. manual agreement on those cases and check whether r drops below 0.90 or bias exceeds 2 frames.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (bias -0.93 frames, 95% LoA -5.33 to +3.47, r=0.98, plus accurate CMVD identification across pathologies) requires that the territory-aware segmentation network correctly assigns perfusion territories and selects the first/last frames on every sequence. Any consistent error in territory definition or frame choice (e.g., due to projection angle, contrast timing, or anatomy outside the three-institution 655-patient set) propagates directly into TMPFC values and therefore into both the agreement statistics and the clinical classification performance. The abstract supplies no segmentation metrics (Dice/IoU per territory), no ablation on frame-selection error, and no external-site testing of the full pipeline, leaving this assumption untested.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript presents DL-TMPFC, a deep learning framework for automated TMPFC quantification from coronary angiography to assess CMVD. It comprises a stenosis detection network to exclude obstructive CAD, a territory-aware segmentation network to identify perfusion territories, and an automatic first/last frame determination module. Validation on a 655-patient multi-center cohort (445 obstructive CAD, 100 confirmed CMVD, 110 controls) reports excellent agreement with manual expert measurements (bias -0.93 frames; 95% LoA -5.33 to +3.47; r=0.98) and accurate CMVD identification across pathologies with continuous severity capture.","tokens_in":2034,"tokens_out":483,"duration_ms":21949,"significance":"If the segmentation and frame-selection components prove robust, the work could enable objective, observer-independent TMPFC measurement in routine angiography, supporting quantitative CMVD risk stratification and clinical translation beyond subjective TIMI grading. The multi-center cohort size and reported agreement statistics provide empirical grounding for the automation claim.","major_comments":[{"comment":"Methods (DL-TMPFC framework description): The territory-aware segmentation network and first/last frame detection are load-bearing for all downstream TMPFC values and agreement metrics, yet no segmentation performance metrics (Dice/IoU per territory), frame-selection error rates, or ablation studies on imaging variations (projection angle, contrast timing) are reported. This leaves the central assumption untested and directly undermines confidence in the bias/LoA/r values and CMVD classification results.","section":"Methods (DL-TMPFC framework)"},{"comment":"Results (validation on 655-patient cohort): No details are provided on training/validation splits, external-site testing of the full pipeline, or handling of edge cases/anatomies outside the three-institution set. Any systematic territory or frame error would propagate into the reported agreement statistics, making the generalizability claim difficult to evaluate.","section":"Results (validation cohort)"}],"minor_comments":[{"comment":"Abstract: The statement that DL-TMPFC 'accurately identified CMVD across a full spectrum of coronary pathologies' would benefit from explicit reference to the specific performance metrics (e.g., sensitivity/specificity or correlation with continuous severity) supporting this claim.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive and detailed review. The comments highlight important aspects of transparency for the DL-TMPFC framework. We respond to each major comment below and indicate planned revisions where appropriate.","responses":[{"response":"We agree that intermediate performance metrics would increase transparency and allow readers to better assess potential error propagation. In the revised manuscript we will report per-territory Dice and IoU scores for the segmentation network, frame-selection accuracy (mean absolute error and percentage of frames within acceptable tolerance), and a focused ablation on projection angle and contrast timing variations using the available multi-center data. The primary clinical validation remains the end-to-end agreement with expert TMPFC measurements, which directly tests the quantity of interest; however, the additional metrics will strengthen the supporting evidence.","revision_made":"yes","referee_comment":"Methods (DL-TMPFC framework): The territory-aware segmentation network and first/last frame detection are load-bearing for all downstream TMPFC values and agreement metrics, yet no segmentation performance metrics (Dice/IoU per territory), frame-selection error rates, or ablation studies on imaging variations (projection angle, contrast timing) are reported. This leaves the central assumption untested and directly undermines confidence in the bias/LoA/r values and CMVD classification results."},{"response":"The 655-patient cohort was collected from three independent institutions specifically to improve diversity. We will expand the Methods section to explicitly state the training/validation split ratios and any site-stratified partitioning used during model development. While the multi-center design already incorporates data from separate sites, we will add a leave-one-site-out analysis where computationally feasible. Edge cases (e.g., anomalous coronary anatomy, suboptimal contrast opacification) will be illustrated with representative examples and failure-mode discussion. These additions will be included in the revision.","revision_made":"partial","referee_comment":"Results (validation on 655-patient cohort): No details are provided on training/validation splits, external-site testing of the full pipeline, or handling of edge cases/anatomies outside the three-institution set. Any systematic territory or frame error would propagate into the reported agreement statistics, making the generalizability claim difficult to evaluate."}],"tokens_in":1547,"tokens_out":477,"duration_ms":22809,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core contribution is a pipeline that detects stenoses, segments perfusion territories, and auto-selects frames to compute TMPFC without manual effort. On 655 patients across three sites it reports tight agreement with expert readers (bias -0.93 frames, limits of agreement roughly -5 to +3, r=0.98) and claims it can flag CMVD across obstructive and non-obstructive cases while giving a continuous severity score.\n\nThat agreement number and the multi-center scale are the parts worth noting. Automating an existing but tedious angiography metric could matter in cath labs where subjective TIMI grading is still common.\n\nThe weak point is exactly what the stress-test note flags: everything rides on the territory-aware segmentation and first/last-frame module working reliably. The abstract supplies no Dice or IoU numbers for the segmentation, no ablation on frame-choice errors, and no external-site test of the full pipeline. Any consistent mistake in territory assignment would move the TMPFC values and therefore the agreement statistics. Without those checks the numbers are hard to interpret.\n\nThis is the sort of applied clinical paper that might interest interventional cardiologists or imaging groups working on microvascular disease. It is coherent on its own terms and shows honest engagement with the clinical gap, so it deserves a serious referee even if the methods section will need expansion on validation and edge cases. I would not desk-reject it.","headline":"The paper automates TMPFC calculation via DL on a 655-patient multi-center set with strong reported agreement, but the abstract leaves the critical segmentation and frame-selection accuracy untested.","tokens_in":2533,"tokens_out":363,"would_cite":false,"duration_ms":15488,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Deep learning framework automates TMPFC calculation to quantify coronary microvascular dysfunction from routine angiography.","keywords":["deep learning","coronary angiography","microvascular dysfunction","TIMI frame count","automated quantification","perfusion assessment","CMVD"],"falsifier":"A new multi-center cohort in which DL-TMPFC values deviate from simultaneous manual expert counts by more than the reported limits of agreement on average.","tokens_in":2789,"feed_emoji":"🫀","tokens_out":642,"duration_ms":18737,"temperature":0.7,"pith_summary":"The paper develops and validates DL-TMPFC, a deep learning system that automatically computes the TIMI Myocardial Perfusion Frame Count from coronary angiograms after first excluding obstructive disease. It combines a stenosis detection network with a territory-aware segmentation network to locate perfusion territories and identify the start and end frames of contrast flow. Tested on 655 patients across three centers, the system matched expert manual counts with bias of -0.93 frames and correlation 0.98, while distinguishing microvascular dysfunction across obstructive, non-obstructive, and control cases and tracking its continuous severity. The automation removes manual calculation steps and observer variability, allowing TMPFC to move from research tool to routine clinical measure. This supplies an objective, angiography-based metric for a condition that currently lacks simple bedside quantification.","feed_headline":"Deep learning automates TMPFC to measure microvascular dysfunction","feed_subtitle":"System matches manual counts with 0.98 correlation and flags dysfunction across all coronary disease types from routine angiograms.","key_machinery":"Territory-aware segmentation network that locates perfusion territories and marks the first and last frames of myocardial contrast appearance for TMPFC computation.","core_discovery":"DL-TMPFC delivers automatic and objective TMPFC values directly from standard angiographic sequences, achieving excellent agreement with manual expert readings and correctly classifying microvascular dysfunction across the full range of coronary artery disease presentations.","pith_inferences":["Hospitals could embed the model in existing cath-lab software so that TMPFC appears automatically in the procedural report.","Longitudinal tracking of the same patient’s TMPFC values before and after therapy becomes practical for monitoring treatment response.","The framework’s exclusion of obstructive disease first may allow combined reporting of epicardial and microvascular status from one study."],"forward_implications":["TMPFC becomes feasible as a standard reportable value on every diagnostic angiogram without added procedure time or staff effort.","Clinicians gain a continuous numeric score rather than a binary yes/no for microvascular dysfunction, supporting graded risk assessment.","Observer-to-observer differences in perfusion grading disappear because the calculation is fully deterministic once the images are acquired.","The same angiographic run used for stenosis assessment now also yields a microvascular metric without extra contrast or radiation."],"fun_headline_variants":["Deep learning automates TMPFC quantification from angiography","DL-TMPFC matches manual TMPFC measurements with high accuracy","Automated framework calculates TMPFC for microvascular dysfunction","AI delivers TMPFC directly from routine angiographic sequences"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The segmentation network will correctly define perfusion territories and mark first and last frames even when imaging angles, patient anatomy, or equipment vary from the training data.","fun_headline_variants_meta":{"raw":{"variants":["Deep learning automates TMPFC quantification from angiography","DL-TMPFC matches manual TMPFC measurements with high accuracy","Automated framework calculates TMPFC for microvascular dysfunction","AI delivers TMPFC directly from routine angiographic sequences"]},"model":"grok-4.3","cost_usd":0.005344,"raw_usage":{"total_tokens":2626,"prompt_tokens":762,"num_sources_used":0,"completion_tokens":60,"cost_in_usd_ticks":53437000,"prompt_tokens_details":{"text_tokens":762,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1804,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":762,"tokens_out":60,"duration_ms":13620,"temperature":1.0,"reasoning_tokens":1804,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T17:53:42.501274+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A new multi-center cohort in which DL-TMPFC values deviate from simultaneous manual expert counts by more than the reported limits of agreement on average.","supporting_citations":[],"review_version":1}