{"id":"3a7195c9-d730-46f3-a0db-c774e0138be8","arxiv_id":"2507.05883","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A deep-learning pipeline using dynamic time warping automatically matches IVUS and OCT coronary images, reaching expert-level agreement in longitudinal and rotational alignment.","lead":"This paper presents a fully automated deep-learning system that aligns two types of coronary artery imaging, ultrasound (IVUS) and optical coherence tomography (OCT), frame by frame and rotationally. It could let researchers fuse complementary views of plaque automatically instead of having experts spend many minutes per vessel.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Longitudinal accuracy may be driven by the dominant normalized-frame-position feature, which mirrors the linear-interpolation ground truth; reported CCC may measure agreement with an interpolation convention rather than true anatomical matching.","rationale":"The paper is a careful integration of established components and has genuine strengths: a large, patient-disjoint test set, expert ground truth, and quantitative comparison against inter-observer variability. The weakest point is precisely the one the reader identified: the normalized-frame-position feature is weighted most heavily in the DTW, and the ground truth itself was generated by linear interpolation between landmarks. This creates a risk that the evaluation measures agreement with the analysts' interpolation convention rather than independent anatomical accuracy. A trivial proportional-index baseline test would settle whether the image-derived features add value beyond that linear prior. The manual segment-of-interest definition additionally qualifies the 'fully automated' claim, but the accuracy concern is the more load-bearing issue for the central claim. Because the reader already issued a CONDITIONAL verdict and requested sensitivity analysis, this stress-test reinforces that verdict rather than moving it. The concern is not an accusation of misconduct; it is a request to demonstrate that the reported concordance is not an artifact of matching the reference's interpolation scheme.","tokens_in":13608,"tokens_out":5993,"duration_ms":74492,"concrete_test":"Recompute the Table 3 longitudinal registration metrics on the 77-vessel test set using a trivial baseline that maps NIRS-IVUS to OCT frames by proportional frame index over the expert-defined segment of interest, without using any image-derived features or DTW. If this baseline achieves a CCC and Williams Index statistically indistinguishable from the reported values, then the headline accuracy is dominated by the linear interpolation prior rather than by the proposed deep-learning features.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is in the interaction between Section 2.5.2 and Section 2.4. The DTW longitudinal alignment weights 'normalized frame position' at 2.5, the highest of the four feature weights, while the expert ground truth was constructed by identifying sparse anatomical landmarks and applying linear interpolation between them. A strong linear position prior therefore tends to reproduce the same piecewise-linear mapping that defines the reference, making the reported CCC>0.99 and Williams Index 0.96 partly self-fulfilling. The evaluation does not establish that the image-derived lumen, side-branch, and calcification features are locating true anatomical correspondence; it may only show that the method agrees with the analysts' interpolation convention. This risk is compounded by Section 2.3, where an expert manually defines the segment of interest using anatomical landmarks; that manual step supplies the endpoints that make normalized frame position meaningful, so the 'fully automated' claim is not yet demonstrated outside that step. If the linear prior fails on data with different pullback speeds, heart-rate-dependent end-diastolic sampling, or imperfect segment matching, the reported concordance could drop substantially. The circumferential result is less directly affected but inherits any longitudinal errors because registration is performed sequentially.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a fully automated deep-learning pipeline for co-registering NIRS-IVUS and OCT pullbacks, using automated detection of lumen borders, side branches, and calcific tissue followed by dynamic time warping for longitudinal alignment and a dynamic programming rotation search for circumferential alignment. The framework is trained and evaluated on 714 vessels from the PACMAN-AMI trial, with a patient-level split into training, validation, and test sets. On the 77-vessel test set, the authors report high concordance with expert analysts (CCC > 0.99 longitudinal, > 0.90 circumferential), a Williams Index of 0.96 and 0.97, and an execution time under 90 seconds per vessel. The paper claims this is the first fully automated framework to overcome limitations of prior semiautomated registration methods.","tokens_in":13817,"tokens_out":2271,"duration_ms":24899,"significance":"If the reported accuracy is genuine, this framework would be a practical and valuable research tool: it is fast, leverages a large expert-annotated dataset, uses a reasonable patient-level evaluation split, and directly addresses a known bottleneck in multimodality intravascular imaging studies. The independent test set and the comparison against inter- and intra-observer variability are strengths. However, the central validity of the longitudinal evaluation is compromised by the design interaction between the dominant 'normalized frame position' feature and the piecewise-linear-interpolation ground truth, so the headline concordance may partly reflect agreement with an interpolation convention rather than with true anatomical matching. The manual definition of the segment of interest in Section 2.3 also weakens the 'fully automated' claim. These concerns must be resolved before the performance claims can be accepted.","major_comments":[{"comment":"The longitudinal registration uses a 'normalized frame position' feature with weight 2.5, the highest among four features, while the expert ground truth in Section 2.4 was created by identifying sparse anatomical landmarks and applying linear interpolation between them to match the remaining frames. A strong linear position prior will tend to reproduce that same piecewise-linear mapping, making the reported CCC > 0.99 and Williams Index 0.96 partly self-confirming. The evaluation does not establish that the image-derived features (lumen area, side branches, calcification) are locating true anatomical correspondence beyond what the position prior alone would achieve. Please provide an ablation study that removes the normalized frame position feature, or an alternative evaluation that compares the DL output against point-wise anatomical landmark matches rather than the interpolated mapping; without this, the longitudinal accuracy claim is not adequately supported.","section":"Section 2.5.2 and Section 2.4"},{"comment":"The paper states in Section 2.3 that an expert analyst defined each segment of interest (SOI) using anatomical landmarks visible in angiography, NIRS-IVUS, and OCT. The automated pipeline then operates only within these manually defined SOIs. As a result, the framework is not 'fully automated' in the sense claimed in the Discussion (Section 4): the manual SOI definition supplies the endpoints that make the normalized frame position feature meaningful, and it may also remove cases or segments where the two pullbacks cover substantially different ranges. Please clarify the exact input required from the operator, and either report the framework's sensitivity to the SOI definition or present results in a setting where the SOI is derived automatically (e.g., from the angiographic pullback ranges). At minimum, the claim should be softened to 'fully automated within a manually defined segment of interest.'","section":"Section 2.3 and Section 4"},{"comment":"The Williams Index is reported as 0.96 and 0.97 with 95% confidence intervals that appear to include 1.00 in Table 3 (e.g., 0.94–1.00 for longitudinal). The text in Section 3.4 interprets this as 'comparable performance,' which is fair, but the abstract and conclusion state the method 'compares favorably' to experts. Since the point estimates are below 1 and the confidence intervals straddle 1, the evidence supports equivalence rather than superiority. Please adjust the wording to avoid overstating the result, and report whether the Wilcoxon comparisons in Section 3.4 are corrected for the multiple comparisons performed (the paper reports four p-values without any multiplicity adjustment).","section":"Section 3.4 and Supplement S1"}],"minor_comments":[{"comment":"The subsection is numbered 2.5.2 but appears after Section 2.6.1 in the text; it should be renumbered (likely 2.6.2) to follow the logical order of the methods.","section":"Section 2.5.2"},{"comment":"There is a typo: 'Sperman correlation coefficient' should be 'Spearman correlation coefficient.'","section":"Section 2.7"},{"comment":"The lumen segmentation training set is described as 61,665 NIRS-IVUS frames in Section 2.6.1, but the abstract states 61,655. Please verify the correct number and use it consistently.","section":"Section 2.6.1"},{"comment":"The figure caption and legend use repeated colors ('red, orange, blue, red') which is likely a typo; one of the 'red' entries should probably be a different color. Please correct for clarity.","section":"Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The core concern is the potential circularity between the dominant normalized-frame-position feature and the linear-interpolation ground truth. This is a scientific-validity issue that a careful ablation or landmark-wise evaluation can address, so I recommend major revision rather than rejection. The authors should also be asked to state clearly the manual SOI input, since the 'fully automated' claim is central to the paper's novelty. The dataset and the split are strengths; the paper is well positioned if these concerns are resolved."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper is worth reading because it tackles a real bottleneck: co-registering IVUS and OCT pullbacks for multimodality plaque analysis. The authors assemble an ensemble of DL models (lumen segmentation, side-branch detection, calcification classification) and feed their outputs into a DTW longitudinal alignment and a dynamic-programming circumferential alignment. On 77 test vessels with expert ground truth from two analysts, they report CCC >0.99 longitudinal and >0.90 circumferential, and Williams Index around 0.96–0.97, at under 90 seconds per vessel. The patient-level split and separate validation set are proper, and the feature extraction accuracies (DSC ~0.96–0.98 lumen; AP 0.58–0.74 side branch; F1 ~0.68–0.88 calcium) are credible. This is a solid integration, more comprehensive than the prior semi-automatic methods, and the demonstration on a large clinical dataset is a step forward.\n\nThe main soft spot is the longitudinal alignment. The 'normalized frame position' feature (frame index normalized by pullback length) is weighted 2.5, the highest of the four longitudinal features. The expert ground truth was built by identifying sparse side-branch/calcification landmarks and then linearly interpolating between them. Strongly weighting normalized frame position essentially forces the DTW path toward a linear mapping, which matches the ground truth's interpolation convention by construction. So the high CCC may be as much about reproducing that convention as about finding true anatomical correspondence. The paper does not report what happens if the position prior is removed or down-weighted, so we don't know how much the image-derived features actually contribute. Also, Section 2.3 describes a manual step: an expert defines the segment of interest using anatomical landmarks, which anchors the endpoints and makes normalized frame position meaningful. That manual step softens the 'fully automated' claim.\n\nThese are not fatal; the circumferential result is less affected, and the overall pipeline is still useful for research workflows. But the evaluation should include a sensitivity analysis of the feature weights, ideally on test data with mismatched pullback lengths or missing landmarks, and the limitations section should state explicitly that the SOI definition remains manual. Code and data release would help independent checking.\n\nI agree with the conditional verdict. The paper deserves thorough peer review, but it needs another revision to support the 'fully automated' and 'accurate matching' claims.","headline":"A serious, well-validated integration of DL feature extraction with DTW for IVUS-OCT co-registration, but the top-weighted position feature and manual SOI step mean the 'fully automated' claim is weaker than presented.","tokens_in":14412,"tokens_out":3378,"would_cite":true,"duration_ms":36463,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A fully automated deep-learning pipeline can co-register IVUS and OCT coronary images to expert-level accuracy in under 90 seconds per vessel.","keywords":["intravascular ultrasound","optical coherence tomography","co-registration","deep learning","dynamic time warping","coronary plaque imaging","multimodality imaging","image registration"],"falsifier":"Run the pipeline on paired pullbacks where the OCT pullback speed differs from the 36 mm/s used here, or where the two catheters cover visibly different segment lengths, and compare against expert landmark-based matching: if the high longitudinal CCC (above 0.99) drops, the normalized-frame-position prior is doing the work rather than the anatomical features. A second check is to examine segments with few side branches and little calcium, where the 1-to-1 DTW constraint and the frame-position prior are nearly the only signals.","tokens_in":1756,"feed_emoji":"🫀","tokens_out":4470,"duration_ms":126313,"temperature":0.7,"pith_summary":"This paper claims that co-registering two complementary intravascular imaging modalities — ultrasound (IVUS) and optical coherence tomography (OCT) — can be made fully automatic, fast, and as accurate as expert human analysts. The authors build a deep-learning pipeline that extracts lumen borders, side-branch origins, and calcific tissue from both image types, then aligns the two pullbacks along the vessel with dynamic time warping and rotationally aligns matching frames with dynamic programming. On a test set of 77 vessels, the automated alignment agreed with expert analysts at concordance levels above 0.99 for longitudinal position and above 0.90 for rotational orientation, with Williams Indices of 0.96 and 0.97, meaning the machine matches experts about as well as experts match each other. The whole pipeline runs in under 90 seconds per vessel. If true, this removes the manual work that currently forces multimodality trials to analyze IVUS and OCT separately, making hybrid plaque characterization feasible on large datasets.","feed_headline":"Automated pipeline matches expert-level IVUS-OCT co-registration","feed_subtitle":"Registers two coronary imaging modalities in 84 seconds with expert-level accuracy, enabling large-scale plaque studies.","key_machinery":"The load-bearing mechanism is a feature-extraction ensemble feeding two alignment algorithms. Three deep-learning components produce the registration features: a previously validated convolutional segmentation network retrained for lumen borders, a region-proposal object detector (Faster R-CNN style) for side-branch origins, and a Polar-UNet — a U-shaped network that operates on polar views of the lumen and uses self-attention in its decoder — to classify which circumferential angles contain calcium. Along the vessel axis, dynamic time warping (DTW) finds corresponding IVUS-OCT frame pairs by minimizing a feature-weighted Euclidean distance over sequences of lumen area, side-branch area, calcification degree, and normalized frame position; circumferentially, a rotation cost matrix is built by circularly sampling features around the lumen center (side-branch angle, calcification angle, lumen eccentricity), and a dynamic programming path with a shape-regularization term picks the rotation per frame. The normalized frame position feature, carrying the largest weight (2.5), is what anchors the longitudinal alignment when anatomical landmarks are sparse.","core_discovery":"The central claim is that a fully automated, multi-stage framework can match NIRS-IVUS and OCT images of the same coronary segment with accuracy indistinguishable from expert manual co-registration. Deep-learning networks first segment lumen borders, detect side-branch origins, and classify calcific arcs in every frame; a feature-weighted dynamic time warping then finds corresponding frames along the pullbacks using lumen area, side-branch area, calcification degree, and normalized frame position; and a dynamic-programming search over a rotation cost matrix aligns the circumferential orientation of the OCT frames to IVUS. The authors report concordance correlation coefficients above 0.99 (longitudinal) and above 0.90 (circumferential) against expert analysts on a test set of 77 vessels, Williams Indices of 0.96 and 0.97, and a runtime under 90 seconds per vessel, and they position the framework as the first to overcome the segmentation, manual-landmark, and validation limitations of earlier co-registration methods.","pith_inferences":["If the normalized-frame-position prior generalizes, the same feature-weighted DTW recipe could be ported to other paired pullback settings — such as different OCT pullback speeds or hybrid catheters — by learning new feature weights per acquisition protocol.","The dependence on side branches and calcific tissue as high-signal landmarks implies that accuracy in long smooth segments with no branches and no calcium would degrade toward the frame-position prior, so performance on such segments is a natural stress test.","The sequential longitudinal-then-circumferential design likely propagates temporal errors into rotational matching; a quantitative comparison against joint optimization would reveal how much error that ordering adds, a step toward the graph-matching approach the authors flag as future work.","A direct generalizability check would apply the pipeline to data from other IVUS or OCT vendors, or to pullbacks with different frame-sampling intervals, which the authors acknowledge is unaddressed."],"forward_implications":["Automated analysis of large multimodality imaging studies becomes feasible, since a vessel that takes an expert analyst about 576 seconds to register is processed by the pipeline in about 84 seconds.","Hybrid assessment of plaque — combining plaque burden and calcium from IVUS with fibrous-cap thickness from OCT at matched cross-sections — can be performed without manual landmark identification.","The framework runs on unsegmented image data, so it can be applied directly to raw pullbacks rather than requiring the separate segmentation step that earlier co-registration methods needed.","Registration results are fully reproducible, removing the inter-observer variability documented between and within expert analysts.","The feature-extraction stage supplies registration signals in both modalities without human annotation, with near-0.9 Dice/F1 performance for lumen, side branch, and calcium detection in IVUS and OCT."],"supporting_citations":[{"why":"Supplies all 230 patients and 714 vessels of PACMAN-AMI trial data used for feature training, validation, and testing.","marker":"(12)"},{"why":"The most robust prior IVUS-OCT co-registration framework, which this work extends and whose dynamic-programming rotation method it follows.","marker":"(17)"},{"why":"Supplies the automated end-diastolic frame detection that selects the NIRS-IVUS frames entering the registration pipeline.","marker":"(22)"},{"why":"The previously validated convolutional segmentation model retrained on this study's NIRS-IVUS and OCT frames for lumen-border extraction.","marker":"(25)"},{"why":"The region-proposal object detection architecture used to locate side-branch origins in bounding boxes.","marker":"(26)"},{"why":"The self-attention mechanism used in the Polar-UNet decoder for calcium arc classification.","marker":"(27)"},{"why":"The dynamic time warping algorithm that produces the longitudinal frame correspondence between the two pullbacks.","marker":"(28)"},{"why":"Supplies the shape-regularization term that constrains rotation changes between consecutive frames in the circumferential path.","marker":"(29)"},{"why":"The statistical methodology defining the Williams Index used to claim the method's performance is comparable to expert analysts.","marker":"(33)"}],"fun_headline_variants":["AI co-registers IVUS and OCT faster than experts","Deep learning matches experts on IVUS-OCT alignment","Fully automated IVUS-OCT co-registration in under 90s","AI matches expert accuracy for IVUS-OCT fusion"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The longitudinal alignment leans hardest on 'normalized frame position' — the assumption that a given frame sits at roughly the same relative spot along both pullbacks — so the match will drift if the two acquisitions cover different segment lengths, use different pullback speeds, or space frames differently.","fun_headline_variants_meta":{"raw":{"variants":["AI co-registers IVUS and OCT faster than experts","Deep learning matches experts on IVUS-OCT alignment","Fully automated IVUS-OCT co-registration in under 90s","AI matches expert accuracy for IVUS-OCT fusion"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000575,"raw_usage":{"total_tokens":2797,"prompt_tokens":1108,"completion_tokens":1689,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":724,"completion_tokens_details":{"reasoning_tokens":1620}},"tokens_in":724,"tokens_out":1689,"duration_ms":12516,"temperature":1.0,"reasoning_tokens":1620,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:16:23.548079+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the pipeline on paired pullbacks where the OCT pullback speed differs from the 36 mm/s used here, or where the two catheters cover visibly different segment lengths, and compare against expert landmark-based matching: if the high longitudinal CCC (above 0.99) drops, the normalized-frame-position prior is doing the work rather than the anatomical features. A second check is to examine segments with few side branches and little calcium, where the 1-to-1 DTW constraint and the frame-position prior are nearly the only signals.","supporting_citations":[],"review_version":1}