REVIEW 3 major objections 7 minor 3 references
Fully automated AI pipeline delivers TAVI measurements in minutes
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-03 10:52 UTC pith:J3NMDSUA
load-bearing objection A genuinely deployable automated TAVI measurement pipeline with solid area/perimeter agreement, but the valve-size prediction is reported from the better of two folds and should not be taken as an unbiased accuracy estimate. the 3 major comments →
TAVI-TEC: An AI-Based Tool for Procedural Planning of Transcatheter Aortic Valve Implantation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, TAVI-TEC automates the full chain: deep-learning segmentation of the left ventricle, aorta, coronary arteries and calcifications; centerline extraction; identification of the annular plane as the plane of minimum cross-sectional area after small rotational perturbations; derivation of annular area, perimeter, diameters, eccentricity, sinus/STJ heights and coronary heights; color-coded vascular access maps; and a multilayer perceptron that predicts implanted valve size from five annular features. In 109 patients the automated annular area agreed with clinician measurement with CCC=0.934, ICC=0.935 and R²=0.881; perimeter with CCC=0.909 and ICC=0.909; diameter agreeme
What carries the argument
The load-bearing mechanism is the automated definition of the annular plane. Starting from the point where the extracted vessel centerline exits the left-ventricle mask, the algorithm places a plane perpendicular to the centerline tangent, then refines it by translating along the centerline and rotating up to 5 degrees about two orthogonal axes, finally selecting the plane with the minimum cross-sectional area. All downstream measurements — area, perimeter, diameters, heights, and the five features feeding the size classifier — are computed relative to this plane, so its accuracy bounds the accuracy of everything downstream.
Load-bearing premise
The central assumption is that the automatically chosen plane of minimum cross-sectional area is the same annular plane an experienced clinician would use for sizing; the paper does not validate this plane independently but only compares final measurements with a single clinician's readings.
What would settle it
Measure the angular and translational difference between TAVI-TEC's automatically selected annular plane and a consensus plane defined by three expert operators from the three leaflet nadirs in the same 109 CT scans. If the automated plane deviates systematically by more than a few degrees or by more than the inter-operator spread, the reported agreement on area and perimeter would be expected to degrade on a different population, and the size predictions would inherit that error. Alternatively, compute segmentation Dice scores against manual ground truth; if segmentation is poor in calcified
If this is right
- TAVI planning could shift from manual multiplanar reformatting to a fully automated first-pass annotation, freeing clinician time.
- The same pipeline could reduce inter-operator variability in annular measurements, a known source of sizing discrepancies.
- Vascular access assessment gains automated curved planar reformations and calcification maps, potentially flagging high-risk access routes earlier in the workflow.
- The 82% sizing accuracy with errors concentrated in adjacent sizes suggests the tool could act as a safety check on manual sizing rather than an autonomous decision-maker.
- If extended to other valve platforms and validated multicentrically, the same architecture could generalize beyond the SAPIEN 3 Ultra.
Where Pith is reading between the lines
- The minimum-area plane rule is a conservative sizing convention; if it were validated against an independent expert consensus, it could itself become a standard definition of the annulus, making automated measurements more reproducible than manual ones.
- Because the segmentation masks are not independently validated, a systematic segmentation bias in calcified or bicuspid anatomy would propagate invisibly into every measurement; testing segmentation quality on challenging cases is a cheap next step.
- The size classifier was trained and tested on 109 patients with only three valve sizes; a testable extension is to add more size categories and report calibration, not just accuracy, on a held-out multicenter cohort.
- The 10% undersizing versus 8% oversizing asymmetry is worth examining clinically, since undersizing risks paravalvular leak while oversizing risks annular rupture; a larger study could determine whether the plane-selection rule biases toward smaller sizes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents TAVI-TEC, a fully automated AI pipeline integrated into a web-based DICOM viewer for pre-procedural TAVI planning. The pipeline performs deep-learning segmentation of cardiovascular structures, calcification detection, centerline extraction, landmark identification, and annular plane definition, yielding annular and aortic root measurements, vascular access maps, and a multilayer perceptron (MLP) classifier for SAPIEN 3 Ultra (S3U) valve-size prediction. In a single-center cohort of 109 patients, the authors report strong agreement with clinician measurements for annular area (CCC=0.934, ICC=0.935, R²=0.881) and perimeter (CCC=0.909, ICC=0.909, R²=0.854), modest agreement for annular diameters, and 82% overall accuracy for valve-size prediction from the better of two cross-validation folds. The paper positions TAVI-TEC as a deployable workflow solution for automated TAVI planning, with runtime 2–6 minutes and no user interaction.
Significance. If the claims hold, TAVI-TEC would be a meaningful contribution to the TAVI planning workflow, potentially reducing operator variability and improving efficiency. The paper's strengths include a fully automated pipeline, integration into a realistic clinical software platform, runtime competitive with prior work, and quantitative agreement metrics for key annular measurements. The authors also transparently acknowledge limitations such as single-center evaluation, single valve platform, lack of formal segmentation validation, and the use of a single expert clinician as reference. However, the central valve-sizing claim rests on a upward-biased evaluation protocol (best fold of two-fold cross-validation, no independent hold-out), and the measurement claims depend on unvalidated segmentation and an unvalidated annular-plane heuristic; these issues must be addressed before the tool's clinical utility can be assessed. The paper is therefore scientifically promising but currently requires substantive revision to support its strongest claims.
major comments (3)
- [§2.2.3, Results, Figure 7] The valve-size prediction accuracy of 82% is reported from 'the best-performing fold selected for final evaluation' of a stratified two-fold cross-validation. With only 109 patients and three classes, this selection systematically overestimates expected performance, and the other fold's metrics are not reported. No independent test set exists, and the ground-truth labels are the sizes implanted by a single center's Heart Team, reflecting local clinical decisions rather than a standardized optimal-size label. The conclusion that TAVI-TEC predicts S3U size with 82% accuracy is therefore not supported as an unbiased estimate. Please report both folds, provide mean and standard deviation across repeated stratified splits, or use a proper hold-out set, and discuss how the choice of fold affects the AUC and accuracy claims.
- [§2.2.1, §4 Discussion] The entire measurement pipeline depends on deep-learning segmentation masks, yet the paper states in the Discussion that 'the accuracy of the deep learning-based segmentation was not formally evaluated.' All annular area, perimeter, diameters, heights, and calcification volumes are derived from these masks, so unvalidated segmentation error propagates into every reported measurement and into the MLP features. The agreement between TAVI-TEC and the clinician's measurements is an indirect check, but it cannot isolate segmentation error from plane-selection error or reveal systemic bias in specific structures (e.g., LVOT, aortic root, coronary arteries). Please provide quantitative segmentation evaluation against manual or expert-derived ground-truth masks (e.g., Dice coefficient, surface distance) for the key structures, or clearly reframe the claims as feasibility results pending such val
- [§2.2.1, §4 Discussion] The annular plane is defined as the plane of minimum cross-sectional area after translations and rotations of up to 5 degrees about the centerline, starting from the point where the centerline exits the left-ventricle mask. This heuristic is central to all downstream measurements and sizing, but it is never validated against an independent standard, such as an expert-annotated annular plane or a consensus of multiple operators. The authors note in the Discussion that discrepancies with manual measurements are 'likely related to differences in annular-plane identification,' which is precisely the unvalidated component. Please provide a direct comparison of plane orientation and position (e.g., angular deviation, center displacement) against expert-defined planes, or present a sensitivity analysis showing that the ±5-degree range covers the expert plane. Without such evidence, the claim th
minor comments (7)
- [Table 2] The table title says 'Average aortic root height measurements (mm)' but the first row is labeled 'Sinus height' and the values are given as 22.3 ± 4.2; the table appears to report heights, not diameters, despite the text referencing 'sinus diameter' in the Discussion. Please clarify the variable names and ensure consistency.
- [Table 3] The table header uses 'interclass correlation coefficient' in the note; the correct term is 'intraclass correlation coefficient' (ICC). Please correct.
- [§2.2.3] The MLP is described as using 'five input features, corresponding to annular measurements,' but the exact features are not specified. Please list the specific measurements (e.g., area, perimeter, Dmax, Dmin, eccentricity) to allow reproducibility.
- [§2.2.3] The MLP hyperparameters (number of neurons, L2 regularization strength, LBFGS convergence criteria) are described qualitatively. For reproducibility, report the exact values or provide code/implementation details.
- [Figure 7C] The text states 'Undersizing ... occurred in 10% of cases, whereas oversizing ... was observed in 8% of cases.' These percentages appear to sum to 18%, leaving 82% exact agreement, but the figure should clarify whether these are percentages of the total cohort or of the test fold, and how many patients fall in each category.
- [Discussion, 5th paragraph] The sentence 'The highest misclassification rate was observed for the 29-mm valve' contradicts the reported F1-score of 0.91 for the 29-mm class, which is the highest among the three classes. Likely the intended statement is about the 26-mm valve (F1=0.73). Please verify.
- [References] Reference [17] contains a corrupted author list ('O. Backer' instead of 'De Backer') and a future year (2026). Please update to the published version.
Circularity Check
Valve-size 'prediction' accuracy is the better of two cross-validation folds — a selected, partly fitted statistic; the measurement-vs-manual claims are externally compared and not circular.
specific steps
-
fitted input called prediction
[Section 2.2.3 (Machine Learning) and Results (MLP performance, Figure 7)]
"Model performance was evaluated using stratified two-fold cross-validation while preserving class distribution across folds. The best-performing fold was selected for final evaluation. ... The multilayer perceptron achieved an overall accuracy of 82% in the best-performing cross-validation fold."
The headline 82% accuracy (and class AUCs) is, by construction, the better of exactly two fold-specific runs (each ~54 patients), with the unselected fold's metrics never reported. Selecting the best fold and presenting it as 'overall accuracy' makes the reported number a maximally fitted statistic rather than an out-of-fold estimate of predictive performance. In addition, the model is trained and evaluated on the same single-center cohort, and its target labels are the implanted S3U sizes, i.e., the local Heart Team's sizing decisions derived from the same annular measurements that form the model's input features; the reported performance is therefore partly a fit to the local decision rule rather than an independent predictive claim.
full rationale
The core measurement claims are not circular: automated annular area, perimeter, and diameters are compared against independently performed clinician-derived manual measurements (CCC, ICC, R2, Bland-Altman), so the agreement statistics are genuine external validations rather than reconstructions of the inputs. No equation in the pipeline defines an output in terms of its own fitted value, and no 'uniqueness theorem' is imported from prior work to force a choice. The self-citations present ([19] for the segmentation modules, [21] for the MLP's earlier comparison against SVM/kNN) are descriptive support and are not load-bearing: the current MLP's performance rests on its own reported results, and the segmentation is explicitly acknowledged not to have been formally evaluated. The one step approaching circularity is the MLP evaluation protocol: the paper states that 'the best-performing fold was selected for final evaluation' and reports '82% overall accuracy in the best-performing cross-validation fold,' so the headline figure is by construction the maximum of two fold fits rather than an unbiased estimate of prediction error. This, together with single-cohort training/evaluation and labels being the locally implanted sizes, makes the 82% figure partly fitted rather than independently predictive. These are statistical-validity and generalizability limitations rather than definitional circularity, and the central measurement-vs-manual comparison retains independent content; score 4 reflects the one partially forced evaluation claim.
Axiom & Free-Parameter Ledger
free parameters (3)
- Calcification threshold (99.5th percentile of HU within ROI) =
99.5th percentile (no absolute HU)
- Annular-plane optimization range (±5 degrees) =
5 degrees
- MLP hyperparameters
axioms (4)
- domain assumption The cited segmentation model [19] generalizes to TAVI CTA scans and yields accurate LV/aorta/coronary masks.
- domain assumption The centerline transition point where the trajectory exits the LV mask marks the anatomical annulus.
- ad hoc to paper The minimum-cross-sectional-area plane after ±5-degree perturbations is clinically equivalent to the operator-selected annular plane.
- domain assumption The implanted S3U valve size is the correct ground-truth label for classifier training and evaluation.
read the original abstract
Computed tomography angiography (CTA) is crucial for preprocedural TAVI planning, providing the anatomical information required for prosthesis sizing and vascular access assessment. As the volume of TAVI procedure increases, improving efficiency and standardizing annotations is becoming essential in clinical practice. This study presents TAVI-TEC, a fully automated artificial intelligence-based framework integrated into a web based DICOM viewer for routine preoperative TAVI planning. Pre-procedural CTA scans from patients undergoing TAVI with SAPIEN 3 Ultra (S3U) prostheses were processed using a fully automated pipeline. Deep learning-based segmentation of cardiovascular structures, calcification detection, centerline extraction, landmark identification, and annular plane definition was implemented to quantify key annular and aortic root measurements and color-coded maps of lumen reduction and vessel diameter for vascular access. A multilayer perceptron classifier was trained to predict prosthesis size prior to the TAVI procedure. Results revealed that TAVI-TEC enabled pre-procedural measurements in approximately 2-6 min. Strong agreement with clinician-derived measurements was observed for annular area (coefficient of concordance, CCC = 0.934; interclass correlation coefficient, ICC = 0.935; R^2 = 0.881) and perimeter (CCC = 0.909; ICC = 0.909; R^2 = 0.854). The valve-size prediction model achieved 82% overall accuracy, with most misclassifications occurring between adjacent prosthesis sizes. Though further multicenter validation and extension to additional measurements and valve platforms are required, the TAVI-TEC methodology may reduce operator variability in pre-TAVI measurements and streamline the preoperative workflows of the Heart Team for decision-making.
Figures
Reference graph
Works this paper leans on
-
[1]
Initially, TAVI was limited to patients at high surgical risk, but indications were progressively expanded to include patients at intermediate and low surgical risk [3, 4, 5]
INTRODUCTION Transcatheter aortic valve implantation (TAVI) has become a less invasive alternative to open-heart surgery for patients with severe aortic stenosis and is increasingly performed worldwide [1, 2]. Initially, TAVI was limited to patients at high surgical risk, but indications were progressively expanded to include patients at intermediate and ...
-
[3]
ANTHEM: AdvaNced Technologies for Human-centrEd Medicine
RESULTS Figure 2 shows screenshots of the TAVI-TEC application implemented in the web-based DICOM_Vision medical imaging software developed by D/Vision Lab. Figure 3 illustrates the output of the TAVI-TEC application, including three-dimensional visualization of the aortic root and corresponding pre-procedural TAVI measurements. These measurements include...
2010
-
[17]
Cannata, I
S. Cannata, I. Sultan, N.M. Van Mieghem, A. Giordano, O. Backer, J. Byrne, D. Tchetche, S. Buccheri, L. Nombela-Franco, R. Campante Teles, M. Barbanti, E. Barbato, I. Amat Santos, D.J. Blackman, F. Maisano, R. Lorusso, K. La Spina, A. Millin, D.E. Kliner, M. van den Dorpel, E. Acerbi, D. Lulic, H. Fayed, C. De Biase, J.F. Chavez Solsol, J. Brito, G. Costa...
2026
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.