{"id":"e7371806-be7d-401e-8d4a-f810ce544f9c","arxiv_id":"2506.00727","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"AdaPR, an A3C-based deep reinforcement learning method with a local coordinate system, achieves plane reformatting and flow quantification in 4D flow MRI comparable to manual expert observers and robust to volume orientation changes.","lead":"This paper introduces AdaPR, a deep reinforcement learning system that places measurement planes in 4D flow MRI scans using a local coordinate system, making it robust to differences in scan orientation and position. It reports flow measurements that match manual expert readings within inter-observer variability, which could reduce time and user-dependence in cardiac flow analysis.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported plane accuracy and flow comparability are both anchored to Observer 1's manual planes, which serve as the reward target (Eq. 7) and the evaluation reference; without an independent multi-observer or phantom consensus, the central claim measures agreement with a single unvalidated standard.","rationale":"The paper's central claim is that AdaPR provides robust, orientation-independent plane reformatting and flow quantification comparable to expert observers. For this claim to hold, the evaluation must reflect anatomical correctness, not merely agreement with one observer's annotation. The reward in Eq. (7) and the primary metrics in Tables 3-4 are both defined relative to Observer 1's planes, so the reported error numbers literally measure how well AdaPR reproduces O1. The inter-observer comparison with O2 provides a reference for human variability, but both observers used the same PC-MRA-based workflow, so a shared systematic error would remain invisible. This concern is load-bearing because it affects both the plane-accuracy headline and the flow-comparability claim: if O1's planes are biased, AdaPR inherits that bias and the 'comparable to experts' conclusion becomes circular. The paper does not validate the manual labels against an independent consensus or a phantom with known ground truth. Other issues, such as the Table 5 inconsistency and the cross-study non-DRL comparison, are real but secondary and easily fixed by corrections or softened wording. The paper has independent support that should be credited: the code is public, the cross-validation protocol is sound, and the A3C-VanillaPR baseline controls for the algorithm change, isolating the benefit of the local coordinate system. The perturbation experiments confirm orientation invariance by construction, since sampling the state in local coordinates makes the policy equivariant to rigid transforms. Given that the O1-dependency is a limitation the authors already partially acknowledge (Section 4.3) and that it can be addressed by a multi-observer consensus study, the appropriate outcome is to maintain the reader's CONDITIONAL verdict rather than escalate to rejection. An independent validation of the manual standard would settle the concern and, if it passes, would considerably strengthen the paper.","tokens_in":19353,"tokens_out":20497,"duration_ms":197688,"concrete_test":"Re-annotate a random subset of 20 of the 88 scans with 3-5 additional experienced observers (blinded to O1/O2) on the same four vessels, and form a consensus target (e.g., geometric median of unit normals and robust mean of centers). Compute AdaPR-to-consensus errors and each observer-to-consensus errors for angular and distance metrics. If AdaPR-to-consensus errors are not significantly larger than the observer-to-consensus spread (or inter-observer consensus error), the single-observer evaluation is adequate and the concern is resolved. If AdaPR is significantly worse against the consensus than individual observers are, then the reported error metrics understate the true plane error and the 'comparable to expert observers' claim would need to be weakened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Equations (7)-(8) define the training reward from the angular and distance error to the target plane (n_T, P_T) placed by Observer 1, and Tables 3-4 evaluate the same errors relative to Observer 1. Thus the headline numbers (6.32° ± 4.15°, 3.40 ± 2.75 mm) are agreement with a single observer's annotation style, not accuracy against an independent ground truth. The inter-observer comparison (O1 vs O2, 4.42° ± 4.48°, R²=0.969) mitigates random label noise but cannot detect a systematic bias shared by both observers, since both used the same PC-MRA-based ParaView protocol. If O1's planes are systematically off (e.g., a consistent tilt or center offset), AdaPR will learn and reproduce that bias, and both the geometry and flow claims would overstate how well the planes match true anatomy. The flow result (no significant difference vs O1/O2, R²=0.972/0.968) is also agreement with the same potentially biased reference; it does not establish that the planes are anatomically correct. The paper acknowledges reliance on manual labels in Section 4.3 but provides no independent validation, so the absolute accuracy component of the central claim is underdetermined.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents AdaPR, a deep reinforcement learning framework for automated plane reformatting in 4D flow MRI. The key idea is to replace the global coordinate system used by prior DRL view-planning methods with a local, plane-centered coordinate system, so that the agent's rotations and translations are defined relative to the current plane rather than the scanner axes. The method is trained with the A3C algorithm using a reward that penalizes angular and distance errors to a target plane placed by an expert observer. Experiments on 88 multi-vendor, multi-institution datasets with four-fold cross-validation report mean angular errors of 6.32° +/- 4.15° and distance errors of 3.40 +/- 2.75 mm relative to Observer 1, with better performance than two VanillaPR (global-coordinate) DRL baselines and maintained accuracy under simulated rigid transformations. Flow measurements from AdaPR planes correlate strongly with both manual observers (R^2 = 0.972 and 0.968), close to inter-observer agreement (R^2 = 0.969).","tokens_in":19690,"tokens_out":9642,"duration_ms":85246,"significance":"If the reported results hold, AdaPR addresses a genuine practical limitation of current DRL-based plane reformatting: the requirement that test volumes share the spatial alignment of the training data. The local-coordinate formulation is simple, conceptually clear, and directly motivated by the multi-scanner, multi-orientation setting of 4D flow MRI. The empirical study is substantial for the field, including patients with congenital heart disease, multiple scanner vendors, sensitivity analyses, and flow-based validation against two observers. The open-source implementation is a further strength. The significance is tempered, however, by the fact that the evaluation standard is a single observer's manual planes used both for training and for the headline accuracy numbers; the paper also contains a clear tabular inconsistency and uses a statistically questionable cross-study comparison. These issues require revision before the claims can be fully accepted.","major_comments":[{"comment":"The target planes placed by Observer 1 serve both as the training reward in Eq. (7) and as the reference for the reported errors (6.32°±4.15°, 3.40±2.75 mm) and the flow correlation analysis. The reported 'accuracy' is therefore a measure of agreement with a single, unvalidated annotation standard. The inter-observer comparison (O1 vs O2) mitigates random label noise, but both observers used the same PC-MRA/ParaView protocol, so a systematic bias common to both would be inherited by AdaPR and would inflate both the geometric and flow agreement claims. Section 4.3 acknowledges the reliance on manual labels, but no independent validation (phantom with known plane geometry, multi-observer consensus, or synthetic planes) is provided. This limitation directly affects the absolute-accuracy component of the central claim; please add such a comparison or explicitly reframe the results as agreement with Observer 1's protocol.","section":"§2.3.2, Eq. (7), §3.1.4, §4.3"},{"comment":"Table 5 has a serious internal inconsistency: for every row, the 'All' column exactly repeats the LPA column (e.g., volunteers angular 5.96±4.16, patients 7.47±3.95, inter-observer volunteers 4.49±4.25). This is not a valid aggregate of the AAo, PA, RPA, and LPA columns. In addition, the text in §3.1.4 reports patient errors of 7.31°±3.98° and 4.36±3.58 mm, which do not match the patient 'All' entries 7.47°±3.95° and 4.41±3.59 mm in Table 5. These discrepancies need to be corrected and the subgroup analysis rechecked, because they directly affect the reported patient-versus-volunteer comparison.","section":"Table 5 and §3.1.4"},{"comment":"The VanillaPR baselines are described as 'trained on data that was pre-aligned to the same orientation and position via rigid registration', but the paper does not state whether the same rigid registration is applied to the test volumes before evaluation, and if so, how it interacts with the sensitivity analysis in §3.1.3, which applies rotations and translations to the input volumes. If the VanillaPR test pipeline omits the pre-registration step, the degradation shown in Figure 2 may reflect a mismatch between training and test conditions rather than an intrinsic limitation of global coordinates; if it includes pre-registration, the perturbation experiment is not a fair test of the full pipeline. Please specify the exact inference-time preprocessing for each algorithm and, ideally, evaluate VanillaPR with its full intended pipeline (including pre-registration) under the same transformations.","section":"§2.3.3 and §3.1.3"},{"comment":"The comparison with non-DRL methods in Table 7 is based on two-sample t-tests applied to published means, standard deviations, and sample sizes from different studies. This is statistically inappropriate because the data come from different populations, acquisition protocols, vessel subsets, and evaluation pipelines, and the test ignores within-study correlation and confounding. The resulting statement that AdaPR significantly 'outperforms' the atlas-based and 3D-CNN methods is therefore not supported by the presented analysis. Please either provide a rigorous matched comparison on the same data or rephrase these results as descriptive benchmark comparisons.","section":"§4.1 and Table 7"}],"minor_comments":[{"comment":"Please clarify how the auxiliary lines are encoded into the state (e.g., as additional input channels or as intensity overlays) and how the network uses them; the current description leaves the implementation ambiguous.","section":"§2.1.2"},{"comment":"The ART ANOVA is applied to repeated measurements from the same subjects across vessels and comparison types; the paper should state whether the model accounts for within-subject correlation or justify why the independence assumption is acceptable.","section":"§2.4.1"},{"comment":"The phrase '(Table 5 Appendix)' appears to be a typo; it should refer to Table 5 or a proper appendix table. Also, the text in §4.3 quotes patient errors as 7.47° and 4.41 mm, which conflicts with the 7.31° and 4.36 mm values in §3.1.4; please reconcile.","section":"§4.3"},{"comment":"Equation (8) defines the reward as a finite difference of the cost, and a terminal reward of 3 is added separately; it would be helpful to state explicitly that the total reward is r_t + 3 at terminal states and how this is incorporated in the advantage calculation, Eq. (11).","section":"§2.1.4"},{"comment":"The action set in Eq. (4) includes rotations about w1 and w2 but no action that rotates the plane about its normal n. If in-plane orientation is irrelevant to the flow measurements, please state that explicitly; otherwise clarify how the auxiliary lines in §2.1.2 are used in the reward or network input.","section":"§2.1.3"}],"recommendation":"major_revision","confidential_remarks":"The paper presents a useful incremental contribution with a clear technical idea, and the multi-center, multi-vendor evaluation is a strength. My main reservation is that the evaluation is anchored to a single observer's manual annotations used both for training and for the headline accuracy numbers; this is not fatal but needs to be addressed, either with an independent validation or with explicitly softened claims. The Table 5 inconsistency and the text mismatch in patient error values suggest the manuscript needs a careful data-check pass. The cross-study statistical comparison in Table 7 is not defensible as a significance test. If these points are resolved, I would support publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Good paper, actually. AdaPR is a real, incremental advance in DRL-based plane reformatting. The local-coordinate action space directly addresses the pre-registration requirement of prior work, and the experiments show it holds up under rigid perturbations. The stress-test worry about Observer 1 being both the reward target and the evaluation reference is real but not fatal: the claim is 'comparable to manual observers', and the inter-observer numbers provide the yardstick. The paper is honest that angular errors are larger than inter-observer for most vessels.\n\nWhat's new: the local coordinate system for plane transitions, plus the A3C twist. Validating on 88 multi-vendor scans, including CHD patients, is a solid step. The robustness plots are convincing: AdaPR stays flat while VanillaPR degrades sharply. Flow agreement (R²=0.972/0.968) is essentially at inter-observer levels (0.969), which is the right benchmark for a tool meant to replace manual work.\n\nThe soft spots: Table 5's 'All' column is a copy-paste of LPA, and the text cites numbers (7.31°, 4.36 mm) that don't match the table (7.47°, 4.41 mm). That's a genuine error. Also, the abstract says 'arbitrary positions and orientations,' but the experiments only test up to 25° and 25 mm—that's a range, not arbitrary. And the comparison with non-DRL methods uses literature means, which is weaker than a direct comparison on the same data.\n\nOn the stress-test note: the lack of phantom validation means we can't exclude shared observer bias, but the clinical gold standard is manual placement, so this is a limitation, not a disqualifier. I think the reader's conditional verdict is right, and the table issue is the main blocker.\n\nTarget audience: medical imaging researchers working on view planning or plane reformatting. It deserves peer review—conditional accept after the corrections.","headline":"AdaPR is a solid incremental advance in DRL plane reformatting; fix the Table 5 inconsistency and tone down the 'arbitrary orientations' claim.","tokens_in":20249,"tokens_out":3782,"would_cite":false,"duration_ms":36019,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that a reinforcement-learning agent can reformat 4D flow MRI planes independently of scan orientation by acting in a plane-local coordinate system, and that the resulting flow measurements match expert observers as…","keywords":["deep reinforcement learning","plane reformatting","4D flow MRI","local coordinate system","A3C","flow quantification","cardiovascular MRI","orientation invariance"],"falsifier":"Run AdaPR on volumes with known ground-truth planes, obtained from a flow phantom with a calibrated double-oblique cut or from synthetic 4D flow data, and compare the recovered planes with the known targets over many orientations. If the angular error systematically exceeds the reported $6.32^\\circ$ or depends on the volume's orientation, the orientation-independence claim would be contradicted; the same test with a consensus of several experts as reference would check the manual ground truth assumption.","tokens_in":19192,"feed_emoji":"🫀","tokens_out":7271,"duration_ms":63973,"temperature":0.7,"pith_summary":"AdaPR tries to make automated plane reformatting for 4D flow MRI independent of how a scan was acquired: instead of moving the plane in the scanner's fixed coordinate axes, a deep reinforcement-learning agent rotates and translates the plane along the plane's own local axes ($\\vec n_t$, $\\vec w^1_t$, $\\vec w^2_t$). On 88 multi-vendor scans, the paper reports a mean angular error of $6.32^\\circ \\pm 4.15^\\circ$ and a distance error of $3.40 \\pm 2.75$ mm relative to the manual planes of one expert, outperforming a global-coordinate DQN baseline and remaining near-constant when volumes are rotated and translated by 5–25 degrees/millimetres. Flow measured on AdaPR planes showed no significant difference from two manual observers, with $R^2 = 0.972$ and $0.968$, comparable to inter-observer agreement ($R^2 = 0.969$). If correct, this makes automated plane placement usable across scanners and institutions without volume pre-registration, by removing the orientation assumption built into earlier DRL view-planning agents.","feed_headline":"Adaptive local coordinates make MRI plane placement match experts","feed_subtitle":"Its measured flows match expert values as well as two experts match each other on 88 scans.","key_machinery":"The load-bearing mechanism is the local orthonormal basis $\\{\\vec n_t, \\vec w^1_t, \\vec w^2_t\\}$: the first axis is the current plane normal and the other two span the plane, with a state sampled as a 3D sub-volume along these axes plus auxiliary lines along $\\vec w^1_t$ and $\\vec w^2_t$ that preserve orientation information. Actions in Eq. (4) are rotations of the normal around the two in-plane axes and translations of the center along the three local axes, which makes the policy equivariant to rigid transformations of the whole volume. The agent's reward is the temporal decrease of the cost $C(t)$ in Eq. (7), with a terminal bonus of 3 when the angle falls below $3^\\circ$ and the distance below 2 mm, and training uses the on-policy A3C algorithm with an LSTM between consecutive states. This combination is what lets the plane be steered in arbitrary scanner coordinates without pre-registration.","core_discovery":"The paper's central claim is that the spatial alignment assumption in prior DRL plane-reformatting methods can be removed by expressing both the state and the actions in a local coordinate system attached to the current plane. At each step, the agent's state is a sub-volume sampled along the local basis $\\{\\vec n_t, \\vec w^1_t, \\vec w^2_t\\}$ centered at $\\vec P_t$, and its actions are a rotation of the normal around $\\vec w^1_t$, a rotation around $\\vec w^2_t$, and translations along $\\vec w^1_t$, $\\vec w^2_t$, and $\\vec n_t$. Because the same action produces the same relative change regardless of the volume's global orientation, the learned policy transfers to arbitrarily posed volumes. The paper demonstrates this with an A3C implementation (AdaPR) trained against a cost $C(t) = (1 - \\vec n_t \\cdot \\vec n_T / (\\|\\vec n_t\\| \\|\\vec n_T\\|)) + \\lambda d(P_t,P_T)$ that balances angular and distance errors, and reports that AdaPR reaches $6.32^\\circ \\pm 4.15^\\circ$ and $3.40 \\pm 2.75$ mm against expert planes, that its error stays within $1.5^\\circ$ and 1.1 mm of baseline under rigid perturbations, and that flow quantification matches inter-observer performance.","pith_inferences":["A stricter falsification of the invariance claim would be a synthetic ground-truth test, where the target plane is generated by a known double-oblique transform of a reference plane; the paper's evaluation only compares against manual planes.","The local-coordinate policy is a general navigation strategy that should transfer to other imaging modalities and other volumetric targeting tasks, such as valve tracking or biopsy planning, but this is an extrapolation beyond the paper's evidence.","The reported flow agreement may be forgiving of angular errors because flow is insensitive to orientation below roughly $25^\\circ$, so the more meaningful clinical test is whether AdaPR's angular performance holds on a larger patient cohort with complex anatomy.","An explicit 'orientation stress test' reporting per-voxel anatomy alignment after rotating the same volume by arbitrary angles would quantify what the paper's axis-aligned perturbations only approximate."],"forward_implications":["AdaPR removes the need to pre-register test volumes to a common spatial framework, so a single trained model can reformat planes in scans acquired in arbitrary orientations and positions from different vendors.","Flow computed from AdaPR planes is statistically indistinguishable from expert flow measurements, with $R^2 = 0.972$ against observer 1 and $R^2 = 0.968$ against observer 2 versus inter-observer $R^2 = 0.969$.","Under rigid rotations and translations of 5–25 degrees and millimetres, AdaPR's average error stays within $1.5^\\circ$ and 1.1 mm of its unperturbed performance, whereas the global-coordinate DQN baseline degrades sharply.","The on-policy A3C formulation converges where the off-policy DQN did not in the larger local-coordinate state space, suggesting that on-policy methods are the practical choice for this task.","Because accuracy held across healthy volunteers and congenital heart disease patients from multiple scanners, the method is expected to generalize to unseen acquisitions of similar anatomy."],"supporting_citations":[{"why":"Supplies the global-coordinate DQN framework and discrete reward that AdaPR compares against and improves.","marker":"[14]"},{"why":"Provides the A3C on-policy algorithm that AdaPR uses for stable training.","marker":"[20]"},{"why":"The direct 3D-CNN plane-placement baseline for 4D flow MRI; its reported observer variability anchors the comparison.","marker":"[13]"},{"why":"Represents the atlas-registration approach and its flow $R^2$ used in the non-DRL comparison.","marker":"[11]"},{"why":"Defines the 3D PC-MRA volume used as the agent's environment.","marker":"[17]"},{"why":"Supplies the CLAHE preprocessing that homogenises scans across vendors.","marker":"[18]"},{"why":"Defines the manual double-oblique plane placement protocol used by the observers.","marker":"[10]"},{"why":"Supports the claim that plane angulation changes over $25^\\circ$ are needed to substantially alter aortic flow, motivating why large angular error can coexist with good flow agreement.","marker":"[32]"}],"fun_headline_variants":["Local coordinates let MRI plane AI ignore image orientation","AI adapts MRI planes to any position, matching expert accuracy","Orientation-independent MRI plane AI matches expert flow values","AdaPR: DRL plane reformatting that generalizes across scanners","MRI flow AI uses local frame to work on arbitrary geometry"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Observer 1's manually placed planes are used as the ground truth in the reward function and as the evaluation reference, so the reported plane accuracy and the 'comparable to observers' flow claim assume that those manual planes are correct and representative of clinical needs.","fun_headline_variants_meta":{"raw":{"variants":["Local coordinates let MRI plane AI ignore image orientation","AI adapts MRI planes to any position, matching expert accuracy","Orientation-independent MRI plane AI matches expert flow values","AdaPR: DRL plane reformatting that generalizes across scanners","MRI flow AI uses local frame to work on arbitrary geometry"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000267,"raw_usage":{"total_tokens":1730,"prompt_tokens":1174,"completion_tokens":556,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":790,"completion_tokens_details":{"reasoning_tokens":474}},"tokens_in":790,"tokens_out":556,"duration_ms":6342,"temperature":1.0,"reasoning_tokens":474,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:59:05.424972+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run AdaPR on volumes with known ground-truth planes, obtained from a flow phantom with a calibrated double-oblique cut or from synthetic 4D flow data, and compare the recovered planes with the known targets over many orientations. If the angular error systematically exceeds the reported $6.32^\\circ$ or depends on the volume's orientation, the orientation-independence claim would be contradicted; the same test with a consensus of several experts as reference would check the manual ground truth assumption.","supporting_citations":[{"cited_title":"Alansary, L","cited_arxiv_id":null,"evidence_quote":"Supplies the global-coordinate DQN framework and discrete reward that AdaPR compares against and improves."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The direct 3D-CNN plane-placement baseline for 4D flow MRI; its reported observer variability anchors the comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Represents the atlas-registration approach and its flow $R^2$ used in the non-DRL comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the 3D PC-MRA volume used as the agent's environment."},{"cited_title":"Publisher: Elsevier","cited_arxiv_id":null,"evidence_quote":"Supplies the CLAHE preprocessing that homogenises scans across vendors."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the manual double-oblique plane placement protocol used by the observers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supports the claim that plane angulation changes over $25^\\circ$ are needed to substantially alter aortic flow, motivating why large angular error can coexist with good flow agreement."}],"review_version":1}