{"id":"2448956b-d7eb-4cd9-ba6f-b6e3d003958e","arxiv_id":"2506.08048","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A physics-regularized per-instance neural network plus finite-element refinement, with optional surgeon line prompts, reduces liver deformation registration error to 2.78 mm on a public phantom benchmark.","lead":"This paper introduces a hybrid neural-biomechanical deformation pipeline and an interactive AR correction mode for aligning preoperative liver models to intraoperative surfaces. On the public Image-to-Physical Liver Registration Sparse Data Challenge it reports 3.42 mm mean target registration error automatically, improving to 2.78 mm with user-drawn prompts.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The phantom 'surpassing SOTA' claim rests on a non-comparable interactive condition and an untested 0.15 mm difference; no interactive baseline or significance test is provided.","rationale":"The reader's verdict is CONDITIONAL and their weakest_assumption concerns point-cloud contamination. I agree the paper is conditionally acceptable, but I find the more load-bearing weakness in the phantom comparison: the superiority claim is based on a human-in-the-loop condition that prior methods were not allowed to use, and the margin over the best prior method is small and untested for significance. The point-cloud contamination concern is real but applies equally to all surface-driven methods and does not undermine the comparative results on the provided phantom data; the interactive baseline and statistical testing directly gate the headline claim. A controlled test that adds the same prompt mechanism to a strong prior baseline, plus a per-case significance test, would settle whether the claimed SOTA advantage is real. Since this is a major revision condition rather than a rejection, I keep the verdict unchanged from the reader's CONDITIONAL.","tokens_in":19359,"tokens_out":8181,"duration_ms":86783,"concrete_test":"Using the public Image-to-Physical Liver Registration Sparse Data Challenge analysis tools, obtain per-case TREs for BiomPINN-PBMs (w/ prompt) and for Yang et al. (2024), and run a paired Wilcoxon signed-rank test on the per-case differences. Then implement the prompt-correction stage of Algorithm 2 on top of Yang et al.'s surface displacement estimates, using the same user annotations and ICP/mutual-nearest-neighbor updates, and compare the resulting TRE. If p ≥ 0.05 or the prompted Yang baseline reaches ≤2.78 mm, the claim that BiomPINN-PBMs surpasses state-of-the-art is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table V is the load-bearing evidence for the headline claim that prompts yield 2.78 mm and 'surpass state-of-the-art.' But the fully automatic BiomPINN-PBMs (w/o prompt) achieves 3.42 ± 0.72 mm, which is worse than Heiselman et al. (3.08 ± 0.85), Mestdagh et al. (3.31 ± 1.86), and Yang et al. (2.93 ± 0.68). The only condition in which the method beats prior work is 'w/ prompt,' where a human draws corrective line annotations on the deformed model and point cloud and can iterate until the surface looks aligned. No prior method is given the same interactive correction mechanism, so the comparison is not apples-to-apples. Moreover, the improvement over the best prior method is only 0.15 mm (2.78 vs. 2.93), well within the reported standard deviations (~0.68 mm), and no paired significance test is reported for the phantom data. If this difference is not statistically significant—or if applying the same prompt-correction procedure to Yang et al. yields a comparable or better TRE—then the paper's central superiority claim is unsupported. This is distinct from the acknowledged point-cloud contamination assumption (Sec. VI-A), which would affect all compared methods equally; the interactive baseline and statistical testing are the load-bearing issues for the claimed advantage.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes a hierarchical per-instance MLP deformation network (BiomPINN) regularized by a finite-element strain-energy term, integrates it with a patient-specific biomechanical model (BiomPINN-PBMs) through the linear system in Eqs. (21)-(24), and adds an interactive AR framework in which surgeons draw line prompts to correct correspondences and re-optimize the deformation. Experiments on a synthetic liver/kidney/prostate dataset show that BiomPINN-PBMs achieves mean TRE statistically equivalent to GMM-FEM while reducing runtime, and experiments on the public Image-to-Physical Liver Registration Sparse Data Challenge phantom dataset report a mean TRE of 3.42 mm without prompts and 2.78 mm with prompts, which the authors claim surpasses all listed prior methods. The paper also reports qualitative in-vivo laparoscopic cases and an end-to-end timing analysis of the interactive loop.","tokens_in":19689,"tokens_out":5248,"duration_ms":58358,"significance":"The synthetic validation is a genuine strength: the paired statistical comparison with GMM-FEM, the runtime measurements, and the Jacobian-determinant field-consistency analysis are concrete and reproducible, and the use of a public challenge dataset with blinded targets is appropriate for benchmarking. The interactive prompt mechanism is novel and clinically motivated, and the timing breakdown of the annotation loop is useful for assessing intraoperative feasibility. However, the headline claim of surpassing state-of-the-art volumetric accuracy currently rests on a non-comparable interactive condition, and the fully automatic variant is not competitive with several prior methods. If the authors add a proper interactive baseline or reframe the claim, the paper would be a solid contribution to non-rigid registration for surgical navigation; as written, the significance is contingent on that additional evidence.","major_comments":[{"comment":"The claim of surpassing state-of-the-art is not supported by the reported comparison. In the fully automatic setting, BiomPINN-PBMs (3.42 ± 0.72 mm) is worse than Heiselman et al. (3.08 ± 0.85 mm), Mestdagh et al. (3.31 ± 1.86 mm), and Yang et al. (2.93 ± 0.68 mm). The only condition in which the method beats prior work is the 'w/ prompt' condition, where no prior method is given the same interactive correction mechanism. Moreover, the improvement over the best prior method is 0.15 mm (2.78 vs. 2.93 mm), which is well within the reported standard deviations, and no paired significance test or per-case analysis is provided for the phantom data. The abstract's statement that the method surpasses state-of-the-art in volumetric accuracy is therefore not an apples-to-apples comparison.","section":"Section V-B, Table V"},{"comment":"The interactive prompt protocol is underspecified, making the 'w/ prompt' result difficult to interpret. The paper does not report the number of prompts per case, the number of interaction iterations, the time spent per case, or the inter-user variability, and the annotators were computer science and biomedical engineering students rather than surgeons. Without a controlled interactive baseline—for example, applying the same line-prompt correction procedure to Yang et al. or to a simple ICP-based correspondence update—the measured improvement cannot be attributed specifically to the proposed framework rather than to the general benefit of additional user-provided correspondences.","section":"Section V-B and Section IV-C"},{"comment":"The residual-based estimate of σ² in Eq. (22) uses the same surface target that appears in the data term of Eq. (21) and is computed from the BiomPINN prediction, but the paper provides no sensitivity analysis for σ² or for the Tikhonov weight β, and β is never reported anywhere in the experiments. Since the linear system in Eq. (24) depends on the product βσ², the numerical results are not fully reproducible, and the claim that the approach avoids iterative regularization tuning should be accompanied by evidence that the results are stable over reasonable ranges of these parameters.","section":"Section III-C, Eq. (22)"}],"minor_comments":[{"comment":"There is a typo in the sentence following Eq. (22): 'Taking the derivative of the derivative of E(u_Ω)' should read 'Taking the derivative of E(u_Ω)'. Additionally, in the definition of C̃, the phrase 'is is the Kronecker product' contains a duplicated 'is'.","section":"Section III-C"},{"comment":"The runtime comparison is asymmetric: GMM-FEM runtime is reported as only the linear-system solve time, whereas BiomPINN-PBMs runtime includes the full per-instance optimization and biomechanical solve. The text claims the setup makes runtimes directly comparable, but the reporting should clarify whether both methods exclude or include correspondence estimation and iterative regularization tuning, since this affects the interpretation of the speedup.","section":"Section V-A, Table III"},{"comment":"The interactive experiments were performed by two students rather than clinicians. The paper should either use the term 'users' instead of 'surgeons' in the experimental protocol or justify why the student annotations are representative of surgical expertise.","section":"Section IV-C"},{"comment":"The human in-vivo evaluation relies solely on a one-sided Chamfer distance to the reconstructed surface and on qualitative inspection; no comparison with existing non-rigid registration methods is provided. The authors should state more explicitly that these results are feasibility demonstrations rather than quantitative performance claims.","section":"Section V-C"}],"recommendation":"major_revision","confidential_remarks":"The synthetic validation and the interactive system are promising, but the paper's central 'surpassing state-of-the-art' claim in the abstract and conclusion is currently supported only by a non-comparable interactive comparison without statistical testing. I would accept a revised version that either adds a proper interactive baseline and significance tests or carefully reframes the claims to state that the interactive framework improves over the automatic variant and that this improvement reaches the range of the best prior methods."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Bottom line: this is a competent hybrid method with a genuinely useful one-pass trick, but the headline phantom result underpromises on statistics and overpromises on comparison fairness. The paper deserves a serious referee and major revision.\n\nWhat is new: per-instance optimization of a neural deformation pyramid with an FE strain-energy regularizer, feeding a biomechanical model whose regularization weight is estimated directly from the residual of the neural prediction (Eq. 22). That bypasses the iterative sigma^2 loop, and on synthetic liver/kidney data it matches GMM-FEM's accuracy at about 40% of the runtime; on prostate it is significantly better. The line-prompt interactive correction is also new, and the incremental formulation in Algorithm 2 is clean. As far as I can tell, the linear system in Eqs. (21)-(24) is correct.\n\nThe paper is honest about several limitations: it states the point-cloud contamination assumption in Sec. VI-A, admits the annotation accuracy is ~6 mm in a small user study, and reports that HCI experiments were done by students. That is more transparent than many papers.\n\nSoft spots, in order of size. First, the phantom 'surpassing state-of-the-art' claim rests on a non-comparable condition. Automatic BiomPINN-PBMs (3.42 mm) is worse than Yang et al. (2.93 mm) and Heiselman (3.08 mm). Only w/ prompt does it hit 2.78 mm, and no prior method is given the same interactive correction. The 0.15 mm gap over Yang et al. is within one standard deviation, and no paired significance test is reported. That's a load-bearing weakness.\n\nSecond, reproducibility: lambda_1 and lambda_2 are tuned per organ, beta is never reported, and the correspondence pre-processing differs across datasets (UTOPIC/GraphSCNet vs. mutual nearest neighbors). Fine for a methods paper, but not exact enough for clinical systems.\n\nThird, the interactive experiments use students rather than surgeons, and the annotation error analysis is a pilot. That is a minor issue given the method is still at the proof-of-concept stage.\n\nWho should read it: people working on deformable registration with sparse intraoperative data, and anyone thinking about human-in-the-loop for surgical AR. It deserves peer review; the right outcome would be major revision, not rejection. The synthetic validation and the one-pass idea are worth keeping; the phantom claim needs a fair interactive baseline and proper statistics.","headline":"Solid hybrid method with a clean one-pass regularization trick, but the phantom SOTA claim needs a fair interactive baseline and a significance test.","tokens_in":20177,"tokens_out":2090,"would_cite":true,"duration_ms":21840,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A hybrid neural–biomechanical solver matches finite-element deformation accuracy while running about 2.5 times faster, and surgeon-drawn prompts lower registration error further.","keywords":["augmented reality","surgical navigation","deformation modeling","non-rigid registration","biomechanical model","physics-informed neural network","human-in-the-loop","target registration error"],"falsifier":"Measure target registration error on the phantom dataset while artificially adding a retractor-shaped patch of points to the intraoperative cloud; if the embedded-bead TRE degrades by more than the roughly 0.6 mm gain that prompts provide, then the clean-surface assumption is the binding limit on real-surgery transfer.","tokens_in":19178,"feed_emoji":"🩺","tokens_out":8872,"duration_ms":88252,"temperature":0.7,"pith_summary":"AR-guided surgery needs a preoperative organ model deformed to match the live anatomy. This paper claims that a hybrid pipeline—a per-instance physics-integrated neural network followed by a one-pass biomechanical solve—can match finite-element deformation accuracy while running about 2.5 times faster, and that letting a surgeon draw corrective prompts on the misaligned surface improves accuracy further. On a public liver-phantom benchmark the automated pipeline reaches 3.42 mm mean target registration error, and prompted interactions reduce this to 2.78 mm, below all prior methods listed in the comparison. If correct, this makes physically plausible deformation modeling fast enough for intraoperative use and gives surgeons a principled way to fix correspondence errors that automated methods cannot resolve.","feed_headline":"Surgeon prompts cut AR liver registration error to 2.78 mm","feed_subtitle":"Hybrid neural-FE solver keeps FEM-level accuracy at a fraction of runtime, and human correction beats prior methods.","key_machinery":"The load-bearing object is the finite-element stiffness matrix $K$ assembled on a tetrahedral mesh of the organ, used twice in different roles. Inside BiomPINN, $K$ appears as a quadratic strain-energy regularizer (via the interpolation matrix $\\Phi$ that maps surface displacements to volumetric ones), penalizing locally inconsistent boundary motion; in the PBM stage, the same $K$ is the elastic term of a Tikhonov-regularized linear system, $(\\Phi^T P \\Phi + \\beta \\sigma^2 K) u_\\Omega = \\Phi^T P b$, whose data weight $\\sigma^2$ is computed directly from BiomPINN's residual rather than by iterative expectation-maximization. This single-pass formulation is what replaces iterative regularization tuning, and the correspondence-refinement loop—prompt lines expanded to local patches, aligned by iterative closest point, and re-matched by mutual nearest neighbors—is what turns surgeon feedback into a biomechanically propagated correction.","core_discovery":"On its own terms, the paper's central claim is that the expensive iterative tuning of a patient-specific biomechanical model can be replaced by a direct prediction from a physics-regularized neural network, with no loss of accuracy. BiomPINN is a four-level MLP deformation pyramid optimized per instance; its displacement field is projected onto a tetrahedral finite-element mesh, and the FE stiffness matrix supplies a strain-energy regularizer at each level. The resulting surface displacement is then fed into the PBM as a residual-based estimate of the data weight, so the volumetric deformation is obtained by solving a single sparse linear system instead of repeatedly re-solving it. On synthetic liver, kidney, and prostate data, the method matches the GMM-FEM baseline's mean target registration error (2.47, 3.40, and 0.95 mm) at runtimes of 1.65, 1.25, and 1.50 seconds versus 4.28, 3.18, and 3.38 seconds. The interactive layer treats surgeon line prompts as corrections to the correspondence matrix, re-optimizes the deformation incrementally, and on the phantom challenge improves TRE from 3.42 mm to 2.78 mm, surpassing all compared prior methods.","pith_inferences":["The annotation phase, not the solver, is the end-to-end bottleneck of a correction cycle, so replacing line-drawing with touchscreen or semi-automated prompts could make the interactive loop viable for continuous, not just on-demand, navigation.","Because the prompt mechanism only edits correspondences, it is a general repair operator: the same interface could be bolted onto any correspondence-driven non-rigid registration algorithm, not just the proposed hybrid.","If the residual-derived data weight estimate generalizes, any finite-element registration that currently re-estimates its regularization weight by outer-loop iteration could adopt a one-pass solve.","The reported mid-air annotation error suggests prompt accuracy, not model capacity, may become the next limiting factor; explicitly visualizing inferred correspondences during annotation could test that directly."],"forward_implications":["Any deformation case processed by BiomPINN-PBMs avoids offline training, so accuracy does not depend on how well synthetic training data match the operating room.","The same pipeline runs in under two seconds per case for liver, kidney, and prostate test cases, making biomechanical correction compatible with intraoperative decision points.","Prompt-based corrections are local updates to the correspondence matrix, so repeated rounds of surgeon guidance accumulate as a chain of incremental deformation fields rather than requiring a full restart.","Retaining the mesh topology after deformation lets preoperative surgical plans be carried onto the intraoperative anatomy in the same coordinate frame."],"supporting_citations":[{"why":"Supplies the GMM-FEM baseline and node-ordering strategy; the central runtime and accuracy equivalence is measured against it.","marker":"[9]"},{"why":"Provides the synthetic data generation protocol and a deep-learning baseline compared on the phantom challenge.","marker":"[5]"},{"why":"Defines the public liver-phantom benchmark with embedded validation beads and evaluation tools.","marker":"[25]"},{"why":"Reports prior method results on the common phantom dataset that the paper's Table V extends.","marker":"[18]"},{"why":"Supplies the coarse-to-fine hierarchical MLP pyramid that BiomPINN adapts for per-instance displacement prediction.","marker":"[23]"},{"why":"Provides the initial overlap-aware correspondence prediction used for the synthetic liver dataset.","marker":"[21]"},{"why":"Refines the liver correspondences and filters outliers before deformation optimization.","marker":"[22]"},{"why":"Is the strongest prior phantom result that the prompted pipeline surpasses.","marker":"[31]"}],"fun_headline_variants":["Surgeon prompts slash AR liver registration error to 2.78 mm","Human-in-the-loop AI trims AR surgical error to under 3 mm","Data-driven biomechanics + surgeon hints hit 2.78 mm overlay accuracy","Interactive AI: surgeon corrections improve AR navigation to 2.78 mm","Efficient AI matches FEM accuracy, surgeon prompts cut error to 2.78 mm"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pipeline assumes the intraoperative point cloud is a clean, accurate capture of the organ surface, without significant contamination from surgical instruments or surrounding tissue; if that fails, the correspondences driving every correction are computed against the wrong geometry.","fun_headline_variants_meta":{"raw":{"variants":["Surgeon prompts slash AR liver registration error to 2.78 mm","Human-in-the-loop AI trims AR surgical error to under 3 mm","Data-driven biomechanics + surgeon hints hit 2.78 mm overlay accuracy","Interactive AI: surgeon corrections improve AR navigation to 2.78 mm","Efficient AI matches FEM accuracy, surgeon prompts cut error to 2.78 mm"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000538,"raw_usage":{"total_tokens":2637,"prompt_tokens":1053,"completion_tokens":1584,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":669,"completion_tokens_details":{"reasoning_tokens":1483}},"tokens_in":669,"tokens_out":1584,"duration_ms":11356,"temperature":1.0,"reasoning_tokens":1483,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:40:54.686631+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure target registration error on the phantom dataset while artificially adding a retractor-shaped patch of points to the intraoperative cloud; if the embedded-bead TRE degrades by more than the roughly 0.6 mm gain that prompts provide, then the clean-surface assumption is the binding limit on real-surgery transfer.","supporting_citations":[{"cited_title":"Biomechanically constrained surface registration: Application to mr-trus fusion for prostate interventions,","cited_arxiv_id":null,"evidence_quote":"Supplies the GMM-FEM baseline and node-ordering strategy; the central runtime and accuracy equivalence is measured against it."},{"cited_title":"Non-rigid volume to surface registration using a data- driven biomechanical model,","cited_arxiv_id":null,"evidence_quote":"Provides the synthetic data generation protocol and a deep-learning baseline compared on the phantom challenge."},{"cited_title":"The image-to- physical liver registration sparse data challenge,","cited_arxiv_id":null,"evidence_quote":"Defines the public liver-phantom benchmark with embedded validation beads and evaluation tools."},{"cited_title":"The image-to-physical liver registration sparse data challenge: Comparison of state-of-the-art using a common dataset,","cited_arxiv_id":null,"evidence_quote":"Reports prior method results on the common phantom dataset that the paper's Table V extends."},{"cited_title":"Non-rigid point cloud registration with neural deformation pyramid,","cited_arxiv_id":null,"evidence_quote":"Supplies the coarse-to-fine hierarchical MLP pyramid that BiomPINN adapts for per-instance displacement prediction."},{"cited_title":"Utopic: Uncertainty-aware overlap prediction network for partial point cloud registration,","cited_arxiv_id":null,"evidence_quote":"Provides the initial overlap-aware correspondence prediction used for the synthetic liver dataset."},{"cited_title":"Deep graph-based spatial consistency for robust non-rigid point cloud registration,","cited_arxiv_id":null,"evidence_quote":"Refines the liver correspondences and filters outliers before deformation optimization."},{"cited_title":"Boundary constraint- free biomechanical model-based surface matching for intraoperative liver deformation correction,","cited_arxiv_id":null,"evidence_quote":"Is the strongest prior phantom result that the prompted pipeline surpasses."}],"review_version":1}