{"id":"29ef1742-b91b-4387-9ab1-ff0a0261e31b","arxiv_id":"2607.13869","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"On five neuro-oncology patients, percentile-based probabilistic IMPT plans improved tumor coverage (up to 0.93 GyRBE) and reduced organ-at-risk doses (up to 33 GyRBE) compared to robust plans, at the cost of much longer optimization times.","lead":"This paper tests a previously proposed probabilistic method for proton therapy planning on five brain-tumor patients, comparing it with current robust planning. The probabilistic plans reduced dose to healthy structures while maintaining or improving tumor coverage, but the study is small and the numerical differences are close to the method's own uncertainty.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"PCE evaluation error band overlaps reported gains; without a direct dose-engine check the trade-off claim is not quantitatively secure.","rationale":"The reader's verdict of CONDITIONAL is appropriate, and my identified concern is closely related to the reader's weakest assumption. However, I sharpen it: the Appendix A error analysis compares two PCEs (Dij→di versus voxel-dose), not the evaluation PCE against the dose engine. The reported error ranges are best read as an upper bound on the optimizer's internal approximation, not as a direct uncertainty on the final evaluated metrics. Yet because those error ranges overlap the claimed target-coverage gains, the paper needs an external validation step before the quantitative trade-off claim can be accepted. I do not see this as a fatal flaw: the feasibility demonstration is credible, the limitations are unusually candid, and the method is mechanistically plausible. The missing piece is a direct dose-engine check on the final plans, which is a concrete and feasible experiment. I therefore keep the conditional verdict rather than moving to acceptance or rejection.","tokens_in":30274,"tokens_out":5134,"duration_ms":52075,"concrete_test":"For each of the five final robust and probabilistic plans, compute direct dose-engine dose distributions for a validation sample of error scenarios drawn from the same Gaussian uncertainty model (e.g., 300–1000 Monte Carlo scenarios, or the 217-point sparse grid used for PCE construction). From these direct doses, recompute the probabilistic metrics in Tables 2–3: 10th percentile D99.8% for CTV and 95th percentile D0.03cc for OARs. If the robust-vs-probabilistic differences for target coverage and summed OAR metrics are reproduced within ±0.2 GyRBE (or within the PCE error bars), the central claim holds; if the differences shift by more than the reported effect sizes, the trade-off conclusion is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—probabilistic planning improves target coverage and OAR trade-offs relative to robust planning—rests on percentile-based metrics evaluated with a polynomial chaos expansion (Section 2.3). Appendix A validates only the Dij→di PCE used inside the optimizer against the voxel-dose PCE used for evaluation, giving mean percentile differences of -0.21 to +0.18 GyRBE and per-voxel ranges from -0.74 to +1.16 GyRBE (Table 6). The evaluation PCE itself is not checked against the actual dose engine for these patients or for the probabilistic plans specifically. Because probabilistic planning may exploit steeper dose gradients and more nonlinear uncertainty responses, PCE error could differ systematically between the robust and probabilistic plans. The reported target-coverage gains (0.19–0.93 GyRBE in Table 2) fall inside or near this error band, while the summed OAR differences are dominated by a few large DVH reductions rather than being consistently reproduced across metrics. Until the evaluation PCE is validated against direct dose-engine percentiles for both arms of the comparison, the headline 'improved probabilistic trade-off' remains potentially an artifact of the approximation rather than a true dosimetric advantage.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a feasibility study of a percentile-based probabilistic IMPT planning approach applied to five neuro-oncological patients. The method approximates voxel-dose percentiles by E[d] ± δ_i · σ̃_i, where the expected dose and an approximate standard deviation are obtained from a polynomial chaos expansion of the dose-influence matrix, and an outer loop updates δ_i until the percentile estimate converges. The probabilistic plans are compared with automated mini-max robust plans. Results: patient 1 achieved the same 10th percentile of D99.8% while reducing summed OAR-related DVH metrics by 19 GyRBE; patients 2–5 improved the 10th percentile of D99.8% by up to 0.93 GyRBE, with summed OAR DVH changes ranging from +8.1 to −33.0 GyRBE. Total optimization times were 44–226 h, reduced to below 10 h for two patients when the outer-loop warm-starting was omitted. The authors conclude that the probabilistic approach yields improved probabilistic trade-offs between target coverage and OAR sparing.","tokens_in":30503,"tokens_out":8042,"duration_ms":74051,"significance":"If the findings are quantitatively secure, the work is a useful step toward direct probabilistic optimization for IMPT: it demonstrates a scalable implementation (diagonal covariance, PCE-based expected values and gradients), probabilistic hard constraints, and an independent automated robust-plan benchmark. The paper is commendably explicit about its limitations, including the mismatch between optimized voxel percentiles and evaluated pDVH metrics, heuristic acceptance probabilities, and the omission of random errors. The open-source PCE toolbox and the detailed convergence analyses are additional strengths. However, the current evidence supports feasibility rather than a proven dosimetric advantage over robust planning: the central plan-quality comparison is not yet anchored to dose-engine ground truth and rests on a small, heterogeneous patient sample.","major_comments":[{"comment":"The central comparison evaluates both plans with PCE-based percentiles, but the validation in Appendix A is only partial. It compares the Dij→di PCE used in optimization with the di PCE used in evaluation, for patient 4 only; the evaluation PCE itself is not checked against the actual dose engine for either arm. The reported voxel-percentile differences have ranges up to −0.74 to +1.16 GyRBE (Table 6), whereas the claimed 10thD99.8 gains are 0.19–0.93 GyRBE (Table 2). Because the same approximation is applied to both plans, systematic differences could bias the plan comparison, especially if probabilistic plans exploit steeper dose gradients and more nonlinear scenario dependence. Please validate the evaluation PCE on the final plans of both arms using direct scenario sampling from the dose engine (or an independent Monte Carlo engine) and report the resulting percentile differences and","section":"Appendix A, Table 6; §3.1–3.2, Table 2"},{"comment":"The group-level claim rests on five patients with no confidence intervals or hypothesis testing, and the group B pattern is heterogeneous. Patient 4 is excluded from the mean coverage gain because it was scaled to identical target coverage; patient 2’s summed OAR DVH metric increases (+8.1 GyRBE) even though target coverage improves. The statement that the probabilistic approach achieves ‘improved trade-offs’ generalizes beyond what n=4/5 with scaled comparisons can support. Please either frame the results strictly as feasibility observations with per-patient reporting, or add uncertainty quantification (e.g., scenario bootstrap) and temper the generalization in the abstract and conclusions.","section":"§3.2, Tables 2–3; §2.5"},{"comment":"The abstract and conclusions state that the method can ‘precisely optimize’ for clinical goals, but the optimized quantities are voxel-wise percentiles, while the reported evaluation metrics are pDVH quantities (10thD99.8%, 95thD0.03cc). Section 4 explicitly acknowledges this mismatch, the heuristic choice of 90%/95% acceptance probabilities, and patient-specific threshold adjustments. This does not invalidate the method, but it means the pDVH improvements are achieved indirectly. I recommend softening the ‘precisely’ wording and clearly distinguishing optimized voxel-percentile goals from the subsequently evaluated pDVH goals.","section":"§2.4, §2.5, Discussion"}],"minor_comments":[{"comment":"The abstract says optimization times were 44–141 h, but Table 4 reports 226 h for patient 2. Please reconcile.","section":"Abstract and §3.3, Table 4"},{"comment":"The row ‘Total optimization time (original)’ lists 20.3 h for patient 4, inconsistent with Table 4 (79.8 h). Likely a typo; please correct.","section":"Appendix B.3, Table 14"},{"comment":"Target structure labels such as ‘CTV (5940 cc)’ appear to use dose units (cGy, corresponding to 59.4 GyRBE) rather than volume units. Please fix the labels for clarity.","section":"Table 6 and Appendix B.2"},{"comment":"Table 12 uses patient identifiers 10, 12, 32, 35 instead of patients 1–5 used elsewhere. Please unify the notation.","section":"Appendix B.3, Table 12"},{"comment":"The phrase ‘eliminating warm-starting (i.e., the outer loop)’ is imprecise: in the no-outer-loop version, δ-factors are still updated within the inner optimization. Please rephrase to distinguish removal of the outer-loop warm-start from removal of δ-factor updating.","section":"§2.4 and Appendix B.3.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest and methodologically interesting. I see no load-bearing circularity in the δ-factor update; the central issue is verification of the PCE-based evaluation against the actual dose engine. The heavy reliance on antecedent work is appropriate given the incremental nature of the contribution, and the five-patient feasibility framing is acceptable if the conclusions are matched to the evidence. A revision that anchors the plan comparison to dose-engine-sampled percentiles for both arms, even on these five patients, would substantially strengthen the claim and make the paper publishable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nThis is a solid feasibility study rather than a definitive win. The genuinely new pieces are: applying the authors' own PCE-based percentile optimizer to five neuro-oncological patients with probabilistic hard constraints, a diagonal-covariance approximation that makes the problem tractable, and a warm-start ablation showing the outer loop can be dropped with 88–94% time savings. The clinical comparisons against Erasmus-iCycle robust plans are new data.\n\nWhat they do well: the method is described in enough detail to reproduce, the patient-specific wish-lists are included, and the limitations section is unusually candid. They flag the mismatch between optimized voxel percentiles and evaluated pDVH metrics, inactive voxels, neglected random errors, and the fact that the robust benchmark is not the actual clinical plan. The δ-factor update in Eq. 15 makes the percentile estimate self-consistent by construction, so there is no load-bearing circularity.\n\nSoft spots: the central claim—that probabilistic planning improves probabilistic target coverage and OAR sparing—rests on five patients with no uncertainty analysis. More importantly, the PCE accuracy check in Appendix A only validates the fast in-optimization PCE against the more expensive evaluation PCE; it does not validate the evaluation PCE against direct dose-engine percentiles for these plans. The reported percentile error range (–0.74 to +1.16 GyRBE) overlaps the target-coverage gains (0.19–0.93 GyRBE), so the coverage improvements could partly be approximation artifact. The OAR reductions (8–33 GyRBE summed) are larger, but they are dominated by a few large DVH changes, and patient 2 actually got worse on the summed OAR metric. The warm-start comparison for patient 5 uses a slightly different constraint setting, though the paper discloses this and patient 4 is clean.\n\nThese are addressable, not fatal. A follow-up with more patients, a direct dose-engine check of the evaluation percentiles for both plan types, and reported confidence intervals would firm up the trade-off claim. The feasibility point stands.\n\nWho it's for: medical physics researchers working on robust or probabilistic IMPT planning. It deserves a serious referee—the methods are clearly described and the limitations are honestly discussed. I would accept it for review, with the expectation that the authors add the missing validation and temper the significance statement.\n\nBottom line: worth citing as a feasibility data point, but do not treat the quantitative gains as established.","headline":"Credible feasibility study of a previously published probabilistic IMPT method, but the headline trade-off gains are not yet quantitatively secure given the PCE error band and n=5.","tokens_in":31076,"tokens_out":2539,"would_cite":true,"duration_ms":23185,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Probability-based proton planning improves coverage and cuts organ dose","keywords":["probabilistic treatment planning","IMPT","intensity-modulated proton therapy","percentile-based optimization","polynomial chaos expansion","setup uncertainty","range uncertainty","robust optimization"],"falsifier":"Recalculate the final probabilistic and robust plans for these five patients using exact, scenario-by-scenario Monte Carlo dose calculation (or a validated independent dose engine) and compare the empirical 10th-percentile D99.8% and OAR DVH metrics; if the probabilistic plan's coverage improvement falls below roughly 0.2 GyRBE or reverses, the central claim is unsupported.","tokens_in":30092,"feed_emoji":"🎯","tokens_out":3932,"duration_ms":32318,"temperature":0.7,"pith_summary":"This paper argues that proton therapy plans for brain and skull-base tumors should be optimized for how likely each error scenario actually is, rather than for the single worst-case scenario. To do this, the authors use a percentile-based objective: instead of asking that every voxel be covered in every error scenario, they ask that the prescribed dose be delivered to the target in at least 90% of scenarios and that organs at risk stay below their limits in at least 95%. In five neuro-oncological patients, the resulting plans matched or improved target coverage while reducing cumulative organ-at-risk dose by up to 33 GyRBE compared with mini-max robust plans. The price is computation: full optimization took 44 to 141 hours, though a simple change (dropping the outer-loop warm-start) cut that below 10 hours for two patients. The approach matters because it turns robustness from an arbitrary scenario choice into a clinically meaningful probability of meeting each goal.","feed_headline":"Probability-based proton plans improve coverage and cut organ dose","feed_subtitle":"Optimizing for real error probabilities instead of worst cases delivered up to 33 GyRBE less organ dose on five patients.","key_machinery":"The core object is the percentile-based objective function, in which the α-th percentile of voxel dose is estimated as the expected dose plus or minus a δ-factor times an approximate standard deviation. The δ-factor is not fixed in advance; an outer loop repeatedly recomputes accurate percentiles from a polynomial chaos expansion of the dose-influence matrix (sampled 100,000 times) and updates δ until convergence, while an inner loop optimizes beam weights for the current percentile estimate. A memory-efficient diagonal covariance representation keeps the optimization tractable.","core_discovery":"The central claim is that a percentile-based, direct probabilistic optimizer can produce IMPT plans that are dosimetrically superior to mini-max robust plans when evaluated with the same probabilistic metrics. For the patient with adequate robust target coverage, the probabilistic plan had identical 10th-percentile D99.8% while lowering summed OAR DVH-measures by 19 GyRBE. For three of four patients whose robust plans could not cover the target enough, the probabilistic plan raised the 10th percentile of D99.8% by up to 0.93 GyRBE while also cutting summed OAR DVH-measures by 15 to 33 GyRBE; the fourth gained coverage at the cost of a within-constraint OAR increase of 8.1 GyRBE. The authors","pith_inferences":["If the reported gains hold under exact (non-PCE) dose calculation, the field's reliance on mini-max robustness and scenario-set selection could shift toward probability-based acceptance criteria, much as margin-based planning was replaced by robust optimization.","The PCE approximation error band (± roughly 1 GyRBE in voxel dose percentiles) overlaps some of the claimed coverage gains (0.19 to 0.93 GyRBE), so an independent Monte Carlo validation would strengthen confidence in the comparison.","The same percentile-optimization machinery could be extended to directly optimize pDVH metrics (e.g., D99.8% or D0.03cc) instead of voxel-wise percentiles, which the authors note is a current mismatch."],"forward_implications":["Probabilistic planning can replace or complement mini-max robust planning for IMPT, giving planners control over the exact probability of meeting each clinical goal.","The reported trade-off gains (up to 33 GyRBE OAR reduction or 0.93 GyRBE coverage gain) would translate to clinically meaningful reductions in the risk of radiation toxicity for neuro-oncological patients.","Optimization times can be brought within a clinically acceptable range (below 10 hours) by updating percentile estimates inside the inner loop instead of via an outer-loop warm-start, suggesting the method is practically deployable.","The approach is directly applicable to other modalities with steep dose gradients, such as photon stereotactic body radiotherapy."],"fun_headline_variants":["Probability-based proton plans cut organ dose up to 33 GyRBE","Probabilistic IMPT planning improves coverage and spares organs","Percentile-based proton planning: better coverage, less OAR dose","Direct probabilistic optimization beats robust in proton planning","Probabilistic proton therapy: target coverage up, organ dose down"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The comparison assumes the polynomial chaos expansion computes dose percentiles accurately enough—within the reported error of roughly ±0.7 to 1.2 GyRBE—so that the 0.19 to 0.93 GyRBE coverage improvements are real plan differences rather than approximation artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Probability-based proton plans cut organ dose up to 33 GyRBE","Probabilistic IMPT planning improves coverage and spares organs","Percentile-based proton planning: better coverage, less OAR dose","Direct probabilistic optimization beats robust in proton planning","Probabilistic proton therapy: target coverage up, organ dose down"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000645,"raw_usage":{"total_tokens":2877,"prompt_tokens":894,"completion_tokens":1983,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":638,"completion_tokens_details":{"reasoning_tokens":1897}},"tokens_in":638,"tokens_out":1983,"duration_ms":23406,"temperature":1.0,"reasoning_tokens":1897,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T03:27:34.128831+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recalculate the final probabilistic and robust plans for these five patients using exact, scenario-by-scenario Monte Carlo dose calculation (or a validated independent dose engine) and compare the empirical 10th-percentile D99.8% and OAR DVH metrics; if the probabilistic plan's coverage improvement falls below roughly 0.2 GyRBE or reverses, the central claim is unsupported.","supporting_citations":[],"review_version":1}