{"id":"9fbceda8-e3a3-45ab-a4bf-326adbd324ff","arxiv_id":"2411.16230","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Interplay-robust optimization, which models the timing between beam delivery and breathing motion during planning, improves target coverage and reduces organ-at-risk doses compared with conventional 4D robust optimization in simulated irregular-breathing lung patients.","lead":"This study tested a new way to plan proton beam treatments for lung cancer patients whose breathing is irregular. It found that modeling the exact timing of the beam delivery against tumor motion improves tumor coverage and lowers dose to healthy organs, which could make treatments faster and more comfortable.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"IPRO's advantage may be an artifact of evaluating on scenarios drawn from the same motion model and phase-start assumption; the claim is conditional on this distribution.","rationale":"The reader's weakest assumption and my stress-test concern coincide: the evaluation distribution mirrors the optimization distribution, with an explicit end-exhale start-phase assumption. This is the most load-bearing condition for the central claim because IPRO's entire rationale is to be robust to motion uncertainty; if the evaluation only samples the same motion model used in optimization, the reported improvements may be an upper bound rather than a realistic estimate. The paper is honest about this limitation, and the internal comparisons are plausible. However, the conclusion that IPRO 'could lead to more efficient mitigation of the interplay effect' for real patients depends on transfer to out-of-distribution motion, which the current experiments do not test. The proposed concrete test would settle whether the advantage survives distribution shift and phase uncertainty. Since the reader already returned CONDITIONAL, my concern does not change the verdict; it reinforces the need for external validation before clinical claims are accepted.","tokens_in":11375,"tokens_out":2928,"duration_ms":30630,"concrete_test":"Re-evaluate the IPRO-49S and 4DRO plans for case 3a using evaluation scenarios built from 4DMRI 2 breathing cycles (the motion pattern used for case 3b), and for case 3b using 4DMRI 1 cycles, while keeping the CT fixed. Additionally, generate evaluation scenarios that start at a uniformly random phase within the breathing cycle rather than at end-exhale. Compare the 5th percentile CTV D98 and the normalized OAR dose improvements. If the IPRO advantage over 4DRO falls below 1 percentage point or reverses, the central claim is not robust to the assumed scenario distribution.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that IPRO/IPRO-1C improve target coverage and OAR sparing over 4DRO for irregularly breathing lung patients. The evaluation (Sec. 2.5) generates 500 scenarios by randomly concatenating held-out breathing cycles from the same 4DMRI-derived s4DCT motion patterns used to generate optimization scenarios, with every cycle started at end-exhale (Sec. 2.4). Thus 'unseen at the optimization stage' means unseen cycles from the same motion distribution, not unseen motion characteristics. The optimization and evaluation sets are statistically matched (Table 1), and the authors explicitly state that the method relies on 'no uncertainty about the breathing cycle phase at the start of the delivery of each beam.' The authors also acknowledge in Sec. 4 that larger discrepancies between optimization and evaluation data are expected to decrease the robustness gain. If real treatment-day motion has different amplitudes, periods, or start phases, the minimax solution found by IPRO may not transfer, and the reported 4.2%/1.7% OAR improvements may shrink or disappear. The paper is internally consistent and transparent, but the headline claim, as phrased, overreaches beyond the narrow, matched-distribution simulation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes and evaluates interplay-robust optimization (IPRO) for pencil beam scanning proton therapy of lung patients with irregular breathing, extending prior work on frequency uncertainty to include amplitude variability. Motion is modeled with synthetic 4DCTs derived from 4DMRI patterns; optimization and evaluation scenarios are generated by randomly concatenating held-out breathing cycles. IPRO and a single-cycle variant IPRO-1C are compared with 4DRO on four synthetic patient cases. The authors report improved near-worst-case target coverage for IPRO and IPRO-1C, and, after equal-coverage dose normalization, reduced OAR doses with average near-worst-case improvements of 4.2% (IPRO-49S) and 1.7% (IPRO-1C-9S).","tokens_in":11624,"tokens_out":6232,"duration_ms":67989,"significance":"If the result holds outside the simulated setting, it is clinically relevant: explicitly modeling the spot-to-phase interplay during optimization could reduce the need for gating or rescanning, and IPRO-1C-9S is claimed to have computational cost comparable to 4DRO. The study has notable strengths: it uses a published delivery-time model, evaluates each plan in 500 motion scenarios as recommended in the literature, and is transparent about its limitations. However, the evidence base is four synthetic cases with evaluation scenarios drawn from the same motion model used for optimization, so the quantitative improvements reported here should be read as conditional on that matched distribution until out-of-distribution robustness is demonstrated.","major_comments":[{"comment":"The central claim that IPRO and IPRO-1C improve target coverage for irregularly breathing lung patients is supported only under a matched-distribution simulation: evaluation scenarios are generated by randomly concatenating held-out breathing cycles from the same s4DCT motion model, and every cycle is assumed to start at end-exhale with no phase uncertainty at beam start (Sec. 2.4). The authors themselves state in Sec. 4 that the case that may compromise IPRO's robustness is when motion states occurring during delivery are not represented in the optimization scenarios. The abstract's phrasing that IPRO 'increased the target coverage for all patient cases' therefore overreaches beyond the simulation setting. Please add an explicit out-of-distribution test, for example evaluating with a random start-phase offset or with breathing amplitude/period distributions deliberately perturbed relative to the optimization set, and temper the conclusions accordingly.","section":"Sec. 2.4 and Sec. 4"},{"comment":"The normalized comparison used to support the OAR-sparing claims is affected by a modeling inconsistency: all evaluation doses are scaled by a constant factor to match the 4DRO 5th-percentile CTV D98, and 'the effect on the time structure was ignored.' Uniformly scaling spot weights changes the delivery time structure, which changes the spot-to-motion-state assignment in Eq. (2) and hence the actual interplay-affected dose that would be delivered. The reported OAR improvements, including the 4.2% average for IPRO-49S, are therefore computed from doses that would not be delivered by the scaled plans. Please re-optimize the normalized plans, include the time-structure effect of the scaling, or quantitatively assess the error introduced by ignoring it.","section":"Sec. 3, Table 2"},{"comment":"The key numerical results are point estimates without confidence intervals or hypothesis tests. The 5th and 95th percentiles over 500 simulated scenarios are random quantities, and differences of 1.7% or even 4.2% could be within scenario sampling noise. Please provide bootstrap confidence intervals or other uncertainty quantification for the reported changes, and state exactly which OAR metrics and patient cases are included in the 'averaging' that produced the 4.2% and 1.7% figures in the abstract.","section":"Table 2 and Sec. 3"},{"comment":"Case 3b* is introduced after observing that the initial IPRO-1C breathing cycle was unrepresentative of the evaluation motion (Sec. 3.4). This post-hoc selection means that the IPRO-1C results in Table 2 are a mixture of a pre-specified case and a repaired case, so they do not provide an unbiased estimate of the method's typical performance when a single cycle is chosen. Combined with the small number of geometries and motion patterns, this limits the strength of the claim that IPRO-1C reliably improves over 4DRO. Please present 3b* explicitly as a post-hoc sensitivity analysis and adjust the corresponding conclusions.","section":"Sec. 3.4, Case 3b*"}],"minor_comments":[{"comment":"The caption states that the whiskers indicate the 5th and 95th percentiles but then repeats the same information; please clarify whether the box edges are quartiles and the whiskers are the stated percentiles, since this differs from standard Tukey boxplots.","section":"Figure 3 caption"},{"comment":"The IPRO-1C scenario grids for period scaling and start-phase shift are described only in prose; a small table or equation would make the scenario construction easier to follow and reproduce.","section":"Sec. 2.7"},{"comment":"The abstract's '4.2 %' and '1.7 %' improvements should specify the exact set of OAR metrics and cases over which the average is taken, because Table 2 shows considerable variation by case and metric.","section":"Abstract and Sec. 3"},{"comment":"The set I^p(x; s) is introduced as 'I p(x; s)' in the text but printed with a superscript in Eq. (1); please unify the notation for readability.","section":"Sec. 2.1, Eq. (1)"},{"comment":"Reference [15] is cited as 's4DCT(MRI)' in the bibliography while the text uses 's4DCT'; please add a clear definition at first use and make the reference title consistent.","section":"References"},{"comment":"The text says the template plan was optimized with 40 iterations of 4DRO and that the 4DRO plan used 40 additional iterations; please clarify whether the total number of iterations for the final 4DRO plan is 80 and how this compares with the 40 iterations used for the other methods.","section":"Sec. 2.6"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is transparent and internally consistent, and the authors correctly identify the main threat to external validity in Sec. 4. My recommendation of major revision is driven by three load-bearing issues: the in-distribution evaluation limits the generalizability of the headline claim, the dose normalization ignores the time-structure effect and thus biases the reported OAR improvements, and the key numerical results lack uncertainty quantification. These are fixable within the scope of the paper: add an out-of-distribution evaluation, re-analyze or explicitly qualify the normalized comparison, and report confidence intervals. I would not recommend rejection, because the core methodological idea is sound and the current evidence, while narrow, is honestly presented."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a solid, transparent simulation study that extends IPRO from frequency-only uncertainty to amplitude-irregular breathing, and it adds a single-cycle variant (IPRO-1C) that looks clinically relevant. Within the stated simulation envelope, the claims hold up. The biggest caveat, which the authors themselves flag, is that the evaluation scenarios are generated from the same motion model and phase-start assumption as the optimization scenarios, so the magnitude of the benefit in real patients is still an open question.\n\nWhat's new here is legitimate: prior IPRO work (Engwall et al.) varied breathing period only; this paper varies amplitude too, uses multi-cycle synthetic 4DCTs from 4DMRI data, and evaluates on held-out cycles. The IPRO-1C idea, generating scenarios by period-scaling and phase-shifting a single cycle, is a practical contribution because it can be done with the data you'd have in a typical clinic. The evaluation design, 500 scenarios per plan and disjoint optimization/evaluation cycle sets, is a clear step up from earlier studies.\n\nThe paper is also honest. The limitations section names the missing ribcage motion, the absence of range/setup errors, and the assumption that each beam starts at end-exhale. The post-hoc case 3b* is presented as an illustration, not hidden. The normalization to equal target coverage is described with its caveat (dose scaling without recomputing the time structure), which is minor since the scaling factors are small.\n\nThe soft spots are proportional. Four synthetic cases and no code or data mean the quantitative results shouldn't be treated as established. The matched-distribution evaluation is the main one: the held-out cycles are statistically similar to the optimization cycles (Table 1), and every cycle starts at end-exhale. If treatment-day breathing has different amplitudes or phase at beam start, the IPRO advantage could shrink. The authors acknowledge exactly this in Section 4. That doesn't undermine the study; it bounds it.\n\nThe spot-to-state assignment heuristic (updated every 10 iterations) is inherited from prior work and is a known limitation, not a new flaw.\n\nWho is this for? Researchers in proton therapy motion management and robust optimization. It's a worthwhile contribution. I'd send it to peer review, asking the referees to focus on whether the evaluation distribution is wide enough and whether the OAR improvements survive a reasonable perturbation of the motion model.","headline":"A transparent simulation study extending IPRO to amplitude-irregular breathing; the advantage over 4DRO holds within the matched-distribution setup, but external motion variability remains untested.","tokens_in":12158,"tokens_out":2264,"would_cite":true,"duration_ms":25514,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper shows that explicitly modeling the time-dependent interference between pencil beam delivery and irregular breathing during plan optimization raises near-worst-case target coverage in lung proton therapy and, at equal coverage…","keywords":["interplay effect","robust optimization","pencil beam scanning","proton therapy","4D dose computation","irregular breathing","lung cancer","synthetic 4DCT"],"falsifier":"Simulate delivery of the IPRO and 4DRO plans using evaluation scenarios drawn from a breathing distribution with, say, 30% larger amplitudes and a uniformly random start phase per beam, and check whether the 5th-percentile CTV D98 of IPRO still exceeds 4DRO; if the margin vanishes or reverses, the claim fails for unrepresented motion states.","tokens_in":11149,"feed_emoji":"🫁","tokens_out":5822,"duration_ms":47135,"temperature":0.7,"pith_summary":"Interplay-robust optimization (IPRO) treats the interference between the scanning beam's arrival times and the patient's breathing as an uncertainty to be optimized against, rather than averaged away. This paper asks whether IPRO still helps when breathing is irregular, varying in both period and amplitude, and when only one breathing cycle is available for planning. On synthetic 4DCTs built from 4DMRI motion patterns, IPRO and a single-cycle variant (IPRO-1C) improved the near-worst-case target coverage (5th percentile CTV D98) for every patient case compared with conventional 4D robust optimization. After scaling all plans to equal target coverage, IPRO with 49 scenarios lowered near-worst-case organ-at-risk dose by an average of 4.2%, and IPRO-1C with 9 scenarios, at the same computational cost as 4DRO, lowered it by 1.7%. If the result holds in the clinic, explicit modeling of interplay during planning could reduce the need for gating, breath-hold, and re-scanning.","feed_headline":"Beam–breathing interplay modeling cuts healthy-tissue dose 4.2%","feed_subtitle":"Robust optimization over many breathing scenarios beats standard 4D planning in lung proton therapy—even with one breathing cycle.","key_machinery":"The central object is the 4D dose computation (4DDC) with phase sorting: each pencil beam spot's dose is computed on the motion state it hits during delivery, deformed to a reference state, and summed. Interplay-robust optimization minimizes a worst-case objective over a set of motion scenarios, each a random concatenation of breathing cycles; the delivery time structure is updated heuristically as spot weights change. IPRO-1C builds its scenario set from one breathing cycle, varied by period-scaling factors and start-phase shifts, to test how much robustness can be recovered when data are limited.","core_discovery":"The authors claim that explicitly modeling the interplay effect, the time-dependent interference between pencil beam spot delivery and breathing motion, during robust optimization produces lung cancer proton plans that are more robust to irregular breathing than plans from phase-averaged 4D robust optimization. Using synthetic 4DCTs with multiple breathing cycles of varying period and amplitude, they show that IPRO, which optimizes against randomly concatenated breathing cycles, and IPRO-1C, which generates scenarios from a single breathing cycle with scaled periods and shifted start phases, both increase the 5th-percentile CTV D98 relative to 4DRO in all four patient cases. After normalizing each plan to match the 4DRO near-worst-case target coverage, the robustly optimized plans spare organs at risk more: IPRO with 49 scenarios reduces the 95th-percentile OAR dose by an average of 4.2%, and IPRO-1C with 9 scenarios by 1.7% while requiring no more dose computations than 4DRO.","pith_inferences":["If the same margin persists under randomized start phases and day-to-day amplitude drift, IPRO could allow fewer re-scans or wider gating windows, shortening treatment times and improving comfort, an implication the authors state as motivation but do not demonstrate.","The 4.2% OAR improvement is likely an optimistic bound because evaluation scenarios are drawn from the same motion model used for optimization; a fairer clinical test would use independent motion data acquired on a different day.","The IPRO-1C result suggests that a single pre-treatment breathing cycle, augmented by period and phase perturbations, may be sufficient to capture the interplay uncertainty; a testable extension would compare IPRO-1C against gated delivery on the same patients."],"forward_implications":["IPRO and IPRO-1C raise the 5th-percentile CTV D98 relative to 4DRO in every patient case studied, with little change to dose homogeneity.","After normalizing to equal target coverage, IPRO with 49 scenarios reduces the near-worst-case (95th percentile) OAR dose by an average of 4.2%, and IPRO-1C with 9 scenarios by 1.7%.","Increasing the scenario count from 9 to 49 typically improves robustness, and using multiple breathing cycles outperforms a single cycle when that single cycle is unrepresentative of the evaluation motion.","IPRO-1C with 9 scenarios uses the same number of motion states and dose computations as 4DRO, so the robustness gain comes at no added computational cost."],"supporting_citations":[{"why":"Introduced 4D dose computation in the optimization loop, the dose engine IPRO builds on.","marker":"[12]"},{"why":"Proposed IPRO for uncertainty in breathing frequency; this paper extends it to amplitude variability.","marker":"[14]"},{"why":"Provided the synthetic 4DCT generation method with multiple breathing cycles used for optimization and evaluation.","marker":"[15]"},{"why":"Supplied the experimental delivery time-structure model for the pencil beam system.","marker":"[18]"},{"why":"Justified the use of 500 evaluation scenarios for statistically accurate interplay analysis.","marker":"[19]"},{"why":"Described the commercial 4DRO optimization used as the baseline comparison.","marker":"[3]"}],"fun_headline_variants":["Interplay-robust optimization reduces OAR dose 4.2% in lung proton plans","Modeling breathing motion improves robust lung proton planning","New IPRO method spares healthy tissue in lung proton therapy","Breathing-adaptive optimization beats standard 4D robust planning","Single-breathing-cycle robust optimization also cuts OAR dose 1.7%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation scenarios are built from the same synthetic motion model and cycle statistics as the optimization scenarios, with every breathing cycle assumed to start at end-exhale; if a patient's real on-couch breathing differs in amplitude, period, or start phase, the reported robustness gain shrinks or disappears.","fun_headline_variants_meta":{"raw":{"variants":["Interplay-robust optimization reduces OAR dose 4.2% in lung proton plans","Modeling breathing motion improves robust lung proton planning","New IPRO method spares healthy tissue in lung proton therapy","Breathing-adaptive optimization beats standard 4D robust planning","Single-breathing-cycle robust optimization also cuts OAR dose 1.7%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000324,"raw_usage":{"total_tokens":1903,"prompt_tokens":1117,"completion_tokens":786,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":733,"completion_tokens_details":{"reasoning_tokens":692}},"tokens_in":733,"tokens_out":786,"duration_ms":35100,"temperature":1.0,"reasoning_tokens":692,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T13:21:11.483836+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Simulate delivery of the IPRO and 4DRO plans using evaluation scenarios drawn from a breathing distribution with, say, 30% larger amplitudes and a uniformly random start phase per beam, and check whether the 5th-percentile CTV D98 of IPRO still exceeds 4DRO; if the margin vanishes or reverses, the claim fails for unrepresented motion states.","supporting_citations":[{"cited_title":"4D robust optimization including uncertainties in time structures can reduce the interplay effect in proton pencil beam scanning radiation therapy","cited_arxiv_id":null,"evidence_quote":"Proposed IPRO for uncertainty in breathing frequency; this paper extends it to amplitude variability."},{"cited_title":"Synthetic 4DCT(MRI) lung phantom generation for 4D radiotherapy and image guidance investigations","cited_arxiv_id":null,"evidence_quote":"Provided the synthetic 4DCT generation method with multiple breathing cycles used for optimization and evaluation."},{"cited_title":"How should we model and evaluate breathing interplay effects in IMPT?","cited_arxiv_id":null,"evidence_quote":"Justified the use of 500 evaluation scenarios for statistically accurate interplay analysis."},{"cited_title":"Treatment planning of scanned proton beams in RaySta- tion","cited_arxiv_id":null,"evidence_quote":"Described the commercial 4DRO optimization used as the baseline comparison."}],"review_version":1}