{"id":"532fa9e7-3329-41f6-a438-05545e34c52d","arxiv_id":"2411.18074","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A greedy energy-layer-optimization method, integrated into an in-house proton treatment planning system, cut delivery times by 18-49% in three clinical cases while passing patient-specific QA.","lead":"Researchers built an in-house treatment planning system for proton therapy that uses a greedy algorithm to reduce the number of energy layers, and tested it on a clinical proton machine. In brain, lung, and abdomen cases, plans passed QA while delivery time dropped by roughly 18-49 percent depending on the case.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Beam-model validation is not independent: IH-TPS is built from RayStation MC data and compared only against RayStation, while measured QA recalculates plans in RayStation, so IH-TPS patient dose accuracy is not established.","rationale":"The reader's weakest assumption is the same one I identify: the beam model is built from RayStation MC and validated against RayStation, so the agreement does not independently establish IH-TPS dose accuracy. The paper has real strengths—machine delivery on a clinical proton system, patient-specific QA, and detailed dosimetric tables—but the QA workflow's comparison to RayStation recalculated dose rather than IH-TPS predicted dose means the experimental measurements do not close the loop on the IH-TPS dose engine.\n\nThe abstract's 62% delivery-time reduction for the brain case is inconsistent with Table 4 (105→54 s is 48.6%), and Tables 2–3 show CI and Dmax degrade under ELO while the text says plan quality is maintained. These internal inconsistencies are real and reinforce the need for careful verification, but they are secondary to the independence gap: even if every reported percentage were exact, the central claim would still require direct evidence that IH-TPS predicts dose in patient-like geometries.\n\nA direct IH-TPS-vs-measured phantom comparison, or an independent Monte Carlo recalculation, would settle the issue. Until then, the CONDITIONAL verdict with moderate confidence remains appropriate; I do not see grounds to reject the work, since the reported engineering results and experimental QA are substantive.","tokens_in":11287,"tokens_out":6251,"duration_ms":58006,"concrete_test":"Perform a dedicated QA comparison in which the ELO plan's delivered spot sequence is measured with MatriXX or ion chambers in a heterogeneous/anthropomorphic phantom, and compare that measured dose directly to the IH-TPS-calculated 3D dose using the same 2%/2mm, 10% threshold gamma criterion, rather than comparing to the RayStation recalculation. Additionally, recompute the same ELO plan in an independent Monte Carlo engine (e.g., TOPAS/Geant4) on the patient CT and compare DVHs and gamma to IH-TPS. If direct IH-TPS-vs-measured gamma is ≥95% and independent MC target/OAR DVH differences are within ~2%, the independence gap is closed and the strongest claim holds. If IH-TPS-vs-measured gamma falls below 95% while RayStation-vs-measured remains ≥95%, the current QA protocol has masked a beam-model error.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim is that an IH-TPS has been experimentally validated for fast IMPT with clinically acceptable plan quality. The load-bearing condition is that the IH-TPS pencil-beam engine accurately predicts dose in patient CT geometries. That condition is not independently tested.\n\nIn §2.1, the beam model is constructed from RayStation Monte Carlo single-spot simulations in a 40×40×40 cm³ water phantom (70–225 MeV at 2.5 MeV intervals). In §3.1, the patient-geometry validation compares IH-TPS dose to RayStation dose—the same system that supplied the base data. In §2.3 and §3.3, ELO plans are exported as spot weight/location CSV files, imported into RayStation for delivery, and recalculated in RayStation; the reported >95% gamma pass rates are measured against RayStation's recalculated 3D dose, not against IH-TPS's own predicted dose.\n\nThus the experimental QA verifies that the spot lists are deliverable and that RayStation models them correctly, but it does not verify that IH-TPS's pencil-beam algorithm—including CT-to-electron-density conversion and beam-parameter interpolation—is accurate in heterogeneous tissue. If the IH-TPS dose engine were biased in patient anatomy, the optimization could select suboptimal or unsafe energy layers and spot weights while all reported metrics (IH-TPS vs RayStation gamma, RayStation-recalculated DVHs, and QA vs RayStation) would still look acceptable. This is a specific, falsifiable gap, not a disagreement with consensus.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports an in-house treatment planning system (IH-TPS) for pencil-beam scanning proton therapy, combining a pencil-beam dose engine with a greedy energy-layer optimization (ELO) algorithm inspired by block orthogonal matching pursuit. The system is tested on three clinical cases (abdomen, brain, lung) by comparing IH-TPS dose distributions with RayStation Monte Carlo, by measuring patient-specific QA on an IBA Proteus ONE machine, and by recording delivery times with and without ELO. The authors report that ELO reduces energy layers and delivery times (e.g., brain: 78 to 40 layers, 105 to 54 s) while preserving plan quality, with QA pass rates above 95% in all cases.","tokens_in":11576,"tokens_out":6190,"duration_ms":53621,"significance":"The work is practically significant if the claims hold: it demonstrates a route to integrate energy-layer optimization into a self-developed TPS, with measured delivery-time reductions on a clinical machine. The patient-specific QA and the direct measurement of delivery times are genuine strengths, as is the explicit use of clinical cases with standard beam angles and DVH objectives. However, the dosimetric validation of the IH-TPS dose engine is not fully independent: the beam model is parameterized from RayStation Monte Carlo data and then compared back against RayStation, leaving the absolute accuracy of the pencil-beam engine in patient geometries without a direct measurement-based check. The plan-quality preservation claim is also weaker than stated when the ELO plans are recalculated in RayStation, as the reported conformity and Dmax metrics show systematic degradation.","major_comments":[{"comment":"","section":"§2.1 and §3.1"},{"comment":"","section":"§3.3 and Table 3"},{"comment":"","section":"§2.3 and §3.3"}],"minor_comments":[{"comment":"","section":"References"},{"comment":"","section":"Equation (2)"},{"comment":"","section":"§2.2.2 algorithm"},{"comment":"","section":"§3.1.1"},{"comment":"","section":"Figure 1 caption"}],"recommendation":"major_revision","confidential_remarks":"The central concern is the circularity of the dose-engine validation: the beam model is derived from RayStation MC and compared back to RayStation, while the QA validates delivery against RayStation's recalculation. The delivery-time and QA pass-rate results are useful and probably correct, but the paper's main claim of an experimentally validated IH-TPS requires a direct measured-to-IH-TPS comparison or a clear limitation statement. The plan-quality preservation claim also needs to be tempered or justified with explicit clinical criteria. These issues are likely addressable with additional experiments (e.g., measuring beam data or recomputing QA against IH-TPS dose) and revised wording, so major revision is appropriate rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is worth a serious look: it reports a working in-house TPS with greedy block orthogonal matching pursuit for energy layer selection, and it validates the plans on a clinical IBA Proteus ONE machine. The measured delivery time reductions (110→90 s, 105→54 s, 178→127 s) and the >95% gamma pass rates in patient-specific QA are real, independent evidence that the ELO plans are deliverable. That is more than most prior ELO papers do, and it is the paper's main contribution.\n\nThe greedy method itself is a reasonable adaptation of compressive-sensing block OMP to the energy-layer-selection problem. The authors are honest in the discussion that plan quality degrades as energy layers are removed, and they show the trade-off in Tables 2 and 3. But the text in the abstract and conclusion says \"maintaining plan quality,\" which is too strong. The objective values increase, CI drops, and Dmax rises for every case and every ELO level; that is degradation, not maintenance, even if it is clinically acceptable. This should be reworded.\n\nThe bigger soft spot is the dose-engine validation. The beam model is built from RayStation Monte Carlo single-spot simulations in water, and then the IH-TPS patient-dose accuracy is compared against RayStation as the benchmark. That is circular. The QA measurements are against RayStation-recalculated doses, not against an independent measurement of the IH-TPS-predicted dose. So the paper establishes that the spot lists are deliverable and that RayStation models them correctly, but it does not establish that the IH-TPS pencil-beam algorithm is accurate in heterogeneous patient geometries. The authors should either validate against measured beam data (which the abstract claims but the methods don't actually show) or against a second independent dose engine.\n\nThere is also a numerical inconsistency: the abstract says a brain case with 78→40 layers gave a 62% reduction in delivery time, but Table 4 shows 105→54 s, a 48.6% reduction. And there is no comparison to any of the previous ELO methods ([10]–[14]), so the claimed advantage over existing approaches is not demonstrated.\n\nThe equations in the preprint are corrupted (symbols missing, garbled display), which made it hard to verify the algorithm details. That is likely a rendering issue, but it still needs fixing.\n\nBottom line: the engineering contribution is real, the delivery-time measurements are solid, and the QA is a step beyond prior work. The dose-engine validation is not independent, the quality-maintenance claim is overstated, and the abstract has a numerical error. These are fixable with a revision. I would send this to peer review.\n\nRecommendation: accept for peer review with major revision. The authors should add independent beam-model validation, correct the abstract, and tone down the plan-quality claims.","headline":"A useful, honestly-reported engineering validation of greedy energy-layer optimization on a clinical proton machine, but the dose-engine validation is benchmarked against its own calibration source.","tokens_in":12127,"tokens_out":1610,"would_cite":true,"duration_ms":16302,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An in-house proton treatment planning system with greedy energy-layer selection can produce IMPT plans that are faster to deliver and still pass clinical QA.","keywords":["proton therapy","intensity-modulated proton therapy","pencil beam scanning","treatment planning system","energy layer optimization","orthogonal matching pursuit","block sparsity","patient-specific quality assurance"],"falsifier":"Measure the integral depth dose and lateral profiles of single spots for a spread of energies on the Proteus ONE with an ionization chamber or film, then compare them to IH-TPS predictions: if gamma pass rates fall below 95% (2 mm, 2%) on those beam-data measurements, the dose engine's machine model is the weak link. Alternatively, run the ELO plans through an independent Monte Carlo engine and a heterogeneous phantom measurement; disagreement there would separate beam-model error from the optimization's benefit.","tokens_in":11059,"feed_emoji":"⚛️","tokens_out":6902,"duration_ms":57740,"temperature":0.7,"pith_summary":"The paper argues that an open, in-house treatment planning system can do what closed commercial platforms make difficult: optimize intensity-modulated proton therapy (IMPT) plans over the number of energy layers, not just spot weights, and then deliver those plans on a real proton machine. The authors built a pencil-beam dose engine matched to an IBA Proteus ONE machine, added a greedy block-selection algorithm that picks a preset number of energy layers, and tested it on abdomen, brain, and lung cases. In their hands the optimized plans used substantially fewer energy layers (for example 40 instead of 78 for the brain), cut measured delivery time from 105 to 54 seconds, and still passed patient-specific quality assurance with gamma pass rates above 95%. If correct, this removes a practical barrier: energy-layer optimization can be experimentally validated and clinically used without waiting for commercial TPS vendors to add it.","feed_headline":"Greedy energy-layer selection cuts proton energy layers by up to 62%","feed_subtitle":"Fewer layers trim delivery time by up to 49% on a clinical proton machine, with QA pass rates above 95%.","key_machinery":"The load-bearing object is the energy-layer optimization formulated as block-sparse inverse planning and solved by a modified orthogonal matching pursuit. The full dose-influence matrix is partitioned into blocks, one block per proton energy, and limiting the number of energies is expressed as the constraint that the beam-intensity vector have at most N nonzero blocks. At each iteration the method picks the energy block whose dose most reduces the residual of the quadratic dose-matching problem, then re-optimizes the full plan (with DVH and minimum-monitor-unit constraints) on the enlarged support, keeping the new energy only if the plan objective actually decreases. This turns a nonconvex combinatorial selection into a sequence of tractable projections.","core_discovery":"The central claim is that a self-contained in-house TPS can combine a clinically matched pencil-beam dose engine with a greedy energy-layer optimization and produce IMPT plans that are both faster to deliver and clinically acceptable. The paper reports that the dose engine's single-spot depth-dose and lateral profiles agree with RayStation's Monte Carlo to gamma pass rates of 96.5% to 100% (2 mm, 2%), and that full plans agree with RayStation at 95.6% to 97.1% (2 mm, 2%). With the greedy selection, the number of energy layers fell from 50 to 30 for the abdomen, 78 to 40 for the brain, and 100 to 70 for the lung, with measured delivery times dropping from 110 to 90 s, 105 to 54 s, and 178 to 127 s respectively. Recalculated in RayStation, the reduced-energy plans retained comparable target coverage and organ-at-risk sparing, and all delivered beams passed the clinical 95% gamma threshold.","pith_inferences":["If the beam model were calibrated directly from measured machine data rather than from RayStation Monte Carlo simulations in water, the greedy selection could plausibly push energy counts even lower than the 30, 40, and 70 values reported, since the paper's own comparison shows the RayStation-recalculated plans degraded slightly relative to the IH-TPS-computed ones.","The same block-greedy idea transfers naturally to other combinatorial choices in radiotherapy, such as beam-angle selection or minimum-monitor-unit constraints; a testable extension is applying it to FLASH dose-rate optimization, where delivery time constraints are tighter.","A stronger validation would compare IH-TPS dose against independent Monte Carlo or measured data in heterogeneous phantoms rather than against RayStation, which supplied the base beam data; such a test would separately quantify dose-engine accuracy and the benefit of the ELO method."],"forward_implications":["Energy-layer reduction is a practical, experimentally verified lever: in the reported cases it cut measured delivery time by roughly 18% to 49%.","Because the TPS is in-house, new optimization constraints such as dose-rate limits or delivery-sequence penalties can be added and validated on a clinical machine without waiting for a commercial vendor.","The greedy selection extends beyond the simplified quadratic problem to full clinical plans with DVH objectives and minimum-monitor-unit constraints.","Plans with fewer energy layers remain deliverable when recalculated in the clinical TPS, meaning the efficiency gain is not an artifact of the in-house dose engine alone."],"supporting_citations":[{"why":"Supplies the pencil-beam dose calculation algorithm that IH-TPS modifies to model the Proteus ONE machine.","marker":"[15]"},{"why":"Provides the orthogonal matching pursuit method, modified here for block-wise greedy energy layer selection.","marker":"[27]"},{"why":"Establishes the block-sparse recovery theory and efficient greedy recovery used to justify selecting energy blocks.","marker":"[26]"},{"why":"Documents that energy-layer switching time dominates total delivery time in cyclotron-based proton systems, motivating the objective.","marker":"[8]"},{"why":"Models spot-scanning delivery time and sequence for compact superconducting synchrocyclotron systems, providing context for the time reduction.","marker":"[9]"},{"why":"Defines the gamma-index technique used to compare IH-TPS dose against RayStation and against measured QA.","marker":"[33]"},{"why":"Formulates the DVH-based plan objectives used in the inverse planning objective function.","marker":"[17]"}],"fun_headline_variants":["In-house TPS with greedy ELO speeds IMPT delivery by 49%","Greedy layer selection cuts proton therapy delivery time by 49%","Greedy energy-layer pruning accelerates IMPT on clinical machine","Custom TPS trims proton plan delivery via greedy layer selection"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that IH-TPS's beam model, built from RayStation Monte Carlo single-spot dose simulations in water, faithfully represents the actual IBA Proteus ONE machine when applied to patient anatomy by the pencil-beam algorithm.","fun_headline_variants_meta":{"raw":{"variants":["In-house TPS with greedy ELO speeds IMPT delivery by 49%","Greedy layer selection cuts proton therapy delivery time by 49%","Greedy energy-layer pruning accelerates IMPT on clinical machine","Custom TPS trims proton plan delivery via greedy layer selection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001551,"raw_usage":{"total_tokens":6275,"prompt_tokens":1092,"completion_tokens":5183,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":708,"completion_tokens_details":{"reasoning_tokens":5106}},"tokens_in":708,"tokens_out":5183,"duration_ms":31683,"temperature":1.0,"reasoning_tokens":5106,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:32:02.486441+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the integral depth dose and lateral profiles of single spots for a spread of energies on the Proteus ONE with an ionization chamber or film, then compare them to IH-TPS predictions: if gamma pass rates fall below 95% (2 mm, 2%) on those beam-data measurements, the dose engine's machine model is the weak link. Alternatively, run the ELO plans through an independent Monte Carlo engine and a heterogeneous phantom measurement; disagreement there would separate beam-model error from the optimization's benefit.","supporting_citations":[{"cited_title":"and Urie, M., 1996","cited_arxiv_id":null,"evidence_quote":"Supplies the pencil-beam dose calculation algorithm that IH-TPS modifies to model the Proteus ONE machine."},{"cited_title":"C., Kuppinger, P., Bolcskei, H., 2010","cited_arxiv_id":null,"evidence_quote":"Establishes the block-sparse recovery theory and efficient greedy recovery used to justify selecting energy blocks."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents that energy-layer switching time dominates total delivery time in cyclotron-based proton systems, motivating the objective."},{"cited_title":"A technique for the quantitative evaluation of dose distributions.MedicalPhysics,25(5),656-61","cited_arxiv_id":null,"evidence_quote":"Defines the gamma-index technique used to compare IH-TPS dose against RayStation and against measured QA."},{"cited_title":"Algorithms and functionality of an intensity modulated radiotherapy optimizationsystem.MedicalPhysics,27(4),701-711","cited_arxiv_id":null,"evidence_quote":"Formulates the DVH-based plan objectives used in the inverse planning objective function."}],"review_version":1}