{"id":"eaa6140e-1ccc-4f66-911e-ab46bc4365dc","arxiv_id":"2608.12656","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":3,"one_line_summary":"Non-coplanar 4π planning substantially lowers pelvic bone marrow doses in cervical cancer, but the headline 23% hematologic toxicity reduction is a model extrapolation that loses statistical significance when model training uncertainty is included.","lead":"This study replanned 114 cervical cancer patients with non-coplanar 4π radiotherapy and used a machine learning model to estimate that the new plans reduce predicted hematologic toxicity risk by 23%. The dosimetric improvements are large, but the clinical benefit is a model extrapolation, and the full uncertainty analysis does not show statistical significance.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's 23% risk-reduction claim is not supported by the paper's own full-pipeline bootstrap: RR 0.80, 95% CI 0.57–1.03, which crosses 1.0 (Section 3.5), while the abstract cites RR 0.77, CI 0.75–0.79.","rationale":"The reader's formal 'weakest assumption' was extrapolation beyond the VMAT training dose range, but the reader's rationale also explicitly notes the contradiction between the abstract's RR 0.77 CI 0.75–0.79 and the paper's own full-pipeline estimate RR 0.80 CI 0.57–1.03. My stress-test identifies that internal statistical inconsistency as the single most load-bearing concern: the paper's central clinical claim is a model-based risk reduction, and the only analysis that properly propagates model-training uncertainty does not reach statistical significance. This is not a matter of disagreeing with consensus; it is a correctness issue within the manuscript's own reported numbers. The dosimetric comparison appears credible and would support a feasibility or hypothesis-generating conclusion, but the abstract's causal-sounding claim of a 'marked reduction' is not supported by the full-pipeline analysis. Because the submitted paper presents this unsupported claim as its headline result, rejection (or at minimum major revision) remains appropriate. The concrete test—rerunning the full-pipeline bootstrap and auditing the origin of the 0.75–0.79 CI—would settle whether the contradiction is merely a reporting artifact or a genuine absence of statistical significance. Given the magnitudes involved, the test will very likely confirm the concern. The verdict is unchanged from the reader's reject decision.","tokens_in":14254,"tokens_out":3791,"duration_ms":36810,"concrete_test":"Re-run the full-pipeline bootstrap from Section 3.5 with a fixed random seed and 10,000 resamples, stratified by toxicity outcome, and report the 95% CI for the risk ratio. If the interval still contains 1.0, the abstract's claim of a significant 23% reduction is unsupported. Separately, independently recompute the 'direct point estimate' CI to identify what procedure produced 0.75–0.79; a point estimate without bootstrapping has no CI, so the interval must come from some model assumption. If it derives from a fixed-model or CLT-based calculation, it does not reflect full-pipeline uncertainty and cannot support the abstract's inferential claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central clinical claim—that 4π planning reduces acute hematologic toxicity by 23%—rests entirely on the risk ratio reported in the abstract: RR 0.77 (95% CI 0.75–0.79). Section 3.5, however, gives three different estimates: full-pipeline bootstrap RR = 0.80 (95% CI 0.57–1.03), fixed-model bootstrap RR = 0.88 (95% CI 0.87–0.90), and a 'direct point estimate without bootstrapping' RR = 0.77 with a CI. The full-pipeline bootstrap is the only analysis that accounts for model-refitting uncertainty, and its 95% interval crosses 1.0, meaning the risk reduction is not statistically significant. The abstract cites only the optimistic 0.77 interval and omits the full-pipeline result. Since the conclusion that 'superior dosimetry should translate to a marked reduction' is explicitly premised on the model-based RR, the headline claim is internally inconsistent with the paper's own uncertainty propagation. Furthermore, the fixed-model interval (0.87–0.90) does not even contain the abstract's point estimate of 0.77, so the reported precision is not reproducible from the disclosed procedures. This statistical inconsistency is load-bearing independently of extrapolation concerns: even if the dose-response model extrapolated perfectly, the appropriate full-uncertainty analysis does not establish a significant benefit.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a retrospective dosimetric planning study of 114 cervical cancer patients treated with volumetric modulated arc therapy (VMAT). The authors develop a radiomics-based machine learning model to predict acute grade ≥2 hematologic toxicity (HT), then replan each patient with non-coplanar 4π radiotherapy and apply the model to estimate the reduction in predicted HT risk. The dosimetric comparison shows substantial reductions in pelvic bone marrow doses under 4π, and the model achieves an AUC of 0.79. The central clinical claim is a 23% reduction in predicted HT risk (RR 0.77, 95% CI 0.75–0.79). The paper also proposes a risk-based triage strategy for selecting patients for 4π planning.","tokens_in":14663,"tokens_out":5977,"duration_ms":53652,"significance":"If valid, the demonstrated ability to reduce pelvic bone marrow doses by 28–52% without compromising target coverage would be clinically valuable in cervical cancer radiotherapy, and a validated model-based prediction of hematologic toxicity reduction would strengthen the case for 4π planning. The study is also notable for its transparent reporting of multiple bootstrap procedures and for combining radiomics with dosimetric and clinical features in a single predictive pipeline. However, the primary clinical claim is not supported by the paper's own full-pipeline uncertainty analysis, the risk model is applied outside its training range, and the stated OAR-sparing benefit is contradicted by several dosimetric comparisons in the appendix. The dosimetric planning results are likely publishable as a dose-reduction study, but the current framing substantially overstates the evidence for clinical toxicity reduction.","major_comments":[{"comment":"The abstract's central claim of a 23% risk reduction (risk ratio 0.77, 95% CI 0.75–0.79) is not supported by the paper's own full-pipeline bootstrap analysis in Section 3.5, which reports RR 0.80 with 95% CI 0.57–1.03, crossing 1.0. The fixed-model bootstrap (RR 0.88, 95% CI 0.87–0.90) does not contain the abstract's point estimate, so the reported precision is not reproducible from the disclosed procedures. The abstract cites only the direct point estimate, which does not account for model-refitting uncertainty, and the conclusion that 'superior dosimetry should translate to a marked reduction' is therefore not internally consistent with the full uncertainty analysis.","section":"Abstract; §3.5"},{"comment":"The model was trained on historical VMAT plans and then applied to 4π dose distributions that lie substantially outside the training range (e.g., all-region pelvic bone V20Gy falls from 0.75 to 0.36, Table F1). As the authors acknowledge in Section 4, this extrapolation means the model 'likely underestimated the clinical benefits,' which is an admission that the 23% risk reduction is a model extrapolation, not an empirically grounded estimate. Because the reported risk ratio is directly computed from the fitted dose-response model, the magnitude of the predicted benefit cannot be validated by the data in this study, and this limitation is load-bearing for the central claim.","section":"§4 (Limitations)"},{"comment":"The abstract claims 4π planning reduces dose 'without compromising target coverage or sparing of other OARs,' but Appendix F Table F2 shows several OAR dose metrics that significantly increase under 4π: rectum V40Gy (0.45 to 0.48, p=0.02), rectum V45Gy (0.29 to 0.35, p=5e-5), and rectum V50Gy (0.08 to 0.16, p=1e-7). Section 3.4 cherry-picks examples of OAR reductions and ignores these increases, so the stated OAR-sparing claim is directly contradicted by the paper's own dosimetric results.","section":"Abstract; §3.4; Appendix F"},{"comment":"The method for computing the 'direct point estimate without bootstrapping' and its 95% CI (0.75–0.79) is not described anywhere in the manuscript. Since this is the estimate quoted in the abstract, the lack of a reproducible procedure is a significant presentation gap that affects the main result.","section":"§3.5"}],"minor_comments":[{"comment":"Section 3 has no subsection 3.2; the text jumps from 3.1 to 3.3, which is confusing for the reader.","section":"Section 3"},{"comment":"References 16 and 19 are identical (Woods et al. 2022); one should be removed or replaced with a different source.","section":"References"},{"comment":"The inclusion criterion 'absence of ≥ grade 1 HT or long-term anemia before radiotherapy' is ambiguous; as written it would exclude all patients with any grade of HT, which contradicts the intent of excluding only pre-existing toxicity. Please rephrase.","section":"Appendix A"},{"comment":"Table C1 has corrupted column headers (e.g., 'MAX DOSE (GY)', 'MIN DOSE TARGET (GY)') and the layout is not readable; please reformat the table.","section":"Appendix C, Table C1"},{"comment":"The appendix states that 100 radiomic features were selected, but the main text does not mention this number; the number of selected features should be given in Methods for completeness.","section":"§2.4 / Appendix B"},{"comment":"The phrase 'targeting OARs' should be 'target OARs' or simply 'OARs' for grammatical correctness.","section":"§3.4"}],"recommendation":"reject","confidential_remarks":"The dosimetric comparison is well executed and could constitute a standalone planning study, but the manuscript's framing around a 23% reduction in predicted HT risk is not supported by the full uncertainty analysis. The abstract's reliance on the direct point estimate while omitting the full-pipeline bootstrap that crosses 1.0 is a reporting problem that should be weighed carefully. The extrapolation issue is acknowledged by the authors but cannot be fixed with the present data. I do not recommend rejecting the dosimetric content per se, but as submitted the central clinical claim is not defensible."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The dosimetric core is solid and worth paying attention to; the headline clinical claim is not supported by the paper's own uncertainty analysis. What is new: applying 4π non-coplanar optimization to cervical cancer, with fully integrated beam-angle and fluence optimization, and showing 28–52% reductions in pelvic bone marrow V10–V40 compared with standard VMAT. That is a much larger marrow sparing than the <10–15% typical of coplanar bone-marrow-sparing trials, and the paired statistics in Table F1 are convincing. Using a radiomics/clinical/dosimetric model to score the plans is a reasonable way to turn dosimetry into a patient-level endpoint, and the discussion is transparent about model limitations.\n\nThe soft spots are in the clinical claim and in the abstract's phrasing. Section 3.5 reports three risk-ratio estimates. The full-pipeline bootstrap, which is the only one that accounts for model refitting, gives 0.80 with 95% CI 0.57–1.03. That interval crosses 1.0. The abstract instead reports the direct point estimate 0.77 (0.75–0.79) and says the reduction is 23%, omitting the full-pipeline result. The paper says “all estimates consistently indicate a reduced HT risk,” which is misleading: the full-pipeline estimate is consistent with no reduction. Also, the fixed-model bootstrap interval (0.87–0.90) does not contain the point estimate 0.77, so the precision reported in the abstract is not reproducible from the described procedure. On top of that, the model was trained on VMAT dose distributions and applied to 4π doses outside the training range; the authors acknowledge this and speculate that it underestimates benefit, but that could go either way. Finally, the abstract says 4π “significantly reduc[es] dose to other pelvic organs at risk,” but Table F2 shows rectum V40, V45, V50 and D1%/D2% are significantly worse in 4π (e.g., V50Gy 0.08 vs 0.16). Target coverage is maintained by normalization, but “without compromising sparing of other OARs” is false as written.\n\nNone of this kills the dosimetric finding. The manuscript deserves a serious referee as a planning and feasibility study, not as evidence of clinical benefit. I would send it out, but request that the abstract and conclusion present the 23% figure as a model-based hypothesis with the full-pipeline uncertainty, and correct the OAR claim. The clinical endpoint needs external or prospective validation before anyone quotes 23%.","headline":"The dosimetric core is real; the headline 23% clinical risk-reduction claim is not supported by the paper's own full-pipeline bootstrap.","tokens_in":15115,"tokens_out":3776,"would_cite":true,"duration_ms":37009,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Replacing coplanar VMAT with 4π non-coplanar planning is predicted to reduce acute hematologic toxicity risk by about 23% in cervical cancer radiotherapy, based on a radiomics-based model applied to replanned dose distributions.","keywords":["cervical cancer","hematologic toxicity","4π non-coplanar radiotherapy","radiomics","machine learning","bone marrow sparing","VMAT","risk prediction"],"falsifier":"A prospective randomized trial in which cervical cancer patients are assigned to VMAT or 4π planning and grade≥2 hematologic toxicity is scored from blood counts would settle it: the model predicts a 23% relative reduction, so an observed reduction clearly outside the 95% confidence interval, or none at all, would falsify the central claim.","tokens_in":14018,"feed_emoji":"🩸","tokens_out":6438,"duration_ms":60256,"temperature":0.7,"pith_summary":"Standard coplanar radiotherapy for cervical cancer irradiates the pelvic bones, which contain much of the body's blood-forming marrow, leading to acute hematologic toxicity. The paper replaces standard coplanar VMAT plans with 4π non-coplanar plans for 114 retrospective cervical cancer patients and measures the dosimetric improvement. It then trains an elastic-net logistic regression model combining CT radiomic, clinical, and dosimetric features to predict grade≥2 acute hematologic toxicity, reaching an AUC of 0.79. Applying this model to the replanned dose distributions predicts a 23% relative reduction in hematologic toxicity risk with 4π planning compared with VMAT (risk ratio 0.77, 95% CI 0.75–0.79). If this prediction holds, 4π planning would be a practical way to reduce a common, dose-limiting side effect of cervical cancer radiotherapy using only standard CT data.","feed_headline":"4π planning cuts predicted blood toxicity by 23%","feed_subtitle":"Non-coplanar replanning lowered pelvic bone dose and predicted grade 2+ blood toxicity by nearly a quarter.","key_machinery":"Two machines carry the argument. The first is the 4π non-coplanar planning engine, an ultra-high-performance parallel (UHPP) optimizer that starts from 1,162 candidate beams spaced 6° apart, prunes to roughly 400 collision-free directions, and splits the cervical PTV into four equal-volume sub-cuboids with separate isocenters so the 20 cm field limit is respected. The second is the toxicity predictor: an elastic-net logistic regression that takes about 100 selected CT radiomic features plus clinical and dosimetric features and outputs a probability of grade≥2 acute hematologic toxicity, trained and evaluated with 5-fold cross-validation. The dose-volume metrics V10Gy through V40Gy in pelvic bone regions are the bridge: they are strongly reduced by 4π and are the features the risk model leans on.","core_discovery":"On its own terms, the paper establishes two linked findings. First, 4π non-coplanar optimization reduces pelvic bone marrow dose-volume metrics substantially relative to coplanar VMAT: normalized total pelvic bone V10Gy, V20Gy, V30Gy, and V40Gy fell from 93%/75%/45%/24% to 67%/36%/24%/16%, with comparable PTV coverage and improved or unchanged OAR doses. Second, a predictive model that combines CT radiomic features with clinical and dosimetric features (AUC 0.79) estimates that these dosimetric gains translate into a 23% reduction in predicted acute grade≥2 hematologic toxicity risk, with a risk ratio of 0.77 (95% CI 0.75–0.79) and odds ratio 0.68 (95% CI 0.65–0.71). The authors explicitly note the model had to extrapolate beyond the VMAT dose range and likely underestimated the clinical benefit because of this extrapolation.","pith_inferences":["Read strictly, the full-pipeline bootstrap interval (RR 0.80, 95% CI 0.57–1.03) means the 23% point estimate is not statistically robust; my read is that the direction of benefit is credible but the magnitude is not yet pinned down.","Given a monotone dose-response, the 23% figure is likely a floor, not a ceiling, because the model was fit on higher-dose VMAT plans and would under-predict the benefit of the much lower 4π doses.","The same two-step recipe—train a radiomics-based toxicity model on historical plans, then score candidate plan geometries—could be applied to other pelvic toxicities or to plan-class library selection in adaptive radiotherapy.","A prospective randomized comparison is the natural next test: select high-risk patients with the model, treat half with 4π and half with VMAT, and compare observed grade≥2 HT rates; that would also reveal whether the model's extrapolation holds outside the VMAT dose range."],"forward_implications":["Bone-marrow dose reductions of 28–52% across V10Gy–V40Gy are several times larger than the <10–15% reductions reported in coplanar bone-marrow-sparing trials, so 4π appears to open a dose regime that conventional coplanar planning cannot reach.","If the predicted 23% relative risk reduction is real, using 4π for cervical cancer would meaningfully lower rates of grade≥2 leukopenia, neutropenia, and related treatment interruptions.","The prediction model can serve as a triage tool: patients predicted to be high-risk under VMAT are offered 4π, while low-risk patients stay on the faster standard VMAT.","Because the model needs only the routine planning CT and dose-volume data, the risk-assessment workflow requires no functional imaging or extra invasive tests."],"supporting_citations":[{"why":"Supplies the UHPP 4π non-coplanar optimization engine used to replan all 114 patients.","marker":"[20]"},{"why":"Prior radiomics-based HT prediction models for cervical cancer; the paper benchmarks its AUC 0.79 against them and adopts their feature paradigm.","marker":"[30, 31]"},{"why":"Meta-analysis of pelvic bone-marrow-sparing trials; provides the <10–15% dose-reduction baseline and the G2+ HT odds-ratio comparison used to contextualize 4π gains.","marker":"[41]"},{"why":"Systematic review establishing the correlation between bone-marrow radiation dose and hematologic toxicity in cervical cancer, the dose-response premise.","marker":"[3]"},{"why":"Identifies the 40 Gy bone-marrow volume threshold associated with HT, a specific dose constraint motivating the V40Gy analysis.","marker":"[23]"},{"why":"Dosimetric predictors of acute HT in cervical cancer (V10Gy/V20Gy), grounding the choice of dose-volume features.","marker":"[38]"},{"why":"NTCP modeling of acute HT in cervical cancer, supporting the dose-response curve behind the risk model.","marker":"[39]"},{"why":"Evidence that CT-visible bone mineral inversely tracks bone-marrow activity, the biological rationale for using CT radiomics as a marrow-reserve proxy.","marker":"[32]"}],"fun_headline_variants":["4π planning reduces marrow dose, predicted blood toxicity 23%","Non-coplanar 4π cuts predicted blood toxicity by 23%","4π radiotherapy lowers pelvic bone dose, predicted toxicity 23%","4π replanning shows 23% lower predicted blood toxicity risk"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central bet is that the risk model, built from past VMAT plans, correctly predicts toxicity for 4π plans whose bone-marrow doses are much lower than anything the model was trained on.","fun_headline_variants_meta":{"raw":{"variants":["4π planning reduces marrow dose, predicted blood toxicity 23%","Non-coplanar 4π cuts predicted blood toxicity by 23%","4π radiotherapy lowers pelvic bone dose, predicted toxicity 23%","4π replanning shows 23% lower predicted blood toxicity risk"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000998,"raw_usage":{"total_tokens":4312,"prompt_tokens":1116,"completion_tokens":3196,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":732,"completion_tokens_details":{"reasoning_tokens":3127}},"tokens_in":732,"tokens_out":3196,"duration_ms":22408,"temperature":1.0,"reasoning_tokens":3127,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T00:02:06.682433+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A prospective randomized trial in which cervical cancer patients are assigned to VMAT or 4π planning and grade≥2 hematologic toxicity is scored from blood counts would settle it: the model predicts a 23% relative reduction, so an observed reduction clearly outside the 95% confidence interval, or none at all, would falsify the central claim.","supporting_citations":[],"review_version":1}