{"id":"bbf0293d-efbf-4d62-88d7-94a819dd021b","arxiv_id":"2504.16805","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In ten brain VMAT plans, the default Monaco CT-to-RED curve gave systematically higher doses than the CIRS phantom curve, with 3 to 5% target and 2 to 4% organ-at-risk differences.","lead":"This study recalculated ten brain VMAT plans with three CT-to-electron density calibration curves and found dose differences of 3 to 5% for targets and 2 to 4% for organs at risk. It concludes that the CIRS phantom-derived curve is most accurate and should replace the default treatment planning system curve.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Accuracy claim rests on unverified ground truth: relative dose differences show sensitivity, not which curve is correct.","rationale":"The reader's weakest_assumption correctly identifies the absence of an independent ground-truth dose measurement as the key gap. I agree with that assessment and with the CONDITIONAL verdict. The paper is a useful sensitivity analysis showing that calibration-curve choice materially changes DVH metrics in brain VMAT, and the statistical methods appear reasonable for comparing three curves. However, the headline 'most accurate' claim cannot be supported by pairwise p-values alone. Because the reader already conditioned acceptance on softer claims or independent verification, my stress-test does not move the verdict; it reinforces the same condition. I do not see an additional internal inconsistency: the recalculations kept beam parameters and MUs constant, the repeated-measures/paired analyses are appropriate for the ten-plan cohort, and the reported dose differences are plausible for HU-to-RED mapping changes. The only path to a definitive accuracy ranking is direct measurement or an equivalent reference-standard calculation. Until then, the paper should be read as demonstrating sensitivity, not establishing which calibration curve is correct.","tokens_in":7173,"tokens_out":3176,"duration_ms":35172,"concrete_test":"Deliver five of the ten brain VMAT plans to an anthropomorphic head phantom with known tissue-equivalent composition, placing a micro-ionization chamber or radiochromic film at PTV and OAR locations. Acquire the planning CT with the same Siemens scanner and 120 kVp protocol, then recalculate the plans with the CIRS, Catphan, and default Monaco curves. Compare each calculation against the measured point doses; the curve with the smallest mean absolute error is the most accurate. If the default curve's error is within 1% of the CIRS curve's error, or if CIRS is not the closest to measurement, the headline claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central conclusion is that the CIRS-derived curve 'provided the most accurate dosimetric estimates and should be prioritized in clinical calibration protocols.' What the data actually establish is that the three curves produce statistically distinguishable DVH metrics: the default curve yields higher PTV and OAR doses than CIRS by 3-5% and 2-4%, respectively, with Catphan intermediate (Sections 3.1-3.3, Table 1). These are relative comparisons. No delivered dose was measured with an independent dosimeter, no reference calculation with known ground-truth densities was performed, and no absolute dose accuracy metric is reported. The assertion that CIRS is 'most accurate' is inferred from the phantom having ten tissue-equivalent inserts (Section 2.2, Discussion), not from agreement between calculated and measured dose. That inference is not logically forced: the observed offset could equally mean the default curve overestimates, the CIRS curve underestimates, or both are wrong in opposite directions. The statistical tests only confirm that the curves differ; they cannot rank the curves by physical accuracy. The conclusion therefore exceeds the evidence. This is the load-bearing weak point because the clinical recommendation and the 'overestimation' framing both depend on knowing which curve is nearer to true dose.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a dosimetric comparison of three CT-to-RED calibration curves (CIRS 062M, Catphan 604, and the Monaco TPS default) applied to ten retrospective brain VMAT plans recalculated in the Monaco treatment planning system. PTV metrics (Dmax, Dmin, Dmean, D95%) and OAR Dmax values were extracted from DVHs and compared using repeated-measures ANOVA/Friedman tests with post-hoc pairwise comparisons. The authors find statistically significant differences between the CIRS-derived and default TPS curves, with the default curve giving higher doses by 3–5% for PTV and 2–4% for OAR Dmax, and conclude that the CIRS-derived curve is the most accurate and should be prioritized in clinical calibration protocols.","tokens_in":7407,"tokens_out":4277,"duration_ms":39936,"significance":"If the reported differences are reliable, the study provides a useful sensitivity analysis showing that calibration curve choice can shift DVH metrics by several percent in brain VMAT, which is clinically relevant against the usual 5% dose tolerance. The use of ten clinical plans and the direct comparison of three established calibration approaches are strengths. However, the central accuracy ranking is not supported by the data as presented: no independent dosimetric measurement or ground-truth reference is used to establish which curve is closer to true delivered dose. The value of the paper lies in quantifying inter-curve variability and its potential clinical implications, not in establishing which calibration curve is physically correct in an absolute sense.","major_comments":[{"comment":"The claim that the CIRS-derived curve \"provided the most accurate dosimetric estimates\" is not supported by the measurements reported. The study compares three calculation pathways with each other; no delivered dose was measured with an independent dosimeter, no reference calculation with known ground-truth densities was performed, and no absolute dose accuracy metric is reported. The observed 3–5% PTV and 2–4% OAR differences are relative differences and could equally indicate that the default curve overestimates, that the CIRS curve underestimates, or that both deviate from true dose in opposite directions. Please either add a direct validation experiment (e.g., phantom measurements with an ionization chamber or film) or reframe the conclusion to state that the curves produce statistically and clinically significant differences, with the default curve giving systematically higher dose estimates than the CIRS curve.","section":"Section 4 and Conclusion"},{"comment":"The reasoning that the CIRS curve is more accurate because the CIRS 062M phantom contains ten tissue-equivalent inserts is an assertion about phantom design, not a measurement of dosimetric accuracy. A larger number of inserts affects the breadth of the fitted calibration curve, but no evidence is presented connecting insert count to agreement with true delivered dose in these patient plans. This inference should be removed or supported by a validation measurement; otherwise the paper should treat CIRS as one of three candidate curves rather than as the reference standard.","section":"Section 2.2 and Discussion"},{"comment":"The manuscript does not report actual dose values or confidence intervals for the comparisons. Table 1 reports p-values and percentages only, and Figure 2 shows \"relative mean dose percent\" without specifying the reference value or the absolute dose scale. To assess clinical significance and to allow reproduction, please report the mean and standard deviation (or median and interquartile range) in dose units for each structure and each curve, together with pairwise differences and 95% confidence intervals.","section":"Table 1 and Figure 2"},{"comment":"The multiple-testing strategy is incompletely described. Fourteen dosimetric parameters are tested, each with an omnibus test and pairwise comparisons, yet the Bonferroni correction appears to be applied only within each parameter's post-hoc tests, not across the full set of parameters. Uncorrected overall p-values such as 0.0035–0.0057 are reported alongside corrected pairwise values. Please specify the total number of comparisons, the correction method applied to the omnibus tests, and whether the p<0.05 threshold is family-wise or per-comparison.","section":"Section 2.5 and Table 1"}],"minor_comments":[{"comment":"The sentence \"by 3 5 for target volumes and 2 4 for some critical organs\" is missing percent signs; it should read \"3–5%\" and \"2–4%\".","section":"Abstract"},{"comment":"The text states that the greatest discrepancy is in Dmax with p<0.001, but Table 1 does not report the magnitude of that discrepancy; please include the corresponding percentage difference at the point of citation.","section":"Section 3.1"},{"comment":"The heading \"3.4. Visual Representation of Dosimetric Differences\" is immediately followed by a duplicate heading \"3.4. Visual Representation of Dosimetric Impact\"; one of the two should be removed.","section":"Section 3.4"},{"comment":"The section labelled \"4. Conclusion\" should be numbered \"5. Conclusion\" because the Discussion is already Section 4.","section":"Section 4"},{"comment":"The caption says \"red for CIRS, black for Catphan, and blue for Default\"; please verify that the printed figure uses these exact colors and that the color legend is unambiguous for readers.","section":"Figure 2 caption"},{"comment":"Reference [12] is cited in the Introduction as \"Duong et al. (2025)\" and in the Discussion as \"Thanh Tai et al. (2025)\"; please standardize the citation style for the same reference throughout the text.","section":"References"},{"comment":"The phantom name is written inconsistently as \"Catphan 604\", \"Catphan 604\", and \"CATPhan\" across the text and references; please unify the spelling.","section":"General"},{"comment":"Given that ten patient plans were used retrospectively, please include a statement on ethics approval or an institutional review board waiver, as is customary for retrospective dosimetric studies.","section":"Section 2.3"}],"recommendation":"major_revision","confidential_remarks":"The core comparison is salvageable and likely of interest to the readership, but the title, abstract, and conclusion currently claim an absolute accuracy ranking that the data cannot support. The paper would be acceptable if the authors either add a validation measurement against a dosimetric ground truth or substantially reframe the conclusions as a relative sensitivity analysis. The multiple-comparison reporting and the absence of absolute dose values should also be addressed during revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Small but genuine empirical comparison. Ten brain VMAT plans were recalculated with three CT-to-RED calibration curves, and the DVH differences are internally consistent and statistically clear. The application to clinical Monaco plans is new in the cited literature, which mostly used phantoms or CBCT. That is a legitimate contribution, modest in scope.\n\nThe soft spot is the accuracy claim. No delivered dose was measured with an independent dosimeter. The data show relative differences: the default curve yields 3-5% higher PTV and 2-4% higher OAR doses than CIRS, with Catphan in between. Those tests confirm the curves differ, not which is closer to true dose. The assertion that CIRS is most accurate because its phantom has ten tissue-equivalent inserts does not follow. The offset could equally mean CIRS underestimates. That is the load-bearing flaw, and it is fixable by softening the conclusion or adding an absolute measurement.\n\nMinor issues: Table 1 reports p-values but no actual dose values or confidence intervals; ten patients is small though acknowledged; there is a duplicated section heading. The self-citation to the earlier scanning-protocol study is fine.\n\nThis paper is for someone doing local QA or commissioning in radiotherapy. It deserves a serious referee because the data are real and the problem is well-defined. I would send it to peer review with major revision and ask for verification of the accuracy claim. I would not cite it in my own work until that is resolved.","headline":"Small but genuine clinical comparison of CT-to-RED curves; the accuracy claim outruns the data without an independent dose measurement.","tokens_in":737,"tokens_out":1333,"would_cite":false,"duration_ms":27124,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The CT-to-RED calibration curve used in brain VMAT changes calculated target doses by 3–5%, and the paper argues the default Monaco curve overestimates relative to a CIRS-derived curve.","keywords":["Brain radiotherapy","CT-to-RED calibration","Relative electron density","Treatment planning system","Dosimetric accuracy","CIRS phantom","Catphan phantom","VMAT"],"falsifier":"Place an independent dosimeter (ionization chamber or film) in a head-equivalent phantom, create identical VMAT plans using the CIRS-derived and default Monaco curves, deliver both, and compare measured to calculated dose; if measured doses match the default-curve calculation within uncertainty while CIRS calculations are 3–5% low, the paper's overestimation claim fails.","tokens_in":7011,"feed_emoji":"🧠","tokens_out":7119,"duration_ms":63390,"temperature":0.7,"pith_summary":"This paper asks whether the curve that converts CT Hounsfield Units into relative electron density changes the calculated dose in brain radiotherapy, and by how much. Recalculating ten brain VMAT plans with three different curves, it finds that the default Monaco curve gives target doses 3–5% higher and organ-at-risk maximum doses 2–4% higher than a curve derived from the CIRS 062M phantom, with the Catphan-derived curve in between. If the CIRS curve is the better representation of brain tissues, then relying on the default curve risks underdosing tumors and overdosing nearby critical structures. The study therefore argues for site-specific, phantom-based calibration for brain plans, a step that could improve accuracy and reduce variability across clinics.","feed_headline":"Default CT curve overestimates brain tumor doses by up to 5%","feed_subtitle":"Phantom-based calibration gives the most accurate brain radiation doses; the Monaco default runs 2–5% high.","key_machinery":"The central object is the CT-to-RED calibration curve, a mapping from each Hounsfield Unit value in the planning CT to a relative electron density used by the Monte Carlo dose engine. The study compares three such curves: one built from a 10-insert tissue-equivalent phantom (CIRS 062M), one from a 7-insert phantom (Catphan 604), and the vendor default preset in the Monaco treatment planning system. Recalculating the same ten VMAT plans with only this mapping changed isolates the effect of the calibration curve on the resulting dose-volume metrics.","core_discovery":"The paper reports that, in ten brain VMAT plans, recalculating with the default Monaco calibration curve produced systematically higher dose metrics than recalculating with a curve derived from the CIRS 062M phantom: PTV Dmax, Dmin, Dmean, and D95% all differed significantly (p < 0.05), with Dmax 3–5% higher, and seven of ten OARs showed significant Dmax increases of 2–4%, including the brainstem (p = 0.003) and optic chiasm (p = 0.005). The Catphan-derived curve sat between the two, with fewer significant differences. The authors conclude that the CIRS-derived curve is the most accurate and should replace generic defaults in brain calibration protocols.","pith_inferences":["Because only one CT scanner, one treatment planning system, and one beam energy were used, the 3–5% magnitude is likely scanner- and protocol-dependent; repeating the recalculation on a second scanner with a different tube voltage would test this.","A direct measurement with an independent dosimeter in a head-equivalent phantom would convert the relative comparison into an absolute accuracy claim and settle whether the default curve overestimates or the CIRS curve underestimates true dose.","The same comparison could be extended to proton therapy, where CT-derived stopping-power ratios are also sensitive to HU-to-RED calibration and where a 3–5% range uncertainty is clinically meaningful.","If the overestimation is systematic, centers using the default curve in multi-center trials may unknowingly bias dose-normalization parameters such as D95% or Dmean, contributing to inter-institutional variability in delivered doses."],"forward_implications":["Using the default Monaco curve instead of a CIRS-derived curve raises calculated PTV Dmax by 3–5% and OAR Dmax by 2–4%, so plans normalized to the same target dose would deliver less physical dose to the tumor and more to organs than intended.","Seven of ten OARs, including the brainstem and optic chiasm, show significantly higher maximum doses with the default curve, raising the predicted risk of neurotoxicity such as brainstem necrosis or optic neuropathy.","The Catphan-derived curve gives intermediate results with fewer significant differences, so it is closer to the CIRS curve than to the default but still not the preferred brain calibration.","Adopting site-specific, phantom-based calibration would reduce inter-institutional variability and help meet the clinical guideline that delivered doses should not deviate by more than 5% from prescription.","Dose-volume histogram comparisons across the three curves show visible deviations in high-gradient regions near critical structures, meaning plan evaluation and treatment approval can depend on the calibration curve chosen."],"supporting_citations":[{"why":"Establishes the clinical tolerance that a 5% dose deviation can reduce tumor control probability by about 20%, which is why the 3–5% shifts reported here matter.","marker":"[2]"},{"why":"Supplies the foundational method of an electron-density calibration phantom for CT-based treatment planning, on which the CIRS phantom curve is built.","marker":"[6]"},{"why":"Provides a prior comparison of Catphan and CIRS phantoms for kV-CBCT dose calculation, reporting larger discrepancies with Catphan and motivating phantom selection.","marker":"[9]"},{"why":"Shows that CT scanning protocol variations affect HU-to-RED calibration and dose calculation, supporting the need for site-specific calibration.","marker":"[12]"},{"why":"Demonstrates the influence of different tube voltages and phantoms on CT-number-to-RED curves, directly supporting the paper's calibration-comparison design.","marker":"[14]"},{"why":"Cited in the discussion as evidence that CIRS phantoms give superior dosimetric accuracy in brain radiotherapy.","marker":"[15]"}],"fun_headline_variants":["Default CT curve inflates brain tumor doses by up to 5%","Phantom-based CT curve most accurate for brain dosimetry","CT calibration curve choice shifts brain VMAT doses by 5%","Monaco default curve overestimates brain doses, phantom curves win","Brain radiotherapy dose errors up to 5% from CT curve selection"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the CIRS phantom's tissue-equivalent inserts truly represent brain tissue electron densities, so its calibration curve is the accuracy benchmark; no independent delivered-dose measurement confirms this.","fun_headline_variants_meta":{"raw":{"variants":["Default CT curve inflates brain tumor doses by up to 5%","Phantom-based CT curve most accurate for brain dosimetry","CT calibration curve choice shifts brain VMAT doses by 5%","Monaco default curve overestimates brain doses, phantom curves win","Brain radiotherapy dose errors up to 5% from CT curve selection"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000309,"raw_usage":{"total_tokens":1717,"prompt_tokens":854,"completion_tokens":863,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":470,"completion_tokens_details":{"reasoning_tokens":773}},"tokens_in":470,"tokens_out":863,"duration_ms":7799,"temperature":1.0,"reasoning_tokens":773,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:54:35.107809+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Place an independent dosimeter (ionization chamber or film) in a head-equivalent phantom, create identical VMAT plans using the CIRS-derived and default Monaco curves, deliver both, and compare measured to calculated dose; if measured doses match the default-curve calculation within uncertainty while CIRS calculations are 3–5% low, the paper's overestimation claim fails.","supporting_citations":[{"cited_title":"Evaluation of Inhomogeneity Correction Performed by Radiotherapy Treatment Planning System","cited_arxiv_id":null,"evidence_quote":"Establishes the clinical tolerance that a 5% dose deviation can reduce tumor control probability by about 20%, which is why the 3–5% shifts reported here matter."},{"cited_title":"An electron density calibration phantom for CT-based treatment planning computers","cited_arxiv_id":null,"evidence_quote":"Supplies the foundational method of an electron-density calibration phantom for CT-based treatment planning, on which the CIRS phantom curve is built."},{"cited_title":"Assessment of the dosimetric accuracies of CATPhan 504 and CIRS 062 using kV-CBCT for performing direct calculations","cited_arxiv_id":null,"evidence_quote":"Provides a prior comparison of Catphan and CIRS phantoms for kV-CBCT dose calculation, reporting larger discrepancies with Catphan and motivating phantom selection."},{"cited_title":"Scanning protocol influence on relative electron Density-CT number calibrations and radiotherapy dose calculation for a Halcyon Li nac","cited_arxiv_id":null,"evidence_quote":"Shows that CT scanning protocol variations affect HU-to-RED calibration and dose calculation, supporting the need for site-specific calibration."},{"cited_title":"The influence of different kVs and phantoms on computed tomography number to relative electron density calibration curve for radiotherapy dose calculation","cited_arxiv_id":null,"evidence_quote":"Demonstrates the influence of different tube voltages and phantoms on CT-number-to-RED curves, directly supporting the paper's calibration-comparison design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Cited in the discussion as evidence that CIRS phantoms give superior dosimetric accuracy in brain radiotherapy."}],"review_version":1}