REVIEW 4 major objections 6 minor 25 references
A Beam's Eye View to Fluence Maps 3D Network for Ultra Fast VMAT Radiotherapy Planning
T0 review · 4 major / 6 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read A single 3D neural network can reconstruct all 180 fluence maps of a one-arc VMAT plan directly from a 3D dose map in under 20 milliseconds, with dose-volume histograms close to those of the input plan.
desk verdict Good architecture study for fast VMAT fluence prediction, with a real dataset-scaling contribution, but the dosimetric validation skips the deliverability step, so the 'ultra-fast planning' claim is not yet proven. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing machinery is the beam's-eye-view (BEV) transform: the 3D dose is projected into the coordinate frame of each of the 180 control points, producing 180 2D dose projections that are stacked into a 3D tensor whose depth axis is the gantry angle. A 3D MedNeXt encoder-decoder with ConvNeXt residual blocks and circular padding over the gantry dimension processes this tensor and regresses all 180 fluence maps at once. The fluence-map targets are computed from Eclipse-optimized MLC positions and monitor units with a forward model that includes leaf leakage and inter-control-point motion but omits tongue-and-groove, and the training loss combines L1 and L2 terms. The depth-direction convolutions let the network exploit translation equivariance along the rotation axis, while the ConvNeXt blocks supply the long-range mixing needed because the total dose is a sum over all control points.
What would settle it
Recompute the dose delivered by the predicted fluence maps through a full leaf-sequencing and delivery simulation that includes tongue-and-groove and real MLC speed and acceleration constraints, then compare per-voxel dose and DVHs with the original Eclipse plan on an independent validation set; a clinically meaningful disagreement would falsify the claim that the predicted maps are directly deliverable and dose-equivalent.
Extended reading notes
Core claim
The paper's central claim is that the inverse mapping from a planned 3D dose to the full set of per-control-point fluence maps of a single-arc VMAT plan can be learned and executed in one forward pass. The evidence is a 3D MedNeXt network that ingests 180 beam's-eye-view dose projections stacked along the gantry-angle dimension and jointly regresses the 180 corresponding fluence maps. On the validation split of 2,266 Eclipse-generated prostate plans the method reaches 28.71 dB PSNR and 0.9712 SSIM, and dose recomputed from the predicted fluence maps with the Acuros dose engine produces DVHs that closely track the target plan's DVHs. The authors take this as support for using the network as a module in ultra-fast inverse planning or as an initialization for an iterative VMAT optimizer.
Load-bearing premise
The load-bearing premise is that the fluence maps computed from the Eclipse-optimized MLC positions and monitor units, via a forward model that includes leaf leakage and inter-control-point motion but ignores tongue-and-groove, together with the Eclipse/Acuros dose engine, accurately represent a deliverable VMAT plan and that the input 3D dose map is exactly the dose produced by those same fluence maps.
Editorial extensions
If this is right
- One forward pass, under 20 ms excluding data loading, produces all 180 fluence maps of a single-arc VMAT plan from a dose map, so inverse planning no longer needs an iterative leaf-sequencing loop for these cases.
- Because the training targets encode MLC constraints from real plans, the predicted maps should be rapidly leaf-sequenceable, making the network a plug-in module in an end-to-end planning pipeline.
- Dataset size is a first-order driver: tripling the training set from about 100 to 400 plans adds 2.1 dB PSNR, and going from 400 to about 1,900 adds another 2.7 dB, so collecting more clinically planned cases is a direct route to better predictions.
- The same network can initialize an established VMAT optimizer such as Varian's Photon Optimization algorithm, potentially reducing the iterations needed to reach a clinical plan.
Reading between the lines
- The dose-to-fluence inverse is likely non-unique, so DVH closeness does not by itself prove the predicted MLC sequence is the deliverable sequence the Eclipse plan would execute; a delivery-simulation or QA test would settle machine compatibility.
- The BEV-stacked representation turns gantry rotation into translation along the depth axis, a transformation that should transfer to multi-arc VMAT, IMRT, or other rotating delivery modalities if training data with the same control-point structure are available.
- The reported lack of benefit from adding CT or contours as extra input channels suggests the 3D dose map alone carries most of the planning information this network can use, at least for prostate cases and this dataset size.
- Because all plans are Varian single-arc prostate, the model's machine specificity is an open question: testing on other linac MLC models, collimator angles, or anatomical sites would reveal how much retraining is needed.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a deep learning method to predict all 180 fluence maps of a single-arc VMAT plan from the 3D dose distribution in one network pass. The dose map is preprocessed into beam's eye view projections at each control point and fed to a 3D MedNeXt architecture. The authors generate additional VMAT plans using Eclipse to enlarge the dataset from 117 to 2266 prostate plans, train with L1+L2 losses, and report PSNR/SSIM improvements over 2D and 3D U-Nets. They also claim that DVHs calculated from the predicted fluence maps are very close to the target DVHs. The central claims are the architectural contribution (BEV transform + 3D convolutions) and dataset scaling, with a downstream goal of ultra-fast VMAT planning.
Significance. If the dosimetric validation were rigorous, the work would be a useful contribution to automated VMAT planning by showing that fluence maps can be predicted in a single fast forward pass and that dataset size is a key driver of accuracy. The paper provides a clear comparison against 2D and 3D U-Net baselines and consistently reports PSNR gains; the large-scale Eclipse-generated dataset is a practical resource. However, the current manuscript does not provide quantitative DVH metrics or a deliverability assessment, so the clinical significance of the reported gains is not established.
major comments (4)
- [II.4 / Results] The manuscript states that MAE in Gy and DVHs were computed for the best method, yet the Results section reports no MAE value and no quantitative DVH statistics (e.g., D95, V95, mean or max dose differences, gamma pass rates). The only evidence for the central claim that 'the resulting DVHs are very close' is four qualitative examples in Fig. 4 and a statement that all DVHs looked similar. This does not substantiate the claim. Please report the numerical MAE/Gy and DVH metrics across the validation set, ideally with confidence intervals.
- [II.4 / Fig. 4] The dose calculation from the predicted fluence maps is not described. Acuros AXB takes MLC leaf positions and MU weights as input, not arbitrary 2D fluence maps; therefore a leaf-sequencing or aperture conversion step is required. The manuscript neither specifies the conversion algorithm nor states whether deliverability constraints (leaf speed, tongue-and-groove, interdigitation) were respected. If the DVHs were computed directly from idealized fluence maps, they bypass the very constraints that define VMAT delivery, and the comparison is not clinically meaningful. This is a load-bearing gap for the ultra-fast planning claim.
- [II.1] The training targets are fluence maps computed with a forward model that includes leaf leakage and inter-control-point motion but ignores tongue-and-groove, while the input dose is computed by Eclipse/Acuros from the MLC positions. The manuscript does not verify that the simplified fluence-to-dose forward model is consistent with the Eclipse/Acuros dose used as input. A large discrepancy would make the supervised task ill-posed (the network would be asked to map a dose that cannot be exactly reproduced by the target fluence representation). Please report a gamma analysis between the dose computed from the target fluence maps and the original Eclipse dose for the training targets.
- [II.4 / Tables 2-3] The results are reported as single point estimates with no error bars, no multi-seed runs, and no statistical significance tests. This makes it impossible to judge whether the smaller differences (e.g., 3D MedNeXt trained on Ecl. 500 vs Ecl. full: 26.01 vs 28.72 dB on Ecl. 500) are meaningful. Additionally, the SSIM values for the 3D U-Net (e.g., 0.2664 in Table 3) appear anomalously low relative to its PSNR (~22 dB) compared with the 2D U-Net (SSIM ~0.89 at similar PSNR), which suggests a possible issue in the SSIM computation or in the 3D U-Net training; please verify.
minor comments (6)
- [II.1] The number of generated plans is given as 'around 2266' in the text and 2266 in Table 1; please use exact numbers consistently.
- [II] The statement 'we are using for the first time a 3D convolutional network architecture' conflicts with the later mention that previous studies (refs 10, 11) used a 3D U-Net to predict fluence maps. Please clarify what is new relative to those works.
- [III] There are several typographical errors: 'the propose 3D network', 'fund that they all looked similar', and 'litterature' should be corrected.
- [II.4] The training details lack the number of epochs, learning rate schedule, and data augmentation, which are needed for reproducibility.
- [Fig. 4] The DVH figure is not labeled with the structures; it is unclear which curves correspond to PTV or OARs. Please add a legend or caption explanation.
- [III] The inference time of 20 ms is reported without specifying the GPU/CPU hardware, input resolution, and batch size; please provide these for reproducibility.
Circularity Check
No significant circularity: the supervised dose-to-fluence inversion is self-consistent and validated against an external dose engine; self-citations are not load-bearing.
full rationale
The paper's derivation chain is a supervised inverse mapping. Section II.1 states the dataset contains the optimized MLC positions and MU values of each control point 'as well as the 3D dose map calculated from these optimized MLC positions and MU values.' Section II.2 computes the training targets by converting MLC positions and MU values into fluence maps using a stated forward model that includes leaf leakage and inter-control-point motion. The network is trained with L1/L2 losses on those fluence maps, and Section II.4 validates on held-out plans by PSNR/SSIM and by recomputing dose with the external Acuros AXB engine for DVH comparison against the input target dose. This is a standard reconstruction loop, not a circular one: the target dose is used as network input but not as a training loss; the DVH metric is not optimized and is computed with an engine external to the network; and the validation set is disjoint from the training set. The reported closeness of DVHs therefore reflects genuine generalization of the learned inverse mapping rather than a quantity forced by construction. The self-citations (refs 6, 17, 22) support companion modules for dose prediction, Eclipse plan generation, and leaf sequencing; none is invoked as a uniqueness theorem or as a mathematical constraint that would make the fluence prediction equivalent to its input. The omitted leaf-sequencing step in the DVH validation is a completeness and deliverability limitation, not a circularity.
Assumptions & free parameters
free parameters (2)
- L1/L2 loss weighting =
not reported
- Network hyperparameters (channel widths, depth, kernel specifics) =
not reported beyond reference to MedNeXt
assumptions (4)
- domain assumption Fluence maps computed from MLC positions and MU values, including leaf leakage and inter-control-point motion but ignoring tongue-and-groove, accurately represent deliverable VMAT radiation.
- domain assumption BEV projections of the 3D dose map preserve enough information for the network to invert the dose calculation.
- domain assumption The Eclipse-generated replans of REQUITE prostate cases are clinically realistic and representative of the target domain.
- domain assumption The inverse mapping from dose to fluence is well-posed enough for a single feedforward network to learn it across 180 control points.
Cite this review
Pith. "Pith review of A Beam's Eye View to Fluence Maps 3D Network for Ultra Fast VMAT Radiotherapy Planning." pith.science (2026). https://pith.science/paper/72TL7NAD
@misc{pith2026250203360,
author = {Pith},
title = {Pith review of: A Beam's Eye View to Fluence Maps 3D Network for Ultra Fast VMAT Radiotherapy Planning},
year = {2026},
howpublished = {\url{https://pith.science/paper/72TL7NAD}},
note = {Machine review of arXiv:2502.03360}
}
read the original abstract
Volumetric Modulated Arc Therapy (VMAT) revolutionizes cancer treatment by precisely delivering radiation while sparing healthy tissues. Fluence maps generation, crucial in VMAT planning, traditionally involves complex and iterative, and thus time consuming processes. These fluence maps are subsequently leveraged for leaf-sequence. The deep-learning approach presented in this article aims to expedite this by directly predicting fluence maps from patient data. We developed a 3D network which we trained in a supervised way using a combination of L1 and L2 losses, and RT plans generated by Eclipse and from the REQUITE dataset, taking the RT dose map as input and the fluence maps computed from the corresponding RT plans as target. Our network predicts jointly the 180 fluence maps corresponding to the 180 control points (CP) of single arc VMAT plans. In order to help the network, we pre-process the input dose by computing the projections of the 3D dose map to the beam's eye view (BEV) of the 180 CPs, in the same coordinate system as the fluence maps. We generated over 2000 VMAT plans using Eclipse to scale up the dataset size. Additionally, we evaluated various network architectures and analyzed the impact of increasing the dataset size. We are measuring the performance in the 2D fluence maps domain using image metrics (PSNR, SSIM), as well as in the 3D dose domain using the dose-volume histogram (DVH) on a validation dataset. The network inference, which does not include the data loading and processing, is less than 20ms. Using our proposed 3D network architecture as well as increasing the dataset size using Eclipse improved the fluence map reconstruction performance by approximately 8 dB in PSNR compared to a U-Net architecture trained on the original REQUITE dataset. The resulting DVHs are very close to the one of the input target dose.
Figures
Reference graph
Works this paper leans on
-
[1]
K. Otto, Volumetric modulated arc therapy: IMRT in a single gantry arc, Medical physics 35 , 310--317 (2008)
work page 2008
-
[2]
H. Liu, B. Sintay, K. Pearman, Q. Shang, L. Hayes, J. Maurer, C. Vanderstraeten, and D. Wiant, Comparison of the progressive resolution optimizer and photon optimizer in VMAT optimization for stereotactic treatments, Journal of applied clinical medical physics 19 , 155--162 (2018)
work page 2018
-
[3]
C. McIntosh, M. Welch, A. McNiven, D. A. Jaffray, and T. G. Purdie, Fully automated treatment planning for head and neck radiotherapy using a voxel-based dose prediction and dose mimicking method, Physics in Medicine & Biology 62 , 5926 (2017)
work page 2017
-
[4]
V. Kearney, J. W. Chan, S. Haaf, M. Descovich, and T. D. Solberg, DoseNet: a volumetric dose prediction algorithm using 3D fully-convolutional neural networks, Physics in Medicine & Biology 63 , 235022 (2018)
work page 2018
-
[5]
B. Zhan, J. Xiao, C. Cao, X. Peng, C. Zu, J. Zhou, and Y. Wang, Multi-constraint generative adversarial network for dose prediction in radiotherapy, Medical Image Analysis 77 , 102339 (2022)
work page 2022
-
[6]
R. Gao, B. Lou, Z. Xu, D. Comaniciu, and A. Kamen, Flexible-Cm GAN: Towards Precise 3D Dose Prediction in Radiotherapy, in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 715--725, 2023
work page 2023
-
[7]
W. Wang et al., Fluence map prediction using deep learning models--direct plan generation for pancreas stereotactic body radiation therapy, Frontiers in artificial intelligence 3 , 68 (2020)
work page 2020
-
[8]
H. Lee, H. Kim, J. Kwak, Y. S. Kim, S. W. Lee, S. Cho, and B. Cho, Fluence-map generation for prostate intensity-modulated radiotherapy planning using a deep-neural-network, Scientific reports 9 , 15671 (2019)
work page 2019
Show all 25 references
-
[9]
W. Wang, Y. Sheng, M. Palta, B. Czito, C. Willett, M. Hito, F.-F. Yin, Q. Wu, Y. Ge, and Q. J. Wu, Deep learning--based fluence map prediction for pancreas stereotactic body radiation therapy with simultaneous integrated boost, Advances in radiation oncology 6 , 100672 (2021)
2021
-
[10]
L. Ma, M. Chen, X. Gu, and W. Lu, Deep learning-based inverse mapping for fluence map prediction, Physics in Medicine & Biology 65 , 235035 (2020)
2020
-
[11]
L. Ma, M. Chen, X. Gu, and W. Lu, Generalizability of deep learning based fluence map prediction as an inverse planning approach, arXiv preprint arXiv:2104.15032 (2021)
2021 arXiv
-
[12]
S. Zhu, A. Maslowski, J. M. Cunningham, E. Kuusela, and I. J. Chetty, 3D Dose-Driven, Automatic VMAT Machine Parameter Generation with Deep Learning, International Journal of Radiation Oncology, Biology, Physics 114 , e581--e582 (2022)
2022
-
[13]
Z. Liu, H. Mao, C.-Y. Wu, C. Feichtenhofer, T. Darrell, and S. Xie, A convnet for the 2020s, in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages 11976--11986, 2022
2022
-
[14]
C i c ek, A
\"O . C i c ek, A. Abdulkadir, S. S. Lienkamp, T. Brox, and O. Ronneberger, 3D U-Net: learning dense volumetric segmentation from sparse annotation, in Medical Image Computing and Computer-Assisted Intervention--MICCAI 2016: 19th International Conference, Athens, Greece, Octob...
2016
-
[15]
K. He, X. Zhang, S. Ren, and J. Sun, Deep residual learning for image recognition, in Proceedings of the IEEE conference on computer vision and pattern recognition , pages 770--778, 2016
2016
-
[16]
P. Seibold et al., REQUITE: a prospective multicentre cohort study of patients undergoing radiotherapy for breast, lung or prostate cancer, Radiotherapy and Oncology 138 , 59--67 (2019)
2019
-
[17]
Gao et al., Automating High Quality RT Planning at Scale, arXiv preprint arXiv:2501.11803 (2025)
R. Gao et al., Automating High Quality RT Planning at Scale, arXiv preprint arXiv:2501.11803 (2025)
2025 arXiv
-
[18]
S. Roy, G. Koehler, C. Ulrich, M. Baumgartner, J. Petersen, F. Isensee, P. F. Jaeger, and K. H. Maier-Hein, Mednext: transformer-driven scaling of convnets for medical image segmentation, in International Conference on Medical Image Computing and Computer-Assisted Intervention...
2023
-
[19]
Dosovitskiy et al., An image is worth 16x16 words: Transformers for image recognition at scale, arXiv preprint arXiv:2010.11929 (2020)
A. Dosovitskiy et al., An image is worth 16x16 words: Transformers for image recognition at scale, arXiv preprint arXiv:2010.11929 (2020)
2020 arXiv
-
[20]
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, Swin transformer: Hierarchical vision transformer using shifted windows, in Proceedings of the IEEE/CVF international conference on computer vision , pages 10012--10022, 2021
2021
-
[21]
Ronneberger, P
O. Ronneberger, P. Fischer, and T. Brox, U-net: Convolutional networks for biomedical image segmentation, in Medical image computing and computer-assisted intervention--MICCAI 2015: 18th international conference, Munich, Germany, October 5-9, 2015, proceedings, part III 18 , p...
2015
-
[22]
Gao, F.-C
R. Gao, F.-C. Ghesu, S. Arberet, S. Basiri, E. Kuusela, M. Kraus, D. Comaniciu, and A. Kamen, Multi-Agent Reinforcement Learning Meets Leaf Sequencing in Radiotherapy, in International Conference on Machine Learning , 2024
2024
-
[23]
V. M. Systems, Eclipse photon and electron algorithms reference guide, 2015
2015
-
[24]
, " * write output.state after.block =
") INTEGERS output.state before.all mid.sentence after.sentence after.block FUNCTION init.state.consts #0 'before.all := #1 'mid.sentence := #2 'after.sentence := #3 'after.block := STRINGS s t FUNCTION output.nonnull 's := output.state mid.sentence = ", " * write output.state...
-
[25]
write newline
" write newline "" before.all 'output.state := FUNCTION fin.entry add.period write newline FUNCTION new.block output.state before.all = 'skip after.block 'output.state := if FUNCTION new.sentence skip FUNCTION not #0 #1 if FUNCTION and 'skip pop #0 if FUNCTION or pop #1 'skip ...
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.