REVIEW 4 major objections 6 minor 31 references
Facial Surgery Preview Based on the Orthognathic Treatment Prediction
T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A scan-only neural pipeline predicts facial shape after jaw-correction surgery more accurately than a published baseline, and blinded experts cannot reliably tell the predictions from real outcomes.
desk verdict Legitimate incremental advance in scan-only orthognathic preview, but the no-plan predictor undercuts the headline claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is FLAME, a parametric head model with nonlinear jaw articulation, used as both encoder and decoder: pre- and post-surgery scans are fitted to FLAME latent codes, and the predicted code difference is decoded back into a mesh with point-to-point correspondences. The predictor itself is a small fully connected network trained with a weighted sum of four losses: mouth-convexity loss (squared deviation beyond 3 mm from the s-line to the upper and lower lip midpoints), asymmetry loss (distances of chin vertex-pair midpoints from a least-squares mid-sagittal plane plus orientation disagreement), latent-code loss, and geometry loss (point positions plus surface normals). A second mechanism is the data-augmentation scheme: a horizontal plane splits the face into an unchanged upper region and a surgically modified lower region, a synthetic face is generated by randomizing the upper-face latent code, and this upper part is stitched to real pre- and post-surgery lower parts to create many plausible training pairs.
What would settle it
Compare the model's single prediction for a set of patients whose preoperative scans are near-identical but whose actual surgeries involved measurably different jaw movements, for example different mandibular advancement distances or different degrees of chin rotation. If the real postoperative outcomes differ from each other by more than roughly the reported Chamfer error (around 2.5 mm) while both are consistent with the same pre-surgery scan, then the scan-only assumption fails and the model can only be producing an averaged outcome.
Extended reading notes
Core claim
The paper's central claim is that postsurgical facial geometry after orthognathic treatment is predictable from the preoperative facial surface alone, once that surface is represented in a parametric model that can articulate the jaw. Concretely, the authors claim that a code-difference predictor operating in FLAME's latent space, trained with (i) a mouth-convexity loss that penalizes lip protrusion beyond a 3 mm tolerance relative to the s-line, (ii) an asymmetry loss measuring chin-point deviation from the mid-sagittal plane, (iii) a latent-code loss, and (iv) a geometry loss on points and normals, produces predictions that are more accurate than the baseline method by both Chamfer and Hausdorff distances. They further claim that their horizontal-split data-augmentation scheme, which synthesizes pre/post pairs by stitching the unchanged upper face to surgically altered lower-face regions, is a major contributor to this accuracy; removing it raises Chamfer distance by 17.46%. Finally, the user-study results are claimed as evidence that the predicted faces are perceptually close enough to real outcomes that neither doctors nor engineers can distinguish them at a statistically significant rate.
Load-bearing premise
The pipeline assumes that the post-surgery facial shape is determined by the preoperative facial surface geometry alone: the same pre-surgery scan is always mapped to one predicted outcome, with no dependence on the specific surgical plan, bone movements, or patient attributes.
Editorial extensions
If this is right
- Patients can be shown a personalized 3D preview during the initial consultation, before any radiation-exposing CBCT or X-ray is taken.
- Because the pipeline needs no bone-movement parameters or manually placed landmarks, it removes clinician input from the prediction step and can run end-to-end in about 25 minutes of training time on a single GPU.
- The 5-fold cross-validated accuracy on 163 real surgery pairs (augmented to 1,330) suggests the model's error is concentrated in small local deviations rather than large outliers, since its Hausdorff distance drops more than its Chamfer distance relative to the baseline.
- Even without synthesized data (133 real samples), the model still edged out the baseline on both metrics, suggesting the loss design carries the method even in data-poor settings.
- Data augmentation is the single largest contributor in the ablation: removing it increased Chamfer distance by 17.46%, so the stitching scheme is essential to the reported accuracy.
Reading between the lines
- Beyond the paper: because the predictor receives only the pre-surgery latent code and never the surgical plan, the model's output is effectively an average of the surgical outcomes in the training set for that facial shape; two patients with the same pre-surgery scan who undergo different jaw movements could not receive different previews, so a plan-conditioned version would be a natural next step
- Beyond the paper: the authors' limitation note that the dataset is Asian-only implies the reported 9.00 mm and 2.50 mm numbers should not be expected to transfer to other populations without retraining on diverse scans, and a direct cross-ethnicity evaluation would settle transferability.
- Beyond the paper: the texture-transfer step, which deforms the original textured scan onto the predicted mesh via barycentric coordinates, means the preview inherits the patient's own skin appearance; this could be repurposed as a shared visual aid during plan discussion, though the paper does not test that workflow.
- Beyond the paper: the user study's near-chance discrimination (specificity around 46%) is a strong perceptual claim, and a larger study with more than five medical professionals and with static plus rotating views would be needed to confirm that indistinguishability generalizes.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents a fully automated pipeline that predicts a 3D postoperative facial mesh from a pre-operative 3D facial scan using FLAME latent codes and a small code-difference predictor. It introduces mouth-convexity and asymmetry losses, latent-code and geometry losses, and a data-augmentation scheme that stitches upper and lower facial halves to synthesize training pairs. The method is evaluated on 163 real orthognathic surgery pairs with 5-fold cross-validation, reporting a mean Hausdorff distance of 9.00 mm and Chamfer distance of 2.50 mm, which are claimed to outperform the LARS baseline. A user study reports that clinicians and engineers cannot reliably distinguish predicted from real postoperative outcomes.
Significance. If the prediction is valid, the method offers a radiation-free, fully automatic consultation tool for orthognathic patients, and the integration of clinically motivated losses and synthetic data augmentation is a useful contribution. The paper includes a 5-fold cross-validated comparison with statistical testing, an ablation study, and a blinded user study, which strengthen the relative-performance claim. However, the absolute accuracy is not clinically anchored, and the plan-dependence of orthognathic outcomes is a major conceptual concern that must be addressed before the central claim is established.
major comments (4)
- [Section 2.1, Eq. (5), Fig. 1; Section 5] The prediction network takes only the pre-operative FLAME latent code as input; no surgical-plan parameters (e.g., jaw movement magnitudes, genioplasty, rotation) are used. Section 5 explicitly acknowledges the model's lack of adjustable parameters for clinicians. Orthognathic outcomes are plan-dependent: the same pre-operative face can undergo different jaw corrections with different soft-tissue results. Therefore, the model can only learn a conditional average over the training distribution, and the preview may not match the patient's actual planned surgery. Table 4 reports only plan-averaged HD/CD, so large plan-dependent systematic errors could be concealed. Please provide an experiment stratifying errors by procedure type (e.g., setback vs. advancement, with/without genioplasty) or include plan parameters as input, and reframe the central claim accordingly.
- [Section 2.1, mouth-convexity and asymmetry losses] The mouth-convexity loss and asymmetry loss explicitly penalize deviations from a normative ideal (3 mm s-line tolerance, perfect chin symmetry). The text states that training data contain residual asymmetry/protrusion and that these losses are intended to 'enhance' the result, meaning the model is trained to produce an idealized outcome rather than necessarily the actual postoperative outcome. Yet the evaluation (Table 4) and user study (Section 3.3) use the real postoperative scan as ground truth. The paper does not reconcile this tension: if the losses bias predictions away from real outcomes, then the 'accurate prediction' claim is ambiguous. Please clarify the intended target (actual outcome vs. idealized correction) and evaluate accordingly, for example by reporting errors against both the real post-op scan and a clinically defined ideal target.
- [Section 3.5, Table 4] The claim of 'high prediction accuracy' is not anchored to any clinical standard. A mean HD of 9.00 mm and CD of 2.50 mm are reported without standard deviations, confidence intervals, or a clinically acceptable error threshold. The t-test in Table 5 establishes only that the method outperforms LARS on these metrics, not that the predictions are clinically accurate. Please report distributional statistics and compare against a clinical reference point, such as the typical inter-operator variability in surgical predictions or a clinically meaningful difference threshold.
- [Section 2.2] The data augmentation procedure is under-specified. The random variable used to conditionally generate F_gen is not defined (the symbol is missing in the text), and there is an apparent inconsistency between 'We modify the upper part of the face while directly copying the lower part' and the earlier description of stitching F^u_gen with the lower parts. Additionally, the assumption that a horizontal plane separates an unchanged upper face from a modified lower face is an unvalidated axiom; if surgery changes the upper lip, nasal base, or other regions above the chosen plane, the synthetic pairs may introduce unrealistic correspondences. Please define the random variable, clarify the stitching process, and provide evidence or a sensitivity analysis for the plane-separation assumption.
minor comments (6)
- [Section 2.2] The sentence 'we conditionally generate a synthetic face F_gen based on the lower part of the postoperative scan F^l_post with a random variable .' has a missing symbol after 'random variable'; please insert the intended notation.
- [Section 3.2.3] The description 'two fully connected modules with a hidden layer of 100 dimensions and input and output layers of 300 dimensions' is ambiguous; please specify the exact architecture (number of layers, activations, and how the two modules are connected).
- [Section 2.1, Eq. (4)] The balancing weight w in the geometry loss is not defined in the text or in the implementation details; please state its value or how it is chosen.
- [Section 3.3] The sentence 'A total of 30 randomly selected images (including A and B) were presented to each participant' is ambiguous; clarify that each participant judged 30 pairs of images (A and B) rather than 30 single images.
- [Section 3.1 vs. Section 5] Section 3.1 states that postoperative scans were recorded at least three months after surgery, while Section 5 says 'at least six months after surgery'; please reconcile this discrepancy.
- [Figure 1 caption] The caption says 'With the help of medical, latent code, and geological types of loss'; 'geological' should be 'geometry' (or 'geometrical').
Circularity Check
No significant circularity: the headline HD/CD numbers are measured against real held-out postoperative scans using metrics that are not the training objective; FLAME, clinical criteria, and the LARS baseline are external; self-citations are not load-bearing.
full rationale
The paper's central claim—that a code-difference predictor trained on 163 real pre/post scan pairs predicts held-out postoperative geometry—is not circular. Section 3.5/Table 4 reports Hausdorff and Chamfer distances against real postoperative 3dMD scans withheld via 5-fold cross-validation (Section 3.2.3). The reported metrics (Eqs. 7-8) are not the training objective: the geometry loss in Eq. 4 is a per-vertex L2 plus surface-normal loss on FLAME correspondences, not the bidirectional nearest-neighbor Chamfer/Hausdorff metrics in Eqs. 7-8, and the mouth-convexity and asymmetry losses in Eqs. 1-2 are clinical regularizers, not fitted parameters renamed as predictions. FLAME [20], Steiner's line [23], the midsagittal-plane criterion [24], and the LARS baseline [4] are external published sources. Self-authored references appear in the FLAME integration discussion ([21] includes author Han, [22] includes authors Zhang and Wang) and in the scanner-accuracy citation [26], but the load-bearing identity of FLAME is established by the external reference [20], and scanner accuracy is not used to define the predicted outputs. The Section 5 limitation—that the model does not accept adjustable surgical-plan parameters and thus may output plan-averaged results—is a clinical validation and generalizability concern, not a circular argument: it does not make the measured HD/CD equal to the training loss or to any fitted parameter. No circular step was found.
Assumptions & free parameters
free parameters (4)
- Loss weights alpha_p, alpha_a, alpha_f, alpha_g =
5000, 5000, 1, 1
- Mouth convexity tolerance =
3 mm
- Augmentation random variable xi =
not specified
- Landmark fitting error threshold =
not specified
assumptions (5)
- domain assumption A patient's postsurgical facial appearance is a function of the preoperative 3D surface geometry alone; the predictor never receives the surgical plan or bone movement parameters.
- domain assumption The FLAME parametric face model provides a latent space in which the pre-to-post surgical change can be approximated by a low-dimensional code difference.
- ad hoc to paper There exists a horizontal plane separating an unchanged upper face from a modified lower face, so synthetic pairs can be made by stitching different patients' upper and lower halves.
- domain assumption Postoperative scans taken at least 3 to 6 months after surgery reflect the stabilized surgical outcome.
- domain assumption Faces are captured in neutral expression and natural head position, so FLAME's expression vector can be omitted and the mid-sagittal plane normal approximated as the x-axis.
Cite this review
Pith. "Pith review of Facial Surgery Preview Based on the Orthognathic Treatment Prediction." pith.science (2026). https://pith.science/paper/WBSMPFQ2
@misc{pith2026241211045,
author = {Pith},
title = {Pith review of: Facial Surgery Preview Based on the Orthognathic Treatment Prediction},
year = {2026},
howpublished = {\url{https://pith.science/paper/WBSMPFQ2}},
note = {Machine review of arXiv:2412.11045}
}
read the original abstract
Orthognathic surgery consultation is essential to help patients understand the changes to their facial appearance after surgery. However, current visualization methods are often inefficient and inaccurate due to limited pre- and post-treatment data and the complexity of the treatment. To overcome these challenges, this study aims to develop a fully automated pipeline that generates accurate and efficient 3D previews of postsurgical facial appearances for patients with orthognathic treatment without requiring additional medical images. The study introduces novel aesthetic losses, such as mouth-convexity and asymmetry losses, to improve the accuracy of facial surgery prediction. Additionally, it proposes a specialized parametric model for 3D reconstruction of the patient, medical-related losses to guide latent code prediction network optimization, and a data augmentation scheme to address insufficient data. The study additionally employs FLAME, a parametric model, to enhance the quality of facial appearance previews by extracting facial latent codes and establishing dense correspondences between pre- and post-surgery geometries. Quantitative comparisons showed the algorithm's effectiveness, and qualitative results highlighted accurate facial contour and detail predictions. A user study confirmed that doctors and the public could not distinguish between machine learning predictions and actual postoperative results. This study aims to offer a practical, effective solution for orthognathic surgery consultations, benefiting doctors and patients.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
S. H. Jeong, J. P. Yun, H.-G. Yeom, H. J. Lim, J. Lee, B. C. Kim, Deep learning based discrimination of soft tissue profiles requiring orthognathic surgery by facial photographs, Scientific Reports 10 (2020) 16235
work page 2020
- [2]
-
[3]
E. D. Rekow, Digital dentistry: The new state of the art—is it disruptive or destructive?, Dental Materials 36 (2020) 9–24
work page 2020
-
[4]
P. G. Knoops, A. Papaioannou, A. Borghi, R. W. Breakey, A. T. Wilson, O. Jeelani, S. Zafeiriou, D. Steinbacher, B. L. Padwa, D. J. Dunaway, et al., A machine learning framework for auto- mated diagnosis and computer-assisted planning in plastic and reconstructive surgery, Scientific reports 9 (2019) 1–12
work page 2019
-
[5]
C. Tanikawa, T. Yamashiro, Development of novel artificial intelligence systems to predict facial morphology after orthog- nathic surgery and orthodontic treatment in japanese patients, Scientific reports 11 (2021) 1–11
work page 2021
-
[6]
R. Ter Horst, H. van Weert, T. Loonen, S. Berg´ e, S. Vinaya- halingam, F. Baan, T. Maal, G. de Jong, T. Xi, Three- dimensional virtual planning in mandibular advancement surgery: Soft tissue prediction based on deep learning, Jour- nal of Cranio-Maxillofacial Surgery 49 (2021) 775–782
work page 2021
-
[7]
N. Chaiprasittikul, B. Thanathornwong, S. Pornprasertsuk- Damrongsri, S. Raocharernporn, S. Maponthong, S. Manopatanakul, Application of a multi-layer percep- tron in preoperative screening for orthognathic surgery, Healthcare Informatics Research 29 (2023) 16–22
work page 2023
-
[8]
J. N. Saeed, A. M. Abdulazeez, D. A. Ibrahim, Automatic facial aesthetic prediction based on deep learning with loss ensembles, Applied Sciences 13 (2023) 9728
work page 2023
Show all 31 references
-
[9]
Park, J.-H
J.-A. Park, J.-H. Moon, J.-M. Lee, S. J. Cho, B.-M. Seo, R. E. Donatelli, S.-J. Lee, Does artificial intelligence predict orthog- nathic surgical outcomes better than conventional linear regres- sion methods?, The Angle Orthodontist (2024)
2024
-
[10]
Kim, J.-S
I.-H. Kim, J.-S. Kim, J. Jeong, J.-W. Park, K. Park, J.-H. Cho, M. Hong, K.-H. Kang, M. Kim, S.-J. Kim, Y.-J. Kim, S.-J. Sung, Y. H. Kim, S.-H. Lim, S.-H. Baek, N. Kim, Orthognathic surgical planning using graph cnn with dual embedding module: External validations with multi-h...
2023
-
[11]
Y. Park, J. Choi, Y. Kim, S. Choi, J. Lee, K. Kim, C. Chung, Deep learning–based prediction of the 3d postorthodontic facial changes, Journal of Dental Research 101 (2022) 1372–1379
2022
-
[12]
Laurinaviˇ cius, R
D. Laurinaviˇ cius, R. Maskeli¯ unas, R. Damaˇ seviˇ cius, Improve- ment of facial beauty prediction using artificial human faces generated by generative adversarial network, Cognitive Com- putation 15 (2023) 998–1015
2023
-
[13]
Cheng, X
M. Cheng, X. Zhang, J. Wang, Y. Yang, M. Li, H. Zhao, J. Huang, C. Zhang, D. Qian, H. Yu, Prediction of orthog- nathic surgery plan from 3d cephalometric analysis via deep learning, BMC Oral Health 23 (2023) 161
2023
-
[14]
Q. Ma, E. Kobayashi, B. Fan, K. Hara, K. Nakagawa, K. Masamune, I. Sakuma, H. Suenaga, Machine-learning-based approach for predicting postoperative skeletal changes for or- thognathic surgical planning, The International Journal of Med- ical Robotics and Computer Assisted Surg...
2022
-
[15]
Sankar, R
H. Sankar, R. Alagarsamy, B. Lal, S. S. Rana, A. Roychoud- hury, A. Agrawal, S. Wankhar, Role of artificial intelligence in treatment planning and outcome prediction of jaw correc- tive surgeries by using 3-d imaging- a systematic review, Oral Surgery, Oral Medicine, Oral Path...
2024
-
[16]
L. Ma, D. Kim, C. Lian, D. Xiao, T. Kuang, Q. Liu, Y. Lang, H. H. Deng, J. Gateno, Y. Wu, et al., Deep simulation of facial appearance changes following craniomaxillofacial bony move- ments in orthognathic surgical planning, in: MICCAI, Springer, 2021, pp. 459–468
2021
-
[17]
Booth, A
J. Booth, A. Roussos, S. Zafeiriou, A. Ponniah, D. Dunaway, A 3d morphable model learnt from 10,000 faces, in: Proceed- ings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 5543–5552
2016
-
[18]
J. Ling, Z. Wang, M. Lu, Q. Wang, C. Qian, F. Xu, Structure- aware editable morphable model for 3d facial detail animation and manipulation, in: Computer Vision–ECCV 2022: 17th European Conference, Tel Aviv, Israel, October 23–27, 2022, Proceedings, Part III, Springer, 2022, p...
2022
-
[19]
Z. Qiu, Y. Li, D. He, Q. Zhang, L. Zhang, Y. Zhang, J. Wang, L. Xu, X. Wang, Y. Zhang, J. Yu, Sculptor: Skeleton-consistent face creation using a learned parametric generator 41 (2022)
2022
-
[20]
T. Li, T. Bolkart, M. J. Black, H. Li, J. Romero, Learning a model of facial shape and expression from 4D scans, ACM Proc. SIGGRAPH Asia 36 (2017) 194:1–194:17
2017
-
[21]
Zheng, J
W. Zheng, J. Zhao, X. Liu, Y. Pan, Z. Gan, H. Han, N. Liu, Flame-based multi-view 3d face reconstruction, in: Advances in Computer Graphics: 40th Computer Graphics International Conference, CGI 2023, Shanghai, China, August 28 – Septem- ber 1, 2023, Proceedings, Part IV, Sprin...
2023
-
[22]
Liang, C
Y. Liang, C. Zhang, J. Zhao, W. Wang, X. Li, Skull-to-Face: Anatomy-Guided 3D Facial Reconstruction and Editing , IEEE Transactions on Visualization & Computer Graphics (5555) 1– 13
-
[23]
C. C. Steiner, The use of cephalometrics as an aid to planning and assessing orthodontic treatment: report of a case, American journal of orthodontics 46 (1960) 721–735
1960
-
[24]
Dobai, Z
A. Dobai, Z. Markella, T. V ´ ızkelety, C. Fouquet, A. Rosta, J. Barab´ as, Landmark-based midsagittal plane analysis in pa- tients with facial symmetry and asymmetry based on cbct anal- 11 ysis tomography, Journal of Orofacial Orthopedics/Fortschritte der Kieferorthop¨ adie 7...
2018
-
[25]
Zhang, S
C. Zhang, S. Bengio, M. Hardt, B. Recht, O. Vinyals, Under- standing deep learning (still) requires rethinking generalization, Communications of the ACM 64 (2021) 107–115
2021
-
[26]
Z. Shan, R. T.-C. Hsung, C. Zhang, J. Ji, W. S. Choi, W. Wang, Y. Yang, M. Gu, B. S. Khambay, Anthropometric accuracy of three-dimensional average faces compared to conventional facial measurements, Scientific Reports 11 (2021) 12254
2021
-
[27]
C. Yu, J. Wang, C. Peng, C. Gao, G. Yu, N. Sang, Bisenet: Bilateral segmentation network for real-time semantic segmen- tation, in: ECCV, 2018, pp. 325–341
2018
-
[28]
Dong, S.-I
X. Dong, S.-I. Yu, X. Weng, S.-E. Wei, Y. Yang, Y. Sheikh, Supervision-by-Registration: An unsupervised approach to im- prove the precision of facial landmark detectors, in: CVPR, 2018, pp. 360–368
2018
-
[29]
Dubuisson, A
M.-P. Dubuisson, A. Jain, A modified hausdorff distance for ob- ject matching, in: Proceedings of 12th International Conference on Pattern Recognition, volume 1, 1994, pp. 566–568 vol.1
1994
-
[30]
Y. Yang, C. Feng, Y. Shen, D. Tian, Foldingnet: Point cloud auto-encoder via deep grid deformation, in: 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 206–215
2018
-
[31]
Starch-Jensen, F
T. Starch-Jensen, F. Hern´ andez-Alfaro, ¨O. Kesmez, R. Gor- gis, A. Valls-Onta˜ n´ on, Accuracy of orthognathic surgical plan- ning using three-dimensional virtual techniques compared with conventional two-dimensional techniques: a systematic review, Journal of Oral & Maxillo...
2023
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.