{"id":"44ac6c00-0526-4925-8303-9d79ffc912d6","arxiv_id":"2412.11045","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A fully automatic FLAME-based pipeline predicts 3D postsurgical facial appearance from a single 3D scan, using mouth-convexity and asymmetry losses plus synthetic data augmentation to beat a published baseline.","lead":"This paper describes a machine learning pipeline that turns a patient's 3D facial scan into a predicted 3D preview of how the face may look after jaw surgery, without requiring CT or X-ray images. The system uses the FLAME face model, custom mouth and asymmetry losses, and synthetic data to improve the preview.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The load-bearing assumption is that post-surgical facial shape is a single-valued function of the pre-operative surface alone; orthognathic outcomes are plan-dependent, so the no-plan predictor may output an average that does not match the actual surgery.","rationale":"Agree with the reader's weakest assumption. The strongest claim is only as strong as the input-output map: if the same pre-op face maps to multiple acceptable post-op faces depending on the surgical plan, the single-output network is at best an average predictor, and the two medical losses add a further normative pull that can move predictions away from the actual outcome. Neither the quantitative comparison (Table 4) nor the user study can detect this, because both aggregate over plans and the user study has no plan label. The FLAME-based construction and the 5-fold CV are reasonable, and the data augmentation is sensible, but they do not address plan dependence. A concrete, feasible test is to condition the predictor on the recorded surgical movement and see whether accuracy improves; the data are already in the authors' treatment records. If the test fails, the central claim should be narrowed to 'typical outcome preview' rather than prediction of a specific treatment; if it passes, the concern is resolved. Other issues (no error bars, no clinical threshold, no code) are real but secondary.","tokens_in":14208,"tokens_out":9094,"duration_ms":90726,"concrete_test":"Using the same 163-case dataset and 5-fold splits, stratify residuals by the actual surgical movement recorded in the treatment notes (mandibular advancement/setback magnitude, maxillary movement, genioplasty). First, regress the signed chin/lip prediction error on the movement vector; a nonzero slope shows the no-plan model misses plan-dependent signal. Then retrain the predictor with the plan vector (or surgery-type indicator) concatenated to the FLAME latent code and compare per-patient CD/HD with the original model by a paired test on identical folds. If the plan-conditioned model is significantly better or errors correlate with movement magnitude, the single-output no-plan model cannot faithfully preview the actual treatment; if neither holds, the concern is resolved.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 2.1 makes the predictor a function of the pre-operative FLAME latent code only (Fig. 1, Eq. 5), and Section 5 explicitly declines adjustable surgical-plan parameters. For orthognathic surgery, the same pre-operative face can undergo different jaw movements (setback vs advancement, with/without genioplasty, differential maxillary impaction), producing different soft-tissue outcomes; the mid-sagittal plane and s-line losses in Eqs. 1-2 further push every prediction toward a single normative ideal. Trained on mixed plans, the code-difference predictor therefore learns a conditional average (plus a normative bias), not the outcome of the patient's actual treatment. Table 4 reports only plan-averaged HD/CD, so a small mean error can conceal large, plan-dependent systematic deviations. This directly threatens the Section 6 central claim 'predicting facial appearance following orthognathic treatment': the treatment itself is not an input. The paper's own limitation paragraph admits the lack of adjustable parameters, but no experiment tests whether outcome variability is explained by pre-op geometry alone.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a fully automated pipeline that predicts a 3D postoperative facial mesh from a pre-operative 3D facial scan using FLAME latent codes and a small code-difference predictor. It introduces mouth-convexity and asymmetry losses, latent-code and geometry losses, and a data-augmentation scheme that stitches upper and lower facial halves to synthesize training pairs. The method is evaluated on 163 real orthognathic surgery pairs with 5-fold cross-validation, reporting a mean Hausdorff distance of 9.00 mm and Chamfer distance of 2.50 mm, which are claimed to outperform the LARS baseline. A user study reports that clinicians and engineers cannot reliably distinguish predicted from real postoperative outcomes.","tokens_in":14481,"tokens_out":8163,"duration_ms":68701,"significance":"If the prediction is valid, the method offers a radiation-free, fully automatic consultation tool for orthognathic patients, and the integration of clinically motivated losses and synthetic data augmentation is a useful contribution. The paper includes a 5-fold cross-validated comparison with statistical testing, an ablation study, and a blinded user study, which strengthen the relative-performance claim. However, the absolute accuracy is not clinically anchored, and the plan-dependence of orthognathic outcomes is a major conceptual concern that must be addressed before the central claim is established.","major_comments":[{"comment":"The prediction network takes only the pre-operative FLAME latent code as input; no surgical-plan parameters (e.g., jaw movement magnitudes, genioplasty, rotation) are used. Section 5 explicitly acknowledges the model's lack of adjustable parameters for clinicians. Orthognathic outcomes are plan-dependent: the same pre-operative face can undergo different jaw corrections with different soft-tissue results. Therefore, the model can only learn a conditional average over the training distribution, and the preview may not match the patient's actual planned surgery. Table 4 reports only plan-averaged HD/CD, so large plan-dependent systematic errors could be concealed. Please provide an experiment stratifying errors by procedure type (e.g., setback vs. advancement, with/without genioplasty) or include plan parameters as input, and reframe the central claim accordingly.","section":"Section 2.1, Eq. (5), Fig. 1; Section 5"},{"comment":"The mouth-convexity loss and asymmetry loss explicitly penalize deviations from a normative ideal (3 mm s-line tolerance, perfect chin symmetry). The text states that training data contain residual asymmetry/protrusion and that these losses are intended to 'enhance' the result, meaning the model is trained to produce an idealized outcome rather than necessarily the actual postoperative outcome. Yet the evaluation (Table 4) and user study (Section 3.3) use the real postoperative scan as ground truth. The paper does not reconcile this tension: if the losses bias predictions away from real outcomes, then the 'accurate prediction' claim is ambiguous. Please clarify the intended target (actual outcome vs. idealized correction) and evaluate accordingly, for example by reporting errors against both the real post-op scan and a clinically defined ideal target.","section":"Section 2.1, mouth-convexity and asymmetry losses"},{"comment":"The claim of 'high prediction accuracy' is not anchored to any clinical standard. A mean HD of 9.00 mm and CD of 2.50 mm are reported without standard deviations, confidence intervals, or a clinically acceptable error threshold. The t-test in Table 5 establishes only that the method outperforms LARS on these metrics, not that the predictions are clinically accurate. Please report distributional statistics and compare against a clinical reference point, such as the typical inter-operator variability in surgical predictions or a clinically meaningful difference threshold.","section":"Section 3.5, Table 4"},{"comment":"The data augmentation procedure is under-specified. The random variable used to conditionally generate F_gen is not defined (the symbol is missing in the text), and there is an apparent inconsistency between 'We modify the upper part of the face while directly copying the lower part' and the earlier description of stitching F^u_gen with the lower parts. Additionally, the assumption that a horizontal plane separates an unchanged upper face from a modified lower face is an unvalidated axiom; if surgery changes the upper lip, nasal base, or other regions above the chosen plane, the synthetic pairs may introduce unrealistic correspondences. Please define the random variable, clarify the stitching process, and provide evidence or a sensitivity analysis for the plane-separation assumption.","section":"Section 2.2"}],"minor_comments":[{"comment":"The sentence 'we conditionally generate a synthetic face F_gen based on the lower part of the postoperative scan F^l_post with a random variable .' has a missing symbol after 'random variable'; please insert the intended notation.","section":"Section 2.2"},{"comment":"The description 'two fully connected modules with a hidden layer of 100 dimensions and input and output layers of 300 dimensions' is ambiguous; please specify the exact architecture (number of layers, activations, and how the two modules are connected).","section":"Section 3.2.3"},{"comment":"The balancing weight w in the geometry loss is not defined in the text or in the implementation details; please state its value or how it is chosen.","section":"Section 2.1, Eq. (4)"},{"comment":"The sentence 'A total of 30 randomly selected images (including A and B) were presented to each participant' is ambiguous; clarify that each participant judged 30 pairs of images (A and B) rather than 30 single images.","section":"Section 3.3"},{"comment":"Section 3.1 states that postoperative scans were recorded at least three months after surgery, while Section 5 says 'at least six months after surgery'; please reconcile this discrepancy.","section":"Section 3.1 vs. Section 5"},{"comment":"The caption says 'With the help of medical, latent code, and geological types of loss'; 'geological' should be 'geometry' (or 'geometrical').","section":"Figure 1 caption"}],"recommendation":"major_revision","confidential_remarks":"The plan-dependence issue is the most serious concern: the central claim of predicting the outcome of orthognathic treatment is not supported when the treatment plan is not an input. However, this is potentially addressable by reframing the contribution as a plan-averaged preview, adding a stratified evaluation, or incorporating plan parameters. The paper is otherwise within scope for CMPB and contains a reproducible experimental design with cross-validation and ablation studies. The self-citations in the FLAME-related references are acceptable but the novelty of the proposed losses relative to standard orthognathic assessment criteria should be clarified."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a legitimate incremental advance in scan-only orthognathic preview, with two genuinely new pieces — the mouth-convexity and asymmetry losses and the upper/lower face stitching augmentation — and the evaluation is honestly done against real postoperative scans. The central claim, though, is bigger than what the architecture can deliver, because the predictor never sees the surgical plan.\n\nThe new stuff is real. The two clinical losses encode surgical goals directly and are differentiable, which is a sensible way to push a FLAME latent-code regressor toward clinically plausible outputs. The stitching augmentation is clever: it preserves the lower-face surgical change while randomizing the upper face, multiplying pairs from a small dataset. The 5-fold CV comparison against LARS is fair enough and shows statistically significant gains (HD 9.00 vs 9.68 mm, CD 2.50 vs 2.77 mm, p<0.05). Ablations show each component helps, with augmentation the biggest contributor. Evaluation uses real post-op scans, so the circularity burden is low.\n\nThe soft spots are real but not fatal. The load-bearing assumption is that post-surgical facial shape is a single-valued function of the pre-op surface alone. The mid-sagittal and s-line losses push every prediction toward a normative ideal, and Section 5 explicitly declines adjustable plan parameters. If the same pre-op face can undergo setback vs advancement, with or without genioplasty, the model can only output one average. The paper's own limitation paragraph admits this, but no experiment tests it. That is the core weakness. Also, absolute accuracy is weakly anchored: no error bars on the HD/CD means, no clinical threshold for what 'accurate' means, and the user study is small (5 medics, 15 engineers) and essentially chance-level — which the authors honestly interpret as indistinguishability, but it is underpowered evidence. No code or data released.\n\nWho this is for: people working on craniofacial surgical prediction or ML-assisted patient consultation. It is a serious paper. I would send it to a competent referee, with the request that the referee ask for plan-conditioned prediction or at least an analysis of outcome variability across plan types, error bars on the metrics, and ideally a release of data/code.","headline":"Legitimate incremental advance in scan-only orthognathic preview, but the no-plan predictor undercuts the headline claim.","tokens_in":14983,"tokens_out":2114,"would_cite":false,"duration_ms":18313,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A scan-only neural pipeline predicts facial shape after jaw-correction surgery more accurately than a published baseline, and blinded experts cannot reliably tell the predictions from real outcomes.","keywords":["computer-aided detection and diagnosis","geometric deep learning","visualization","orthognathic surgery","facial appearance prediction","FLAME parametric model","data augmentation","3D face reconstruction"],"falsifier":"Compare the model's single prediction for a set of patients whose preoperative scans are near-identical but whose actual surgeries involved measurably different jaw movements, for example different mandibular advancement distances or different degrees of chin rotation. If the real postoperative outcomes differ from each other by more than roughly the reported Chamfer error (around 2.5 mm) while both are consistent with the same pre-surgery scan, then the scan-only assumption fails and the model can only be producing an averaged outcome.","tokens_in":14025,"feed_emoji":"🦷","tokens_out":9622,"duration_ms":80407,"temperature":0.7,"pith_summary":"This paper tries to establish that a fully automated pipeline can predict a patient's facial appearance after orthognathic (jaw-correction) surgery from a single preoperative 3D facial scan, with no X-ray, CT, surgical plan, or manual landmarks. The authors represent pre- and post-surgery faces as latent codes of the FLAME parametric head model, then train a small network to predict the code difference, guided by two medically motivated losses (mouth-convexity and asymmetry) together with latent-code and geometry losses. On 163 real surgery pairs, augmented to 1,330 samples, the method reports a mean Hausdorff distance of 9.00 mm and Chamfer distance of 2.50 mm, improving on the baseline's 9.68 mm and 2.77 mm. A blinded user study found that doctors and engineers could not reliably distinguish machine-generated previews from real surgical outcomes. If these results hold, orthognathic consultations can offer patients a fast, radiation-free preview that may lower anxiety and support shared decision-making.","feed_headline":"Scan-only AI previews jaw surgery outcomes within 2.5 mm","feed_subtitle":"Doctors and engineers could not reliably distinguish machine-generated faces from real surgical results.","key_machinery":"The load-bearing object is FLAME, a parametric head model with nonlinear jaw articulation, used as both encoder and decoder: pre- and post-surgery scans are fitted to FLAME latent codes, and the predicted code difference is decoded back into a mesh with point-to-point correspondences. The predictor itself is a small fully connected network trained with a weighted sum of four losses: mouth-convexity loss (squared deviation beyond 3 mm from the s-line to the upper and lower lip midpoints), asymmetry loss (distances of chin vertex-pair midpoints from a least-squares mid-sagittal plane plus orientation disagreement), latent-code loss, and geometry loss (point positions plus surface normals). A second mechanism is the data-augmentation scheme: a horizontal plane splits the face into an unchanged upper region and a surgically modified lower region, a synthetic face is generated by randomizing the upper-face latent code, and this upper part is stitched to real pre- and post-surgery lower parts to create many plausible training pairs.","core_discovery":"The paper's central claim is that postsurgical facial geometry after orthognathic treatment is predictable from the preoperative facial surface alone, once that surface is represented in a parametric model that can articulate the jaw. Concretely, the authors claim that a code-difference predictor operating in FLAME's latent space, trained with (i) a mouth-convexity loss that penalizes lip protrusion beyond a 3 mm tolerance relative to the s-line, (ii) an asymmetry loss measuring chin-point deviation from the mid-sagittal plane, (iii) a latent-code loss, and (iv) a geometry loss on points and normals, produces predictions that are more accurate than the baseline method by both Chamfer and Hausdorff distances. They further claim that their horizontal-split data-augmentation scheme, which synthesizes pre/post pairs by stitching the unchanged upper face to surgically altered lower-face regions, is a major contributor to this accuracy; removing it raises Chamfer distance by 17.46%. Finally, the user-study results are claimed as evidence that the predicted faces are perceptually close enough to real outcomes that neither doctors nor engineers can distinguish them at a statistically significant rate.","pith_inferences":["Beyond the paper: because the predictor receives only the pre-surgery latent code and never the surgical plan, the model's output is effectively an average of the surgical outcomes in the training set for that facial shape; two patients with the same pre-surgery scan who undergo different jaw movements could not receive different previews, so a plan-conditioned version would be a natural next step","Beyond the paper: the authors' limitation note that the dataset is Asian-only implies the reported 9.00 mm and 2.50 mm numbers should not be expected to transfer to other populations without retraining on diverse scans, and a direct cross-ethnicity evaluation would settle transferability.","Beyond the paper: the texture-transfer step, which deforms the original textured scan onto the predicted mesh via barycentric coordinates, means the preview inherits the patient's own skin appearance; this could be repurposed as a shared visual aid during plan discussion, though the paper does not test that workflow.","Beyond the paper: the user study's near-chance discrimination (specificity around 46%) is a strong perceptual claim, and a larger study with more than five medical professionals and with static plus rotating views would be needed to confirm that indistinguishability generalizes."],"forward_implications":["Patients can be shown a personalized 3D preview during the initial consultation, before any radiation-exposing CBCT or X-ray is taken.","Because the pipeline needs no bone-movement parameters or manually placed landmarks, it removes clinician input from the prediction step and can run end-to-end in about 25 minutes of training time on a single GPU.","The 5-fold cross-validated accuracy on 163 real surgery pairs (augmented to 1,330) suggests the model's error is concentrated in small local deviations rather than large outliers, since its Hausdorff distance drops more than its Chamfer distance relative to the baseline.","Even without synthesized data (133 real samples), the model still edged out the baseline on both metrics, suggesting the loss design carries the method even in data-poor settings.","Data augmentation is the single largest contributor in the ablation: removing it increased Chamfer distance by 17.46%, so the stitching scheme is essential to the reported accuracy."],"supporting_citations":[{"why":"Supplies the baseline method and its reported Chamfer and Hausdorff numbers that the paper must outperform.","marker":"[4]"},{"why":"Documents a prior network that needs CBCT data and suffers overfitting on fewer than 100 patients, motivating the scan-only design.","marker":"[5]"},{"why":"Describes a prior method requiring clinician-specified bone movement at inference, which the paper explicitly avoids.","marker":"[16]"},{"why":"Defines the FLAME parametric head model with jaw articulation that serves as the encoder-decoder for the latent codes.","marker":"[20]"},{"why":"Defines the s-line and the 3 mm lip-deviation tolerance used to construct the mouth-convexity loss.","marker":"[23]"},{"why":"Defines the landmark-based mid-sagittal plane analysis used to construct the asymmetry loss.","marker":"[24]"},{"why":"Validates the sub-0.2 mm accuracy of the facial scanning system used to collect the dataset.","marker":"[26]"},{"why":"Supplies the bilateral segmentation network used to remove hair and other non-facial elements from scans before fitting.","marker":"[27]"},{"why":"Supplies the registration-based landmark annotation method used for automatic 68-landmark labeling.","marker":"[28]"}],"fun_headline_variants":["One scan predicts post-surgery face within 2.5 mm","AI forecasts jaw surgery look from single facial scan","Single-scan AI matches actual post-op faces","Machine learning previews surgery outcome from one image","Predicting jaw surgery results from a single scan"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pipeline assumes that the post-surgery facial shape is determined by the preoperative facial surface geometry alone: the same pre-surgery scan is always mapped to one predicted outcome, with no dependence on the specific surgical plan, bone movements, or patient attributes.","fun_headline_variants_meta":{"raw":{"variants":["One scan predicts post-surgery face within 2.5 mm","AI forecasts jaw surgery look from single facial scan","Single-scan AI matches actual post-op faces","Machine learning previews surgery outcome from one image","Predicting jaw surgery results from a single scan"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000405,"raw_usage":{"total_tokens":2135,"prompt_tokens":998,"completion_tokens":1137,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":614,"completion_tokens_details":{"reasoning_tokens":1062}},"tokens_in":614,"tokens_out":1137,"duration_ms":7401,"temperature":1.0,"reasoning_tokens":1062,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:21:37.653071+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare the model's single prediction for a set of patients whose preoperative scans are near-identical but whose actual surgeries involved measurably different jaw movements, for example different mandibular advancement distances or different degrees of chin rotation. If the real postoperative outcomes differ from each other by more than roughly the reported Chamfer error (around 2.5 mm) while both are consistent with the same pre-surgery scan, then the scan-only assumption fails and the model can only be producing an averaged outcome.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the baseline method and its reported Chamfer and Hausdorff numbers that the paper must outperform."},{"cited_title":"Tanikawa, T","cited_arxiv_id":null,"evidence_quote":"Documents a prior network that needs CBCT data and suffers overfitting on fewer than 100 patients, motivating the scan-only design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes a prior method requiring clinician-specified bone movement at inference, which the paper explicitly avoids."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the FLAME parametric head model with jaw articulation that serves as the encoder-decoder for the latent codes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the s-line and the 3 mm lip-deviation tolerance used to construct the mouth-convexity loss."},{"cited_title":"Dobai, Z","cited_arxiv_id":null,"evidence_quote":"Defines the landmark-based mid-sagittal plane analysis used to construct the asymmetry loss."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Validates the sub-0.2 mm accuracy of the facial scanning system used to collect the dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the bilateral segmentation network used to remove hair and other non-facial elements from scans before fitting."},{"cited_title":"Dong, S.-I","cited_arxiv_id":null,"evidence_quote":"Supplies the registration-based landmark annotation method used for automatic 68-landmark labeling."}],"review_version":1}