{"id":"94e4d25f-f6c0-45e9-84f6-26b9436ca4fb","arxiv_id":"2607.16631","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"TANS-FO maps clinical text to lattice insole geometry and uses a GNN stress surrogate in a closed loop, reporting R²=0.94 vs FEA and a 34.7% surrogate-predicted peak-pressure reduction over parametric CAD.","lead":"This paper presents a prototype pipeline that turns written clinical prescriptions for foot orthoses into 3D-printable insole designs, using a text-aligned neural network and a fast stress-prediction model in place of slow finite-element simulations. The authors report large surrogate-predicted pressure reductions, but the headline numbers are computational estimates, not measured clinical outcomes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline 34.7% reduction rests on a surrogate that is both optimizer and evaluator; its R²=0.942 was measured on generic foot meshes, not on the optimized lattice designs, and its known high-curvature underestimation may inflate the reported gain.","rationale":"The reader's CONDITIONAL verdict is appropriate. The paper is candid that all performance figures are surrogate-predicted and that clinical efficacy is not claimed; nevertheless, the surrogate-predicted 34.7% reduction is presented as the main quantitative result. The most load-bearing assumption is that the GNN surrogate remains accurate on the designs it produces in the closed loop. The R²=0.942 against Abaqus is real independent evidence for the surrogate's general accuracy on a generic test set, and the paper's honesty about the high-curvature underestimation is creditable. However, that same admission creates a concrete risk: if optimized designs are richer in high-curvature features than the generic test meshes, the optimizer can exploit the surrogate's systematic bias to report reductions that would not survive an Abaqus re-run. This is not an accusation of fraud or sloppiness; it is a standard closed-loop validation gap. The proposed Abaqus re-run of final optimized designs is the minimal check that would settle whether the headline is robust. Until that is done, the CONDITIONAL verdict should stand, and the paper should either supply the FEA verification or present the headline as a purely surrogate-internal metric with no claim of external validity. The self-referential embedding update in Algorithm 1 is a secondary concern but does not change this recommendation.","tokens_in":17191,"tokens_out":4160,"duration_ms":42847,"concrete_test":"Take the 50 Male 18–40 test designs from Table 5 (code/data are released in the GitHub/Kaggle repositories) and re-run each final TANS-FO mesh and its parametric-CAD counterpart through Abaqus quasi-static FEA using the same loading protocol. Compute the peak-pressure reduction on the Abaqus outputs and compare it with the surrogate-predicted 34.7%/21.4% values. If the Abaqus-verified difference is below the reported 95% CI lower bound (11.5 points), or if the GNN error on the optimized designs is directionally biased (under-prediction at high-curvature regions), the headline claim is not supported. Also record the GNN-vs-Abaqus error separately on optimized designs to check distribution shift.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—34.7% peak-pressure reduction over parametric CAD (§4.1.4, Table 5)—is generated and evaluated by the same GNN surrogate within the closed loop (Algorithm 1, Eqs. (2)–(4)). The only external grounding is the surrogate-vs-Abaqus agreement R²=0.942, but that is reported on 'an independent test mesh set of 500 irregular foot models' (§4.1.2), which is not shown to be representative of the optimized TANS-FO lattice orthoses. The paper itself notes the GNN 'tends to slightly underestimate peak stress in high-curvature regions ... by 2–4%' (§4.1.4). If the optimized designs contain more high-curvature features (density gradients, lattice-strut intersections) than the generic test meshes, the surrogate's systematic under-prediction would disproportionately reduce the predicted peak pressure for TANS-FO designs, inflating the reported 13.3-point improvement over the parametric CAD baseline. No final optimized design is re-run through Abaqus; the limitation section explicitly says final designs 'should always be verified by a high-fidelity FEA solver before fabrication' (§5), but this verification is not performed for the reported results. The clinical embedding feedback in Algorithm 1 (AttentionModulation) adds a further self-referential element, but the surrogate-evaluator confounding is the more direct threat to the headline.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents TANS-FO, a modular research-prototype pipeline that translates unstructured clinical text into customized foot-orthosis geometry. The system couples a Text-Aligned Neural Surrogate (TANS) that projects clinical-text embeddings onto a lattice-density field with a Graph Neural Network (GNN) surrogate that predicts plantar stress in real time under quasi-static loading. The authors report a GNN-vs-Abaqus agreement of R²=0.942 on 500 foot meshes, a surrogate-predicted peak-pressure reduction of 34.7% over parametric CAD on the Male 18–40 cohort, a fit error of 0.42 mm, and an exploratory two-week VAS observation (n=12, 6.4→2.1). The paper is explicitly framed as a design-support prototype, not a clinically validated device, with human-in-the-loop review and regulatory clearance explicitly disclaimed.","tokens_in":17580,"tokens_out":5292,"duration_ms":56881,"significance":"If the surrogate-based performance gains survive independent FEA verification, this work would represent a useful contribution to automated orthotic design: it offers a reproducible open dataset (PicoFoot-5K), open-sourced code and weights, and a surrogate that is externally benchmarked against Abaqus. The authors are commendably transparent about the quasi-static scope, the exploratory nature of the clinical observation, and the need for FEA verification before fabrication. However, the central quantitative claim—the 34.7% peak-pressure reduction—is currently supported only by the same GNN that is used to select the designs, which is a significant validation gap.","major_comments":[{"comment":"The headline 34.7% peak-pressure reduction is computed with the same GNN surrogate that is used to select the optimized designs in the closed loop (Algorithm 1, steps 4–11; Eqs. (2)–(4)). The paper’s own error analysis (§4.1.4) states that the GNN underestimates peak stress in high-curvature regions by 2–4%. If the optimized TANS-FO lattice orthoses contain more high-curvature features (e.g., density gradients, strut intersections) than the generic test meshes, the surrogate’s systematic under-prediction would disproportionately reduce the predicted peak pressure for TANS-FO designs, potentially inflating the reported 13.3-point improvement over the parametric CAD baseline. No final optimized design is re-run through Abaqus or experimentally measured; the limitation section (§5) states that final designs should be FEA-verified before fabrication, but this verification is not performed fo","section":"§4.1.4, Table 5, Algorithm 1, Eqs. (2)–(4)"},{"comment":"The R²=0.942 external validation of the GNN surrogate is carried out on 'an independent test mesh set of 500 irregular foot models' (§4.1.2), not on the distribution of TANS-FO-generated orthoses (lattice structures with spatially varying density). Thus the surrogate’s accuracy in the actual design space of the pipeline is not established. To support the central claim, the authors need to validate the surrogate on a holdout set of TANS-FO outputs using Abaqus ground truth, or at least demonstrate that the 500 test meshes are representative of the optimized designs in terms of curvature, lattice density, and mesh topology. Without this, the external benchmark does not transfer to the reported 34.7% reduction.","section":"§4.1.2, §4.1.4"},{"comment":"The ablation study compares TANS-FO with and without GNN feedback, but the 'Peak Press. Red.' column is, per the table note and §5, surrogate-predicted. If the same GNN is used to evaluate both the full system and the 'w/o GNN' variant, the ablation may be circular: the improvement attributed to the GNN feedback could partly reflect the surrogate reporting lower pressures for designs it selected itself, given its documented high-curvature underestimation. The authors should clarify which evaluator is used for each row and, ideally, re-evaluate the ablated variants with independent FEA to confirm the 34.7% vs 21.4% difference is not an artifact of the surrogate’s bias.","section":"§4.2, Table 4"},{"comment":"The 'semantic-physics alignment' loss L_align measures a learned distance between a text-embedding projection (F_proj) and a learned function of GNN stress outputs (G_attn). This is an embedding-space alignment, not a physics-based constraint; the phrase 'rigorous semantic-physics alignment' overstates the physical grounding. The GNN surrogate itself is trained on Abaqus data, but the alignment loss itself does not encode any continuum-mechanics law. This is a terminology/scope issue rather than a fatal flaw, but the claims should be calibrated to what is actually achieved.","section":"§3.6, Eq. (4)"}],"minor_comments":[{"comment":"The paragraph 'GNN error spatial pattern' appears twice verbatim. Please remove the duplicate.","section":"§4.1.4"},{"comment":"The 'Time (s)' column compares computation-only time for TANS-FO with time that includes physical fabrication steps for manual/CAD baselines. This makes the time comparison misleading; the table note acknowledges it, but the column header should be labeled more explicitly (e.g., 'compute time' vs 'total workflow time').","section":"Table 5 note"},{"comment":"The text states that meshes were 'voxelized at a resolution of 1.0 mm,' but the GNN operates on an irregular graph. Clarify whether voxelization is used only for preprocessing/augmentation or as an intermediate representation that conflicts with the graph abstraction.","section":"§4.1.2"},{"comment":"The error figures are stated as 'up to 4.8%' and '2–4%' underestimation in high-curvature regions. These should be reconciled and reported with confidence intervals, as the directional bias is central to the validation concern.","section":"§4.1.4"},{"comment":"Several references are incomplete (e.g., ref. 22 lacks volume/page numbers; refs. 32 and 33 have inconsistent formatting). Please check against the journal style.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is transparent and provides a publicly available dataset and code, which is a strength. The central concern is the circular use of the GNN surrogate as both optimizer and evaluator for the headline 34.7% figure. This is fixable within the manuscript's scope by re-running final designs through Abaqus (or by substantially hedging the claim and quantifying the bias). I also note that some references appear questionable (e.g., ref. 22 with missing details); the editor may want to verify citation integrity. The paper is otherwise a reasonable research-prototype description, but the current validation gap prevents acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe two things you should know about this paper before reading it: the integrated text-to-density closed loop with a GNN stress surrogate is real and new, and the authors are unusually honest about what it is not—a clinically validated device. The headline 34.7% peak-pressure reduction is explicitly labeled as surrogate-predicted, and the paper repeatedly warns against conflating it with measured outcomes.\n\nThe genuinely new contribution is the coupling of clinical text embeddings to lattice-density fields via cross-attention, with a GNN surrogate inside the optimization loop for real-time biomechanical feedback. The individual components are standard—GAT message passing, DeepSDF losses, cross-attention—but the integrated system for orthotic design, plus the released PicoFoot-5K dataset and open-source code, make it a useful engineering contribution. The R²=0.942 agreement with Abaqus on 500 irregular foot meshes is a legitimate external benchmark for the surrogate itself.\n\nThe soft spots are in proportion. The main one is that the same GNN both selects designs and evaluates the 34.7% improvement. That confounding is real: the R² is measured on foot meshes, not necessarily on the optimized lattice geometries, and the paper itself notes the GNN underestimates peak stress in high-curvature regions by 2–4%. If the generated lattices have more high-curvature features than the validation set, the reported improvement could be inflated. The paper states final designs should be verified by high-fidelity FEA, but doesn't do it for the headline result. That's a legitimate criticism, but not a fatal one for a prototype—the authors have clearly framed the work as a design-support tool, and the limitation is acknowledged in the text. One small editorial issue: section 4.1.4 has a duplicated paragraph on GNN error spatial pattern.\n\nWho gets value: researchers working on generative medical design, orthotic CAD automation, or surrogate-guided optimization. It's not the paper to cite for clinical efficacy or dynamic gait, and the authors don't claim it is.\n\nI'd send it to peer review. A good referee should ask them to re-run a subset of final optimized designs through Abaqus or provide a sensitivity analysis showing the 34.7% holds under plausible surrogate error. That's an actionable request, not a rejection.","headline":"Honest, well-scoped engineering prototype; the 34.7% headline is surrogate-on-surrogate and should be treated as provisional until an independent FEA run confirms it.","tokens_in":18087,"tokens_out":2986,"would_cite":false,"duration_ms":31037,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A closed-loop generative pipeline maps clinical text to 3D printable foot orthoses, using a GNN stress surrogate to predict a 34.7% peak-pressure reduction over parametric CAD.","keywords":["foot orthoses","generative design","graph neural networks","semantic-physics alignment","lattice structures","plantar pressure","finite element surrogate","closed-loop design"],"falsifier":"Run the final TANS-FO STL designs through Abaqus (or instrumented in-shoe pressure mapping) on the same 50-subject male 18–40 cohort and compare peak pressure against the parametric CAD baseline; if the observed reduction falls far below 34.7% or the fit error exceeds ~0.42 mm, the central claim fails. The paper does not perform this check.","tokens_in":17051,"feed_emoji":"🦶","tokens_out":6475,"duration_ms":63038,"temperature":0.7,"pith_summary":"The paper sets out to solve a 'semantic-physics misalignment' in foot orthosis design: clinical prescriptions like 'offload the first metatarsal head' are not deterministically mapped to 3D geometry, so design remains manual and slow. It presents TANS-FO, a modular pipeline in which a text-aligned neural surrogate projects clinical text onto a lattice-density field, and a graph neural network predicts plantar stress in real time, closing the loop. On the Male 18–40 cohort, the system is reported to achieve a surrogate-predicted peak-pressure reduction of 34.7% over parametric CAD, a fit error of 0.42 mm, and a design-to-validation time of about 52 seconds. The GNN surrogate matches an Abaqus reference solver with R²=0.94 under quasi-static loading. The authors are explicit that these are computational surrogate results, not clinical outcomes, and the device is a research prototype.","feed_headline":"AI pipeline predicts 34.7% peak-pressure cut for custom insoles","feed_subtitle":"Clinical text becomes a 3D-printed insole in under a minute, with a GNN surrogate predicting the offload — pending FEA verification.","key_machinery":"The Text-Aligned Neural Surrogate (TANS) is the central object: a cross-attention mechanism that projects a 512-dimensional clinical-text embedding (from a BioBERT-based parser) onto a 128-dimensional per-node field over the orthotic mesh, with dimensions 1–64 encoding offloading intensity, 65–96 support stiffness, and 97–128 boundary smoothness, which together determine local Gyroid lattice density. The second pillar is the GNN surrogate — a 4-layer graph attention network with 4 heads, operating on ~12,000-node mesh graphs — which predicts node-level stress in real time, substituting for FEA. Its output feeds back through an alignment loss to modulate the embedding, closing the loop; a hum","core_discovery":"The central claim is that semantic-physics alignment can be achieved by a cross-attention layer that maps clinical-text embeddings to a 128-dimensional nodal field encoding offloading intensity, support stiffness, and boundary smoothness, which in turn modulates Gyroid lattice relative density. With a graph attention network providing near-instant stress predictions, the generative loop can be closed in minutes, and the resulting designs show a surrogate-predicted 34.7% peak-pressure reduction over parametric CAD on the male 18–40 cohort, with a fit error of 0.42 mm and GNN–Abaqus agreement of R²=0.94. If correct, this means the manual, expert-dependent translation from prescription to ortho","pith_inferences":["If an independent Abaqus re-run of the final designs confirms a comparable offloading gain, the same closed-loop architecture could be transplanted to other patient-specific devices — ankle-foot orthoses, insoles for offloading diabetic ulcers — where the text-to-geometry mapping problem is analogous.","The reported 34.7% is a surrogate-predicted number; the paper does not re-run final designs through FEA. Until that check is done, readers should treat the figure as an upper-bound estimate of the real-world offloading gain.","The GNN's systematic 2–4% underestimation of peak stress at high-curvature regions suggests that for high-risk patients (e.g., diabetic foot), the surrogate should be paired with a safety margin or mandatory FEA verification; the paper's human-in-the-loop gate is a start but not a quantitative safety factor.","The PicoFoot-5K dataset omits children under 15, BMI>35, and severe deformities — precisely the populations most likely to need custom orthoses. Applying the pipeline to those groups would test whether the alignment generalizes beyond the training distribution."],"forward_implications":["Design-to-validation time per orthosis drops from hours or days to roughly 52 seconds on a modern GPU, allowing rapid exploration of alternative clinical directives.","Adding GNN feedback improved fit error on the male 18–40 cohort from 1.24 mm to 0.42 mm and peak-pressure reduction from 21.4% to 34.7%, showing the closed-loop surrogate meaningfully guides the generator.","The random-text ablation shows the alignment is not an artifact of anatomical priors alone: corrupting the clinical text degrades performance to near baseline, so the semantic signal is doing real work.","The pipeline outputs a manufacturing-ready STL of a TPU lattice insole, so a fabricated-device workflow is, in principle, a direct next step, subject to the human review gate.","Because the GNN surrogate runs in seconds, final designs can still be routed to a high-fidelity FEA for verification, separating exploration from certification."],"fun_headline_variants":["AI text-to-insole pipeline: 34.7% peak-pressure cut, predicted","Closed-loop design: clinical text to 3D insole in minutes","GNN surrogate predicts 34.7% peak-pressure drop in custom insoles","Semantic-physics alignment predicts 34.7% peak-pressure cut","Research prototype: AI-generated insoles show 34.7% offload"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The final performance figures rest on the assumption that the GNN surrogate's predicted stress fields are accurate enough to certify the 34.7% peak-pressure reduction; the optimized designs were not re-run through full FEA, and the GNN is known to underestimate peak stress in high-curvature regions by 2–4%.","fun_headline_variants_meta":{"raw":{"variants":["AI text-to-insole pipeline: 34.7% peak-pressure cut, predicted","Closed-loop design: clinical text to 3D insole in minutes","GNN surrogate predicts 34.7% peak-pressure drop in custom insoles","Semantic-physics alignment predicts 34.7% peak-pressure cut","Research prototype: AI-generated insoles show 34.7% offload"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000788,"raw_usage":{"total_tokens":3365,"prompt_tokens":851,"completion_tokens":2514,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":595,"completion_tokens_details":{"reasoning_tokens":2412}},"tokens_in":595,"tokens_out":2514,"duration_ms":17205,"temperature":1.0,"reasoning_tokens":2412,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T20:23:09.947800+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the final TANS-FO STL designs through Abaqus (or instrumented in-shoe pressure mapping) on the same 50-subject male 18–40 cohort and compare peak pressure against the parametric CAD baseline; if the observed reduction falls far below 34.7% or the fit error exceeds ~0.42 mm, the central claim fails. The paper does not perform this check.","supporting_citations":[],"review_version":1}