{"id":"e0171104-583c-481d-bad8-43ac8ff25429","arxiv_id":"2607.26234","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"FORGE-SIM directly optimizes a six-patch B-spline boundary representation from sparse images and demonstrates isogeometric heat and modal simulations on the reconstructed model.","lead":"FORGE-SIM reconstructs a watertight, multi-patch B-spline CAD model directly from sparse posed photos, then runs thermal and modal simulations on that same model. It is a step toward automatic image-to-simulation digital twins, bypassing manual CAD reconstruction and mesh repair.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Modal results on real scenes are undermined by the outer-surface-as-mid-surface shell approximation, which is not small for the Lion scene (d/L≈6.7%) and is unquantified.","rationale":"The reader identified the mid-surface approximation as part of the weakest assumption, along with manual scale/material parameters. I focus on the mid-surface offset because it is a systematic modeling error that directly affects the modal analysis claim and is explicitly acknowledged but dismissed with a justification that is numerically false for the Lion scene. The synthetic validation uses thin shells where the offset is negligible, so it cannot detect this error; the real-world modal results are qualitative only. This does not invalidate the entire paper—thermal simulation is unaffected, and the mid-surface issue is addressable by offsetting the spline—so the CONDITIONAL verdict stands, but the condition should explicitly require quantifying or correcting this approximation. Manual scale/material parameters are acknowledged limitations and less load-bearing because they are user inputs that could be refined; the mid-surface error is inherent to the claimed pipeline as presented. I agree with the reader's overall CONDITIONAL assessment, hence UNCHANGED, but with partial agreement on the specific weakest assumption.","tokens_in":49188,"tokens_out":7277,"duration_ms":68192,"concrete_test":"Using the synthetic Suzanne scene (known geometry and materials), run the FEM mid-surface modal solve on the ground-truth mesh offset inward by d/2, and compare against the FEM solve on the outer surface used in the paper, for thickness values d = 0.002, 0.005, 0.01, 0.02 m (matching the Lion's d/L ratio). If the median eigenfrequency shift at d/L≈6.7% exceeds ~3%, the real-world modal results in Fig. 4 are not reliable. Alternatively, re-run the IGA modal solve on the Lion reconstruction with the spline offset inward by d/2 along the normal and compare to the un-offset result; a shift of >5% in the reported frequencies would confirm the concern.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that FORGE-SIM produces models 'of sufficiently high quality to enable ... modal analysis' is quantitatively validated only on synthetic thin shells where the reconstructed outer surface is used as the Reissner-Mindlin mid-surface. Supp. Sec. 6.2 acknowledges this approximation, claiming the thickness is 'small compared to the size of the manifold.' For the Lion real scene, d=0.02 m and the bounding-box scale L=0.3 m (Table 4), so d/L≈6.7%, which is not small. For a shell, the bending and membrane stiffness depend on the mid-surface geometry; using the outer surface as the mid-surface changes the effective radii of curvature by d/2. Eigenfrequencies of curved shells scale with curvature (roughly f ∝ sqrt(E d^2/(ρ R^4)) for bending-dominated modes, so a ~6.7% change in effective radius can shift frequencies by several percent to tens of percent. The synthetic end-to-end modal validation (Table 1) uses d=0.002 m on objects of scale ~0.3 m (d/L≈0.7%), where the offset is negligible. Therefore the synthetic agreement (2.5–8.7% frequency error) cannot support the real-world modal numbers in Fig. 4, which have no ground truth and are only described as 'plausible.' The paper's own statement in Supp. 6.2 is false for the Lion scene, and the error is unquantified.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes FORGE-SIM, a pipeline that reconstructs a six-patch cubic B-spline boundary representation and associated scalar fields (thermal, semantic) directly from sparse posed RGB/thermal images, and then uses the same spline basis for isogeometric heat and Reissner-Mindlin shell modal analyses. The core claims are that this yields watertight, smooth, simulation-ready geometry without manual geometry authoring, repair, or meshing, and that the resulting models are of sufficiently high quality for thermal simulation and modal analysis. Evidence includes synthetic end-to-end comparisons against independent FEM solves on ground-truth meshes (1.1--6.9% heat-field error, 2.5--8.7% median eigenfrequency error, MAC 0.32--0.53), solver-agreement checks on identical geometry (0.3--0.5%), reconstruction/NVS benchmarks, and qualitative real-world thermal and modal demonstrations.","tokens_in":49564,"tokens_out":5709,"duration_ms":60428,"significance":"If the claims hold, this is a valuable contribution at the vision--simulation interface: it demonstrates that a differentiable multi-patch B-spline representation can be optimized from images and consumed directly by IGA solvers, avoiding the usual CAD-repair/meshing pipeline. The paper is strong in several concrete ways: it reports a genuine two-solver agreement study on identical geometry, uses multiple complementary metrics, transparently reports field-transfer artifacts and floor corrections, and provides detailed derivations of the heat and shell formulations. The synthetic end-to-end validation is credible and shows a monotonic relationship between reconstruction quality and simulation error. The main weakness is the real-world modal validation, which is only qualitative and rests on the outer-surface-as-mid-surface shell approximation and manually fixed scale/material parameters; this supports a major revision rather than acceptance.","major_comments":[{"comment":"The shell reduction treats the reconstructed outer surface as the Reissner-Mindlin mid-surface. The paper justifies this by saying \"the thickness is small compared to the size of the manifold,\" but for the Lion scene Table 4/5 gives d = 0.02 m and L = 0.3 m, i.e. d/L ≈ 6.7%, which is not small. The synthetic modal validation (Table 1, Fig. S3) uses d = 0.002 m on objects of scale roughly 0.3 m (d/L ≈ 0.7%), so it does not probe this regime. Since shell stiffness and mass depend on mid-surface curvature, using the outer surface shifts effective radii by d/2 and can change eigenfrequencies by several percent or more. This error is unquantified, so the real-scene modal frequencies in Fig. 4 are not supported by the synthetic evidence. The authors should either reconstruct/offset a proper mid-surface, add a quantitative sensitivity analysis for d/L, or explicitly restrict the modal claim to","section":"Supplementary Sec. 6.2; Tables 4--5; Sec. 2.1.2"},{"comment":"The real-world modal results depend on a manually estimated metric scale factor L and hand-assigned material parameters E, nu, rho, and d. The paper states these are estimated to give \"physically plausible\" frequencies, but no ground truth is available and the method for estimating L is not described in Sec. 5.6 despite being referenced there. A sensitivity analysis over L and the material parameters is needed to show that the reported frequencies and the \"sufficiently high quality for modal analysis\" claim are robust; otherwise the real-scene modal demonstration is only a plausibility statement, not validation.","section":"Sec. 2.1.2, Sec. 5.6, Table 5, Fig. 4"},{"comment":"The real-world heat simulations are described only qualitatively (\"physically coherent thermal diffusion\"). In contrast to the synthetic end-to-end heat validation, which is quantitative, the real-scene thermal results have no measured or reference temperature field to compare against. The low assigned temperature of 0.1 in unseen regions (Sec. 5.1.3) may dominate the visual result. The paper should either provide a quantitative proxy for real-scene thermal fidelity or clearly state that real-scene heat results are demonstrations of numerical stability rather than accuracy.","section":"Sec. 2.1.1, Fig. 4"}],"minor_comments":[{"comment":"The abstract says the pipeline runs \"without manual intervention,\" but Sec. 1 later qualifies that per-scene inputs include the refinement schedule and physical material parameters, and Sec. 3 notes that metric scale is fixed via an estimated calibration factor. Please harmonize the wording to avoid overclaiming full autonomy.","section":"Abstract / Sec. 1"},{"comment":"The floor-corrected field discrepancy is computed as sqrt(cross^2 - floor^2). This assumes independence between the cross-field discrepancy and the floor. Since they are computed from the same transfer process, the independence assumption should be justified or the raw values and floors should be used as the primary reporting.","section":"Table 1 / Supp. Sec. 7.6"},{"comment":"The caption should explicitly state that real-world modal frequencies depend on the manually estimated scale and material parameters, and that the results are illustrative, not validated against ground truth.","section":"Fig. 4 caption"},{"comment":"The mode numbering starts at 6 because modes 0--5 are rigid-body modes. This should be stated in the main text or caption so readers do not misread the fundamental frequency as mode 0.","section":"Table S4 caption"},{"comment":"The sentence \"We omit the results presented in the main text concerning the relative heat errors...\" is confusing; it likely means the section provides supplementary detail rather than omitting results. Please rephrase.","section":"Supp. Sec. 7.6"},{"comment":"The scene label \"Building A Spring\" appears to be a typo and should be \"Building A\".","section":"Table S2"}],"recommendation":"major_revision","confidential_remarks":"The paper is technically substantial and the synthetic validation is convincing, but the real-scene modal claim is the load-bearing part of the abstract and it is not quantitatively supported. I would not reject: the issue can be addressed by additional sensitivity studies and/or by carefully narrowing the claims. The authors should also provide the scale-estimation procedure and consider reconstructing a mid-surface for thicker shells."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou should read this one if you care about connecting image-based reconstruction to simulation. FORGE-SIM optimizes a six-patch cubic B-spline B-rep directly from sparse posed images, then runs IGA heat and Reissner-Mindlin modal solves on that same geometry. That integration is new, as far as I know: previous work optimized triangle meshes or rendered parametric geometry, but nobody has gone end-to-end from images to IGA simulation on the reconstructed spline without manual CAD authoring.\n\nThe synthetic validation is the strongest part. They compare IGA on the reconstruction against an independent FEM solver on ground-truth meshes, and they first check solver agreement on identical geometry (0.3–0.5% across the eight scenes). The end-to-end errors (1.1–6.9% heat field, 2.5–8.7% median eigenfrequency) track reconstruction error, not solver noise. They also report honest limitations: genus-0 only, over-smoothing, manual scale and material parameters for real scenes. That's a solid, well-written engineering paper.\n\nThe soft spots are real but not fatal. The headline field-discrepancy metric is floor-corrected (they subtract a round-trip transfer artifact), which is transparent but does make the main number a bit friendlier. No code is shipped yet, only a promise. And the modal results on real scenes use the reconstructed outer surface as the shell mid-surface; for the Lion, thickness/scale is ~7%, which is not in the regime where the synthetic d/L≈0.7% validation applies. The paper acknowledges this in the supplement but doesn't quantify the error. Since the real-scene modal results are only qualitative ('plausible'), this undermines the 'enables modal analysis' claim only weakly, but it should be tightened.\n\nThe 'comparable or superior' NVS framing is also a bit strong—on real scenes, 3DGS often has higher PSNR/SSIM. That said, NVS is not the paper's point.\n\nWho should read it: anyone building digital twins from images, or working on spline-based inverse rendering. It deserves a serious referee: it's a conditional accept with requested revisions, not a desk reject.\n\nMy bottom line: engage with it, ask for code and for an error analysis of the mid-surface approximation, but don't let the real-world modal caveat obscure the genuinely new pipeline.\n\nLet me know your thoughts,\n[Your name]","headline":"Genuinely new end-to-end pipeline from sparse images to IGA simulation via spline B-reps, with credible synthetic validation; real-world modal results lean on an unquantified mid-surface approximation and the NVS 'comparable or superior' claim is a bit strong.","tokens_in":50049,"tokens_out":4043,"would_cite":true,"duration_ms":38298,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["65D17","65N30"],"pacs":[],"model":"deepseek-v4-flash","headline":"FORGE-SIM claims the first direct image-to-simulation pipeline: a watertight, smooth boundary representation optimized from sparse photos, on which isogeometric heat and modal solves match finite-element ground truth within 1–9%.","keywords":["boundary representation","isogeometric analysis","B-spline","sparse view reconstruction","thermal simulation","modal analysis","Reissner-Mindlin shell","watertight geometry"],"falsifier":"Take one of the synthetic test objects (known ground truth), reconstruct at native scale, but deliberately set the shell thickness to a non-negligible fraction of the curvature radius (e.g., d ≈ 0.1R) and compare the IGA modal spectrum against a converged 3D FEM solve on the true volume geometry; if the frequency error grows far above the reported 2.5–8.7% band, the mid-surface reduction — not geometry recovery — is the binding assumption. Alternatively, mis-specify the per-scene scale factor by 10% and observe an approximately proportional shift in all reported eigenfrequencies.","tokens_in":49058,"feed_emoji":"📐","tokens_out":5788,"duration_ms":56201,"temperature":0.7,"pith_summary":"The paper's central claim is that a six-patch cubic B-spline boundary representation can be optimized directly from sparse posed RGB images, yielding a watertight, smooth geometry that a physics solver can consume without manual CAD repair or meshing. The authors validate this by running heat-flow and Reissner–Mindlin modal analysis with isogeometric analysis (IGA) on the reconstructed spline and comparing against finite-element simulation on the ground-truth mesh: temperature fields agree within 1.1–6.9% and median eigenfrequencies within 2.5–8.7%. They also show that observation-derived fields—thermal state, semantic material classes—can be inpainted in the same spline basis so the solver uses them directly, and the model can be exported as standard CAD formats. A sympathetic reader would care because it removes the long-standing manual bottleneck between image capture and numerical simulation, enabling as-built digital twins for thermal and structural assessment.","feed_headline":"From sparse photos to thermal and modal simulation in one pass","feed_subtitle":"Heat and vibration solvers run directly on the reconstructed spline—no manual CAD modeling or meshing in between.","key_machinery":"The load-bearing object is a six-patch cubed-sphere B-spline boundary representation: a closed, watertight surface built from six tensor-product cubic B-splines with shared edge/corner control points and near-G1 continuity enforced by a normal-alignment loss plus a Willmore-energy curvature penalty. The shape is optimized by alternating an L-BFGS pass over the global control points with a texture pass, using gradients from a differentiable renderer that evaluates a tessellation of the spline rather than the spline itself; a coarse-to-fine h-refinement schedule increases knot spans as optimization proceeds. The same B-spline basis is then used to represent inpainted thermal and material field","core_discovery":"The discovery is that optimizing the spline representation itself, rather than reconstructing a mesh or implicit field and converting it, produces geometry that is simultaneously accurate for novel-view synthesis and numerically valid for simulation. FORGE-SIM is, to the authors' knowledge, the first framework to demonstrate shape optimization of a multi-patch B-spline surface directly from images and to run simulation on the resulting reconstruction without manual geometry authoring, repair, or meshing. The key quantitative result is the end-to-end agreement: for the best-reconstructed shapes, IGA on the recovered B-spline boundary representation reproduces FEM heat fields to within about 1","pith_inferences":["If the thin-shell reduction is pushed beyond its regime—for instance, a plush object with thickness 0.02 m against a 0.3 m scale, or surfaces with sharp curvature—the modal frequencies will be biased by treating the reconstructed outer surface as the mid-surface; a direct 3D FEM comparison would reveal this bias, which the paper explicitly accepts in the supplementary.","The per-scene requirement of a hand-set metric scale and material constants means the pipeline is not yet fully autonomous; recovering absolute scale from thermal/RGB fusion or known references would close the loop.","The cubed-sphere six-patch topology restricts reconstructions to genus-0 objects; extending to adaptive spline families such as T-splines or hierarchical B-splines—as the paper itself suggests—is a natural test of whether the parameterization, not the optimization, is the ceiling.","A testable extension: run the pipeline on a synthetic object with known ground-truth volume and deliberately varied thickness-to-curvature ratios to calibrate the range of validity of the mid-surface approximation."],"forward_implications":["If correct, image capture becomes a viable entry point for simulation-driven workflows such as structural health monitoring and thermal inspection, skipping the manual CAD and meshing pipeline.","The simulation-fidelity error tracks reconstruction quality, implying that improving geometric reconstruction—not solver refinement—is the next lever for end-to-end accuracy.","Because the same spline basis serves both the renderer and the IGA solver, any field that can be rendered (temperature, material class, and so on) can be baked into the model and consumed by a PDE solve without format conversion.","The framework's compatibility with both IGA and conventional finite-element workflows (via standard CAD export) means it can slot into existing engineering toolchains rather than requiring new solvers."],"fun_headline_variants":["Sparse photos to simulation-ready splines in one shot","Skip the mesh: spline geometry straight from images","Image-based reconstruction that runs thermal and modal solvers natively","Direct spline optimization from sparse RGB images for simulation","From a few photos to heat and vibration analysis without CAD"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The reconstructed outer surface is treated as the mid-surface of a thin shell whose thickness is small compared with the object's size and curvature, and the metric scale plus material constants are supplied by hand rather than recovered; if the thickness is not small (e.g., Lion with d = 0.02 m vs scale L = 0.3 m) or the scale is wrong, the modal frequencies will shift accordingly.","fun_headline_variants_meta":{"raw":{"variants":["Sparse photos to simulation-ready splines in one shot","Skip the mesh: spline geometry straight from images","Image-based reconstruction that runs thermal and modal solvers natively","Direct spline optimization from sparse RGB images for simulation","From a few photos to heat and vibration analysis without CAD"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000178,"raw_usage":{"total_tokens":1123,"prompt_tokens":725,"completion_tokens":398,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":469,"completion_tokens_details":{"reasoning_tokens":318}},"tokens_in":469,"tokens_out":398,"duration_ms":4037,"temperature":1.0,"reasoning_tokens":318,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T00:24:30.432992+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take one of the synthetic test objects (known ground truth), reconstruct at native scale, but deliberately set the shell thickness to a non-negligible fraction of the curvature radius (e.g., d ≈ 0.1R) and compare the IGA modal spectrum against a converged 3D FEM solve on the true volume geometry; if the frequency error grows far above the reported 2.5–8.7% band, the mid-surface reduction — not geometry recovery — is the binding assumption. Alternatively, mis-specify the per-scene scale factor by 10% and observe an approximately proportional shift in all reported eigenfrequencies.","supporting_citations":[],"review_version":1}