{"id":"4cbf1713-c8ad-4cda-8d73-9b9b435f4a44","arxiv_id":"2507.21350","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"DEM-NeRF is a pipeline that builds a 3D model from multi-view images with NeRF and predicts hyperelastic deformation with a deep energy method, but it presents no accuracy validation.","lead":"DEM-NeRF combines a neural radiance field with a physics-informed deep energy network to reconstruct an object from images and predict its elastic deformation. The paper claims real-time, accurate simulation, but reports no quantitative comparison against ground truth or prior methods.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No quantitative accuracy evaluation: the claim of reliable deformation predictions rests on figures and a runtime table, with no displacement error or ground-truth comparison; the time-independent FDEM in Eq. (1) also cannot represent the claimed hyperelastodynamics.","rationale":"The reader's rejection is justified. I looked for a point where the central claim could survive despite missing metrics, but none exists: runtime speed without accuracy cannot establish reliable deformation prediction. The strongest condition for the claim is equivalence between the DEM output and the actual deformation field, and this is never checked. I partially agree with the reader's weakest assumption: the unreported constitutive and load parameters are a serious reproducibility and correctness risk, but the broader absence of any quantitative evaluation is more load-bearing because even with fully reported parameters one cannot tell whether the method works. The time-independent FDEM also contradicts the dynamic language in the abstract, reinforcing the concern. A single displacement-error comparison against the shown ground truth would settle the issue. Since this concern aligns with the existing REJECT verdict, no verdict adjustment is needed.","tokens_in":10075,"tokens_out":3679,"duration_ms":48900,"concrete_test":"Reconstruct the T-bar experiment with the reported (or newly disclosed) values of Young's modulus, Poisson's ratio, torque magnitude, and fixed boundary setup, solve Eq. (14) with the DEM network, and compute the per-vertex mean and maximum Euclidean displacement error between the predicted deformed surface and the ground-truth deformed T-bar shown in Figure 6, normalized by the object's bounding-box size. Compare the same error for PAC-NeRF and PIE-NeRF. If the error is not small or is not reported, the central reliability claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of reliable deformation predictions in hyperelastodynamics and loading scenarios is not supported by any quantitative accuracy result. Section IV-A and Table IV-A report only computational and rendering times; no displacement error, strain error, or comparison against the ground-truth deformed T-bar in Figure 6, or against PAC-NeRF or PIE-NeRF, is given. The predicted field can be wrong even if the energy minimization in Eq. (9) or Eq. (14) is exact, because the model assumes a compressible Neo-Hookean quasi-static energy with fixed boundary and traction loads, while the material parameters E and nu and the torque magnitudes are never reported. Moreover, FDEM in Eq. (1) maps only (x, y, z) to u with no time input, so the claimed spatiotemporal or hyperelastodynamic capability is not present in the stated architecture. The minimum condition for the central claim to hold, demonstrated accuracy of the predicted displacement field against measured deformation, is entirely missing.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"DEM-NeRF proposes a neuro-symbolic pipeline that reconstructs a 3D solid from multi-view images with NGP-NeRF, samples it into particles, and predicts deformation with a Deep Energy Method network minimizing a Neo-Hookean strain-energy functional. The paper claims high-quality reconstructions and reliable deformation predictions in hyperelastodynamic and loading scenarios, with a runtime advantage over PAC-NeRF and FEM. The experimental section applies the pipeline to a T-bar under torque, but reports only computational and rendering times together with qualitative renderings; no quantitative accuracy metrics, material parameters, or load values are provided.","tokens_in":10229,"tokens_out":4444,"duration_ms":53486,"significance":"If substantiated, the direction is significant: combining NeRF-based geometry acquisition with an energy-constrained network could enable real-time, label-free deformation simulation for digital twins, soft robotics, and biomechanics. The paper has two genuine strengths: the DEM loss is an explicit physical energy minimization rather than a fitted surrogate, and the runtime figures (about 20 s DEM training plus 1 s prediction) are potentially attractive. However, the central claims of reliable prediction and hyperelastodynamic capability are not backed by any quantitative evaluation in the current manuscript. The evidence is limited to images and a timing table, so the contribution, as presented, is not yet established.","major_comments":[{"comment":"The paper reports only computational and rendering times; there is no displacement error, strain error, or any quantitative comparison of the predicted T-bar deformation against the ground-truth deformed shape in Figure 6 or against FEM, PAC-NeRF, or PIE-NeRF. Since the abstract and Section I claim \"reliable deformation predictions\" and \"high-quality reconstructions,\" the absence of any accuracy metric leaves the central contribution unsubstantiated. The lack of ablations for the boundary weight Wu, particle count, and sampling method compounds this problem.","section":"Section IV-A, Table IV-A"},{"comment":"The total potential energy Π in Eq. (9) is minimized as a quasi-static equilibrium, with no kinetic-energy or inertial term, and FDEM in Eq. (1) maps (x,y,z) directly to u with no time coordinate. Thus the architecture cannot represent the \"hyperelastodynamics\" or the \"spatiotemporal representation\" claimed in the abstract and Impact Statement. A model that assumes quasi-static Neo-Hookean equilibrium is not a dynamical model, and the claim of dynamic prediction is not supported by the stated equations.","section":"Section III-D, Eqs. (1), (9)-(14)"},{"comment":"The Young's modulus E, Poisson's ratio ν, torque magnitudes, body force fb, and boundary weight Wu are never reported. These quantities determine the predicted displacement field through the energy in Eq. (9); without them the experiment is not reproducible, and if they were hand-tuned to match the observed deformation, that tuning should be disclosed. The manuscript provides no evidence that the assumed compressible Neo-Hookean model with the chosen loads corresponds to the physical T-bar experiment.","section":"Section IV-A, Eqs. (7)-(9)"}],"minor_comments":[{"comment":"References [4] and [26] are the same Karniadakis et al. paper, and references [5] and [25] are the same Raissi et al. paper; the duplicate entries should be consolidated.","section":"References"},{"comment":"The text contains \"particle-could representation,\" which should be \"particle-cloud representation.\"","section":"Section III-A"},{"comment":"The traction boundary set is written with \"NP i\" and the traction term uses t, while Eq. (9) writes \\bar{t}; these notations should be made consistent and the typo corrected.","section":"Eq. (14)"},{"comment":"The manuscript mentions \"I.C./B.C.\" but initial conditions are never defined or used in the quasi-static formulation; the abbreviation should be removed or the initial conditions specified.","section":"Section IV-A"},{"comment":"The performance table has no caption and no table number; it should be given a proper number and a caption defining the hardware, the meaning of training time, and which stages are included in each reported number.","section":"Section IV-A, Table"},{"comment":"The comparison of random, Poisson disc, and uniform mesh sampling in Figure 7 does not state the particle spacing or radius used for each method; specifying these parameters would make the qualitative comparison meaningful.","section":"Section IV-B"}],"recommendation":"reject","confidential_remarks":"I see no evidence of self-citation inflation or circularity: the DEM loss is a genuine energy minimization and the method is defined independently of the target deformation. The primary problem is evidentiary and architectural. The paper's headline claims are not tested quantitatively, and the quasi-static, time-independent formulation contradicts the claimed hyperelastodynamic capability. These issues are load-bearing and would require substantial new experiments and likely a reformulation of the model before the claims could be evaluated."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper takes the PAC-NeRF/PIE-NeRF recipe—reconstruct a solid with NeRF, then simulate deformation with a physics-based neural solver—and swaps MPM for the deep energy method (DEM). That specific combination is new, and the DEM formulation is described cleanly: only first-order derivatives are needed, which is a real computational advantage over strong-form PINNs. If the speed numbers hold, the pipeline is also much faster to train than PAC-NeRF (~20s DEM training vs ~600s MPM), which is a meaningful selling point.\n\nThat's where the credit ends. The paper never demonstrates that the predicted deformation is accurate. The only quantitative results are a runtime table. There is no displacement error, no strain error, no comparison against the ground-truth deformed T-bar in Figure 6, and no accuracy comparison against PAC-NeRF or PIE-NeRF. The central claim of \"reliable deformation predictions\" is simply asserted. This is a load-bearing absence: if the simulation is wrong, the whole pipeline is just a NeRF with a fancy solver.\n\nThe modeling frame also overreaches. The abstract promises hyperelastodynamics, but the FDEM network in Eq. (1) has no time input, and Eq. (9) is a quasi-static energy minimization without inertial terms. The experiment is a static torque load. That's fine as a quasi-static showcase, but the paper should not claim dynamic capability. Relatedly, the material parameters (E, ν) and the torque magnitudes are never reported, so the experiment is not reproducible, and the loss weight W_u is unspecified.\n\nThe novelty is incremental, as the reader notes, but that alone wouldn't sink it. The sinker is the missing evaluation.\n\nGiven the internal contradiction between the stated scope and the actual model, and the lack of any accuracy metric, I don't think this paper is ready for publication. But the core idea—DEM as the physics constraint inside a NeRF pipeline—is worth exploring, and with a proper evaluation (error against ground truth, material parameter reporting, and a time-dependent extension if they want to claim dynamics) it could become a solid contribution. I'd send it to review rather than desk-reject, because the combination is new and the community could benefit from a careful critique. My own verdict on the current version: reject and resubmit after major revision.","headline":"A plausible but unevaluated swap of MPM for DEM inside a NeRF+physics pipeline; the central accuracy claim is unsupported and the dynamics language overreaches.","tokens_in":10796,"tokens_out":2600,"would_cite":false,"duration_ms":30063,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A neural radiance field combined with a physics-driven network reconstructs a 3D object and predicts its elastic deformation directly from images.","keywords":["neuro-symbolic","Neural Radiance Field","Physics-Informed Neural Network","Deep Energy Method","hyperelastic deformation","solid reconstruction","real-time simulation","Neo-Hookean"],"falsifier":"Recreate the T-bar experiment with the same multi-view images and independently measure the actual 3D displacement field (for example with digital image correlation or tracked markers), then compare it with DEM-NeRF's prediction using the paper's unreported material and load parameters; a mismatch beyond rendering noise, or any observable time-dependent motion that the quasi-static energy minimization cannot represent, would undercut the 'reliable deformation prediction' claim.","tokens_in":9838,"feed_emoji":"🧊","tokens_out":5478,"duration_ms":54473,"temperature":0.7,"pith_summary":"The paper proposes DEM-NeRF, a neuro-symbolic pipeline that reconstructs a deformable solid from sparse multi-view images and then predicts its deformation under load, without explicit geometry, labeled data, or specialized hardware. The reconstruction stage uses a neural radiance field to turn images into a 3D particle cloud, and the simulation stage uses a physics-informed network that minimizes the total potential energy of a compressible Neo-Hookean (a standard hyperelastic) solid. The authors claim this runs in about 45 seconds of NeRF training, 20 seconds of DEM training, and roughly one second per prediction, which would make interactive real-time elastic simulation practical. The attraction is that it sidesteps the usual bottleneck of numerical solvers, which require complete geometric models and known boundary conditions.","feed_headline":"NeRF plus physics net predicts solid deformation in one second.","feed_subtitle":"A neuro-symbolic pipeline reconstructs a 3D object from photos and simulates its bending without geometry or labeled data.","key_machinery":"The load-bearing mechanism is the Deep Energy Method, in which a neural network acts as a global shape function mapping each particle's rest position $\\mathbf{X}$ to a displacement $\\mathbf{u}$, so the deformation gradient and the strain energy are computed from first-order derivatives of the network output. Instead of enforcing the strong-form equilibrium equation as standard physics-informed neural networks do, DEM minimizes the total potential energy $\\Pi$ of Eq. (9) — strain energy minus the work of body forces and tractions — together with a boundary-condition mean-squared-error term, which the authors argue needs only first-order differentiation and therefore converges faster. Around this sit the NGP-NeRF reconstruction, which supplies the initial geometry, and the particle-sampling step (random, Poisson disc, or uniform mesh), which converts that geometry into the DEM training domain.","core_discovery":"DEM-NeRF's central claim is that coupling NGP-NeRF for image-based geometry with a Deep Energy Method (DEM) network for quasi-static Neo-Hookean elasticity produces high-quality 3D reconstruction and reliable deformation prediction directly from image sequences, in hyperelastodynamic and loading scenarios. The pipeline reconstructs the undeformed T-bar as a mesh, samples it into a particle cloud, applies a fixed base and a torque formed by two opposing forces, and trains the DEM network by minimizing the total potential energy of the solid plus a boundary-condition penalty. The predicted displacement field is fed back into the NeRF renderer to generate new views of the deformed object. The paper reports DEM training of about 20 seconds and prediction of about 1 second, against roughly 600 seconds for the MPM-based PAC-NeRF baseline and 1,140 seconds for full-order FEM, and calls DEM-NeRF the first neuro-symbolic method for real-time solid deformation reconstruction and prediction.","pith_inferences":["The authors stop short of reporting any quantitative accuracy metric (error against measured displacement or rendered image); a head-to-head error comparison with PAC-NeRF and FEM would be the natural next test of the 'reliable prediction' claim.","The same architecture can be turned into an inverse method: by treating Young's modulus and Poisson's ratio as trainable parameters, DEM-NeRF could identify material properties from video alone, something the paper mentions as motivation but does not implement.","The abstract's 'hyperelastodynamics' suggests time-dependent behavior, but Eq. (9) is a quasi-static energy minimization with no inertial terms; adding kinetic energy and time stepping would be needed to justify that label.","A testable extension is to demonstrate the pipeline on a second object with a genuinely complex boundary (e.g., a teapot) using Poisson disc sampling, which the paper argues should outperform uniform sampling but does not evaluate."],"forward_implications":["Real-time interactive simulation of elastic solids becomes feasible from ordinary camera images, since the reported per-prediction cost is about one second after short training.","The method removes the need for explicit geometric models, manually specified meshes, or known hard-boundary conditions, which are the main obstacles to applying FEM and BEM to real scenes.","Because it uses images plus physics rather than labeled displacement data, it can be applied to new objects and loading setups without collecting ground-truth deformation fields.","If the quasi-static Neo-Hookean assumption holds, the same pipeline extends naturally to any object whose material can be described by a strain energy function, including soft robotics and biomedical tissue models."],"supporting_citations":[{"why":"Supplies the neural radiance field formulation that represents the object's geometry and appearance from images.","marker":"[16]"},{"why":"Provides the NGP-NeRF instant graphics primitives used for fast reconstruction and NeRF-to-mesh conversion.","marker":"[30]"},{"why":"Introduces the Deep Energy Method that DEM-NeRF adopts to solve finite-deformation hyperelasticity by potential-energy minimization.","marker":"[31]"},{"why":"Is the PAC-NeRF baseline (MPM-based physics-augmented NeRF) that DEM-NeRF compares against for training time.","marker":"[19]"},{"why":"Gives the PINN loss formulation with equilibrium residual and boundary terms that the paper contrasts with DEM's first-order-only energy minimization.","marker":"[32]"},{"why":"Supports training the network to fit boundary conditions, which DEM-NeRF uses for its displacement boundary loss.","marker":"[33]"},{"why":"Supplies the full-order FEM baseline whose 1,140 second computation time appears in the performance comparison.","marker":"[34]"}],"fun_headline_variants":["Neuro-symbolic NeRF predicts solid deformation in a second","No geometry needed: NeRF plus physics net simulates deformation","Real-time solid deformation from sparse photos via DEM-NeRF","DEM-NeRF: From images to deformation predictions in one second","Photo-based 3D simulation without explicit geometry in seconds"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole pipeline assumes the T-bar's deformation is exactly a compressible Neo-Hookean quasi-static response with manually chosen Young's modulus, Poisson's ratio, and torque loads, and none of those values are reported in the paper.","fun_headline_variants_meta":{"raw":{"variants":["Neuro-symbolic NeRF predicts solid deformation in a second","No geometry needed: NeRF plus physics net simulates deformation","Real-time solid deformation from sparse photos via DEM-NeRF","DEM-NeRF: From images to deformation predictions in one second","Photo-based 3D simulation without explicit geometry in seconds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000687,"raw_usage":{"total_tokens":3123,"prompt_tokens":965,"completion_tokens":2158,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":581,"completion_tokens_details":{"reasoning_tokens":2072}},"tokens_in":581,"tokens_out":2158,"duration_ms":16167,"temperature":1.0,"reasoning_tokens":2072,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T12:51:01.735909+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recreate the T-bar experiment with the same multi-view images and independently measure the actual 3D displacement field (for example with digital image correlation or tracked markers), then compare it with DEM-NeRF's prediction using the paper's unreported material and load parameters; a mismatch beyond rendering noise, or any observable time-dependent motion that the quasi-static energy minimization cannot represent, would undercut the 'reliable deformation prediction' claim.","supporting_citations":[{"cited_title":"NeRF: Representing scenes as neural radiance fields for view synthesis,","cited_arxiv_id":null,"evidence_quote":"Supplies the neural radiance field formulation that represents the object's geometry and appearance from images."},{"cited_title":"Instant neural graphics primitives with a multiresolution hash encoding,","cited_arxiv_id":null,"evidence_quote":"Provides the NGP-NeRF instant graphics primitives used for fast reconstruction and NeRF-to-mesh conversion."},{"cited_title":"A deep energy method for finite deformation hyperelasticity,","cited_arxiv_id":null,"evidence_quote":"Introduces the Deep Energy Method that DEM-NeRF adopts to solve finite-deformation hyperelasticity by potential-energy minimization."},{"cited_title":"The mixed deep energy method for resolv- ing concentration features in finite strain hyperelasticity,","cited_arxiv_id":null,"evidence_quote":"Gives the PINN loss formulation with equilibrium residual and boundary terms that the paper contrasts with DEM's first-order-only energy minimization."},{"cited_title":"Physics-informed deep learning for computational elastodynamics without labeled data,","cited_arxiv_id":null,"evidence_quote":"Supports training the network to fit boundary conditions, which DEM-NeRF uses for its displacement boundary loss."},{"cited_title":"Reduced order modeling for nonlinear structural analysis using Gaussian process regression,","cited_arxiv_id":null,"evidence_quote":"Supplies the full-order FEM baseline whose 1,140 second computation time appears in the performance comparison."}],"review_version":1}