{"id":"de480027-7793-4943-9862-e0d0b07e4fa0","arxiv_id":"2506.04120","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A hybrid 3D Gaussian splatting plus explicit mesh representation, optimized end-to-end with differentiable rendering and physics, reconstructs objects and calibrates robot poses from imperfect real-world RGB trajectories.","lead":"This paper presents a pipeline that turns raw video from a low-cost robot into a simulation-ready 3D scene, combining photorealistic Gaussian splatting with explicit object meshes. It jointly fixes robot and camera poses while reconstructing objects, which matters because real robot data is usually too noisy for standard reconstruction tools.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reported hold-out PSNR uses test-time pose alignment (Appendix D.4), so the novel-view numbers do not validate the claimed calibrated poses; calibration quality on unseen views is unmeasured.","rationale":"I read the paper as a systems contribution; the novel SplatMesh representation is plausible, and the simulation-to-simulation calibration results in Appendix D.1 provide partial independent support. The reader's weakest_assumption (fixed kinematic model) is a real external risk, but the paper explicitly frames its setup as 'reasonably accurate geometry and kinematics,' so a robustness test would be informative rather than decisive. The test-time pose alignment is more pointed because it is an internal inconsistency: the evaluation protocol undercuts the calibration claim in the central real-world experiment. The reader's rationale did list 'held-out PSNR uses pose alignment' as a weakness, so my agreement is partial; I elevate it to the primary concern because it is self-admitted and directly affects the main quantitative evidence. The 'physical parameters' overclaim is secondary: the abstract overreaches, but the core reconstruction pipeline can be assessed without that component. I recommend keeping the CONDITIONAL verdict, now conditioned specifically on re-evaluation without test-time pose alignment and on releasing code/data for independent verification.","tokens_in":14082,"tokens_out":6108,"duration_ms":56353,"concrete_test":"Re-run the six-object real evaluation of Table 3, rendering held-out views with the poses from the training-time optimization only (no Appendix D.4 alignment step). If average PSNR drops by more than ~1.5 dB, or if the alignment step moves the estimated wrist-camera poses by more than a few millimeters, the novel-view metric does not support the pose-calibration claim and the verdict should require a recalibrated evaluation. This check needs code/data; currently neither is released.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is that the paper's headline claims—joint recovery of photorealistic novel views and annotation-free robot/camera pose calibration from imperfect robot data—are not jointly tested. Appendix D.4 states explicitly that the PSNR values in Table 3 are computed 'after an additional optimization step, to align the camera poses for the held-out views.' That is a test-time fit of the quantities the method is supposed to have calibrated. If the training-time pose estimates were accurate, this step would be unnecessary. The reported PSNR therefore measures the quality of the SplatMesh representation given test-fitted poses, not the generalization of the calibration. The gap between PSNR with and without this alignment is an unquantified measure of calibration error, and Table 3's comparison against the Proprio-only ablation conflates representation quality with pose-calibration quality. The abstract also claims refinement of 'physical parameters,' but Appendix B.2 states the work 'focus[es] on object reconstruction and kinematics,' and no experiment recovers mass, friction, or contact parameters. These are scope mismatches between the stated contribution and the evidence. The pose-alignment issue is the most load-bearing because it directly affects the central quantitative evidence on real data.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a hybrid scene representation called SplatMesh, which combines 3D Gaussian Splatting for appearance with an explicit triangle mesh for geometry, and an end-to-end optimization pipeline that uses differentiable rendering and differentiable MuJoCo (MJX) physics to jointly refine object geometry, appearance, robot joint angles, and camera extrinsics directly from raw RGB images and proprioceptive robot states. The authors validate the approach on a simulated dataset of 64 YCB objects and on real ALOHA 2 bimanual robot trajectories involving six YCB objects, reporting novel-view synthesis metrics, geometry reconstruction errors, and a 3D asset generation pipeline via CAT3D. The central claim is that a single pipeline can recover metric object meshes, photorealistic novel views, and annotation-free robot/camera pose calibration from imperfect, low-cost robot data without additional data collection.","tokens_in":14383,"tokens_out":4703,"duration_ms":44323,"significance":"If the claims hold, the framework is a valuable contribution to real-to-sim: it unifies appearance and physics-ready geometry in a single differentiable representation, requires no separate data collection or manual calibration, and demonstrates results on a low-cost ALOHA 2 platform with onboard RGB sensors. The simulation results on 64 YCB objects are clean and the method outperforms NeRFacto and 3DGS on several novel-view metrics while also producing explicit meshes. The real-data reconstructions are qualitatively plausible and the comparison against a TRELLIS baseline is informative. The paper also ships a reproducible experimental setup and ablation studies. However, the headline claims about pose calibration and physical parameter refinement are not fully supported by the presented evidence, and several evaluation choices make the reported numbers optimistic.","major_comments":[{"comment":"The real-world novel-view PSNR values in Table 3 are computed after an additional optimization step that aligns the camera poses for the held-out views, as explicitly stated in Appendix D.4. This means the reported PSNR measures the quality of the SplatMesh representation given test-time fitted poses, not the generalization of the poses calibrated at training time. Since 'annotation-free robot pose calibration' is one of the three headline contributions, the evaluation conflates representation quality with calibration quality. Please report PSNR without the test-time alignment and, if possible, quantify the pose error on held-out views (e.g., by evaluating the optimized camera extrinsics against a reference) so that the calibration claim is directly tested.","section":"Appendix D.4 / Table 3"},{"comment":"The abstract and Section 4.1 claim joint refinement of 'physical parameters' and use of 'differentiable physics,' but Appendix B.2 states the work focuses on 'object reconstruction and kinematics,' and no experiment recovers or validates mass, friction, contact, or other physical parameters. The real and simulated experiments involve static posed objects; no trajectory with contacts or object dynamics is shown. Either add a dynamics experiment (for example, releasing or pushing a reconstructed object and comparing simulated motion against real observations) or revise the claims to say 'kinematic parameters' and 'differentiable kinematics' where appropriate. As written, the scope mismatch is load-bearing because 'physical parameters' is part of the advertised contribution.","section":"Abstract / Section 4.1 / Appendix B.2"},{"comment":"The reported novel-view metrics in Table 2 use object-specific regularization weights (lambda_LL in [0.1, 1.0] and lambda_E in [0.01, 0.1]) that appear tuned per object. If these weights were selected using the test set or with access to the ground-truth meshes, the reported PSNR is optimistic and the comparison against NeRFacto and 3DGS is unfair. Please clarify the weight selection protocol (for example, a split-off validation set or a fixed schedule) or report results with a single fixed weight, as done in the 'Ours w/o mesh reg.' ablation.","section":"Section 5.1.2 / Table 2"},{"comment":"The calibration experiments in Appendix D.1 add zero-mean Gaussian noise to joint angles, but the central motivation is robustness to 'imperfect' and 'inaccurate' robot models, which includes systematic errors such as wrong link lengths, joint offsets, or camera intrinsics. The paper states the setup requires a simulator with 'reasonably accurate geometry and kinematics, but not perfect,' yet no experiment perturbs the kinematic structure itself. A concrete stress test would be to modify link lengths or joint offsets in the simulated ALOHA model and measure whether the end-to-end fit still recovers accurate poses and object geometry. Without such a test, the method's robustness to systematic model mismatch—a key part of the real-to-sim claim—is not established.","section":"Appendix D.1 / Section 3.2"},{"comment":"The real-world evaluation is limited to six objects with no repeated runs, no error bars, and no report of variance across random initializations or different ALOHA trajectories. Given the stochastic nature of gradient-based optimization and the mask pipeline, a single run per object is insufficient to support the quantitative geometry and PSNR comparisons in Table 3, especially for the claimed improvement over the Proprio-only baseline. Please provide multiple runs (or a variance estimate) and state the number of seeds used; this is important because the real-data evidence is the main support for the paper's central claim.","section":"Section 5.2 / Table 3"}],"minor_comments":[{"comment":"The conclusion contains a typo: 'the the feasibility' should be 'the feasibility.'","section":"Section 6"},{"comment":"There is a typo in the first sentence: 'optimizaiton' should be 'optimization.'","section":"Appendix A.1"},{"comment":"The sentence 'shows results for novel-viewy synthesis' contains a typo; 'novel-viewy' should be 'novel-view.'","section":"Appendix D.3"},{"comment":"SuGaR is listed twice: reference [26] and reference [36] refer to the same paper by Guédon and Lepetit; please consolidate.","section":"References"},{"comment":"The ablation description mentions 'a uniform weight (lambda_L1 = 0.1)' but no L1 regularization term is defined in the loss equations in Appendix B.2. Please clarify whether this refers to the photometric loss weight (L_photo).","section":"Section 5.1.2"},{"comment":"The text says the framework is 'open source,' but no code repository or link is provided; please include the project URL or state where the code will be released.","section":"Section 5.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely an early arXiv version from a strong group, and the core idea is promising. The largest concern is the test-time pose alignment in Appendix D.4, which directly undermines the calibration claim as evaluated. The scope mismatch between the advertised 'physical parameters' and the actually optimized quantities should also be resolved. I would encourage the editor to request a revision that closes these gaps rather than rejecting, as the simulation results and the qualitative real-world reconstructions suggest the method has merit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The one thing you should know: this is a solid systems paper with a genuinely new representation, but the headline real-world numbers are softer than the abstract suggests. The held-out PSNR on real objects is computed after an additional test-time pose alignment step (Appendix D.4), which means those numbers validate the representation given oracle-like poses, not the calibration. And the abstract's claim about refining 'physical parameters' is not backed by experiments—Appendix B.2 says the work focuses on object reconstruction and kinematics, and no mass/friction/contact identification is demonstrated.\n\nWhat is actually new: the SplatMesh representation. Tying 3D Gaussians to the faces of an explicit mesh—so geometry and appearance are optimized jointly and the mesh is directly usable in a physics engine—is a real step beyond prior radiance-field real-to-sim work. Coupling that with differentiable MuJoCo MJX physics and joint refinement of robot joint angles and camera extrinsics from imperfect trajectories is a practical contribution. The simulation study on 64 YCB objects is clean: Chamfer distance 0.073 mm^2, PSNR 30.91, beating NeRFacto and 3DGS, with ablations that show the Laplacian and surfel constraints matter. The real ALOHA2 experiments convincingly show that without pose optimization the geometry collapses, and the recovered meshes are qualitatively decent.\n\nThe soft spots are real but not fatal. The test-time pose alignment is the biggest: it decouples the novel-view numbers from the calibration claim. There are no error bars on the six real objects, the per-object regularization weights appear to be tuned on the test set, and no code or data are released. The pose calibration itself is only validated in simulation (Appendix D.1), not on the real data. Still, the central claim—that this hybrid representation can be optimized end-to-end from imperfect robot data—holds up in simulation and is plausibly correct on real data.\n\nThis is for people working on real-to-sim, 3DGS for robotics, or differentiable rendering pipelines. It deserves a serious referee; I'd want the authors to either report PSNR without test-time alignment or validate calibration error on held-out views, add error bars, and fix the 'physical parameters' language. But the core idea is sound and the paper is honest about its limitations (fixed topology, local minima, rigid objects, no relighting). Engage with it.","headline":"A genuinely useful new hybrid representation (mesh + surface-bound Gaussians) and an honest end-to-end pipeline, but the real-world novel-view numbers are weakened by test-time pose alignment and the abstract oversells 'physical parameters.'","tokens_in":14933,"tokens_out":2181,"would_cite":true,"duration_ms":21544,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"One hybrid representation lets a single differentiable pipeline recover object meshes, appearance, and robot and camera poses directly from raw, imperfect robot trajectories.","keywords":["3D Gaussian Splatting","differentiable rendering","differentiable physics","MuJoCo","real-to-sim","robot calibration","mesh reconstruction","ALOHA 2"],"falsifier":"Deliberately corrupt the robot model with a known systematic error — for example, lengthen one arm link by 2 cm or add a fixed 0.05 rad offset to one joint encoder — then run the pipeline on a real trajectory and measure the recovered object mesh against ground truth. If the joint-angle and camera-extrinsic parameterization cannot absorb systematic (as opposed to random) error, the recovered geometry will be measurably biased, which directly tests the paper's premise that the simulator needs only 'reasonably accurate' geometry and kinematics.","tokens_in":13885,"feed_emoji":"🤖","tokens_out":15220,"duration_ms":127445,"temperature":0.7,"pith_summary":"The paper tries to establish that real-to-sim — creating a physically usable simulation from real robot observations — can be solved as a single differentiable optimization instead of a chain of separate reconstruction, calibration, and meshing steps. Its central object is SplatMesh, a hybrid representation that couples the photorealistic appearance of 3D Gaussian Splatting with explicit triangle meshes suited to physics simulation, with Gaussians constrained to lie on the mesh surface. Because both the renderer and the MuJoCo physics backend are differentiable, pixel-level photometric errors propagate back through the whole system to mesh vertices, camera extrinsics, and robot joint angles. If this works, a robot that simply films an object while moving its arms can yield, from its own imperfect onboard video, a metric-scale object mesh, photorealistic novel views, and calibrated robot and camera poses, with no extra data collection, no structure-from-motion preprocessing, and no manual annotation. That would make high-fidelity simulation assets cheap enough to generate from ordinary robot operation, which is what makes large-scale robot learning from simulation practical.","feed_headline":"Robot video becomes a physics scene in one end-to-end fit","feed_subtitle":"Hybrid Gaussians-plus-mesh recovers shape, appearance, and robot pose from noisy RGB alone.","key_machinery":"SplatMesh, a hybrid scene representation in which a deformable triangle mesh fixes the geometry and 3D Gaussians carry the appearance. Gaussian means are sampled on mesh faces via barycentric weights, so appearance follows the mesh as it deforms; covariance is kept axis-aligned to the face normal with the normal-direction scale clamped near zero, turning each Gaussian into a surface element rather than a volumetric blob. The same differentiable pipeline then couples the 3D Gaussian Splatting rasterizer with MuJoCo's JAX-based MJX physics, so losses on rendered RGB, silhouettes, and estimated surface normals, together with Laplacian mesh smoothing, back-propagate in one pass to mesh vertices, Gaussian parameters, camera extrinsics, and robot joint angles.","core_discovery":"The central claim is that visual and physical reconstruction of a dynamic robot scene should be one problem, not two: the same scene is simultaneously a set of 3D Gaussians for rendering and a triangle mesh for simulation, and the two are coupled because Gaussian means live on mesh faces through barycentric coordinates while Gaussian covariance stays axis-aligned to the face normal with near-zero normal extent. Minimizing a weighted sum of photometric, silhouette, normal-consistency, and Laplacian-smoothing losses — through differentiable rasterization and differentiable MuJoCo kinematics — jointly refines object geometry, appearance, robot joint angles, and camera poses directly from raw RGB and proprioception. On a low-cost ALOHA 2 bi-manual platform with two fixed and two wrist-mounted RGB cameras, the authors report reconstructed YCB object meshes with square-root Chamfer distances between 3.1 and 7.4 mm, novel-view PSNR between 21.6 and 25.5 dB versus 14.0 to 20.2 dB without pose calibration, and, in simulation, tool-center-point error roughly halved under joint-angle noise. The same framework doubles as an asset generator, turning single images or text prompts into textured meshes importable into MuJoCo.","pith_inferences":["I read the calibration results as implying an untested corollary: the same optimization could serve as a passive health monitor for a robot fleet, with slowly growing joint-offset residuals flagging kinematic wear or mounting drift over time — the paper does not test this.","The fixed-topology limitation (a sphere-initialized mesh stays topologically sphere-like) suggests a concrete extension the authors mention only as future work: initialize from a coarse multi-view or single-view estimate instead of a sphere, which would extend the method to objects with handles or holes.","Because the method deliberately avoids segmenting the robot arm and instead lets the Menagerie model explain the scene, I expect calibration to degrade gracefully rather than catastrophically as background clutter grows, since the model already explains the arm's appearance and the photometric loss will favor poses that keep it aligned."],"forward_implications":["A single optimization pass recovers metric object meshes, photorealistic novel views, and robot and camera calibration at once, from raw RGB and proprioception alone, with no dedicated scanning session, COLMAP, or manual pose annotation.","Because recovered objects come out at metric scale with 6-DoF pose in the robot workspace, the meshes are simulation-ready as produced, and where a simulator cannot render Gaussians directly, the appearance can be baked into a standard texture map.","On the simulated YCB benchmark the full method outperforms both NeRFacto and vanilla 3DGS at the same 15,000-iteration budget (PSNR 30.91 vs 30.29 and 26.97), and the ablations show that both mesh regularization and the surface-aligned (surfel) constraint are needed for that margin.","Jointly optimizing poses makes the reconstruction itself feasible on the low-cost platform: freezing the cameras at nominal values leaves real-object geometry essentially unconverged (proprio-only Chamfer errors of 11.7 to 18.9 mm) and drops held-out PSNR by roughly 5 to 8 dB."],"supporting_citations":[{"why":"3D Gaussian Splatting — supplies the differentiable rasterizer through which photometric gradients reach the Gaussians and mesh vertices.","marker":"[2]"},{"why":"MuJoCo (2012) — the physics engine the pipeline builds on, providing simulation of the robot and objects.","marker":"[3]"},{"why":"Brax / MJX — the JAX-based differentiable physics implementation used to propagate gradients through robot kinematics.","marker":"[5]"},{"why":"MuJoCo Menagerie — the open-source ALOHA 2 simulation model that serves as the imperfect, 'reasonably accurate' robot prior being calibrated.","marker":"[24]"},{"why":"ALOHA 2 (2024) — the low-cost bi-manual platform whose onboard RGB and proprioceptive recordings form the real-world dataset.","marker":"[30]"},{"why":"YCB object set — provides the objects for the simulated benchmark and the six real-world reconstruction props.","marker":"[29]"},{"why":"SAM 2 — produces the object segmentation masks that supervise the silhouette loss on real scenes.","marker":"[28]"},{"why":"Fine-tuned diffusion normal estimator — supplies per-frame surface normals used in the normal-consistency loss.","marker":"[27]"},{"why":"Gaussian surfels — the surface-aligned Gaussian formulation the paper adapts, clamping covariance in the normal direction; the ablation shows dropping this constraint hurts both geometry and view quality.","marker":"[25]"},{"why":"TRELLIS — the single-view 3D generation baseline used as a comparison point for real geometry reconstruction, aligned to ground truth using privileged scale and pose information.","marker":"[32]"}],"fun_headline_variants":["One fit turns robot video into mesh plus physics","Hybrid Gaussians and meshes build sim from raw robot footage","Real-to-sim from imperfect robot data in one end-to-end pass","Jointly refine mesh, pose, and physics from raw robot video","Gaussians for looks, meshes for physics, one fit from noisy video"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The pipeline assumes the supplied MuJoCo model of the ALOHA 2 robot is accurate enough in link lengths, joint offsets, and kinematics that every real-robot error can be absorbed by optimizing joint angles and camera extrinsics alone; systematic kinematic errors in that model would silently bias the recovered scene.","fun_headline_variants_meta":{"raw":{"variants":["One fit turns robot video into mesh plus physics","Hybrid Gaussians and meshes build sim from raw robot footage","Real-to-sim from imperfect robot data in one end-to-end pass","Jointly refine mesh, pose, and physics from raw robot video","Gaussians for looks, meshes for physics, one fit from noisy video"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000617,"raw_usage":{"total_tokens":2898,"prompt_tokens":1009,"completion_tokens":1889,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":625,"completion_tokens_details":{"reasoning_tokens":1798}},"tokens_in":625,"tokens_out":1889,"duration_ms":13341,"temperature":1.0,"reasoning_tokens":1798,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:46:49.204313+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Deliberately corrupt the robot model with a known systematic error — for example, lengthen one arm link by 2 cm or add a fixed 0.05 rad offset to one joint encoder — then run the pipeline on a real trajectory and measure the recovered object mesh against ground truth. If the joint-angle and camera-extrinsic parameterization cannot absorb systematic (as opposed to random) error, the recovered geometry will be measurably biased, which directly tests the paper's premise that the simulator needs only 'reasonably accurate' geometry and kinematics.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Brax / MJX — the JAX-based differentiable physics implementation used to propagate gradients through robot kinematics."},{"cited_title":"Zakka, Y","cited_arxiv_id":null,"evidence_quote":"MuJoCo Menagerie — the open-source ALOHA 2 simulation model that serves as the imperfect, 'reasonably accurate' robot prior being calibrated."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"SAM 2 — produces the object segmentation masks that supervise the silhouette loss on real scenes."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gaussian surfels — the surface-aligned Gaussian formulation the paper adapts, clamping covariance in the normal direction; the ablation shows dropping this constraint hurts both geometry and view quality."}],"review_version":1}