{"id":"20ab71f8-39ff-4e31-9861-ae2edbba6ad3","arxiv_id":"2607.21471","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"FutureSurf, a new benchmark for held-out future surface reconstruction, shows deformation-MLP methods leave a 2-6.6× future-surface gap while rendering quality stays flat.","lead":"This paper introduces a benchmark and dataset that tests whether dynamic 3D reconstruction methods can predict a scene's surface after the observed video ends. It finds that current methods are 2-6.6× less accurate on future geometry, and that image-quality metrics do not track this error.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Decoupling claim inferred from within-method time-series, not cross-method comparison; NVS metrics may still track geometry across methods.","rationale":"The paper's central novel contribution is the rendering–surface decoupling: that future rendering quality is not sufficient evidence of future-surface quality. The evidence for this is §7's rank correlations between LPIPS and CD, computed per-frame across future frames within each of six scenes for one backbone (DG-Mesh). This establishes that for this method, over time, the two errors are not monotonically related. However, the field standardly uses NVS metrics to compare different models, not to track a single model's error over time. A strong cross-method correlation could exist even if within-method temporal trajectories are uncorrelated. The paper provides no cross-method comparison; the second backbone (Deformable-3DGS) is used only for the gap result, not the decoupling, and shares the same temporal model. Thus the headline conclusion that 'the novel-view-synthesis metrics the field reports do not track future geometry' is an overstatement of what is demonstrated. This is load-bearing because the title and abstract present the decoupling as a primary finding. The benchmark itself, as a controlled evaluation protocol with exact ground truth and falsification controls, remains valuable and well-constructed; the issue is in the generalization of one of its headline claims. The reader flagged a related overgeneralization (the shared deformation-MLP family) but not this specific evidential gap, so my agreement is partial. The verdict remains CONDITIONAL because the issue can be fixed by adding cross-method evidence or tempering the abstract; it does not undermine the entire contribution. I therefore do not change the reader's verdict.","tokens_in":11728,"tokens_out":6774,"duration_ms":68127,"concrete_test":"Run the released benchmark on at least two additional methods that output per-frame meshes (or multiple training runs with different seeds/hyperparameters of existing methods) on the six asset scenes. Compute future PSNR/LPIPS and future CD for each method-scene combination using the released scorer. Then compute the Spearman rank correlation of LPIPS vs CD across these method-scene pairs. If |ρ| > 0.5, the decoupling claim as stated would be falsified; if |ρ| remains < 0.2, the claim is supported. Also report the same across the two existing backbones on scenes where both have surface proxies.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The claim that future NVS metrics do not track future geometry (§7) is based on Spearman correlations computed per-frame across future times within each scene, for a single backbone (DG-Mesh). This shows only that, for one method, rendering quality and surface error do not move together over the prediction horizon. But the field uses NVS metrics to compare methods, not to track a single model in time. Without cross-method evidence, it is logically possible that better-rendering methods also have better future surfaces, even if within-method temporal fluctuations are uncorrelated. No such cross-method analysis is provided; the two backbones share a deformation-MLP temporal model and the decoupling is reported only on DG-Mesh. Thus the conclusion that 'the novel-view-synthesis metrics the field reports do not track future geometry' overstates the evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"FutureSurf proposes a standardized benchmark and dataset for evaluating dynamic surface reconstruction at held-out future times. It contributes eight analytically defined controlled motions (five surface-changing, three falsification controls) with exact per-frame ground-truth meshes, a 75/25 train/future split, a method-agnostic per-frame Chamfer protocol, a ground-truth-side recoverability oracle, benchmark card, and Croissant metadata. Using DG-Mesh and Deformable-3DGS as backbones, the paper reports a 2.0–6.6× future/observed Chamfer gap on asset scenes and 2.7–4.1× on the controlled motions, with falsification controls behaving as designed; a per-motion oracle recovers four of five constructed futures using simple matched rules; and a within-method analysis finds weak rank correlation between future rendering quality (LPIPS) and future-surface CD (mean |ρ|=0.13), which the paper interprets as a rendering–surface decoupling. The authors release splits, scoring code, and metadata. The paper is explicitly framed as a diagnostic benchmark rather than a new reconstruction method.","tokens_in":11915,"tokens_out":5198,"duration_ms":51234,"significance":"If taken as scoped, this is a valuable contribution: it makes an unmeasured quantity—future-surface mesh accuracy—measurable, and provides exact ground truth, falsification controls, and reproducible scripts. The recoverability oracle is honestly scoped as an optimistic, ground-truth-side reference, and the paper carefully separates removable gauge error from non-rigid error. The main limitations are the breadth of tested temporal models (both backbones share a deformation-MLP temporal family) and the within-method nature of the decoupling evidence; these limit the generality of the headline 'Future Rendering ≠ Future Surface' but do not undermine the benchmark itself. The release of code, splits, and a CPU-run oracle is a concrete strength for reproducibility.","major_comments":[{"comment":"The abstract and Contribution 2 state that 'four of five recoverable from observed motion by a fixed rule.' This is contradicted by §6: Table 6 shows the best rule varies per motion (harmonic K=2 for wave/compound, cubic for stretch, quadratic for accel, velocity for bulge) and the text explicitly says 'No single per-vertex extrapolation rule covers the suite.' The oracle establishes recoverability only under per-motion matched rule families, not a fixed rule. This phrasing is load-bearing because it is used to argue that the 2.7–4.1× gap reflects a backbone limitation rather than intrinsic unknowability. Please rephrase to 'per-motion matched rules' or 'simple rules' and align the abstract/contributions with §6.","section":"Abstract and §1 (Contribution 2) vs §6"},{"comment":"The decoupling conclusion is based on per-frame Spearman correlations computed within each DG-Mesh scene over future frames (mean |ρ(LPIPS,CD)| = 0.13). This supports only the within-method statement that, for DG-Mesh, rendering quality does not track surface error over the prediction horizon. It does not support the broader claim that 'the NVS metrics the field reports do not track future geometry,' which is a cross-method claim. No cross-method ranking analysis is provided; the second backbone is not included in the decoupling analysis, and §1 itself states the decoupling is reported on DG-Mesh. Please either add a cross-method analysis (e.g., rank-correlating method-level future PSNR/LPIPS with future CD across methods/scenes) or restrict the claim to the tested DG-Mesh backbone.","section":"§7, Table 4"}],"minor_comments":[{"comment":"Spacing typos throughout: e.g., 'a monocular orbit camera,200frames, the first75%' (§3) and 'gap:2.7–4.1×by mean' (§5). These are likely LaTeX artifacts but should be cleaned before publication.","section":"§3, §5"},{"comment":"The column header 'Non-rigid' should explicitly say it is the Sim(3)-gauge-removed future gap; otherwise a reader may confuse it with a non-rigid motion class. The caption explains it, but the header itself is ambiguous.","section":"Table 4"},{"comment":"The oracle CD is bounding-box-normalized linear Chamfer, which the text says is on a different scale from Table 5's absolute mesh CD. It would help to state in the table caption that the ratios vs. the freeze reference are the primary comparison, not the absolute CD values.","section":"§6, Table 6"}],"recommendation":"major_revision","confidential_remarks":"The paper's empirical core is sound and the release appears carefully done. My concerns are scope overstatements in the abstract/contribution list and in §7; both are fixable with rewording or additional analysis. I would not reject. The recoverability 'fixed rule' phrasing is likely a wording error rather than a scientific flaw, but it must be corrected because it currently contradicts §6. The decoupling claim needs either a cross-method analysis or a clear restriction to the tested backbone."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Let me give you the short version: this paper creates a benchmark (FutureSurf) that scores reconstructed surface meshes at held-out future times, with exact ground-truth meshes and falsification controls. That is a real gap in dynamic 3D reconstruction evaluation. The protocol is clean: method-agnostic, fixed 75/25 split, released scoring code, and an observed-window precondition so bad future results are not just bad reconstruction. The three falsification controls behave as designed, which gives me confidence in the metric.\n\nWhat is genuinely new: the held-out-future surface protocol, the controlled analytic motions that isolate temporal factors, and the ground-truth-side recoverability oracle. The oracle is scoped honestly as optimistic, and it shows that most of the constructed futures are recoverable by simple per-vertex rules fitted to observed ground truth. So the large gaps (2.7-4.1x) in the controlled study do point to a limitation of the tested deformation-MLP backbones, not intrinsic unknowability.\n\nSoft spots, in order of severity. First, the abstract overstates the recoverability result: it says four of five recoverable from observed motion by a fixed rule, but the table uses per-motion best rule. A fixed rule across all motions would be different; this is a wording issue but it matters. Second, the rendering-surface decoupling claim is weaker than advertised. The evidence is per-frame Spearman correlations across future frames for DG-Mesh only. That shows within one method, rendering quality and surface error do not move together over the horizon. It does not show the NVS metrics used to compare methods do not track geometry. Cross-method comparisons are missing. The stress-test concern is valid. Both backbones share the same temporal model, so the gap might be family-specific. The authors acknowledge this in Section 8 but still frame the gap as the headline result.\n\nMinor: the CD tables lack error bars. For a ratio metric, that would be worth adding.\n\nMy overall take: the benchmark contribution is solid and the body is carefully scoped. The central claimed gap is supported within the stated scope. The paper deserves peer review; an editor should send it out with a request to fix the abstract and either add a temporally distinct backbone or temper the generalization. I would cite this if I were doing dynamic reconstruction evaluation.","headline":"FutureSurf is a genuinely useful evaluation contribution with a clean protocol, but the abstract overstates two results and the decoupling claim needs cross-method evidence.","tokens_in":716,"tokens_out":3550,"would_cite":true,"duration_ms":37782,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Future rendering quality does not track future surface accuracy in dynamic-scene reconstruction; a new benchmark measures the gap.","keywords":["future-surface reconstruction","dynamic scenes","benchmark","extrapolation","Chamfer distance","rendering-surface decoupling","novel view synthesis","diagnostic evaluation"],"falsifier":"Train a dynamic reconstruction method with a temporally distinct representation (e.g., an explicit 4D grid or physics-based simulation) on the released FutureSurf splits; if its future Chamfer gap drops to near unity (future/observed ≈ 1) on the recoverable controlled motions and its per-frame |ρ(LPIPS,CD)| rises substantially, the paper's claim of a general field-wide decoupling would be refuted.","tokens_in":11618,"feed_emoji":"📏","tokens_out":6013,"duration_ms":49690,"temperature":0.7,"pith_summary":"This paper argues that the standard way of evaluating dynamic-scene reconstruction—checking how well a model renders held-out future frames—misses the geometry that deployment actually needs. It introduces a controlled benchmark with analytically defined motions and exact future ground-truth meshes, and shows that two dynamic-Gaussian backbones reconstruct the observed surface accurately but degrade 2.0–6.6× when scored on the future surface mesh. The rendering metrics (PSNR, LPIPS) are weakly rank-correlated with future surface error (mean |ρ|=0.13), so a future rendering can look plausible while the surface is wrong where it moves. If correct, the field's novel-view-synthesis metrics are not a proxy for future-surface quality, and future-time reconstruction needs separate geometric evaluation and stronger temporal inductive bias.","feed_headline":"Future rendering misses future geometry, benchmark shows","feed_subtitle":"Held-out future surfaces are 2–6.6× worse than observed, while rendering metrics stay flat: NVS metrics miss geometry.","key_machinery":"The benchmark's load-bearing instrument is the future/observed gap: the ratio of per-frame bidirectional Chamfer distance between extracted and ground-truth meshes on the held-out future window to the same quantity on the observed window. It converts 'how good is the future surface' into a normalized diagnostic that isolates extrapolation failure from interpolation quality, and it is paired with three falsification controls (a surface-invariant twist, a rigid-rotation gauge control, and a frozen-future stop control) designed to expose a broken metric or alignment. A second instrument, the ground-truth-side recoverability oracle, fits simple and learned per-vertex temporal rules to observed g","core_discovery":"The central discovery is a measured decoupling: for the tested dynamic-surface backbones, future rendering quality and future-surface accuracy do not move together. On six animated asset scenes and a suite of eight controlled motions with exact future ground truth, per-frame future Chamfer distance is 2.0–6.6× the observed-window error, and per-frame |ρ(LPIPS,CD)| averages 0.13, with a linear fit explaining under 6% of variance. The future error is structured, concentrating where the surface moves, and it persists after removing a global Sim(3) gauge, so it is genuine non-rigid shape error. The paper frames this as evidence that the field's standard evaluation protocol measures the wrong qua","pith_inferences":["A natural next test is to train a temporally distinct representation—for example, an explicit 4D grid, a physics-based simulator, or a neural-SDF flow—on the released splits; if its future/observed gap drops to near unity on the recoverable controlled motions, the paper's family-level conclusion would be refined into an architecture-specific one.","The decoupling result implies that prior dynamic-scene forecasting papers that report only rendering metrics may be unknowingly releasing methods whose future geometry is poor; re-evaluating them with the released mesh protocol would be a high-value, low-cost extension.","The benchmark's synthetic-only design could be extended toward real captures by using a fitted proxy for future ground truth (e.g., a high-fidelity offline reconstruction) or depth sensors, though the paper's point that exact future GT requires analytic motion suggests a hybrid evaluation may be needed.","If the gap generalizes to other architecture families, it would motivate treating future-surface accuracy as a first-class benchmark axis alongside rendering, and could drive new training objectives that include temporal-extrapolation regularization or drift penalties on static futures."],"forward_implications":["The standard practice of reporting PSNR/SSIM/LPIPS on future frames does not certify future-surface quality; a separate per-frame mesh metric is needed for deployment claims.","Any method evaluated only inside the observed window may overstate its usefulness for future-time tasks such as AR overlays, robot interaction, and anticipatory planning.","The gap persists across two deformation-MLP backbones and six varied scenes, suggesting the limitation is not scene-specific but tied to the temporal model family.","The error's concentration where the surface moves implies that future-surface failure is predictable in location, and that motion-aware diagnostics (per-vertex maps) should accompany scores.","The recoverability oracle shows that some futures are known in principle from observed motion; the gap on those is a representation or extrapolation issue, not an information-theoretic one."],"fun_headline_variants":["Future geometry 2-6.6x worse than observed","NVS metrics miss future surface errors","Future surfaces diverge even when predictable","Benchmark exposes gap in future surface reconstruction"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The paper frames the measured future-surface gap and rendering–surface decoupling as a property of dynamic-scene reconstruction in general, but both tested backbones share the same time-conditioned deformation-MLP temporal model; if a temporally distinct representation extrapolates accurately, the gap would be an artifact of that architecture family rather than a field-wide finding.","fun_headline_variants_meta":{"raw":{"variants":["Future geometry 2-6.6x worse than observed","NVS metrics miss future surface errors","Future surfaces diverge even when predictable","Benchmark exposes gap in future surface reconstruction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000194,"raw_usage":{"total_tokens":1263,"prompt_tokens":891,"completion_tokens":372,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":635,"completion_tokens_details":{"reasoning_tokens":327}},"tokens_in":635,"tokens_out":372,"duration_ms":4190,"temperature":1.0,"reasoning_tokens":327,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T07:19:30.172145+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a dynamic reconstruction method with a temporally distinct representation (e.g., an explicit 4D grid or physics-based simulation) on the released FutureSurf splits; if its future Chamfer gap drops to near unity (future/observed ≈ 1) on the recoverable controlled motions and its per-frame |ρ(LPIPS,CD)| rises substantially, the paper's claim of a general field-wide decoupling would be refuted.","supporting_citations":[],"review_version":1}