{"id":"9a100e99-6504-464e-bd12-2a878af57c97","arxiv_id":"2508.13228","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"PreSem-Surf combines RGB, depth, and semantic labels with an MLP pre-rendering mechanism to reconstruct scene surfaces, reporting best C-L1, F-score, and IoU on seven synthetic scenes.","lead":"This paper presents PreSem-Surf, a neural method that reconstructs 3D surfaces from RGB-D camera sequences by adding semantic object information and a pre-rendering step. Its value is faster, more accurate surface reconstruction for applications like robotics and augmented reality.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Practical-applicability claim rests solely on seven synthetic scenes; no real-world RGB-D validation is reported.","rationale":"The reader's weakest assumption identifies exactly the same fragile premise: the synthetic-only evaluation does not support the practical-applicability conclusion. I agree with that assessment. The abstract itself provides no real-world evidence, and the method's components are plausibly sensitive to real-sensor depth characteristics and noisy semantic labels. This is a load-bearing concern because the headline claim includes 'practical applicability,' not merely synthetic benchmark performance. However, the reader's verdict is already UNVERDICTED with low confidence due to insufficient information, so my concern does not move the verdict; it reinforces the existing UNVERDICTED status. A concrete real-world evaluation would settle whether the concern actually lands: if the method retains its C-L1/F-score/IoU advantages on real sensor data with noisy semantics, the applicability claim survives; otherwise it must be weakened to synthetic-domain-only. I found no internal inconsistency from the abstract alone, and I am not raising an objection about disagreement with consensus; the issue is purely an evidentiary gap in the reported experiments.","tokens_in":658,"tokens_out":3054,"duration_ms":33056,"concrete_test":"Re-run the full evaluation on at least 10 held-out real-world RGB-D sequences with ground-truth surfaces (e.g., ScanNet test scenes or TUM RGB-D) using raw sensor depth and noisy semantic predictions from a standard pretrained segmenter, comparing the same baselines across all six metrics. If the C-L1/F-score/IoU advantages do not persist outside the synthetic domain, the paper's applicability claim should be restricted to synthetic scenes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's conclusion that PreSem-Surf demonstrates 'effectiveness and practical applicability' is supported only by experiments on seven synthetic scenes. Real RGB-D input differs from synthetic rendering in ways that bear directly on the two proposed components: depth sensors produce missing values, edge noise, and systematic biases that clean renderers do not generate, and semantic supervision in practice comes from a pretrained segmenter with label noise, not from ground-truth annotations. Because the SG-MLP sampling / PR-MLP voxel pre-rendering is motivated by separating noise from local details and the progressive semantic modeling depends on semantic input quality, the reported gains in C-L1, F-score, and IoU could be artifacts of the synthetic domain. The six evaluation metrics are all computed on synthetic scenes, and the abstract reports no real-world validation, so the practical-applicability claim is underdetermined. This is an external-validity gap rather than an observed internal inconsistency.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes PreSem-Surf, a NeRF-based method for RGB-D surface reconstruction that integrates RGB, depth, and semantic information, introducing an SG-MLP sampling structure and a PR-MLP for voxel pre-rendering, plus progressive semantic modeling. The abstract reports state-of-the-art results on seven synthetic scenes across six metrics and claims practical applicability and reduced training time. The provided manuscript contains no detailed methods or experimental results beyond this abstract, so the main claims cannot be verified.","tokens_in":807,"tokens_out":3994,"duration_ms":39186,"significance":"If the performance claims are correct, PreSem-Surf would be a meaningful step toward fast semantic-aware RGB-D reconstruction, and the reported gains in C-L1, F-score, and IoU suggest the two proposed mechanisms are effective in the synthetic setting. The paper deserves consideration, but the current evidence base is too thin to judge its contribution.","major_comments":[{"comment":"The final sentence claims 'practical applicability' based solely on seven synthetic scenes. Real RGB-D inputs contain missing depth values, edge noise, and systematic sensor biases, and practical semantic supervision comes from pretrained segmenters with label noise. Neither the SG-MLP/PR-MLP components nor the progressive semantic modeling is tested under these conditions, so the external validity claim is underdetermined.","section":"Abstract"},{"comment":"The performance claim ('best performance in C-L1, F-score, and IoU, while maintaining competitive results in NC, Accuracy, and Completeness') is unreviewable without specification of baselines, dataset details, error bars, ablations, and hyperparameters. At present no experimental setup is described.","section":"Abstract"},{"comment":"The 'short time' and 'reducing training time' claims are not accompanied by any measured runtime, convergence comparison, or timing table, so the efficiency advantage over baselines cannot be assessed.","section":"Abstract"},{"comment":"No ablation isolates the contribution of SG-MLP, PR-MLP, or progressive semantic modeling; without such ablations, the attribution of the reported gains to the proposed components is not established.","section":"Abstract"}],"minor_comments":[{"comment":"The acronym 'SG-MLP' is not expanded; please define it at first use.","section":"Abstract"},{"comment":"The metric 'C-L1' is not defined; please provide its formula or a reference.","section":"Abstract"},{"comment":"The abstract does not state which baselines are compared; listing them would improve clarity.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"The version of the manuscript made available to me contains only the abstract; the body, figures, and tables are missing. I was unable to assess the methods or the experimental section. The editor should confirm that the full submission was sent for review. The external-validity concern about synthetic-only evaluation is substantial."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the abstract for PreSem-Surf describes a semantic-guided NeRF variant with two new pieces—SG-MLP sampling and PR-MLP voxel pre-rendering—plus progressive semantic modeling. That's a plausible direction, and the reported wins on C-L1, F-score, and IoU across seven synthetic scenes are worth checking. But the paper's own conclusion that this demonstrates 'practical applicability' is undercut by having no real-world RGB-D validation. That's the main soft spot.\n\nWhat looks good: the method isn't just a tweaked loss; it restructures the sampling and rendering path to use semantic cues earlier, which could plausibly help separate noise from geometry. The progressive semantic modeling idea to cut training time is also sensible. If the full paper has clean ablations and compares against a reasonable baseline set, this could be a useful engineering contribution to the NeRF-surface-reconstruction crowd.\n\nWhere I'd push back: the stress-test note gets it right—this is an external-validity gap, not an internal inconsistency. Real depth sensors give you holes, edge noise, and bias; semantic labels come from a pretrained segmenter with errors. None of that appears in clean synthetic renders. So the reported gains might not transfer. The abstract gives no real-world experiments, no discussion of sensor noise, and no ablations showing how much each component contributes. That doesn't kill the paper, but it does mean the 'practical applicability' sentence needs to be dialed back or supported.\n\nIf the full text is as the abstract suggests, I'd send it out. A referee can ask for a real-world dataset or at least a careful discussion of why the synthetic results would hold. The architecture is concrete enough to be worth a look. I just wouldn't cite it yet for any application claim.","headline":"Reasonable architectural ideas for semantic NeRF surface reconstruction, but the practical-applicability claim rides entirely on seven synthetic scenes.","tokens_in":1298,"tokens_out":1876,"would_cite":false,"duration_ms":19031,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Semantic NeRF variant out-reconstructs RGB-D baselines on seven scenes.","keywords":["neural radiance field","RGB-D surface reconstruction","semantic modeling","voxel pre-rendering","preconditioning MLP","progressive semantics","Chamfer-L1 distance","F-score"],"falsifier":"Compare PreSem-Surf against the same baselines on real RGB-D sequences that contain sensor noise, occlusions, and imperfect depth, using the same six metrics; if Chamfer-L1, F-score, and IoU no longer beat the baselines by the reported margins, the paper's practical-applicability claim is false.","tokens_in":480,"feed_emoji":"🧊","tokens_out":5280,"duration_ms":49628,"temperature":0.7,"pith_summary":"PreSem-Surf is a neural surface reconstruction method built on the Neural Radiance Field idea: it takes RGB-D video and reconstructs the scene as a surface, adding semantic labels as an extra guide. The paper argues that letting the network name what it sees—walls, chairs, floors—while it learns geometry helps it separate real surface detail from depth noise. To make that work quickly, the method pre-renders a coarse voxel representation with a preconditioning multilayer perceptron and then refines semantics progressively. On seven synthetic scenes, the paper reports the best results for Chamfer-L1 distance, F-score, and intersection-over-union, with competitive normal consistency, accuracy, and completeness, and shorter training time. If true, this makes semantic awareness a practical ingredient for fast, accurate RGB-D surface reconstruction.","feed_headline":"Semantic NeRF variant out-reconstructs RGB-D baselines on seven scenes","feed_subtitle":"Pre-rendering voxels with a semantic MLP speeds training and sharpens surface geometry.","key_machinery":"The mechanism that carries the argument is voxel pre-rendering: an SG-MLP sampling structure combined with a PR-MLP, or preconditioning multilayer perceptron, renders a coarse voxel-level view of the scene before the full surface field is computed. This early rendering lets the model absorb global scene structure and distinguish noise from local detail. Progressive semantic modeling accompanies it, extracting semantic information at increasing precision levels, which reduces training time while improving scene understanding. Together these two components are what the paper credits for the reported gains.","core_discovery":"The central claim is that jointly modeling RGB, depth, and semantic information inside a NeRF-style surface reconstruction pipeline improves both reconstruction quality and training speed. PreSem-Surf introduces an SG-MLP sampling structure and a PR-MLP preconditioning multilayer perceptron for voxel pre-rendering, so the model captures scene-level information before committing to fine detail; progressive semantic modeling then feeds labels at increasing precision instead of all at once. Reported on seven synthetic scenes with six metrics, the method achieves the best Chamfer-L1, F-score, and IoU and remains competitive on normal consistency, accuracy, and completeness. The authors present this as evidence that the method is effective and practically applicable.","pith_inferences":["A natural next test is real-world RGB-D data: if the gain shrinks under sensor noise, occlusions, and imperfect depth, the practical-applicability claim would need to be scaled back.","The semantic-preconditioning idea could transfer to monocular depth reconstruction, where the depth signal is noisier and scene semantics could carry more weight.","Replacing semantic labels with automatically generated pseudo-labels would test whether the gains come from the semantic signal itself or from the progressive training schedule."],"forward_implications":["If the reported results hold, RGB-D surface reconstruction can be made faster by adding a semantic preconditioning stage, without sacrificing geometric accuracy.","Semantic labels can serve as a regularizer that suppresses depth noise in NeRF-based surface fitting.","The six-metric evaluation on seven scenes gives future RGB-D reconstruction work a direct baseline to compare against.","Progressive semantic modeling offers a concrete way to trade training time against scene understanding in neural reconstruction."],"supporting_citations":[],"fun_headline_variants":["Semantic NeRF speeds surface reconstruction while boosting precision","SG-MLP pre-rendering cuts training time and sharpens RGB-D surfaces","Semantic MLP and progressive modeling: faster, finer 3D surfaces","NeRF with semantic pre-rendering tops RGB-D surface benchmarks","Progressive semantic NeRF reconstructs scenes faster and better"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the seven synthetic scenes represent real-world RGB-D conditions well enough that the accuracy and speed gains carry over to actual sensor data.","fun_headline_variants_meta":{"raw":{"variants":["Semantic NeRF speeds surface reconstruction while boosting precision","SG-MLP pre-rendering cuts training time and sharpens RGB-D surfaces","Semantic MLP and progressive modeling: faster, finer 3D surfaces","NeRF with semantic pre-rendering tops RGB-D surface benchmarks","Progressive semantic NeRF reconstructs scenes faster and better"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000359,"raw_usage":{"total_tokens":1892,"prompt_tokens":843,"completion_tokens":1049,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":459,"completion_tokens_details":{"reasoning_tokens":972}},"tokens_in":459,"tokens_out":1049,"duration_ms":8184,"temperature":1.0,"reasoning_tokens":972,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:21:37.587669+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare PreSem-Surf against the same baselines on real RGB-D sequences that contain sensor noise, occlusions, and imperfect depth, using the same six metrics; if Chamfer-L1, F-score, and IoU no longer beat the baselines by the reported margins, the paper's practical-applicability claim is false.","supporting_citations":[],"review_version":1}