{"id":"ad77c960-6d04-48c2-ad8e-0363b5c5d15b","arxiv_id":"2411.12635","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A dual-stream Mamba plus depth network for single-view 3D reconstruction reports state-of-the-art 3D-FRONT scores, but its experimental evidence is inconsistent and missing controlled baselines.","lead":"M3D is a neural-network system that reconstructs a 3D object from a single photo by processing color and depth in separate branches and merging them with a state-space model. It reports large gains on the 3D-FRONT benchmark, but the paper's tables are internally inconsistent and lack fair comparisons, so the state-of-the-art claim is not established.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table III's SOTA comparison is not controlled: no dataset, resolution, or metric protocol is stated, and M3D's CD drops from 6.60 (Table I) to 0.0066 while PSNR jumps ~10 dB above every listed baseline, indicating a different benchmark; without a shared protocol the headline claim collapses.","rationale":"The reader's weakest assumption identifies Table III comparability, and I agree. The central claim is SOTA performance, and SOTA must be measured against the strongest currently relevant baselines; Table III is where that comparison happens. If the baselines are not run under the same protocol, the table is meaningless, regardless of how plausible the architecture is. The numbers are strongly indicative of a mismatched benchmark: M3D's CD changes by three orders of magnitude between Table I and Table III, and the PSNR gap is far beyond typical method-to-method variation. The paper provides no dataset name for Table III, no code URL despite claiming public release in the abstract, and no error bars or significance tests. These are not style nits; they bear directly on whether the claimed SOTA is evidence or a formatting artifact. The internal inconsistency between the 36.9% CD improvement claimed in Section I and Table I and the 53.9% improvement in Table V further weakens the empirical case. A conditional acceptance might be possible if the authors supplied reproducible code and a unified evaluation protocol, but as written the evidence does not support the headline. Therefore the reader's REJECT verdict remains appropriate.","tokens_in":11957,"tokens_out":4134,"duration_ms":38433,"concrete_test":"Request the authors' code and evaluation scripts; re-run OpenLRM, InstantMesh, CRM, Unique3D, and Wonder3D+ISOMER under the exact protocol used for M3D on the same test split, with the same CD normalization, rendering resolution, and PSNR formula. If the numbers in Table III cannot be reproduced in a single shared evaluation, the SOTA claim is void.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is state-of-the-art single-view 3D reconstruction. The most load-bearing evidence for that claim is Table III, which compares M3D with OpenLRM, InstantMesh, CRM, Unique3D, and Wonder3D+ISOMER. That table lists no dataset, no input resolution, no rendering protocol, no metric definitions, and no training or evaluation status for the baselines. The numbers themselves signal an uncontrolled comparison: the same M3D model reports CD 6.60 in Table I on 3D-FRONT but CD 0.0066 in Table III; the listed baselines' PSNR values (18.0-20.1 dB) are roughly 10 dB below M3D's 30.04 dB, a gap that is implausible on identical data and typical of different renderings, different CD normalization, or different image and depth scale. Because Table III is the only evidence against modern large reconstruction models, and its comparability premise is asserted only by placing numbers side by side, the SOTA conclusion is unsupported. Additionally, ablation Table V reports a 53.9% CD improvement at epoch 120, which conflicts with the 36.9% improvement claimed from Table I, reinforcing that the reported effect sizes are not internally consistent.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes M3D, a single-view 3D reconstruction method that uses a dual-stream architecture: an RGB stream built on selective state-space models with residual and attention components, and a depth stream based on pre-trained monocular depth maps. The two streams are fused with an implicit SDF representation and volumetric rendering, trained with a two-stage loss. The paper claims state-of-the-art performance on 3D-FRONT, with headline gains of 36.9% in Chamfer Distance, 13.3% in F-Score, and 5.5% in Normal Consistency over SSR, and also reports favorable comparisons with recent large reconstruction models. The contribution is primarily architectural: balancing global and local feature extraction through a dual-stream SSM design while injecting depth information.","tokens_in":12199,"tokens_out":9731,"duration_ms":93229,"significance":"The architectural idea is credible and the ablation sequence directly targets the stated design choices: enhanced shallow residual blocks, a depth branch, and an SSM-based deep feature module. Training hyperparameters and loss weights are disclosed, which aids reproducibility. However, the central SOTA claim is not supported by the current evidence: the comparison with recent models in Table III is not a controlled experiment, and several internal inconsistencies in the reported tables make the quantitative claims unreliable. The qualitative figures cannot compensate for these problems. If the authors can supply a common evaluation protocol and reconcile the tables, the method would be a plausible contribution, but as written the empirical case is not established.","major_comments":[{"comment":"The comparison in Table III is not controlled and cannot support the central SOTA claim. The table does not state the dataset, input resolution, rendering protocol, mesh extraction settings, CD normalization, or whether the baselines were retrained or evaluated under the same conditions as M3D. The same model reports CD 6.60 on 3D-FRONT in Table I but CD 0.0066 in Table III, and the baseline PSNR values all lie about 10 dB below M3D's 30.04. These signatures strongly suggest that the numbers were not produced on the same benchmark. Since Table III is the only comparison against OpenLRM, InstantMesh, CRM, Unique3D, and Wonder3D, the state-of-the-art conclusion is unsupported unless all baselines are rerun under one common protocol and that protocol is specified in the paper.","section":"Section IV-B, Table III"},{"comment":"The ablation results in Table V contradict the headline improvements. Table V reports M3D's CD as 7.61 at epoch 120 and the SSR baseline as 16.51, which is a 53.9% improvement, while Section I and Table I report CD 6.60 versus 10.45, i.e., 36.9%. The same mismatch appears for F-Score (32.7% vs 13.3%) and NC (12.7% vs 5.5%). If the two baselines are different training checkpoints, evaluation conditions, or splits, that must be stated; otherwise the paper's reported effect sizes are internally inconsistent and the ablation conclusions are not reliable.","section":"Section V-B, Table V; Section I; Table I"},{"comment":"The 'mean' column in Table I is not reproducible from the eight per-category numbers shown. For example, the arithmetic mean of the SSR CD row is 11.32, not the reported 10.45; for M3D it is 8.88, not 6.60; several F-Score and NC rows show similar mismatches. If the means are weighted by object frequency or by some other factor, the weighting must be given; without this, the percentage improvements quoted in the abstract and Section I cannot be verified.","section":"Table I"},{"comment":"Table II compares M3D with Zero-1-to-3 and Shape-E without stating the dataset or evaluation protocol. The reported M3D values (CD 6.60, F-Score 80.85, NC 0.901) match Table I's 3D-FRONT numbers, but no evidence is provided that Zero-1-to-3 and Shape-E were evaluated on the same 3D-FRONT split with the same metrics. This comparison is load-bearing because it is used to argue superiority over methods with large-scale pre-trained priors.","section":"Section IV-B, Table II"}],"minor_comments":[{"comment":"The paper states that code and dataset are available at 'this URL', but no URL appears anywhere in the manuscript; please provide the actual repository link.","section":"Abstract and Section I"},{"comment":"There is a sentence fragment beginning 'blocks [7, 8] for shallow feature extraction contribute to a 56% reduction in CD'; the preceding text appears to have been cut and should be rewritten.","section":"Section IV-B"},{"comment":"The symbol beta is used both as the learnable SDF-to-density temperature in Eq. (6) and as a loss weight in Eq. (13); please use distinct symbols.","section":"Equations (6) and (13)"},{"comment":"The caption states that F-Score values are divided by 100 for visualization, but the F-Score values in Table V are already in [0,1]; please clarify the scaling so the axes can be interpreted.","section":"Figure 4 caption"},{"comment":"The column headers 'CD ↓ (51.0%)', 'F-Score ↑ (33.6%)', and 'NC ↑ (10.3%)' are unexplained; if these are aggregation weights or other factors, define them in the caption.","section":"Table I"},{"comment":"References [2] and [18] share the same title and are cited in similar contexts; please verify that they are distinct works and update the citation labels and reference list accordingly.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The required changes are substantial: the authors must rerun all external baselines under a common protocol, reconcile Table V with Table I, explain the Table I means, and make code and checkpoints available. If the authors cannot deliver a controlled comparison, the paper should be rejected rather than accepted after minor edits. The architectural idea is plausible, but the current numerical evidence is too unreliable to publish."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is an engineering paper that combines known components into a new system for single-view 3D reconstruction: a Mamba-style selective SSM as a dual-stream RGB feature extractor, a Depth Anything depth branch, and SDF-based volumetric rendering. The architecture is new in its cited literature, and the ablation study is structured and shows monotonic gains from each module. That part is reasonably solid.\n\nThe punchline is that the experimental evidence for the headline SOTA claim does not hold up. Table III, the only comparison against modern large reconstruction models (OpenLRM, InstantMesh, CRM, Unique3D, Wonder3D+ISOMER), lists no dataset, no resolution, no metric definitions, and no training/eval status for the baselines. The numbers are implausible as a controlled comparison: M3D's PSNR is 30.04 while every baseline sits at 18–20, a gap that strongly suggests different benchmarks. CD values are also scaled by 1000 between Table I (6.60) and Table III (0.0066), which the paper never explains. Even within the paper, the claimed improvements drift: the abstract/Table I give 36.9% CD improvement over SSR, the ablation Table V shows 53.9% for the same baseline; NC is 5.5% in Table I and 12.7% in Table V. These could reflect a re-implemented baseline, but the paper doesn't say so.\n\nThere's also a missing technical detail that matters: Depth Anything outputs affine-invariant depth, and the paper never explains how those maps are aligned to the SDF or the ground truth before the depth consistency loss is applied. Without that, the depth branch's contribution is unclear. Minor sloppiness is pervasive: no code URL despite claiming code is public, a mis-cited reference for LDIF, vague loss-weight scheduling.\n\nWhat is worth preserving: the dual-stream SSM-plus-depth design is a reasonable contribution to this task, and the ablations indicate each element helps. But the empirical case is not there. The authors need to rerun baselines on a shared protocol, state units and alignment procedures, and reconcile their tables. As written, I'd reject. If the authors meaningfully revise, the idea might become publishable.\n\nFor peer review: I'd still send it to referees because the system is concrete and the flaws are identifiable and fixable, but I'd expect a negative or major-revision outcome. For your own work, I wouldn't cite it in this form.","headline":"Plausible architecture, unsupported SOTA claim: Table III lacks protocol and inner numbers conflict.","tokens_in":12798,"tokens_out":5750,"would_cite":false,"duration_ms":56775,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A dual-stream architecture that combines a selective state-space model with a depth branch is claimed to achieve state-of-the-art single-view 3D reconstruction, with a 36.9% lower Chamfer Distance and 13.3% higher F-score than the SSR…","keywords":["single-view 3D reconstruction","implicit neural representation","selective state space model","depth estimation","dual-stream feature extraction","3D-FRONT dataset","signed distance function","3D vision"],"falsifier":"Run M3D and the listed comparison methods on one shared 3D-FRONT test image using identical mesh extraction and metric code, and check whether M3D's Chamfer Distance is around 6.60 (as in the main table) or around 0.0066 (as in the large-model comparison); reproducing either value, and seeing whether the baselines reproduce their published values, settles which comparison is meaningful.","tokens_in":11694,"feed_emoji":"🧊","tokens_out":7006,"duration_ms":65738,"temperature":0.7,"pith_summary":"This paper tries to establish that a single-view 3D reconstruction network can achieve high-fidelity indoor object geometry by processing RGB and depth in two separate streams and fusing them with a selective state-space model. On the 3D-FRONT dataset, the proposed M3D framework reports a Chamfer Distance of 6.60, an F-score of 80.85, and a Normal Consistency of 0.901, corresponding to a 36.9% CD improvement, 13.3% F-score gain, and 5.5% NC gain over the SSR baseline. The motivation is that CNNs capture local details but miss global context, while transformers capture global context at quadratic cost; the paper argues a selective SSM plus explicit depth cues gets both. If correct, the design gives a practical recipe for preserving fine details and occlusion robustness in indoor scene reconstruction for VR, robotics, and driving.","feed_headline":"Dual-stream state-space model lifts single-view 3D fidelity","feed_subtitle":"M3D reports a 36.9% lower Chamfer Distance and a 13.3% higher F-score than the SSR baseline on 3D-FRONT.","key_machinery":"M3D's load-bearing mechanism is the dual-stream Selective Attention Module combined with a depth-driven geometric stream. In the RGB stream, high-dimensional features are channel-split: one half is processed by a selective state-space model (a sequence model that scans tokens in linear time and selectively retains relevant context), the other by convolutional residual blocks, and the two are merged and refined by self-attention; this iteration is repeated with fewer layers to refine global and local features. In parallel, a pretrained monocular depth estimator produces a depth map offline, and the depth feature is combined with shallow RGB features via generalized addition and bilinear interpolation onto 2D-projected 3D points, then concatenated with the point coordinate and passed through an MLP. The implicit SDF is converted to density via a learnable $\\beta$ for differentiable volume rendering, and the loss combines 3D SDF, color, depth consistency, and normal consistency terms.","core_discovery":"On the paper's own terms, the central discovery is that the global/local feature tradeoff in single-view 3D reconstruction is best resolved by a dual-stream design: one stream runs a Selective Attention Module that splits features and processes half with a selective state-space model and half with convolutional residual blocks before self-attention, while a parallel stream injects an offline-estimated depth map through generalized addition and bilinear projection onto 3D points. The fused RGB-depth feature is decoded by an implicit signed-distance-field representation with volumetric rendering and supervised by geometry, color, depth, and normal losses. The paper reports that the depth branch alone reduces Chamfer Distance by 28.6% relative to the SSR baseline at epoch 120, and the full M3D reduces it by 53.9%, with the selective SSM credited for smoother normals and faster convergence. The claimed result is state-of-the-art reconstruction fidelity on the 3D-FRONT benchmark across CD, F-score, NC, PSNR, and IoU.","pith_inferences":["Because the depth stream leans on a specific pretrained estimator, M3D's fidelity gains may track the accuracy of that estimator; swapping in other monocular depth models would test how much of the improvement is depth quality versus architecture.","The same dual-stream fusion could transfer to multi-view or video reconstruction, where depth is easier to obtain and the SSM's linear scan could aggregate temporal context.","The 3D-FRONT results are for indoor furniture categories; applying M3D to outdoor or deformable objects would clarify whether the benefit is specific to indoor geometry."],"forward_implications":["If the reported gains hold, separating RGB and depth streams is a more effective way to use monocular depth in neural implicit reconstruction than joint single-stream feature extraction.","The linear-time selective SSM component suggests the framework can scale to higher image resolutions than transformer-only backbones without quadratic attention cost.","Offline depth estimation keeps training costs low while still providing geometric cues, so the architecture is feasible on a single GPU with batch size 30.","Each added component (enhanced residual blocks, depth branch, selective attention) contributes positively in ablations, implying the gains are compositional rather than from one module alone."],"supporting_citations":[{"why":"Supplies the selective state-space model that forms the core of the RGB feature stream.","marker":"[3]"},{"why":"Provides the 3D-FRONT dataset and the training/evaluation splits used for the main results.","marker":"[6]"},{"why":"Is the SSR baseline whose setup M3D extends and against which the headline improvements are computed.","marker":"[35]"},{"why":"Supplies the offline monocular depth maps used by the depth stream.","marker":"[41]"},{"why":"Provides the implicit neural representation and volumetric rendering used for decoding.","marker":"[24]"},{"why":"Motivates the hybrid SSM-CNN-self-attention Selective Attention Module.","marker":"[7]"},{"why":"Provides the ResNet backbone used for shallow features in M3D and in the baseline.","marker":"[9]"}],"fun_headline_variants":["Dual-stream state-space model with depth lifts single-view 3D fidelity","Selective state-space dual stream with depth beats SSR on 3D-FRONT","Depth-boosted dual-stream SSM cuts Chamfer Distance 53.9%","State-space dual-stream with depth: new SOTA for single-view 3D","M3D: dual selective SSM branches plus depth for high-fidelity 3D"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim assumes the comparison numbers for other methods in Table III were produced on the same evaluation set, resolution, and metric implementation as M3D's numbers; if they were not, the reported state-of-the-art advantage is not established.","fun_headline_variants_meta":{"raw":{"variants":["Dual-stream state-space model with depth lifts single-view 3D fidelity","Selective state-space dual stream with depth beats SSR on 3D-FRONT","Depth-boosted dual-stream SSM cuts Chamfer Distance 53.9%","State-space dual-stream with depth: new SOTA for single-view 3D","M3D: dual selective SSM branches plus depth for high-fidelity 3D"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000551,"raw_usage":{"total_tokens":2619,"prompt_tokens":927,"completion_tokens":1692,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":1583}},"tokens_in":543,"tokens_out":1692,"duration_ms":11012,"temperature":1.0,"reasoning_tokens":1583,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T17:19:50.884631+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run M3D and the listed comparison methods on one shared 3D-FRONT test image using identical mesh extraction and metric code, and check whether M3D's Chamfer Distance is around 6.60 (as in the main table) or around 0.0066 (as in the large-model comparison); reproducing either value, and seeing whether the baselines reproduce their published values, settles which comparison is meaningful.","supporting_citations":[{"cited_title":"3d-front: 3d furnished rooms with layouts and semantics","cited_arxiv_id":null,"evidence_quote":"Provides the 3D-FRONT dataset and the training/evaluation splits used for the main results."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the SSR baseline whose setup M3D extends and against which the headline improvements are computed."},{"cited_title":"Srinivasan, Matthew Tancik, Jonathan T","cited_arxiv_id":null,"evidence_quote":"Provides the implicit neural representation and volumetric rendering used for decoding."}],"review_version":1}