{"id":"85c21886-645f-4283-b104-e5cbf1515115","arxiv_id":"2504.17812","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":2.0,"correctness_risk":"low","formal_verification":"none","parameter_count":7,"one_line_summary":"The thesis presents FlowCapsules, RobustNeRF, and SpotLessSplats, demonstrating that unsupervised object-based learning with motion and geometric consistency improves segmentation and 3D reconstruction.","lead":"This thesis compiles three published works on unsupervised object discovery and robust 3D reconstruction. It shows how motion cues and geometric consistency can segment objects and ignore distractors, improving real-world capture robustness.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Robustness claim has a hidden outlier-fraction ceiling: median-trimmed masks (Eqs. 4.8-4.10) and SLS's residual-bootstrap labels require distractors to occupy well under half the frame; Fig. 4.17 already shows degradation at 44% occupancy.","rationale":"The reader identified the central assumption as spatial coherence of photometric outliers, which is accurate but incomplete. The deeper issue is the percentile-based trimming mechanism: Eqs. 4.8-4.10 depend on the median residual labeling a majority of pixels as inliers, and Fig. 4.17 shows the default threshold fails when distractors occupy 44% of pixels. This is a concrete, load-bearing limitation of the stated robustness claim in casual captures. SpotLessSplats uses the same residual-quantile masks as weak supervision for its semantic MLP, so it inherits the risk unless the semantic features can overcome the contaminated labels. The published contributions remain valid in the regimes they were designed and tested for, and the thesis does include explicit limitation statements for other failure modes. However, because the high-occupancy ceiling is hidden rather than acknowledged, and because a single concrete experiment would settle whether SLS inherits it, the thesis-level claim should be accepted conditionally on that scoping or on a positive high-occupancy result.","tokens_in":51719,"tokens_out":11587,"duration_ms":115974,"concrete_test":"Use the existing Kubric hard generator from Fig. 4.17 and train SLS-mlp and SLS-agg with default hyperparameters on variants where distractor occupancy is 0.3, 0.4, 0.5, and 0.6 of all pixels; measure PSNR on the provided distractor-free test set and IoU of the inferred masks against the known distractor masks. If PSNR or mask IoU degrades steeply as occupancy passes about 0.4, or if SLS does not clearly beat RobustNeRF in that regime, the thesis's robustness claim must be explicitly scoped to minority-outlier captures.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"RobustNeRF's trimming is built on a per-image median residual (Eq. 4.8) and a patch-level inlier-fraction threshold T_R=0.6 (Eq. 4.10). This construction is only sound while the distractor pixel fraction p is comfortably below 1-T_R, i.e. below roughly 0.4. The hard Kubric setup in Fig. 4.17 has p=0.44, and the accompanying text states that any T_R above 0.5 gives worse results; with the default T_R=0.6 the method is already past its operating point. The thesis's 'robust 3D modelling in a casual capture setup' claim is therefore not established once a transient object such as a nearby pedestrian occupies more than about 40% of a frame. The limitation sections (Sec. 4.5, Sec. 5.5.4) acknowledge small distractors and soft shadows, but they do not acknowledge this high-occupancy ceiling. SpotLessSplats does not automatically escape it: the MLP weak labels (Eqs. 5.7-5.9) are RobustNeRF-style quantile masks from Eq. 5.2, and if the median falls inside the distractor residual distribution, the bootstrap labels are contaminated; semantic clustering may propagate rather than correct the error. The load-bearing assumption is not merely that distractors are spatially coherent photometric outliers, but that they are a minority of pixels, and this second condition is neither stated nor tested for SLS.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This PhD thesis bundles three published projects (FlowCapsules, RobustNeRF, SpotLessSplats) under the thesis that unsupervised, object-based representations improve robustness in image understanding and 3D reconstruction. Chapter 3 learns part capsules from motion via a self-supervised flow-rendering loss; Chapter 4 removes transient distractors in NeRF training through a trimmed, smoothed, patch-aggregated robust loss; Chapter 5 adapts this idea to 3D Gaussian Splatting using semantic features from a pre-trained diffusion model, with clustering or an MLP for outlier masks, plus utilization-based pruning. The empirical evaluation covers synthetic and real-world datasets (Geo, Exercise, Kubric, D2NeRF, RobustNeRF's own captures, NeRF On-the-go) with held-out clean test views, and consistently reports gains over mip-NeRF 360, D2NeRF, NeRF On-the-go, NeRF-HuGS, and vanilla 3DGS.","tokens_in":52112,"tokens_out":2438,"duration_ms":26309,"significance":"If the results hold, the thesis makes a solid contribution: RobustNeRF and SpotLessSplats provide practical, unsupervised ways to handle transient distractors in casual captures, and FlowCapsules demonstrates motion-supervised part discovery with shape completion under occlusion. The empirical work is extensive and mostly well-controlled, including paired distractor/distractor-free captures and held-out test views, and each chapter is an already-peer-reviewed publication. The methods are described in enough detail to reimplement, and several limitations (small distractors, soft shadows, semantically similar instances) are acknowledged explicitly. The reuse of published material makes the thesis a coherent anthology rather than a single new derivation, but its central claims are empirical and evaluated on external benchmarks as well as new datasets.","major_comments":[{"comment":"The central robustness claim of Chapter 4 has a hidden outlier-occupancy ceiling that is visible in the paper's own ablation. The trimming rule (4.8) uses a per-image median residual, the patch rule (4.10) requires at least 60% of a 16x16 neighborhood to be inliers, and the text around Fig. 4.17 states that in the hard Kubric setting (44% outlier pixels) any TR above 50% gives worse results. With the default TR=0.6, the method is already past its operating point on that scene, and the reconstruction degrades substantially. The limitation paragraphs in Sec. 4.5 mention small distractors and statistical inefficiency but do not acknowledge this high-occupancy ceiling, which is directly load-bearing for the claim of robust 3D modelling in casual captures. I recommend either adding a precise condition on distractor pixel fraction, or reporting a robustness curve (as in Fig. 4.17) for the natural scenes and stating the operating range explicitly in the conclusions.","section":"Sec. 4.4.5 / Fig. 4.17"},{"comment":"The SpotLessSplats masking pipeline inherits the same median-based ceiling from RobustNeRF. The weak labels U and L in Eqs. (5.8)-(5.9) are quantile masks from Eq. (5.2), which relies on a per-image/global median of residual magnitudes. If more than half the pixels in the histogram are distractors, the median falls inside the distractor residual distribution and the bootstrap labels are contaminated; the semantic clustering or MLP can then propagate the error rather than correct it. The thesis reports results on NeRF On-the-go 'high' occlusion scenes, but it does not quantify the outlier pixel fraction in those scenes, nor does it test SLS on the 44%-occupancy hard Kubric setup. Please add an explicit statement of the required inlier-majority condition, and include an experiment with controlled outlier fractions (e.g., the Kubric easy/medium/hard settings) for both SLS-agg and SLS-mlp.","section":"Sec. 5.4.1, Eqs. (5.2)-(5.4) and Fig. 5.8"}],"minor_comments":[{"comment":"The row 'No 6-Layer' in the occlusion-inductive-bias ablation is ambiguous: 'No' appears to mean 'without depth ordering', but the column header does not say so, and the reader must infer this from the surrounding text. Please rename the row to 'No depth ordering'.","section":"Sec. 3.5.4, Table 3.3"},{"comment":"There is a typo '3DSG' in the sentence 'renders the entire image in a single forward pass'; it should read '3DGS'.","section":"Sec. 5.3"},{"comment":"The phrase 'detecting and removal of the object of interest from the input images' in the abstract is confusing, because the 3D chapters remove distractors, not objects of interest. Please rephrase to 'detecting and removing transient distractors'.","section":"Abstract and Sec. 1.1"},{"comment":"The D2NeRF hyperparameter tuning is reported as 'Config 1' being best, but the tuning is only done on two of the four datasets (Statue and Crab). This is acknowledged in the text, but it would help to state explicitly that the comparison to D2NeRF is therefore not a fully tuned baseline on all scenes.","section":"Sec. 4.4.2 / Fig. 4.14"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a compilation of three peer-reviewed papers, and each chapter is individually strong. The main risk is that the thesis-level claim about robust 3D reconstruction in casual captures overstates the operating range of the robust estimators: the median-trimming construction and the T_R=0.6 threshold require distractors to occupy well under half of the pixels, and this is already visible in Fig. 4.17 but not acknowledged in the limitations. I would ask the author to either characterize this operating range quantitatively or soften the central claim accordingly. The proposed fixes are local and do not invalidate the published results, so major revision seems appropriate rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nIf you are considering this thesis for its new results, there are none: it is an explicit compilation of FlowCapsules, RobustNeRF, and SpotLessSplats, and the author says so in Section 1.1. The value is in having the three papers under one cover, with an expanded related-work chapter. That is a real value for a PhD thesis, but not for a peer-reviewed contribution.\n\nWhat is actually good: the RobustNeRF chapter is well done and has been influential. The idea of treating distractors as outliers in NeRF optimization, with a trimmed least-squares loss and spatial smoothing, is simple and effective. The empirical work is thorough: new datasets with paired clean/cluttered captures, comparison against D2NeRF and mip-NeRF 360, and ablations of the smoothing and patch terms. SpotLessSplats extends this sensibly to 3D Gaussian Splatting by adding semantic features from a diffusion model, and the ablations are careful. The methods are described well enough to reimplement.\n\nThe soft spots are real but not fatal. The stress-test note is correct: the residual-based trimming in RobustNeRF (Eqs. 4.8-4.10) only works when the distractor occupies well under the threshold T_R in each patch. The paper actually shows this in Figure 4.17: at 44% occupancy, any T_R above 50% degrades, and the default T_R=0.6 is past its operating point. The text reports this but frames it as a hyperparameter sensitivity, not as a limitation of the method's core assumption. The thesis's conclusion says RobustNeRF \"performs well on scenes with distractors,\" without mentioning the high-occupancy ceiling. SpotLessSplats' weak labels still come from residual quantiles, so it inherits a version of the problem, though the semantic clustering can mitigate it.\n\nThe other weakness is the self-citation density, but that is not a problem in a thesis where the chapters are the author's own published papers.\n\nOverall: this is a solid synthesis for a committee. If you expect new experimental results or theory, you will be disappointed. If you want a careful record of three reasonably influential methods, read it. I would put it to a serious referee only to check the limitation statements; the underlying papers have already survived review.\n\nRecommendation: engage with it for the RobustNeRF and SpotLessSplats chapters, not for novelty.","headline":"A solid thesis compiling three peer-reviewed papers; the robust-3D methods are genuinely useful, but the robustness claim has a hidden high-occupancy ceiling that the thesis never acknowledges.","tokens_in":52593,"tokens_out":2771,"would_cite":false,"duration_ms":25106,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Unsupervised object representations—learned from motion in 2D and geometric consistency in 3D—can segment objects of interest and ignore dynamic distractors, yielding cleaner 3D reconstruction from casual captures.","keywords":["unsupervised object segmentation","capsule networks","self-supervised motion learning","robust estimation","neural radiance fields","3D Gaussian splatting","distractor removal","casual capture 3D reconstruction"],"falsifier":"Capture a scene with a soft shadow cast by a moving person plus a fine static texture of similar residual magnitude; if the robust loss either leaves the shadow in or removes the texture, the spatial-coherence assumption fails.","tokens_in":51515,"feed_emoji":"🧩","tokens_out":12530,"duration_ms":105341,"temperature":0.7,"pith_summary":"This thesis argues that explicit object representations learned without supervision—by motion in 2D and geometric consistency in 3D—are the key to robust computer vision. It shows that a capsule network can parse a single image into movable parts using video pairs as training data, and that the same parts can be used for unsupervised classification and segmentation. It then shows that treating transient objects as outliers in a robust optimization makes neural radiance fields and 3D Gaussian splatting reconstruct clean static scenes from casual, distractor-filled captures. If these claims hold, 3D reconstruction from everyday photographs no longer requires clean data, manual masks, or class-specific detectors, and image understanding gains viewpoint invariance and compositional generality.","feed_headline":"Learning objects without labels cleans up 3D captures","feed_subtitle":"Motion in 2D and geometric consistency in 3D teach networks to ignore transient objects.","key_machinery":"The machinery that carries the argument is the object hypothesis itself: a capsule is a part descriptor $c_k = (s_k, \\theta_k, d_k)$ holding a canonical shape vector, a pose transformation, and a depth scalar; combined with a neural implicit decoder $D_\\omega$, it yields a visible mask $\\Lambda_k^+$ that explains a portion of the flow field $\\Phi(u)=\\sum_k \\Lambda_k^+(u)[T_k(u)-u]$. In 3D, the corresponding mechanism is a trimmed estimator with a spatial-coherence prior: residuals $\\epsilon(r)$ are thresholded at their median, blurred with a $3\\times3$ kernel, and aggregated over $8\\times8$ patches to set binary inlier/outlier weights $W(r)$, which are then used in an iteratively reweighted least-squares loss. SpotLessSplats replaces the RGB-residual classifier with a semantic-feature clustering or an MLP $H(F;\\theta)$ trained on weak labels derived from lower/upper residual thresholds, and adds utilization-based pruning to stabilize Gaussian splatting. These mechanisms turn 'objectness' into a computable, self-supervising signal.","core_discovery":"The thesis's central claim is that a network with an explicit object-based representation—parts that carry shape, pose, and depth in 2D; photometrically consistent regions in 3D—can learn to segment scenes without labels and thereby become more robust. In FlowCapsules, motion between video frames serves as the training signal: an encoder parses a single image into primary capsules, and a decoder renders their shapes; the capsules' poses and visibility masks produce a flow field that warps one frame toward the next, and optimizing this flow teaches the network which image regions are movable objects. In RobustNeRF, distractors are treated as outliers in a trimmed least-squares objective: pixels with high residual, spatially smoothed and patch-aggregated, are excluded from the NeRF loss. In SpotLessSplats, the same idea is transferred to 3D Gaussian splatting, but outlier detection uses semantic features from a text-to-image diffusion model rather than raw color, so distractors that share the background color are still masked.","pith_inferences":["Beyond the paper's demonstrations, the learned-outlier view could be pointed at transient phenomena the thesis does not test—specular highlights, rain, lens flare—as long as they are spatially coherent and photometrically inconsistent.","A testable extension is to use the FlowCapsules motion cue to generate pseudo-masks for video frames and feed those masks to a 3D reconstructor, closing the loop between 2D object learning and 3D robustness without any labels.","Because the thesis shows residual trimming is statistically inefficient on clean data, one could switch between robust and non-robust losses based on an online estimate of the clutter fraction, preserving quality when no distractors are present."],"forward_implications":["Casual captures—with pedestrians, pets, or moving shadows—can be reconstructed into clean 3D models without per-image labeling, because the optimizer itself learns which pixels to ignore.","A single robust-loss recipe transfers from NeRF-style volumetric rendering to Gaussian splatting, so distractor handling need not be re-engineered for each new scene representation.","Unsupervised part representations give viewpoint invariance and shape completion under occlusion, improving classification and segmentation on cluttered data.","Reconstruction quality degrades gracefully as the fraction of cluttered training images grows: on one scene, RobustNeRF stays above 31 dB PSNR while the base mip-NeRF 360 drops from 33 to 25 dB.","Utilization-based pruning removes floaters and cuts the number of Gaussians by a factor of 2 to 4.5 with little quality loss, even on clean scenes."],"supporting_citations":[{"why":"Supplies the psychological premise that infants group visual arrays by common motion, which FlowCapsules exploits as a self-supervision signal.","marker":"Spelke, 1990"},{"why":"PSD is the motion-based part-discovery baseline FlowCapsules compares against on part segmentation.","marker":"Xu et al., 2019"},{"why":"NeRF is the volumetric representation whose photometric-consistency assumption RobustNeRF relaxes.","marker":"Mildenhall et al., 2020b"},{"why":"mip-NeRF 360 is the base model and strongest baseline whose loss RobustNeRF replaces.","marker":"Barron et al., 2022b"},{"why":"Trimmed least squares is the robust-estimation recipe behind RobustNeRF's median-thresholded outlier weights.","marker":"Chetverikov et al., 2002"},{"why":"3D Gaussian Splatting is the explicit scene representation SpotLessSplats adapts to robust training.","marker":"Kerbl et al., 2023"},{"why":"RobustNeRF is the prior robust-masking method SpotLessSplats extends and compares against.","marker":"Sabour et al., 2023"},{"why":"NeRF On-the-go provides the casual-capture dataset and method SpotLessSplats benchmarks against.","marker":"Ren et al., 2024"},{"why":"Shows that diffusion-model features transfer to segmentation and keypoint tasks, the semantic-feature source SpotLessSplats uses for outlier detection.","marker":"Tang et al., 2023"}],"fun_headline_variants":["Unsupervised object learning gives robust 3D","Motion and geometry teach label-free 3D cleaning","No labels: transient objects vanish in 3D"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the objects worth keeping are exactly the regions that move together in 2D and look inconsistent across views in 3D, so motion and residual-based trimming can separate them from static background detail without supervision.","fun_headline_variants_meta":{"raw":{"variants":["Unsupervised object learning gives robust 3D","Motion and geometry teach label-free 3D cleaning","No labels: transient objects vanish in 3D"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001041,"raw_usage":{"total_tokens":4364,"prompt_tokens":917,"completion_tokens":3447,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":533,"completion_tokens_details":{"reasoning_tokens":3397}},"tokens_in":533,"tokens_out":3447,"duration_ms":21914,"temperature":1.0,"reasoning_tokens":3397,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:06:47.752085+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Capture a scene with a soft shadow cast by a moving person plus a fine static texture of similar residual magnitude; if the robust loss either leaves the shadow in or removes the texture, the spatial-coherence assumption fails.","supporting_citations":[{"cited_title":"P., and Hariharan, B","cited_arxiv_id":null,"evidence_quote":"Shows that diffusion-model features transfer to segmentation and keypoint tasks, the semantic-feature source SpotLessSplats uses for outlier detection."}],"review_version":1}