{"id":"c5b8b6ae-1845-4df3-aae7-ac575c888c3b","arxiv_id":"1908.08185","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A projector-camera system estimates dense 3D geometry and per-point spectral reflectance from off-the-shelf hardware by modeling geometric shading in a Lambertian reflectance optimization.","lead":"This paper builds a low-cost 3D scanner from an ordinary projector and RGB camera that captures both shape and true spectral color of small objects. It combines self-calibrating structured-light 3D reconstruction with a shading-aware spectral reflectance model, demonstrating accurate reflectance recovery on standard color charts and sculptures.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Color-chart evaluation confounds the geometric shading model with multi-view averaging; the central photometric contribution is not isolated.","rationale":"The reader's stated weakest assumption is the Lambertian/no-ambient/no-interreflection shading model, which is acknowledged in the paper and mainly affects concave regions. The reader's rationale also notes the single-view versus multi-view comparison confound, and I believe that confound is the more load-bearing issue for the central claim. The paper is a well-engineered system, and the mathematical derivation of Eq. 10 is internally consistent, but the experimental design does not isolate whether the geometric shading model or the use of four projector-camera pairs produces the reported improvement. This is a validation gap rather than a demonstrated error, so it does not justify rejection; it strengthens the need for a controlled ablation before accepting the photometric contribution as established. A concrete ablation comparing the proposed model with and without the geometric shading factor, while controlling the number of pairs, would settle whether the central claim holds. If the ablation shows the shading model is responsible, the paper's contribution stands; if not, the claim should be substantially weakened.","tokens_in":13310,"tokens_out":7246,"duration_ms":85039,"concrete_test":"On the colorchart data, run four variants using the same reflectance cost (Eqs. 11-12): (A) the proposed full model with all four pairs; (B) the same optimization but with the shading factor s_k replaced by a per-pair constant, or preferably a per-pair smooth shading field estimated without using the estimated geometry; (C) the proposed model using only pair 4; (D) baselines [18] and [5] applied independently to each pair, with their reflectance estimates averaged over the four pairs. If the RMSE of B or D is close to A while C is much worse, then multi-view averaging, not Eq. 10, drives the reported improvement. This test directly isolates the paper's main photometric contribution.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4.3 compares the proposed method using all four projector-camera pairs against the baselines [18] and [5] using only one pair (pair 4), as shown in Fig. 8(c)-(e). This conflates two distinct factors: the proposed geometry-based shading factor of Eq. 10 and the benefit of averaging observations from four view/light positions through the rendering cost E_ren in Eq. 12. The central claim that Eq. 10 enables recovery of 'inherent' spectral reflectance is therefore not established by the reported RMSE comparison. If the single-view baselines are each biased by shading in different directions, averaging four independent estimates could reduce RMSE even without any geometric model. The paper's own Section 4.3 states that without shading, its accuracy would be similar to existing methods, so isolating the shading model is essential to the novelty claim. The Lambertian/no-interreflection limitation is acknowledged by the authors and appears primarily in concave regions (Fig. 9(d)); it is real but less threatening to the central contribution than this evaluation confound.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents Pro-Cam SSfM, an acquisition system that uses an off-the-shelf RGB camera and an off-the-shelf projector for simultaneous dense 3D reconstruction and per-point spectral reflectance estimation. The camera and projector are alternately moved around the object; structured-light codes are matched across all cameras and projectors and fed to an SfM pipeline with weighted bundle adjustment (Eq. 1). The estimated projector poses and reconstructed 3D points and normals are then used in a physically-motivated shading model (Eq. 10) inside a linear least-squares reflectance estimation with a spectral smoothness regularizer (Eqs. 11-13). Experiments on a color chart and two real objects report denser point clouds than the prior method [30] and lower spectral-reflectance RMSE than two single-view baselines.","tokens_in":13434,"tokens_out":4744,"duration_ms":49884,"significance":"If the central claims hold, this is a practically useful low-cost spectral-3D scanning system: it avoids dedicated multispectral cameras or light sources, and it replaces fitted or ignored shading with a geometry-based shading factor derived from the inverse-square law and a Lambertian assumption. The derivation of Eq. (10) is sound and the reflectance optimization is a standard regularized linear inverse problem. The paper also offers a concrete demonstration that projector poses estimated during SfM can be reused for photometric modeling, which is a nice integration of geometric and photometric pipelines. However, the empirical validation currently conflates the proposed shading model with multi-view averaging, so the central claim that Eq. (10) enables recovery of inherent spectral reflectance is not yet established. The paper does not release code or data, which limits reproducibility, but the experimental protocol is otherwise reasonably standard.","major_comments":[{"comment":"The comparison is unequal: the proposed method minimizes E_ren over all four projector-camera pairs, whereas the single-view baselines [18] and [5] are applied to projector-camera pair 4 only. This conflates the geometric shading factor of Eq. (10) with the benefit of averaging four independent observations, since each single-view estimate carries a different shading bias and averaging them could reduce RMSE even without a geometric model. To support the central claim, the authors should add an ablation: (i) the proposed method with the shading factor disabled (or set to 1) using the same four pairs; (ii) the proposed method using only pair 4; and, ideally, (iii) the baselines applied to all four pairs and averaged. The paper's own statement in Section 4.3 that without shading its accuracy would be similar to existing methods underscores that isolating the shading model is essential to the novelty claim.","section":"Section 4.3, Fig. 8(d)-(e), Eq. (12)"},{"comment":"The model assumes Lambertian reflectance, negligible ambient light, and no interreflection. These are not merely boundary conditions: because the shading factor in Eq. (10) is computed deterministically from geometry rather than fitted, any violation of these assumptions is absorbed into the estimated reflectance, biasing the 'inherent spectral reflectance' claim. The paper acknowledges that interreflections cause errors in concave skirt regions of the clay sculpture, but it does not quantify the resulting reflectance bias. The authors should either add a controlled experiment on a concave or non-Lambertian target and report reflectance error as a function of concavity or material gloss, or explicitly restrict the claim to weakly interreflective matte scenes.","section":"Section 3.3.3, Fig. 9(d), Concluding Remarks"},{"comment":"The band selection is performed on the same color chart used for the final RMSE evaluation: for each number of bands the paper reports the 'selected best band set' chosen by evaluating all possible band sets, which is a test-set selection and can overstate accuracy. The authors should use a separate validation set or a fixed predefined band set, or at least report the performance of a fixed band set (e.g., the six-band combination) on held-out patches or objects.","section":"Section 4.3, Fig. 7"}],"minor_comments":[{"comment":"The first sentence contains the typo 'ﬁst' and should read 'first'.","section":"Concluding Remarks"},{"comment":"The word 'Lambartian' should be 'Lambertian'.","section":"Section 3.3.3"},{"comment":"The statement that Ceres is used to solve the 'non-linear optimization problem' of Eq. (11) is misleading because with known shading factors the problem is linear in the coefficient vector α_k; please clarify or justify the nonlinearity.","section":"Section 4.1, Eq. (11)"},{"comment":"Please clarify how baseline [5] is adapted to this setup: [5] is designed for single RGB images, so it should be stated whether it receives the RGB response under one illumination or the full 21-band vector, since that determines the fairness of the comparison.","section":"Section 4.3"},{"comment":"The threshold for connecting features from different projectors is defined in pixels but the image resolution and scale are not stated; please report the resolution and discuss how the 0.5-pixel threshold affects the number and accuracy of correspondences.","section":"Section 3.2.1"},{"comment":"Please add axis labels and units to the plots of camera spectral sensitivity and illumination spectral power distributions; the current figure is difficult to interpret without them.","section":"Figure 4"},{"comment":"The clay sculpture and stuffed toy results are presented qualitatively; please provide quantitative spectral-reflectance RMSE against ground truth for these objects, or state explicitly that only qualitative evaluation is intended.","section":"Section 4.4"},{"comment":"The smoothness term E_ssm should be written with an explicit norm and the operator D should be defined precisely (e.g., a second-difference matrix) so that the dimension of the regularization is unambiguous.","section":"Eq. (13)"}],"recommendation":"major_revision","confidential_remarks":"The core physical derivation is sound and the system concept is valuable, so I do not see grounds for rejection. The main risk is the evaluation confound identified in major comment 1, which can be addressed with additional experiments without changing the method. The self-citation to [30] is appropriate because that paper is the direct geometric baseline. I would not require a fully controlled psychophysical study, but the shading-model ablation is essential before the 'inherent spectral reflectance' claim can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, this is a well-engineered system that adds spectral reflectance estimation to a self-calibrating projector-camera 3D scanner. The core idea — computing a per-point shading factor from the reconstructed geometry rather than baking shading into the reflectance — is physically sound and clearly derived. Second, the main quantitative comparison does not isolate that idea: the proposed method uses all four projector-camera pairs, while the baselines use only one, so the reported RMSE gain could partly come from multi-view averaging. Read the paper for the system and the derivation, but treat the numerical claim as unproven as reported.\n\nWhat is new: including projectors as features in the SfM pipeline roughly doubles the point count over the authors' earlier work, and the shading-aware spectral estimation model is a reasonable extension of [18]. The shading factor in Eq. 10 is computed from geometry, not fitted, and the visibility handling is sensible. The relighting demonstrations on the sculpture are visually convincing; the baked-in shading is visibly reduced.\n\nSoft spots. The evaluation confound is the main one. Fig. 8 compares the all-four-pair proposed method against pair-4-only baselines. The authors even state that without shading their accuracy would be similar to existing methods, so the experiment should have included a multi-view baseline that averages the single-view estimates. Without that, the RMSE numbers don't demonstrate the shading model's contribution. There is also some tuning on the evaluation chart: the band set in Fig. 7 is selected by lowest RMSE on the same chart, and gamma, wp, and basis count are empirically chosen without a sensitivity analysis. No code or data is released, which hurts reproducibility. The Lambertian and no-interreflection assumptions are acknowledged and mainly affect concave regions, as the authors note.\n\nWho this is for: people building practical spectral 3D acquisition systems, or anyone working on photometric estimation under active light. The paper deserves a serious referee; the flaw is fixable with a proper ablation and a multi-view baseline. I would send it to review rather than desk-reject.","headline":"Solid system paper with a clean shading model, but the main experiment confounds multi-view averaging with the proposed geometric shading term.","tokens_in":14025,"tokens_out":2102,"would_cite":true,"duration_ms":22004,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A projector and camera moving around an object can recover both its 3D shape and the per-point light spectrum.","keywords":["spectral reflectance","structure from motion","structured light","projector-camera system","multispectral imaging","3D reconstruction","self-calibration","relighting"],"falsifier":"Capture a matte, concave object (such as the inside of a cup) with known spectral reflectance using this system and compare the estimated reflectance in the concave region against a spectrometer reading. A systematic bias that grows toward the concavity, where interreflections are strongest, while convex regions match, would falsify the no-interreflection assumption; a glossy object whose estimated reflectance deviates would falsify the Lambertian assumption.","tokens_in":13057,"feed_emoji":"📽️","tokens_out":5310,"duration_ms":45992,"temperature":0.7,"pith_summary":"The paper claims that a single off-the-shelf projector and a standard RGB camera, moved alternately around an object, can acquire both a dense 3D model and the inherent spectral reflectance of every reconstructed surface point. The projector plays a dual role: it projects gray-code patterns for self-calibrating multi-view structured-light reconstruction, and uniform color illuminations for multispectral observations. The central new idea is a spectral reflectance estimation model that divides out a geometry-based shading factor, computed from each 3D point's position, its normal, and the estimated projector location, so that shading and cast shadows are not baked into the recovered reflectance. If correct, this makes full spectral 3D scanning practical with inexpensive hardware and no pre-calibration.","feed_headline":"A projector and camera recover 3D shape and per-point spectra","feed_subtitle":"A self-calibrating system estimates each point's spectral reflectance while removing shading and shadow effects.","key_machinery":"The load-bearing object is the shading model of Eq. (10): s_k = (p_pro - p_k)/||p_pro - p_k||^3 · n_k, which expresses the fraction of projected light that reaches the camera as the dot product of the normalized lighting direction with the surface normal, scaled by inverse-square distance attenuation. It turns the projector's estimated position into a geometric prediction of shading, so the optimization over reflectance coefficients does not have to absorb per-view brightness variations. The second piece is the weighted bundle adjustment (Eqs. 1–2), which gives projector reprojection errors a larger weight (wp=100) so that the estimated projector poses, which feed the shading model, stay stable.","core_discovery":"The paper's central claim is that the shading factor at a surface point can be computed directly from the estimated geometry as s_k = (p_pro - p_k)/||p_pro - p_k||^3 · n_k, combining the Lambertian cosine law with the inverse-square falloff of a nearby point light source. Substituting this into a linear rendering model y = s C^T L B α lets the system solve for spectral reflectance coefficients α per 3D point while eliminating the shading and shadow effects that single-view methods 'bake in.' The authors demonstrate on real objects (a color chart, a clay sculpture, a stuffed toy) that the resulting reflectance curves match spectrometer ground truth and support relighting under novel light directions and spectra.","pith_inferences":["The residual between the model's predicted and observed intensities could be mined as a signal for non-Lambertian or interreflection effects, pointing toward a future extension that estimates a per-point BRDF or global illumination rather than assuming Lambertian shading.","The system's reliance on a known illumination spectrum (measured by a spectrometer) could be relaxed by jointly estimating the projector's spectral power distribution, turning the hardware into a fully self-contained scanner.","Because the shading model uses the projector as a calibrated moving light source, Pro-Cam SSfM is effectively a self-calibrating photometric-stereo setup; the same geometric term could be used for normal refinement or for estimating the projector's radiometric falloff.","The method's per-point reflectance estimates, once corrected for interreflections, could serve as ground truth for single-image spectral recovery methods, which currently train on 'baked-in' shading."],"forward_implications":["Dense spectral 3D acquisition becomes possible with off-the-shelf projector and camera hardware, removing the need for calibrated multispectral cameras or light sources.","The recovered reflectance is a property of the surface, not of the lighting or viewpoint, so the same model supports relighting under arbitrary light directions and spectra.","Including projector pixel correspondences in the SfM bundle adjustment roughly doubles the reconstructed point-cloud density compared with camera-only correspondences.","The shading-aware estimation reduces the number of spectral bands needed: a six-band subset of the 21 available light-camera combinations already reaches near-minimal error."],"supporting_citations":[{"why":"Supplies the self-calibrating multi-view structured light acquisition procedure that Pro-Cam SSfM extends to include projector features.","marker":"[30]"},{"why":"Provides the projector-based spectral reflectance recovery baseline and the basis-function spectral model used in the rendering equation.","marker":"[18]"},{"why":"Provides the Munsell spectral reflectance dataset from which the eight basis functions are computed by PCA.","marker":"[41]"},{"why":"Supplies the structure-from-motion and bundle adjustment pipeline used to estimate camera and projector poses and 3D points.","marker":"[43]"},{"why":"Introduces the spectral smoothness term (second-order derivative) and the multiplexed illumination model used in the cost optimization.","marker":"[40]"},{"why":"Shows how camera sensitivity and illumination spectra can be preliminarily estimated, supporting the known-CT/L assumption.","marker":"[48]"},{"why":"Serves as the single-view spectral-recovery baseline that fails under shading, highlighting the contribution of the geometric shading model.","marker":"[5]"},{"why":"Used to compute surface normals from the reconstructed point cloud, needed for the shading factor.","marker":"[20]"}],"fun_headline_variants":["Projector-camera pair yields 3D shape and spectral reflectance","Self-calibrating projector and camera recover reflectance per point","Off-the-shelf projector-camera system for 3D and spectral reflectance","Projector geometry strips shadows to recover true reflectance"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The system assumes every surface reflects light diffusely (Lambertian), ambient light is negligible, and light does not interreflect between surface points; any deviation from these is absorbed into the estimated spectral reflectance.","fun_headline_variants_meta":{"raw":{"variants":["Projector-camera pair yields 3D shape and spectral reflectance","Self-calibrating projector and camera recover reflectance per point","Off-the-shelf projector-camera system for 3D and spectral reflectance","Projector geometry strips shadows to recover true reflectance"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000651,"raw_usage":{"total_tokens":2942,"prompt_tokens":855,"completion_tokens":2087,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":471,"completion_tokens_details":{"reasoning_tokens":2016}},"tokens_in":471,"tokens_out":2087,"duration_ms":14216,"temperature":1.0,"reasoning_tokens":2016,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:46:32.783889+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Capture a matte, concave object (such as the inside of a cup) with known spectral reflectance using this system and compare the estimated reflectance in the concave region against a spectrometer reading. A systematic bias that grows toward the concavity, where interreflections are strongest, while convex regions match, would falsify the no-interreflection assumption; a glossy object whose estimated reflectance deviates would falsify the Lambertian assumption.","supporting_citations":[{"cited_title":"Robust, precise, and calibration-free shape acquisition with an off- the-shelf camera and projector","cited_arxiv_id":null,"evidence_quote":"Supplies the self-calibrating multi-view structured light acquisition procedure that Pro-Cam SSfM extends to include projector features."},{"cited_title":"Fast spectral reﬂectance recovery using DLP projector","cited_arxiv_id":null,"evidence_quote":"Provides the projector-based spectral reflectance recovery baseline and the basis-function spectral model used in the rendering equation."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Munsell spectral reflectance dataset from which the eight basis functions are computed by PCA."},{"cited_title":"Schonberger and Jan-Michael Frahm","cited_arxiv_id":null,"evidence_quote":"Supplies the structure-from-motion and bundle adjustment pipeline used to estimate camera and projector poses and 3D points."},{"cited_title":"Grossberg, and Shree K","cited_arxiv_id":null,"evidence_quote":"Introduces the spectral smoothness term (second-order derivative) and the multiplexed illumination model used in the cost optimization."},{"cited_title":"Brown, Marc Pollefeys, and Seon Joo Kim","cited_arxiv_id":null,"evidence_quote":"Shows how camera sensitivity and illumination spectra can be preliminarily estimated, supporting the known-CT/L assumption."},{"cited_title":"Sparse recovery of hyper- spectral signal from natural RGB images","cited_arxiv_id":null,"evidence_quote":"Serves as the single-view spectral-recovery baseline that fails under shading, highlighting the contribution of the geometric shading model."},{"cited_title":"Surface reconstruction from unor- ganized points","cited_arxiv_id":null,"evidence_quote":"Used to compute surface normals from the reconstructed point cloud, needed for the shading factor."}],"review_version":1}