{"id":"494f8032-5ce7-456e-9561-8a77fafd5a36","arxiv_id":"2412.09772","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A polarized reflectance field capture pipeline estimates per-pixel normals and spatially varying BRDF parameters for objects from diffuse to glossy, with artifact removal and whole-sequence optimization.","lead":"The paper describes a capture system that photographs objects under hundreds of polarized one-light-at-a-time conditions, then separates diffuse and specular reflection and estimates per-pixel surface normals, albedo, roughness, and anisotropy. If the accuracy claims hold, it offers a practical path to digitizing glossy and diffuse real objects for relighting and for training inverse-rendering methods.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The sub-0.1mm/px normal claim is not secured: Eqs. 2/3 and 11/12 assume specular light stays linearly polarized and diffuse response equals n·ω_i, so Fresnel rotation, roughness depolarization, and nonuniform Li bias normals; no independent benchmark tests this.","rationale":"The reader's weakest assumption identifies exactly the same load-bearing point: the normal refinement treats cross-polarized OLAT observations as uniform Lambertian responses, while polarization leakage, Fresnel effects, and interreflection violate that model. I agree with that assessment and with the REJECT verdict, because the paper's own evidence is insufficient to support the central quantitative claim. The paper does have genuine strengths: a substantial hardware setup, a full OLAT polarized reflectance field, a plausible preprocessing scheme for overexposure, and qualitative ablations that show visible improvements. Those count as real engineering contributions. However, the claimed 0.1mm/px accuracy is never defined numerically, no external geometry or material ground truth is used, and the only quantitative-looking table reports solver gradient norms rather than reconstruction error. The relighting reference is generated by weighting the same captured OLAT frames that were used for fitting, so it is a consistency check, not an independent validation. My proposed synthetic test directly isolates the weakest assumption: with known normals and a known BRDF, any bias from Fresnel polarization rotation or nonuniform illumination becomes measurable. This is the single most load-bearing concern because every downstream map is refit conditioned on the normals, so an error there invalidates the entire SVBRDF output, not just one component. The verdict should remain REJECT, or at most CONDITIONAL pending such a test, but since the reader already recommends rejection, I mark it UNCHANGED.","tokens_in":18794,"tokens_out":5723,"duration_ms":71428,"concrete_test":"Render a synthetic polarized OLAT sequence with a physically based renderer (e.g., Mitsuba 3) for a known ground-truth curved object with a smooth dielectric clear coat and/or anisotropic GGK metal, modeling the polarizer and analyzer as linear polarizers in the light path. Run the paper's entire pipeline on these synthetic images without modification and compare recovered n_d and n_s against the ground-truth normals. If the angular or projected geometric error exceeds the claimed 0.1mm/px, or if error grows with surface glossiness or incidence angle, the Lambertian polarization model in Eqs. 11–12 is the cause. A secondary check is to repeat with per-light incident radiance made nonuniform; if the normals shift, the missing radiometric calibration is a separate load-bearing issue.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—surface normals within 0.1mm/px and accurate SVBRDF over a wide material range—rests on the polarization separation (Eqs. 2–3) and the Lambertian normal refinement (Eqs. 11–12). Eq. 2 is derived in the supplement by passing unpolarized light through a polarizer and then an analyzer. But in the actual capture the incident light is already linearly polarized, and Fresnel reflection from a specular surface does not preserve the incident polarization axis: the s and p reflection coefficients differ, so a dielectric at off-normal incidence reflects light with a rotated polarization axis, and rough or multilayer surfaces partially depolarize. The crossed analyzer therefore transmits part of the specular lobe, and the subtraction Is = 2I∥ − 2I⊥ no longer isolates specular. Equation 11 then assumes every cross-polarized observation is a Lambertian response Li(n·ω_i), and Eq. 12 maximizes correlation against that template; any specular leakage or uneven Li (the paper reports no radiometric calibration of the 346 LEDs) shifts the optimum normal. Since ρd, ρs, σ, and γ are all refit using these normals (Eqs. 14–17), the bias propagates through the entire output. The paper's validation—qualitative renderings and a relighting comparison whose reference is synthesized from the same OLAT images used in fitting—cannot detect this bias. Thus the 0.1mm/px claim is not established by the evidence presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a polarized reflectance-field capture and optimization pipeline for acquiring surface normals and spatially varying reflectance (diffuse/specular albedo, specular variance, anisotropy, roughness) of real, non-planar objects. The capture uses a Light Stage with 346 linearly polarized one-light-at-a-time (OLAT) directions and eight synchronized cameras; diffuse/specular separation is performed from cross- and parallel-polarized observations. The pipeline removes overexposure, inter-reflection, and lens-flare artifacts, synthesizes spherical gradient illuminations from the OLAT data, optimizes diffuse and specular normal maps via cross-correlation with a Lambertian template, and refits albedo, roughness, and anisotropy from those normals. Validation is mainly qualitative: ablations, rendered comparisons, and a relighting test whose reference is itself synthesized from the captured OLAT images. The abstract claims surface normals within 0.1 mm/px, but no ground-truth or error-metric experiment supports that number.","tokens_in":19138,"tokens_out":8263,"duration_ms":92376,"significance":"If the claims were substantiated, the contribution would be significant: a full-SVBRDF-plus-normal capture solution for concave, glossy, multi-material objects, together with a planned release of polarized reflectance-field data, would be a valuable resource for inverse rendering and material acquisition. The pipeline is coherently organized: dense OLAT capture, analytic artifact suppression, and coupled optimization are clearly specified, and the supplementary material provides derivations for the albedo initialization. However, the manuscript currently does not provide evidence for the headline precision claim, and the polarization model as derived omits the surface's effect on the polarization state of reflected light. The value of the contribution therefore remains conditional on additional, independent validation and on either correcting or quantitatively justifying the polarization assumptions.","major_comments":[{"comment":"The central claim of \"highly accurate surface normals (within 0.1mm/px)\" is never tested. No ground-truth geometry, synthetic render with known normals, precision-machined phantom, or independent sensor is used; Figs. 7, 16, and 17 are qualitative ablations, Fig. 10 is a qualitative comparison against [35] with no error metric, and Table 3 reports only solver gradient norms, not measurement error against any reference. The pixel-to-millimeter mapping needed to interpret 0.1 mm/px is also not defined. A dedicated validation with known geometry or synthetic scenes, plus a per-pixel error metric (angular error, RMS normal error, or equivalent), is required to support the abstract's quantitative claim.","section":"Abstract; §5"},{"comment":"The relighting validation is self-referential. The reference image under HDRI lighting is synthesized by weighting the same captured OLAT images that are used to fit the material parameters (Eqs. 12-17). Matching this reference can only attest to internal consistency of the fitting procedure, not to physical accuracy of the measured normals or reflectance. A valid test would compare against real photographs under independent illumination, use a held-out subset of OLAT directions for prediction, or render a synthetic object with known SVBRDF and compare against ground truth.","section":"§5 (Relighting, Fig. 9)"},{"comment":"The separation Id = 2I⊥ and Is = 2I∥ − 2I⊥ is derived in the supplement with a Mueller-calculus proof that places no reflective surface between the polarizer and the analyzer. In the actual capture, the incident light is already linearly polarized, and Fresnel reflection at non-normal incidence changes the polarization state because the s and p reflection coefficients differ; rough or multilayer surfaces additionally depolarize. The crossed analyzer therefore transmits part of the specular lobe, and the subtraction in Eq. (3) does not generally isolate the specular component. The paper should either derive the correct Mueller-matrix reflectance model including Fresnel rotation and depolarization, or provide a quantitative calibration (for example, a dielectric sphere at Brewster angle) showing that the leakage is negligible over the claimed material range.","section":"§4.1, Eqs. (2)-(3); Supplement A"},{"comment":"The diffuse normal refinement assumes every cross-polarized OLAT observation is a Lambertian response proportional to n·ωi with uniform incident radiance Li across all 346 lights. No radiometric calibration of the LEDs is reported, and the constraints n·ωi ≥ 0 remove interreflection and lens-flare candidates but do not remove specular leakage or nonuniform-incidence bias. Because the optimized normal enters the subsequent albedo, roughness, and anisotropy fits (Eqs. 14-17), any bias propagates through the entire SVBRDF output. A synthetic sensitivity study - render surfaces with known normals and BRDFs, then deliberately add polarization leakage and nonuniform Li and rerun the optimization - is needed to bound the resulting normal error.","section":"§4.3, Eqs. (11)-(12)"}],"minor_comments":[{"comment":"The algorithm initializes δ as the mean of the entire OLAT sequence, while the surrounding text describes δ as an ambient value with δ ≪ ε; the mean of a signal containing specular peaks is not generally an ambient value, and the threshold ε is never given a concrete value or calibration procedure.","section":"Algorithm 1"},{"comment":"The constraint is written as n × ωi ≥ 0; throughout the main paper the intended constraint is the dot product n · ωi ≥ 0 (Eqs. 12-17). The cross-product expression is dimensionally inconsistent with a scalar inequality and should be corrected.","section":"Supplement B"},{"comment":"In the caption, the specular occlusion map is labeled τd; given the notation in Table 1 and the surrounding text, this should presumably be τs.","section":"Fig. 4"},{"comment":"The related-work section states that the capture is based on linearly polarized spherical gradient illumination [35], while the method actually captures OLAT and synthesizes gradient patterns offline (Eqs. 8-9). This distinction should be stated clearly in the method overview to avoid confusion about what is physically captured.","section":"§2.3 and §4.1"},{"comment":"The unit \"0.1mm/px\" needs a precise definition: is it a per-pixel normal perturbation expressed as surface displacement at the object's depth, and for what camera resolution and object distance? Without this definition the quantitative claim is not interpretable.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":"The paper describes a potentially useful capture pipeline, and the planned data release is commendable, but the current manuscript does not establish its central quantitative claim. The missing ground-truth validation and the polarization-model issue are substantial: if the authors cannot supply independent validation experiments and either fix or quantitatively justify the polarization separation, the manuscript should not be accepted. I chose major_revision rather than reject because the pipeline concept is sound in outline and the identified defects are, in principle, addressable with additional experiments and a corrected physical model."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a real system contribution: a multi-view polarized OLAT rig (346 lights, 8 cameras) with per-pixel optimization over the full sequence to estimate SVBRDF and surface normals. The artifact-removal heuristics (overexposure, interreflection, lens flare) and the whole-sequence optimization go beyond prior polarized gradient illumination work, and the qualitative renderings are convincing. The promised data release is a real plus. But the headline claim—'highly accurate surface normals (within 0.1mm/px)'—is not established. There is no ground-truth comparison, no numeric comparison to prior methods, and the relighting reference is synthesized from the same OLAT images used to fit the material parameters. That makes the validation partly self-referential. The solver comparison in the supplement measures convergence error, not accuracy, so it does not fill this gap.\n\nThe physical model also has soft spots. Equations 2/3 assume specular reflection preserves polarization and diffuse is unpolarized, which breaks down for rough or multilayer surfaces; Eq. 11/12 assume Lambertian response with uniform incident radiance, but no radiometric calibration of the 346 LEDs is reported. These are common assumptions in this line of work and may be acceptable for a narrower material range, but the paper claims everything from diffuse to highly glossy, and the bias would propagate through albedo, roughness, and anisotropy. The stress-test note is right about the mechanism; I just would not call it fatal on its own—the damning problem is the absence of independent validation.\n\nThe paper is honest about some limitations—no geometry measurement, and interreflections remain hard—but that honesty actually undercuts the abstract's broad claim.\n\nMy bottom line: this is a novel, well-engineered capture system that deserves a serious referee and probably heavy revision. It needs a numerical benchmark against external geometry (e.g., laser scan or structured light), an independent relighting reference, released data and code, and calibration details including the values of ε, M, and κ. If the authors add that, the paper will be a useful reference for practical polarized OLAT capture.\n\nRecommendation: accept for peer review, with major revision expected.","headline":"A serious and novel polarized OLAT capture system, but the headline normal-accuracy claim is not supported by any independent measurement; deserves peer review with major revision expected.","tokens_in":19665,"tokens_out":4303,"would_cite":false,"duration_ms":46218,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single polarized reflectance-field capture yields 0.1mm/px normals and full SVBRDF maps for diffuse-to-glossy objects.","keywords":["polarized reflectance field","SVBRDF acquisition","surface normal estimation","spherical gradient illumination","diffuse-specular separation","Ward BRDF","light stage capture","OLAT"],"falsifier":"Capture a deeply concave, glossy object (for example, a polished metal bowl) whose true surface normals are known from a laser scanner, run the pipeline, and compare measured normals inside the concavity: if the claimed accuracy holds, cross-polarized intensities must follow the Lambertian cosine law even where interreflections are strong, so any normal error that grows with concavity or glossiness would show the load-bearing assumption is violated.","tokens_in":18539,"feed_emoji":"📸","tokens_out":7187,"duration_ms":66724,"temperature":0.7,"pith_summary":"The paper proposes a complete capture pipeline in which a light stage records the object under 346 one-light-at-a-time directions through cross- and parallel-polarized light, producing a polarized reflectance field. Its central claim is that statistical analysis of this sequence removes overexposure, interreflection, and lens flare artifacts, and that an optimization over the whole image collection refines per-pixel surface normals to within 0.1mm/px while fitting albedo, specular, roughness, and anisotropy. This matters because current polarization-based methods either separate diffuse and specular reflectance well or handle complex non-convex objects, but not both, and glossy objects with concave geometry remain a failure mode for studio capture. A working version of this pipeline would give realistic-rendering and inverse-rendering research a practical source of measured materials and a dataset of polarized reflectance fields for training.","feed_headline":"Polarized light capture yields normals and reflectance in one pass","feed_subtitle":"A 346-direction polarized scan separates diffuse and specular maps for materials from matte to mirror-like.","key_machinery":"The load-bearing identity is the polarization separation $I_d=2I_\\perp$ and $I_s=2I_\\parallel-2I_\\perp$, which converts one hardware pass into two reflectance sequences. The optimization that carries the accuracy claim is normalized cross-correlation between the observed radiance vector and the predicted Lambertian irradiance vector $\\nu_k\\, \\mathbf n\\cdot \\boldsymbol\\omega_i^k$ for the diffuse normal, and the corresponding reflection-direction vector for the specular normal, both subject to $\\mathbf n\\cdot \\boldsymbol\\omega_i^k \\ge 0$ to suppress interreflection and lens flare. The Ward BRDF, with a 2D Gaussian lobe in the half-vector frame, is the model that turns the refined normal into roughness and anisotropy estimates.","core_discovery":"At the core is an argument that a polarized OLAT sequence can be treated as a per-pixel signal, turning material capture into a sequence of statistical estimation steps. Cross-polarized frames isolate the diffuse reflection, and the difference between parallel- and cross-polarized frames isolates the specular reflection, with the separation factors derived through Mueller calculus and Malus's law. Overexposure is removed by treating each pixel's lighting-response curve as a signal and clipping anomalous pulses; interreflection and lens flare are suppressed through the constraint that the active light direction and the surface normal have nonnegative dot product; and occlusion is estimated from a visibility-weighted integral. Initial normals are synthesized from spherical gradient patterns, then refined by maximizing normalized cross-correlation between the observed intensity vector and the predicted Lambertian irradiance vector, or the specular reflection-direction vector for the specular normal. The same refined normals feed least-squares fits of the Ward model to obtain roughness and anisotropy, with albedo refit afterward. The paper's claim is that this produces physically consistent maps for objects that previously defeated polarization-based capture, from matte to mirrorlike.","pith_inferences":["An untested extension is to apply the same per-pixel correlation refinement to non-polarized OLAT data, as long as specular contamination is removed by another means; the cost function itself does not require polarization hardware.","The paper's quantitative claim is a normal accuracy figure, but most validation is visual; a systematic comparison against an independent geometric ground truth on concave and multilayer objects would be the natural next step and would bound how the Lambertian assumption degrades with interreflection strength.","Because every albedo, roughness, and anisotropy map is refit using the optimized normals, the method's failure mode should appear first at grazing angles on brushed or anisotropic surfaces, where Fresnel effects break the cross-polarized Lambertian model; a targeted test on brushed metal at grazing views would reveal this.","The authors' own note that they do not currently measure geometry suggests the practical ceiling of the method on severely concave objects; integrating multiview geometry could turn the hard interreflection constraint into a learned or model-based correction."],"forward_implications":["Objects with clear coats, metals, and diffuse bases can be captured in a single studio session, producing separate diffuse and specular albedo and normal maps rather than a mixed estimate.","Physically based renderings built from the measured maps reproduce reference photographs under HDRI and area lighting, including anisotropic highlights that stay linear on cylindrical surfaces.","The released polarized reflectance fields of captured objects can serve as supervised training data for inverse-rendering and relighting methods.","The preprocessing and optimization steps remove overexposure, interreflection, and lens flare artifacts that static gradient-illumination capture leaves in albedo and normals.","The measured maps are view-consistent across the eight capture cameras, so the method supports multiview acquisition of shape with complex reflectance."],"supporting_citations":[{"why":"Supplies the polarized spherical-gradient illumination formulation that initializes the normals and serves as the main qualitative baseline.","marker":"[35]"},{"why":"Provides the anisotropic Ward BRDF model used to fit roughness and anisotropy from the specular lobe.","marker":"[44]"},{"why":"Derives specular roughness and anisotropy from second-order spherical gradient illumination, the approach the paper adapts to its OLAT sequence.","marker":"[19]"},{"why":"Establishes multiview capture with polarized spherical gradient illumination, the direct lineage of the acquisition setup.","marker":"[21]"},{"why":"Shows circularly polarized illumination for view-dependent specular capture, the contrast that motivates the paper's linear-polarization separation.","marker":"[20]"},{"why":"Supplies the Stokes/Mueller formalism used in the appendix to prove the diffuse-specular separation factors.","marker":"[9]"},{"why":"Is the continuous spherical-harmonic method that fails on non-convex objects, defining the gap the paper targets.","marker":"[43]"}],"fun_headline_variants":["Polarized scans separate diffuse and specular in one pass","One polarized capture yields normals, albedo, roughness, anisotropy","Polarized reflectance fields capture normals and BRDF together","Polarized light untangles reflection to capture material and shape","Polarized OLAT gives per-pixel normals and reflectance maps"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that under cross-polarized light every surface point reflects like a uniform Lambertian surface, so the observed brightness across lighting directions is proportional to the cosine of the angle between each light direction and the surface normal; if polarization leaks, Fresnel effects, or interreflections break that proportionality, the optimized normals are biased and the bias propagates into every later map.","fun_headline_variants_meta":{"raw":{"variants":["Polarized scans separate diffuse and specular in one pass","One polarized capture yields normals, albedo, roughness, anisotropy","Polarized reflectance fields capture normals and BRDF together","Polarized light untangles reflection to capture material and shape","Polarized OLAT gives per-pixel normals and reflectance maps"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000223,"raw_usage":{"total_tokens":1482,"prompt_tokens":992,"completion_tokens":490,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":608,"completion_tokens_details":{"reasoning_tokens":401}},"tokens_in":608,"tokens_out":490,"duration_ms":5586,"temperature":1.0,"reasoning_tokens":401,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:44:32.632842+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Capture a deeply concave, glossy object (for example, a polished metal bowl) whose true surface normals are known from a laser scanner, run the pipeline, and compare measured normals inside the concavity: if the claimed accuracy holds, cross-polarized intensities must follow the Lambertian cosine law even where interreflections are strong, so any normal error that grows with concavity or glossiness would show the load-bearing assumption is violated.","supporting_citations":[{"cited_title":"Rapid acqui- sition of specular and diffuse normal maps from polarized spherical gradient illumination","cited_arxiv_id":null,"evidence_quote":"Supplies the polarized spherical-gradient illumination formulation that initializes the normals and serves as the main qualitative baseline."},{"cited_title":"Measuring and modeling anisotropic re- flection","cited_arxiv_id":null,"evidence_quote":"Provides the anisotropic Ward BRDF model used to fit roughness and anisotropy from the specular lobe."},{"cited_title":"Estimating specular roughness and anisotropy from second order spherical gradient illumina- tion","cited_arxiv_id":null,"evidence_quote":"Derives specular roughness and anisotropy from second-order spherical gradient illumination, the approach the paper adapts to its OLAT sequence."},{"cited_title":"Multiview face cap- ture using polarized spherical gradient illumination","cited_arxiv_id":null,"evidence_quote":"Establishes multiview capture with polarized spherical gradient illumination, the direct lineage of the acquisition setup."},{"cited_title":"Circularly polarized spherical illumi- nation reflectometry","cited_arxiv_id":null,"evidence_quote":"Shows circularly polarized illumination for view-dependent specular capture, the contrast that motivates the paper's linear-polarization separation."},{"cited_title":"Field guide to polarization","cited_arxiv_id":null,"evidence_quote":"Supplies the Stokes/Mueller formalism used in the appendix to prove the diffuse-specular separation factors."},{"cited_title":"Acquiring reflectance and shape from continuous spheri- cal harmonic illumination","cited_arxiv_id":null,"evidence_quote":"Is the continuous spherical-harmonic method that fails on non-convex objects, defining the gap the paper targets."}],"review_version":1}