{"id":"8af72316-117c-4036-84a1-7c4a2f8db247","arxiv_id":"2508.14411","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"The paper introduces and evaluates the first real-world dataset for inverse rendering with a display-camera system, including a baseline method that recovers normals more accurately than prior methods.","lead":"Researchers built a display-camera setup and captured the first public real-world dataset for display-based inverse rendering, with 16 objects, stereo polarization images, and ground-truth geometry. The dataset lets others test and improve inverse rendering methods that use LCD screens as programmable light sources.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Point-superpixel model error is asserted, not measured; Table 3 'outperforming' claim is internally contradicted by SRSH's higher PSNR/SSIM.","rationale":"The reader correctly targeted the superpixel point-light assumption as the weakest assumption, and I agree that the qualitative sphere experiment does not prove its validity for all materials. That is the paper's main correctness risk: the dataset's claim to support 'precise' evaluation depends on the image formation model being accurate for the actual objects. However, the reader's framing underplays how directly the paper's own Table 3 undermines the abstract's 'outperforming' claim—this is not a nuance but a contradiction. The dataset itself remains novel and potentially useful, and the paper provides calibration details, ground-truth scans, and a reproducible baseline; those are real strengths. But the finite-area systematic error is quantified nowhere, and the headline comparison is not consistently true across all reported metrics. Given the supplemental does include a backlight-invariance test and a point-light sphere test, the authors are aware of the concern and could address it with a targeted computation; a CONDITIONAL verdict with the finite-area re-rendering test is the honest adjustment. I do not see a reason to reject the dataset claim, but the 'outperforming' statement and the point-light error need explicit resolution before the paper can be accepted as-is.","tokens_in":20350,"tokens_out":1437,"duration_ms":13624,"concrete_test":"Re-run the baseline and SRSH on all 16 objects using the released code, replacing Eq. 2's point-light term with an integral over each 240×240 superpixel's finite area (using the calibrated display geometry), and compare normal MAE and relighting PSNR/SSIM. If the finite-area rendering changes the baseline's metrics by more than ~1 dB PSNR or ~2° MAE, the point-light approximation is a dominant source of error; if SRSH still exceeds the baseline on relighting, the 'outperforming' claim must be qualified.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central artifact—the dataset—is built on an image formation model (Eq. 2) that treats each 240×240-pixel superpixel as a point light with 1/d^2 falloff. The supplement's only evidence for this 'minimal impact' is Figure S7, a qualitative glossy-sphere image comparison with no error metric, no object at the actual capture distance, and no specular-roughness range. Section 6's own Table 3 shows SRSH [37] achieves higher relighting PSNR (41.28 vs 39.33) and SSIM (0.9895 vs 0.9821); the abstract's 'outperforming state-of-the-art inverse rendering methods' is supportable only for normal MAE (20.94 vs 25.25), not for relighting. The point-light approximation could systematically bias both the dataset's near-field calibration and the baseline's reconstruction of highly specular objects; Table 6's 480×480-superpixel experiment demonstrates the approximation's failure mode (area-light behavior breaks conventional methods), but no quantitative error bound is given for the actual 240×240 configuration on the captured objects.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a real-world dataset for display-camera inverse rendering, built from a calibrated LCD display and stereo polarization cameras. It captures 16 objects with varied geometry and reflectance under one-light-at-a-time (OLAT) superpixel patterns, provides ground-truth geometry from structured-light scanning, and supports synthesis of arbitrary display patterns and noise by linear superposition. The paper also evaluates several photometric stereo and inverse rendering methods and proposes a simple baseline based on photometric stereo initialization, stereo depth, and a basis-BRDF optimization with a point-light near-field image formation model.","tokens_in":20643,"tokens_out":4283,"duration_ms":50875,"significance":"If the claims hold, the dataset is a valuable public benchmark for a practically attractive but under-served imaging configuration: programmable LCD illumination with polarization-based diffuse/specular separation. The paper's strengths include a detailed radiometric and geometric calibration procedure, stereo polarization captures, ground-truth scanned geometry, a linear-synthesis capability with controllable noise, and a broad evaluation across calibrated and uncalibrated methods. The point-light superpixel model and the claimed superiority of the baseline are two areas where the current evidence is not yet sufficient, and both affect how the dataset and baseline should be used.","major_comments":[{"comment":"The claim that the baseline is 'outperforming state-of-the-art inverse rendering methods' is not supported by Table 3 as written. There, SRSH [37] achieves higher relighting PSNR (41.28 vs 39.33) and SSIM (0.9895 vs 0.9821) than the proposed baseline; the baseline wins only in normal MAE (20.94 vs 25.25). The abstract, introduction, and Section 6 discussion should either restrict the superiority claim to normal accuracy or provide a broader metric-by-metric discussion instead of the current unqualified statement.","section":"Abstract; Section 6, Table 3"},{"comment":"The image formation model treats each 240x240-pixel superpixel as a point light with 1/d^2 falloff. The supplement's only support for 'minimal impact' is Fig. S7, a qualitative glossy-sphere image set with no error metric, no stated object distance or roughness range, and no test at the actual 50 cm capture distance. Section 6, Table 6 shows the approximation's failure mode at 480x480 superpixels, but no error bound is provided for the 240x240 configuration. Please add a quantitative validation, e.g., comparing the point-light model against an area-light integral for a calibrated sphere over the distances and roughness values in the dataset, and report the resulting bias in normals, roughness, and relighting PSNR.","section":"Section 3, Eq. (2); Supplement Section 5, Fig. S7"},{"comment":"The comparison protocol for Table 3 is underspecified. The 'Patterns' row shows that the proposed baseline uses both multiplexed and OLAT inputs, while SRSH, DPIR, and IIR use OLAT only; the text says the 144 OLAT images are divided into training and testing sets with a 5:1 ratio but does not state whether SRSH and the other methods receive all 144 images or only the training subset, nor how relighting PSNR is computed on held-out patterns. This ambiguity affects the interpretation of 'outperforming'. Please specify the exact input for each method (number and type of patterns, training/test split, held-out patterns) and, if possible, add a like-for-like comparison with the same M patterns for all methods.","section":"Section 6, Table 3 and text"}],"minor_comments":[{"comment":"The main text says the LCD emits vertically polarized light, while the supplement says 'each pixel emits horizontally linearly-polarized light.' Reconcile the statement.","section":"Section 3 vs. Supplement Section 1"},{"comment":"The object name 'OBJET' appears to be a typo for 'OBJECT'; please correct for consistency.","section":"Figure 2 and Table 2"},{"comment":"The text 'Robustness without Stereo Imaging' says the uniform-depth baseline 'outperforms previous methods, with relighting PSNR 38.8 and normal MAE 28.29, as shown in Table 3,' but Table 3 does not contain these numbers. Add a dedicated table or remove the citation.","section":"Supplement Section 5"},{"comment":"The column heading 'low res. 32-inch Default(M = 32)' is difficult to parse. Define each configuration explicitly (superpixel size, display size, number of superpixels).","section":"Table 6"},{"comment":"There are several OCR-style spacing issues such as 'OLA T' instead of 'OLAT' and 'Y ujin' / 'V arious'. Please correct typographical issues.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The dataset contribution is likely to be useful to the community, and the calibration effort is solid. The main issue is that the paper's headline claim of outperforming state-of-the-art inverse rendering is contradicted by its own Table 3, and the point-light assumption that underlies both the dataset and the baseline needs quantitative validation. These are fixable with rewording and additional experiments, but they are load-bearing for the presentation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The headline: this is the first real-world display-camera inverse rendering dataset, and that is worth having. The paper's stronger claim—that the baseline 'outperforms state-of-the-art inverse rendering methods'—only holds for normal accuracy, not for relighting, and the point-light validation is much thinner than the forward model needs. Neither issue kills the dataset; both are fixable.\n\nWhat is new and good: they built a display-camera rig with stereo polarization cameras, calibrated display backlight and nonlinearity, estimated superpixel positions via mirror-based checkerboard, captured 16 objects with varied materials, and provide ground-truth geometry from structured-light scanning. They also show how to synthesize arbitrary display patterns by linear combination of OLAT images. That synthesis is useful. The evaluation of existing photometric stereo and inverse rendering methods on this setup is honest about current failures, and the proposed baseline is simple and runs in 150 seconds. The decision to release code and data is the right call. The paper does not try to hide the trade-off between SRSH's relighting quality and its own better normals; the abstract just overreaches.\n\nSoft spots, in order of importance. First, Eq. 2 models every 240x240 display superpixel as a point light with inverse-square falloff. The supplement's Figure S7 is the only evidence that this matters little, and it is a qualitative glossy-sphere comparison with no numbers, no object at the actual 50-cm capture distance, and no specular roughness range. The dataset's synthetic relighting and the baseline both inherit this approximation, so a systematic error here would propagate into any benchmark that uses the dataset. Table 6's 480x480 experiment shows the area-light transition is real for large superpixels, which makes the lack of a quantitative check at 240x240 more noticeable. This is a moderate concern, not a fatal one—it is measurable and should be measured.\n\nSecond, the 'outperforming' claim. Table 3 has SRSH at 41.28 dB PSNR and 0.9895 SSIM against the baseline's 39.33 and 0.9821 on OLAT; the baseline wins only on normal MAE. The paper's own caption says SRSH relights well, so the abstract should say the baseline improves normals, or qualify relighting as competitive. This is a wording fix.\n\nThird, ground-truth mesh-to-image alignment via mutual information is reported without a quantification of alignment error. That is a minor gap for a dataset that ships depth and normal maps as ground truth.\n\nWho should read this: anyone doing inverse rendering with displays, near-field lighting, or polarized capture. It gives the community a missing testbed and a reference point. I'd send it to review, but with a request for the forward-model validation and a more exact contribution statement.","headline":"First real display-camera inverse rendering dataset, worth engaging; baseline claim and point-light validation need fixing before acceptance.","tokens_in":21152,"tokens_out":4119,"would_cite":true,"duration_ms":43253,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces the first real-world dataset for display-based inverse rendering, with 16 objects captured under 144 one-light-at-a-time LCD patterns, stereo polarized views, and scanned ground-truth geometry.","keywords":["display inverse rendering","LCD illumination","polarization imaging","photometric stereo","BRDF estimation","near-field lighting","relighting dataset","ground-truth geometry"],"falsifier":"Capture a glossy sphere of known BRDF under the same OLAT patterns and compare the actual pixel intensities with those rendered by the point-light model for every superpixel; a residual that grows with superpixel angular size or with surface gloss would falsify the point-light assumption. Independently, synthesize a multiplexed pattern from OLAT images using Eq. 4 and physically capture that pattern: if the residual exceeds the stated noise and clipping model on several objects, the linear-synthesis claim fails.","tokens_in":20257,"feed_emoji":"💡","tokens_out":7583,"duration_ms":82822,"temperature":0.7,"pith_summary":"The central claim is that an LCD display, used as a programmable near-field light source, can support practical inverse rendering, and that no public real-world dataset existed to test this. The paper builds and calibrates a display-camera rig, captures 16 objects with diverse materials under 144 superpixel OLAT patterns, and provides structured-light ground-truth geometry. It then shows that captured OLAT images can be linearly recombined to synthesize arbitrary display patterns and noise levels, and proposes a simple differentiable-rendering baseline that reconstructs normals and basis-BRDF reflectance. On this dataset the baseline reports better relighting and normal accuracy than existing photometric stereo and inverse rendering methods, positioning the dataset as a benchmark for display-camera inverse rendering.","feed_headline":"One LCD display becomes 144 lights for inverse rendering","feed_subtitle":"16 real objects, stereo polarized views, scanned geometry: a testbed for display-based inverse rendering.","key_machinery":"The load-bearing machinery is the display-camera image formation model: captured intensity is a clipped sum over N display superpixels of BRDF times cosine falloff times superpixel intensity divided by squared distance, plus Gaussian noise (Eq. 2). Calibrated backlight and gamma (Eq. 1) make each OLAT capture a linear basis, so arbitrary patterns are synthesized by Eq. 4. For reconstruction, the key object is the basis-BRDF representation: spatially varying reflectance is a weighted sum of analytic Cook-Torrance BRDFs, which regularizes the sparse light-view angular sampling inherent to displays and is optimized together with per-pixel normals.","core_discovery":"The paper's core contribution is the first real-world dataset for display inverse rendering, together with the validation that display-camera capture works for this task. The image formation model treats each 240×240-pixel LCD superpixel as a calibrated near-field point light source, accounts for spatially varying backlight and display nonlinearity, and separates diffuse and specular components using polarization. Because transport is linear, an image under any display pattern is a weighted sum of OLAT images plus noise, so the dataset supports simulation without recapturing. A baseline that optimizes per-pixel normals and a weighted sum of Cook-Torrance basis BRDFs by differentiable renderi","pith_inferences":["The linear-synthesis property also makes the dataset a natural testbed for illumination estimation: an algorithm can be asked to recover the 144-dimensional display pattern from a single image and be scored against the known synthesis weights.","The point-light superpixel assumption sets a practical ceiling on spatial lighting resolution; extending the image formation model to finite-area emitters or deconvolving the superpixel footprint is a direct next step that the supplement's sphere experiment only partially validates.","Since the baseline uses depth only for initialization and geometric regularization, the data could support joint normal-depth-reflectance refinement studies, including stability when stereo input is degraded.","The finding that backlight is invariant to superpixel intensity suggests a simple dark-frame subtraction strategy that might transfer to other LCD displays, lowering the cost of reproducing the capture setup."],"forward_implications":["Display-camera inverse rendering finally has a public real-world benchmark with ground-truth geometry, so methods can be compared on physical captures rather than synthetic data.","Because arbitrary display patterns and noise levels can be synthesized offline from OLAT images, researchers can test new pattern designs without re-running the capture hardware.","The evaluation identifies the main bottlenecks of display inverse rendering—limited light-view angular sampling and near-field attenuation—and shows that modeling attenuation improves relighting quality.","As few as two learned multiplexed patterns support competitive photometric stereo, indicating that faster acquisition with fewer display patterns is achievable.","Polarization-separated diffuse images improve normal accuracy for some methods, suggesting further gains from exploiting LCD polarization."],"supporting_citations":[{"why":"Supplies the superpixel display parameterization, mirror-based display calibration, and learned display patterns used throughout the dataset and baseline.","marker":"[10]"},{"why":"Provides the point-based differentiable rendering approach and basis-BRDF ideas the baseline builds on, and serves as a comparison method.","marker":"[11]"},{"why":"Supplies the interpretable basis-BRDF representation used by the baseline for reflectance regularization.","marker":"[12]"},{"why":"Primary photometric stereo comparator in the evaluation and the method used to probe behavior across display configurations.","marker":"[27]"},{"why":"Comparison inverse rendering method whose normal accuracy is benchmarked against the baseline on this dataset.","marker":"[37]"},{"why":"Provides the RAFT-Stereo depth initialization used by the baseline before optimization.","marker":"[45]"},{"why":"Mitsuba3 is used to render ground-truth depth maps, normal maps, and masks from the scanned meshes.","marker":"[28]"},{"why":"Checkerboard method used for stereo camera intrinsic and extrinsic calibration.","marker":"[81]"},{"why":"Mutual-information registration aligns the structured-light scan to the captured images for ground-truth geometry.","marker":"[14]"},{"why":"Rusinkiewicz coordinates used to analyze the limited light-view angular coverage of the display-camera setup.","marker":"[58]"}],"fun_headline_variants":["First real-world dataset for display-based inverse rendering","LCD screen as 144 lights: stereo polarized inverse rendering dataset","Display-camera dataset: 16 objects, OLAT patterns, scanned geometry","Real-world inverse rendering dataset from display-camera system","Polarized display lights enable diffuse-specular separation in dataset"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The whole capture and reconstruction pipeline treats each 240×240-pixel display region as a point light source with inverse-square falloff; if the finite size and angular extent of those regions noticeably bias the lighting model, the calibrated lighting and all reconstructed normals and reflectance would be systematically off.","fun_headline_variants_meta":{"raw":{"variants":["First real-world dataset for display-based inverse rendering","LCD screen as 144 lights: stereo polarized inverse rendering dataset","Display-camera dataset: 16 objects, OLAT patterns, scanned geometry","Real-world inverse rendering dataset from display-camera system","Polarized display lights enable diffuse-specular separation in dataset"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000169,"raw_usage":{"total_tokens":1094,"prompt_tokens":731,"completion_tokens":363,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":475,"completion_tokens_details":{"reasoning_tokens":278}},"tokens_in":475,"tokens_out":363,"duration_ms":4474,"temperature":1.0,"reasoning_tokens":278,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T18:33:33.016709+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Capture a glossy sphere of known BRDF under the same OLAT patterns and compare the actual pixel intensities with those rendered by the point-light model for every superpixel; a residual that grows with superpixel angular size or with surface gloss would falsify the point-light assumption. Independently, synthesize a multiplexed pattern from OLAT images using Eq. 4 and physically capture that pattern: if the residual exceeds the stated noise and clipping model on several objects, the linear-synthesis claim fails.","supporting_citations":[{"cited_title":"Cnn-ps: Cnn-based photometric stereo for general non-convex surfaces","cited_arxiv_id":null,"evidence_quote":"Provides the point-based differentiable rendering approach and basis-BRDF ideas the baseline builds on, and serves as a comparison method."},{"cited_title":"Universal photometric stereo network using global lighting contexts","cited_arxiv_id":null,"evidence_quote":"Supplies the interpretable basis-BRDF representation used by the baseline for reflectance regularization."},{"cited_title":"Diligent-pi: Photometric stereo for planar surfaces with rich details- benchmark dataset and beyond","cited_arxiv_id":null,"evidence_quote":"Primary photometric stereo comparator in the evaluation and the method used to probe behavior across display configurations."},{"cited_title":"From shading to local shape","cited_arxiv_id":null,"evidence_quote":"Mitsuba3 is used to render ground-truth depth maps, normal maps, and masks from the scanned meshes."},{"cited_title":"Large scale multi-view stereopsis evaluation","cited_arxiv_id":null,"evidence_quote":"Mutual-information registration aligns the structured-light scan to the captured images for ground-truth geometry."}],"review_version":1}