{"id":"2b608c27-7d39-4520-a0dd-90d0fc5b514b","arxiv_id":"2607.25371","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A transformer-based method jointly recovers reflectance, shading, specularity, and illuminant from a single non-Lambertian hyperspectral image, with a new annotated dataset.","lead":"The authors introduce a deep-learning method that separates a hyperspectral image into material reflectance, shading, specular highlights, and light spectrum without extra sensors. They also release a new real-world dataset for non-Lambertian objects and a generator for controlled synthetic training data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Absolute-MSE reflectance metric ignores a genuine DRM scale ambiguity: S and g can be jointly rescaled without changing the image, so the reported leaderboard is gauge-dependent, contradicting Sec. VI-B's claim that reflectance admits no scale ambiguity under the joint constraint.","rationale":"The reader identified the synthetic-data distribution as the weakest assumption, and also noted the absolute-MSE reflectance metric as a secondary concern. I agree that the PISG synthetic evaluation is a limitation, but the scale ambiguity is more load-bearing because it is an internal mathematical error that directly invalidates the primary quantitative comparison, independent of whether the data is synthetic or real. The paper explicitly claims 'reflectance admits no scale ambiguity under the joint constraint' (Sec. VI-B), which is contradicted by a simple gauge transformation of the DRM. This affects the interpretation of every reflectance MSE number in Table I and the conclusion that DichroicFormer outperforms baselines by large margins. The concern is not about consensus or external assumptions; it is about the correctness of the evaluation protocol. Even on the authors' own synthetic data, absolute MSE is not a physically meaningful fidelity measure for reflectance unless the gauge is fixed by prior knowledge unavailable from a single image. The paper could be revised by adopting scale-invariant metrics for reflectance or by explicitly defining and justifying the gauge, but as written the central SOTA claim is not supported. The reader's verdict of CONDITIONAL is appropriate; the revision should address this gauge issue in addition to the synthetic-data generalization concern.","tokens_in":23768,"tokens_out":10416,"duration_ms":102802,"concrete_test":"Analytically verify the zero mode: substitute S' = S/α, g' = α g into Eq. (1) and confirm I is unchanged for arbitrary α. Then take a PISG test image and apply this transformation to DichroicFormer's predicted reflectance and shading (keeping L and k fixed); recompute the absolute MSE and show that the error changes while the image is exactly reproduced, proving the metric is gauge-dependent. Alternatively, re-run Table I using a scale-invariant reflectance metric (e.g., si-MSE) and check whether the reported 5×/10× margins persist; if they shrink, the leaderboard is an artifact of gauge choice.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section VI-B states: 'For reflectance, which admits no scale ambiguity under the joint constraint, we report absolute MSE and SSIM...' This claim is analytically false. In the DRM (Eq. 1), the transformation g(u) → α g(u), S(u,λ) → S(u,λ)/α for any α>0 leaves I = g L S + k L unchanged, with L and k untouched. The inversion formulas (Eqs. 4–5) preserve this ambiguity: given a valid (P_d, P_r) pair, (P_d, P_r/α) with g derived from P_d/P_r (scaled by α) yields another exact decomposition of the same image with the same normalized L. Thus reflectance is determined only up to a global scale even when L is known and L∞=1. Providing ground-truth illuminant to baselines (Sec. VI-A) does not remove this ambiguity because the transformation leaves L unchanged. Consequently, the reported absolute-MSE gap (0.007 vs 0.189) may largely reflect gauge alignment to the PISG's conventions rather than physical decomposition fidelity, undermining the central SOTA claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes DichroicFormer, a single-image hyperspectral intrinsic decomposition method for non-Lambertian scenes under the dichromatic reflection model (DRM). The core idea is to reformulate the recovery of four coupled components (reflectance, shading factor, specular coefficient, illuminant spectrum) as estimation of two spectral–spatial variables: the diffuse term P_d and the reflectance P_r. Illuminant L, shading g, and specular coefficient k are then derived in closed form. The architecture is a dual-scale network with a global-stage invariant-driven module using spectral-gradient ratios and a local-stage specularity-guided attention module. The paper also introduces the CITE real-world dataset with component-level annotations and the PISG synthetic generator. Quantitative evaluation on PISG-generated test pairs is claimed to show state-of-the-art performance under both scale-invariant and scale-coupled metrics, with qualitative results on CITE and KAUST real HSIs.","tokens_in":24101,"tokens_out":9440,"duration_ms":104908,"significance":"If the claims hold, the paper makes several valuable contributions: a useful algebraic reduction of DRM inversion, a new public dataset (CITE), a controllable physically motivated generator (PISG), and a carefully ablated architecture. The closed-form derivations for L, g, and k are parameter-free and transparent, and the dual-scale design with photometrically invariant descriptors is well motivated. The ablations are internally consistent and isolate contributions of the proposed modules. However, the paper's central quantitative claim is weakened by two issues: the absolute reflectance metric is gauge-dependent because of a genuine scale ambiguity in the DRM, and all quantitative results are obtained on PISG-synthesized test images rather than on real captures. These issues do not necessarily invalidate the method, but they require correction or additional evidence before the claimed state-of-the-art performance can be accepted.","major_comments":[{"comment":"The statement \"For reflectance, which admits no scale ambiguity under the joint constraint, we report absolute MSE and SSIM\" is analytically incorrect under the DRM. In Eq. (1), the transformation g(u) -> alpha g(u), S(u,lambda) -> S(u,lambda)/alpha leaves I(u,lambda) unchanged for any alpha>0, with L and k untouched. This is an exact symmetry of the image-formation model. The closed-form inversion in Eqs. (4)-(5) preserves this ambiguity rather than removing it. Consequently, absolute MSE for reflectance depends on the gauge in which each method happens to output S, and the large gap in Table I (0.007 vs 0.189) may substantially reflect gauge alignment of the learned estimator to the PISG training convention rather than physical decomposition fidelity. The authors should either adopt a canonical gauge (e.g., normalize each predicted reflectance by a fixed scalar) and apply it to all met","section":"Sec. VI-B, Eq. (1)"},{"comment":"All quantitative results in Tables I-III are produced on PISG-generated test pairs: \"the PISG generates 240 test pairs from the held-out objects\" (Sec. VI-A), using the same DRM-plus-clipping synthesis model (Eq. (16)) as the training distribution. Real CITE captures are evaluated only qualitatively (Sec. VI-D.1, VI-D.2), and the KAUST experiment is also qualitative. Thus the central SOTA claim is not yet supported by evidence on real sensor data, where noise, misalignment, polarizer imperfections, and deviations from the DRM may substantially change performance. I recommend adding a quantitative evaluation on real CITE images using the component annotations obtained from Eq. (15), or explicitly redrawing the quantitative claims to the synthetic distribution. At minimum, the paper should report how the PISG test distribution relates to the real CITE captures (e.g., how much clipping, noi","section":"Secs. V-B, VI-A, VI-C; Eq. (16)"},{"comment":"The paper replaces the analytic relation P_s = I - P_d with a learned specular prediction P_L_s from SGAM, and the two are decoupled: \"the analytic pathway ties the final specular output directly to the diffuse estimate\" and the specular branch is independently supervised and detached. This means the final outputs do not have to satisfy P_L_d + P_L_s = I, so the closed-form derivations of L, g, and k in Eqs. (4)-(5), which assume a consistent P_s, may not be anchored to the observed image. The recomposition panels in Fig. 8 are visual only. I recommend reporting a quantitative recomposition error, e.g., ||I - (P_L_d + P_L_s)||, on the test set and on real captures, and stating explicitly whether the derived photometric components are required to satisfy the DRM exactly or only approximately.","section":"Sec. IV-C; Eq. (3)"}],"minor_comments":[{"comment":"The notation \"P^G_d = I + R^G_s\" appears inconsistent with the standard relation P_d = I - P_s. If R^G_s is a negative specular residual, this should be stated; otherwise the sign convention is confusing.","section":"Sec. IV-A"},{"comment":"The expression \"R = I_perp ⊙ R_coat / I_coat\" mixes element-wise multiplication and division in an ambiguous way. Please use explicit element-wise division or bracket notation.","section":"Sec. V-A, Eq. (15)"},{"comment":"The approximation rho_1 ≈ rho_2 ≈ 1 is justified by a small spectral sampling interval. Since the paper subsamples to 36 bands over 440-700 nm, the effective interval is roughly 7-8 nm. Please state whether the SR/CSR approximation was validated at this subsampled resolution, or whether the descriptors are computed on the full 80-band data before subsampling.","section":"Sec. IV-B, Eqs. (9)-(12)"},{"comment":"The training illuminant pool is described as containing 30 spectra, while the test set uses five illuminant spectra. Please clarify whether these five are seen during training or are held-out, and how they were selected.","section":"Sec. VI-A"},{"comment":"The sentence \"DichroicFormer achieves the best performance on every evaluated component under both metric families\" is too strong given the gauge-dependence of the absolute MSE metric. Please qualify it in view of the scale ambiguity discussed above.","section":"Sec. VI-C"},{"comment":"Several formal analyses are deferred to the Supplementary Material, including the formal argument about scale relationships. Since the main text makes strong claims about scale ambiguity, the key statement should at least be summarized in the main text or the supplementary material should be provided with the submission.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is technically rich and the architectural innovations are well supported by ablations, but the main quantitative claim rests on two fixable problems: an incorrect assertion about reflectance scale ambiguity and synthetic-only quantitative evaluation. If the authors add real-capture quantitative results and adopt a defensible gauge alignment or scale-invariant metric, the paper could become acceptable. I do not see a fundamental flaw in the inversion algebra or the network design, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, the genuinely new pieces are real: a complete blind single-image non-Lambertian HID framework that recovers all four DRM components, and CITE, the first public real-world HID dataset with specular-component annotations. The acquisition protocol (parallel/cross/coated captures) is clever and the dataset alone is a real contribution. Second, the quantitative evaluation has a load-bearing soft spot: all numbers come from PISG-synthesized test pairs generated from the same measured component pool used for training, and the real-data evaluation is qualitative only. That is a distribution assumption, not a math error, but it means the reported margins over baselines should not be taken at face value until tested on independent real captures with ground truth.\n\nThe inversion reformulation (reduce to estimating P_d and P_r, derive L, g, k in closed form) is simple algebra and it is correct under the stated DRM assumptions. The SR/CSR descriptors are a reasonable extension of known spectral-gradient invariants to unknown illumination. The architecture is competent, and the ablations are internally consistent. The paper is honest about the PISG limitation in Sec. VI-A, which counts in its favor.\n\nWhere the paper is soft: Sec. VI-B claims reflectance has no scale ambiguity under the joint constraint, and that is flat wrong. In the DRM, g(u)→αg(u), S(u,λ)→S(u,λ)/α leaves I unchanged for any α>0, with L and k untouched. The inversion formulas preserve this ambiguity. So the absolute-MSE reflectance numbers are gauge-dependent, and the headline gap (0.007 vs 0.189) may partly reflect gauge alignment to the PISG's conventions rather than physical fidelity. This is a genuine analytical error in a section that is supposed to justify the evaluation protocol. It does not invalidate the network or the dataset, but it does undermine the absolute-MSE claim. The reader's stress-test note lands; the paper's own statement contradicts its own equations.\n\nAlso minor: the dataset itself is not yet public (only announced), so the claimed benchmark value is conditional on release. The baselines are given ground-truth illuminant while DichroicFormer estimates it, which is favorable to the baselines, but the margin is large enough that this is a minor point, not a fatal one.\n\nWho this is for: people working in hyperspectral intrinsic decomposition, specular separation, or spectral reflectance estimation. The CITE dataset and the PISG are worth engaging with even if the method's quantitative claims need scrutiny. I would send this to a serious referee: the framing is coherent, the dataset is a real resource, and the scale-ambiguity issue is fixable in revision. The paper is not incoherent on its own terms, so serious_thinker is yes.","headline":"A solid non-Lambertian HID framework with a genuinely useful dataset, but the quantitative leaderboard is built on synthetic data from the authors' own generator and the reflectance scale-ambiguity claim in Sec. VI-B is analytically wrong.","tokens_in":24544,"tokens_out":704,"would_cite":true,"duration_ms":9700,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Hyperspectral intrinsic decomposition in non-Lambertian scenes can be reduced to estimating two spectral–spatial variables, with illuminant, shading, and specularity derived in closed form, making complete blind decomposition from one image","keywords":["hyperspectral intrinsic decomposition","dichromatic reflection model","non-Lambertian scenes","specular reflection","reflectance recovery","illuminant estimation","DichroicFormer","intrinsic image decomposition"],"falsifier":"Acquire a set of real objects with the same cross-polarization measurement protocol used to build the paper's dataset, but under illuminants and specular configurations outside the synthetic generator's sampling pools, compute pixel-wise ground truth for reflectance, shading, and specular coefficient, and check whether DichroicFormer's roughly 5–13× MSE advantages over the baselines persist; if the margins vanish, the synthetic training distribution—not the inversion math—is doing the work.","tokens_in":23704,"feed_emoji":"🌈","tokens_out":6507,"duration_ms":59374,"temperature":0.7,"pith_summary":"The paper sets out to show that non-Lambertian hyperspectral intrinsic decomposition—separating a hyperspectral image into material reflectance plus photometric effects of shading, specular reflection, and illuminant—does not require estimating four heterogeneous components. It reformulates the dichromatic reflection model so that only two spectral–spatial variables, the diffuse term and the reflectance, need to be predicted; the illuminant spectrum, shading factor, and specular coefficient then follow in closed form. On this basis it builds DichroicFormer, a dual-scale network whose global stage uses photometrically invariant gradient-ratio descriptors to preserve intrinsic boundaries and whose local stage uses specularity-guided attention to refine strong-highlight and clipped regions. Supported by a new real-world annotated dataset and a controllable synthetic generator, the paper reports that DichroicFormer outperforms existing intrinsic-decomposition baselines on every evaluated component under both standard and scale-coupled metrics, making complete blind decomposition from a single image plausible.","feed_headline":"Four components collapse to two estimates; the rest is derived","feed_subtitle":"One hyperspectral image recovers reflectance, shading, specularity, and illuminant blind.","key_machinery":"The load-bearing identity is the two-variable inversion of the dichromatic reflection model: with P_d(u,λ)=g(u)L(λ)P_r(u,λ) and P_s=I−P_d, the four coupled components are re-expressed as two spectral–spatial targets, and closed-form averaging gives the illuminant spectrum, shading factor, and specular coefficient. Two derived photometrically invariant descriptors carry the global stage: the Spectral-Gradient Ratio (SR), which suppresses illuminant and specular terms while retaining shading-modulated reflectance differences, and the Cross Spectral-Gradient Ratio (CSR), which depends only on reflectance; they are injected as boundary priors aligned to the diffuse and reflectance subnetworks. T","core_discovery":"Under the neutral-interface dichromatic reflection model, an observed hyperspectral pixel is I(u,λ)=g(u)L(λ)S(u,λ)+k(u)L(λ). The key algebraic observation is that defining the diffuse term P_d=gL S and keeping reflectance P_r=S turns the image into I=P_d+P_s, so the specular term is just I−P_d. Estimating the two spectral–spatial targets P_d and P_r therefore determines the full factorization: the illuminant spectrum is recovered by spatially averaging P_d/P_r plus P_s with peak normalization, and the shading factor and specular coefficient are recovered by spectral averaging. This reduction removes the dimensional mismatch that forced earlier methods to estimate components separately, and t","pith_inferences":["If the inversion reduction is as domain-simplifying as it appears, the same P_d/P_r parametrization could be applied to RGB or multispectral intrinsic decomposition, since the closed-form steps are wavelength-agnostic; this is an extension the paper does not claim.","The large gap between si-MSE and rs-MSE for the baselines suggests that per-component scale alignment has been systematically hiding relative-scale errors; adopting shared-scale metrics more widely in intrinsic imaging would likely re-rank existing methods—an editorial inference, not a paper claim.","The SR/CSR descriptors assume high spectral sampling density so the illuminant is locally flat; a natural stress test is whether performance degrades gracefully on coarse-band or strongly peaked-illuminant sensors, which the paper leaves to future work.","The generator's independent control over diffuse and specular scaling could be used to map robustness curves, such as error versus specular fraction or clipping severity, giving practitioners a principled way to choose acquisition exposure settings."],"forward_implications":["A single hyperspectral image, without depth, multi-view, or controlled illumination hardware, can in principle yield a complete dichromatic decomposition: reflectance, shading, specular coefficient, and illuminant spectrum.","Because the illuminant, shading, and specular coefficient are derived rather than predicted, their relative scales are automatically consistent with the image formation model, which is why the paper introduces shared-scale rs-MSE metrics.","Decoupling the specular output from the analytic I−P_d path allows the method to handle saturated and clipped specular highlights without letting clipping corrupt the diffuse estimate.","The synthetic generator's independent control over illuminant spectra and diffuse–specular balance enables training and testing under conditions real captures cannot exhaust, while the new real-world annotations provide a reference for non-Lambertian evaluation.","The two-variable reformulation is not network-specific: any estimator of P_d and P_r inherits the closed-form recovery of the photometric components."],"fun_headline_variants":["Two estimates recover four components in hyperspectral scenes","Hyperspectral intrinsic decomposition reduced to two variables","Blind non-Lambertian HSI decomposition via two targets","One image, two variables: full intrinsic factorization","Two targets derive four components from a single HSI"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The quantitative results rest on the assumption that the paper's synthetic training-and-test renderer—a simple linear mixture of diffuse and specular terms followed by clipping—faithfully captures real imaging, since real-capture evaluation is only qualitative.","fun_headline_variants_meta":{"raw":{"variants":["Two estimates recover four components in hyperspectral scenes","Hyperspectral intrinsic decomposition reduced to two variables","Blind non-Lambertian HSI decomposition via two targets","One image, two variables: full intrinsic factorization","Two targets derive four components from a single HSI"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000517,"raw_usage":{"total_tokens":2385,"prompt_tokens":829,"completion_tokens":1556,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":1480}},"tokens_in":573,"tokens_out":1556,"duration_ms":13297,"temperature":1.0,"reasoning_tokens":1480,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T02:37:06.637226+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Acquire a set of real objects with the same cross-polarization measurement protocol used to build the paper's dataset, but under illuminants and specular configurations outside the synthetic generator's sampling pools, compute pixel-wise ground truth for reflectance, shading, and specular coefficient, and check whether DichroicFormer's roughly 5–13× MSE advantages over the baselines persist; if the margins vanish, the synthetic training distribution—not the inversion math—is doing the work.","supporting_citations":[],"review_version":1}