{"id":"2c235470-546f-402c-88df-86c09582057a","arxiv_id":"2507.07333","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A virtual try-on framework approximates Kubelka-Munk optical blending with a Taylor expansion, enabling realistic foundation previews using only e-commerce product data.","lead":"The authors present a phone-friendly way to simulate how foundation makeup will look on a given face, using only product images and coverage labels from online stores. The system applies a physics-based color blending model and a fast approximation to it, and it reports more realistic results than standard alpha blending.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The realism claim rests on an unvalidated hand-set mapping from coverage labels to scattering coefficients (30/20/10% of R∞ at D=0.25); if that mapping is wrong, both the KM approximation and the alpha-blending comparison inherit the error.","rationale":"The reader's strongest claim and weakest assumption align. The Taylor approximation itself is internally validated by Table 3 (SSIM 0.99, Delta E 0.22), so I do not challenge Eq. (12)-(15) as a computational shortcut. However, the shortcut and vanilla integration share the same unmeasured S input; internal agreement cannot validate the physical calibration. The 10-subject LPIPS comparison is too small to substantiate 'outperforms' by itself, but the more fundamental issue is that the comparison is conditional on S. Since the missing calibration is directly testable and the rest of the pipeline is plausibly sound, the appropriate verdict remains conditional: the paper should not be rejected, but the realism claim cannot be accepted until the coverage-to-S mapping is checked against physical measurements or a sensitivity analysis.","tokens_in":10900,"tokens_out":7029,"duration_ms":83804,"concrete_test":"Measure the reflectance spectra of a small set of real foundations with full/medium/low coverage labels (or use published measured data such as Doi et al. [8]) on a black substrate at the same layer thickness D used in the paper, compute R/R∞ per wavelength for each coverage class, and compare with the assumed 30/20/10%. Then rerun the Table 4 LPIPS comparison using the measured S values instead of the hand-set ones; if the ranking against alpha blending changes or the full-coverage synthesized images become visibly more opaque, the hand-set mapping is load-bearing and must be replaced or calibrated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The end-to-end realism claim depends on Eq. (7) through the scattering coefficient S, and S is set in Sec. 4.3 by the assertion 'for fixed D=0.25, we set the reflectance R of full coverage foundation to be 30% of R∞, 20% for medium coverage and 10% for low coverage,' with no measurement, citation, or sensitivity study. S controls both the foundation's own reflectance R_m and its transmittance T_m, so any error in these fractions changes the blended color and opacity. The Table 4 comparison to alpha blending therefore evaluates a particular S calibration as much as it evaluates KM; a different calibration could erase or reverse the reported LPIPS margin. The assumption is also physically questionable: a 'full coverage' foundation is defined by high opacity, yet no argument connects 30% of R∞ to that definition, so the mapping may systematically understate opacity for the products the framework claims to support. Because the paper's stated scalability is 'solely depending on the product information available on e-commerce sites,' this unmeasured mapping is the load-bearing point: the coverage label alone does not determine S without an additional, unvalidated relationship.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an end-to-end virtual try-on framework for foundation makeup. The pipeline performs face semantic segmentation, intrinsic image decomposition into albedo, shading, and specular highlight, and then blends foundation color with skin albedo using Kubelka-Munk (KM) theory. The authors introduce a piecewise-linear reflectance estimation from sRGB values, a coverage-to-scattering-coefficient mapping, and a first-order Taylor expansion that approximates the KM blending in XYZ space to avoid per-wavelength spectral integration. They validate the approximation against full spectral integration (SSIM 0.99, Delta E 0.22, LPIPS 8.5e-5) and compare the end-to-end result against alpha blending on a small set of real before/after makeup images.","tokens_in":11106,"tokens_out":3914,"duration_ms":44917,"significance":"If the coverage-to-scattering calibration is sound, this is a practically useful contribution: the Taylor approximation is internally validated against full spectral integration, runs roughly 5x faster than vanilla KM blending, and the overall pipeline is plausible for e-commerce virtual try-on. The paper correctly identifies and attacks a real scalability barrier of KM-based makeup rendering. The main significance is limited by two load-bearing weaknesses: the hand-set 30/20/10% reflectance mapping in Sec. 4.3 is unvalidated, and the real-image comparison in Table 4 rests on only 10 images with no statistical test. The approximation result itself is credible and well supported by the reported metrics.","major_comments":[{"comment":"The coverage-to-scattering mapping is the central load-bearing assumption of the realism claim. The sentence 'for fixed D=0.25, we set the reflectance R of full coverage foundation to be 30% of R∞, 20% for medium coverage and 10% for low coverage' is an empirical assertion with no measurement, citation, or sensitivity study. Since S determines Rm and Tm through Eqs. (4)-(5), and Rm and Tm fully control the blended reflectance in Eq. (7), any error in these fractions changes both color and opacity of the synthesized result. Consequently, the Table 4 comparison to alpha blending evaluates this particular calibration as much as it evaluates KM theory; a different calibration could erase or reverse the reported LPIPS margin. The additional free parameter D=0.25 (with no unit or justification) has the same problem. Please provide either a calibration against measured foundation-on-skin data, a sensitivity analysis over the R/R∞ fractions and D, or a clearly stated limitation that the method requires per-product calibration before it can be claimed to work 'solely depending on the product information available on e-commerce sites.'","section":"Sec. 4.3, Eqs. (3)-(7)"},{"comment":"The claim that the framework 'outperforms other techniques' on real after-makeup images is supported only by mean LPIPS over 10 images (0.226 vs. 0.236 and 0.231), with no error bars, confidence intervals, per-subject results, or paired statistical test. The margin is small, and the paper itself acknowledges that lighting direction and intensity may differ between before and after images. With n=10, this evidence is insufficient to establish a robust advantage over alpha blending. Please report the per-subject LPIPS values and a paired test (e.g., Wilcoxon signed-rank), and ideally add a colorimetric metric on color-checker-normalized skin regions.","section":"Sec. 5.2, Table 4"},{"comment":"The derivation of the Taylor approximation is internally consistent, and the validation in Table 3 is strong. However, the approximation is validated only against the same KM model with the same inputs; it does not validate the realism of the KM inputs themselves. Thus the Table 3 numbers (SSIM 0.99, Delta E 0.22) confirm that the approximation preserves the KM output, but the end-to-end realism claim still inherits the unvalidated coverage mapping and D value from Sec. 4.3. This should be stated explicitly in the conclusion and abstract, or the missing calibration must be supplied.","section":"Sec. 4.4, Eq. (15)"}],"minor_comments":[{"comment":"The phrase 'reflectance and transmittance of the foundation foundation' contains a duplicated word and should read 'foundation layer.'","section":"Sec. 3.2"},{"comment":"The column header 'Ment2015' is a typo for 'Meng2015' (reference [24], Meng et al.).","section":"Table 1"},{"comment":"The text says 'we only need to integrate over the values for wavelength from 500 to 600' for the X channel, but Eq. (15) is written for the Z channel and the X-channel derivation is not shown. Please present the X-channel piecewise integration explicitly.","section":"Sec. 4.4"},{"comment":"The figure caption refers to 'black lines' and a 'green dotted line,' but the printed figure may not distinguish these clearly; please use distinct markers and annotate the three-piece approximation, including the flat 600-700 nm segment.","section":"Sec. 4.2, Fig. 6"},{"comment":"The phrase 'solely depending on the product information available on e-commerce sites' is too strong given that Sec. 4.3 introduces a hand-set coverage-to-S mapping and a fixed D. If those parameters are not calibrated, please soften the claim to 'depending on product information plus an assumed coverage-opacity relationship.'","section":"Abstract and Sec. 1"},{"comment":"The statement that the method achieves a better LPIPS score '100% of the time' is not a meaningful statistic by itself; please report the number of paired comparisons and the distribution of per-image differences.","section":"Sec. 5.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is an industrial application paper with a modest novelty core: a Taylor approximation to KM blending that is internally validated and a coverage heuristic that is not. The unvalidated 30/20/10% mapping is the main correctness risk; the authors should be pushed to either calibrate it or clearly scope the claim. The 10-image real-world evaluation is too small for the strength of the claim; a paired statistical test and per-subject results should be required. If the authors can provide a sensitivity analysis or calibration, the paper may be acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The core trick here is real. The first-order Taylor expansion of the KM blending equation in XYZ space, checked against full spectral integration (SSIM 0.99, Delta E 0.22, ~5x faster), is a legitimate contribution. It's not circular—they derive it from the KM equations and validate it against the brute-force result. The sRGB-to-reflectance estimation with a three-piece linear function is also a reasonable engineering choice, and the intrinsic decomposition pipeline is coherent.\n\nWhat's shaky is the end-to-end realism claim. The scattering coefficient S is set by a hand-picked mapping: for D=0.25, full coverage means R=30% of R∞, medium 20%, low 10%. No measurement, no citation, no sensitivity study backs this up. S controls both reflectance and transmittance of the foundation layer, so if the real product deviates from these fractions, the blended color and opacity will be off. The table comparing to alpha blending on 10 subjects (LPIPS 0.226 vs 0.236) therefore measures the calibration as much as the model. The margin is small and there are no error bars. It's the kind of result that could flip with a different mapping.\n\nThere are other, more minor issues: the evaluation is on ten selfies, no code or data released, and the reflectance estimation is a rough approximation. But those are addressable.\n\nWho gets value from this? People working on physics-based VTO, especially in beauty AR, and anyone thinking about cheap approximations to Kubelka-Munk. The Taylor expansion could be useful beyond foundation. The coverage mapping, if validated, would make the whole pipeline more credible.\n\nI'd send it to peer review. The core approximation deserves scrutiny and the paper is honest about what it does—the mapping is explicitly stated as an empirical assumption, not hidden. A good referee can push for a sensitivity study or a small measurement validation. That said, I wouldn't cite it for the realism claim as it stands, but I'd point to the approximation result.","headline":"The Taylor approximation is real and validated; the coverage-to-scattering mapping is a hand-picked guess that carries the realism claim.","tokens_in":11672,"tokens_out":1900,"would_cite":false,"duration_ms":20056,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A first-order Taylor expansion of the Kubelka-Munk blend equation, evaluated directly in XYZ color space, reproduces full spectral integration for foundation-on-skin rendering while cutting blend time from 0.925 s to 0.17 s per face.","keywords":["virtual try-on","foundation makeup","Kubelka-Munk theory","Taylor expansion","reflectance estimation","intrinsic image decomposition","color blending","e-commerce scalability"],"falsifier":"Measure the reflectance of thin layers of commercial foundations at controlled thickness on known skin-tone substrates; if the inferred scattering coefficient for a given coverage label varies across shades or brands beyond noise, the fixed 30/20/10 mapping is false. The Taylor approximation itself can be checked by sampling skin and foundation reflectances from a wide color gamut and comparing the XYZ output against full spectral integration; any color pair with $\\Delta$ E CIE2000 well above 1 would falsify the no-noticeable-loss claim.","tokens_in":1901,"feed_emoji":"💄","tokens_out":6987,"duration_ms":129805,"temperature":0.7,"pith_summary":"This paper tries to show that virtual try-on for foundation makeup can be both realistic and scalable: realistic because synthesis is driven by Kubelka-Munk light-transport theory rather than alpha blending, and scalable because the only inputs needed are the product images and coverage labels already found on e-commerce pages. The central move is to approximate the KM blend equation with a first-order Taylor expansion in XYZ space, so that the expensive per-wavelength spectral integration is replaced by a direct combination of skin and foundation color channels. If the approximation holds, foundation blending drops from 0.925 s to 0.17 s per face with near-identical output, and the end-to-end system beats alpha blending on real after-makeup photos. The paper's realism claim ultimately rests on a hand-set mapping from coverage level to opacity.","feed_headline":"Foundation try-on runs 5x faster via Taylor shortcut to Kubelka-Munk","feed_subtitle":"KM blending matches full spectral integration with SSIM 0.99 and needs only e-commerce product data.","key_machinery":"The load-bearing object is the two-layer Kubelka-Munk reflectance formula for a finite foundation layer over opaque skin, $R = R_m + T_m^2 R_s/(1 - R_m R_s)$, which sums the infinite series of internal reflections between foundation and skin. The paper expands the second term $f(R_m,R_s)=T_m^2 R_s/(1-R_m R_s)$ in a first-order Taylor series around midpoints $r_m,r_s$ chosen per spectral segment. With the two partial derivatives $f_{R_m}=t R_s^2/(1-R_m R_s)^2$ and $f_{R_s}=t/(1-R_m R_s)^2$, replacing the spectral integral by moments of the color-matching functions turns blending into a linear combination of the skin and foundation XYZ values plus a constant; the result is then reconstructed from $X,Y,Z$. A piecewise-linear reflectance model with breakpoints at 400, 500, 600, and 700 nm supplies the spectra from sRGB alone, and the coverage-derived scattering coefficient fixes the layer transmittance $T_m$.","core_discovery":"The paper's central claim is that Eq. (7), the Kubelka-Munk blend of a finite foundation layer over skin, can be replaced by a first-order Taylor expansion evaluated directly in XYZ color space with no perceptible loss: the term $f(R_m,R_s) = T_m^2 R_s/(1 - R_m R_s)$ is expanded as $f(R_m,R_s) \\approx f_{R_m}(r_m,r_s) R_m + f_{R_s}(r_m,r_s) R_s + \\text{const}$, so the per-wavelength integration over the visible spectrum can be skipped. Combined with a piecewise-linear sRGB-to-reflectance conversion and a coverage-based scattering coefficient, this makes the whole pipeline depend only on product images and coverage labels from e-commerce pages. In the paper's experiments the approximation matches full spectral integration with SSIM 0.99, $\\Delta$ E CIE2000 0.22, and LPIPS 8.5e-5, cuts KM blending time from 0.925 s to 0.17 s per face, and produces LPIPS 0.226 against real after-makeup photos versus 0.236 for $\\alpha$ blending.","pith_inferences":["A likely consequence is that the same Taylor-in-XYZ trick transfers to any two-layer KM compositing problem, such as lipstick, concealer, or sunscreen rendering, wherever the blend term is smooth in the two reflectances; the bottleneck would shift from integration to reflectance estimation.","Because the approximation matches full integration to SSIM 0.99, any remaining realism gap versus real photos probably comes from the sRGB-to-spectrum reflectance model and the coverage-to-S mapping, not from the Taylor step; improving those inputs should be the next lever.","The fixed 30/20/10 coverage fractions suggest a calibration experiment: measure real foundation layers on known substrates and fit the scattering coefficient per product; if it varies within a coverage class by shade, adding shade-dependent opacity would improve realism beyond the current single scalar.","The comparison against alpha blending at LPIPS 0.226 versus 0.236 is close enough that the practical advantage may be less about raw fidelity and more about removing per-product tuning and preserving features; a larger before/after dataset would sharpen the claim."],"forward_implications":["Foundation VTO becomes practical on phones: per-face KM blending drops from 0.925 s to 0.17 s on a laptop-class CPU, with preprocessing also reduced from 11.17 s to 6.85 s.","New products can be added without lab measurements: any foundation with an sRGB product image and a coverage label can be blended using the pipeline.","The Taylor approximation is visually indistinguishable from full KM spectral integration in the reported metrics (SSIM 0.99, Delta E CIE2000 0.22, LPIPS 8.5e-5), so the faster path does not sacrifice color realism.","On real after-makeup photos, the method scores LPIPS 0.226, beating plain alpha blending (0.236) and alpha blending on decomposed albedo (0.231), and it avoids the per-product alpha hand-tuning.","The intrinsic decomposition plus semantic mask lets the system preserve facial features such as eyebrows and eyelashes, which makeup-transfer models tend to alter."],"supporting_citations":[{"why":"Establishes the Kubelka-Munk reflectance of an opaque layer, used to derive the absorption-to-scattering ratio from foundation product color.","marker":"[18]"},{"why":"Gives the two-layer blended reflectance formula in Eq. (7) that the Taylor approximation is designed to replace.","marker":"[17]"},{"why":"Provides the finite-thickness reflectance and transmittance equations from which scattering coefficient and layer opacity are derived.","marker":"[16]"},{"why":"Supplies measured foundation reflectance spectra used to justify the piecewise-linear reflectance approximation.","marker":"[8]"},{"why":"Defines the CIE XYZ color matching functions used to convert reflectance spectra into perceived color and to build the XYZ-space Taylor shortcut.","marker":"[10]"},{"why":"Provides the reference skin reflectance dataset used to evaluate the skin reflectance estimation method.","marker":"[5]"},{"why":"Supplies the intrinsic image decomposition formulation separating albedo, shading, and specular highlight before blending.","marker":"[20]"},{"why":"Defines alpha blending, the standard baseline that the KM-based method is compared against and reported to beat.","marker":"[33]"}],"fun_headline_variants":["5x faster foundation try-on via Taylor KM shortcut","Taylor expansion speeds Kubelka-Munk try-on 5x","Virtual foundation try-on matches full spectrum with SSIM 0.99","Scalable foundation VTO: e-commerce data only, 5x faster","Kubelka-Munk Taylor shortcut cuts blend time 5x"],"cache_read_input_tokens":13824,"weakest_assumption_plain":"The load-bearing premise is that a foundation's coverage label fixes its opacity: full, medium, and low coverage are assumed to reflect 30%, 20%, and 10% of the fully opaque reflectance at a fixed thickness; real products might not follow that mapping.","fun_headline_variants_meta":{"raw":{"variants":["5x faster foundation try-on via Taylor KM shortcut","Taylor expansion speeds Kubelka-Munk try-on 5x","Virtual foundation try-on matches full spectrum with SSIM 0.99","Scalable foundation VTO: e-commerce data only, 5x faster","Kubelka-Munk Taylor shortcut cuts blend time 5x"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000467,"raw_usage":{"total_tokens":2321,"prompt_tokens":928,"completion_tokens":1393,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":544,"completion_tokens_details":{"reasoning_tokens":1301}},"tokens_in":544,"tokens_out":1393,"duration_ms":12613,"temperature":1.0,"reasoning_tokens":1301,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:43:32.624220+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the reflectance of thin layers of commercial foundations at controlled thickness on known skin-tone substrates; if the inferred scattering coefficient for a given coverage label varies across shades or brands beyond noise, the fixed 30/20/10 mapping is false. The Taylor approximation itself can be checked by sampling skin and foundation reflectances from a wide color gamut and comparing the XYZ output against full spectral integration; any color pair with $\\Delta$ E CIE2000 well above 1 would falsify the no-noticeable-loss claim.","supporting_citations":[{"cited_title":"An article on optics of paint layers.Z","cited_arxiv_id":null,"evidence_quote":"Establishes the Kubelka-Munk reflectance of an opaque layer, used to derive the absorption-to-scattering ratio from foundation product color."},{"cited_title":"New contributions to the optics of intensely light-scattering materials","cited_arxiv_id":null,"evidence_quote":"Gives the two-layer blended reflectance formula in Eq. (7) that the Taylor approximation is designed to replace."},{"cited_title":"Springer Berlin, Heidelberg, 1969","cited_arxiv_id":null,"evidence_quote":"Provides the finite-thickness reflectance and transmittance equations from which scattering coefficient and layer opacity are derived."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies measured foundation reflectance spectra used to justify the piecewise-linear reflectance approximation."},{"cited_title":"Harris and I.L","cited_arxiv_id":null,"evidence_quote":"Defines the CIE XYZ color matching functions used to convert reflectance spectra into perceived color and to build the XYZ-space Taylor shortcut."},{"cited_title":"Refer- ence data set of human skin reflectance, 2017","cited_arxiv_id":null,"evidence_quote":"Provides the reference skin reflectance dataset used to evaluate the skin reflectance estimation method."},{"cited_title":"Simulating makeup through physics-based manipulation of intrinsic image layers","cited_arxiv_id":null,"evidence_quote":"Supplies the intrinsic image decomposition formulation separating albedo, shading, and specular highlight before blending."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines alpha blending, the standard baseline that the KM-based method is compared against and reported to beat."}],"review_version":1}