REVIEW 3 major objections 6 minor 79 references
Monocular Facial Appearance Capture in the Wild
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A short monocular video of a head rotation is enough to recover studio-grade facial appearance maps, without assumptions on lighting.
desk verdict A genuinely new visibility-modulated split-sum shading model for monocular face capture, honest about its limits, but with quantitative claims that outrun the evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the occlusion-aware split-sum shading model. The standard split-sum approximation factorizes the rendering equation into a precomputed BRDF integral and a prefiltered environment map; the paper modifies the second factor to include a view-dependent visibility term $\tilde V(x,\omega_r)$, computed as a Monte Carlo average of the binary visibility $V(x,\omega_k)$ weighted by the BRDF normal distribution. This removes baked-in self-shadowing from the albedo without giving up the efficiency of split-sum lookups, while diffuse light transport is ray-traced to stay accurate.
What would settle it
Render a synthetic face with known ground-truth albedo, specular, roughness, and environment using a full path tracer; run the method on the rendered monocular head-rotation video with noiseless poses, and measure the albedo error in self-occluded regions such as the nose crease, under-chin, and eye sockets. If the recovered albedo still contains residual baked-in shadow beyond the reported error scale, or if adding realistic pose noise of about two degrees shifts the recovered albedo by more than the skin-tone ambiguity, the central claim of physically correct diffuse/specular separation under tracking errors fails.
Extended reading notes
Core claim
The central discovery is that the failure of existing monocular appearance capture to separate diffuse and specular reflectance stems largely from ignoring self-occlusion in the shading model, and that a tractable correction exists. The paper proposes a visibility-modulated split-sum approximation: the prefiltered environment map term is multiplied by a view-dependent visibility factor, estimated by Monte Carlo sampling of the BRDF's normal distribution. For the typically low specular roughness of skin the approximation is accurate, while diffuse shading is handled by explicit ray tracing with multiple importance sampling. Jointly optimizing geometry, albedo, specular intensity, roughness, and environment lighting with this model produces relightable appearance maps that approach studio multi-view capture quality.
Load-bearing premise
The method trusts the per-frame head poses and neck rotations estimated by the initial monocular 3DMM tracking; if those poses are wrong, the appearance decomposition is impaired substantially, and the inverse-rendering stage never corrects them.
Editorial extensions
If this is right
- A standard camera on a tripod can produce relightable face assets for VFX and games, cutting the cost and complexity of studio capture.
- Because no lighting assumption is made, the method works outdoors in sun or shadow, indoors, and under mixed illumination—capture can happen on a film set or at home.
- The recovered albedo and specular maps can be fed directly into modern skin shaders for relighting under novel environments.
- The visibility-modulated split-sum term is a general rendering approximation that can be dropped into other inverse-rendering pipelines, for example for glossy objects, as a cheap way to add self-shadowing.
- For the research community, the result shifts the in-the-wild appearance-capture bottleneck from lighting assumptions to the quality of the initial monocular tracking.
Reading between the lines
- The paper's comparison implies that any method which ignores self-occlusion will bake area shadows into albedo; an immediate testable extension is to verify on a public synthetic dataset with known ground truth whether the recovered albedo is invariant across different capture environments, since the paper itself notes skin-tone recovery is not guaranteed.
- The explicit reliance on external 3DMM tracking suggests a natural next step: end-to-end refinement of poses inside the inverse-rendering loop with temporal smoothness, which the authors tried and found jittery—a robust pose-differentiable rendering scheme could close the remaining gap.
- Because the visibility approximation becomes exact for mirror-like surfaces and degrades with roughness, the method's scope is implicitly limited to materials with low specular roughness; a quantitative roughness ceiling could be measured by testing on synthetic objects with controlled roughness.
- Applying the same capture protocol to the same subject under two different lighting conditions and checking that albedo maps agree would directly test the disentanglement claim, since the environment map absorbs lighting differences.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents an inverse-rendering method that recovers facial geometry, diffuse albedo, specular intensity, and specular roughness from a monocular video of a head rotation in an arbitrary static environment. The optimization jointly solves for mesh vertices, texture maps, and an environment map using a differentiable rasterizer and ray tracing. The main technical contribution is a visibility-modulated split-sum shading model (Eqs. 6-7) that accounts for self-occlusion in the specular term. The authors report qualitative relighting results, comparisons with FLARE, NextFace, and SunStage, and a synthetic ablation, and they claim that the recovered appearance maps approach the fidelity of studio-based multi-view capture.
Significance. If the claims hold, this is a practically important step: it could replace controlled multi-view or light-stage capture for relightable facial assets with a short monocular sequence under arbitrary static illumination. The visibility formulation directly addresses a known limitation of split-sum inverse rendering, and the qualitative results, especially the relighting comparisons, are compelling. The main strengths are the clear problem formulation, the explicit treatment of visibility, and the practical capture protocol. However, the central decomposition claim is not yet quantitatively validated: the synthetic evaluation uses a Lambertian material and does not exercise the specular model, and the quantitative metrics are computed on the optimization frames. With additional targeted validation of specular separation and generalization, the contribution would be solid; in its current form the evidence is not sufficient for the advertised claims.
major comments (3)
- [§3.2, §4.3, Fig. 6] The central claim that the shading model 'correctly separates the diffuse and specular components' is not supported by the quantitative evaluation. Equation (6) is explicitly an approximation valid for r << 1, and the authors restrict it to the specular lobe; human skin specular roughness is not necessarily in this regime, so the recovered specular intensity and roughness can be biased unless this is tested. The synthetic experiment in §4.3 uses a Lambertian material (supplement §9.2) and reports only diffuse albedo and geometry errors, so it never provides ground-truth validation of specular intensity, specular roughness, or the visibility model itself. I ask for a synthetic evaluation with a non-Lambertian BRDF with known specular parameters and environment, reporting errors in the recovered specular maps and visibility, and, if feasible, a real-world cross-check under a second illumination.
- [Table 1, §9.1] The quantitative comparison in Table 1 is computed over frames used in the optimization (supplement §9.1 explicitly states this), so the reported errors largely measure fitting ability rather than generalization or correct decomposition. A low render error on training frames is compatible with an incorrect decomposition, as the FLARE comparison in Fig. 3 itself shows: similar final renders can accompany near-zero specular. Please report reconstruction errors on held-out frames of the same sequence and, where possible, on novel views or relit images, and add a metric that directly evaluates the decomposition, such as albedo and specular consistency across held-out poses.
- [§5, supplement §10] The method's in-the-wild robustness rests on the accuracy of the external monocular tracking: head poses and neck rotations are inputs and are not refined in the inverse rendering stage. The manuscript acknowledges that pose inaccuracies 'impair the reconstruction quality of our method substantially' and the supplemental shows distorted results from a poor fit. Because this tracking pipeline is outside the method, the paper should quantify sensitivity to pose error, for example by perturbing ground-truth poses in the synthetic dataset and reporting geometry and albedo error, and state more precisely under what pose-error range the claimed fidelity holds.
minor comments (6)
- [§4.2] The phrase 'completely physically-based' overstates the case, since Eq. (6) is an approximation and Eq. (9) is a heuristic regularizer; I suggest softening this wording.
- [§3.2, Eq. (7)] The notation D(n, ωk, ωr, r) is not defined precisely; please clarify that D is the Beckmann normal distribution function and specify how the samples ωk are drawn.
- [Table 1] The table reports averages over all subjects but the number of subjects and per-subject variance are not given; adding these would make the comparison more informative.
- [Supplement §8] There is a typo: 'a learning of 0.1' should read 'a learning rate of 0.1'.
- [References] Several reference entries contain stray trailing numbers (e.g., [4], [37], [62]) and inconsistent formatting; these should be cleaned up.
- [§7.1] The capture protocol restricts rotation to 20-30 degrees and deliberately avoids large side views; this is a practical limitation that should be mentioned in the main text rather than only in the supplement.
Circularity Check
Training-frame metrics make the quantitative comparison partly circular; the core shading derivation is self-contained.
-
fitted input called prediction
[Section 4.2, Table 1; Supplementary Section 9.1]
"For fairness, we compute errors only on skin regions and average over all subjects in the dataset. Our method prevails in all metrics. ... The statistics in Table 1 are averaged over frames used in the optimization, i.e., all the frames for our method, FLARE, SunStage, and only three frames for NextFace."
Table 1 is presented as the paper's quantitative superiority evidence, but the metrics are computed on the very frames used to minimize the image reconstruction loss Limg in Eq. 10. The model is fit to those images, so low reconstruction error on those images is partly forced by construction; it demonstrates training-set fit, not an independent prediction of the appearance decomposition or its generalization under relighting. The synthetic ground-truth evaluation uses a Lambertian material, so it never validates specular intensity/roughness; the qualitative relighting comparisons are the only independent check, leaving the headline quantitative support partly circular.
full rationale
The core derivation is not circular. The shading model begins from the standard rendering equation (Eq. 2), adopts the external split-sum approximation (Karis; Munkberg et al.), and derives a visibility-modulated form (Eqs. 6-7) with the limitation "introduces large errors for rough surfaces" disclosed. The recovered maps are outputs of an optimization over the image loss and regularizers, not quantities defined in terms of the target claim. Self-citations to Chandran et al. [10, 11] and Riviere et al. [55] are data sources, a landmark detector, and a studio reference capture; they do not carry the argument for the occlusion-aware shading model, so they are not load-bearing circularity. The concrete circular element is Table 1, which reports reconstruction errors on the same frames used for optimization, as the supplement states. The paper's limitations (head-pose sensitivity, extreme shadows, skin-tone ambiguity, Lambertian-only synthetic specular evaluation) are correctness risks and missing evidence, not additional circular steps. Weighted overall, this is partial circularity in the quantitative evaluation rather than in the derivation.
Assumptions & free parameters
free parameters (6)
- λgeo =
19
- λLap =
10
- λdiffuse =
0.01
- λrough =
0.1
- λlight =
0.1
- λmask =
0.1
assumptions (6)
- domain assumption The split-sum approximation (Karis 2013) adequately models specular reflection of environment lighting.
- ad hoc to paper The visibility-modulated split-sum approximation in Eq. 6 is accurate enough for the specular component during optimization.
- domain assumption The initial 3DMM tracking provides accurate head poses and a reasonable initial mesh.
- domain assumption Lighting is static and expression is constant during the capture sequence.
- domain assumption The PCA face basis from Chandran et al. [10] spans the identities and expressions needed.
- domain assumption The Lambertian diffuse plus Kelemen and Szirmay-Kalos specular BRDF is suitable for skin.
Cite this review
Pith. "Pith review of Monocular Facial Appearance Capture in the Wild." pith.science (2026). https://pith.science/paper/FJUMLG52
@misc{pith2026241212765,
author = {Pith},
title = {Pith review of: Monocular Facial Appearance Capture in the Wild},
year = {2026},
howpublished = {\url{https://pith.science/paper/FJUMLG52}},
note = {Machine review of arXiv:2412.12765}
}
read the original abstract
We present a new method for reconstructing the appearance properties of human faces from a lightweight capture procedure in an unconstrained environment. Our method recovers the surface geometry, diffuse albedo, specular intensity and specular roughness from a monocular video containing a simple head rotation in-the-wild. Notably, we make no simplifying assumptions on the environment lighting, and we explicitly take visibility and occlusions into account. As a result, our method can produce facial appearance maps that approach the fidelity of studio-based multi-view captures, but with a far easier and cheaper procedure.
Figures
Figures from the paper (11 more)
Reference graph
Works this paper leans on
-
[1]
High-res facial appearance cap- ture from polarized smartphone images
Dejan Azinovi ´c, Olivier Maury, Christophe Hery, Matthias Nießner, and Justus Thies. High-res facial appearance cap- ture from polarized smartphone images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16836–16846, 2023. 2, 1
work page 2023
-
[2]
Spark: Self-supervised personalized real-time monocular face capture
Kelian Baert, Shrisha Bharadwaj, Fabien Castan, Benoit Maujean, Marc Christie, Victoria Abrevaya, and Adnane Boukhayma. Spark: Self-supervised personalized real-time monocular face capture. arXiv preprint arXiv:2409.07984 ,
-
[3]
High-quality single-shot capture of fa- cial geometry
Thabo Beeler, Bernd Bickel, Paul Beardsley, Bob Sumner, and Markus Gross. High-quality single-shot capture of fa- cial geometry. In ACM SIGGRAPH 2010 papers, pages 1–9
work page 2010
-
[4]
Flare: Fast learning of animatable and relightable mesh avatars
Shrisha Bharadwaj, Yufeng Zheng, Otmar Hilliges, Michael J Black, and Victoria Fernandez-Abrevaya. Flare: Fast learning of animatable and relightable mesh avatars. arXiv preprint arXiv:2310.17519, 2023. 2, 3, 5, 6, 1
arXiv 2023
-
[5]
Gs 3: Efficient relighting with triple gaussian splatting
Zoubin Bi, Yixin Zeng, Chong Zeng, Fan Pei, Xiang Feng, Kun Zhou, and Hongzhi Wu. Gs 3: Efficient relighting with triple gaussian splatting. In SIGGRAPH Asia 2024 Confer- ence Papers, 2024. 2
work page 2024
-
[6]
A morphable model for the synthesis of 3d faces
V olker Blanz and Thomas Vetter. A morphable model for the synthesis of 3d faces. In Seminal Graphics Papers: Pushing the Boundaries, Volume 2, pages 157–164. 2023. 2
2023
-
[7]
Nerd: Neural reflectance decomposition from image collections
Mark Boss, Raphael Braun, Varun Jampani, Jonathan T Bar- ron, Ce Liu, and Hendrik Lensch. Nerd: Neural reflectance decomposition from image collections. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, pages 12684–12694, 2021. 2
work page 2021
-
[8]
Neural-pil: Neu- ral pre-integrated lighting for reflectance decomposition
Mark Boss, Varun Jampani, Raphael Braun, Ce Liu, Jonathan Barron, and Hendrik Lensch. Neural-pil: Neu- ral pre-integrated lighting for reflectance decomposition. Advances in Neural Information Processing Systems , 34: 10691–10704, 2021. 2
work page 2021
Show all 79 references
-
[9]
Samurai: Shape and material from uncon- strained real-world arbitrary image collections
Mark Boss, Andreas Engelhardt, Abhishek Kar, Yuanzhen Li, Deqing Sun, Jonathan Barron, Hendrik Lensch, and Varun Jampani. Samurai: Shape and material from uncon- strained real-world arbitrary image collections. Advances in Neural Information Processing Systems, 35:26389–26403,
-
[10]
Semantic deep face models
Prashanth Chandran, Derek Bradley, Markus Gross, and Thabo Beeler. Semantic deep face models. In 2020 interna- tional conference on 3D vision (3DV), pages 345–354. IEEE,
2020
-
[11]
Continuous landmark detection with 3d queries
Prashanth Chandran, Gaspard Zoss, Paulo Gotardo, and Derek Bradley. Continuous landmark detection with 3d queries. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 16858– 16867, 2023. 2, 1
2023
-
[12]
Learn- ing to predict 3d objects with an interpolation-based differ- entiable renderer
Wenzheng Chen, Huan Ling, Jun Gao, Edward Smith, Jaakko Lehtinen, Alec Jacobson, and Sanja Fidler. Learn- ing to predict 3d objects with an interpolation-based differ- entiable renderer. Adv. Neural Inf. Process. Syst., 32, 2019. 3
2019
-
[13]
Differentiable display photo- metric stereo
Seokjun Choi, Seungwoo Yoon, Giljoo Nam, Seungyong Lee, and Seung-Hwan Baek. Differentiable display photo- metric stereo. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 11831– 11840, 2024. 2
2024
-
[14]
Acquiring the reflectance field of a human face
Paul Debevec, Tim Hawkins, Chris Tchou, Haarm-Pieter Duiker, Westley Sarokin, and Mark Sagar. Acquiring the reflectance field of a human face. In Proceedings of the 27th annual conference on Computer graphics and interac- tive techniques, pages 145–156, 2000. 2
2000
-
[15]
Practical face reconstruction via differentiable ray tracing
Abdallah Dib, Gaurav Bharaj, Junghyun Ahn, C ´edric Th´ebault, Philippe Gosselin, Marco Romeo, and Louis Chevallier. Practical face reconstruction via differentiable ray tracing. In Computer Graphics Forum, pages 153–164. Wiley Online Library, 2021. 2, 5, 6, 7, 1
2021
-
[16]
Towards high fidelity monocular face reconstruction with rich reflectance using self-supervised learning and ray trac- ing
Abdallah Dib, Cedric Thebault, Junghyun Ahn, Philippe- Henri Gosselin, Christian Theobalt, and Louis Chevallier. Towards high fidelity monocular face reconstruction with rich reflectance using self-supervised learning and ray trac- ing. In Proceedings of the IEEE/CVF Internati...
2021
-
[17]
S2f2: Self-supervised high fidelity face reconstruction from monocular image
Abdallah Dib, Junghyun Ahn, Cedric Thebault, Philippe- Henri Gosselin, and Louis Chevallier. S2f2: Self-supervised high fidelity face reconstruction from monocular image. arXiv preprint arXiv:2203.07732, 2022. 1
2022 arXiv
-
[18]
Mosar: Monocular semi-supervised model for avatar reconstruction using differentiable shading
Abdallah Dib, Luiz Gustavo Hafemann, Emeline Got, Trevor Anderson, Amin Fadaeinejad, Rafael MO Cruz, and Marc- Andr´e Carbonneau. Mosar: Monocular semi-supervised model for avatar reconstruction using differentiable shading. In Proceedings of the IEEE/CVF Conference on Compute...
2024
-
[19]
Adam: A method for stochastic opti- mization
P Kingma Diederik. Adam: A method for stochastic opti- mization. (No Title), 2014. 1
2014
-
[20]
Shi- nobi: Shape and illumination using neural object decompo- sition via brdf optimization in-the-wild
Andreas Engelhardt, Amit Raj, Mark Boss, Yunzhi Zhang, Abhishek Kar, Yuanzhen Li, Deqing Sun, Ricardo Martin Brualla, Jonathan T Barron, Hendrik Lensch, et al. Shi- nobi: Shape and illumination using neural object decompo- sition via brdf optimization in-the-wild. In Proceedin...
2024
-
[21]
Diffusion reflectance map: Single-image stochastic inverse rendering of illumination and reflectance
Yuto Enyo and Ko Nishino. Diffusion reflectance map: Single-image stochastic inverse rendering of illumination and reflectance. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) ,
-
[22]
Learning an animatable detailed 3d face model from in-the- wild images
Yao Feng, Haiwen Feng, Michael J Black, and Timo Bolkart. Learning an animatable detailed 3d face model from in-the- wild images. ACM Transactions on Graphics (ToG), 40(4): 1–13, 2021. 2
2021
-
[23]
Relightable 3d gaussian: Real-time point cloud relighting with brdf decomposition and ray trac- ing
Jian Gao, Chun Gu, Youtian Lin, Hao Zhu, Xun Cao, Li Zhang, and Yao Yao. Relightable 3d gaussian: Real-time point cloud relighting with brdf decomposition and ray trac- ing. arXiv preprint arXiv:2311.16043, 2023. 2
2023 arXiv
-
[24]
Morphable face models-an open framework
Thomas Gerig, Andreas Morel-Forster, Clemens Blumer, Bernhard Egger, Marcel Luthi, Sandro Sch ¨onborn, and Thomas Vetter. Morphable face models-an open framework. In 2018 13th IEEE international conference on automatic 9 face & gesture recognition (FG 2018) , pages 75–82. IEEE,
2018
-
[25]
Multiview face cap- ture using polarized spherical gradient illumination
Abhijeet Ghosh, Graham Fyffe, Borom Tunwattanapong, Jay Busch, Xueming Yu, and Paul Debevec. Multiview face cap- ture using polarized spherical gradient illumination. In Pro- ceedings of the 2011 SIGGRAPH Asia Conference, pages 1– 10, 2011. 2
2011
-
[26]
Practical dynamic facial ap- pearance modeling and acquisition
Paulo Gotardo, J ´er´emy Riviere, Derek Bradley, Abhijeet Ghosh, and Thabo Beeler. Practical dynamic facial ap- pearance modeling and acquisition. ACM Transactions on Graphics (ToG), 37(6):1–13, 2018. 2
2018
-
[27]
Learning a 3d mor- phable face reflectance model from low-cost data
Yuxuan Han, Zhibo Wang, and Feng Xu. Learning a 3d mor- phable face reflectance model from low-cost data. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 8598–8608, 2023. 2
2023
-
[28]
High-quality facial geometry and appearance capture at home
Yuxuan Han, Junfeng Lyu, and Feng Xu. High-quality facial geometry and appearance capture at home. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 697–707, 2024. 2, 3
2024
-
[29]
Shape, Light, and Material Decomposition from Im- ages using Monte Carlo Rendering and Denoising
Jon Hasselgren, Nikolai Hofmann, and Jacob Munkberg. Shape, Light, and Material Decomposition from Im- ages using Monte Carlo Rendering and Denoising. arXiv:2206.03380, 2022. 2
2022 arXiv
-
[30]
Mitsuba 3 renderer, 2022
Wenzel Jakob, S ´ebastien Speierer, Nicolas Roussel, Merlin Nimier-David, Delio Vicini, Tizian Zeltner, Baptiste Nicolet, Miguel Crespo, Vincent Leroy, and Ziyi Zhang. Mitsuba 3 renderer, 2022. https://mitsuba-renderer.org. 2
2022
-
[31]
The rendering equation
James T Kajiya. The rendering equation. In Proceedings of the 13th annual conference on Computer graphics and inter- active techniques, pages 143–150, 1986. 3
1986
-
[32]
Real shading in unreal engine
Brian Karis and Epic Games. Real shading in unreal engine
-
[33]
Physically Based Shading Theory Practice, 4(3):1,
Proc. Physically Based Shading Theory Practice, 4(3):1,
-
[34]
A unified approach to prefiltered envi- ronment maps
Jan Kautz, Pere-Pau V ´azquez, Wolfgang Heidrich, and Hans-Peter Seidel. A unified approach to prefiltered envi- ronment maps. In Rendering Techniques 2000: Proceedings of the Eurographics Workshop in Brno, Czech Republic, June 26–28, 2000 11, pages 185–196. Springer, 2000. 3
2000
-
[35]
Modnet: Real-time trimap-free portrait mat- ting via objective decomposition
Zhanghan Ke, Jiayu Sun, Kaican Li, Qiong Yan, and Ryn- son WH Lau. Modnet: Real-time trimap-free portrait mat- ting via objective decomposition. In Proceedings of the AAAI Conference on Artificial Intelligence , pages 1140– 1147, 2022. 4
2022
-
[36]
A microfacet based coupled specular-matte brdf model with importance sampling
Csaba Kelemen and Laszlo Szirmay-Kalos. A microfacet based coupled specular-matte brdf model with importance sampling. In Eurographics (short presentations), 2001. 3
2001
-
[37]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42 (4), 2023. 2
2023
-
[38]
Modular primitives for high-performance differentiable rendering
Samuli Laine, Janne Hellsten, Tero Karras, Yeongho Seol, Jaakko Lehtinen, and Timo Aila. Modular primitives for high-performance differentiable rendering. ACM Transac- tions on Graphics, 39(6), 2020. 2, 1
2020
-
[39]
Avatarme: Realistically ren- derable 3d facial reconstruction” in-the-wild”
Alexandros Lattas, Stylianos Moschoglou, Baris Gecer, Stylianos Ploumpis, Vasileios Triantafyllou, Abhijeet Ghosh, and Stefanos Zafeiriou. Avatarme: Realistically ren- derable 3d facial reconstruction” in-the-wild”. In Proceed- ings of the IEEE/CVF conference on computer visio...
2020
-
[40]
Practical and scalable desktop- based high-quality facial capture
Alexandros Lattas, Yiming Lin, Jayanth Kannan, Ekin Ozturk, Luca Filipi, Giuseppe Claudio Guarnera, Gaurav Chawla, and Abhijeet Ghosh. Practical and scalable desktop- based high-quality facial capture. In European Conference on Computer Vision, pages 522–537. Springer, 2022. 2
2022
-
[41]
Differentiable monte carlo ray tracing through edge sampling
Tzu-Mao Li, Miika Aittala, Fr ´edo Durand, and Jaakko Lehti- nen. Differentiable monte carlo ray tracing through edge sampling. ACM Trans. Graph. (Proc. SIGGRAPH Asia), 37 (6):222:1–222:11, 2018. 2
2018
-
[42]
Photorealistic object insertion with diffusion-guided inverse rendering
Ruofan Liang, Zan Gojcic, Merlin Nimier-David, David Acuna, Nandita Vijaykumar, Sanja Fidler, and Zian Wang. Photorealistic object insertion with diffusion-guided inverse rendering. arXiv preprint, 2024. 2
2024
-
[43]
Gs-ir: 3d gaussian splatting for inverse rendering
Zhihao Liang, Qi Zhang, Ying Feng, Ying Shan, and Kui Jia. Gs-ir: 3d gaussian splatting for inverse rendering. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 21644–21653, 2024. 2
2024
-
[44]
Diffusion posterior illumination for ambiguity-aware inverse rendering
Linjie Lyu, Ayush Tewari, Marc Habermann, Shun- suke Saito, Michael Zollh ¨ofer, Thomas Leimk ¨uehler, and Christian Theobalt. Diffusion posterior illumination for ambiguity-aware inverse rendering. ACM Transactions on Graphics, 42(6), 2023. 2
2023
-
[45]
T. M. MacRobert. Spherical harmonics: An elementary trea- tise on harmonic functions and their applications, 1928. 3
1928
-
[46]
Efficient rendering of spatial bi-directional reflectance distribution functions
David K McAllister, Anselmo Lastra, and Wolfgang Hei- drich. Efficient rendering of spatial bi-directional reflectance distribution functions. In Proceedings of the ACM SIG- GRAPH/EUROGRAPHICS conference on Graphics hard- ware, pages 79–88, 2002. 3
2002
-
[47]
Nerf: Representing scenes as neural radiance fields for view syn- thesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. Communications of the ACM , 65(1):99–106, 2021. 2
2021
-
[48]
Instant neural graphics primitives with a mul- tiresolution hash encoding
Thomas M ¨uller, Alex Evans, Christoph Schied, and Alexan- der Keller. Instant neural graphics primitives with a mul- tiresolution hash encoding. ACM transactions on graphics (TOG), 41(4):1–15, 2022. 2
2022
-
[49]
Extracting Triangular 3D Models, Materials, and Light- ing From Images
Jacob Munkberg, Jon Hasselgren, Tianchang Shen, Jun Gao, Wenzheng Chen, Alex Evans, Thomas M¨uller, and Sanja Fi- dler. Extracting Triangular 3D Models, Materials, and Light- ing From Images. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognitio...
2022
-
[50]
GPU Gems 3
Hubert Nguyen. GPU Gems 3 . Addison-Wesley Profes- sional, 2007. 3
2007
-
[51]
Large steps in inverse rendering of geometry
Baptiste Nicolet, Alec Jacobson, and Wenzel Jakob. Large steps in inverse rendering of geometry. ACM Transactions on Graphics (TOG), 40(6):1–13, 2021. 3, 4
2021
-
[52]
Relightify: Re- lightable 3d faces from a single image via diffusion models
Foivos Paraperas Papantoniou, Alexandros Lattas, Stylianos Moschoglou, and Stefanos Zafeiriou. Relightify: Re- lightable 3d faces from a single image via diffusion models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 8806–8817, 2023. 2 10
2023
-
[53]
Optix: a general purpose ray tracing engine
Steven G Parker, James Bigler, Andreas Dietrich, Heiko Friedrich, Jared Hoberock, David Luebke, David McAllis- ter, Morgan McGuire, Keith Morley, Austin Robison, et al. Optix: a general purpose ray tracing engine. Acm transac- tions on graphics (tog), 29(4):1–13, 2010. 4, 1
2010
-
[54]
Versatile head alignment with adaptive ap- pearance priors
Shenhan Qian. Versatile head alignment with adaptive ap- pearance priors. 2024. 2, 1
2024
-
[55]
Neu- ral shading fields for efficient facial inverse rendering
Gilles Rainer, Lewis Bridgeman, and Abhijeet Ghosh. Neu- ral shading fields for efficient facial inverse rendering. In Computer Graphics Forum, page e14943. Wiley Online Li- brary, 2023. 2, 1
2023
-
[56]
Single-shot high-quality facial geometry and skin appearance capture
J ´er´emy Riviere, Paulo FU Gotardo, Derek Bradley, Abhijeet Ghosh, and Thabo Beeler. Single-shot high-quality facial geometry and skin appearance capture. ACM Trans. Graph., 39(4):81, 2020. 1, 2, 6
2020
-
[57]
Relightable gaussian codec avatars
Shunsuke Saito, Gabriel Schwartz, Tomas Simon, Junxuan Li, and Giljoo Nam. Relightable gaussian codec avatars. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, pages 130–141, 2024. 2, 4
2024
-
[58]
Litnerf: Intrinsic radiance decomposition for high-quality view syn- thesis and relighting of faces
Kripasindhu Sarkar, Marcel C B ¨uhler, Gengyan Li, Daoye Wang, Delio Vicini, J ´er´emy Riviere, Yinda Zhang, Sergio Orts-Escolano, Paulo Gotardo, Thabo Beeler, et al. Litnerf: Intrinsic radiance decomposition for high-quality view syn- thesis and relighting of faces. InSIGGRAP...
2023
-
[59]
William A. P. Smith, Alassane Seck, Hannah Dee, Bernard Tiddeman, Joshua Tenenbaum, and Bernhard Egger. A mor- phable face albedo model. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 5011–5020, 2020. 2
2020
-
[60]
Nerv: Neural reflectance and visibility fields for relighting and view synthesis
Pratul P Srinivasan, Boyang Deng, Xiuming Zhang, Matthew Tancik, Ben Mildenhall, and Jonathan T Barron. Nerv: Neural reflectance and visibility fields for relighting and view synthesis. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pag...
2021
-
[61]
Optimally combining sampling techniques for monte carlo rendering
Eric Veach and Leonidas J Guibas. Optimally combining sampling techniques for monte carlo rendering. In Proceed- ings of the 22nd annual conference on Computer graphics and interactive techniques, pages 419–428, 1995. 4
1995
-
[62]
Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction
Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. arXiv preprint arXiv:2106.10689, 2021. 2
2021 arXiv
-
[63]
Sunstage: Portrait reconstruction and re- lighting using the sun as a light stage
Yifan Wang, Aleksander Holynski, Xiuming Zhang, and Xuaner Zhang. Sunstage: Portrait reconstruction and re- lighting using the sun as a light stage. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 20792–20802, 2023. 1, 2, 5, 6, 7
2023
-
[64]
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Si- moncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004. 6
2004
-
[65]
Deferredgs: Decoupled and editable gaussian splatting with deferred shading
Tong Wu, Jia-Mu Sun, Yu-Kun Lai, Yuewen Ma, Leif Kobbelt, and Lin Gao. Deferredgs: Decoupled and editable gaussian splatting with deferred shading. arXiv preprint arXiv:2404.09412, 2024. 2
2024 arXiv
-
[66]
Chen Xi, Peng Sida, Yang Dongchen, Liu Yuan, Pan Bowen, Lv Chengfei, and Zhou. Xiaowei. Intrinsicanything: Learn- ing diffusion priors for inverse rendering under unknown il- lumination. arxiv: 2404.11593, 2024. 2
2024 arXiv
-
[67]
Renerf: Re- lightable neural radiance fields with nearfield lighting
Yingyan Xu, Gaspard Zoss, Prashanth Chandran, Markus Gross, Derek Bradley, and Paulo Gotardo. Renerf: Re- lightable neural radiance fields with nearfield lighting. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 22581–22591, 2023. 2
2023
-
[68]
Artist-friendly re- lightable and animatable neural heads
Yingyan Xu, Prashanth Chandran, Sebastian Weiss, Markus Gross, Gaspard Zoss, and Derek Bradley. Artist-friendly re- lightable and animatable neural heads. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2457–2467, 2024. 2
2024
-
[69]
Accu- rate translucent material rendering under spherical gaussian lights
Ling-Qi Yan, Yahan Zhou, Kun Xu, and Rui Wang. Accu- rate translucent material rendering under spherical gaussian lights. Comput. Graph. Forum, 31(7):2267–2276, 2012. 3
2012
-
[70]
Towards practical capture of high-fidelity relightable avatars
Haotian Yang, Mingwu Zheng, Wanquan Feng, Haibin Huang, Yu-Kun Lai, Pengfei Wan, Zhongyuan Wang, and Chongyang Ma. Towards practical capture of high-fidelity relightable avatars. In SIGGRAPH Asia 2023 Conference Proceedings, 2023. 2
2023
-
[71]
PhySG: Inverse rendering with spherical gaussians for physics-based material editing and relighting
Kai Zhang, Fujun Luan, Qianqian Wang, Kavita Bala, and Noah Snavely. PhySG: Inverse rendering with spherical gaussians for physics-based material editing and relighting. In The IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), 2021. 2, 3
2021
-
[72]
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shecht- man, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recogni- tion, pages 586–595, 2018. 6
2018
-
[73]
Ner- factor: Neural factorization of shape and reflectance under an unknown illumination
Xiuming Zhang, Pratul P Srinivasan, Boyang Deng, Paul De- bevec, William T Freeman, and Jonathan T Barron. Ner- factor: Neural factorization of shape and reflectance under an unknown illumination. ACM Transactions on Graphics (ToG), 40(6):1–18, 2021. 2
2021
-
[74]
Neuface: Realistic 3d neural face rendering from multi-view images
Mingwu Zheng, Haiyu Zhang, Hongyu Yang, and Di Huang. Neuface: Realistic 3d neural face rendering from multi-view images. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 16868– 16877, 2023. 2, 3
2023
-
[75]
Gs-ror: 3d gaussian splatting for reflective object relighting via sdf pri- ors
Zuo-Liang Zhu, Beibei Wang, and Jian Yang. Gs-ror: 3d gaussian splatting for reflective object relighting via sdf pri- ors. arXiv e-prints, pages arXiv–2406, 2024. 2 11 Monocular Facial Appearance Capture in the Wild Supplementary Material In this supplementary document we beg...
2024
-
[76]
Capture Protocol We record videos at 25 fps using a Canon EOS 1200D cam- era fixed on a tripod
Dataset Details 7.1. Capture Protocol We record videos at 25 fps using a Canon EOS 1200D cam- era fixed on a tripod. Depending on the distance to the sub- ject, the camera is mounted with either a 35mm or 60mm lens and we use the corresponding focal length as a known parameter...
-
[77]
We then uniformly sample around 250 frames to use in the inverse rendering
Implementation Details We capture 500 to 800 frames for each subject. We then uniformly sample around 250 frames to use in the inverse rendering. We did not observe a performance gain or drop when using all of the frames. The images are cropped to 1K resolution. We also solve ...
-
[78]
Implementation of the Related Methods Next we describe the steps performed to run the compar- isons
Experiment Details 9.1. Implementation of the Related Methods Next we describe the steps performed to run the compar- isons. We use the default parameters of FLARE [4]. We noticed that the FLARE geometry is very bumpy, hence we tried setting a larger weight on the Laplacian me...
-
[79]
Visualization of the Visibility
Additional Results Next we present additional results for the ablation and show some challenging situations for our algorithm. Visualization of the Visibility. First we provide addi- tional visualizations of the view-dependent visibility under different roughness in Fig. 12. T...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.