REVIEW 5 major objections 6 minor 90 references
GraphicsDreamer: Image to 3D Generation with Physical Consistency
T0 review · 5 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read GraphicsDreamer produces a relightable 3D mesh with full PBR maps from a single input image.
desk verdict Solid engineering with a useful six-domain PBR-aware pipeline, but the closed-form SG/BRDF derivation is flawed as written and the evaluation is too thin to support the SOTA claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is the simplified physically based rendering equation: outgoing radiance is the integral over the hemisphere of incoming spherical-Gaussian lighting times a Disney BRDF approximation times the cosine term, all expressed as spherical Gaussians so the integral has a closed form (Eqs. 7–15). A spherical Gaussian is a directional function $\mu e^{\lambda(v\cdot p - 1)}$, with lobe axis $p$, sharpness $\lambda$, and amplitude $\mu$. A 16-lobe SG mixture represents the environment light; the BRDF's normal distribution term is wrapped into one SG; the cosine factor is a fixed SG; and the product of SGs integrates analytically. This rendering equation appears twice: once as supervision inside the multi-view diffusion model, where it ties albedo, roughness, metallic, and normal predictions back to the color image, and again in the inverse-rendering stage, where a material MLP predicts those properties at surface points and the same equation is enforced during optimization. A secondary mechanism is the mixed surface representation: an implicit SDF provides differentiable geometry, and explicit z-buffer-guided interpolation between sign-change sample pairs yields smooth intersection points compatible with surface reflection.
What would settle it
Take an input image containing sharp, high-frequency reflections or interreflections (for example, a polished metal object under a point light) and compare the predicted albedo, roughness, and metallic maps against ground-truth captures made under measured environment lighting; if the rendering-equation supervision forces a wrong decomposition, visible as albedo bleeding in highlights or relighting artifacts, then the SG lighting model is too weak and the central claim fails.
Extended reading notes
Core claim
On its own terms, the paper's discovery is that modeling a 3D asset as a joint distribution over six aligned image domains — color, normal, depth, albedo, roughness, and metallic — and supervising that distribution with the rendering equation yields multi-view predictions consistent enough to reconstruct a high-quality surface mesh. The color domain is treated as primary; cross-attention from color queries to the other domains aligns geometry and materials, and the closed-form PBR integral (Eqs. 7–15) forces the four material and geometry maps to jointly reproduce the observed color. The same PBR constraint is re-applied during inverse rendering, where a mixed implicit-SDF and explicit-surface representation keeps optimization stable while resolving surface intersections smoothly. The result is a clean-topology, UV-unwrapped, fully textured mesh with the best novel-view PSNR (27.93) and the best Chamfer distance (0.0231) and volume IoU (0.5779) on GSO among the compared baselines.
Load-bearing premise
The load-bearing premise is that a 16-lobe spherical-Gaussian mixture plus the closed-form Disney BRDF approximation captures the relevant lighting and reflection behavior in both the training renders and real single input images.
Editorial extensions
If this is right
- After automated quad remeshing and UV unwrapping, generated assets can be imported directly into standard graphics engines without manual cleanup.
- Because albedo, roughness, and metallic are decoupled from lighting, the same mesh can be relit under different environment maps with consistent material appearance.
- Novel-view synthesis improves to 27.93 PSNR and 0.937 SSIM, above the compared RGB-only and RGB-normal baselines, and reconstruction achieves the best Chamfer distance and volume IoU on GSO.
- Highlights, shadows, metallic surfaces, and even some transparent materials are separated from true surface color, as shown in the predicted albedo and metallic channels.
- PBR supervision in the diffusion stage gives the inverse-rendering stage richer pseudo ground truth, reducing the ambiguity that plagues color-only multi-view reconstruction.
Reading between the lines
- If the 16-SG lighting model holds, the same PBR-constrained diffusion objective could be trained on images captured under arbitrary real environments, not just synthetic renders, and should improve material decomposition on in-the-wild photos.
- The rendering-equation supervision could act as a self-consistency loss for video or multi-image input, chaining frames and letting lighting be estimated jointly rather than per image.
- A direct test of the method's ceiling is material-editing benchmarks: replacing estimated albedo while keeping roughness and metallic should produce plausible appearance changes if the decomposition is truly physical.
- The explicit-surface sampling trick may generalize to other volumetric or SDF-based reconstruction pipelines by reducing the number of ray-marching iterations needed near the surface.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. GraphicsDreamer proposes a two-stage pipeline for image-to-3D generation that integrates physically based rendering (PBR) constraints into both multi-view synthesis and geometry reconstruction. The first stage extends a Wonder3D-style cross-domain diffusion model to jointly predict six domains: color, normal, depth, albedo, roughness, and metallic. A PBR lighting model, using a 16-lobe spherical-Gaussian environment map and a simplified Disney BRDF, is used to supervise consistency among these domains. The second stage reconstructs a surface mesh from the generated pseudo ground truth via a mixed implicit/explicit SDF representation, with a material MLP and the same PBR model for inverse rendering. A final asset-enhancement step performs quadrangulation, UV unwrapping, and baking to produce engine-ready assets. The paper reports improved novel-view synthesis PSNR/SSIM and geometry metrics on the GSO dataset, with qualitative relighting results.
Significance. If the technical derivation and evaluation are sound, the paper addresses a practically important gap: generating 3D assets that are directly usable in graphics engines, with clean topology, UVs, and complete PBR material maps. The idea of embedding PBR conditions into both diffusion and inverse rendering is timely and relevant to the graphics and vision communities. The results are promising: the method achieves the best reported novel-view PSNR (27.93) and geometry (CD 0.0231, IoU 0.5779) among the compared baselines, and the qualitative figures show convincing material separation and relighting. However, the central derivation in Sec. 3.3 contains a potential variable inconsistency in the SG-BRDF formulation, and the quantitative evaluation lacks error bars, significance tests, ablations, and material-quality metrics. These issues currently prevent the reader from confirming the claimed 'physical consistency' advantage.
major comments (5)
- [Sec. 3.3, Eq. (12)-(14)] The closed-form SG integration claim is not justified as written. Eq. (12) approximates the NDF D(h) as an SG in the half-vector h with lobe axis p_w = 2(ω_o·n)n − ω_o; for fixed ω_o, D(h) is maximized at h = n, not at the reflected view direction p_w. Eq. (14) then writes f_s ≈ Gs(h; p_w, λ_w/(4|ω_o·n|), F0G0μ_w), i.e., an SG in h, while the lighting in Eq. (11) and the cosine factor in Eq. (15) are SGs in ω_i. A product of SGs in different variables is not an SG in ω_i, so the integral in Eq. (7) cannot be evaluated by the cited closed-form formulas without an explicit change of variables and its Jacobian. The standard SG-BRDF representation (e.g., Wang et al., 2009) expresses the specular lobe as an SG in ω_i centered at p_w. Please either correct Eq. (14) to use ω_i as the SG variable, or provide the complete change-of-variable derivation showing how the product becomes an SG in ω_i. Since no code is released, this ambiguity blocks verification of the core physical-consistency contribution.
- [Sec. 4.2-4.4, Tables 1 and 2] The quantitative evaluation reports single-point metrics with no error bars, variance, or significance tests. The geometry results in Table 2 are extremely close: against Wonder3D, Chamfer distance improves from 0.0237 to 0.0231 (about 2.5% relative), and volume IoU from 0.5762 to 0.5779 (about 0.3% relative). Without per-sample statistics or a paired significance test, these differences are not demonstrably meaningful. The novel-view PSNR gain over Wonder3D (27.93 vs. 24.11) is larger, but still lacks confidence intervals. Please report the number of test assets, mean±std or per-sample distributions, and appropriate significance tests (e.g., paired t-test or Wilcoxon signed-rank).
- [Sec. 3.1 and Sec. 3.3] No ablation isolates the contribution of the PBR condition, which is the paper's central claim. There is no experiment that removes the PBR supervision from either the diffusion stage or the inverse-rendering stage, nor an ablation varying the number of SG lobes, nor an ablation of the mixed surface representation. Consequently, the observed improvements could come from the additional domains, the cross-domain attention, or the reconstruction details rather than from the PBR conditioning. Please include at least a variant without the PBR rendering loss and report the same metrics; this experiment is essential to substantiate the 'physical consistency' claim.
- [Sec. 4.5] The material and relighting claims are supported only by qualitative figures (Figs. 7 and 8). Since the paper emphasizes complete PBR maps and reliable relighting, the absence of any quantitative material fidelity metric is a significant gap. The GSO dataset contains ground-truth 3D assets; please report, for example, albedo/roughness/metallic estimation error against the ground-truth material maps used for the GSO renderings, or a relighting consistency measure that compares re-rendered images under novel environment maps to held-out ground-truth renderings.
- [Sec. 3.1] The text states that extending Wonder3D's cross-domain attention from two to six domains 'preserves the prior knowledge of the pretrained model and supports fast convergence and robust generalization,' but no experimental evidence for this specific claim is provided. No convergence curves or ablation comparing two-domain versus six-domain training are shown. Please either provide supporting experiments or temper the claim accordingly.
minor comments (6)
- [Abstract and Introduction] Several grammatical errors should be corrected, e.g., 'significantly lags in industrial application' and 'Traditional 3D modeling processes heavily relies on manual labor.'
- [Sec. 3.2, Eq. (4)] The word 'Neus' should be 'NeuS' for consistency with the citation [72].
- [Fig. 2 caption] 'our method will product appealing 3D assets' should read 'will produce appealing 3D assets.'
- [Sec. 3.4] 'weunwrap the UVs' is missing a space; it should be 'we unwrap the UVs.'
- [References] Several reference entries include stray page-number suffixes (e.g., '[3] Blender Online Community ... 2024. 2, 7, 1' and '[12] Objaverse ... 2023. 3, 7, 1, 2'), which appear to be citation-page remnants from the source files. These should be cleaned up.
- [Sec. 4.3] The qualitative claim that Wonder3D 'sometimes produces distorted geometries and struggles with complex structures' may be true, but it is not directly supported by the presented quantitative tables; consider adding per-category breakdowns or failure-case images.
Circularity Check
No significant circularity: the PBR supervision uses standard external rendering/BRDF/SG models, and the reported results are benchmarked against the external GSO dataset.
full rationale
The derivation chain is self-contained rather than circular. The paper's central technical component is the PBR rendering-equation supervision of Sec. 3.3, which combines the standard rendering equation (Eq. 7), the Disney BRDF (Eqs. 8-9), and spherical-Gaussian representations of lighting, BRDF, and cosine terms (Eqs. 10-15). These are external, established models cited from the literature, not quantities fitted to the paper's own outputs. The 16-SG lighting mixture (Eq. 11) is a representational assumption, not a parameter fitted to the predicted images and then renamed as a prediction. The consistency loss between predicted color and predicted normal/albedo/roughness/metallic is a self-supervision constraint, but it does not make the benchmark numbers forced: novel-view PSNR/SSIM (Table 1) and Chamfer distance/volume IoU (Table 2) are measured against held-out GSO ground truth, and the comparison methods are external baselines. There are no load-bearing self-citations, no imported uniqueness theorems, and no ansatz smuggled in via citation that itself rests on this paper's claims. One could raise a mathematical-correctness concern about Eq. 14's use of SGs in different variables (the half-vector h versus the incident direction ω_i), but that is a correctness risk about whether the closed-form integral is valid as written, not a circularity in which a result reduces by construction to its own inputs.
Assumptions & free parameters
free parameters (6)
- Number of spherical Gaussian lobes for lighting (N) =
16
- Number of camera views (K) =
6
- Number of uniform samples between surface-crossing points (d) =
8
- Probability of interpolating a single surface point (p) =
0.5
- Objaverse filtering count (approx.) =
32,000
- Number of ray samples for SDF sign tests (N) =
64
assumptions (5)
- domain assumption The rendering equation (Eq. 7) with the simplified Disney BRDF (Eqs. 8-14) is an adequate physical model of image formation for the objects and lighting considered.
- domain assumption A mixture of 16 spherical Gaussians can represent the relevant environment lighting for training and test images.
- domain assumption The filtered Objaverse subset (about 32k objects) is large and diverse enough for zero-shot generalization to unseen images.
- ad hoc to paper The Wonder3D-style cross-domain attention and the pretrained Stable Diffusion prior remain effective when extended from two to six domains.
- domain assumption The NeuS assumption that ray weight distributions are unimodal holds on the generated sparse-view images.
Cite this review
Pith. "Pith review of GraphicsDreamer: Image to 3D Generation with Physical Consistency." pith.science (2026). https://pith.science/paper/KNEL434K
@misc{pith2026241214214,
author = {Pith},
title = {Pith review of: GraphicsDreamer: Image to 3D Generation with Physical Consistency},
year = {2026},
howpublished = {\url{https://pith.science/paper/KNEL434K}},
note = {Machine review of arXiv:2412.14214}
}
read the original abstract
Recently, the surge of efficient and automated 3D AI-generated content (AIGC) methods has increasingly illuminated the path of transforming human imagination into complex 3D structures. However, the automated generation of 3D content is still significantly lags in industrial application. This gap exists because 3D modeling demands high-quality assets with sharp geometry, exquisite topology, and physically based rendering (PBR), among other criteria. To narrow the disparity between generated results and artists' expectations, we introduce GraphicsDreamer, a method for creating highly usable 3D meshes from single images. To better capture the geometry and material details, we integrate the PBR lighting equation into our cross-domain diffusion model, concurrently predicting multi-view color, normal, depth images, and PBR materials. In the geometry fusion stage, we continue to enforce the PBR constraints, ensuring that the generated 3D objects possess reliable texture details, supporting realistic relighting. Furthermore, our method incorporates topology optimization and fast UV unwrapping capabilities, allowing the 3D products to be seamlessly imported into graphics engines. Extensive experiments demonstrate that our model can produce high quality 3D assets in a reasonable time cost compared to previous methods.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[71]
All-frequency rendering of dynamic, spatially- varying reflectance
Jiaping Wang, Peiran Ren, Minmin Gong, John Snyder, and Baining Guo. All-frequency rendering of dynamic, spatially- varying reflectance. In Proceedings of SIGGRAPH Asia , pages 1–10, 2009. 4, 6, 7
work page 2009
-
[1]
Barrow and J
Harry G. Barrow and J. Martin Tenenbaum. Recovering in- trinsic scene characteristics from images. Computer Vision Systems, pages 3–26, 1978. 3, 4
1978
-
[2]
Intrinsic images in the wild
Sean Bell, Kavita Bala, and Noah Snavely. Intrinsic images in the wild. ACM TOG, 33(4):159:1–159:12, 2014. 4
2014
-
[3]
Blender - a 3d modelling and rendering package
Blender Online Community. Blender - a 3d modelling and rendering package. http://www.blender.org, 2024. 2, 7, 1
2024
-
[4]
Sf3d: Stable fast 3d mesh reconstruction with uv-unwrapping and illumination disentanglement
Mark Boss, Zixuan Huang, Aaryaman Vasishta, and Varun Jampani. Sf3d: Stable fast 3d mesh reconstruction with uv-unwrapping and illumination disentanglement. arXiv preprint, 2024. 8
2024
-
[5]
Physically based shading at disney
Brent Burley and Walt Disney Animation Studios. Physically based shading at disney. In SIGGRAPH, pages 1–7, 2012. 6, 7, 1
2012
-
[6]
Intrinsic image decomposi- tion via ordinal shading
Chris Careaga and Ya ˘gız Aksoy. Intrinsic image decomposi- tion via ordinal shading. ACM TOG, 43(1), 2023. 4
2023
-
[7]
Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation
Rui Chen, Yongwei Chen, Ningxin Jiao, and Kui Jia. Fantasia3d: Disentangling geometry and appearance for high-quality text-to-3d content creation. arXiv preprint arXiv:2303.13873, 2023. 2, 3
arXiv 2023
Show all 90 references
-
[8]
Intrinsicanything: Learning diffusion priors for inverse rendering under unknown illumi- nation
Xi Chen, Sida Peng, Dongchen Yang, Yuan Liu, Bowen Pan, Chengfei Lv, and Xiaowei Zhou. Intrinsicanything: Learning diffusion priors for inverse rendering under unknown illumi- nation. In ECCV, 2024. 4
2024
-
[9]
3dtopia-xl: High-quality 3d pbr asset generation via primitive diffusion
Zhaoxi Chen, Jiaxiang Tang, Yuhao Dong, Ziang Cao, Fangzhou Hong, Yushi Lan, Tengfei Wang, Haozhe Xie, Tong Wu, Shunsuke Saito, Liang Pan, Dahua Lin, and Zi- wei Liu. 3dtopia-xl: High-quality 3d pbr asset generation via primitive diffusion. arXiv preprint arXiv:2409.12957, 2024. 8
2024 arXiv
-
[10]
Sdfusion: Multimodal 3d shape completion, reconstruction, and generation
Yen-Chi Cheng, Hsin-Ying Lee, Sergey Tulyakov, Alexan- der G Schwing, and Liang-Yan Gui. Sdfusion: Multimodal 3d shape completion, reconstruction, and generation. In CVPR, 2023. 2, 3
2023
-
[11]
Cook and Kenneth E
Robert L. Cook and Kenneth E. Torrance. A reflectance model for computer graphics. In SIGGRAPH, pages 307– 316, 1981. 7
1981
-
[12]
Objaverse: A universe of annotated 3d objects
Matt Deitke, Dustin Schwenk, Jordi Salvador, Luca Weihs, Oscar Michel, Eli VanderBilt, Ludwig Schmidt, Kiana Ehsani, Aniruddha Kembhavi, and Ali Farhadi. Objaverse: A universe of annotated 3d objects. In CVPR, 2023. 3, 7, 1, 2
2023
-
[13]
Tore: Token reduction for efficient human mesh re- covery with transformer
Zhiyang Dou, Qingxuan Wu, Cheng Lin, Zeyu Cao, Qiangqiang Wu, Weilin Wan, Taku Komura, and Wenping Wang. Tore: Token reduction for efficient human mesh re- covery with transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 15143– 15155, 2023. 1
2023
-
[14]
Google scanned objects: A high- quality dataset of 3d scanned household items
Laura Downs, Anthony Francis, Nate Koenig, Brandon Kin- man, Ryan Hickman, Krista Reymann, Thomas B McHugh, and Vincent Vanhoucke. Google scanned objects: A high- quality dataset of 3d scanned household items. In ICRA,
-
[15]
Hyperdiffusion: Generating implicit neural fields with weight-space diffusion
Ziya Erkoc ¸, Fangchang Ma, Qi Shan, Matthias Nießner, and Angela Dai. Hyperdiffusion: Generating implicit neural fields with weight-space diffusion. arXiv preprint arXiv:2303.17015, 2023. 2
2023 arXiv
-
[16]
One-2-3-45
Hugging Face. One-2-3-45. https://huggingface. co/spaces/One-2-3-45/One-2-3-45 , 2023. 3
2023
-
[17]
Get3d: A generative model of high quality 3d tex- tured shapes learned from images
Jun Gao, Tianchang Shen, Zian Wang, Wenzheng Chen, Kangxue Yin, Daiqing Li, Or Litany, Zan Gojcic, and Sanja Fidler. Get3d: A generative model of high quality 3d tex- tured shapes learned from images. NeurIPS, 2022. 2, 3
2022
-
[18]
Relightable 3d gaussian: Real-time point cloud relighting with brdf decomposition and ray trac- ing
Jian Gao, Chun Gu, Youtian Lin, Hao Zhu, Xun Cao, Li Zhang, and Yao Yao. Relightable 3d gaussian: Real-time point cloud relighting with brdf decomposition and ray trac- ing. ECCV, 2024. 4
2024
-
[19]
A survey on intrinsic images: Delv- ing deep into lambert and beyond
Elena Garces, Carlos Rodriguez-Pardo, Dan Casas, and Jorge Lopez-Moreno. A survey on intrinsic images: Delv- ing deep into lambert and beyond. IJCV, 130(3):836–868,
-
[20]
K. Guo, P. Lincoln, P. Davidson, J. Busch, X. Yu, M. Whalen, G. Harvey, S. Orts-Escolano, R. Pandey, J. Dour- garian, D. Tang, A. Tkach, A. Kowdle, E. Cooper, M. Dou, S. Fanello, G. Fyffe, C. Rhemann, J. Taylor, P. Debevec, and S. Izadi. The relightables: V olumetric performan...
-
[21]
3dgen: Triplane latent diffusion for textured mesh generation
Anchit Gupta, Wenhan Xiong, Yixin Nie, Ian Jones, and Bar- las O˘guz. 3dgen: Triplane latent diffusion for textured mesh generation. arXiv preprint arXiv:2303.05371, 2023. 2, 3
2023 arXiv
-
[22]
Ray- marching distance fields with cuda
Avelina Hadji-Kyriacou and Ognjen Arandjelovi ´c. Ray- marching distance fields with cuda. Electronics, 10(22),
-
[23]
Hasselgren, N
J. Hasselgren, N. Hofmann, and J. Munkberg. Shape, light, and material decomposition from images using monte carlo rendering and denoising. In NeurIPS, 2022. 4 9
2022
-
[24]
Escap- ing plato’s cave: 3d shape from adversarial rendering
Philipp Henzler, Niloy J Mitra, and Tobias Ritschel. Escap- ing plato’s cave: 3d shape from adversarial rendering. In ICCV, pages 9984–9993, 2019. 2, 3
2019
-
[25]
Jingwei Huang, Yichao Zhou, Matthias Niessner, Jonathan Richard Shewchuk, and Leonidas J. Guibas. QuadriFlow: A Scalable and Robust Method for Quadran- gulation. Computer Graphics Forum, 37, 2018. 7
2018
-
[26]
Zero-shot text-guided object genera- tion with dream fields
Ajay Jain, Ben Mildenhall, Jonathan T Barron, Pieter Abbeel, and Ben Poole. Zero-shot text-guided object genera- tion with dream fields. InProceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 867–876, 2022. 3
2022
-
[27]
Shap-e: Generat- ing conditional 3d implicit functions
Heewoo Jun and Alex Nichol. Shap-e: Generat- ing conditional 3d implicit functions. arXiv preprint arXiv:2305.02463, 2023. 1
2023 arXiv
-
[28]
James T. Kajiya. The rendering equation. In Proceedings of the 13th Annual Conference on Computer Graphics and Interactive Techniques, 1986. 6
1986
-
[29]
In- trinsic image diffusion for indoor single-view material esti- mation
Peter Kocsis, Vincent Sitzmann, and Matthias Niessner. In- trinsic image diffusion for indoor single-view material esti- mation. In CVPR, 2024. 4
2024
-
[30]
Shading annotations in the wild
Balazs Kovacs, Sean Bell, Noah Snavely, and Kavita Bala. Shading annotations in the wild. In CVPR, pages 850–859,
-
[31]
Stanford-orb: a real-world 3d object inverse rendering benchmark
Zhengfei Kuang, Yunzhi Zhang, Hong-Xing Yu, Samir Agar- wala, Elliott Wu, Jiajun Wu, et al. Stanford-orb: a real-world 3d object inverse rendering benchmark. 2023. 4
2023
-
[32]
Land and John J
Edwin H. Land and John J. McCann. Lightness and retinex theory. Journal of the Optical Society of America , 61(1):1– 11, 1971. 4
1971
-
[33]
Inverse ren- dering for complex indoor scenes: Shape, spatially-varying lighting and svbrdf from a single image
Zhengqin Li, Mohammad Shafiei, Ravi Ramamoorthi, Kalyan Sunkavalli, and Manmohan Chandraker. Inverse ren- dering for complex indoor scenes: Shape, spatially-varying lighting and svbrdf from a single image. In CVPR, 2020. 4
2020
-
[34]
Inverse ren- dering for complex indoor scenes: Shape, spatially-varying lighting and svbrdf from a single image
Zhengqin Li, Mohammad Shafiei, Ravi Ramamoorthi, Kalyan Sunkavalli, and Manmohan Chandraker. Inverse ren- dering for complex indoor scenes: Shape, spatially-varying lighting and svbrdf from a single image. In CVPR, pages 2472–2481, 2020. 4
2020
-
[35]
Openrooms: An open framework for photorealistic indoor scene datasets
Zhengqin Li, Ting-Wei Yu, Shen Sang, Sarah Wang, Meng Song, Yuhan Liu, Yu-Ying Yeh, Rui Zhu, Nitesh Gun- davarapu, Jia Shi, Sai Bi, Hong-Xing Yu, Zexiang Xu, Kalyan Sunkavalli, Milo ˇs Ha ˇsan, Ravi Ramamoorthi, and Manmohan Chandraker. Openrooms: An open framework for photore...
2021
-
[36]
Daniel Lichy, Jiaye Wu, Soumyadip Sengupta, and David W. Jacobs. Shape and material capture at home. InCVPR, 2021. 4
2021
-
[37]
Magic3d: High-resolution text-to-3d content creation
Chen-Hsuan Lin, Jun Gao, Luming Tang, Towaki Takikawa, Xiaohui Zeng, Xun Huang, Karsten Kreis, Sanja Fidler, Ming-Yu Liu, and Tsung-Yi Lin. Magic3d: High-resolution text-to-3d content creation. In CVPR, 2023. 2, 3
2023
-
[38]
One-2-3-45++: Fast single image to 3d objects with consistent multi-view generation and 3d diffusion
Minghua Liu, Ruoxi Shi, Linghao Chen, Zhuoyang Zhang, Chao Xu, Xinyue Wei, Hansheng Chen, Chong Zeng, Ji- ayuan Gu, and Hao Su. One-2-3-45++: Fast single image to 3d objects with consistent multi-view generation and 3d diffusion. arXiv preprint arXiv:2311.07885, 2023. 7, 8
2023 arXiv
-
[39]
Zero-1-to-3: Zero-shot one image to 3d object
Ruoshi Liu, Rundi Wu, Basile Van Hoorick, Pavel Tok- makov, Sergey Zakharov, and Carl V ondrick. Zero-1-to-3: Zero-shot one image to 3d object. In ICCV, 2023. 1, 3, 7, 8
2023
-
[40]
Syncdreamer: Gen- erating multiview-consistent images from a single-view im- age
Yuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long, Lingjie Liu, Taku Komura, and Wenping Wang. Syncdreamer: Gen- erating multiview-consistent images from a single-view im- age. arXiv preprint arXiv:2309.03453, 2023. 1, 2, 3, 7, 8
2023 arXiv
-
[41]
Meshdiffu- sion: Score-based generative 3d mesh modeling
Zhen Liu, Yao Feng, Michael J Black, Derek Nowrouzezahrai, Liam Paull, and Weiyang Liu. Meshdiffu- sion: Score-based generative 3d mesh modeling. In ICLR,
-
[42]
Adaptive sur- face normal constraint for depth estimation
Xiaoxiao Long, Cheng Lin, Lingjie Liu, Wei Li, Christian Theobalt, Ruigang Yang, and Wenping Wang. Adaptive sur- face normal constraint for depth estimation. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion (ICCV), pages 12849–12858, 2021. 1
2021
-
[43]
Wonder3d: Sin- gle image to 3d using cross-domain diffusion.arXiv preprint arXiv:2310.15008, 2023
Xiaoxiao Long, Yuan-Chen Guo, Cheng Lin, Yuan Liu, Zhiyang Dou, Lingjie Liu, Yuexin Ma, Song-Hai Zhang, Marc Habermann, Christian Theobalt, et al. Wonder3d: Sin- gle image to 3d using cross-domain diffusion.arXiv preprint arXiv:2310.15008, 2023. 2, 3, 4, 5, 7, 8
-
[44]
Diffusion probabilistic models for 3d point cloud generation
Shitong Luo and Wei Hu. Diffusion probabilistic models for 3d point cloud generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages 2837–2845, 2021. 2, 3
2021
-
[45]
Hemispherical gaussians for accurate light integration
Julia Meder and Beat Br ¨uderlin. Hemispherical gaussians for accurate light integration. In Computer Vision and Graphics, pages 3–15, 2018. 4, 6, 7
2018
-
[46]
Realfusion: 360deg reconstruction of any object from a single image
Luke Melas-Kyriazi, Iro Laina, Christian Rupprecht, and Andrea Vedaldi. Realfusion: 360deg reconstruction of any object from a single image. In CVPR, 2023. 2
2023
-
[47]
Nerf: Representing scenes as neural radiance fields for view syn- thesis
Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. In ECCV, 2020. 3
2020
-
[48]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. In ECCV, 2020. 4, 5, 7
2020
-
[49]
Clip-mesh: Generating textured meshes from text using pretrained image-text models
Nasir Mohammad Khalid, Tianhao Xie, Eugene Belilovsky, and Tiberiu Popa. Clip-mesh: Generating textured meshes from text using pretrained image-text models. InSIGGRAPH Asia Conference Papers, pages 1–8, 2022. 3
2022
-
[50]
Diffrf: Rendering-guided 3d radiance field diffusion
Norman M ¨uller, Yawar Siddiqui, Lorenzo Porzi, Samuel Rota Bulo, Peter Kontschieder, and Matthias Nießner. Diffrf: Rendering-guided 3d radiance field diffusion. In CVPR, 2023. 3
2023
-
[51]
G. Nam, J. H. Lee, D. Gutierrez, and M. H. Kim. Practi- cal svbrdf acquisition of 3d objects with unstructured flash photography. ACM TOG, 37(6):267:1–12. 4
-
[52]
Point-e: A system for generat- ing 3d point clouds from complex prompts
Alex Nichol, Heewoo Jun, Prafulla Dhariwal, Pamela Mishkin, and Mark Chen. Point-e: A system for generat- ing 3d point clouds from complex prompts. arXiv preprint arXiv:2212.08751, 2022. 2, 3
2022 arXiv
-
[53]
Differentiable volumetric rendering: Learn- ing implicit 3d representations without 3d supervision
Michael Niemeyer, Lars Mescheder, Michael Oechsle, and Andreas Geiger. Differentiable volumetric rendering: Learn- ing implicit 3d representations without 3d supervision. In CVPR, pages 3504–3515, 2020. 4 10
2020
-
[54]
A survey of inverse ren- dering problems
Gustavo Patow and Xavier Pueyo. A survey of inverse ren- dering problems. Computer Graphics Forum, 2003. 3, 4
2003
-
[55]
Physically Based Rendering: From Theory to Implementation
Matt Pharr, Wenzel Jakob, and Greg Humphreys. Physically Based Rendering: From Theory to Implementation. Morgan Kaufmann Publishers Inc., 3rd edition, 2016. 4, 6
2016
-
[56]
Dreamfusion: Text-to-3d using 2d diffusion
Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Milden- hall. Dreamfusion: Text-to-3d using 2d diffusion. In ICLR,
-
[57]
Magic123: One image to high-quality 3d object generation using both 2d and 3d diffusion priors.arXiv preprint arXiv:2306.17843,
Guocheng Qian, Jinjie Mai, Abdullah Hamdi, Jian Ren, Aliaksandr Siarohin, Bing Li, Hsin-Ying Lee, Ivan Sko- rokhodov, Peter Wonka, Sergey Tulyakov, et al. Magic123: One image to high-quality 3d object generation using both 2d and 3d diffusion priors.arXiv preprint arXiv:2306.17843,
-
[58]
Richdreamer: A generalizable normal-depth diffusion model for detail richness in text-to- 3d
Lingteng Qiu, Guanying Chen, Xiaodong Gu, Qi Zuo, Mu- tian Xu, Yushuang Wu, Weihao Yuan, Zilong Dong, Liefeng Bo, and Xiaoguang Han. Richdreamer: A generalizable normal-depth diffusion model for detail richness in text-to- 3d. In Proceedings of the IEEE/CVF Conference on Com- ...
-
[59]
Learn- ing transferable visual models from natural language super- vision
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learn- ing transferable visual models from natural language super- vision. In ICML, 2021. 3
2021
-
[60]
High-resolution image syn- thesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj¨orn Ommer. High-resolution image syn- thesis with latent diffusion models. In CVPR, 2022. 3
2022
-
[61]
Photorealistic text-to-image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to-image diffusion models with deep language understanding. NeurIPS, 2022. 3
2022
-
[62]
Relightable gaussian codec avatars
Shunsuke Saito, Gabriel Schwartz, Tomas Simon, Junxuan Li, and Giljoo Nam. Relightable gaussian codec avatars. In CVPR, 2024. 4
2024
-
[63]
Zero123++: a single image to consistent multi-view dif- fusion base model
Ruoxi Shi, Hansheng Chen, Zhuoyang Zhang, Minghua Liu, Chao Xu, Xinyue Wei, Linghao Chen, Chong Zeng, and Hao Su. Zero123++: a single image to consistent multi-view dif- fusion base model. arXiv preprint arXiv:2310.15110, 2023. 3
-
[64]
Mvdream: Multi-view diffusion for 3d gen- eration
Yichun Shi, Peng Wang, Jianglong Ye, Mai Long, Kejie Li, and Xiao Yang. Mvdream: Multi-view diffusion for 3d gen- eration. arXiv preprint arXiv:2308.16512, 2023. 2, 3, 5
2023 arXiv
-
[65]
Sloan, Jan Kautz, and James Snyder
Peter-P. Sloan, Jan Kautz, and James Snyder. Precomputed radiance transfer for real-time rendering in dynamic, low- frequency lighting environments. InSIGGRAPH, pages 527– 536, 2002. 7
2002
-
[66]
Viewset diffusion:(0-) image-conditioned 3d gener- ative models from 2d data.arXiv preprint arXiv:2306.07881,
Stanislaw Szymanowicz, Christian Rupprecht, and Andrea Vedaldi. Viewset diffusion:(0-) image-conditioned 3d gener- ative models from 2d data.arXiv preprint arXiv:2306.07881,
-
[67]
Make-it-3d: High-fidelity 3d creation from a single image with diffusion prior, 2023
Junshu Tang, Tengfei Wang, Bo Zhang, Ting Zhang, Ran Yi, Lizhuang Ma, and Dong Chen. Make-it-3d: High-fidelity 3d creation from a single image with diffusion prior, 2023. 3
2023
-
[68]
The magic of the z-buffer: A sur- vey
Theoharis Theoharis, Georgios Papaioannou, and Evaggelia- Aggeliki Karabassi. The magic of the z-buffer: A sur- vey. In The 9-th International Conference in Central Europe on Computer Graphics, Visualization and Computer Vision, pages 379–386, 2001. 6, 1
2001
-
[69]
Consistent view synthesis with pose-guided diffusion models
Hung-Yu Tseng, Qinbo Li, Changil Kim, Suhib Alsisan, Jia- Bin Huang, and Johannes Kopf. Consistent view synthesis with pose-guided diffusion models. In CVPR, 2023. 3
2023
-
[70]
Score jacobian chaining: Lifting pretrained 2d diffusion models for 3d generation
Haochen Wang, Xiaodan Du, Jiahao Li, Raymond A Yeh, and Greg Shakhnarovich. Score jacobian chaining: Lifting pretrained 2d diffusion models for 3d generation. In CVPR,
-
[72]
Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction
Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. Neus: Learning neural implicit surfaces by volume rendering for multi-view reconstruction. NeurIPS, 2021. 4, 5
2021
-
[73]
Rodin: A generative model for sculpting 3d digital avatars using diffusion
Tengfei Wang, Bo Zhang, Ting Zhang, Shuyang Gu, Jianmin Bao, Tadas Baltrusaitis, Jingjing Shen, Dong Chen, Fang Wen, Qifeng Chen, et al. Rodin: A generative model for sculpting 3d digital avatars using diffusion. In CVPR, 2023. 2, 3
2023
-
[74]
Image quality assessment: from error visibility to structural similarity
Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. TIP, 2004. 8
2004
-
[75]
Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distilla- tion
Zhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao, Chongxuan Li, Hang Su, and Jun Zhu. Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distilla- tion. arXiv preprint arXiv:2305.16213, 2023. 2
2023 arXiv
-
[76]
Consistent123: Improve consistency for one image to 3d object synthesis
Haohan Weng, Tianyu Yang, Jianan Wang, Yu Li, Tong Zhang, CL Chen, and Lei Zhang. Consistent123: Improve consistency for one image to 3d object synthesis. arXiv preprint arXiv:2310.08092, 2023. 3
2023 arXiv
-
[77]
Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling
Jiajun Wu, Chengkai Zhang, Tianfan Xue, Bill Freeman, and Josh Tenenbaum. Learning a probabilistic latent space of object shapes via 3d generative-adversarial modeling. 29,
-
[78]
Dream3d: Zero-shot text-to-3d synthesis using 3d shape prior and text-to-image diffusion models
Jiale Xu, Xintao Wang, Weihao Cheng, Yan-Pei Cao, Ying Shan, Xiaohu Qie, and Shenghua Gao. Dream3d: Zero-shot text-to-3d synthesis using 3d shape prior and text-to-image diffusion models. In CVPR, 2023. 3
2023
-
[79]
Instantmesh: Efficient 3d mesh generation from a single image with sparse-view large reconstruction models
Jiale Xu, Weihao Cheng, Yiming Gao, Xintao Wang, Shenghua Gao, and Ying Shan. Instantmesh: Efficient 3d mesh generation from a single image with sparse-view large reconstruction models. arXiv preprint arXiv:2404.07191 ,
-
[80]
Multiview neu- ral surface reconstruction by disentangling geometry and ap- pearance
Lior Yariv, Yoni Kasten, Dror Moran, Meirav Galun, Matan Atzmon, Basri Ronen, and Yaron Lipman. Multiview neu- ral surface reconstruction by disentangling geometry and ap- pearance. NeurIPS, 33, 2020. 4, 6
2020
-
[81]
Consistent-1-to-3: Consistent image to 3d view syn- thesis via geometry-aware diffusion models
Jianglong Ye, Peng Wang, Kejie Li, Yichun Shi, and Heng Wang. Consistent-1-to-3: Consistent image to 3d view syn- thesis via geometry-aware diffusion models. arXiv preprint arXiv:2310.03020, 2023. 3 11
2023 arXiv
-
[82]
Lion: Latent point diffusion models for 3d shape generation
Xiaohui Zeng, Arash Vahdat, Francis Williams, Zan Gojcic, Or Litany, Sanja Fidler, and Karsten Kreis. Lion: Latent point diffusion models for 3d shape generation. In NeurIPS,
-
[83]
3dshape2vecset: A 3d shape representation for neu- ral fields and generative diffusion models
Biao Zhang, Jiapeng Tang, Matthias Niessner, and Peter Wonka. 3dshape2vecset: A 3d shape representation for neu- ral fields and generative diffusion models. In SIGGRAPH,
-
[84]
Zhang, Y
J. Zhang, Y . Yao, S. Li, J. Liu, T. Fang, D. McKinnon, Y . Tsin, and L. Quan. Neilf++: Inter-reflectable light fields for geometry and material estimation. In ICCV, 2023. 4
2023
-
[85]
PhySG: Inverse rendering with spherical gaussians for physics-based material editing and relighting
Kai Zhang, Fujun Luan, Qianqian Wang, Kavita Bala, and Noah Snavely. PhySG: Inverse rendering with spherical gaussians for physics-based material editing and relighting. In CVPR, 2021. 4, 6
2021
-
[86]
Efficientdreamer: High-fidelity and robust 3d cre- ation via orthogonal-view diffusion prior
Minda Zhao, Chaoyi Zhao, Xinyue Liang, Lincheng Li, Zeng Zhao, Zhipeng Hu, Changjie Fan, and Xin Yu. Efficientdreamer: High-fidelity and robust 3d cre- ation via orthogonal-view diffusion prior. arXiv preprint arXiv:2308.13223, 2023. 3
2023 arXiv
-
[87]
Glosh: Global- local spherical harmonics for intrinsic image decomposition
Hao Zhou, Xiang Yu, and David Jacobs. Glosh: Global- local spherical harmonics for intrinsic image decomposition. In ICCV, 2019. 4
2019
-
[88]
3d shape generation and completion through point-voxel diffusion
Linqi Zhou, Yilun Du, and Jiajun Wu. 3d shape generation and completion through point-voxel diffusion. In Proceed- ings of the IEEE/CVF International Conference on Com- puter Vision, pages 5826–5835, 2021. 2
2021
-
[89]
Learning-based inverse rendering of complex indoor scenes with differentiable monte carlo raytracing
Jingsen Zhu, Fujun Luan, Yuchi Huo, Zihao Lin, Zhihua Zhong, Dianbing Xi, Rui Wang, Hujun Bao, Jiaxiang Zheng, and Rui Tang. Learning-based inverse rendering of complex indoor scenes with differentiable monte carlo raytracing. In SIGGRAPH Asia, 2022. 4
2022
-
[90]
Irisformer: Dense vision transform- ers for single-image inverse rendering in indoor scenes
Rui Zhu, Zhengqin Li, Janarbek Matai, Fatih Porikli, and Manmohan Chandraker. Irisformer: Dense vision transform- ers for single-image inverse rendering in indoor scenes. In CVPR, 2022. 4 12 GraphicsDreamer: Image to 3D Generation with Physical Consistency Supplementary Materi...
2022
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.