Pith. sign in

REVIEW 4 major objections 5 minor 1 cited by

TSGS: Improving Gaussian Splatting for Transparent Surface Reconstruction via Normal and De-lighting Priors

T0 review · 4 major / 5 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Transparent surfaces get accurate 3D geometry by learning geometry separately from appearance and extracting first-surface depth with a sliding-window rule.

desk verdict Solid engineering paper with a useful synthetic dataset, but the headline 3 mm error bound and the central depth-extraction mechanism are both weaker than claimed. read the letter →

arxiv 2504.12799 v2 pith:726KO6JV submitted 2025-04-17 cs.CV

classification cs.CV
keywords 3DGaussiansplattingtransparentsurfacereconstructionfirst-surfacedepthextractionaccumulatedtransmittancealpha-blendingnormalandde-lightingpriorsanisotropicsphericalGaussiansTransLabdataset
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that the transparency-depth dilemma in 3D Gaussian Splatting—where alpha-blending weights tuned for realistic appearance push depth estimates of transparent objects behind the true surface—can be resolved without abandoning the 3DGS framework. It introduces TSGS, which learns geometry from de-lighted images and normal priors in a first stage, then refines appearance with anisotropic specular modeling while freezing opacity in a second stage. For depth, it replaces full alpha-blended depth with a sliding-window maximum-weight search over accumulated transmittance, extracting the first surface along each ray. On the new TransLab dataset of laboratory glassware, TSGS reports a 37.3% lower chamfer distance and an 8.0% higher F1 score than the strongest baseline, with a 0.41 dB PSNR gain, showing geometry and photorealistic transparency can be recovered together. This matters because robotic lab manipulation needs millimeter-accurate positions of transparent vessels.

What carries the argument

The load-bearing mechanism is the maximum-weight sliding-window first-surface extractor. For each pixel ray, after sorting intersecting Gaussians by depth, it computes accumulated transmittance $T_i$ and blending weight $T_i\alpha_i$, keeps only the segment with $T_i$ between $T_{\text{end}}$ and $T_{\text{start}}$, slides a window of fixed size $\delta_t$ (3 mm) over candidate Gaussians, and selects $W^* = \arg\max_j \sum_{i\in W_j} T_i\alpha_i$. The first-surface depth is $D_{\text{first}}(\mathbf{u}) = (\sum_{i\in W^*} T_i\alpha_i \hat{d}_i(\mathbf{u})) / (\sum_{i\in W^*} T_i\alpha_i)$, using per-Gaussian plane depths rather than center depths. This is what isolates the first surface from background transmission and averages out floaters; the two-stage training, geometry on de-lighted images with normal priors followed by appearance refinement with opacity frozen and anisotropic spherical Gaussians, supplies the opacity field that makes the window meaningful.

What would settle it

Construct a controlled scene with a flat transparent slab of known position, varying surface opacity and background texture contrast, and compare TSGS's extracted first-surface depth against the known slab distance: if depth error exceeds the 3 mm window size whenever the slab's alpha contribution is low, the window-containment assumption is violated. A simpler version is a filled beaker with an immersed object, where multi-layer transparency is present and the single-layer assumption predicts a biased first-surface depth.

Watch

Extended reading notes

Core claim

On its own terms, the discovery is that the opacity field learned purely for appearance still contains a reliable first-surface signal: the accumulated transmittance $T_i$ drops where Gaussians representing the first surface contribute, and the sum of $T_i\alpha_i$ spikes there. TSGS locates this by restricting attention to the ray segment where $T_i$ lies between thresholds, sliding a fixed-size window along the sorted Gaussians, and picking the window with the largest total $T_i\alpha_i$ weight. Depth is then the weighted average of per-Gaussian plane depths $\hat{d}_i(\mathbf{u}) = d_i/(\mathbf{n}_i \cdot \mathbf{v}_\mathbf{u})$ inside that window. Weighting by $T_i\alpha_i$ within the window suppresses floaters that corrupt nearest-depth methods, while restricting to the window excludes background Gaussians seen through transparency that corrupt standard $\alpha$-blended depth. The window size doubles as an error bound, set to 3 mm in the paper, so the reported accuracy figures rest on the window's ability to contain the true surface.

Load-bearing premise

The load-bearing premise is that the ray segment where the sliding-window sum of accumulated-transmittance-weighted alpha values is largest actually contains the true first surface; if a weak first-surface signal relative to transmitted background or poor threshold choices breaks that containment, the extracted depth is biased.

Editorial extensions

If this is right

  • Transparent object geometry becomes extractable from standard appearance-optimized Gaussians at inference time, without a separate depth network or ray-tracing pass.
  • With the 3 mm window bound, the resulting depth error is controlled enough for millimeter-precision robotic manipulation of laboratory glassware, on the paper's reported results.
  • Freezing opacity after geometry learning prevents appearance refinement from eroding shape accuracy, so visual fidelity and geometry are not forced to trade off within a single optimization.
  • On TransLab, the full pipeline yields CD 1.85 and F1 0.95, compared with 2.95 and 0.88 for the strongest baseline, while keeping 105 FPS rendering and a training time near 0.8 hours.
  • The same geometry stage transfers to opaque objects: on DTU, TSGS reports a mean chamfer distance of 0.51, competitive with or better than the compared surface reconstruction methods.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editor's inference: The sliding-window rule could be applied as a drop-in depth estimator to other volume-rendered transparent-object reconstructions, since it only needs accumulated-transmittance weights that those renderers already compute.
  • Editor's inference: The stated 3 mm error bound is conditional on the maximum-weight window containing the true first surface; a stress test crossing single-layer versus nested transparent objects, such as liquid inside a beaker, would show how often that containment fails.
  • Editor's inference: Because the ablation shows the normal prior is the largest single geometry contributor, the method's accuracy is likely sensitive to the quality of that diffusion-derived prior; testing on scenes with heavy occlusion or unusual glassware would quantify this sensitivity.
  • Editor's inference: A direct extension would be multi-window extraction to recover the second surface, turning TSGS from a single-shell reconstruction into a layered refractive reconstruction; the paper's own limitation section points at this direction.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. TSGS addresses transparent surface reconstruction with 3D Gaussian Splatting by separating geometry learning from appearance refinement. Stage 1 optimizes geometry with de-lighted hybrid images, normal priors from StableNormal, a transparency-attribute loss, and flatten regularization; Stage 2 fixes opacity and adds anisotropic spherical Gaussian (ASG) appearance modeling supervised by the original images. At inference, a first-surface depth is extracted by restricting the ray to an accumulated-transmittance interval, selecting a sliding window that maximizes the sum of T_i alpha_i, and averaging the Gaussian plane depths in that window. The method is evaluated on a new synthetic TransLab dataset of eight laboratory scenes with ground-truth meshes, on DTU, and on a limited subset of ClearPose; it reports a 37.3% chamfer-distance reduction and an 8.0% F1 improvement over PGSR on TransLab, plus a 0.41 dB PSNR gain.

Significance. The paper targets a real problem, the transparency-depth dilemma in 3DGS, and its proposed pipeline is coherent and reasonably motivated. The authors release code and a dataset, and they include a comparison against baselines augmented with the same normal priors (Table 8), which strengthens the attribution of the gains. The two-stage strategy with frozen opacity is a sensible way to prevent appearance optimization from corrupting geometry, and the use of masked normal and de-light priors is carefully designed. However, the headline geometric improvement is measured on a synthetic dataset, the signature first-surface extraction contributes only a small amount in the ablation (Table 4), and the claimed 3 mm error bound in Appendix A.3 is not actually proven. These issues temper the contribution, although they are addressable within the manuscript's scope.

major comments (4)
  1. [Appendix A.3, Eq. (18)] The statement that the depth error is 'naturally bounded' by the window size, at most delta_t = 3 mm, is not established. The bound is valid only if the true first surface lies inside the selected window W*, but the maximum-weight criterion is a heuristic and is not shown to guarantee this; for a highly transparent foreground with small alpha_i relative to a more opaque background, the maximum-weight window can sit on the transmitted background. Please remove or replace this bound with an empirical error distribution against the ground-truth depth maps available in TransLab, and state explicitly that the assumption of window containment is what the method relies on.
  2. [Table 4, Section 3.4] The ablation shows that replacing the first-surface extraction with unbiased depth changes chamfer distance only from 1.85 to 1.89 and leaves the F1 score unchanged, a small effect relative to the headline gain over PGSR (1.85 vs. 2.95). This does not support the description of the maximum-weight window extraction as the load-bearing mechanism behind the reported improvement. Please provide a targeted analysis, such as per-transparent-mask errors or a breakdown on the scenes with the strongest transparency effects, and adjust the wording in Sections 1 and 3.4 to reflect the measured contribution.
  3. [Section 4.1, Section 3.4] Several core hyperparameters are not reported anywhere: T_start, T_end, theta_T, theta_n, and the sliding-window size is given only in the appendix (3 mm). Without these values, without a sensitivity analysis, and without variance across random seeds or initialization, the TransLab results cannot be fully reproduced or assessed for stability. Please report the settings and include error bars or seed variance for the main quantitative tables.
  4. [Appendix A.4] The paper honestly acknowledges the single-layer transparency assumption, but this limitation should be connected to the benchmark claims. TransLab is a synthetic dataset and, as described, does not include the multi-layer refractive cases (e.g., liquid inside a beaker) that motivate the lab-manipulation application, so the 37.3% chamfer-distance improvement may not transfer to those cases. Please scope the claims accordingly and, if feasible, add a TransLab variant with such multi-layer scenes.
minor comments (5)
  1. [Section 1] There is a typo in the Introduction: 'chamber distance' should be 'chamfer distance'; the same misspelling appears in the abstract-related text and should be checked throughout.
  2. [Section 4.1 / Appendix A.1] The abstract and Section 4.1 describe TransLab as close to 'realistic conditions,' while Appendix A.1 states that it is a synthetic benchmark rendered with Blender's PBR engine; the main text should state clearly that the dataset is synthetic.
  3. [Section 3.4] In the phrase 'T_i guadually decreases,' 'guadually' should be 'gradually'; please also proofread the equations for formatting issues such as the garbled subscripts in Eq. (5) and the surrounding text.
  4. [Table 6 / Section 4.3] The ClearPose evaluation uses only one scene per set, subsamples one frame every 100, and reports a unidirectional chamfer distance; this is a weak basis for the claim of real-world robustness and should be described as a preliminary result rather than a definitive validation.
  5. [Appendix A.6] The comparison with NU-NeRF on TransLab (Table 9) reports that NU-NeRF collapses to a spherical shape; please provide the qualitative evidence in the appendix or state the convergence criterion used, since a single failure mode does not by itself establish method superiority.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper's claims rest on external priors and independent ground-truth evaluation, not on self-referential fits.

full rationale

The derivation chain is self-contained against independent benchmarks. Stage 1 geometry uses normal and de-light priors from StableNormal/StableDelight [75] and a transparency mask from Grounded-SAM; these are external models, not outputs of TSGS. Stage 2 freezes opacity and adds ASG appearance following Spec-Gaussian [73] and PGSR [9], which are prior external works. The first-surface depth extraction (Eq. 18) is presented as a heuristic ('We operate under the hypothesis...'), not as a fitted prediction; its accuracy is measured against ground-truth meshes on TransLab, DTU, and ClearPose, so the headline 37.3% CD reduction is not a renamed training objective. The paper's own limitation (Appendix A.4) concedes the single-layer-transparency assumption, and the 3 mm error bound in Appendix A.3 is logically incomplete because it presupposes that the true first surface lies in the selected window W*; however, this is a correctness gap in an auxiliary error analysis, not a circular reduction of an output to an input. The only self-citation is reference [84] (same authors) in a related-work list of diffusion priors; it is not load-bearing for any claimed result. No fitted parameter is relabeled as a prediction, and no asserted uniqueness theorem or author-imported constraint forces the claimed outcome.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The method's headroom over baselines depends on several hand-set thresholds that are not reported and on external diffusion priors. The least justified assumption is that the maximum-weight window contains the first surface; the paper's own limitation section restricts this to single-layer transparency. No new physical entities are introduced.

free parameters (5)
  • Sliding window size delta_t = 3 mm
    Set by hand in Appendix A.3; the final depth is a weighted average inside the window, and the claimed theoretical error bound equals this value.
  • Transmittance search thresholds T_start and T_end = Not reported
    These define the ray segment used to search for the first surface in Section 3.4; no numerical values are given anywhere in the paper.
  • Transparency attribute threshold theta_T = Not reported
    Used in Section 3.2 to decide whether a Gaussian is transparent; the value is never specified.
  • Normal prior mask threshold theta_n = Not reported
    Used in Equation 8 to mask unreliable normal priors; the value is never specified.
  • Loss weights lambda_t, lambda_n, lambda_f, lambda_r = 0.1, 0.1, 100, 0.2
    Reported in Sections 3.2 and 3.3; chosen by hand and likely tuned on the TransLab benchmark itself.
assumptions (5)
  • domain assumption Known camera poses are available from SfM or SLAM.
    Stated in Section 2; all multi-view reconstruction and TSDF fusion rely on this.
  • domain assumption StableNormal and StableDelight provide sufficiently accurate priors after masking.
    Section 3.2 supervises geometry with these external diffusion-model outputs; if the priors are wrong in unmasked regions, the geometry is wrong.
  • domain assumption PGSR flattened Gaussians and plane depth give usable per-Gaussian surface geometry.
    Sections 3.1 and 3.4 adopt PGSR's flatten regularization and Gaussian plane depth; the first-surface depth averages these plane depths.
  • ad hoc to paper The maximum-weight sliding window contains the true first surface.
    Core heuristic of Section 3.4; the paper provides no proof and its own limitation section says multi-layer transparency is not handled.
  • domain assumption TransLab synthetic renderings are representative of real transparent surfaces.
    The main quantitative claims are evaluated only on this synthetic Blender dataset; real-world ClearPose results are limited to four subsampled scenes in the appendix.

how reviews work

0 comments
Cite this review

Pith. "Pith review of TSGS: Improving Gaussian Splatting for Transparent Surface Reconstruction via Normal and De-lighting Priors." pith.science (2026). https://pith.science/paper/726KO6JV

@misc{pith2026250412799,
  author       = {Pith},
  title        = {Pith review of: TSGS: Improving Gaussian Splatting for Transparent Surface Reconstruction via Normal and De-lighting Priors},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/726KO6JV}},
  note         = {Machine review of arXiv:2504.12799}
}
abstract

Reconstructing transparent surfaces is essential for tasks such as robotic manipulation in labs, yet it poses a significant challenge for 3D reconstruction techniques like 3D Gaussian Splatting (3DGS). These methods often encounter a transparency-depth dilemma, where the pursuit of photorealistic rendering through standard $\alpha$-blending undermines geometric precision, resulting in considerable depth estimation errors for transparent materials. To address this issue, we introduce Transparent Surface Gaussian Splatting (TSGS), a new framework that separates geometry learning from appearance refinement. In the geometry learning stage, TSGS focuses on geometry by using specular-suppressed inputs to accurately represent surfaces. In the second stage, TSGS improves visual fidelity through anisotropic specular modeling, crucially maintaining the established opacity to ensure geometric accuracy. To enhance depth inference, TSGS employs a first-surface depth extraction method. This technique uses a sliding window over $\alpha$-blending weights to pinpoint the most likely surface location and calculates a robust weighted average depth. To evaluate the transparent surface reconstruction task under realistic conditions, we collect a TransLab dataset that includes complex transparent laboratory glassware. Extensive experiments on TransLab show that TSGS achieves accurate geometric reconstruction and realistic rendering of transparent objects simultaneously within the efficient 3DGS framework. Specifically, TSGS significantly surpasses current leading methods, achieving a 37.3% reduction in chamfer distance and an 8.0% improvement in F1 score compared to the top baseline. The code and dataset are available at https://longxiang-ai.github.io/TSGS/.

Figures

Figures reproduced from arXiv: 2504.12799 by the authors.

Figure 1
Figure 1. (a) We introduce TransLab, a novel dataset specifically designed for evaluating transparent object reconstruction. [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. The transparency-depth dilemma in Gaussian Splat [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Pipeline of TSGS. (a) The two-stage training process. In Stage 1, 3D Gaussians are optimized using geometric priors [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (10 more)
Figure 4
Figure 4. Figure 4: Normal prior mask. Normal priors can be inaccurate [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Comparison of normal maps derived from differ [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Comparison with other methods on the TransLab [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Additional qualitative reconstruction results on various scenes from the TransLab dataset. Our method (TSGS) [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]
Figure 8
Figure 8. Figure 8: Additional qualitative reconstruction results on various scenes from the TransLab dataset. Our method (TSGS) [PITH_FULL_IMAGE:figures/full_fig_p015_8.png]
Figure 9
Figure 9. Figure 9: Additional qualitative reconstruction results on scenes from the DTU dataset. This demonstrates the effectiveness of [PITH_FULL_IMAGE:figures/full_fig_p016_9.png]
Figure 10
Figure 10. Figure 10: Additional qualitative reconstruction results on scenes from the DTU dataset (Continued). This demonstrates the [PITH_FULL_IMAGE:figures/full_fig_p017_10.png]
Figure 11
Figure 11. Figure 11: Additional qualitative reconstruction results on scenes from the DTU dataset (Continued). This demonstrates the [PITH_FULL_IMAGE:figures/full_fig_p018_11.png]
Figure 12
Figure 12. Figure 12: Additional qualitative reconstruction results on scenes from the DTU dataset (Continued). This demonstrates the [PITH_FULL_IMAGE:figures/full_fig_p019_12.png]
Figure 13
Figure 13. Figure 13: Additional qualitative reconstruction results on scenes from the DTU dataset (Continued). This demonstrates the [PITH_FULL_IMAGE:figures/full_fig_p020_13.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Mixed-Primitive-based Gaussian Splatting Method for Surface Reconstruction

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MP-GS combines Gaussian ellipses, line segments, and triangles as splatting primitives and reports state-of-the-art Chamfer distance on DTU and F1 on Tanks and Temples.

Reference graph

Works this paper leans on

92 extracted references · 41 canonical work pages · cited by 1 Pith paper

  1. [1]

    Connelly Barnes, Eli Shechtman, Adam Finkelstein, and Dan B Goldman. 2009. PatchMatch: A randomized correspondence algorithm for structural image edit- ing. ACM Trans. Graph. 28, 3 (2009), 24

  2. [2]

    Jonathan T Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P Srinivasan. 2021. Mip-nerf: A multiscale repre- sentation for anti-aliasing neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 5855–5864

  3. [3]

    Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P

    Jonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P. Srinivasan. 2021. Mip-NeRF: A Multiscale Repre- sentation for Anti-Aliasing Neural Radiance Fields. ICCV (2021)

  4. [4]

    Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Peter Hedman. 2022. Mip-nerf 360: Unbounded anti-aliased neural radiance fields. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 5470–5479

  5. [5]

    Jonathan T Barron, Ben Mildenhall, Dor Verbin, Pratul P Srinivasan, and Pe- ter Hedman. 2023. Zip-nerf: Anti-aliased grid-based neural radiance fields. In Proceedings of the IEEE/CVF International Conference on Computer Vision . 19697– 19705

  6. [6]

    Jean-Daniel Boissonnat. 1984. Geometric structures for three-dimensional shape representation. ACM Trans. Graph. 3, 4 (Oct. 1984), 266–286. doi:10.1145/357346. 357349

  7. [7]

    Carlos Campos, Richard Elvira, Juan J Gómez Rodríguez, José MM Montiel, and Juan D Tardós. 2021. Orb-slam3: An accurate open-source library for visual, visual–inertial, and multimap slam. IEEE Transactions on Robotics 37, 6 (2021), 1874–1890

  8. [8]

    Frédéric Cazals and Joachim Giesen. 2006. Delaunay triangulation based sur- face reconstruction. In Effective computational geometry for curves and surfaces . Springer, 231–276

Show all 92 references
  1. [9]

    Danpeng Chen, Hai Li, Weicai Ye, Yifan Wang, Weijian Xie, Shangjin Zhai, Nan Wang, Haomin Liu, Hujun Bao, and Guofeng Zhang. 2024. PGSR: Planar- based Gaussian Splatting for Efficient and High-Fidelity Surface Reconstruction. IEEE Transactions on Visualization and Computer Gra...

  2. [10]

    Danpeng Chen, Nan Wang, Runsen Xu, Weijian Xie, Hujun Bao, and Guofeng Zhang. 2021. Rnin-vio: Robust neural inertial navigation aided visual-inertial odometry in challenging scenes. In 2021 IEEE International Symposium on Mixed and Augmented Reality (ISMAR). IEEE, 275–283

  3. [11]

    Danpeng Chen, Shuai Wang, Weijian Xie, Shangjin Zhai, Nan Wang, Hujun Bao, and Guofeng Zhang. 2022. Vip-slam: An efficient tightly-coupled rgb-d visual inertial planar slam. In 2022 International Conference on Robotics and Automation (ICRA). IEEE, 5615–5621

  4. [12]

    Xiaotong Chen, Huijie Zhang, Zeren Yu, Anthony Opipari, and Odest Chad- wicke Jenkins. 2022. ClearPose: Large-scale Transparent Object Dataset and;Benchmark. In Computer Vision - ECCV 2022: 17th European Conference, Tel A viv, Israel, October 23-27, 2022, Proceedings, Part VII...

  5. [13]

    Yiwen Chen, Tong He, Di Huang, Weicai Ye, Sijin Chen, Jiaxiang Tang, Xin Chen, Zhongang Cai, Lei Yang, Gang Yu, et al . 2024. MeshAnything: Artist- Created Mesh Generation with Autoregressive Transformers. arXiv preprint arXiv:2406.10163 (2024)

  6. [14]

    Kai Cheng, Xiaoxiao Long, Kaizhi Yang, Yao Yao, Wei Yin, Yuexin Ma, Wenping Wang, and Xuejin Chen. 2024. GaussianPro: 3D Gaussian splatting with progres- sive propagation. In Proceedings of the 41st International Conference on Machine Learning (Vienna, Austria) (ICML’24). JMLR...

  7. [15]

    Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese

    Christopher B. Choy, Danfei Xu, JunYoung Gwak, Kevin Chen, and Silvio Savarese

  8. [16]

    Brian Curless and Marc Levoy. 1996. A volumetric method for building complex models from range images. In SIGGRAPH. 303–312

  9. [17]

    Nianchen Deng, Zhenyi He, Jiannan Ye, Budmonde Duinkharjav, Praneeth Chakravarthula, Xubo Yang, and Qi Sun. 2022. Fov-nerf: Foveated neural radi- ance fields for virtual reality. IEEE Transactions on Visualization and Computer Graphics 28, 11 (2022), 3854–3864

  10. [18]

    Weijian Deng, Dylan Campbell, Chunyi Sun, Shubham Kanitkar, Matthew Shaf- fer, and Stephen Gould. 2024. Differentiable Neural Surface Refinement for Transparent Objects. In CVPR

  11. [19]

    Herbert Edelsbrunner and Ernst P. Mücke. 1994. Three-dimensional alpha shapes. ACM Trans. Graph. 13, 1 (Jan. 1994), 43–72. doi:10.1145/174462.156635

  12. [20]

    Haoqiang Fan, Hao Su, and Leonidas J. Guibas. 2017. A Point Set Generation Network for 3D Object Reconstruction from a Single Image. In IEEE Conference on Computer Vision and Pattern Recognition . 2463–2471

  13. [21]

    Yasutaka Furukawa, Carlos Hernández, et al. 2015. Multi-view stereo: A tutorial. Foundations and trends® in Computer Graphics and Vision 9, 1-2 (2015), 1–148

  14. [22]

    Fangzhou Gao, Lianghao Zhang, Li Wang, Jiamin Cheng, and Jiawan Zhang

  15. [23]

    Peng Gao, Le Zhuo, Dongyang Liu, Ruoyi Du, Xu Luo, Longtian Qiu, Yuhang Zhang, Chen Lin, Rongjie Huang, Shijie Geng, Renrui Zhang, Junlin Xi, Wenqi Shao, Zhengkai Jiang, Tianshuo Yang, Weicai Ye, He Tong, Jingwen He, Yu Qiao, and Hongsheng Li. 2024. Lumina-T2X: Transforming Te...

  16. [24]

    Antoine Guédon and Vincent Lepetit. 2024. SuGaR: Surface-Aligned Gaussian Splatting for Efficient 3D Mesh Reconstruction and High-Quality Mesh Rendering. CVPR (2024)

  17. [25]

    Binbin Huang, Zehao Yu, Anpei Chen, Andreas Geiger, and Shenghua Gao. 2024. 2D Gaussian Splatting for Geometrically Accurate Radiance Fields. In SIGGRAPH 2024 Conference Papers . Association for Computing Machinery. doi:10.1145/ 3641519.3657428

  18. [26]

    Chenxi Huang, Yuenan Hou, Weicai Ye, Di Huang, Xiaoshui Huang, Binbin Lin, Deng Cai, and Wanli Ouyang. 2024. NeRF-Det++: Incorporating Semantic Cues and Perspective-aware Depth Supervision for Indoor Multi-View 3D Detection. arXiv preprint arXiv:2402.14464 (2024)

  19. [27]

    Rasmus Jensen, Anders Dahl, George Vogiatzis, Engin Tola, and Henrik Aanæs

  20. [28]

    Qing Jiang, Feng Li, Zhaoyang Zeng, Tianhe Ren, Shilong Liu, and Lei Zhang

  21. [29]

    Yingwenqi Jiang, Jiadong Tu, Yuan Liu, Xifeng Gao, Xiaoxiao Long, Wenping Wang, and Yuexin Ma. 2024. GaussianShader: 3D Gaussian Splatting with Shading Functions for Reflective Surfaces. In2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) . 5322–5332. ...

  22. [30]

    Michael Kazhdan, Matthew Bolitho, and Hugues Hoppe. 2006. Poisson surface reconstruction. In Proceedings of the Fourth Eurographics Symposium on Geometry Processing (Cagliari, Sardinia, Italy) (SGP ’06). Eurographics Association, Goslar, DEU, 61–70

  23. [31]

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis

  24. [32]

    Jeongyun Kim, Jeongho Noh, Dong-Guw Lee, and Ayoung Kim. 2025. TranSplat: Surface Embedding-guided 3D Gaussian Splatting for Transparent Object Manip- ulation. arXiv:2502.07840 [cs.CV] https://arxiv.org/abs/2502.07840

  25. [33]

    Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C. Berg, Wan-Yen Lo, Piotr Dollár, and Ross Girshick. 2023. Segment Anything. arXiv:2304.02643 (2023)

  26. [34]

    Arno Knapitsch, Jaesik Park, Qian-Yi Zhou, and Vladlen Koltun. 2017. Tanks and temples: Benchmarking large-scale scene reconstruction. ACM Transactions on Graphics (ToG) 36, 4 (2017), 1–13

  27. [35]

    Kiriakos N Kutulakos and Steven M Seitz. 2000. A theory of shape by space carving. International journal of computer vision 38 (2000), 199–218

  28. [36]

    In ACM Transactions on Graphics, Vol

    3D Gaussian Splatting for Real-Time Radiance Field Rendering. In ACM Transactions on Graphics, Vol. 42

  29. [37]

    Maxime Lhuillier and Long Quan. 2005. A quasi-dense approach to surface reconstruction from uncalibrated images. IEEE transactions on pattern analysis and machine intelligence 27, 3 (2005), 418–433

  30. [38]

    Congcong Li, Jin Wang, Xiaomeng Wang, Xingchen Zhou, Wei Wu, Yuzhi Zhang, and Tongyi Cao. 2025. Car-GS: Addressing Reflective and Transparent Surface Challenges in 3D Car Reconstruction. arXiv:2501.11020 [cs.CV] https://arxiv. org/abs/2501.11020

  31. [39]

    Hai Li, Xingrui Yang, Hongjia Zhai, Yuqian Liu, Hujun Bao, and Guofeng Zhang

  32. [40]

    Hai Li, Weicai Ye, Guofeng Zhang, Sanyuan Zhang, and Hujun Bao. 2020. Saliency guided subdivision for single-view mesh reconstruction. In 2020 International Conference on 3D Vision (3DV) . IEEE, 1098–1107

  33. [41]

    David Levin. 2004. Mesh-independent surface interpolation. In Geometric model- ing for scientific visualization . Springer, 37–49

  34. [42]

    Chen-Hsuan Lin, Chen Kong, and Simon Lucey. 2018. Learning Efficient Point Cloud Generation for Dense 3D Object Reconstruction. InConference on Artificial Intelligence. 7114–7121

  35. [43]

    Lingjie Liu, Jiatao Gu, Kyaw Zaw Lin, Tat-Seng Chua, and Christian Theobalt

  36. [44]

    Shilong Liu, Zhaoyang Zeng, Tianhe Ren, Feng Li, Hao Zhang, Jie Yang, Chunyuan Li, Jianwei Yang, Hang Su, Jun Zhu, et al . 2023. Grounding dino: Marrying dino with grounded pre-training for open-set object detection. arXiv preprint arXiv:2303.05499 (2023)

  37. [45]

    Lorensen and Harvey E

    William E. Lorensen and Harvey E. Cline. 1987. Marching cubes: A high resolution 3D surface construction algorithm. SIGGRAPH Comput. Graph. 21, 4 (Aug. 1987), 163–169. doi:10.1145/37402.37422

  38. [46]

    Jiahui Lyu, Bojian Wu, Dani Lischinski, Daniel Cohen-Or, and Hui Huang. 2020. Differentiable refraction-tracing for mesh reconstruction of transparent objects. ACM Trans. Graph. 39, 6, Article 195 (Nov. 2020), 13 pages. doi:10.1145/3414685. 3417815

  39. [47]

    Zhaoshuo Li, Thomas Müller, Alex Evans, Russell H Taylor, Mathias Unberath, Ming-Yu Liu, and Chen-Hsuan Lin. 2023. Neuralangelo: High-fidelity neural surface reconstruction. In Proceedings of the IEEE/CVF Conference on Computer MM ’25, October 27–31, 2025, Dublin, Ireland Ming...

  40. [48]

    Ben Mildenhall, Pratul P Srinivasan, Matthew Tancik, Jonathan T Barron, Ravi Ramamoorthi, and Ren Ng. 2021. Nerf: Representing scenes as neural radiance fields for view synthesis. Commun. ACM 65, 1 (2021), 99–106

  41. [49]

    Yuhang Ming, Weicai Ye, and Andrew Calway. 2022. idf-slam: End-to-end rgb-d slam with neural implicit mapping and deep feature tracking. arXiv preprint arXiv:2209.07919 (2022)

  42. [50]

    Nicolas Moenne-Loccoz, Ashkan Mirzaei, Or Perel, Riccardo de Lutio, Janick Mar- tinez Esturo, Gavriel State, Sanja Fidler, Nicholas Sharp, and Zan Gojcic. 2024. 3D Gaussian Ray Tracing: Fast Tracing of Particle Scenes. ACM Transactions on Graphics and SIGGRAPH Asia (2024)

  43. [51]

    Pierre Moulon, Pascal Monasse, and Renaud Marlet. 2012. Adaptive Structure from Motion with a Contrario Model Estimation. In Proceedings of the Asian Computer Vision Conference (ACCV 2012) . Springer Berlin Heidelberg, 257–270. doi:10.1007/978-3-642-37447-0_20

  44. [52]

    Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. 2022. In- stant neural graphics primitives with a multiresolution hash encoding. ACM transactions on graphics (TOG) 41, 4 (2022), 1–15

  45. [53]

    Richard A Newcombe, Shahram Izadi, Otmar Hilliges, David Molyneaux, David Kim, Andrew J Davison, Pushmeet Kohi, Jamie Shotton, Steve Hodges, and Andrew Fitzgibbon. 2011. Kinectfusion: Real-time dense surface mapping and tracking. In 2011 10th IEEE international symposium on mi...

  46. [54]

    Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger

    Lars M. Mescheder, Michael Oechsle, Michael Niemeyer, Sebastian Nowozin, and Andreas Geiger. 2019. Occupancy Networks: Learning 3D Reconstruction in Function Space. In IEEE Conference on Computer Vision and Pattern Recognition . 4460–4470

  47. [55]

    Ben Poole, Ajay Jain, Jonathan T Barron, and Ben Mildenhall. 2022. Dreamfusion: Text-to-3d using 2d diffusion. arXiv preprint arXiv:2209.14988 (2022)

  48. [56]

    Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman Rädle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junting Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao-Yuan Wu, Ross Girshick, Piotr Dollár, and Christoph Feichtenh...

  49. [57]

    Tianhe Ren, Qing Jiang, Shilong Liu, Zhaoyang Zeng, Wenlong Liu, Han Gao, Hongjie Huang, Zhengyu Ma, Xiaoke Jiang, Yihao Chen, Yuda Xiong, Hao Zhang, Feng Li, Peijun Tang, Kent Yu, and Lei Zhang. 2024. Grounding DINO 1.5: Advance the "Edge" of Open-Set Object Detection. arXiv:...

  50. [58]

    Tianhe Ren, Shilong Liu, Ailing Zeng, Jing Lin, Kunchang Li, He Cao, Jiayu Chen, Xinyu Huang, Yukang Chen, Feng Yan, Zhaoyang Zeng, Hao Zhang, Feng Li, Jie Yang, Hongyang Li, Qing Jiang, and Lei Zhang. 2024. Grounded SAM: Assembling Open-World Models for Diverse Visual Tasks. ...

  51. [59]

    Paul-Edouard Sarlin, Cesar Cadena, Roland Siegwart, and Marcin Dymczyk. 2019. From Coarse to Fine: Robust Hierarchical Localization at Large Scale. In CVPR

  52. [60]

    Johannes Lutz Schönberger and Jan-Michael Frahm. 2016. Structure-from-motion revisited. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR) . 4104–4113

  53. [61]

    Newcombe, and Steven Lovegrove

    Jeong Joon Park, Peter Florence, Julian Straub, Richard A. Newcombe, and Steven Lovegrove. 2019. DeepSDF: Learning Continuous Signed Distance Functions for Shape Representation. In IEEE Conference on Computer Vision and Pattern Recognition. 165–174

  54. [62]

    Jia-Mu Sun, Tong Wu, Ling-Qi Yan, and Lin Gao. 2024. NU-NeRF: Neural Recon- struction of Nested Transparent Objects with Uncontrolled Capture Environment. ACM Trans. Graph. 43, 6, Article 262 (Nov. 2024), 14 pages. doi:10.1145/3687757

  55. [63]

    Jiaxiang Tang, Jiawei Ren, Hang Zhou, Ziwei Liu, and Gang Zeng. 2023. Dream- gaussian: Generative gaussian splatting for efficient 3d content creation. arXiv preprint arXiv:2309.16653 (2023)

  56. [64]

    Dongqing Wang, Tong Zhang, and Sabine Süsstrunk. 2023. NEMTO: Neural Environment Matting for Novel View and Relighting Synthesis of Transparent Objects. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 317–327

  57. [65]

    Fangjinhua Wang, Silvano Galliani, Christoph Vogel, Pablo Speciale, and Marc Pollefeys. 2021. Patchmatchnet: Learned multi-view patchmatch stereo. In Pro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition . 14194–14203

  58. [66]

    Nanyang Wang, Yinda Zhang, Zhuwen Li, Yanwei Fu, Wei Liu, and Yu-Gang Jiang. 2018. Pixel2Mesh: Generating 3D Mesh Models from Single RGB Images. In European Conference on Computer Vision , Vol. 11215. 55–71

  59. [67]

    Peng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt, Taku Komura, and Wenping Wang. 2021. NeuS: learning neural implicit surfaces by volume rendering for multi-view reconstruction. In Proceedings of the 35th International Conference on Neural Information Processing Systems (N...

  60. [68]

    Johannes Lutz Schönberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm. 2016. Pixelwise View Selection for Unstructured Multi-View Stereo. In European Conference on Computer Vision (ECCV)

  61. [69]

    Qi Wu, Janick Martinez Esturo, Ashkan Mirzaei, Nicolas Moenne-Loccoz, and Zan Gojcic. 2025. 3DGUT: Enabling Distorted Cameras and Secondary Rays in Gaussian Splatting. Conference on Computer Vision and Pattern Recognition (CVPR) (2025)

  62. [70]

    Tianhao Walter Wu, Fangcheng Zhong, Gernot Riegler, Shimon Vainer, Jiankang Deng, Cengiz Oztireli, et al . [n. d.]. 𝛼surf: Implicit surface reconstruction for semi-transparent and thin objects with decoupled geometry and opacity. In International Conference on 3D Vision 2025

  63. [71]

    Haozhe Xie, Hongxun Yao, Xiaoshuai Sun, Shangchen Zhou, and Shengping Zhang. 2019. Pix2Vox: Context-Aware 3D Reconstruction From Single and Multi-View Images. In IEEE/CVF International Conference on Computer Vision . 2690–2698

  64. [72]

    Qiangeng Xu, Zexiang Xu, Julien Philip, Sai Bi, Zhixin Shu, Kalyan Sunkavalli, and Ulrich Neumann. 2022. Point-nerf: Point-based neural radiance fields. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition . 5438–5448

  65. [73]

    Ziyi Yang, Xinyu Gao, Yangtian Sun, Yihua Huang, Xiaoyang Lyu, Wen Zhou, Shaohui Jiao, Xiaojuan Qi, and Xiaogang Jin. 2024. Spec-gaussian: Anisotropic view-dependent appearance for 3d gaussian splatting. arXiv preprint arXiv:2402.15870 (2024)

  66. [74]

    Lior Yariv, Jiatao Gu, Yoni Kasten, and Yaron Lipman. 2021. Volume Rendering of Neural Implicit Surfaces. In Advances in Neural Information Processing Systems . 4805–4815

  67. [75]

    Changchang Wu. 2013. Towards linear-time incremental structure from motion. In 2013 International Conference on 3D Vision-3DV 2013 . IEEE, 127–134

  68. [76]

    Weicai Ye, Shuo Chen, Chong Bao, Hujun Bao, Marc Pollefeys, Zhaopeng Cui, and Guofeng Zhang. 2023. IntrinsicNeRF: Learning Intrinsic Neural Radiance Fields for Editable Novel View Synthesis. In Proceedings of the IEEE/CVF International Conference on Computer Vision

  69. [77]

    Weicai Ye, Hai Li, Tianxiang Zhang, Xiaowei Zhou, Hujun Bao, and Guofeng Zhang. 2021. SuperPlane: 3D plane detection and description from a single image. In 2021 IEEE Virtual Reality and 3D User Interfaces (VR) . IEEE, 207–215

  70. [78]

    Zongxin Ye, Wenyu Li, Sidun Liu, Peng Qiao, and Yong Dou. 2024. AbsGS: Recovering Fine Details in 3D Gaussian Splatting. In Proceedings of the 32nd ACM International Conference on Multimedia (Melbourne VIC, Australia) (MM ’24). Association for Computing Machinery, New York, NY...

  71. [79]

    Zehao Yu, Torsten Sattler, and Andreas Geiger. 2024. Gaussian Opacity Fields: Efficient and Compact Surface Reconstruction in Unbounded Scenes. arXiv preprint arXiv:2404.10772 (2024)

  72. [80]

    Haoran Zhang, Junkai Deng, Xuhui Chen, Fei Hou, Wencheng Wang, Hong Qin, Chen Qian, and Ying He. 2025. From transparent to opaque: rethinking neural implicit surfaces with𝛼-NeuS. In Proceedings of the 38th International Conference on Neural Information Processing Systems (Vanc...

  73. [81]

    Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang

  74. [82]

    Chongjie Ye, Lingteng Qiu, Xiaodong Gu, Qi Zuo, Yushuang Wu, Zilong Dong, Liefeng Bo, Yuliang Xiu, and Xiaoguang Han. 2024. StableNormal: Reducing Diffusion Variance for Stable and Sharp Normal. ACM Transactions on Graphics (TOG) (2024)

  75. [83]

    Hong-Kai Zhao, Stanley Osher, and Ronald Fedkiw. 2001. Fast surface reconstruc- tion using the level set method. InProceedings of the IEEE Workshop on Variational and Level Set Methods . IEEE, 194–201

  76. [84]

    Dewei Zhou, Mingwei Li, Zongxin Yang, and Yi Yang. 2025. DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models. arXiv:2503.12885 [cs.CV] https://arxiv.org/abs/2503.12885 TSGS: Improving Gaussian Splatting for Transparent Surface Reconstruct...

  77. [90]

    Cryer, and M

    Ruo Zhang, Ping-Sing Tsai, J.E. Cryer, and M. Shah. 1999. Shape-from-shading: a survey. IEEE Transactions on Pattern Analysis and Machine Intelligence 21, 8 (1999), 690–706. doi:10.1109/34.784284

  78. [2014]

    In Proceedings of the IEEE conference on computer vision and pattern recognition

    Large scale multi-view stereopsis evaluation. In Proceedings of the IEEE conference on computer vision and pattern recognition . 406–413

  79. [2016]

    In European Conference on Computer Vision , Vol

    3D-R2N2: A Unified Approach for Single and Multi-view 3D Object Recon- struction. In European Conference on Computer Vision , Vol. 9912. 628–644

  80. [2018]

    In Proceedings of the IEEE conference on computer vision and pattern recognition

    The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition . 586–595

  81. [2020]

    In Advances in Neural Information Processing Systems

    Neural Sparse Voxel Fields. In Advances in Neural Information Processing Systems. 15651–15663

  82. [2022]

    IEEE Transactions on Visualization and Computer Graphics (2022)

    Vox-surf: Voxel-based implicit surface representation. IEEE Transactions on Visualization and Computer Graphics (2022)

  83. [2023]

    In SIGGRAPH Asia 2023 Conference Papers (Sydney, NSW, Australia) (SA ’23)

    Transparent Object Reconstruction via Implicit Differentiable Refraction Rendering. In SIGGRAPH Asia 2023 Conference Papers (Sydney, NSW, Australia) (SA ’23). Association for Computing Machinery, New York, NY, USA, Article 57, 11 pages. doi:10.1145/3610548.3618236

  84. [2024]

    arXiv:2403.14610 [cs.CV]

    T-Rex2: Towards Generic Object Detection via Text-Visual Prompt Synergy. arXiv:2403.14610 [cs.CV]

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.