Pith. sign in

REVIEW 3 major objections 4 minor 93 references

BEAM: Bridging Physically-based Rendering and Gaussian Modeling for Relightable Volumetric Video

T0 review · 3 major / 4 minor · reviewed 2026-08-08 · deepseek-v4-flash

Pith's one-line read BEAM turns multi-view RGB footage of a moving person into relightable 4D Gaussians that render in standard CG engines.

desk verdict BEAM's real contribution is the pipeline integration for CG-friendly relightable 4D Gaussians, but the material decomposition has an unquantified interreflection error that the paper admits but does not evaluate. read the letter →

arxiv 2502.08297 v2 pith:T4QUL4SB submitted 2025-02-12 cs.GR cs.CV

classification cs.GRcs.CV
keywords volumetricvideorelightable4DGaussiansphysically-basedrenderingmaterialdecompositionGaussianraytracingambientocclusionhumanperformancecapturedeferredshading
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

BEAM is a pipeline that produces relightable volumetric video from multi-view RGB footage of a human performance. The central claim is that, once the dynamic geometry is tracked and several material properties are fixed, the remaining physically-based rendering attributes (ambient occlusion and base color) can be solved directly from the input images through accumulated ray-traced visibility rather than through expensive inverse rendering. The pivotal equation is $L_o = \rho L_D^o + L_S^o$, which separates outgoing radiance into a diffuse part proportional to base color $\rho$ and a specular part independent of it. This makes the recovered asset usable in existing CG pipelines, supporting real-time deferred shading at 100 FPS in 1080p and offline ray tracing. If the claim holds, it removes the need for controlled light-stage capture or neural relighting networks and gives artists a compact, editable, relightable representation of a moving person.

What carries the argument

The load-bearing object is the simplified Disney BRDF rendering equation, $L_o(\mathbf{x},\boldsymbol{\omega}_o) = \rho(\mathbf{x}) L_D^o(\mathbf{x},\boldsymbol{\omega}_o) + L_S^o(\mathbf{x},\boldsymbol{\omega}_o)$, where the diffuse residue $L_D^o$ and specular residue $L_S^o$ are evaluated by Monte Carlo integration over an environment map weighted by the visibility $V_{\mathrm{env}}(\mathbf{x},\boldsymbol{\omega}_i) = \prod_i (1 - o_i G_i(\mathbf{x}_i))$ accumulated along rays through the Gaussian cloud. Because the diffuse term is proportional to base color $\rho$ and the specular term is independent of it, the same visibility computation yields both ambient occlusion and an algebraic solution for $\rho$ from observed pixel radiance. The dynamic geometry is supplied by a dual Gaussian representation with joint Gaussians for motion tracking and skin Gaussians for surface detail, rendered through a geometry-aware rasterizer that emits depth and normal maps.

What would settle it

Run BEAM on a synthetic human with a known ground-truth base color and a deliberately strong self-occluding pose (for example, forearms held close to the chest) under an environment map containing a bright colored region; if the recovered base color or ambient occlusion in shadowed areas changes when that bright region is moved or recolored, the distant-light, no-interreflection assumption is falsified.

Watch

Extended reading notes

Core claim

The paper establishes that a 4D Gaussian sequence augmented with per-splat base color, roughness, and ambient occlusion can be relit under arbitrary environment lighting after a single capture. After tracking joint and skin Gaussians with normal-consistency constraints, BEAM assigns roughness from a generative material model, then computes 2D ambient-occlusion and base-color maps in input views. The key step is a simplification of the rendering equation, $L_o = \rho L_D^o + L_S^o$, where the diffuse and specular residues are computed by Monte Carlo integration over visibility from a tailored Gaussian ray tracer, allowing the base color to be recovered by subtraction from the measured pixel radiance. These 2D maps are then optimized into the corresponding dense Gaussian attributes, and the view-dependent color attribute is discarded, leaving a compact PBR representation.

Load-bearing premise

The whole material solve assumes that the only illumination is a known, distant environment map and that light bouncing off the person's own body can be ignored, so any significant self-reflection gets misattributed to albedo or occlusion.

Editorial extensions

If this is right

  • A single multi-view RGB capture of a moving person yields a 4D asset that can be relit under any distant environment map in real time via deferred shading at 100 FPS in 1080p, or offline with ray tracing.
  • The final Gaussian representation drops the view-dependent SH color attribute, leaving position, orientation, scale, opacity, base color, roughness, and ambient occlusion, which reduces storage and render load in CG engines.
  • Material decomposition remains stable when the number of input cameras is reduced from 50 to 20 views, based on the paper's synthetic-data ablation.
  • On synthetic data with known ground truth, BEAM reports higher PSNR, SSIM, and lower LPIPS for ambient occlusion, base color, and relighting than static Gaussian relighting methods and a parametric-body avatar method.
  • The pipeline supports scene and lighting editing in a CG engine, and the resulting 4D sequences can be deployed to VR headsets for immersive rendering and interaction.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not investigate non-human scenes, but the same visibility-first material solve should apply to any predominantly diffuse or dielectric dynamic object captured under known distant lighting, so the method is likely to generalize beyond human performance capture.
  • Because the rendering equation used here drops interreflection, a direct stress test would be a synthetic subject with known albedo in a close self-contact pose under a bright colored environment source; if the recovered base color shifts toward the source color, that failure mode is confirmed.
  • The 2D-to-3D baking strategy means achievable albedo and ambient-occlusion detail is capped by the visibility sampling rate and the denoiser, not by the Gaussian representation itself; higher-frequency materials would require more samples per pixel or a learned visibility estimate.
  • Allowing base color to vary per Gaussian over time means the representation can also encode time-varying appearance changes such as skin tone shifts, an option the paper leaves unexplored.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper presents BEAM, a pipeline for producing relightable 4D Gaussian sequences from multi-view RGB input with known environment lighting. Geometry is obtained by combining DualGS-based performance tracking with RaDe-GS-style geometry-aware rasterization; roughness is generated by CLAY's material diffusion module; ambient occlusion and base color are estimated per input view by evaluating visibility with a Gaussian ray tracer (Eqs. 4-8) and solving a simplified Disney BRDF rendering equation (Eq. 5). The resulting Gaussians support deferred real-time rendering in Unity and offline ray tracing. Claims are supported by qualitative results on six captured performances, quantitative synthetic comparisons against Relightable-3DGS, GS-IR, and MeshAvatar (Table 1), a camera-count ablation (Table 2), and a user study.

Significance. If the material decomposition were validated, this would be a practically significant contribution: it would provide a CG-friendly, relightable volumetric-video workflow with real-time 1080p rendering at 100 FPS and compatibility with standard engines. The core equations are standard inverse rendering, and the pipeline choices are reasonable and clearly described. However, the central decomposition rests on assumptions that the paper itself concedes in Section 6: direct-only environment illumination and the use of a generative model for roughness. The quantitative validation is thin, and the reported user-study percentages are internally inconsistent. These gaps must be addressed before the intrinsic-material and relightability claims can be considered established.

major comments (3)
  1. [Sec. 3.2, Eqs. (5)-(8); Sec. 6 Limitations] The base-color solve in Eq. (5) uses L_D_o and L_S_o computed with L(x,omega_i)=V_env(x,omega_i)L_env(omega_i), which accounts only for direct environment light. In a real dome capture of a human, interreflections in concave regions such as armpits, clothing folds, and under the chin add multiply-scattered radiance to the measured numerator L_o while being absent from the computed denominator L_D_o. The solver therefore bakes that indirect radiance into rho and, through the coupled AO optimization, into A. This is an unquantified failure mode of the central decomposition; Section 6 explicitly admits the approximation but no experiment measures its magnitude. Please add a synthetic experiment with known interreflection (or otherwise quantify the effect) and show that the recovered base color and AO, as well as the relit output, remain correct in concave regions.
  2. [Sec. 3.2, 'Roughness'; Eq. (7)] Roughness is generated by the CLAY material diffusion module and assigned to each Gaussian by nearest-pixel UV projection. It is not derived from the captured multi-view images under the known environment map, nor is it validated against any measured reflectance. Since r enters the specular integral L_S_o in Eq. (7), and since Eq. (5) obtains base color by subtracting L_S_o from measured L_o, errors in the hallucinated roughness directly contaminate the recovered base color and AO. Please either validate the generated roughness against synthetic ground truth or provide a sensitivity analysis showing that plausible roughness errors have a small effect on the final relighting result.
  3. [Sec. 5.1 and 5.3, Tables 1-2 and user study] The quantitative evidence is thin and partly inconsistent. Table 1 reports metrics over only two synthetic sequences without standard deviations or per-sequence breakdowns, and metrics are computed only inside the human bounding box. The user study states that 30 users participated but reports 95.65% and 87% preference; with 30 users, 95.65% is not attainable, so either the participant count or the percentages need correction. Additionally, Table 2 reports an 'Ours' relighting PSNR of 26.21 while Table 1 reports an 'Ours' relighting PSNR of 26.57 for what appears to be the same 50-camera synthetic setup; the relationship between the two tables should be clarified. These issues undercut the claim of consistent superiority and should be fixed with proper statistics.
minor comments (4)
  1. [Sec. 3.1, after Eq. (2)] The text says 'E_norm is introduced' but the energy term defined above is E_normal; please unify the notation.
  2. [Sec. 1, related work citations] In the introduction, citations such as 'Li [43]' and 'High-quality FVV [9]' should use the same style as the rest of the text, adding 'et al.' where there are multiple authors.
  3. [Sec. 3.2, AO/base-color optimization] The sentence 'initialize the AO attributes with zeros and the base color attributes with RGB values' is vague; please specify the initialization (for example, from the input-view SH color, from a per-view bake, or from a constant).
  4. [Sec. 5.2, Figs. 7 and 8] The ablations for the ray-origin offset and the sampling strategy are qualitative only; adding quantitative metrics, even for a few frames, would make the reported choices easier to evaluate.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the material decomposition is an explicit inverse-rendering solve from measured radiance and a known environment map, with relighting evaluated against external synthetic ground truth.

full rationale

The central derivation chain in Sec. 3.2 is a direct inverse-rendering computation rather than a circular restatement of its inputs. Base color is solved from Eq. 5, rho = (L_o - L_S_o)/L_D_o, after computing L_D_o and L_S_o by Monte Carlo integration of Eqs. 6 and 7 using visibility from Eq. 8, the captured environment map, and the estimated roughness. This is an explicit physical inversion: the measured outgoing radiance L_o is the input, and the extracted base color is the unknown solved from a simplified rendering equation. It is not a fitted parameter renamed as a prediction; the relighting evaluation uses Blender-rendered synthetic ground truth from RenderPeople meshes (Sec. 5.1), providing an external benchmark independent of the training views. The paper's use of prior work by the same authors, DualGS [35] for tracking and CLAY [80] for roughness generation, is component reuse rather than load-bearing circularity: those components are published external systems, and the core material decomposition does not invoke a self-citation to justify its own result. The limitation acknowledged in Sec. 6, that the rendering equation is approximated and indirect illumination is disregarded, is a physical-accuracy caveat rather than a circularity: it means residual interreflection may be baked into base color or AO, but this does not make the stated derivation equivalent to its inputs by construction. No equation reduces to its own output, no fitted quantity is relabeled as a prediction, and no uniqueness claim is imported from the authors' prior work. The derivation is self-contained given the stated direct-illumination and distant-environment assumptions, so the circularity score is 0.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The extended Gaussian attribute vector in Eq. 3 is a representation change, not a new physical entity. No new particles, forces, mediators, or dimensions are introduced. The load-bearing imports are assumptions about lighting, geometry, and the diffusion-generated roughness.

free parameters (5)
  • Geometry optimization loss weights (lambda_color, lambda_smooth, lambda_temp, lambda_normal) = 1, 0.001, 0.00005, 0.03
    Hand-set coefficients in Eq. 2 that balance photometric, smoothness, temporal, and normal terms. They are chosen by the authors, not derived.
  • Ray origin offset epsilon for material ray tracing = 0.02
    Offset applied to ray origins along surface normals to avoid self-occlusion. The Section 5.2 ablation shows strong sensitivity, with artifacts at 0 and 0.1, so the value is tuned on captured human data.
  • Monte Carlo sample counts for AO and base color = 50 spp for AO, 100 spp for base color
    Selected as a balance between noise and computation; the ablation compares lower/higher counts and denoising, confirming the choice is empirical.
  • Early ray termination threshold T = 0.0001
    Computational stopping criterion for the visibility product in Eq. 8; a practical approximation threshold.
  • Target number of Gaussians = about 140,000
    Selected via ablation on the virtual dataset as the best balance between rendering quality and geometric fidelity.
assumptions (5)
  • domain assumption Human-centric scenes have negligible metallic component, so metallic is set to zero.
    Section 3 intro reduces the PBR attribute set to roughness, AO, base color, and normal. This fails for non-human or metallic objects, but the claim is scoped to human performances.
  • domain assumption Capture lighting is a known, distant environment map with no interreflections from the body.
    Section 3.2, Eqs. 5-7 ignore indirect illumination and assume the environment map from bracketed DSLR photos stitched in PTGui is correct. The Limitations section concedes this introduces errors in the decoupled materials.
  • domain assumption Ray-Gaussian intersections occur at the maximum Gaussian response, giving correct depth and normals.
    Section 3.1 and Eq. 8 use this assumption both for the normal consistency loss and for ray-traced visibility in material estimation.
  • ad hoc to paper Roughness from the CLAY material diffusion module is a valid, time-invariant material property for each Gaussian.
    Section 3.2 assigns roughness by projecting a generated texture onto the mesh and taking nearest-pixel values for Gaussians. The roughness is never validated against measured material properties.
  • domain assumption TSDF fusion of rendered depth maps yields a mesh accurate enough for UV mapping and roughness transfer.
    Section 3.1 uses the mesh from depth fusion as the bridge to CLAY texture generation and subsequent material baking; errors in this mesh propagate to materials.

how reviews work

0 comments
Cite this review

Pith. "Pith review of BEAM: Bridging Physically-based Rendering and Gaussian Modeling for Relightable Volumetric Video." pith.science (2026). https://pith.science/paper/T4QUL4SB

@misc{pith2026250208297,
  author       = {Pith},
  title        = {Pith review of: BEAM: Bridging Physically-based Rendering and Gaussian Modeling for Relightable Volumetric Video},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/T4QUL4SB}},
  note         = {Machine review of arXiv:2502.08297}
}
read the original abstract

Volumetric video enables immersive experiences by capturing dynamic 3D scenes, enabling diverse applications for virtual reality, education, and telepresence. However, traditional methods struggle with fixed lighting conditions, while neural approaches face trade-offs in efficiency, quality, or adaptability for relightable scenarios. To address these limitations, we present BEAM, a novel pipeline that bridges 4D Gaussian representations with physically-based rendering (PBR) to produce high-quality, relightable volumetric videos from multi-view RGB footage. BEAM recovers detailed geometry and PBR properties via a series of available Gaussian-based techniques. It first combines Gaussian-based human performance tracking with geometry-aware rasterization in a coarse-to-fine optimization framework to recover spatially and temporally consistent geometries. We further enhance Gaussian attributes by incorporating PBR properties step by step. We generate roughness via a multi-view-conditioned diffusion model, and then derive AO and base color using a 2D-to-3D strategy, incorporating a tailored Gaussian-based ray tracer for efficient visibility computation. Once recovered, these dynamic, relightable assets integrate seamlessly into traditional CG pipelines, supporting real-time rendering with deferred shading and offline rendering with ray tracing. By offering realistic, lifelike visualizations under diverse lighting conditions, BEAM opens new possibilities for interactive entertainment, storytelling, and creative visualization.

Figures

Figures reproduced from arXiv: 2502.08297 by the authors.

Figure 1
Figure 1. We present BEAM, a novel pipeline that bridges [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. We propose a novel BEAM method to produce Gaussian sequences. We first use tracking results of joint Gaussians [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Gallery of our results. We present some real-time rendering results under HDRI settings, which deliver high-fidelity [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: We demonstrate scene and lighting editing within Unity, and immersive viewing using VR headsets. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: We present the results of different rendering tech [PITH_FULL_IMAGE:figures/full_fig_p006_5.png]
Figure 6
Figure 6. Figure 6: Qualitative comparisons of our method against R-3DGS [ [PITH_FULL_IMAGE:figures/full_fig_p007_6.png]
Figure 9
Figure 9. Figure 9: Ablation study on the number of geometry-aware [PITH_FULL_IMAGE:figures/full_fig_p008_9.png]
Figure 8
Figure 8. Figure 8: Ablation on sampling strategies for material decom [PITH_FULL_IMAGE:figures/full_fig_p008_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

93 extracted references · 67 canonical work pages

  1. [1]

    Mark Boss, Raphael Braun, Varun Jampani, Jonathan T Barron, Ce Liu, and Hendrik Lensch. 2021. Nerd: Neural reflectance decomposition from image collections. InProceedings of the IEEE/CVF International Conference on Computer Vision. 12684–12694

  2. [2]

    Aljaz Bozic, Pablo Palafox, Michael Zollhofer, Justus Thies, Angela Dai, and Matthias Nießner. 2021. Neural Deformation Graphs for Globally-consistent Non-rigid Reconstruction. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 1450–1459

  3. [3]

    Brent Burley and Walt Disney Animation Studios. 2012. Physically-based shading at disney. InAcm Siggraph, Vol. 2012. vol. 2012, 1–7

  4. [4]

    Diogo Carbonera Luvizon, Vladislav Golyanik, Adam Kortylewski, Marc Haber- mann, and Christian Theobalt. 2024. Relightable Neural Actor with Intrinsic Decomposition and Pose Control. InEuropean Conference on Computer Vision (Milan, Italy). Springer-Verlag, 465–483

  5. [5]

    Charles-Félix Chabert, Per Einarsson, Andrew Jones, Bruce Lamond, Wan-Chun Ma, Sebastian Sylwan, Tim Hawkins, and Paul Debevec. 2006. Relighting human locomotion with flowed reflectance fields. InACM SIGGRAPH 2006 Sketches. 76–es

  6. [6]

    Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. 2022. Tensorf: Tensorial radiance fields. InEuropean Conference on Computer Vision. Springer, 333–350

  7. [7]

    Yushuo Chen, Zerong Zheng, Zhe Li, Chao Xu, and Yebin Liu. 2024. MeshAvatar: Learning High-quality Triangular Human Avatars from Multi-view Videos. In European Conference on Computer Vision

  8. [8]

    Zhaoxi Chen and Ziwei Liu. 2022. Relighting4d: Neural relightable human from videos. InEuropean Conference on Computer Vision. Springer, 606–623

Show all 93 references
  1. [9]

    Alvaro Collet, Ming Chuang, Pat Sweeney, Don Gillett, Dennis Evseev, David Calabrese, Hugues Hoppe, Adam Kirk, and Steve Sullivan. 2015. High-quality streamable free-viewpoint video.ACM Transactions on Graphics (TOG)34, 4 (2015), 69

  2. [10]

    R. L. Cook and K. E. Torrance. 1982. A Reflectance Model for Computer Graphics. ACM Trans. Graph.1, 1 (Jan. 1982), 7–24. doi:10.1145/357290.357293

  3. [11]

    Brian Curless and Marc Levoy. 1996. A volumetric method for building com- plex models from range images. InProceedings of the 23rd annual conference on Computer graphics and interactive techniques. 303–312

  4. [12]

    Paul Debevec. 2012. The light stages and their applications to photoreal digital actors.SIGGRAPH Asia2, 4 (2012), 1–6

  5. [13]

    Michael Deering, Stephanie Winner, Bic Schediwy, Chris Duffy, and Neil Hunt

  6. [14]

    Zheng Ding, Xuaner Zhang, Zhihao Xia, Lars Jebe, Zhuowen Tu, and Xiuming Zhang. 2023. Diffusionrig: Learning personalized priors for facial appearance editing. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12736–12746

  7. [15]

    Mingsong Dou, Sameh Khamis, Yury Degtyarev, Philip Davidson, Sean Ryan Fanello, Adarsh Kowdle, Sergio Orts Escolano, Christoph Rhemann, David Kim, Jonathan Taylor, Pushmeet Kohli, Vladimir Tankovich, and Shahram Izadi. 2016. Fusion4D: real-time performance capture of challengi...

  8. [16]

    Yuanxing Duan, Fangyin Wei, Qiyu Dai, Yuhang He, Wenzheng Chen, and Bao- quan Chen. 2024. 4d-rotor gaussian splatting: towards efficient novel view synthesis for dynamic scenes. InACM SIGGRAPH 2024 Conference Papers. 1–11

  9. [17]

    Sara Fridovich-Keil, Giacomo Meanti, Frederik Rahbæk Warburg, Benjamin Recht, and Angjoo Kanazawa. 2023. K-planes: Explicit radiance fields in space, time, and appearance. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 12479–12488

  10. [18]

    Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. 2022. Plenoxels: Radiance fields without neural net- works. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5501–5510

  11. [19]

    Jian Gao, Chun Gu, Youtian Lin, Zhihao Li, Hao Zhu, Xun Cao, Li Zhang, and Yao Yao. 2025. Relightable 3d gaussians: Realistic point cloud relighting with brdf decomposition and ray tracing. InEuropean Conference on Computer Vision. Springer, 73–89

  12. [20]

    Mathieu Garon, Pierre-Olivier Boulet, Jean-Philippe Doiron, Luc Beaulieu, and Jean-François Lalonde. 2016. Real-time high resolution 3D data on the HoloLens. In2016 IEEE International Symposium on Mixed and Augmented Reality (ISMAR- Adjunct). IEEE, 189–191

  13. [21]

    Chun Gu, Xiaofei Wei, Zixuan Zeng, Yuxuan Yao, and Li Zhang. 2024. IRGS: Inter-Reflective Gaussian Splatting with 2D Gaussian Ray Tracing.arXiv preprint arXiv:2412.15867(2024)

  14. [22]

    Kaiwen Guo, Peter Lincoln, Philip Davidson, Jay Busch, Xueming Yu, Matt Whalen, Geoff Harvey, Sergio Orts-Escolano, Rohit Pandey, Jason Dourgarian, et al. 2019. The relightables: Volumetric performance capture of humans with realistic relighting.ACM Transactions on Graphics (T...

  15. [23]

    Marc Habermann, Lingjie Liu, Weipeng Xu, Gerard Pons-Moll, Michael Zollhoefer, and Christian Theobalt. 2023. Hdhumans: A hybrid approach for high-fidelity digital humans.Proceedings of the ACM on Computer Graphics and Interactive Techniques6, 3 (2023), 1–23

  16. [24]

    Tim Hawkins, Jonathan Cohen, and Paul Debevec. 2001. A photometric approach to digitizing cultural artifacts. InProceedings of the 2001 conference on Virtual reality, archeology, and cultural heritage. 333–342

  17. [25]

    Qiang Hu, Qihan He, Houqiang Zhong, Guo Lu, Xiaoyun Zhang, Guangtao Zhai, and Yanfeng Wang. 2025. Varfvv: View-adaptive real-time interactive free- view video streaming with edge computing.IEEE Journal on Selected Areas in Communications(2025)

  18. [26]

    Qiang Hu, Zihan Zheng, Houqiang Zhong, Sihua Fu, Li Song, Xiaoyun Zhang, Guangtao Zhai, and Yanfeng Wang. 2025. 4DGC: Rate-Aware 4D Gaussian Compression for Efficient Streamable Free-Viewpoint Video. InProceedings of the Computer Vision and Pattern Recognition Conference. 875–885

  19. [27]

    Qiang Hu, Houqiang Zhong, Zihan Zheng, Xiaoyun Zhang, Zhengxue Cheng, Li Song, Guangtao Zhai, and Yanfeng Wang. 2025. VRVVC: Variable-Rate NeRF- Based Volumetric Video Compression. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 39. 3563–3571

  20. [28]

    Yi-Hua Huang, Yang-Tian Sun, Ziyi Yang, Xiaoyang Lyu, Yan-Pei Cao, and Xiao- juan Qi. 2024. Sc-gs: Sparse-controlled gaussian splatting for editable dynamic scenes. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 4220–4230

  21. [29]

    Matthias Innmann, Michael Zollhöfer, Matthias Nießner, Christian Theobalt, and Marc Stamminger. 2016. Volumedeform: Real-time volumetric non-rigid reconstruction. InEuropean Conference on Computer Vision. Springer, 362–379

  22. [30]

    Intel Corporation. 2025. Intel®Open Image Denoise. https://www. openimagedenoise.org/

  23. [31]

    Mustafa Işık, Martin Rünz, Markos Georgopoulos, Taras Khakhulin, Jonathan Starck, Lourdes Agapito, and Matthias Nießner. 2023. Humanrf: High-fidelity neural radiance fields for humans in motion.ACM Transactions on Graphics (TOG)42, 4 (2023), 1–12

  24. [32]

    Yuheng Jiang, Suyi Jiang, Guoxing Sun, Zhuo Su, Kaiwen Guo, Minye Wu, Jingyi Yu, and Lan Xu. 2022. NeuralHOFusion: Neural Volumetric Rendering Under Human-Object Interactions. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition. 6155–6165

  25. [33]

    Yujiao Jiang, Qingmin Liao, Zhaolong Wang, Xiangru Lin, Zongqing Lu, Yuxi Zhao, Hanqing Wei, Jingrui Ye, Yu Zhang, and Zhijing Shao. 2024. SMPLX-Lite: A Realistic and Drivable Avatar Benchmark with Rich Geometry and Texture Annotations. In2024 IEEE International Conference on ...

  26. [34]

    Yuheng Jiang, Zhehao Shen, Chengcheng Guo, Yu Hong, Zhuo Su, Yingliang Zhang, Marc Habermann, and Lan Xu. 2025. RePerformer: Immersive Human- centric Volumetric Videos from Playback to Photoreal Reperformance. InPro- ceedings of the Computer Vision and Pattern Recognition Conf...

  27. [35]

    Yuheng Jiang, Zhehao Shen, Yu Hong, Chengcheng Guo, Yize Wu, Yingliang Zhang, Jingyi Yu, and Lan Xu. 2024. Robust Dual Gaussian Splatting for Immersive Human-centric Volumetric Videos.ACM Transactions on Graphics (TOG)43, 6 (2024), 1–15

  28. [36]

    Yuheng Jiang, Zhehao Shen, Penghao Wang, Zhuo Su, Yu Hong, Yingliang Zhang, Jingyi Yu, and Lan Xu. 2024. Hifi4g: High-fidelity human performance rendering via compact gaussian splatting. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 19734–19745

  29. [37]

    Yingwenqi Jiang, Jiadong Tu, Yuan Liu, Xifeng Gao, Xiaoxiao Long, Wenping Wang, and Yuexin Ma. 2024. Gaussianshader: 3d gaussian splatting with shading functions for reflective surfaces. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5322–5332

  30. [38]

    Yuheng Jiang, Kaixin Yao, Zhuo Su, Zhehao Shen, Haimin Luo, and Lan Xu

  31. [39]

    James T Kajiya. 1986. The rendering equation. InProceedings of the 13th annual conference on Computer graphics and interactive techniques. 143–150

  32. [40]

    Yoshihiro Kanamori and Yuki Endo. 2018. Relighting humans: occlusion-aware inverse rendering for full-body human images.ACM Transactions on Graphics (Dec 2018), 1–11. doi:10.1145/3272127.3275104

  33. [41]

    Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, and George Drettakis

  34. [42]

    Jason Lawrence, Ryan Overbeck, Todd Prives, Tommy Fortes, Nikki Roth, and Brett Newman. 2024. Project starline: A high-fidelity telepresence system. In ACM SIGGRAPH 2024 Emerging Technologies. 1–2

  35. [43]

    Hao Li, Linjie Luo, Daniel Vlasic, Pieter Peers, Jovan Popović, Mark Pauly, and Szymon Rusinkiewicz. 2012. Temporally coherent completion of dynamic shapes. ACM Trans. Graph.31, 1, Article 2 (Feb. 2012), 11 pages

  36. [44]

    MM ’25, October 27–31, 2025, Dublin, Ireland Yu Hong et al

    3d gaussian splatting for real-time radiance field rendering.ACM Transac- tions on Graphics (ToG)42, 4 (2023), 1–14. MM ’25, October 27–31, 2025, Dublin, Ireland Yu Hong et al

  37. [45]

    Zhe Li, Zerong Zheng, Lizhen Wang, and Yebin Liu. 2024. Animatable Gaussians: Learning Pose-dependent Gaussian Maps for High-fidelity Human Avatar Mod- eling. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

  38. [46]

    Zhihao Liang, Qi Zhang, Ying Feng, Ying Shan, and Kui Jia. 2024. Gs-ir: 3d gaussian splatting for inverse rendering. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 21644–21653

  39. [47]

    Zhan Li, Zhang Chen, Zhong Li, and Yi Xu. 2024. Spacetime gaussian feature splatting for real-time dynamic view synthesis. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 8508–8520

  40. [48]

    Jonathon Luiten, Georgios Kopanas, Bastian Leibe, and Deva Ramanan. 2024. Dynamic 3d gaussians: Tracking by persistent dynamic view synthesis. In2024 International Conference on 3D Vision (3DV). IEEE, 800–809

  41. [49]

    Srinivasan, Matthew Tancik, Jonathan T

    Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. 2020. NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis. InComputer Vision – ECCV 2020, Andrea Vedaldi, Horst Bischof, Thomas Brox, and Jan-Michael Frah...

  42. [50]

    Matthew Loper, Naureen Mahmood, Javier Romero, Gerard Pons-Moll, and Michael J Black. 2023. SMPL: A skinned multi-person linear model. InSeminal Graphics Papers: Pushing the Boundaries, Volume 2. 851–866

  43. [51]

    Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. 2022. Instant Neural Graphics Primitives with a Multiresolution Hash Encoding.ACM Trans. Graph.41, 4, Article 102 (July 2022), 15 pages. doi:10.1145/3528223.3530127

  44. [52]

    New House Internet Services B.V. 2025. PTGui Stitching Software. https://ptgui. com

  45. [53]

    Nicolas Moenne-Loccoz, Ashkan Mirzaei, Or Perel, Riccardo de Lutio, Janick Mar- tinez Esturo, Gavriel State, Sanja Fidler, Nicholas Sharp, and Zan Gojcic. 2024. 3D Gaussian Ray Tracing: Fast Tracing of Particle Scenes.ACM Transactions on Graphics and SIGGRAPH Asia(2024)

  46. [54]

    Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang, Qing Shuai, Hujun Bao, and Xiaowei Zhou. 2021. Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. In Proceedings of the IEEE/CVF Conference on Computer Visi...

  47. [55]

    Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer

  48. [56]

    Richard A Newcombe, Dieter Fox, and Steven M Seitz. 2015. Dynamicfusion: Reconstruction and tracking of non-rigid scenes in real-time. InProceedings of the IEEE conference on computer vision and pattern recognition. 343–352

  49. [57]

    Ruizhi Shao, Zerong Zheng, Hanzhang Tu, Boning Liu, Hongwen Zhang, and Yebin Liu. 2023. Tensor4d: Efficient neural 4d decomposition for high-fidelity dynamic reconstruction and rendering. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 16632–16642

  50. [58]

    Yahao Shi, Yanmin Wu, Chenming Wu, Xing Liu, Chen Zhao, Haocheng Feng, Jingtuo Liu, Liangjun Zhang, Jian Zhang, Bin Zhou, et al. 2023. Gir: 3d gaussian in- verse rendering for relightable scene factorization.arXiv preprint arXiv:2312.05133 (2023)

  51. [59]

    Miroslava Slavcheva, Maximilian Baust, Daniel Cremers, and Slobodan Ilic. 2017. Killingfusion: Non-rigid 3d reconstruction without correspondences. InProceed- ings of the IEEE Conference on Computer Vision and Pattern Recognition. 1386– 1395

  52. [60]

    RenderPeople. 2025. RenderPeople. https://renderpeople.com/

  53. [61]

    Pratul P Srinivasan, Boyang Deng, Xiuming Zhang, Matthew Tancik, Ben Milden- hall, and Jonathan T Barron. 2021. Nerv: Neural reflectance and visibility fields for relighting and view synthesis. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition...

  54. [62]

    Guoxing Sun, Xin Chen, Yizhang Chen, Anqi Pang, Pei Lin, Yuheng Jiang, Lan Xu, Jingyi Yu, and Jingya Wang. 2021. Neural free-viewpoint performance rendering under complex human-object interactions. InProceedings of the 29th ACM International Conference on Multimedia. 4651–4660

  55. [63]

    Guoxing Sun, Rishabh Dabral, Heming Zhu, Pascal Fua, Christian Theobalt, and Marc Habermann. 2025. Real-time Free-view Human Rendering from Sparse-view RGB Videos using Double Unprojected Textures. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni...

  56. [64]

    Miroslava Slavcheva, Maximilian Baust, and Slobodan Ilic. 2018. Sobolevfusion: 3d reconstruction of scenes undergoing free non-rigid motion. InProceedings of the IEEE conference on computer vision and pattern recognition. 2646–2655

  57. [65]

    Xin Suo, Yuheng Jiang, Pei Lin, Yingliang Zhang, Minye Wu, Kaiwen Guo, and Lan Xu. 2021. NeuralHumanFVV: Real-Time Neural Volumetric Human Performance Rendering using RGB Cameras. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 6226–6237

  58. [66]

    Tajima, Y

    D. Tajima, Y. Kanamori, and Y. Endo. 2021. Relighting Humans in the Wild: Monocular Full-Body Human Relighting with Domain Adaptation.Computer Graphics Forum(Oct 2021), 205–216. doi:10.1111/cgf.14414

  59. [67]

    Edgar Tretschk, Ayush Tewari, Vladislav Golyanik, Michael Zollhöfer, Christoph Lassner, and Christian Theobalt. 2021. Non-rigid neural radiance fields: Recon- struction and novel view synthesis of a dynamic scene from monocular video. In Proceedings of the IEEE/CVF Internation...

  60. [68]

    Jiakai Sun, Han Jiao, Guangyuan Li, Zhanjie Zhang, Lei Zhao, and Wei Xing

  61. [69]

    Penghao Wang, Zhirui Zhang, Liao Wang, Kaixin Yao, Siyuan Xie, Jingyi Yu, Minye Wu, and Lan Xu. 2024. Vˆ 3: Viewing Volumetric Videos on Mobiles via Streamable 2D Dynamic Gaussians.ACM Transactions on Graphics (TOG)43, 6 (2024), 1–13

  62. [70]

    Andreas Wenger, Andrew Gardner, Chris Tchou, Jonas Unger, Tim Hawkins, and Paul Debevec. 2005. Performance relighting and reflectance transformation with time-multiplexed illumination.ACM Transactions on Graphics (TOG)24, 3 (2005), 756–764

  63. [71]

    Tim Weyrich, Wojciech Matusik, Hanspeter Pfister, Bernd Bickel, Craig Donner, Chien Tu, Janet McAndless, Jinho Lee, Addy Ngan, Henrik Wann Jensen, et al

  64. [72]

    Wenqi Xian, Jia-Bin Huang, Johannes Kopf, and Changil Kim. 2021. Space-time neural irradiance fields for free-viewpoint video. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 9421–9431

  65. [73]

    Liao Wang, Qiang Hu, Qihan He, Ziyu Wang, Jingyi Yu, Tinne Tuytelaars, Lan Xu, and Minye Wu. 2023. Neural Residual Radiance Fields for Streamably Free- Viewpoint Videos. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 76–87

  66. [74]

    Zeyu Yang, Zijie Pan, Xiatian Zhu, Li Zhang, Yu-Gang Jiang, and Philip H. S. Torr. 2024. 4D Gaussian Splatting: Modeling Dynamic Scenes with Native 4D Primitives. https://api.semanticscholar.org/CorpusID:275134199

  67. [75]

    Yao Yao, Jingyang Zhang, Jingbo Liu, Yihang Qu, Tian Fang, David McKin- non, Yanghai Tsin, and Long Quan. 2022. Neilf: Neural incident light field for physically-based material estimation. InEuropean Conference on Computer Vision. Springer, 700–716

  68. [76]

    Tao Yu, Zerong Zheng, Kaiwen Guo, Pengpeng Liu, Qionghai Dai, and Yebin Liu. 2021. Function4D: Real-time Human Volumetric Capture from Very Sparse Consumer RGBD Sensors. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 5746–5756

  69. [77]

    Chong Zeng, Guojun Chen, Yue Dong, Pieter Peers, Hongzhi Wu, and Xin Tong

  70. [78]

    Chong Zeng, Yue Dong, Pieter Peers, Youkang Kong, Hongzhi Wu, and Xin Tong

  71. [79]

    Zhen Xu, Sida Peng, Chen Geng, Linzhan Mou, Zihan Yan, Jiaming Sun, Hujun Bao, and Xiaowei Zhou. 2024. Relightable and animatable neural avatar from sparse-view video. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 990–1000

  72. [80]

    Longwen Zhang, Ziyu Wang, Qixuan Zhang, Qiwei Qiu, Anqi Pang, Haoran Jiang, Wei Yang, Lan Xu, and Jingyi Yu. 2024. CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D Assets.ACM Transactions on Graphics (TOG)43, 4 (2024), 1–20

  73. [81]

    Xiuming Zhang, Pratul P Srinivasan, Boyang Deng, Paul Debevec, William T Freeman, and Jonathan T Barron. 2021. Nerfactor: Neural factorization of shape and reflectance under an unknown illumination.ACM Transactions on Graphics (ToG)40, 6 (2021), 1–18

  74. [82]

    Yang Zheng, Qingqing Zhao, Guandao Yang, Wang Yifan, Donglai Xiang, Flo- rian Dubost, Dmitry Lagun, Thabo Beeler, Federico Tombari, Leonidas Guibas, et al. 2025. Physavatar: Learning the physics of dressed 3d avatars from visual observations. InEuropean Conference on Computer ...

  75. [83]

    Zerong Zheng, Xiaochen Zhao, Hongwen Zhang, Boning Liu, and Yebin Liu

  76. [84]

    InACM SIGGRAPH 2023 Conference Proceedings

    Relighting neural radiance fields with shadow and highlight hints. InACM SIGGRAPH 2023 Conference Proceedings. 1–11

  77. [86]

    InSpecial Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers ’24 (SIGGRAPH ’24)

    DiLightNet: Fine-grained Lighting Control for Diffusion-based Image Gen- eration. InSpecial Interest Group on Computer Graphics and Interactive Techniques Conference Conference Papers ’24 (SIGGRAPH ’24). ACM, 1–12

  78. [87]

    Baowen Zhang, Chuan Fang, Rakesh Shrestha, Yixun Liang, Xiaoxiao Long, and Ping Tan. 2024. RaDe-GS: Rasterizing Depth in Gaussian Splatting.arXiv preprint arXiv:2406.01467(2024)

  79. [92]

    AvatarRex: Real-time Expressive Full-body Avatars.ACM Transactions on Graphics (TOG)42, 4 (2023)

  80. [93]

    Heming Zhu, Fangneng Zhan, Christian Theobalt, and Marc Habermann. 2024. Trihuman: a real-time and controllable tri-plane representation for detailed hu- man geometry and appearance synthesis.ACM Transactions on Graphics44, 1 (2024), 1–17

  81. [1988]

    InProceedings of the 15th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH ’88)

    The triangle processor and normal vector shader: a VLSI system for high performance graphics. InProceedings of the 15th Annual Conference on Computer Graphics and Interactive Techniques (SIGGRAPH ’88). 21–30. doi:10.1145/54852. 378468

  82. [2006]

    ACM Transactions on Graphics (ToG)25, 3 (2006), 1013–1024

    Analysis of human faces using a measurement-based skin reflectance model. ACM Transactions on Graphics (ToG)25, 3 (2006), 1013–1024

  83. [2021]

    InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    D-nerf: Neural radiance fields for dynamic scenes. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 10318–10327

  84. [2023]

    InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    Instant-NVR: Instant Neural Volumetric Rendering for Human-Object Interactions From Monocular RGBD Stream. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 595–605

  85. [2024]

    InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

    3dgstream: On-the-fly training of 3d gaussians for efficient streaming of photo-realistic free-viewpoint videos. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 20675–20685

Pith tools

Reviewed August 8, 2026 · model on record in the stance chip above.