REVIEW 4 major objections 8 minor 2 cited by
SignSplat: Rendering Sign Language via Gaussian Splatting
T0 review · 4 major / 8 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read The paper claims that anchoring regularized 3D Gaussians to an SMPL-X mesh renders sign language with higher fidelity than existing human avatar methods, especially for hands and face.
desk verdict Solid engineering contribution to avatar rendering, but the sign-language SOTA claim rests on a private 24-frame test set and author-provided baseline fits; the public benchmark results are the load-bearing evidence. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is an SMPL-X mesh with 3D Gaussians anchored in canonical space; a shared 1D convolutional network maps each canonical point and its pose transform to a feature embedding, from which separate networks predict Gaussian scales, rotations, opacities, and colors. The argument is carried by three interacting mechanisms: mesh densification and adaptive control that clone and prune Gaussians on mesh faces rather than in free space, hard constraints on hand degrees of freedom and vertex displacement limits that keep the fitted mesh anatomically plausible, and variance-based regularization over mesh-face neighborhoods that prevents splat parameters from overfitting training views. Together these mechanisms keep the high-frequency hand and face detail intact during novel pose and novel view synthesis.
What would settle it
Render a held-out signing sequence for which ground-truth hand motion is available from a motion-capture corpus, then compute hand-region PSNR and keypoint reprojection error for this method versus the closest sign-language baseline. If hand-region error does not track the initial SMPL-X hand fitting error, or if the method fails whenever the hand estimator fails, the central claim would need revision.
Extended reading notes
Core claim
On its own terms, the paper's discovery is that mesh-anchored Gaussian splatting can be made to work for sign language if the Gaussian parameters are regularized strongly enough, and that the usual tricks of the avatar-rendering literature are insufficient without such regularization. The method attaches 3D Gaussians to a densified SMPL-X mesh, predicts per-point attributes through a shared convolutional feature embedding, and constrains scales, opacities, colors, and mesh displacements differently for body, face, and hands. With those constraints, the paper reports the highest PSNR, SSIM, and LPIPS among compared methods on the NeuMan and X-Humans benchmarks, and on a six-view sign-language test set it reports a clear margin over prior sign-language Gaussian splatting and over generic human avatars. The claim is therefore that careful regularization, not more cameras or a larger model, is what unlocks subtle hand and face motion in few-view avatar rendering.
Load-bearing premise
The whole pipeline inherits the accuracy of the single-view SMPL-X pose and hand fitting: if the hand pose estimator misplaces fingers in fast, self-occluded signing, the mesh-anchored Gaussians cannot recover correct appearance.
Editorial extensions
If this is right
- Novel sign sequences can be generated from a library of single-view glosses without multi-view capture, because the SMPL-X mesh anchor supplies the 3D prior the splats need.
- Sign language production can move from 2D pose interpolation to 3D mesh interpolation in SMPL-X parameter space, removing scale and depth ambiguities and reducing finger-blending artifacts during gloss transitions.
- Benchmark suites for human avatar rendering should include highly articulated signing sequences, since methods that look comparable on walking and dancing separate clearly on hand and face motion.
- Accumulating the image loss over multiple poses before backpropagation is a stabilizer for avatar training, and the ablations indicate it improves accuracy on its own.
Reading between the lines
- Because the framework's only hard requirement is a mesh prior with known skinning, an analogous regularized-anchor design could transfer to other articulated subjects such as four-legged or virtual characters; the paper does not test this.
- The delayed activation of spherical harmonics points to a view-consistency bottleneck: a future version could replace the delay with an explicit multi-view consistency loss and likely recover texture detail earlier, which is a natural extension rather than a claim of the paper.
- The method's success on six-view sign data suggests that temporal variability in a sequence can substitute for camera diversity; a direct test would train from a single view and compare hand-region fidelity against the six-view model.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SignSplat, a Gaussian-splatting avatar model anchored to SMPL-X meshes, targeting expressive human motion such as sign language. The method lifts single- or multi-view video to SMPL-X with hand/face refinement, attaches canonical Gaussians to mesh vertices, and introduces two technical components: (i) a shared 1D CNN over canonical mesh points to predict pose-dependent Gaussian attributes and vertex displacements, and (ii) a mesh-face densification/pruning strategy with neighborhood variance regularization. The authors also propose a 3D gloss-stitching procedure that interpolates SMPL-X parameters via SLERP with an adaptive frame count. Experiments report state-of-the-art on NeuMan and X-Humans, and a claimed significant gain on a private six-view sign-language sequence (323 training frames, 24 test frames) over GaussianAvatar, SplattingAvatar, ExAvatar, and EVA. An additional ablation on NeuMan and a qualitative sign-stitching comparison are provided.
Significance. If the claims hold, the paper would advance photorealistic avatar rendering for a socially important domain where hand and face fidelity matters more than gross body motion, and it would provide a practical 3D alternative to 2D-GAN sign language production. The method is a reasonable engineering contribution: it adapts existing mesh-anchored Gaussian-splatting ideas (SplattingAvatar, ExAvatar, GaussianAvatar) with sequence-level regularization and a mesh-aware densification scheme, and the NeuMan/X-Humans results are consistent with the state of the art. The benchmark tables are a useful check that the framework did not sacrifice generic performance. The significance is, however, tempered by the evaluation: the central sign-language claim rests on a small private test set with no statistical support, and the fairness of the baselines' setup is open to question. The sign-stitching contribution is only shown qualitatively. These issues make the headline contribution currently unverified rather than clearly validated.
major comments (4)
- [Sec. 4.2, Table 3] The sign-language comparison is based on a private capture with 323 training frames and 24 test frames, and no confidence intervals, repeated runs, or significance tests are reported. The PSNR margin over EVA is 0.58 dB (33.769 vs. 33.193), which is within the run-to-run variation typical of Gaussian-splatting optimization; the abstract's claim of "significantly outperform" is therefore not supported by the evidence presented. The authors should either release the test data and fits or report variance over multiple optimization runs and a hypothesis test.
- [Sec. 4.2, Table 3] The statement "We estimated SMPL-X and camera parameters using our proposed single-view fitting framework" introduces a fairness risk for all baselines. Since every compared method is mesh-anchored, providing them with meshes produced by the authors' pipeline (OSX, HaMer, and the authors' 2D reprojection fitting) can differentially favor SignSplat, whose displacement limits and regularization were designed around those fits. The authors should either use each method's native fitting procedure, or demonstrate that the relative rankings are stable across fitting variations.
- [Sec. 4.4, Table 4] The ablation table is internally contradictory: removing adaptive densification yields higher PSNR (24.8416 vs. 24.6137) while the text states that "integrating all changes yields the best accuracy"; only SSIM and LPIPS actually improve with densification. Furthermore, the PSNR values in Table 4 (around 24.6) are far below the 35.47 reported for the same NeuMan dataset in Table 1. The dataset, subset, or evaluation protocol for Table 4 must be clarified, and the contribution of adaptive control needs to be presented honestly.
- [Sec. 4.5, Fig. 5] The sign-stitching contribution (listed as contribution 3 in Sec. 1.2) is evaluated only with a single qualitative comparison against a skeleton-based GAN approach. No quantitative metric, user study, or coverage of failure cases is provided, so the claimed "improved smoothness and continuity" is not established. A quantitative comparison (e.g., pose-error or image metrics, or at least a controlled perceptual study) is necessary to support this contribution.
minor comments (8)
- [Sec. 4.2, Tables 1-2] The benchmark tables report no variance or repeated runs; some gaps are small (e.g., sequence 00087 in Table 2, 32.17 vs. 32.01 dB), so the authors should report standard deviations or at least clarify whether the differences are stable across runs.
- [Sec. 3.3, Eq. (3)] The shared 1D CNN architecture and the exact input features (xc and φ(xc,S)) are not described in enough detail to reproduce; please provide the network depth, channel sizes, and how the canonical/observation coordinates are concatenated.
- [Sec. 3.4, Eq. (5)] The criteria for selecting splats to densify or prune (view-space gradient magnitude, opacity, and scale thresholds) are not given; even a reference to the default hyperparameters of 3DGS would help reproducibility.
- [Sec. 3.5, Eq. (6)] The regularization strength (weight) for the variance term is not reported, and the interaction between the "neighborhood radius" and the per-part (hand/face/body) constraints is not specified quantitatively.
- [Sec. 3.1] For the single-view case, the paper says "approximated camera parameters" are used in the 2D reprojection fitting, but does not say how the focal length or principal point are initialized; since NeuMan sequences are monocular, this detail matters for reproducibility.
- [Sec. 3.5] The sentence "3DGS [49] apply as-isometric-as-possible regularization" appears to cite the wrong reference; [49] is 3DGS-Avatar, whereas the original 3DGS is [24].
- [Sec. 4.2 and 4.3] The formatted dataset/method names contain spurious spaces, e.g., "N EUMAN" and "EV A"; these should be fixed to "NeuMan" and "EVA".
- [References [57] and [58]] References [57] and [58] both cite the same paper (Shen et al., X-Avatar) with slightly different formatting; this duplicate citation should be consolidated.
Circularity Check
No circularity: the rendering pipeline and its external benchmarks are self-contained; self-citations are domain background, not load-bearing.
full rationale
The paper's derivation chain is empirical and self-contained: 2D keypoints (MMPose) and segmentation (SAM) are lifted to SMPL-X via OSX with HaMer hand refinement and 2D reprojection fitting (Sec. 3.1); mesh-anchored Gaussians are optimized against L1/D-SSIM/LPIPS image losses (Sec. 4.1); adaptive densification/pruning and variance regularizers constrain this optimization (Secs. 3.4-3.5). No predicted quantity is defined as the parameter being optimized: the reported PSNR/SSIM/LPIPS values come from rendering held-out frames with frozen appearance, not from the fitted SMPL-X parameters themselves. The NeuMan protocol's test-time SMPL-X fitting is a standard pose-fitting step and does not fit the Gaussian appearance parameters that determine the metrics. The only self-citations are to the group's prior sign-language production work (e.g., references 52, 62, 63), used as a methodological starting point for gloss-based stitching and as a baseline; these are published, domain-specific, and not the source of the central rendering claim. The private six-view sign test is small and uses author-provided SMPL-X fits for all baselines, but this affects evaluation fairness and generalizability, not circularity: the comparison target is image appearance, and all methods receive the same fitted geometry. The conclusion's caveat that regularization effectiveness depends on input data is a stated limitation, not a circular step. No equation in the paper defines an output in terms of its own target, and no fitted scalar is renamed a prediction. Therefore there is no significant circularity.
Assumptions & free parameters
free parameters (5)
- densification thresholds (view-space gradient, opacity, scale)
- per-part displacement limits for mesh vertices
- neighborhood radius and variance weights for regularization
- spherical harmonics degree schedule
- sign stitching frame-count constant
assumptions (4)
- domain assumption SMPL-X provides an adequate parametric body model for sign language, including hands and face
- domain assumption Pre-trained OSX and HaMer estimators, plus 2D reprojection fitting, yield mesh parameters accurate enough for sign language
- ad hoc to paper A shared 1D CNN over canonical mesh points can represent appearance variations across poses
- domain assumption Image reconstruction losses (L1, D-SSIM, LPIPS) are sufficient to drive the optimization for sign language rendering
Cite this review
Pith. "Pith review of SignSplat: Rendering Sign Language via Gaussian Splatting." pith.science (2026). https://pith.science/paper/RJ4XPBIM
@misc{pith2026250502108,
author = {Pith},
title = {Pith review of: SignSplat: Rendering Sign Language via Gaussian Splatting},
year = {2026},
howpublished = {\url{https://pith.science/paper/RJ4XPBIM}},
note = {Machine review of arXiv:2505.02108}
}
read the original abstract
State-of-the-art approaches for conditional human body rendering via Gaussian splatting typically focus on simple body motions captured from many views. This is often in the context of dancing or walking. However, for more complex use cases, such as sign language, we care less about large body motion and more about subtle and complex motions of the hands and face. The problems of building high fidelity models are compounded by the complexity of capturing multi-view data of sign. The solution is to make better use of sequence data, ensuring that we can overcome the limited information from only a few views by exploiting temporal variability. Nevertheless, learning from sequence-level data requires extremely accurate and consistent model fitting to ensure that appearance is consistent across complex motions. We focus on how to achieve this, constraining mesh parameters to build an accurate Gaussian splatting framework from few views capable of modelling subtle human motion. We leverage regularization techniques on the Gaussian parameters to mitigate overfitting and rendering artifacts. Additionally, we propose a new adaptive control method to densify Gaussians and prune splat points on the mesh surface. To demonstrate the accuracy of our approach, we render novel sequences of sign language video, building on neural machine translation approaches to sign stitching. On benchmark datasets, our approach achieves state-of-the-art performance; and on highly articulated and complex sign language motion, we significantly outperform competing approaches.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 2 Pith papers
-
SignSparK: Efficient Multilingual Sign Language Production via Sparse Keyframe Learning
Sparse keyframe-conditioned Conditional Flow Matching produces fluid, articulate 3D sign language motion across four languages while enabling precise Keyframe-to-Pose editing.
-
Using Sign Language Production as Data Augmentation to enhance Sign Language Translation
Adding synthetic sign-language data produced by stitching, a GAN, or Gaussian splatting to the training set improves sign-language translation, with the largest gains for skeleton-pose models.
Reference graph
Works this paper leans on
-
[1]
Bbc-oxford british sign language dataset, 2021
Samuel Albanie, G ¨ul Varol, Liliane Momeni, Hannah Bull, Triantafyllos Afouras, Himel Chowdhury, Neil Fox, Bencie Woll, Rob Cooper, Andrew McParland, and Andrew Zisser- man. Bbc-oxford british sign language dataset, 2021. 5
work page 2021
-
[2]
Video based reconstruction of 3d people models
Thiemo Alldieck, Marcus Magnor, Weipeng Xu, Christian Theobalt, and Gerard Pons-Moll. Video based reconstruction of 3d people models. InIEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 8387–8397,
-
[3]
imghum: Implicit generative models of 3d human shape and articulated pose
Thiemo Alldieck, Hongyi Xu, and Cristian Sminchisescu. imghum: Implicit generative models of 3d human shape and articulated pose. 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 5441–5450, 2021. 1
work page 2021
-
[4]
Scape: shape completion and animation of people
Dragomir Anguelov, Praveen Srinivasan, Daphne Koller, Se- bastian Thrun, Jim Rodgers, and James Davis. Scape: shape completion and animation of people. ACM Trans. Graph., 24(3):408–416, 2005. 1
work page 2005
-
[5]
J.A. Bangham, S.J. Cox, R. Elliott, J.R.W. Glauert, I. Mar- shall, S. Rankov, and M. Wells. Virtual signing: capture, ani- mation, storage and transmission-an overview of the visicast project. In IEE Seminar on Speech and Language Processing for Disabled and Elderly People (Ref. No. 2000/025), pages 6/1–6/7, 2000. 3
work page 2000
-
[6]
Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P
Jonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P. Srinivasan. Mip-nerf: A multiscale representation for anti-aliasing neu- ral radiance fields. In 2021 IEEE/CVF International Confer- ence on Computer Vision (ICCV) , pages 5835–5844, 2021. 1
work page 2021
-
[7]
Animatable neural radiance fields from monocular rgb videos, 2021
Jianchuan Chen, Ying Zhang, Di Kang, Xuefei Zhe, Linchao Bao, Xu Jia, and Huchuan Lu. Animatable neural radiance fields from monocular rgb videos, 2021. 6
work page 2021
-
[8]
Vasileios Choutas, Georgios Pavlakos, Timo Bolkart, Dim- itrios Tzionas, and Michael J. Black. Monocular expressive body regression through body-driven attention. In European Conference on Computer Vision (ECCV), 2020. 1
work page 2020
Show all 70 references
-
[9]
Openmmlab pose estimation tool- box and benchmark
MMPose Contributors. Openmmlab pose estimation tool- box and benchmark. https://github.com/open- mmlab/mmpose, 2020. 3
2020
-
[10]
Cox, Michael Lincoln, Judy Tryggvason, Melanie Nakisa, Mark Wells, Marcus Tutt, and Sanja Abbott
Stephen J. Cox, Michael Lincoln, Judy Tryggvason, Melanie Nakisa, Mark Wells, Marcus Tutt, and Sanja Abbott. Tessa, a system to aid communication with deaf people. pages 205– 212, 2002. ASSETS 2002, Fifth International ACM SIG- CAPH Conference on Assistive Technologies ; Confe...
2002
-
[11]
The dicta-sign wiki: Enabling web communication for the deaf
E Efthimiou, SE Fotinea, T Hanke, J Glauert, R Bowden, A Braffort, C Collet, P Maragos, and F Lefebvre-Albaret. The dicta-sign wiki: Enabling web communication for the deaf. Lecture Notes in Computer Science (including sub- series Lecture Notes in Artificial Intelligence and L...
2012
-
[12]
Plenoxels: Radiance fields without neural networks
Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. Plenoxels: Radiance fields without neural networks. In CVPR, 2022. 1
2022
-
[13]
Vid2avatar: 3d avatar reconstruction from videos in the wild via self-supervised scene decomposition
Chen Guo, Tianjian Jiang, Xu Chen, Jie Song, and Otmar Hilliges. Vid2avatar: 3d avatar reconstruction from videos in the wild via self-supervised scene decomposition. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023. 2, 6
2023
-
[14]
Expres- sive gaussian human avatars from monocular rgb video
Hezhen Hu, Zhiwen Fan, Tianhao Wu, Yihan Xi, Seoyoung Lee, Georgios Pavlakos, and Zhangyang Wang. Expres- sive gaussian human avatars from monocular rgb video. In NeurIPS, 2024. 3, 7
2024
-
[15]
GaussianAvatar: Towards realistic human avatar modeling from a single video via animatable 3D gaussians
Liangxiao Hu, Hongwen Zhang, Yuxiang Zhang, Boyao Zhou, Boning Liu, Shengping Zhang, and Liqiang Nie. GaussianAvatar: Towards realistic human avatar modeling from a single video via animatable 3D gaussians. arXiv preprint arXiv:2312.02134, 2023. 6
2023 arXiv
-
[16]
Gaussianavatar: Towards realistic human avatar model- ing from a single video via animatable 3d gaussians
Liangxiao Hu, Hongwen Zhang, Yuxiang Zhang, Boyao Zhou, Boning Liu, Shengping Zhang, and Liqiang Nie. Gaussianavatar: Towards realistic human avatar model- ing from a single video via animatable 3d gaussians. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (C...
2024
-
[17]
GauHuman: Artic- ulated Gaussian Splatting from Monocular Human Videos
Shoukang Hu, Tao Hu, and Ziwei Liu. GauHuman: Artic- ulated Gaussian Splatting from Monocular Human Videos . In 2024 IEEE/CVF Conference on Computer Vision and Pat- tern Recognition (CVPR), pages 20418–20431, Los Alami- tos, CA, USA, 2024. IEEE Computer Society. 3
2024
-
[18]
In- stantavatar: Learning avatars from monocular video in 60 seconds
Tianjian Jiang, Xu Chen, Jie Song, and Otmar Hilliges. In- stantavatar: Learning avatars from monocular video in 60 seconds. 2023. 6
2023
-
[19]
In- stantAvatar: Learning avatars from monocular video in 60 seconds
Tianjian Jiang, Xu Chen, Jie Song, and Otmar Hilliges. In- stantAvatar: Learning avatars from monocular video in 60 seconds. In CVPR, 2023. 6
2023
-
[20]
NeuMan: Neural human radiance field from a single video
Wei Jiang, Kwang Moo Yi, Golnoosh Samei, Oncel Tuzel, and Anurag Ranjan. NeuMan: Neural human radiance field from a single video. In ECCV, 2022. 2, 6
2022
-
[21]
Hifi4g: High-fidelity human performance rendering via compact gaussian splatting
Yuheng Jiang, Zhehao Shen, Penghao Wang, Zhuo Su, Yu Hong, Yingliang Zhang, Jingyi Yu, and Lan Xu. Hifi4g: High-fidelity human performance rendering via compact gaussian splatting. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) , ...
2024
-
[22]
Black, David W
Angjoo Kanazawa, Michael J. Black, David W. Jacobs, and Jitendra Malik. End-to-end recovery of human shape and pose. In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 7122–7131, 2018. 1
2018
-
[23]
Neu- ral 3d mesh renderer
Hiroharu Kato, Yoshitaka Ushiku, and Tatsuya Harada. Neu- ral 3d mesh renderer. InProceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018. 1
2018
-
[24]
3d gaussian splatting for real-time radiance field rendering
Bernhard Kerbl, Georgios Kopanas, Thomas Leimk ¨uhler, and George Drettakis. 3d gaussian splatting for real-time radiance field rendering. ACM Transactions on Graphics, 42 (4), 2023. 1, 2, 5
2023
-
[25]
Mitra, and Helmut Pottmann
Martin Kilian, Niloy J. Mitra, and Helmut Pottmann. Geo- metric modeling in shape space. ACM Trans. Graph., 26(3): 64–es, 2007. 5
2007
-
[26]
Kingma and Jimmy Ba
Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. CoRR, abs/1412.6980, 2014. 6 9
2014 arXiv
-
[27]
Hugs: Human gaussian splats,
Muhammed Kocabas, Rick Chang, James Gabriel, Oncel Tuzel, and Anurag Ranjan. Hugs: Human gaussian splats,
-
[28]
Six challenges for neural machine translation
Philipp Koehn and Rebecca Knowles. Six challenges for neural machine translation. InProceedings of the First Work- shop on Neural Machine Translation, pages 28–39, Vancou- ver, 2017. Association for Computational Linguistics. 5
2017
-
[29]
Generalizable human gaussians for sparse view synthesis
Youngjoong Kwon, Baole Fang, Yixing Lu, Haoye Dong, Cheng Zhang, Francisco Vicente Carrasco, Albert Mosella- Montoro, Jianjin Xu, Shingo Takagi, Daeil Kim, Aayush Prakash, and Fernando De la Torre. Generalizable human gaussians for sparse view synthesis. In Computer Vision – E...
2024
-
[30]
Gart: Gaussian articulated template mod- els
Jiahui Lei, Yufu Wang, Georgios Pavlakos, Lingjie Liu, and Kostas Daniilidis. Gart: Gaussian articulated template mod- els. In Proceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition (CVPR), pages 19876– 19887, 2024. 2, 5
2024
-
[31]
Tianye Li, Timo Bolkart, Michael. J. Black, Hao Li, and Javier Romero. Learning a model of facial shape and ex- pression from 4D scans. ACM Transactions on Graphics, (Proc. SIGGRAPH Asia), 36(6):194:1–194:17, 2017. 4
2017
-
[32]
One-stage 3d whole-body mesh recovery with component aware transformer
Jing Lin, Ailing Zeng, Haoqian Wang, Lei Zhang, and Yu Li. One-stage 3d whole-body mesh recovery with component aware transformer. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 21159–21168, 2023. 3
2023
-
[33]
Neural actor: Neural free-view synthesis of human actors with pose con- trol
Lingjie Liu, Marc Habermann, Viktor Rudnev, Kripasindhu Sarkar, Jiatao Gu, and Christian Theobalt. Neural actor: Neural free-view synthesis of human actors with pose con- trol. ACM Trans. Graph.(ACM SIGGRAPH Asia), 2021. 2
2021
-
[34]
Gva: Reconstructing vivid 3d gaussian avatars from monocular videos
Xinqi Liu, Chenming Wu, Jialun Liu, Xing Liu, Chen Zhao, Haocheng Feng, Errui Ding, and Jingdong Wang. Gva: Reconstructing vivid 3d gaussian avatars from monocular videos. Arxiv, 2024. 2
2024
-
[35]
Matthew Loper, Naureen Mahmood, Javier Romero, Ger- ard Pons-Moll, and Michael J. Black. SMPL: A skinned multi-person linear model. ACM Trans. Graphics (Proc. SIGGRAPH Asia), 34(6):248:1–248:16, 2015. 1, 2, 4
2015
-
[36]
Srinivasan, Matthew Tancik, Jonathan T
Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. Nerf: Representing scenes as neural radiance fields for view syn- thesis. In ECCV, 2020. 1
2020
-
[37]
Expressive whole-body 3D gaussian avatar
Gyeongsik Moon, Takaaki Shiratori, and Shunsuke Saito. Expressive whole-body 3D gaussian avatar. In ECCV, 2024. 2, 4, 5, 6, 7
2024
-
[38]
Instant neural graphics primitives with a multires- olution hash encoding
Thomas M ¨uller, Alex Evans, Christoph Schied, and Alexan- der Keller. Instant neural graphics primitives with a multires- olution hash encoding. ACM Trans. Graph. , 41(4):102:1– 102:15, 2022. 1
2022
-
[39]
Ahmed A A Osman, Timo Bolkart, and Michael J. Black. STAR: A sparse trained articulated human body regressor. In European Conference on Computer Vision (ECCV), pages 598–613, 2020. 1
2020
-
[40]
Ash: Animatable gaussian splats for efficient and photoreal human rendering
Haokai Pang, Heming Zhu, Adam Kortylewski, Christian Theobalt, and Marc Habermann. Ash: Animatable gaussian splats for efficient and photoreal human rendering. In Pro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 1165–1175, 2024. 2
2024
-
[41]
Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin- Brualla, and Steven M
Keunhong Park, Utkarsh Sinha, Peter Hedman, Jonathan T. Barron, Sofien Bouaziz, Dan B Goldman, Ricardo Martin- Brualla, and Steven M. Seitz. Hypernerf: A higher- dimensional representation for topologically varying neural radiance fields. ACM Trans. Graph., 40(6), 2021. 2
2021
-
[42]
Georgios Pavlakos, Vasileios Choutas, Nima Ghorbani, Timo Bolkart, Ahmed A. A. Osman, Dimitrios Tzionas, and Michael J. Black. Expressive body capture: 3D hands, face, and body from a single image. In Proceedings IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), pa...
2019
-
[43]
Reconstruct- ing hands in 3D with transformers
Georgios Pavlakos, Dandan Shan, Ilija Radosavovic, Angjoo Kanazawa, David Fouhey, and Jitendra Malik. Reconstruct- ing hands in 3D with transformers. In CVPR, 2024. 3
2024
-
[44]
Ani- matable neural radiance fields for modeling dynamic human bodies
Sida Peng, Junting Dong, Qianqian Wang, Shangzhan Zhang, Qing Shuai, Xiaowei Zhou, and Hujun Bao. Ani- matable neural radiance fields for modeling dynamic human bodies. pages 14294–14303, 2021. 2
2021
-
[45]
Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans
Sida Peng, Yuanqing Zhang, Yinghao Xu, Qianqian Wang, Qing Shuai, Hujun Bao, and Xiaowei Zhou. Neural body: Implicit neural representations with structured latent codes for novel view synthesis of dynamic humans. In CVPR,
-
[46]
Henriques, and Christian Rupprecht
Lorenza Prospero, Abdullah Hamdi, Joao F. Henriques, and Christian Rupprecht. Gst: Precise 3d human body from a single image with gaussian splatting transformers, 2024. 2
2024
-
[47]
D-NeRF: Neural Radiance Fields for Dynamic Scenes
Albert Pumarola, Enric Corona, Gerard Pons-Moll, and Francesc Moreno-Noguer. D-NeRF: Neural Radiance Fields for Dynamic Scenes. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, 2020. 2
2020
-
[48]
Gaus- sianavatars: Photorealistic head avatars with rigged 3d gaus- sians
Shenhan Qian, Tobias Kirschstein, Liam Schoneveld, Davide Davoli, Simon Giebenhain, and Matthias Nießner. Gaus- sianavatars: Photorealistic head avatars with rigged 3d gaus- sians. arXiv preprint arXiv:2312.02069, 2023. 2, 4, 5
2023 arXiv
-
[49]
3DGS-Avatar: Animatable avatars via deformable 3D gaussian splatting
Zhiyin Qian, Shaofei Wang, Marko Mihajlovic, Andreas Geiger, and Siyu Tang. 3DGS-Avatar: Animatable avatars via deformable 3D gaussian splatting. In CVPR, 2024. 5, 6
2024
-
[50]
Sam 2: Segment anything in images and videos
Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu, Ronghang Hu, Chaitanya Ryali, Tengyu Ma, Haitham Khedr, Roman R¨adle, Chloe Rolland, Laura Gustafson, Eric Mintun, Junt- ing Pan, Kalyan Vasudev Alwala, Nicolas Carion, Chao- Yuan Wu, Ross Girshick, Piotr Doll´ar, and Christoph Feic...
2024 arXiv
-
[51]
Javier Romero, Dimitrios Tzionas, and Michael J. Black. Embodied hands: Modeling and capturing hands and bod- ies together. ACM Transactions on Graphics, (Proc. SIG- GRAPH Asia), 36(6), 2017. 4
2017
-
[52]
Everybody sign now: Translating spoken language to photo realistic sign language video
Ben Saunders, Necati Cihan Camgoz, and Richard Bowden. Everybody sign now: Translating spoken language to photo realistic sign language video. 2020. 3, 7
2020
-
[53]
Progressive transformers for end-to-end sign language pro- duction, 20200823 - 20200828
Ben Saunders, Necati Cihan Camg ¨oz, and Richard Bowden. Progressive transformers for end-to-end sign language pro- duction, 20200823 - 20200828. 1, 3 10
-
[54]
Structure-from-motion revisited
Johannes Lutz Sch ¨onberger and Jan-Michael Frahm. Structure-from-motion revisited. In Conference on Com- puter Vision and Pattern Recognition (CVPR), 2016. 3
2016
-
[55]
Pixelwise view selection for un- structured multi-view stereo
Johannes Lutz Sch ¨onberger, Enliang Zheng, Marc Pollefeys, and Jan-Michael Frahm. Pixelwise view selection for un- structured multi-view stereo. In European Conference on Computer Vision (ECCV), 2016. 3
2016
-
[56]
SplattingAvatar: Realistic Real-Time Human Avatars with Mesh-Embedded Gaussian Splatting
Zhijing Shao, Zhaolong Wang, Zhuang Li, Duotun Wang, Xiangru Lin, Yu Zhang, Mingming Fan, and Zeyu Wang. SplattingAvatar: Realistic Real-Time Human Avatars with Mesh-Embedded Gaussian Splatting. In Computer Vision and Pattern Recognition (CVPR), 2024. 2, 4, 5, 7
2024
-
[57]
X-avatar: Ex- pressive human avatars.Computer Vision and Pattern Recog- nition (CVPR), 2023
Kaiyue Shen, Chen Guo, Manuel Kaufmann, Juan Zarate, Julien Valentin, Jie Song, and Otmar Hilliges. X-avatar: Ex- pressive human avatars.Computer Vision and Pattern Recog- nition (CVPR), 2023. 6
2023
-
[58]
X- Avatar: Expressive human avatars
Kaiyue Shen, Chen Guo, Manuel Kaufmann, Juan Jose Zarate, Julien Valentin, Jie Song, and Otmar Hilliges. X- Avatar: Expressive human avatars. In CVPR, 2023. 2, 7
2023
-
[59]
A-nerf: articulated neural radiance fields for learn- ing human shape, appearance, and pose
Shih-Yang Su, Frank Yu, Michael Zollh ¨ofer, and Helge Rhodin. A-nerf: articulated neural radiance fields for learn- ing human shape, appearance, and pose. In Proceedings of the 35th International Conference on Neural Information Processing Systems, Red Hook, NY , USA, 2024. C...
2024
-
[60]
Youtube-sl-25: A large- scale, open-domain multilingual sign language parallel cor- pus
Garrett Tanzer and Biao Zhang. Youtube-sl-25: A large- scale, open-domain multilingual sign language parallel cor- pus. ArXiv, abs/2407.11144, 2024. 5
2024 arXiv
-
[61]
BodyNet: V olu- metric inference of 3D human body shapes
G ¨ul Varol, Duygu Ceylan, Bryan Russell, Jimei Yang, Ersin Yumer, Ivan Laptev, and Cordelia Schmid. BodyNet: V olu- metric inference of 3D human body shapes. In ECCV, 2018. 1
2018
-
[62]
Se- lect and reorder: A novel approach for neural sign lan- guage production
Harry Walsh, Ben Saunders, and Richard Bowden. Se- lect and reorder: A novel approach for neural sign lan- guage production. In Proceedings of the 2024 Joint Interna- tional Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) , pages 1...
2024
-
[63]
Sign stitching: A novel approach to sign language produc- tion, 20241125 - 20241128
Harry Thomas Walsh, Ben Saunders, and Richard Bowden. Sign stitching: A novel approach to sign language produc- tion, 20241125 - 20241128. 3
-
[64]
Neural rendering of humans in novel view and pose from monocular video, 2023
Tiantian Wang, Nikolaos Sarafianos, Ming-Hsuan Yang, and Tony Tung. Neural rendering of humans in novel view and pose from monocular video, 2023. 2
2023
-
[65]
Hu- manNeRF: Free-viewpoint rendering of moving people from monocular video
Chung-Yi Weng, Brian Curless, Pratul P Srinivasan, Jonathan T Barron, and Ira Kemelmacher-Shlizerman. Hu- manNeRF: Free-viewpoint rendering of moving people from monocular video. In CVPR, 2022. 6
2022
-
[66]
Srinivasan, Jonathan T
Chung-Yi Weng, Brian Curless, Pratul P. Srinivasan, Jonathan T. Barron, and Ira Kemelmacher-Shlizerman. Hu- manNeRF: Free-viewpoint rendering of moving people from monocular video. In Proceedings of the IEEE/CVF Confer- ence on Computer Vision and Pattern Recognition (CVPR) , ...
2022
-
[67]
H- nerf: neural radiance fields for rendering and temporal recon- struction of humans in motion
Hongyi Xu, Thiemo Alldieck, and Cristian Sminchisescu. H- nerf: neural radiance fields for rendering and temporal recon- struction of humans in motion. In Proceedings of the 35th International Conference on Neural Information Processing Systems, Red Hook, NY , USA, 2024. Curra...
2024
-
[68]
The unreasonable effectiveness of deep features as a perceptual metric
Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In CVPR, 2018. 6
2018
-
[69]
Bagautdinov, Shunsuke Saito, Michael Zollhofer, Justus Thies, and Javier Romero
Wojciech Zielonka, Timur M. Bagautdinov, Shunsuke Saito, Michael Zollhofer, Justus Thies, and Javier Romero. Driv- able 3d gaussian avatars. ArXiv, abs/2311.08581, 2023. 3 11
2023 arXiv
-
[2018]
CVPR Spotlight Paper. 2
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.