REVIEW 4 cited by
3D-Consistent Human Avatars with Sparse Inputs via Gaussian Splatting and Contrastive Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Existing approaches for human avatar generation--both NeRF-based and 3D Gaussian Splatting (3DGS) based--struggle with maintaining 3D consistency and exhibit degraded detail reconstruction, particularly when training with sparse inputs. To address this challenge, we propose CHASE, a novel framework that achieves dense-input-level performance using only sparse inputs through two key innovations: cross-pose intrinsic 3D consistency supervision and 3D geometry contrastive learning. Building upon prior skeleton-driven approaches that combine rigid deformation with non-rigid cloth dynamics, we first establish baseline avatars with fundamental 3D consistency. To enhance 3D consistency under sparse inputs, we introduce a Dynamic Avatar Adjustment (DAA) module, which refines deformed Gaussians by leveraging similar poses from the training set. By minimizing the rendering discrepancy between adjusted Gaussians and reference poses, DAA provides additional supervision for avatar reconstruction. We further maintain global 3D consistency through a novel geometry-aware contrastive learning strategy. While designed for sparse inputs, CHASE surpasses state-of-the-art methods across both full and sparse settings on ZJU-MoCap and H36M datasets, demonstrating that our enhanced 3D consistency leads to superior rendering quality.
Forward citations
Cited by 4 Pith papers
-
MoGA: 3D Generative Avatar Prior for Monocular Gaussian Avatar Reconstruction
Fitting a learned 3D Gaussian avatar prior to six diffusion-hallucinated views reconstructs an animatable, high-fidelity avatar from a single image.
-
Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular Videos
Deblur-Avatar reconstructs sharp, animatable human avatars from motion-blurred monocular video by optimizing SMPL start and end poses and averaging rendered virtual frames inside 3D Gaussian Splatting.
-
Efficient Physics Simulation for 3D Scenes via MLLM-Guided Gaussian Splatting
PhysSplat uses an MLLM plus a trained distribution network to estimate material properties of 3D objects and simulate their motion with MPM, claiming realistic dynamics in about two minutes.
-
SMAP: Self-supervised Motion Adaptation for Physically Plausible Humanoid Whole-body Control
SMAP uses a vector-quantized periodic autoencoder to adapt human motion into physically plausible humanoid motion, then distills an RL teacher policy into a student policy for whole-body control.
Discussion (0). Continue with ORCID to comment.