REVIEW 6 cited by
RodinHD: High-Fidelity 3D Avatar Generation with Diffusion Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
We present RodinHD, which can generate high-fidelity 3D avatars from a portrait image. Existing methods fail to capture intricate details such as hairstyles which we tackle in this paper. We first identify an overlooked problem of catastrophic forgetting that arises when fitting triplanes sequentially on many avatars, caused by the MLP decoder sharing scheme. To overcome this issue, we raise a novel data scheduling strategy and a weight consolidation regularization term, which improves the decoder's capability of rendering sharper details. Additionally, we optimize the guiding effect of the portrait image by computing a finer-grained hierarchical representation that captures rich 2D texture cues, and injecting them to the 3D diffusion model at multiple layers via cross-attention. When trained on 46K avatars with a noise schedule optimized for triplanes, the resulting model can generate 3D avatars with notably better details than previous methods and can generalize to in-the-wild portrait input.
Forward citations
Cited by 6 Pith papers
-
Synthetic Prior for Few-Shot Drivable Head Avatar Inversion
A few-shot inversion method that fine-tunes a synthetic-data-trained VQ-VAE Gaussian head prior to reconstruct a drivable avatar from three input images, improving novel view and expression synthesis over SOTA methods.
-
Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis
A video-to-4D model that encodes mesh animations into compact Gaussian variation latents and diffuses them conditioned on the video and a canonical Gaussian splat.
-
Arc2Avatar: Generating Expressive 3D Avatars from a Single Image via ID Guidance
Arc2Avatar generates expressive 3D head avatars from a single image by distilling a LoRA-fine-tuned Arc2Face model into 3D Gaussian splats anchored to a FLAME mesh, enabling blendshape expressions.
-
CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion Models
CAP4D combines a morphable multi-view diffusion model with 3D Gaussian splatting to build animatable 4D head avatars from 1 to 100 reference images, claiming state-of-the-art results.
-
Towards High-Fidelity 3D Portrait Generation with Rich Details by Cross-View Prior-Aware Diffusion
A three-stage pipeline with hybrid multi-view conditioning and anchor-noise resampling produces 3D portraits with sharper textures than prior single-image baselines.
-
Generative AI for Character Animation: A Comprehensive Survey of Techniques, Applications, and Future Directions
A comprehensive survey that unifies generative AI techniques for character animation across facial, gesture, motion, and 3D asset generation, with a shared taxonomy and resource list.
Discussion (0). Continue with ORCID to comment.