Pith. sign in

REVIEW 6 cited by

RodinHD: High-Fidelity 3D Avatar Generation with Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.06938 v2 pith:J5JEE6ZC submitted 2024-07-09 cs.CV

classification cs.CV
keywords avatarsdetailsportraitdecoderdiffusiongeneratehigh-fidelityimage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We present RodinHD, which can generate high-fidelity 3D avatars from a portrait image. Existing methods fail to capture intricate details such as hairstyles which we tackle in this paper. We first identify an overlooked problem of catastrophic forgetting that arises when fitting triplanes sequentially on many avatars, caused by the MLP decoder sharing scheme. To overcome this issue, we raise a novel data scheduling strategy and a weight consolidation regularization term, which improves the decoder's capability of rendering sharper details. Additionally, we optimize the guiding effect of the portrait image by computing a finer-grained hierarchical representation that captures rich 2D texture cues, and injecting them to the 3D diffusion model at multiple layers via cross-attention. When trained on 46K avatars with a noise schedule optimized for triplanes, the resulting model can generate 3D avatars with notably better details than previous methods and can generalize to in-the-wild portrait input.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Synthetic Prior for Few-Shot Drivable Head Avatar Inversion

    cs.CV 2025-01 conditional novelty 7.0 of 10

    A few-shot inversion method that fine-tunes a synthetic-data-trained VQ-VAE Gaussian head prior to reconstruct a drivable avatar from three input images, improving novel view and expression synthesis over SOTA methods.

  2. Gaussian Variation Field Diffusion for High-fidelity Video-to-4D Synthesis

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A video-to-4D model that encodes mesh animations into compact Gaussian variation latents and diffuses them conditioned on the video and a canonical Gaussian splat.

  3. Arc2Avatar: Generating Expressive 3D Avatars from a Single Image via ID Guidance

    cs.CV 2025-01 conditional novelty 6.0 of 10

    Arc2Avatar generates expressive 3D head avatars from a single image by distilling a LoRA-fine-tuned Arc2Face model into 3D Gaussian splats anchored to a FLAME mesh, enabling blendshape expressions.

  4. CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    CAP4D combines a morphable multi-view diffusion model with 3D Gaussian splatting to build animatable 4D head avatars from 1 to 100 reference images, claiming state-of-the-art results.

  5. Towards High-Fidelity 3D Portrait Generation with Rich Details by Cross-View Prior-Aware Diffusion

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A three-stage pipeline with hybrid multi-view conditioning and anchor-noise resampling produces 3D portraits with sharper textures than prior single-image baselines.

  6. Generative AI for Character Animation: A Comprehensive Survey of Techniques, Applications, and Future Directions

    cs.CV 2025-04 conditional novelty 3.0 of 10

    A comprehensive survey that unifies generative AI techniques for character animation across facial, gesture, motion, and 3D asset generation, with a shared taxonomy and resource list.

Pith tools