Pith. sign in

REVIEW 5 cited by

AvatarCLIP: Zero-Shot Text-Driven Generation and Animation of 3D Avatars

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.08535 v1 pith:YDCVNS5V submitted 2022-05-17 cs.CV

classification cs.CV
keywords avataravatarclipgenerationanimationavatarsgeometryhumantexture
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

3D avatar creation plays a crucial role in the digital age. However, the whole production process is prohibitively time-consuming and labor-intensive. To democratize this technology to a larger audience, we propose AvatarCLIP, a zero-shot text-driven framework for 3D avatar generation and animation. Unlike professional software that requires expert knowledge, AvatarCLIP empowers layman users to customize a 3D avatar with the desired shape and texture, and drive the avatar with the described motions using solely natural languages. Our key insight is to take advantage of the powerful vision-language model CLIP for supervising neural human generation, in terms of 3D geometry, texture and animation. Specifically, driven by natural language descriptions, we initialize 3D human geometry generation with a shape VAE network. Based on the generated 3D human shapes, a volume rendering model is utilized to further facilitate geometry sculpting and texture generation. Moreover, by leveraging the priors learned in the motion VAE, a CLIP-guided reference-based motion synthesis method is proposed for the animation of the generated 3D avatar. Extensive qualitative and quantitative experiments validate the effectiveness and generalizability of AvatarCLIP on a wide range of avatars. Remarkably, AvatarCLIP can generate unseen 3D avatars with novel animations, achieving superior zero-shot capability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing

    cs.CV 2025-05 conditional novelty 6.0 of 10

    InteractAnything combines LLM-based relation reasoning, inpainting-based object affordance parsing, and force-closure-inspired optimization to synthesize zero-shot 3D human-object interactions from text and an object mesh.

  2. SplatPainter: Interactive Authoring of 3D Gaussians from 2D Edits via Test-Time Training

    cs.CV 2025-12 conditional novelty 5.0 of 10

    A test-time-trained feedforward model that propagates 2D edits onto 3D Gaussian attributes at interactive speeds.

  3. Half-Physics: Enabling Kinematic 3D Human Model with Physical Interactions

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Half physics converts kinematic SMPL-X poses into velocities that drive a physics engine, preserving the original motion when contact-free and giving physically correct responses when collisions occur.

  4. TextMesh4D: Zero-shot Text-to-4D Mesh Generation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    TextMesh4D generates text-conditioned dynamic meshes by combining a Jacobian Deformation Field, video score distillation, and a local-global semantic regularizer in a zero-shot pipeline.

  5. Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward

    cs.CV 2025-05 conditional novelty 3.0 of 10

    A survey paper reviews multimodal generative AI and autoregressive LLMs for text-driven human motion generation, with comparative tables of models, datasets, and metrics.

Pith tools