REVIEW 5 cited by
AvatarCLIP: Zero-Shot Text-Driven Generation and Animation of 3D Avatars
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
3D avatar creation plays a crucial role in the digital age. However, the whole production process is prohibitively time-consuming and labor-intensive. To democratize this technology to a larger audience, we propose AvatarCLIP, a zero-shot text-driven framework for 3D avatar generation and animation. Unlike professional software that requires expert knowledge, AvatarCLIP empowers layman users to customize a 3D avatar with the desired shape and texture, and drive the avatar with the described motions using solely natural languages. Our key insight is to take advantage of the powerful vision-language model CLIP for supervising neural human generation, in terms of 3D geometry, texture and animation. Specifically, driven by natural language descriptions, we initialize 3D human geometry generation with a shape VAE network. Based on the generated 3D human shapes, a volume rendering model is utilized to further facilitate geometry sculpting and texture generation. Moreover, by leveraging the priors learned in the motion VAE, a CLIP-guided reference-based motion synthesis method is proposed for the animation of the generated 3D avatar. Extensive qualitative and quantitative experiments validate the effectiveness and generalizability of AvatarCLIP on a wide range of avatars. Remarkably, AvatarCLIP can generate unseen 3D avatars with novel animations, achieving superior zero-shot capability.
Forward citations
Cited by 5 Pith papers
-
InteractAnything: Zero-shot Human Object Interaction Synthesis via LLM Feedback and Object Affordance Parsing
InteractAnything combines LLM-based relation reasoning, inpainting-based object affordance parsing, and force-closure-inspired optimization to synthesize zero-shot 3D human-object interactions from text and an object mesh.
-
SplatPainter: Interactive Authoring of 3D Gaussians from 2D Edits via Test-Time Training
A test-time-trained feedforward model that propagates 2D edits onto 3D Gaussian attributes at interactive speeds.
-
Half-Physics: Enabling Kinematic 3D Human Model with Physical Interactions
Half physics converts kinematic SMPL-X poses into velocities that drive a physics engine, preserving the original motion when contact-free and giving physically correct responses when collisions occur.
-
TextMesh4D: Zero-shot Text-to-4D Mesh Generation
TextMesh4D generates text-conditioned dynamic meshes by combining a Jacobian Deformation Field, video score distillation, and a local-global semantic regularizer in a zero-shot pipeline.
-
Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation: A Way Forward
A survey paper reviews multimodal generative AI and autoregressive LLMs for text-driven human motion generation, with comparative tables of models, datasets, and metrics.
Discussion (0). Continue with ORCID to comment.