Pith. sign in

REVIEW 7 cited by

LHM: Large Animatable Human Reconstruction Model from a Single Image in Seconds

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.10625 v1 pith:Q75QZC3E submitted 2025-03-13 cs.CV cs.AI

classification cs.CVcs.AI
keywords humanreconstructionanimatablefeaturesimagelargemodelability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Animatable 3D human reconstruction from a single image is a challenging problem due to the ambiguity in decoupling geometry, appearance, and deformation. Recent advances in 3D human reconstruction mainly focus on static human modeling, and the reliance of using synthetic 3D scans for training limits their generalization ability. Conversely, optimization-based video methods achieve higher fidelity but demand controlled capture conditions and computationally intensive refinement processes. Motivated by the emergence of large reconstruction models for efficient static reconstruction, we propose LHM (Large Animatable Human Reconstruction Model) to infer high-fidelity avatars represented as 3D Gaussian splatting in a feed-forward pass. Our model leverages a multimodal transformer architecture to effectively encode the human body positional features and image features with attention mechanism, enabling detailed preservation of clothing geometry and texture. To further boost the face identity preservation and fine detail recovery, we propose a head feature pyramid encoding scheme to aggregate multi-scale features of the head regions. Extensive experiments demonstrate that our LHM generates plausible animatable human in seconds without post-processing for face and hands, outperforming existing methods in both reconstruction accuracy and generalization ability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    TryOnCrafter is the first DiT-based framework for camera-controllable video virtual try-on via a renderable 4D try-on proxy distilled from 2D priors into 3DGS avatar animated with SMPL-X.

  2. SONG: A Photorealistic 3D Gaussian Simulation Platform for Benchmarking Social Navigation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    SONG, a benchmark with 1,000 photorealistic Gaussian scenes, 500 animated human avatars, and 500 difficulty-graded episodes, finds current vision-based social navigation policies succeed below 22% in easy episodes and...

  3. Parametric Gaussian Human Model: Generalizable Prior for Efficient and Realistic Human Avatar Modeling

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A pretrained parametric Gaussian prior with a UV-aligned latent identity map and disentangled multi-head U-Net creates animatable human avatars from monocular video in about 20 minutes per subject.

  4. AdaHuman: Animatable Detailed 3D Human Generation with Compositional Multiview Diffusion

    cs.CV 2025-05 conditional novelty 6.0 of 10

    AdaHuman generates high-fidelity, reposable 3D avatars from a single image by adding pose conditioning and local body-part refinement to multiview diffusion with Gaussian splat reconstruction.

  5. AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A diffusion model animates a character into arbitrary dynamic backgrounds by conditioning on a rendered 3D-avatar video, reframing open-domain animation as a restoration problem.

  6. Impact-driven Context Filtering For Cross-file Code Completion

    cs.SE 2025-08 unverdicted novelty 5.0 of 10

    The manuscript's abstract claims a new code-completion filtering method, yet the body contains an unrelated 3D animation paper, leaving the claimed work unverifiable.

  7. Multi-View Face and Gesture Animation with Dynamic Gaussians

    cs.CV 2026-08 conditional novelty 4.0 of 10

    Combining separate face and hand models with a parametric body and Gaussian splatting enables multi-view-consistent upper-body avatars that can be re-animated with new expressions and gestures.

Pith tools