Pith. sign in

REVIEW 11 cited by

LHM: Large Animatable Human Reconstruction Model from a Single Image in Seconds

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.10625 v1 pith:Q75QZC3E submitted 2025-03-13 cs.CV cs.AI

LHM: Large Animatable Human Reconstruction Model from a Single Image in Seconds

classification cs.CV cs.AI
keywords humanreconstructionanimatablefeaturesimagelargemodelability
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Animatable 3D human reconstruction from a single image is a challenging problem due to the ambiguity in decoupling geometry, appearance, and deformation. Recent advances in 3D human reconstruction mainly focus on static human modeling, and the reliance of using synthetic 3D scans for training limits their generalization ability. Conversely, optimization-based video methods achieve higher fidelity but demand controlled capture conditions and computationally intensive refinement processes. Motivated by the emergence of large reconstruction models for efficient static reconstruction, we propose LHM (Large Animatable Human Reconstruction Model) to infer high-fidelity avatars represented as 3D Gaussian splatting in a feed-forward pass. Our model leverages a multimodal transformer architecture to effectively encode the human body positional features and image features with attention mechanism, enabling detailed preservation of clothing geometry and texture. To further boost the face identity preservation and fine detail recovery, we propose a head feature pyramid encoding scheme to aggregate multi-scale features of the head regions. Extensive experiments demonstrate that our LHM generates plausible animatable human in seconds without post-processing for face and hands, outperforming existing methods in both reconstruction accuracy and generalization ability.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy

    cs.CV 2026-06 unverdicted novelty 7.0

    TryOnCrafter is the first DiT-based framework for camera-controllable video virtual try-on via a renderable 4D try-on proxy distilled from 2D priors into 3DGS avatar animated with SMPL-X.

  2. HumANDiff: Articulated Noise Diffusion for Motion-Consistent Human Video Generation

    cs.CV 2026-04 unverdicted novelty 7.0

    HumANDiff improves motion consistency in human video generation by sampling diffusion noise on an articulated human body template and adding joint appearance-motion prediction plus a geometric consistency loss.

  3. Large-scale Codec Avatars: The Unreasonable Effectiveness of Large-scale Avatar Pretraining

    cs.CV 2026-04 unverdicted novelty 7.0

    Pretraining on 1M wild videos followed by post-training on curated data yields high-fidelity feedforward 3D avatars that generalize across identities, clothing, and lighting with emergent relightability and loose-garm...

  4. VRGaussianAvatar: Integrating 3D Gaussian Avatars into VR

    cs.CV 2026-02 conditional novelty 7.0

    VRGaussianAvatar enables real-time full-body 3D Gaussian Splatting avatars in VR from HMD tracking alone via inverse kinematics and binocular batching for efficient stereo rendering, outperforming mesh baselines in pe...

  5. SONG: A Photorealistic 3D Gaussian Simulation Platform for Benchmarking Social Navigation

    cs.RO 2026-07 conditional novelty 6.0

    SONG, a benchmark with 1,000 photorealistic Gaussian scenes, 500 animated human avatars, and 500 difficulty-graded episodes, finds current vision-based social navigation policies succeed below 22% in easy episodes and...

  6. PointSplat: Compact Gaussian Splatting via Human-Centric Prediction

    cs.CV 2026-06 unverdicted novelty 6.0

    PointSplat infers compact Gaussian splats directly in 3D space from input point sets via ray casting and Point-Image Transformer to reduce inter-view redundancy and improve novel-view quality for humans.

  7. TryOnCrafter: Unleashing Camera Trajectories for Realistic Video Virtual Try-on via a Renderable 4D Try-on Proxy

    cs.CV 2026-06 conditional novelty 6.0

    A 4D try-on proxy (3DGS avatar + SMPL-X + background points) anchors a DiT so virtual try-on videos can follow arbitrary camera trajectories with consistent garments and scene structure.

  8. High-Fidelity Single-Image Head Modeling with Industry-Grade Topology

    cs.CV 2026-05 unverdicted novelty 6.0

    A single-image head reconstruction method uses coarse-to-fine optimization with normal consistency, landmarks, and geometry-aware constraints on curvature and conformality to produce meshes with industry-grade topolog...

  9. Visually-grounded Humanoid Agents

    cs.CV 2026-04 unverdicted novelty 6.0

    A coupled world-agent framework uses 3D Gaussian reconstruction and first-person RGB-D perception with iterative planning to enable goal-directed, collision-avoiding humanoid behavior in novel reconstructed scenes.

  10. Generator-Refiner-Examiner: A Tri-Module Data Augmentation Framework for 3D Human Avatar Learning from Monocular Videos

    cs.CV 2026-05 unverdicted novelty 5.0

    TrioMan is a tri-module data augmentation framework using a Generator for pose/camera perturbations, a Refiner with one-step diffusion, and an Examiner with dual-branch attention to improve 3D avatar learning from mon...

  11. Human Interaction-Aware 3D Reconstruction from a Single Image

    cs.CV 2026-04 unverdicted novelty 5.0

    HUG3D uses group-instance multi-view diffusion and physics-based optimization to create physically plausible 3D reconstructions of interacting people from a single image.