Pith. sign in

REVIEW 3 cited by

HAMSt3R: Human-Aware Multi-view Stereo 3D Reconstruction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2508.16433 v1 pith:42CXOTFN submitted 2025-08-22 cs.CV

HAMSt3R: Human-Aware Multi-view Stereo 3D Reconstruction

classification cs.CV
keywords humanscenemulti-viewreconstructiondensegeometryhamst3rhuman-centric
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Recovering the 3D geometry of a scene from a sparse set of uncalibrated images is a long-standing problem in computer vision. While recent learning-based approaches such as DUSt3R and MASt3R have demonstrated impressive results by directly predicting dense scene geometry, they are primarily trained on outdoor scenes with static environments and struggle to handle human-centric scenarios. In this work, we introduce HAMSt3R, an extension of MASt3R for joint human and scene 3D reconstruction from sparse, uncalibrated multi-view images. First, we exploit DUNE, a strong image encoder obtained by distilling, among others, the encoders from MASt3R and from a state-of-the-art Human Mesh Recovery (HMR) model, multi-HMR, for a better understanding of scene geometry and human bodies. Our method then incorporates additional network heads to segment people, estimate dense correspondences via DensePose, and predict depth in human-centric environments, enabling a more comprehensive 3D reconstruction. By leveraging the outputs of our different heads, HAMSt3R produces a dense point map enriched with human semantic information in 3D. Unlike existing methods that rely on complex optimization pipelines, our approach is fully feed-forward and efficient, making it suitable for real-world applications. We evaluate our model on EgoHumans and EgoExo4D, two challenging benchmarks con taining diverse human-centric scenarios. Additionally, we validate its generalization to traditional multi-view stereo and multi-view pose regression tasks. Our results demonstrate that our method can reconstruct humans effectively while preserving strong performance in general 3D reconstruction tasks, bridging the gap between human and scene understanding in 3D vision.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. UniCon3R: Unified Contact-aware 4D Human-Scene Reconstruction from Monocular Video

    cs.CV 2026-04 unverdicted novelty 7.0

    UniCon3R reconstructs 3D human motion and scene geometry from monocular video by inferring human-scene contacts and using them as an active corrective prior to enforce physical plausibility.

  2. UniCon3R: Unified Contact-aware 4D Human-Scene Reconstruction from Monocular Video

    cs.CV 2026-04 unverdicted novelty 6.0

    UniCon3R infers 4D contact from pose and geometry to correct human mesh and scene alignment in monocular video, yielding more physically plausible joint reconstructions than prior feed-forward methods.

  3. Understanding the Impact of Geometric Foundation Models on Vision-Language-Action Models

    cs.CV 2026-05 unverdicted novelty 3.0

    The paper quantifies the geometric gap in current VLAs via linear probing and compares three architectures for injecting geometry from GFMs while analyzing impacts of data, cameras, and reconstruction quality.