Pith. sign in

REVIEW 9 cited by

DUSt3R: Geometric 3D Vision Made Easy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.14132 v3 pith:FXQ3Y7FP submitted 2023-12-21 cs.CV

classification cs.CV
keywords cameradust3rreconstructiontasksvisiondeptheasyestimation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-view stereo reconstruction (MVS) in the wild requires to first estimate the camera parameters e.g. intrinsic and extrinsic parameters. These are usually tedious and cumbersome to obtain, yet they are mandatory to triangulate corresponding pixels in 3D space, which is the core of all best performing MVS algorithms. In this work, we take an opposite stance and introduce DUSt3R, a radically novel paradigm for Dense and Unconstrained Stereo 3D Reconstruction of arbitrary image collections, i.e. operating without prior information about camera calibration nor viewpoint poses. We cast the pairwise reconstruction problem as a regression of pointmaps, relaxing the hard constraints of usual projective camera models. We show that this formulation smoothly unifies the monocular and binocular reconstruction cases. In the case where more than two images are provided, we further propose a simple yet effective global alignment strategy that expresses all pairwise pointmaps in a common reference frame. We base our network architecture on standard Transformer encoders and decoders, allowing us to leverage powerful pretrained models. Our formulation directly provides a 3D model of the scene as well as depth information, but interestingly, we can seamlessly recover from it, pixel matches, relative and absolute camera. Exhaustive experiments on all these tasks showcase that the proposed DUSt3R can unify various 3D vision tasks and set new SoTAs on monocular/multi-view depth estimation as well as relative pose estimation. In summary, DUSt3R makes many geometric 3D vision tasks easy.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Princeton365: A Diverse Dataset with Accurate Camera Pose

    cs.CV 2025-06 conditional novelty 7.0 of 10

    Princeton365 is a 365-video SLAM/NVS benchmark with board-calibrated millimeter-accurate 6-DoF poses, a new scale-aware optical-flow error metric, and an NVS benchmark of fully non-Lambertian 360-degree scans.

  2. Splat and Replace: 3D Reconstruction with Repetitive Elements

    cs.GR 2025-06 conditional novelty 7.0 of 10

    Repetitive objects in 3D scenes are registered into a shared Gaussian representation that propagates well-observed geometry and appearance to poorly observed instances, improving rendered novel views.

  3. Future Rendering $\neq$ Future Surface: A Benchmark and Dataset for Dynamic Surface Reconstruction Beyond the Observed Window

    cs.CV 2026-07 conditional novelty 6.0 of 10

    FutureSurf, a new benchmark for held-out future surface reconstruction, shows deformation-MLP methods leave a 2-6.6× future-surface gap while rendering quality stays flat.

  4. PACE: Polar Axis-Conditioned Estimation for PairUAV Relative Localization

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A shared image-pair network beats a single-head baseline by giving heading and range their own decoder readouts—PACE's raw model scores 0.002460 on the PairUAV hidden test.

  5. X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras

    cs.CV 2026-07 unverdicted novelty 6.0 of 10

    X-Lens fuses arbitrary calibrated fisheye and pinhole views into real-time metric depth at 41 FPS with a 0.04B-parameter model and a new 266K-frame synthetic dataset.

  6. VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment

    cs.CV 2026-03 conditional novelty 5.5 of 10

    A VGGT-based pointwise reprojection reward with geometry-aware sampling improves video geometric consistency via SFT/DPO and causal test-time search.

  7. Quo Vadis, World Modeling?

    cs.CV 2026-08 conditional novelty 5.0 of 10

    An agent-centric reframing of world modeling, replacing physical state prediction with 'information transitions' organized into six proxy functions and three empowerment levels.

  8. GLidE-SLAM: GL-Accelerated Indirect-Direct Embedded SLAM

    cs.RO 2026-07 conditional novelty 5.0 of 10

    GLidE-SLAM moves pose-only photometric tracking to OpenGL ES compute shaders, reporting up to 9x faster frame rates than ORB-SLAM2 on embedded platforms with comparable ATE on TUM and EuRoC sequences.

  9. Improving Video Diffusion Transformer Training by Multi-Feature Fusion and Alignment from Self-Supervised Vision Encoders

    cs.CV 2025-09 conditional novelty 5.0 of 10

    Matching video diffusion transformer tokens to concatenated DINOv2 and SAM2 features improves FVD/FID and speeds convergence, e.g., 400K-step fusion beats 1M-step baseline on UCF-101.

Pith tools