Pith. sign in

REVIEW 3 cited by

MVSFormer: Multi-View Stereo by Learning Robust Image Features and Temperature-based Depth

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.02541 v3 pith:PAZBAWOQ submitted 2022-08-04 cs.CV

classification cs.CV
keywords mvsformerfeaturefpnslearningvitsattentioncompetitiveefficient
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Feature representation learning is the key recipe for learning-based Multi-View Stereo (MVS). As the common feature extractor of learning-based MVS, vanilla Feature Pyramid Networks (FPNs) suffer from discouraged feature representations for reflection and texture-less areas, which limits the generalization of MVS. Even FPNs worked with pre-trained Convolutional Neural Networks (CNNs) fail to tackle these issues. On the other hand, Vision Transformers (ViTs) have achieved prominent success in many 2D vision tasks. Thus we ask whether ViTs can facilitate feature learning in MVS? In this paper, we propose a pre-trained ViT enhanced MVS network called MVSFormer, which can learn more reliable feature representations benefited by informative priors from ViT. The finetuned MVSFormer with hierarchical ViTs of efficient attention mechanisms can achieve prominent improvement based on FPNs. Besides, the alternative MVSFormer with frozen ViT weights is further proposed. This largely alleviates the training cost with competitive performance strengthened by the attention map from the self-distillation pre-training. MVSFormer can be generalized to various input resolutions with efficient multi-scale training strengthened by gradient accumulation. Moreover, we discuss the merits and drawbacks of classification and regression-based MVS methods, and further propose to unify them with a temperature-based strategy. MVSFormer achieves state-of-the-art performance on the DTU dataset. Particularly, MVSFormer ranks as Top-1 on both intermediate and advanced sets of the highly competitive Tanks-and-Temples leaderboard.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. D^2-4DGS: Dual-Depth Guided Sparse-Camera 4D Gaussian Splatting

    cs.CV 2026-08 conditional novelty 6.0 of 10

    D²-4DGS aligns monocular and multi-view depths, uses their agreement as verified anchors to guide densification, pruning, and depth supervision, and reports the best PSNR in all nine sparse-camera settings tested.

  2. 2K Retrofit: Entropy-Guided Efficient Sparse Refinement for High-Resolution 3D Geometry Prediction

    cs.CV 2026-03 conditional novelty 5.5 of 10

    Entropy-guided sparse refinement upgrades frozen low-resolution geometric foundation models to accurate 2K depth and pointmap outputs at a fraction of full-resolution cost.

  3. Online 3D Gaussian Splatting Modeling with Novel View Selection

    cs.CV 2025-08 conditional novelty 5.0 of 10

    During online Gaussian splatting SLAM, training extra on non-keyframes that view the most uncertain Gaussians improves model completeness over keyframe-only training.

Pith tools