Pith. sign in

REVIEW 3 major objections 5 minor 74 references

SPFSplatV2: Efficient Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read SPFSplatV2 claims that a feed-forward 3D Gaussian splatting model trained without ground-truth camera poses—using only rendering and reprojection losses—matches or beats pose-required and supervised pose-free methods on novel view…

desk verdict Solid incremental extension of SPFSplat with strong empirical results, but the 'self-supervised' claim outruns the evidence: most performance comes from supervised MASt3R/VGGT initialization. read the letter →

arxiv 2509.17246 v2 pith:FYTW5EMR submitted 2025-09-21 cs.CV

classification cs.CV
keywords 3DGaussiansplattingself-supervisedlearningpose-freenovelviewsynthesissparse-viewreconstructioncameraposeestimationmaskedattentionreprojectionlossfeed-forward
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that ground-truth camera poses are not needed to train a feed-forward 3D reconstruction model for novel view synthesis. It proposes a single network that, given a few unposed images, predicts both 3D Gaussian primitives and camera poses in one shared canonical space, and is trained by rendering target views with the estimated poses plus a pixel-reprojection loss. The authors claim this self-supervised, pose-free pipeline is stable enough to beat pose-required methods such as pixelSplat and MVSplat, supervised pose-free methods such as NoPoSplat, and earlier self-supervised pose-free methods, both in-domain and out-of-domain, while also producing competitive relative camera poses. If true, the result matters because it removes the main bottleneck to training on large, diverse, unannotated video and photo collections.

What carries the argument

The load-bearing mechanism is a masked multi-view decoder with learnable pose tokens. During training, target images enter the same forward pass as context images, but cross-attention masks prevent context tokens from seeing target tokens, so the reconstructed Gaussians depend only on the context views; target tokens attend to all views so the pose head can estimate the target pose. The second mechanism is the reprojection loss, which projects each predicted Gaussian center back into its own image with the estimated pose and penalizes pixel displacement, acting as a differentiable geometric constraint without ground-truth poses.

What would settle it

Train the full pipeline from randomly initialized weights with the DUSt3R point-cloud distillation warm-up deleted, leaving only the rendering and reprojection losses; if the model then collapses to near-identity poses and blurred renderings (pose AUC near zero), the claim that ground-truth poses are unnecessary for this training objective is falsified for the from-scratch regime.

Watch

Extended reading notes

Core claim

The central discovery is that joint optimization of 3D Gaussians and camera poses, with a shared backbone and a masked attention decoder, makes self-supervised pose-free training practical for sparse views. By letting context tokens attend only to context tokens while target tokens attend to everything, the model prevents target-view information from leaking into the reconstruction, yet still uses the target image to estimate its pose for the rendering loss. A pixel-wise reprojection loss on the predicted Gaussian centers and context poses supplies the geometric constraint that keeps the Gaussians pixel-aligned and the training stable. The paper reports that the resulting SPFSplatV2 and the larger VGGT-based SPFSplatV2-L outperform prior work across overlap regimes on RealEstate10K and ACID, generalize zero-shot to unseen datasets, and surpass many geometric-supervision methods in relative pose estimation.

Load-bearing premise

The claim that training is pose-free rests on the assumption that initializing from MASt3R or VGGT pretrained weights—models trained with pose or depth supervision—does not count as using ground-truth poses; Table XI shows most of the performance disappears under random initialization.

Editorial extensions

If this is right

  • Training no longer requires SfM pose preprocessing, so the same framework can be trained directly on larger and more diverse unposed video and photo collections; the paper shows that adding DL3DV to training improves pose estimation on multiple benchmarks.
  • The training paradigm transfers across reconstruction architectures: the same losses and masked-attention scheme work on a MASt3R-style backbone and a VGGT-style backbone, suggesting the self-supervised objective, not a specific network design, is what removes the pose dependency.
  • Because estimated poses and reconstructed Gaussians are jointly optimized, rendering with the predicted poses is competitive with methods that use evaluation-time pose alignment, which decouples rendering quality from pose accuracy.
  • A single model with multi-view dropout handles different numbers and spatial arrangements of context views, so practitioners do not need to train separate models for two-view versus many-view inputs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the strongest reading of the results is not that poses are unnecessary, but that the rendering-plus-reprojection objective can replace explicit pose labels during fine-tuning, provided the network starts from weights that already encode geometric priors.
  • Editorial inference: the masked-attention design may be the key reusable idea beyond splatting—any feed-forward reconstruction task that needs target-view information for supervision but not for the representation could use the same attention mask.
  • Editorial inference: a direct testable extension is to train the same pipeline on an unposed dataset that has no SfM poses at all, using only the photometric and reprojection losses, and measure whether the pose-estimation gains in Table VIII persist or saturate.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces SPFSplatV2, a feed-forward framework for novel view synthesis from unposed sparse multi-view images using 3D Gaussian splatting. The method uses a shared ViT encoder and a masked multi-view decoder to jointly predict pixel-aligned Gaussians in a canonical space and relative camera poses, trained with a rendering loss and a reprojection loss, without any ground-truth pose supervision in the objective. Two variants are presented: SPFSplatV2 built on a MASt3R-style architecture and SPFSplatV2-L built on a VGGT-style architecture. Experiments on RealEstate10K, ACID, DTU, DL3DV, and ScanNet++ report state-of-the-art results in novel view synthesis, cross-dataset generalization, and relative pose estimation, together with efficiency and ablation studies, including an ablation that compares against training with ground-truth poses and an initialization analysis.

Significance. The reported results, if they hold, would be a meaningful advance for pose-free generalizable 3D reconstruction: the masked-attention single-branch design is a clean solution to the target-leakage problem, the reprojection loss is shown to be essential for stability, and the framework's compatibility with two different reconstruction backbones suggests the training paradigm is portable. The paper is also commendably transparent in running the ground-truth-pose and initialization ablations in Tables X and XI, which most pose-free papers would omit. The main significance caveat is that the method delivers its strongest performance when initialized from MASt3R or VGGT, models that were themselves trained with pose or depth supervision; the contribution is therefore best understood as a pose-free fine-tuning objective for strong geometric backbones, not as evidence that such geometry can be learned from image-level losses alone.

major comments (3)
  1. [§IV-C, §IV-D, §V, Table XI] The abstract and Section IV-C state that the method works 'despite the absence of pose supervision' and 'despite no geometry priors during training,' but Section V concedes that the method 'benefits from the priors provided by supervised models such as MASt3R and VGGT,' and Table XI shows that random initialization with a DUSt3R distillation warm-up drops RE10K PSNR from 26.157 to 22.394. In addition, Section IV-D states that training solely with a photometric loss without ground-truth geometric supervision makes it difficult to learn Gaussians in canonical space, and that the warm-up is 'essential.' These statements put the 'no pose supervision' claim in a materially narrower form: the method fine-tunes pose- or depth-supervised reconstruction backbones with a pose-free objective. I recommend that the claims in the abstract, Section IV-C, and the conclusion be reworded to state exactly this, and that the contribution be framed as a pose-free fine-tuning paradigm rather than purely image-supervised training from scratch.
  2. [§IV-D, Table XI] The random-initialization ablation is not a clean test of fully self-supervised training because the reported 'Random' row uses a warm-up phase with a DUSt3R point-cloud distillation loss for the first 10,000 steps. As the text itself says, this supervision is essential; without it the photometric-only objective fails to learn canonical Gaussians. The table and surrounding text should therefore label this setting as 'random init + DUSt3R distillation warm-up,' and the limitation should be acknowledged in the conclusion. As reported, the row conflates two effects: removal of pretrained weights and addition of geometric distillation, so it does not by itself establish what a purely image-supervised run would achieve.
  3. [§IV-C, Tables I–IV] All quantitative results are reported as single numbers with no standard errors, confidence intervals, or repeated-seed information. Several headline comparisons are very close: for example, Table I lists SPFSplatV2* at 26.157 dB versus SPFSplat* at 25.845 dB, and Table II lists SPFSplatV2* at 26.809 dB versus SPFSplat* at 26.796 dB. Given that the central claim is state-of-the-art performance, the paper should at least report results over multiple seeds with variance, or otherwise justify that the differences are not within run-to-run noise. This is particularly important for the zero-shot cross-dataset tables, which are reported on what appear to be single evaluation passes.
minor comments (5)
  1. [Table I] In the Splatt3R row, the average SSIM and LPIPS values (0.337 and 0.596) appear to be swapped relative to the overlap-specific columns; please check and correct.
  2. [§IV-C, Relative Pose Estimation] The sentence 'Despite no geometry priors during training' is inconsistent with the initialization discussion in Section IV-D and Section V; please revise it to avoid contradicting the paper's own limitation statement.
  3. [§IV-C, Cross-Dataset Generalization] The statement 'SPFSplatV2-L consistently outperforms SPFSplatV2 both with and without pose alignment' is not supported by Table III on ACID with pose alignment, where SPFSplatV2* reports PSNR 26.802 versus SPFSplatV2-L* 26.680; please qualify the claim by metric or dataset.
  4. [§II, Related Work] The phrase 'according to the architecuture of CroCo' contains a typo: 'architecuture' should be 'architecture.'
  5. [§IV-B, Implementation Details] Please state explicitly whether the multi-view dropout strategy is applied in all main experiments and ablations, since it affects the comparability of the reported numbers.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the losses and benchmarks are external and measured; the supervised-pretraining dependency narrows the 'self-supervised' claim but does not reduce any result to its inputs.

full rationale

SPFSplatV2 is an empirical systems paper. The training objectives are the rendering loss (Eq. 13) and the reprojection loss (Eq. 14), both of which optimize the network against ground-truth target images and input pixel coordinates. All headline claims are evaluated on external benchmarks (RE10K, ACID, DTU, DL3DV, ScanNet++) against baselines from other groups, and no equation defines a predicted quantity in terms of the quantity it is claimed to predict. The paper's self-citation of SPFSplat is used as a baseline and starting point for architectural changes, with improvements measured independently, so it is not load-bearing. The one significant caveat is the scope of the 'self-supervised' label: Section IV-D discloses that random initialization requires a DUSt3R point-cloud distillation warm-up, and Section V states that the method 'benefits from the priors provided by supervised models such as MASt3R and VGGT,' with Table XI showing a 3.76 dB PSNR drop without such pretrained initialization. This means the strong results are largely inherited from pose/geometry-supervised pretraining, and the abstract's 'absence of pose supervision' is true of the fine-tuning losses but not of the full training pipeline. That is an important interpretation and attribution limitation, but it is not circularity: the SOTA numbers are externally benchmarked, the losses do not reduce by construction to their inputs, and no fitted parameter is renamed as a prediction.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

The method introduces no new physical entities or measurements. Its scientific contribution is an empirical architecture and training recipe. The load-bearing inputs are two hand-set loss weights and the pretrained models, which are not derived within the paper.

free parameters (2)
  • LPIPS loss weight gamma = 0.05
    Set by hand in Eq. 13 to balance L2 and LPIPS in the rendering loss; the value is not derived and is not swept in any ablation.
  • reprojection loss weight = 0.001
    Set by hand in the total loss, mentioned in Section IV-B. No sensitivity analysis is reported for this coefficient.
assumptions (3)
  • domain assumption The canonical coordinate frame anchored at the first context view I_1 fixes the gauge ambiguity.
    Section III-A sets I_1 as global frame with pose [U|0]. The method relies on this to define relative poses and pixel-aligned Gaussians; if the first view is not well conditioned the reconstruction is defined around it.
  • domain assumption RE10K and ACID ground-truth poses from SfM (COLMAP) are accurate enough for evaluation and for supervising baselines.
    Section IV-A states camera poses are obtained via SfM and follows official splits. If these poses are noisy, the comparisons against pose-supervised baselines could be affected.
  • ad hoc to paper MASt3R and VGGT pretrained weights provide useful geometric priors and are loaded as initialization.
    Section IV-B initializes the encoder/decoder/heads from MASt3R or VGGT. Table XI shows random initialization degrades performance substantially, so the method's outcome depends on these external, pose-supervised pretrained models.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SPFSplatV2: Efficient Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views." pith.science (2026). https://pith.science/paper/FYTW5EMR

@misc{pith2026250917246,
  author       = {Pith},
  title        = {Pith review of: SPFSplatV2: Efficient Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse Views},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/FYTW5EMR}},
  note         = {Machine review of arXiv:2509.17246}
}
read the original abstract

We introduce SPFSplatV2, an efficient feed-forward framework for 3D Gaussian splatting from sparse multi-view images, requiring no ground-truth poses during training or inference. The framework employs a shared feature extraction backbone to jointly predict 3D Gaussian primitives and camera poses in a canonical space from unposed inputs. To enable efficient and accurate pose estimation, we introduce a masked attention mechanism for target-view pose prediction and a reprojection loss that enforces pixel-aligned Gaussian primitives, providing stronger geometric constraints. We further demonstrate the compatibility of our training framework with different reconstruction architectures, resulting in two model variants. Remarkably, despite the absence of pose supervision, our method achieves state-of-the-art performance in both in-domain and out-of-domain novel view synthesis, even under extreme viewpoint changes and limited image overlap. It also surpasses many methods that rely on geometric supervision in relative pose estimation. By eliminating dependence on ground-truth poses, our method offers the scalability to leverage larger and more diverse datasets. Code and pretrained models will be available on our project page: https://ranrhuang.github.io/spfsplatv2/.

Figures

Figures reproduced from arXiv: 2509.17246 by the authors.

Figure 1
Figure 1. Comparison of three typical training pipelines for sparse-view 3D reconstruction in novel view synthesis. For simplicity, the image rendering loss on the rendered target view is omitted. (a) Pose-required methods rely on ground-truth poses for both 3D scene reconstruction and target-view rendering. (b) Supervised pose-free methods requires no ground-truth poses for reconstruction but still rely on ground-truth poses… view at source ↗
Figure 2
Figure 2. Training pipeline of SPFSplatV2. A shared backbone with three specialized heads simultaneously predicts Gaussian centers, additional Gaussian [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Architecture comparison of SPFSplatV2 and SPFSplatV2-L. SPF [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Comparison of cross-attention in (a) SPFSplat and (b) SPFSplatV2/ [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Qualitative comparison on RE10K (top three rows) and ACID (bottom three rows). Our method 1) better handles extreme viewpoint changes and [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 6
Figure 6. Figure 6: Qualitative comparison on cross-dataset generalization. All methods are trained on RE10K and evaluated on ACID and DTU, DL3DV, and ScanNet++ [PITH_FULL_IMAGE:figures/full_fig_p009_6.png]
Figure 7
Figure 7. Figure 7: Comparison of 3D Gaussians and rendered results. Red and green denote context and target camera poses, respectively. Rendered images and depth [PITH_FULL_IMAGE:figures/full_fig_p010_7.png]
Figure 8
Figure 8. Figure 8: 3D Gaussians from smartphone without intrinsics and rendered image. [PITH_FULL_IMAGE:figures/full_fig_p013_8.png]
Figure 9
Figure 9. Figure 9: Failure cases of SPFSplatV2. Blurriness and artifacts occur in occluded [PITH_FULL_IMAGE:figures/full_fig_p013_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

74 extracted references · 26 canonical work pages

  1. [1]

    Nerf: Representing scenes as neural radiance fields for view synthesis,

    B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng, “Nerf: Representing scenes as neural radiance fields for view synthesis,”Communications of the ACM, vol. 65, no. 1, pp. 99–106, 2021

  2. [2]

    3d gaussian splatting for real-time radiance field rendering

    B. Kerbl, G. Kopanas, T. Leimk ¨uhler, and G. Drettakis, “3d gaussian splatting for real-time radiance field rendering.”ACM Trans. Graph., vol. 42, no. 4, pp. 139–1, 2023

  3. [3]

    Mvsnerf: Fast generalizable radiance field reconstruction from multi-view stereo,

    A. Chen, Z. Xu, F. Zhao, X. Zhang, F. Xiang, J. Yu, and H. Su, “Mvsnerf: Fast generalizable radiance field reconstruction from multi-view stereo,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 14 124–14 133

  4. [4]

    Murf: multi-baseline radiance fields,

    H. Xu, A. Chen, Y . Chen, C. Sakaridis, Y . Zhang, M. Pollefeys, A. Geiger, and F. Yu, “Murf: multi-baseline radiance fields,” inPro- ceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 20 041–20 050

  5. [5]

    pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d recon- struction,

    D. Charatan, S. L. Li, A. Tagliasacchi, and V . Sitzmann, “pixelsplat: 3d gaussian splats from image pairs for scalable generalizable 3d recon- struction,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 19 457–19 467

  6. [6]

    Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images,

    Y . Chen, H. Xu, C. Zheng, B. Zhuang, M. Pollefeys, A. Geiger, T.- J. Cham, and J. Cai, “Mvsplat: Efficient 3d gaussian splatting from sparse multi-view images,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 370–386

  7. [7]

    No pose, no problem: Surprisingly simple 3d gaussian splats from sparse unposed images,

    B. Ye, S. Liu, H. Xu, X. Li, M. Pollefeys, M.-H. Yang, and S. Peng, “No pose, no problem: Surprisingly simple 3d gaussian splats from sparse unposed images,” inThe Thirteenth International Conference on Learning Representations, 2025. [Online]. Available: https://openreview.net/forum?id=P4o9akekdf

  8. [8]

    Splatt3r: Zero- shot gaussian splatting from uncalibrated image pairs,

    B. Smart, C. Zheng, I. Laina, and V . A. Prisacariu, “Splatt3r: Zero- shot gaussian splatting from uncalibrated image pairs,”arXiv preprint arXiv:2408.13912, 2024

Show all 74 references
  1. [9]

    Gs-lrm: Large reconstruction model for 3d gaussian splatting,

    K. Zhang, S. Bi, H. Tan, Y . Xiangli, N. Zhao, K. Sunkavalli, and Z. Xu, “Gs-lrm: Large reconstruction model for 3d gaussian splatting,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 1–19

  2. [10]

    Lgm: Large multi-view gaussian model for high-resolution 3d content creation,

    J. Tang, Z. Chen, X. Chen, T. Wang, G. Zeng, and Z. Liu, “Lgm: Large multi-view gaussian model for high-resolution 3d content creation,” in European Conference on Computer Vision. Springer, 2024, pp. 1–18

  3. [11]

    Grm: Large gaussian reconstruction model for efficient 3d reconstruction and generation,

    Y . Xu, Z. Shi, W. Yifan, H. Chen, C. Yang, S. Peng, Y . Shen, and G. Wetzstein, “Grm: Large gaussian reconstruction model for efficient 3d reconstruction and generation,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 1–20

  4. [13]

    LEAP: Liberate sparse- view 3d modeling from camera poses,

    H. Jiang, Z. Jiang, Y . Zhao, and Q. Huang, “LEAP: Liberate sparse- view 3d modeling from camera poses,” inThe Twelfth International Conference on Learning Representations, 2024. [Online]. Available: https://openreview.net/forum?id=KPmajBxEaF

  5. [14]

    PF-LRM: Pose-free large reconstruction model for joint pose and shape prediction,

    P. Wang, H. Tan, S. Bi, Y . Xu, F. Luan, K. Sunkavalli, W. Wang, Z. Xu, and K. Zhang, “PF-LRM: Pose-free large reconstruction model for joint pose and shape prediction,” inThe Twelfth International Conference on Learning Representations, 2024. [Online]. Available: https://open...

  6. [15]

    Upfu- sion: Novel view diffusion from unposed sparse view observations,

    B. R. Nagoor Kani, H.-Y . Lee, S. Tulyakov, and S. Tulsiani, “Upfu- sion: Novel view diffusion from unposed sparse view observations,” in European Conference on Computer Vision (ECCV), 2024

  7. [16]

    Scene representation transformer: Geometry-free novel view synthesis through set-latent scene representations,

    M. S. Sajjadi, H. Meyer, E. Pot, U. Bergmann, K. Greff, N. Radwan, S. V ora, M. Lu ˇci´c, D. Duckworth, A. Dosovitskiyet al., “Scene representation transformer: Geometry-free novel view synthesis through set-latent scene representations,” inProceedings of the IEEE/CVF Con- fer...

  8. [17]

    Unifying correspondence pose and nerf for generalized pose-free novel view synthesis,

    S. Hong, J. Jung, H. Shin, J. Yang, S. Kim, and C. Luo, “Unifying correspondence pose and nerf for generalized pose-free novel view synthesis,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 20 196–20 206

  9. [18]

    Dbarf: Deep bundle-adjusting generalizable neural radiance fields,

    Y . Chen and G. H. Lee, “Dbarf: Deep bundle-adjusting generalizable neural radiance fields,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 24–34

  10. [19]

    Barf: Bundle-adjusting neural radiance fields,

    C.-H. Lin, W.-C. Ma, A. Torralba, and S. Lucey, “Barf: Bundle-adjusting neural radiance fields,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 5741–5751

  11. [20]

    Pf3plat: Pose-free feed-forward 3d gaussian splatting,

    S. Hong, J. Jung, H. Shin, J. Han, J. Yang, C. Luo, and S. Kim, “Pf3plat: Pose-free feed-forward 3d gaussian splatting,”arXiv preprint arXiv:2410.22128, 2024. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2025 14

  12. [21]

    Selfsplat: Pose-free and 3d prior-free generalizable 3d gaussian splat- ting,

    G. Kang, J. Yoo, J. Park, S. Nam, H. Im, S. Shin, S. Kim, and E. Park, “Selfsplat: Pose-free and 3d prior-free generalizable 3d gaussian splat- ting,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 22 012–22 022

  13. [22]

    No pose at all: Self-supervised pose-free 3d gaussian splatting from sparse views,

    R. Huang and K. Mikolajczyk, “No pose at all: Self-supervised pose-free 3d gaussian splatting from sparse views,”arXiv preprint arXiv:2508.01171, 2025

  14. [23]

    Grounding image matching in 3d with mast3r,

    V . Leroy, Y . Cabon, and J. Revaud, “Grounding image matching in 3d with mast3r,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 71–91

  15. [24]

    Vggt: Visual geometry grounded transformer,

    J. Wang, M. Chen, N. Karaev, A. Vedaldi, C. Rupprecht, and D. Novotny, “Vggt: Visual geometry grounded transformer,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 5294– 5306

  16. [25]

    Instant neural graphics primitives with a multiresolution hash encoding,

    T. M ¨uller, A. Evans, C. Schied, and A. Keller, “Instant neural graphics primitives with a multiresolution hash encoding,”ACM transactions on graphics (TOG), vol. 41, no. 4, pp. 1–15, 2022

  17. [26]

    K-planes: Explicit radiance fields in space, time, and appearance,

    S. Fridovich-Keil, G. Meanti, F. R. Warburg, B. Recht, and A. Kanazawa, “K-planes: Explicit radiance fields in space, time, and appearance,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 12 479–12 488

  18. [27]

    Tensorf: Tensorial radiance fields,

    A. Chen, Z. Xu, A. Geiger, J. Yu, and H. Su, “Tensorf: Tensorial radiance fields,” inEuropean conference on computer vision. Springer, 2022, pp. 333–350

  19. [28]

    Plenoxels: Radiance fields without neural networks,

    S. Fridovich-Keil, A. Yu, M. Tancik, Q. Chen, B. Recht, and A. Kanazawa, “Plenoxels: Radiance fields without neural networks,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 5501–5510

  20. [29]

    Dust3r: Geometric 3d vision made easy,

    S. Wang, V . Leroy, Y . Cabon, B. Chidlovskii, and J. Revaud, “Dust3r: Geometric 3d vision made easy,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 20 697–20 709

  21. [30]

    Relpose: Predicting prob- abilistic relative rotation for single objects in the wild,

    J. Y . Zhang, D. Ramanan, and S. Tulsiani, “Relpose: Predicting prob- abilistic relative rotation for single objects in the wild,” inEuropean Conference on Computer Vision. Springer, 2022, pp. 592–611

  22. [31]

    Relpose++: Recov- ering 6d poses from sparse-view observations,

    A. Lin, J. Y . Zhang, D. Ramanan, and S. Tulsiani, “Relpose++: Recov- ering 6d poses from sparse-view observations,” in2024 International Conference on 3D Vision (3DV). IEEE, 2024, pp. 106–115

  23. [32]

    Posediffusion: Solving pose estimation via diffusion-aided bundle adjustment,

    J. Wang, C. Rupprecht, and D. Novotny, “Posediffusion: Solving pose estimation via diffusion-aided bundle adjustment,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 9773–9783

  24. [33]

    Sparf: Neural radiance fields from sparse and noisy poses,

    P. Truong, M.-J. Rakotosaona, F. Manhardt, and F. Tombari, “Sparf: Neural radiance fields from sparse and noisy poses,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 4190–4200

  25. [34]

    Nope-nerf: Optimising neural radiance field with no pose prior,

    W. Bian, Z. Wang, K. Li, J.-W. Bian, and V . A. Prisacariu, “Nope-nerf: Optimising neural radiance field with no pose prior,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 4160–4169

  26. [35]

    Colmap- free 3d gaussian splatting,

    Y . Fu, S. Liu, A. Kulkarni, J. Kautz, A. A. Efros, and X. Wang, “Colmap- free 3d gaussian splatting,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 20 796–20 805

  27. [36]

    Flowcam: Training gener- alizable 3d radiance fields without camera poses via pixel-aligned scene flow,

    C. Smith, Y . Du, A. Tewari, and V . Sitzmann, “Flowcam: Training gener- alizable 3d radiance fields without camera poses via pixel-aligned scene flow,”Advances in Neural Information Processing Systems, vol. 36, pp. 1476–1488, 2023

  28. [37]

    Lightglue: Local feature matching at light speed,

    P. Lindenberger, P.-E. Sarlin, and M. Pollefeys, “Lightglue: Local feature matching at light speed,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 17 627–17 638

  29. [38]

    U-net: Convolutional networks for biomedical image segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” inMedical image computing and computer-assisted intervention–MICCAI 2015: 18th international con- ference, Munich, Germany, October 5-9, 2015, proceedings, part III 18. ...

  30. [39]

    High- resolution image synthesis with latent diffusion models,

    R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High- resolution image synthesis with latent diffusion models,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 10 684–10 695

  31. [40]

    Hartley and A

    R. Hartley and A. Zisserman,Multiple view geometry in computer vision. Cambridge university press, 2003

  32. [41]

    Structure-from-motion revisited,

    J. L. Schonberger and J.-M. Frahm, “Structure-from-motion revisited,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 4104–4113

  33. [42]

    Distinctive image features from scale-invariant keypoints,

    D. G. Lowe, “Distinctive image features from scale-invariant keypoints,” International journal of computer vision, vol. 60, pp. 91–110, 2004

  34. [43]

    Speeded-up robust features (surf),

    H. Bay, A. Ess, T. Tuytelaars, and L. Van Gool, “Speeded-up robust features (surf),”Comput. Vis. Image. Und., vol. 110, no. 3, pp. 346– 359, 2008

  35. [44]

    Orb: An efficient alternative to sift or surf,

    E. Rublee, V . Rabaud, K. Konolige, and G. Bradski, “Orb: An efficient alternative to sift or surf,” inProc. IEEE Int. Conf. Comput. Vision. (ICCV). Ieee, 2011, pp. 2564–2571

  36. [45]

    Random sample consensus: a paradigm for model fitting with applications to image analysis and automated car- tography,

    M. FISCHLER AND, “Random sample consensus: a paradigm for model fitting with applications to image analysis and automated car- tography,”Commun. ACM, vol. 24, no. 6, pp. 381–395, 1981

  37. [46]

    Triangulation,

    R. I. Hartley and P. Sturm, “Triangulation,”Computer vision and image understanding, vol. 68, no. 2, pp. 146–157, 1997

  38. [47]

    Bundle adjustment—a modern synthesis,

    B. Triggs, P. F. McLauchlan, R. I. Hartley, and A. W. Fitzgibbon, “Bundle adjustment—a modern synthesis,” inInternational workshop on vision algorithms. Springer, 1999, pp. 298–372

  39. [48]

    Superpoint: Self- supervised interest point detection and description,

    D. DeTone, T. Malisiewicz, and A. Rabinovich, “Superpoint: Self- supervised interest point detection and description,” inProceedings of the IEEE conference on computer vision and pattern recognition workshops, 2018, pp. 224–236

  40. [49]

    D2-Net: A Trainable CNN for Joint Detection and Description of Local Features,

    M. Dusmanu, I. Rocco, T. Pajdla, M. Pollefeys, J. Sivic, A. Torii, and T. Sattler, “D2-Net: A Trainable CNN for Joint Detection and Description of Local Features,” inProceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2019

  41. [50]

    Drkf: Distilled rotated kernel fusion for efficient rotation invariant descriptors in local feature matching,

    R. Huang, J. Cai, C. Li, Z. Wu, X. Liu, and Z. Chai, “Drkf: Distilled rotated kernel fusion for efficient rotation invariant descriptors in local feature matching,” in2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2023, pp. 1885–1892

  42. [51]

    Aslfeat: Learning local features of accurate shape and localization,

    Z. Luo, L. Zhou, X. Bai, H. Chen, J. Zhang, Y . Yao, S. Li, T. Fang, and L. Quan, “Aslfeat: Learning local features of accurate shape and localization,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 6589–6598

  43. [52]

    Superglue: Learning feature matching with graph neural networks,

    P.-E. Sarlin, D. DeTone, T. Malisiewicz, and A. Rabinovich, “Superglue: Learning feature matching with graph neural networks,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2020, pp. 4938–4947

  44. [53]

    Loftr: Detector- free local feature matching with transformers,

    J. Sun, Z. Shen, Y . Wang, H. Bao, and X. Zhou, “Loftr: Detector- free local feature matching with transformers,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2021, pp. 8922–8931

  45. [54]

    Ba-net: Dense bundle adjustment network,

    C. Tang and P. Tan, “Ba-net: Dense bundle adjustment network,”arXiv preprint arXiv:1806.04807, 2018

  46. [55]

    Vggsfm: Visual geometry grounded deep structure from motion,

    J. Wang, N. Karaev, C. Rupprecht, and D. Novotny, “Vggsfm: Visual geometry grounded deep structure from motion,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2024, pp. 21 686–21 697

  47. [56]

    Mv-dust3r+: Single-stage scene reconstruction from sparse views in 2 seconds,

    Z. Tang, Y . Fan, D. Wang, H. Xu, R. Ranjan, A. Schwing, and Z. Yan, “Mv-dust3r+: Single-stage scene reconstruction from sparse views in 2 seconds,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 5283–5293

  48. [57]

    Fast3r: Towards 3d reconstruction of 1000+ images in one forward pass,

    J. Yang, A. Sax, K. J. Liang, M. Henaff, H. Tang, A. Cao, J. Chai, F. Meier, and M. Feiszli, “Fast3r: Towards 3d reconstruction of 1000+ images in one forward pass,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 21 924–21 935

  49. [58]

    Croco: Self- supervised pre-training for 3d vision tasks by cross-view completion,

    P. Weinzaepfel, V . Leroy, T. Lucas, R. Br ´egier, Y . Cabon, V . Arora, L. Antsfeld, B. Chidlovskii, G. Csurka, and J. Revaud, “Croco: Self- supervised pre-training for 3d vision tasks by cross-view completion,” Advances in Neural Information Processing Systems, vol. 35, pp. ...

  50. [59]

    Flare: Feed-forward geometry, appearance and camera estimation from uncalibrated sparse views,

    S. Zhang, J. Wang, Y . Xu, N. Xue, C. Rupprecht, X. Zhou, Y . Shen, and G. Wetzstein, “Flare: Feed-forward geometry, appearance and camera estimation from uncalibrated sparse views,” inProceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 21 936– 21 947

  51. [60]

    An image is worth 16x16 words: Transformers for image recognition at scale,

    A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” inInternational Conference on Learning ...

  52. [61]

    Vision transformers for dense prediction,

    R. Ranftl, A. Bochkovskiy, and V . Koltun, “Vision transformers for dense prediction,” inProceedings of the IEEE/CVF international conference on computer vision, 2021, pp. 12 179–12 188

  53. [62]

    Accelerated coordi- nate encoding: Learning to relocalize in minutes using rgb and poses,

    E. Brachmann, T. Cavallari, and V . A. Prisacariu, “Accelerated coordi- nate encoding: Learning to relocalize in minutes using rgb and poses,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 5044–5053

  54. [63]

    On the continuity of rotation representations in neural networks,

    Y . Zhou, C. Barnes, J. Lu, J. Yang, and H. Li, “On the continuity of rotation representations in neural networks,” inProceedings of the JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2025 15 IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 5745–5753

  55. [64]

    The unreasonable effectiveness of deep features as a perceptual metric,

    R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 586–595

  56. [65]

    Scene coordinate reconstruction: Posing of image collections via incremental learning of a relocalizer,

    E. Brachmann, J. Wynn, S. Chen, T. Cavallari, ´A. Monszpart, D. Tur- mukhambetov, and V . A. Prisacariu, “Scene coordinate reconstruction: Posing of image collections via incremental learning of a relocalizer,” inEuropean Conference on Computer Vision. Springer, 2024, pp. 421– 440

  57. [66]

    Stereo magnification: learning view synthesis using multiplane images,

    T. Zhou, R. Tucker, J. Flynn, G. Fyffe, and N. Snavely, “Stereo magnification: learning view synthesis using multiplane images,”ACM Transactions on Graphics (TOG), vol. 37, no. 4, pp. 1–12, 2018

  58. [67]

    Infinite nature: Perpetual view generation of natural scenes from a single image,

    A. Liu, R. Tucker, V . Jampani, A. Makadia, N. Snavely, and A. Kanazawa, “Infinite nature: Perpetual view generation of natural scenes from a single image,” inProceedings of the IEEE/CVF Inter- national Conference on Computer Vision, 2021, pp. 14 458–14 467

  59. [68]

    Roma: Robust dense feature matching,

    J. Edstedt, Q. Sun, G. B ¨okman, M. Wadenb¨ack, and M. Felsberg, “Roma: Robust dense feature matching,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 19 790–19 800

  60. [69]

    Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision,

    L. Ling, Y . Sheng, Z. Tu, W. Zhao, C. Xin, K. Wan, L. Yu, Q. Guo, Z. Yu, Y . Luet al., “Dl3dv-10k: A large-scale scene dataset for deep learning-based 3d vision,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 22 160–22 169

  61. [70]

    Large scale multi-view stereopsis evaluation,

    R. Jensen, A. Dahl, G. V ogiatzis, E. Tola, and H. Aanæs, “Large scale multi-view stereopsis evaluation,” inProceedings of the IEEE conference on computer vision and pattern recognition, 2014, pp. 406–413

  62. [71]

    Scannet++: A high- fidelity dataset of 3d indoor scenes,

    C. Yeshwanth, Y .-C. Liu, M. Nießner, and A. Dai, “Scannet++: A high- fidelity dataset of 3d indoor scenes,” inProceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 12–22

  63. [72]

    Image quality assessment: from error visibility to structural similarity,

    Z. Wang, A. C. Bovik, H. R. Sheikh, and E. P. Simoncelli, “Image quality assessment: from error visibility to structural similarity,”IEEE transactions on image processing, vol. 13, no. 4, pp. 600–612, 2004

  64. [73]

    Dinov2: Learning robust visual features without supervision,

    M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Noubyet al., “Dinov2: Learning robust visual features without supervision,”arXiv preprint arXiv:2304.07193, 2023

  65. [74]

    Real-time deep pose estimation with geodesic loss for image-to-template rigid registration,

    S. S. M. Salehi, S. Khan, D. Erdogmus, and A. Gholipour, “Real-time deep pose estimation with geodesic loss for image-to-template rigid registration,”IEEE transactions on medical imaging, vol. 38, no. 2, pp. 470–481, 2018

  66. [75]

    Croco v2: Improved cross-view completion pre-training for stereo matching and optical flow,

    P. Weinzaepfel, T. Lucas, V . Leroy, Y . Cabon, V . Arora, R. Br ´egier, G. Csurka, L. Antsfeld, B. Chidlovskii, and J. Revaud, “Croco v2: Improved cross-view completion pre-training for stereo matching and optical flow,” inProceedings of the IEEE/CVF International Conference ...

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.