Pith. sign in

REVIEW 2 major objections 2 minor 53 references

BA-T: An Iterative Transformer for Two-View Bundle Adjustment

T0 review · 2 major / 2 minor · reviewed 2026-06-28 · grok-4.3

Pith's one-line read BA-T replaces deep attention stacks with a single repeatable lightweight layer that performs bundle-adjustment style updates for two-view 3D reconstruction.

desk verdict BA-T frames bundle adjustment as a single repeatable transformer layer for two-view reconstruction and claims big parameter savings, but the math link to classical BA is not shown. read the letter →

arxiv 2606.03287 v1 pith:UXSID53T submitted 2026-06-02 cs.CV

classification cs.CV
keywords bundleadjustmentiterativetransformertwo-viewreconstructioncross-viewconsistencylightweightdecoder3Dposerefinement
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes BA-T, an iterative Transformer that treats classical bundle adjustment as a repeatable structured update process inside token space. Instead of stacking many attention layers, it uses one lightweight layer to refine latent residuals between poses and local geometry. The central goal is to obtain stronger cross-view consistency and progressive accuracy gains without the parameter cost of conventional decoders. A reader would care if this shows that geometric structure can substitute for depth in feed-forward reconstruction models. Experiments indicate that accuracy keeps rising across iterations while decoder size stays at 16 percent of larger baselines.

What carries the argument

The BA-T layer: a single lightweight transformer layer that executes BA-style structured updates as a repeatable operation in token space.

What would settle it

Measure whether reconstruction error and cross-view consistency continue to improve after multiple iterations on a held-out two-view benchmark or plateau at the level of a single pass.

Watch

Extended reading notes

Core claim

BA-T implements bundle adjustment as an iterative information propagation process between poses and local geometry realized as a single lightweight repeatable layer in implicit token space. This layer refines predictions from latent residuals rather than relying on deep cross-view attention stacks, producing progressive improvements in pose and reconstruction accuracy together with stronger cross-view consistency.

Load-bearing premise

One lightweight layer can faithfully carry out the structured geometric updates of bundle adjustment inside implicit token representations.

Editorial extensions

If this is right

  • Pose and point accuracy increase with each additional iteration of the BA-T layer.
  • Cross-view consistency exceeds that obtained from conventional deep decoder stacks.
  • Performance matches or exceeds substantially larger models while using 16 percent of their decoder parameters.
  • The architecture supplies a compact structural alternative to depth-heavy attention for accurate 3D reconstruction.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same lightweight update layer could be stacked or adapted for three or more input views without redesigning the core mechanism.
  • Training might converge faster if the BA-style residual update is initialized from classical bundle-adjustment solutions on the same data.
  • Runtime cost in real-time pipelines could drop further if the number of iterations is made input-dependent rather than fixed.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 2 minor

Summary. The paper proposes BA-T, an iterative Transformer for two-view bundle adjustment that draws from classical BA as an information-propagation process between poses and local geometry. It replaces heavy decoder stacks with a single repeatable lightweight layer that performs BA-style structured updates in implicit token space by refining predictions from latent residuals. The central claims are that this yields progressive gains in pose/reconstruction accuracy across iterations, stronger cross-view consistency than conventional decoders, and performance matching or exceeding much larger models while using only 16% of their decoder parameters.

Significance. If the claimed structural equivalence to BA holds and is shown to be non-circular, the result would supply a compact, parameter-efficient architectural primitive for multi-view geometry that could replace depth-heavy attention in feed-forward 3D reconstruction pipelines. The explicit promise of public code strengthens reproducibility.

major comments (2)
  1. [Abstract / §3 (method)] Abstract and method description: the claim that the repeatable lightweight layer 'implements BA-style structured updates' and refines 'based on latent residual' is load-bearing for all consistency and efficiency assertions, yet no equations are supplied that map the layer operations (attention, residual, or token interactions) onto classical BA quantities such as the normal equations, Schur complement, or explicit pose-point information propagation. Without this mapping it remains possible that observed gains arise from iteration count or residual connections alone.
  2. [Abstract / §4 (experiments)] Experimental section: the abstract asserts progressive accuracy improvement, stronger consistency, and parameter-efficient superiority, but the provided text supplies no dataset names, baseline architectures, error metrics (e.g., rotation/translation error, reprojection), ablation controls on layer depth versus iteration count, or statistical significance tests. These details are required to substantiate the cross-model comparison at 16% decoder parameters.
minor comments (2)
  1. [Abstract] The abstract states that code will be released at a GitHub URL; confirming the repository contains the exact training and evaluation scripts used for the reported numbers would aid verification.
  2. [§3] Notation for pose and point tokens should be introduced once with explicit dimensionality before the layer description to avoid ambiguity in the implicit token space.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for the constructive comments. We address each major comment below and indicate the planned revisions.

read point-by-point responses
  1. Referee: [Abstract / §3 (method)] Abstract and method description: the claim that the repeatable lightweight layer 'implements BA-style structured updates' and refines 'based on latent residual' is load-bearing for all consistency and efficiency assertions, yet no equations are supplied that map the layer operations (attention, residual, or token interactions) onto classical BA quantities such as the normal equations, Schur complement, or explicit pose-point information propagation. Without this mapping it remains possible that observed gains arise from iteration count or residual connections alone.

    Authors: We agree that the absence of an explicit mapping leaves the structural claim open to the interpretation raised. The layer is motivated by viewing BA as iterative information propagation between poses and points, realized via attention and residuals in token space, but the manuscript does not derive or equate the operations to the normal equations or Schur complement. In revision we will add a concise subsection in §3 that supplies a conceptual correspondence (e.g., how cross-view attention approximates pose-point message passing and how the residual step parallels the BA update), while acknowledging it is an implicit rather than algebraic equivalence. This will clarify the intended source of the observed gains. revision: yes

  2. Referee: [Abstract / §4 (experiments)] Experimental section: the abstract asserts progressive accuracy improvement, stronger consistency, and parameter-efficient superiority, but the provided text supplies no dataset names, baseline architectures, error metrics (e.g., rotation/translation error, reprojection), ablation controls on layer depth versus iteration count, or statistical significance tests. These details are required to substantiate the cross-model comparison at 16% decoder parameters.

    Authors: We accept that the experimental reporting must be expanded for the claims to be fully substantiated. The current manuscript text does not enumerate the required specifics. In the revised version we will augment §4 with explicit dataset names, baseline architectures, the precise error metrics, ablations that isolate iteration count from layer depth, and any statistical tests performed, thereby supporting the progressive improvement and 16 % parameter-efficiency statements. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: architectural claim stands independent of inputs

full rationale

The paper presents BA-T as an iterative transformer layer inspired by classical bundle adjustment for structured updates in token space. No equations, fitted parameters, or self-citations are shown that reduce the claimed consistency gains or parameter efficiency to a definitional equivalence or statistical forcing. The derivation chain consists of an inspiration step followed by experimental validation; the layer is not shown to be equivalent to its inputs by construction, nor does any load-bearing premise collapse to a prior self-citation. This is the common case of a self-contained architectural proposal.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Only the abstract is available; no explicit free parameters, axioms, or invented entities are identifiable from the provided text.

how reviews work

0 comments
Cite this review

Pith. "Pith review of BA-T: An Iterative Transformer for Two-View Bundle Adjustment." pith.science (2026). https://pith.science/paper/UXSID53T

@misc{pith2026260603287,
  author       = {Pith},
  title        = {Pith review of: BA-T: An Iterative Transformer for Two-View Bundle Adjustment},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/UXSID53T}},
  note         = {Machine review of arXiv:2606.03287}
}
read the original abstract

Feed-forward models for 3D reconstruction have achieved strong performance using deep cross-view attention to exchange information across images. However, these approaches often depend on heavy decoder stacks and lack a structured mechanism for geometry refinement, resulting in poor multi-view consistency. We address this by drawing inspiration from classical bundle adjustment (BA), which can be viewed as an iterative information propagation process between poses and local geometry. Inspired by BA, we propose BA-T, an iterative Transformer that implements BA-style structured updates as a repeatable layer in implicit token space. Instead of relying on deep attention stacks, BA-T refines predictions based on latent residual by a single lightweight layer. Experiments demonstrate that BA-T progressively improves pose and reconstruction accuracy across iterations, achieves stronger cross-view consistency than conventional decoders, and matches or surpasses substantially larger models while using only 16% of their decoder parameters. BA-T provides a compact, efficient, and structural alternative to depth-heavy attention, enabling accurate 3D reconstruction within a lightweight architecture. The code will be made publicly at https://github.com/zhangganlin/BA-T.

Figures

Figures reproduced from arXiv: 2606.03287 by the authors.

Figure 1
Figure 1. Overview of BA-T. Given input images, BA-T performs iterative updates on camera and local geometry tokens using a compact, reusable BA-T layer in latent space. ←− indicates error between GT poses and estimated poses (blue and pink), which gradually decreases, and red boxes highlight regions progressively refined across iterations. The 3D point error maps visualize the per-view 3D point errors. Abstract Feed-forward … view at source ↗
Figure 2
Figure 2. Overview of the BA-T pipeline. BA-T takes camera tokens (from learnable initialization) and local geometry tokens (from the image encoder) as input and refines them iteratively. At each step, it performs a BA-inspired implicit refinement step by transforming geometry tokens across camera spaces, matching correspondences, and computing latent residuals. Camera tokens and per-view geometry tokens are refined via the C… view at source ↗
Figure 3
Figure 3. Token-level correspondence response. Given a local geometry token from one view, its attention scores highlight responses in the other view. Correct responses are observed for both ambiguous (left, middle) and distinctive (right) regions. 3.3 Implementation of General Functions in BA-T Inspired by the general form of BA, we design the following components to support iterative refinement and information exchange betw… view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Visualization of latent residuals. Residuals are computed in the latent space of view b. Their magnitude decreases across refinement iterations, indicating increasingly accurate estimates. the query and the residual tokens act as the context to be aggregated, ∆c (k) a→…
Figure 5
Figure 5. Figure 5: (right) Iterative refinement behavior. BA-T (green curve) consistently outperforms ViSTA† w/ iterative training (red curve) and converges within 3 ∼ 4 refinement steps. Rot. AUC (deg) Trans. AUC (m) Geometry Decoder Size # Method Iter @5°↑ @10°↑ @20°↑ @0.05↑ @0.10↑ @0.…
Figure 6
Figure 6. Figure 6: (left) Qualitative reconstruction results. We visualize reconstructed geometry and 3D points (in local frame) error maps at iteration 1 and iteration 4. Red boxes highlight noticeable misalignments in the 1st iteration, which are corrected in the 4th iteration, as indi…
Figure 7
Figure 7. Figure 7: Multiview reconstruction results. Estimated poses and reconstructed scenes are visualized across iterations for a 4-view input setup, demonstrating BA-T’s ability to handle multi-view settings. The green frustums indicate GT camera poses. The red boxes and decreasing t…
Figure 8
Figure 8. Figure 8: Reconstruction comparison. Both methods are initialized from BA-T (iter = 0). Droid￾SLAM uses ground-truth intrinsics, whereas BA-T does not. BA-T achieves stronger performance with fewer iterations and faster runtime, benefiting from its more expressive latent space. …
Figure 9
Figure 9. Figure 9: Visualization of camera-conditioned geometry transformation. Given two input views, a and b, we visualize the point map regressed from the geometry tokens ga of view a as colored point clouds in the coordinate frame of view a. We also visualize the point map regressed …
Figure 10
Figure 10. Figure 10: More qualitative 4-view results. 17 [PITH_FULL_IMAGE:figures/full_fig_p017_10.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

53 extracted references · 9 canonical work pages

  1. [1]

    In: ECCV

    Agarwal, S., Snavely, N., Seitz, S.M., Szeliski, R.: Bundle adjustment in the large. In: ECCV . pp. 29–42. Springer (2010)

  2. [2]

    In: ACCV

    Alismail, H., Browning, B., Lucey, S.: Photometric bundle adjustment for vision-based slam. In: ACCV . pp. 324–341. Springer (2016)

  3. [3]

    In: ECCV

    Avetisyan, A., Xie, C., Howard-Jenkins, H., Yang, T.Y ., Aroudj, S., Patra, S., Zhang, F., Frost, D., Holland, L., Orme, C., et al.: Scenescript: Reconstructing scenes with an autoregressive structured language model. In: ECCV . pp. 247–263. Springer (2024)

  4. [4]

    In: NeurIPS (2021)

    Baruch, G., Chen, Z., Dehghan, A., Dimry, T., Feigin, Y ., Fu, P., Gebauer, T., Joffe, B., Kurz, D., Schwartz, A., Shulman, E.: ARKitScenes - a diverse real-world dataset for 3D indoor scene understanding using mobile RGB-D data. In: NeurIPS (2021)

  5. [5]

    In: CVPR

    Cabon, Y ., Stoffl, L., Antsfeld, L., Csurka, G., Chidlovskii, B., Revaud, J., Leroy, V .: MUSt3R: Multi-view network for stereo 3D reconstruction. In: CVPR. pp. 1050–1060 (2025)

  6. [6]

    (eds.): SLAM Handbook

    Carlone, L., Kim, A., Barfoot, T., Cremers, D., Dellaert, F. (eds.): SLAM Handbook. From Localization and Mapping to Spatial Intelligence. Cambridge University Press (2026)

  7. [7]

    In: CVPR (2017)

    Dai, A., Chang, A.X., Savva, M., Halber, M., Funkhouser, T., Nießner, M.: ScanNet: Richly- annotated 3D reconstructions of indoor scenes. In: CVPR (2017)

  8. [8]

    ACM TOG (2017)

    Dai, A., Nießner, M., Zollöfer, M., Izadi, S., Theobalt, C.: BundleFusion: Real-time globally consistent 3D reconstruction using on-the-fly surface re-integration. ACM TOG (2017)

Show all 53 references
  1. [9]

    IEEE TPAMI29(6), 1052–1067 (2007)

    Davison, A.J., Reid, I.D., Molton, N.D., Stasse, O.: MonoSLAM: Real-time single camera SLAM. IEEE TPAMI29(6), 1052–1067 (2007)

  2. [10]

    In: CVPR

    Dong, S., Wang, S., Liu, S., Cai, L., Fan, Q., Kannala, J., Yang, Y .: Reloc3r: Large-scale training of relative camera pose regression for generalizable, fast, and accurate visual localization. In: CVPR. pp. 16739–16752 (2025)

  3. [11]

    In: ICLR (2021)

    Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al.: An image is worth 16x16 words: Transformers for image recognition at scale. In: ICLR (2021)

  4. [12]

    IEEE TPAMI40(3), 611–625 (2017)

    Engel, J., Koltun, V ., Cremers, D.: Direct sparse odometry. IEEE TPAMI40(3), 611–625 (2017)

  5. [13]

    In: ECCV

    Engel, J., Schöps, T., Cremers, D.: LSD-SLAM: Large-scale direct monocular SLAM. In: ECCV . pp. 834–849. Springer (2014)

  6. [14]

    In: 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)

    Gao, X., Wang, R., Demmel, N., Cremers, D.: LDSO: Direct sparse odometry with loop closure. In: 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). pp. 2198–2204. IEEE (2018)

  7. [15]

    In: ICCV

    Hagemann, A., Knorr, M., Stiller, C.: Deep geometry-aware camera self-calibration from video. In: ICCV . pp. 3438–3448 (October 2023)

  8. [16]

    Cambridge university press (2003)

    Hartley, R., Zisserman, A.: Multiple view geometry in computer vision. Cambridge university press (2003)

  9. [17]

    In: International Conference on 3D Vision (3DV)

    Keetha, N., Müller, N., Schönberger, J., Porzi, L., Zhang, Y ., Fischer, T., Knapitsch, A., Zauss, D., Weber, E., Antunes, N., Luiten, J., Lopez-Antequera, M., Bulò, S.R., Richardt, C., Ramanan, D., Scherer, S., Kontschieder, P.: MapAnything: Universal feed-forward metric 3D r...

  10. [18]

    In: ECCV

    Leroy, V ., Cabon, Y ., Revaud, J.: Grounding image matching in 3D with MASt3R. In: ECCV . pp. 71–91. Springer (2024)

  11. [19]

    arXiv preprint arXiv:2511.10647 (2025)

    Lin, H., Chen, S., Liew, J.H., Chen, D.Y ., Li, Z., Shi, G., Feng, J., Kang, B.: Depth Anything 3: Recovering the visual space from any views. arXiv preprint arXiv:2511.10647 (2025)

  12. [20]

    In: ECCV

    Liu, S., Gao, Y ., Zhang, T., Pautrat, R., Schönberger, J.L., Larsson, V ., Pollefeys, M.: Robust incremental structure-from-motion with hybrid features. In: ECCV . pp. 249–269. Springer (2024)

  13. [21]

    arXiv preprint arXiv:1711.05101 (2017) 10

    Loshchilov, I., Hutter, F.: Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101 (2017) 10

  14. [22]

    IEEE transactions on robotics31(5), 1147–1163 (2015)

    Mur-Artal, R., Montiel, J.M.M., Tardos, J.D.: ORB-SLAM: A versatile and accurate monocular SLAM system. IEEE transactions on robotics31(5), 1147–1163 (2015)

  15. [23]

    Springer (2006)

    Nocedal, J., Wright, S.J.: Numerical optimization. Springer (2006)

  16. [24]

    In: ECCV

    Pan, L., Baráth, D., Pollefeys, M., Schönberger, J.L.: Global structure-from-motion revisited. In: ECCV . pp. 58–77. Springer (2024)

  17. [25]

    In: ICCV

    Peebles, W., Xie, S.: Scalable diffusion models with transformers. In: ICCV . pp. 4195–4205 (2023)

  18. [26]

    arXiv preprint arXiv:2602.14021 (2026)

    Qian, S., Zhang, G., Wu, S., Cremers, D.: Flow4R: Unifying 4d reconstruction and tracking with scene flow. arXiv preprint arXiv:2602.14021 (2026)

  19. [27]

    In: ICCV (2021)

    Reizenstein, J., Shapovalov, R., Henzler, P., Sbordone, L., Labatut, P., Novotny, D.: Common objects in 3D: Large-scale learning and evaluation of real-life 3D category reconstruction. In: ICCV (2021)

  20. [28]

    In: CVPRW (2025)

    Sandström, E., Zhang, G., Tateno, K., Oechsle, M., Niemeyer, M., Zhang, Y ., Patel, M., Van Gool, L., Oswald, M., Tombari, F.: Splat-SLAM: Globally optimized rgb-only SLAM with 3D gaussians. In: CVPRW (2025)

  21. [29]

    In: CVPR

    Schonberger, J.L., Frahm, J.M.: Structure-from-motion revisited. In: CVPR. pp. 4104–4113 (2016)

  22. [30]

    In: CVPR

    Shotton, J., Glocker, B., Zach, C., Izadi, S., Criminisi, A., Fitzgibbon, A.: Scene coordinate regression forests for camera relocalization in RGB-D images. In: CVPR. pp. 2930–2937 (2013)

  23. [31]

    arXiv preprint arXiv:1906.05797 (2019)

    Straub, J., Whelan, T., Ma, L., Chen, Y ., Wijmans, E., Green, S., Engel, J.J., Mur-Artal, R., Ren, C., Verma, S., Clarkson, A., Yan, M., Budge, B., Yan, Y ., Pan, X., Yon, J., Zou, Y ., Leon, K., Carter, N., Briales, J., Gillingham, T., Mueggler, E., Pesqueira, L., Savva, M.,...

  24. [32]

    In: IROS (Oct 2012)

    Sturm, J., Engelhard, N., Endres, F., Burgard, W., Cremers, D.: A benchmark for the evaluation of RGB-D SLAM systems. In: IROS (Oct 2012)

  25. [33]

    In: ICLR (2019)

    Tang, C., Tan, P.: BA-Net: Dense bundle adjustment network. In: ICLR (2019)

  26. [34]

    In: ECCV

    Teed, Z., Deng, J.: RAFT: Recurrent all-pairs field transforms for optical flow. In: ECCV . pp. 402–419. Springer (2020)

  27. [35]

    NeurIPS34, 16558–16569 (2021)

    Teed, Z., Deng, J.: DROID-SLAM: Deep visual SLAM for monocular, stereo, and RGB-D cameras. NeurIPS34, 16558–16569 (2021)

  28. [36]

    NeurIPS36, 39033–39051 (2023)

    Teed, Z., Lipson, L., Deng, J.: Deep patch visual odometry. NeurIPS36, 39033–39051 (2023)

  29. [37]

    In: International workshop on vision algorithms

    Triggs, B., McLauchlan, P.F., Hartley, R.I., Fitzgibbon, A.W.: Bundle adjustment—a modern synthesis. In: International workshop on vision algorithms. pp. 298–372. Springer (1999)

  30. [38]

    In: 3DV (2025)

    Wang, H., Agapito, L.: 3D reconstruction with spatial memory. In: 3DV (2025)

  31. [39]

    In: CVPR (2025)

    Wang, J., Chen, M., Karaev, N., Vedaldi, A., Rupprecht, C., Novotny, D.: VGGT: Visual geometry grounded transformer. In: CVPR (2025)

  32. [40]

    In: CVPR

    Wang, J., Karaev, N., Rupprecht, C., Novotny, D.: VGGSfM: Visual geometry grounded deep structure from motion. In: CVPR. pp. 21686–21697 (2024)

  33. [41]

    In: CVPR

    Wang, Q., Zhang, Y ., Holynski, A., Efros, A.A., Kanazawa, A.: Continuous 3D perception model with persistent state. In: CVPR. pp. 10510–10522 (2025)

  34. [42]

    In: CVPR

    Wang, S., Leroy, V ., Cabon, Y ., Chidlovskii, B., Revaud, J.: DUSt3R: Geometric 3D vision made easy. In: CVPR. pp. 20697–20709 (2024)

  35. [43]

    Wang, Y ., Zhou, J., Zhu, H., Chang, W., Zhou, Y ., Li, Z., Chen, J., Pang, J., Shen, C., He, T.: π3: Scalable permutation-equivariant visual geometry learning (2025)

  36. [44]

    arXiv preprint arXiv:2510.08575 (2025)

    Xu, H., Barath, D., Geiger, A., Pollefeys, M.: Resplat: Learning recurrent gaussian splats. arXiv preprint arXiv:2510.08575 (2025)

  37. [45]

    In: CVPR

    Yang, N., Stumberg, L.v., Wang, R., Cremers, D.: D3VO: Deep depth, deep pose and deep uncertainty for monocular visual odometry. In: CVPR. pp. 1281–1292 (2020)

  38. [46]

    In: ECCV

    Yang, N., Wang, R., Stuckler, J., Cremers, D.: Deep virtual stereo odometry: Leveraging deep depth prediction for monocular direct sparse odometry. In: ECCV . pp. 817–833 (2018) 11

  39. [47]

    In: ICCV (2023)

    Yeshwanth, C., Liu, Y .C., Nießner, M., Dai, A.: ScanNet++: A high-fidelity dataset of 3D indoor scenes. In: ICCV (2023)

  40. [48]

    arXiv preprint arXiv:2509.16909 (2025)

    Yuan, Y ., Chen, Z., Li, K., Wang, W., Zhao, H.: SLAM-Former: Putting SLAM into one transformer. arXiv preprint arXiv:2509.16909 (2025)

  41. [49]

    arXiv preprint arXiv:2509.01584 (2025)

    Zhang, G., Qian, S., Wang, X., Cremers, D.: ViSTA-SLAM: Visual SLAM with symmetric two-view association. arXiv preprint arXiv:2509.01584 (2025)

  42. [50]

    arXiv preprint arXiv:2403.19549 (2024)

    Zhang, G., Sandström, E., Zhang, Y ., Patel, M., Van Gool, L., Oswald, M.R.: GlORIE- SLAM: Globally optimized RGB-only implicit encoding point cloud SLAM. arXiv preprint arXiv:2403.19549 (2024)

  43. [51]

    arXiv preprint arXiv:2411.17982 (2024)

    Zhang, W., Cheng, Q., Skuddis, D., Zeller, N., Cremers, D., Haala, N.: HI-SLAM2: Geometry- aware gaussian SLAM for fast monocular scene reconstruction. arXiv preprint arXiv:2411.17982 (2024)

  44. [52]

    Zhang, W., Sun, T., Wang, S., Cheng, Q., Haala, N.: HI-SLAM: Monocular real-time dense mapping with hybrid implicit fields. IEEE Robotics and Automation Letters9(2), 1548–1555 (2023) 12 BA-T: An Iterative Transformer for Two-View Bundle Adjustment Appendix A Architecture and T...

  45. [53]

    BA-Net [33], RAFT [34], ReSplat [44], and BA-T all share an iterative refinement flavor, but the differences are substantial

    [39] [19] [17] [18] [42] [49] (k=1/2/3/4) Running time↓(ms) 160.93 128.71 78.14 41.13 93.25 66.47 37.96 21.40 / 24.65 / 28.32 /30.92 Decoder size↓(M) / 605 765 171 227 227 11338 Peak GPU mem↓(GB) 2.43 4.93 6.52 3.23 2.72 2.02 1.761.32 E Discussion on Iteration and Refinement B...

Pith tools

Reviewed June 28, 2026 · model on record in the stance chip above.