{"id":"e4b65e80-54df-4771-9b4f-f76117bf119c","arxiv_id":"2504.17636","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A systematic benchmark shows structureless visual localization methods that use explicit classical geometry (local triangulation, semi-generalized pose solvers) clearly outperform neural relative-pose regression, and approach the accuracy of 3D-model-based localization.","lead":"This paper benchmarks methods that localize a camera in a known place using only a database of photos with known positions, without building a 3D model. It finds that classical geometry-based methods beat recent neural pose-regression methods, and that the best structureless methods are nearly as accurate as standard 3D-model-based systems.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Regression-vs-geometry claim rests on one dataset and a self-admittedly suboptimal MASt3R variant; the abstract's 'wide margin' overstates the evidence.","rationale":"The paper delivers a careful and useful benchmark of non-regression structureless families; Tables 3-6 support the relative ordering among pose triangulation, semi-generalized solvers, and local triangulation under the stated protocol. The load-bearing weakness is in the abstract's headline contrast with pose regression. That contrast relies on a single dataset, Aachen Day-Night, and on a representative (MASt3R pose align) that the authors themselves flag as suboptimal because the released implementation cannot consume known intrinsics and reference poses. The paper's own Sec. 2 also claims regression methods can excel under low-overlap conditions, yet no low-overlap or additional-dataset regression evaluation is provided. This makes the 'wide margin' statement more general than the evidence. The reader's conditional verdict already identified essentially this concern; my analysis agrees and sharpens it by connecting the single-dataset gap to the low-overlap claim in Sec. 2. A targeted multi-dataset regression evaluation plus a known-pose MASt3R variant would settle whether the finding is robust. This does not change the reader's verdict: acceptance should remain conditional on addressing the regression evidence gap.","tokens_in":26126,"tokens_out":6986,"duration_ms":69849,"concrete_test":"Run the regression baselines (MASt3R pose align, MASt3R depth + P3P, Ess. mat. (MASt3R poses), Reloc3r) on Extended CMU Seasons and at least one NAVER scene using the same top-10 EigenPlaces retrieval and the same pose-error thresholds as the non-regression methods. Additionally, implement a 'known-pose' variant of MASt3R pose align that feeds database intrinsics and reference poses into the optimization, or replaces Kabsch alignment with pose-graph refinement constrained by reference poses, and recompute the Aachen day/night results. If regression recalls on the additional datasets remain far below E5+1 and Local triangulation, and the known-pose variant does not substantially narrow the Aachen gap, the paper's ranking holds. If either condition fails, the 'wide margin' claim should be qualified as setup- and dataset-dependent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central comparative claim that classical geometry-based structureless methods outperform pose-regression methods by a wide margin rests entirely on Table 1, which covers only Aachen Day-Night v1.1. The paper explicitly states in Sec. 4.1: 'Given the inaccurate pose estimates observed for the Aachen dataset, coupled with long run-times, we did not evaluate the pose regression-based approaches on other datasets.' Thus the family-level ranking is not tested on Extended CMU Seasons or the NAVER indoor scenes, where retrieval overlap, scene geometry, and lighting conditions differ substantially. This is not just a completeness issue: Sec. 2 itself notes that recent regression methods 'can significantly outperform classical approaches' under low visual overlap, a regime absent from the top-10-retrieval Aachen evaluation. Moreover, the main regression representative, MASt3R pose align, is acknowledged in Sec. 3 to be suboptimal because the released implementation cannot use known intrinsics and database poses; the authors describe it as measuring how well existing methods work 'out of the box.' A regression pipeline that exploited the known reference poses and intrinsics, which is the standard structureless setting, could close a material part of the gap. Therefore the abstract's unqualified 'wide margin' overstates what the current evidence supports.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a review and experimental comparison of structureless visual localization methods, which estimate a query camera pose from a database of images with known poses and no stored 3D model. The authors compare pose triangulation (Ess. mat., LazyLoc), semi-generalized relative pose estimation (E5+1, E3+1), local SfM-on-the-fly triangulation, and relative pose regression (MASt3R-based variants, Reloc3r) on Aachen Day-Night v1.1, Extended CMU Seasons, and NAVER indoor datasets. The evaluation uses multiple sparse and dense feature matchers and two depth estimators, and reports localization recall at standard thresholds, runtimes, and a comparison with structure-based baselines. The main claims are that more classical geometric reasoning tends to improve pose accuracy, that local triangulation and E5+1 are the most accurate structureless options, that regression-based methods underperform on the tested Aachen setup, and that structureless methods are competitive with, but slightly less accurate than, structure-based methods.","tokens_in":26430,"tokens_out":8325,"duration_ms":74597,"significance":"If the comparative findings hold, the paper would be a useful reference for practitioners and a benchmark for structureless localization research. Its strengths are the breadth of the evaluation across three datasets and multiple feature matchers, the consistent evaluation protocol, and the explicit acknowledgment of limitations, including the suboptimal out-of-the-box MASt3R pose alignment and the single-dataset regression evaluation. The paper does not introduce a new method, and its headline claim about regression-based methods is broader than the current evidence; the contribution is therefore primarily a well-organized empirical study whose conclusions need to be scoped precisely.","major_comments":[{"comment":"The central claim that classical geometry-based structureless methods outperform pose-regression methods 'by a wide margin' is supported only on Aachen Day-Night v1.1. Tab. 1 is the only regression evaluation; Sec. 4.1 states that the regression approaches were not evaluated on other datasets because of inaccurate pose estimates and long run-times. Since Sec. 2 itself notes that recent regression methods 'can significantly outperform classical approaches' under conditions of little visual overlap, and the Aachen top-10-retrieval setup is not such a regime, the abstract's unqualified margin claim overstates the evidence. The MASt3R pose align representative is further acknowledged in Sec. 3 to be suboptimal because the released implementation cannot use known intrinsics and database poses, so the family-level ranking is at least partly a statement about out-of-the-box pipelines. I recommend either adding regression evaluations on the other benchmarks or restricting the claim to Aachen and clearly attributing it to the evaluated variants.","section":"Abstract; Sec. 4.1, Tab. 1"},{"comment":"The accuracy-runtime trade-off conclusion relies on runtimes measured with a different configuration than the accuracy tables. Tab. 2 reports runtimes with pre-computed SuperPoint features matched by LightGlue, while the per-method best setups in Tabs. 3-6 use RoMa or MASt3R for most methods (e.g., E5+1 and Local triangulation use RoMa on Aachen; LazyLoc uses MASt3R on Aachen). Because feature matching is a dominant cost in these pipelines, the claims that 'E5+1 offers a better trade-off between pose accuracy and run-time' and that 'LazyLoc ... provides the fastest run-times' are not justified by consistent measurements. Runtimes should be reported for the exact configurations used in the accuracy tables, or the trade-off claims should be explicitly scoped to the SP+LG configuration.","section":"Sec. 4.2, Tabs. 2-6"}],"minor_comments":[{"comment":"The sentence 'E5+1 and E3+1 directly compute the query pose w.r.t. multiple database images, which further improves performance' is contradicted by E3+1's results on Aachen Day-Night (Tab. 3) and on the Extended CMU park scene (Tab. 4), where E3+1 is worse than several pose-triangulation baselines. Since the paragraph immediately discusses this scene dependence, please name E5+1 in the opening sentence or move the caveat before the claim.","section":"Sec. 4.2, Tabs. 3-4"},{"comment":"The sentence 'Given the inaccurate pose estimates observed for the Aachen dataset' should specify 'inaccurate pose estimates of the regression-based approaches', because as written it could be misread as a statement about the Aachen dataset itself.","section":"Sec. 4.1"},{"comment":"The claim that the MASt3R limitation 'applies to all other 3D reconstruction approaches based on the relative pose regression [132, 129, 40]' cites Reloc3r [40], which is a relative pose regression method but not a 3D reconstruction approach; the citation should be adjusted to avoid implying Reloc3r builds a 3D model.","section":"Sec. 3"},{"comment":"There are several typos and formatting inconsistencies: 'NA VER' should be 'NAVER', 'RoMA' should be 'RoMa', 'ocal depth errors' is missing 'l', 'MAST3R depth' in the Fig. 3 caption should be 'MASt3R depth', and 'theE5+1' in Sec. 4.2 needs a space.","section":"Throughout"},{"comment":"The phrase 'methods using less geometric reasoning can offer a better performance' is ambiguous; it should say 'better runtime performance' or 'a better accuracy-runtime trade-off'.","section":"Sec. 1, contribution (c)"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is an honest and useful empirical survey; the main revision is to align the headline claims with the evidence actually presented. I do not see citation-fairness concerns: the authors' own prior work appears as baselines but is not privileged, and Ess. mat. is among the weaker structureless baselines. The paper fits the scope of the journal, and the requested changes should be achievable within a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is the first systematic cross-family comparison of structureless visual localization, and it earns its place as a reference benchmark. The geometry-family ranking — E5+1 and local triangulation ahead of essential-matrix triangulation and LazyLoc — is well supported by consistent experiments on Aachen, Extended CMU, and NAVER. The E5+1 baseline is a real contribution: prior structureless papers omitted it, and it turns out to offer the best accuracy/runtime tradeoff. The ablations over sparse/dense matchers and depth sources are thorough, runtimes are reported, and the comparison with Hloc/MeshLoc makes the structureless-vs-structure gap concrete.\n\nThe soft spot is the regression-vs-geometry claim, which the abstract states as a 'wide margin' win for classical pose estimation. That claim rests on regression results only on Aachen Day-Night — the paper says it skipped other datasets due to inaccuracy and runtime — and the main regression representative, MASt3R pose align, is self-admittedly suboptimal because the released implementation cannot use known intrinsics and database poses. The paper's own related work notes that recent regression methods can outperform classical approaches under low visual overlap, a regime not tested here. So the family-level verdict on regression is plausible but not established across the benchmark suite. The conclusion text is properly hedged ('from our evaluation'); the abstract is not.\n\nOther soft spots are minor: best-per-dataset selection is optimistic but transparent, and there is no code release, which limits reproducibility of the ablation grids. Authors do flag the suboptimal MASt3R variant clearly, and the core geometric comparisons are not affected by these issues.\n\nThis paper deserves peer review. The benchmark is valuable for practitioners choosing among structureless methods and for researchers working on relative pose regression. I'd ask the authors to either add regression results on at least one more dataset (even a subset) or rephrase the abstract to scope the regression claim to Aachen and note the suboptimal baseline. With that revision, it becomes a solid reference.","headline":"Useful first systematic benchmark of structureless localization, but the abstract's 'wide margin' regression claim rests on a single dataset and a self-admittedly suboptimal baseline.","tokens_in":26971,"tokens_out":2611,"would_cite":true,"duration_ms":24153,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper's benchmark shows that structureless visual localization is most accurate when query poses come from classical multi-view geometry rather than from neural pose regression.","keywords":["visual localization","structureless localization","relative pose estimation","pose regression","image retrieval","essential matrix","local triangulation","benchmark"],"falsifier":"Re-run the comparison with a regression pipeline that is allowed to use known intrinsics and reference poses during alignment, or that is fine-tuned to be scale-aware, and evaluate it on the other outdoor and indoor datasets used in the paper; if its localization recall matches or exceeds the E5+1 and local-triangulation results reported here, the paper's main accuracy ordering fails.","tokens_in":25951,"feed_emoji":"📍","tokens_out":6280,"duration_ms":55690,"temperature":0.7,"pith_summary":"This paper maps out a family of visual localization methods that avoid storing a 3D scene model, representing the scene only as database images with known poses. It compares four families—pose triangulation from pairwise relative poses, semi-generalized relative pose estimation, locally triangulating 3D points on the fly, and relative pose regression—across outdoor and indoor benchmarks. The central finding is that the more classical geometry a method uses, the more accurate it tends to be, and that recent regression-based approaches trail the classical families by a wide margin on the dataset where they were tested. The practical stakes are that structureless methods, which are easy to update as scenes change, can come close to the accuracy of model-based pipelines, making them a viable alternative for self-driving and augmented-reality systems.","feed_headline":"Geometry beats pose regression in structureless localization","feed_subtitle":"Classical geometric solvers stay accurate without a 3D scene model, rivaling structure-based pipelines","key_machinery":"The comparison is organized around how much explicit geometric reasoning each structureless pipeline applies to 2D-2D matches between the query and retrieved database images. The key mechanisms are: pairwise essential-matrix estimation with rotation averaging and translation triangulation; the E5+1 semi-generalized solver, which computes the query pose with respect to two database images at once (five matches to one image fix the relative pose, one match to the other fixes scale); on-the-fly triangulation of 3D points followed by a P3P absolute pose solver inside RANSAC; and regression models that predict depths or relative poses directly. The measured quantities are localization recall at standard pose-error thresholds and average runtime per query.","core_discovery":"The paper's central claim is that among structureless localization approaches, accuracy follows the amount of geometric reasoning in the pipeline. Pose triangulation, which fuses pairwise relative poses, is consistently less accurate than semi-generalized relative pose estimation, where the query pose is computed jointly with respect to two database images and the translation scale is recovered. Locally triangulating a small 3D model per query and then applying absolute pose estimation yields the best accuracy among structureless methods, effectively catching up with a mesh-based structure-based pipeline on the Aachen day split and remaining only a few points behind at night. Relative pose regression, represented by MASt3R-based pipelines and Reloc3r, performs worst on Aachen Day-Night, with even the best regression variant notably behind the classical five-point essential-matrix method. The paper also reports that the accuracy-versus-runtime trade-off favors the E5+1 solver, since local triangulation is most accurate but slower.","pith_inferences":["If the regression gap is mostly an artifact of unmodified inference, a version of MASt3R that consumes known intrinsics and reference poses could close much of the reported margin; the paper itself flags this as a suboptimal aspect of the released implementation.","Because the regression families were only evaluated on Aachen Day-Night, the accuracy ordering may not transfer to indoor or seasonal-change datasets, where day-night appearance shifts are less central.","A natural next experiment is a hybrid: use regression models to propose dense matches, then feed those matches to E5+1 or local triangulation; the ablation data suggest matches, not pose predictions, are where learned models already help.","The scene-dependent failure of the depth-aware E3+1 variant hints that adaptive selection of a depth-aware solver—trusting monocular depth only when scene structure is close to the camera—could beat either fixed choice."],"forward_implications":["Practitioners who value accuracy can use local on-the-fly triangulation as a structureless method that rivals a structure-based mesh pipeline on the Aachen day split.","Practitioners who value runtime can use the E5+1 semi-generalized solver, which keeps most of the accuracy of local triangulation at a fraction of the runtime.","Released relative pose regression models are not yet a drop-in replacement for geometric solvers in this setting: on Aachen Day-Night they trail even the pairwise essential-matrix baseline.","The best feature type is method- and scene-dependent, so a robust structureless system should expose matcher choice rather than hard-code one feature.","Structureless localization's flexibility—adding or removing database images—can be obtained with only a modest accuracy loss relative to structure-based state of the art."],"supporting_citations":[{"why":"Supplies the essential-matrix pose triangulation method and code that the paper reimplements as the pairwise baseline.","marker":"[147]"},{"why":"LazyLoc, the motion-averaging pose triangulation method compared as a state-of-the-art representative.","marker":"[39]"},{"why":"Introduces the E5+1 semi-generalized relative pose solver that anchors the best accuracy-runtime trade-off.","marker":"[143]"},{"why":"Provides the 3-point-with-depth relative pose solver used for Ess. mat. (3Pt+depth) and E3+1.","marker":"[37]"},{"why":"The MASt3R model is the main relative pose regression representative, evaluated as pose alignment, depth plus P3P, and pose input.","marker":"[65]"},{"why":"Reloc3r, the other regression-based method, evaluated via pose triangulation from predicted relative poses.","marker":"[40]"},{"why":"DUSt3R, the regression approach underlying MASt3R's depth and pose prediction.","marker":"[132]"},{"why":"RoMa dense matcher, which the ablations show is frequently the best feature choice for geometric methods.","marker":"[42]"},{"why":"EigenPlaces image-level descriptor used for retrieval by all evaluated structureless pipelines.","marker":"[14]"},{"why":"Hloc, the structure-based pipeline whose recall numbers anchor the comparison between structureless and structure-based methods.","marker":"[97]"}],"fun_headline_variants":["Geometry beats regression in structureless localization","Classical geometric solvers win in structureless pose estimation","Structureless localization: geometric reasoning outperforms regression","Local triangulation bests pose regression without a 3D model","Flexible localization: geometric solvers rival structure-based accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ranking of pose regression as the least accurate family assumes that the released MASt3R pipeline, which cannot exploit the known intrinsics and poses of database images, is a fair representative of that family.","fun_headline_variants_meta":{"raw":{"variants":["Geometry beats regression in structureless localization","Classical geometric solvers win in structureless pose estimation","Structureless localization: geometric reasoning outperforms regression","Local triangulation bests pose regression without a 3D model","Flexible localization: geometric solvers rival structure-based accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000292,"raw_usage":{"total_tokens":1724,"prompt_tokens":986,"completion_tokens":738,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":602,"completion_tokens_details":{"reasoning_tokens":661}},"tokens_in":602,"tokens_out":738,"duration_ms":6728,"temperature":1.0,"reasoning_tokens":661,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:35:03.720652+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the comparison with a regression pipeline that is allowed to use known intrinsics and reference poses during alignment, or that is fine-tuned to be scale-aware, and evaluate it on the other outdoor and indoor datasets used in the paper; if its localization recall matches or exceeds the E5+1 and local-triangulation results reported here, the paper's main accuracy ordering fails.","supporting_citations":[{"cited_title":"To Learn or Not to Learn: Visual Localization from Essential Matrices","cited_arxiv_id":null,"evidence_quote":"Supplies the essential-matrix pose triangulation method and code that the paper reimplements as the pairwise baseline."},{"cited_title":"Structure From Motion Using Structure-Less Resection","cited_arxiv_id":null,"evidence_quote":"Introduces the E5+1 semi-generalized relative pose solver that anchors the best accuracy-runtime trade-off."},{"cited_title":"DUSt3R: Geometric 3D Vision Made Easy","cited_arxiv_id":null,"evidence_quote":"DUSt3R, the regression approach underlying MASt3R's depth and pose prediction."},{"cited_title":"From Coarse to Fine: Robust Hierarchical Localization at Large Scale","cited_arxiv_id":null,"evidence_quote":"Hloc, the structure-based pipeline whose recall numbers anchor the comparison between structureless and structure-based methods."}],"review_version":1}