Pith. sign in

REVIEW 3 major objections 5 minor 79 references

Aerial-ground Cross-modal Localization: Dataset, Ground-truth, and Benchmark

T0 review · 3 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read The paper introduces a large-scale benchmark that matches ground-level street imagery to airborne laser scanning point clouds in Wuhan, Hong Kong, and San Francisco, using an indirect ground-truth pipeline reported at 9–16 cm checkpoint acc

desk verdict A genuinely useful ground-image-to-ALS dataset, but the 9–16 cm ground-truth accuracy claim is a self-consistency measure until validated against an independent reference. read the letter →

arxiv 2509.07362 v1 pith:VV5UW7HC submitted 2025-09-09 cs.RO

classification cs.RO
keywords aerial-groundcross-modallocalizationairbornelaserscanning(ALS)mobilemappingsystemimage-to-point-cloudregistration6-DoFground-truthposesposegraphoptimizationurbanbenchmarkmulti-sensorfusion
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The authors set out to establish that airborne laser scanning (ALS) data—publicly available over many cities—can serve as a reliable prior map for visual localization, provided researchers have a dataset with trustworthy image poses. Their load-bearing contribution is a new large-scale dataset pairing ground-level mobile-mapping images with ALS point clouds across three dense urban areas. Because direct image-to-ALS registration is too hard, they generate ground-truth image poses indirectly: align vehicle-mounted LiDAR submaps to the ALS via ground segmentation, façade reconstruction, and pose-graph optimization, then transfer the refined trajectory to the camera through a rigid calibration. Checkpoint evaluations report average errors of 0.09–0.16 m. On this benchmark, current image-to-point-cloud localization methods perform poorly, especially the state-of-the-art fine-registration methods, showing that aerial-ground cross-modal localization is a live open problem.

What carries the argument

The load-bearing mechanism is the tight multi-sensor pose graph whose nodes are full LiDAR-frame states (pose, velocity, IMU biases) and whose edges include an aerial-ground constraint: ICP registration of MLS submaps to ALS-derived façade and ground features. That aerial-ground residual, together with loop, IMU, odometry, and GNSS factors, converts the globally referenced ALS into an absolute anchor that corrects drift in the ground trajectory. Compatibility between MLS and ALS is manufactured by extracting planar ground seeds and by projecting roof boundaries onto the ground to complete façades barely visible from above.

What would settle it

Take one sequence from each city and survey a set of camera locations with an independent instrument not involved in the optimization—total-station ground-control targets, or high-accuracy RTK-GNSS positions collected after the fact—then compare those surveyed poses to the published ground-truth poses. If disagreement substantially exceeds the reported 0.09–0.16 m average, the claimed ground-truth accuracy is not supported; if it matches, the indirect MLS-to-ALS pipeline is validated as an independent check.

Watch

Extended reading notes

Core claim

The central claim is that accurate 6-DoF ground-truth poses for ground images relative to ALS point clouds can be obtained without direct image-to-point-cloud registration. The method first registers mobile laser scanning (MLS) submaps to pre-georeferenced ALS data using ground planes and reconstructed building façades as cross-modal features, then solves a multi-sensor pose graph that fuses loop closures, aerial-ground ICP residuals, IMU pre-integration, odometry, and GNSS constraints. The rigid LiDAR-camera mounting then transfers the optimized MLS trajectory to the image stream. Across the three cities, average checkpoint errors are 0.09 m in Wuhan, 0.11 m in Hong Kong, and 0.16 m in San

Load-bearing premise

The load-bearing premise is that manual checkpoint alignment on the same MLS and ALS clouds used to build the pose-graph constraints gives an independent accuracy estimate; if that alignment or the underlying ALS georeferencing is biased, the claimed 9–16 cm ground-truth error inherits the bias.

Editorial extensions

If this is right

  • The public dataset gives researchers a standardized way to train and evaluate image-to-point-cloud localization across a genuine aerial-ground platform gap, not just same-platform scenarios.
  • If the reported checkpoint accuracy holds, the dataset supplies ground-truth image poses at centimeter-to-decimeter level for three dense urban environments, enabling reliable benchmarking and supervision.
  • Global I2P methods that use projected range images as a proxy bridge the modality gap better than methods consuming raw point clouds, indicating a productive design direction.
  • Fine-grained I2P registration methods that work on vehicle-borne point clouds do not transfer to ALS-based maps, so the benchmark defines a concrete unsolved subproblem.
  • The trajectory refinement also improves the underlying MLS point-cloud quality, meaning the dataset can support point-cloud quality assessment beyond localization.
  • The performance gap between ALS and vehicle-borne input suggests that airborne prior maps remain under-exploited: with better cross-modal features, ALS-based localization could scale to any city with government LiDAR coverage.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the manual checkpoints are selected on the same MLS and ALS clouds that the pose graph already aligned, the reported 9–16 cm figures are best read as consistency checks rather than independent accuracy bounds; true absolute error could be larger if the ALS georeferencing itself carries systematic bias.
  • The benchmark’s fine-registration failures might not be purely a modality-gap problem: the learning-based baselines were originally trained on denser, structure-rich vehicle LiDAR, so the benchmark could also be used to isolate how much of the failure is due to scene geometry versus platform discrepancy.
  • A testable extension is change-aware evaluation: several ALS datasets were collected years apart from the ground imagery, so temporal change can be measured per sequence, and descriptors or registration methods could be scored separately in changed versus stable regions.
  • The ground-truth generation strategy generalizes beyond this dataset: any city with public ALS and a mobile-mapping trajectory could receive the same treatment, making ALS-based localization benchmarks feasible government-data scale rather than research-collection scale.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper introduces a dataset for aerial-ground cross-modal localization, combining ground-level imagery from mobile mapping systems with airborne laser scanning (ALS) point clouds in Wuhan, Hong Kong, and San Francisco. The ground-truth 6-DoF image poses are generated indirectly: mobile LiDAR (MLS) submaps are aligned to ALS point clouds via ground segmentation, façade reconstruction, and multi-sensor pose-graph optimization that fuses loop-closure, aerial-ground ICP, IMU, GNSS, and odometry factors. The optimized MLS trajectory is then transferred to the camera stream through extrinsic calibration. The authors report average checkpoint errors of 0.09–0.16 m for three evaluation sequences (Table 2) and use the dataset to benchmark three global I2P localization methods (AE-Spherical, LIPLoc, SaliencyI2PLoc) and three fine-registration methods (DeepI2P, CorrI2P, CoFiI2P). LIPLoc performs best on global retrieval; the fine-registration baselines fail, so an SfM+VINS+ICP pipeline is used as an alternative baseline.

Significance. If the ground-truth accuracy claim is trustworthy, this dataset would be a valuable resource for evaluating image-to-point-cloud localization across aerial and ground platforms, a setting that is underrepresented in existing benchmarks. The paper contributes a large-scale multi-city dataset, an indirect alignment pipeline that avoids laborious ground surveying, and a reproducible evaluation of several state-of-the-art methods. The main significance, however, hinges on the reliability of the claimed 9–16 cm pose accuracy. The current validation is not independent of the optimization that generated the poses, so the absolute accuracy of the dataset remains unestablished. If the validation is strengthened or the claim appropriately qualified, the benchmark would still be useful for comparing methods, but the paper's central quantitative claim currently exceeds what the evidence supports.

major comments (3)
  1. [§5.2 and Eq. (8)] The quantitative validation of the ground-truth poses is not independent of the optimization that produced them. The checkpoint evaluation in Section 5.2 manually aligns selected feature points between the same MLS and ALS point clouds that define the aerial-ground constraints in Eq. (8). If the ALS frame carries a global translation or rotation error, or if the façade-completion procedure biases the extracted ALS features, both the optimized trajectory and the checkpoint alignment are affected in the same way, so the residuals in Table 2 cannot detect such errors. The reported 0.09–0.16 m errors should therefore be described as self-consistency residuals, not absolute accuracy. To support the abstract's accuracy claim, the authors need an independent reference (surveyed ground control points, total-station measurements, or an independent high-accuracy GNSS/INS trajectory) and a report o
  2. [§3.3 and Table 2] The claimed ground-truth accuracy is below the stated accuracy of the ALS reference. Table 2 reports average checkpoint errors of 0.16 m (California), 0.11 m (Hong Kong), and 0.09 m (Wuhan), while Section 3.3 gives ALS horizontal accuracies of 0.12 m, 0.3 m, and 0.2 m for the same sites. Since the pose-graph optimization references the MLS trajectory to the ALS frame (Eq. 8), and since the GNSS factor (Eq. 12) in dense urban canyons is typically too weak to override a global ALS shift, the absolute pose error cannot in general be smaller than the ALS georeferencing error. The checkpoint metric as defined in Section 5.2 excludes such global shifts. The paper should either provide external absolute measurements or explicitly restate the claimed accuracy as 'relative to the ALS frame' and quantify the propagation of ALS georeferencing uncertainty into the final poses.
  3. [§5.2] The checkpoint evaluation is under-reported. The number of checkpoints per sequence, the selection criteria, the manual alignment procedure, and the operator variability are not given; only average/min/max errors for three sequences are listed in Table 2. With such a small and unquantified sample, a 9–16 cm average error is not a statistically robust accuracy certificate. This matters because Table 2 is the only quantitative support for the central ground-truth claim.
minor comments (5)
  1. [Table 3] The table is garbled: the column headers repeat 'HK' and the train/evaluation counts for the three datasets are not aligned with the rows. Please reformat it so that the split sizes for California, Hong Kong, and Wuhan are unambiguous.
  2. [§5.2] The phrase 'According to2' (before 'in the dense urban datasets') should be a proper cross-reference to Table 2.
  3. [§3.4] The definition of 'patches of 100 m2' is ambiguous; please specify whether this is 100 m × 100 m or another shape, and clarify how patches overlap between adjacent images.
  4. [§6.2] The fine-localization section reports that all three learning-based methods 'failed to produce reliable correspondences' but then does not report their quantitative results; the SfM+VINS+ICP baseline in Table 6 is then presented as the benchmark. Please make explicit that Table 6 is a baseline sanity check rather than a comparison of the three selected methods.
  5. [§6.1] The AE-Spherical training uses positive/negative distance thresholds of 20 m and 40 m without sensitivity analysis; given the strong effect of these thresholds on place-recognition recall, this is worth a sentence of justification.

Circularity Check

1 steps flagged · score 6.0 of 10

Ground-truth accuracy claim is a self-fit: checkpoint errors (Table 2) measure the same MLS–ALS alignment that Eq. (8) optimizes, not an independent absolute check.

  1. fitted input called prediction [Section 4.3, Eq. (8); Section 5.2 and Table 2]
    "Aerial-ground constraints are derived by registering MLS submaps to ALS point clouds using Iterative Closest Point (ICP). The alignment residual is: r_i^aerial = [log(R̂^{-1}_aerial R_i); t_i − t̂_aerial] (8). ... several checkpoints were manually selected and aligned between the MLS and ALS point clouds to evaluate the accuracy of the trajectory. ... the average checkpoint errors are 11cm and 9cm ... California Bay Bridge yields a higher average error of 0.16m."

    The pose-graph cost includes the aerial-ground residual (Eq. 8), which is minimized by registering each MLS submap to the ALS cloud via ICP. Section 5.2 then validates the optimized trajectory by manually aligning checkpoints between the same MLS and ALS point clouds and reporting the mismatch as checkpoint error (Table 2). This is the same MLS–ALS alignment that Eq. (8) fits, so the 0.09–0.16 m errors are post-fit residuals of the aerial-ground constraint, not an independent measurement of absolute pose accuracy. A systematic ALS georeferencing error (0.12–0.3 m horizontal, Section 3.3) would appear both in the optimized GT poses and in the checkpoints, and would be invisible in Table 2. Hence the claimed accuracy reduces, by construction, to the residual of the very alignment used to def

full rationale

The paper's central contribution is a dataset with accurate 6-DoF ground-truth poses obtained by aligning MLS submaps to ALS point clouds through pose-graph optimization. The load-bearing numerical claim is the 0.09–0.16 m checkpoint accuracy reported in Table 2. That validation is not independent: the pose graph minimizes aerial-ground residuals (Eq. 8) that register the MLS submaps to the ALS cloud, and the Section 5.2 checkpoints are manually selected and aligned between those same MLS and ALS point clouds. The resulting errors are therefore residuals of the same alignment the optimization was designed to fit, so they cannot detect systematic bias in the ALS reference frame. The pose graph does also fuse IMU, GNSS, odometry, and loop-closure factors, so the trajectory is not purely a fit to the checkpoint targets; nevertheless, the absolute accuracy claim is not separately established. The I2P benchmark evaluations themselves are not circular—they compare external algorithms against the generated GT—but they inherit any bias in that GT. No load-bearing self-citation or uniqueness-theorem circularity is present; the main issue is the self-referential validation of the GT accuracy. The score reflects that this is a partial, material circularity in the central accuracy claim rather than a fully forced derivation.

Assumptions & free parameters 6 free parameters · 4 assumptions · 0 invented entities

The central claim, that the dataset provides reliable 6-DoF ground-truth poses, rests on the accuracy of the public ALS data, the extrinsic calibration of the source datasets, and the validity of the manual checkpoint evaluation. These are external inputs the paper trusts rather than derives.

free parameters (6)
  • planarity threshold for roof classification = 0.5
    Section 4.2: supervoxels with planarity > 0.5 are classified as building roof points, controlling which ALS points constrain the pose graph; hand-chosen, not tuned to independent data.
  • verticality threshold for roof classification = 0.3
    Section 4.2: used with planarity to extract facades; hand-chosen threshold that affects the aerial-ground feature set.
  • ALS downsampling resolution = 0.5 m
    Section 3.4: point clouds downsampled to 0.5 m spacing before patching; influences registration precision.
  • patch size = 100 m^2
    Section 3.4: ALS divided into 100 m^2 patches; affects retrieval difficulty in benchmarks.
  • image sampling interval = 0.5 m
    Section 3.4: images sampled every 0.5 m along trajectory; determines number of pairs.
  • positive/negative distance thresholds for AE-Spherical training = 20 m and 40 m
    Section 6.1: used to define positives and negatives in contrastive training; directly affects reported recall.
assumptions (4)
  • domain assumption The public ALS point clouds are accurately georeferenced at the stated accuracies (HK 0.1-0.3 m, SF ~0.12-0.2 m, Wuhan ~0.2 m)
    Section 3.3: The ground-truth poses are defined in the ALS mapping frame, so any systematic error in ALS georeferencing propagates directly into the image poses.
  • domain assumption The camera-LiDAR extrinsics and camera intrinsics from UrbanNav, UrbanLoco, and WHU-Helmet are accurate enough to transfer optimized LiDAR poses to images
    Section 3.4 and Eq. (1): image poses are computed by applying the fixed extrinsics to the optimized LiDAR trajectory; calibration error shifts every image pose.
  • domain assumption The initial trajectories from SPAN-CPT (UrbanNav/UrbanLoco) and WHU-Helmet are close enough for ICP and pose graph optimization to converge to the global optimum
    Section 4.3: the aerial-ground residuals (Eq. 8) are local ICP alignments; poor initialization could trap the optimization in local minima, especially in the hilly San Francisco data.
  • domain assumption Manual checkpoint alignment between the optimized MLS clouds and ALS clouds gives an unbiased estimate of absolute pose accuracy
    Section 5.2: the reported errors (Table 2) rely on human picking of corresponding corners; this is not validated against any independent absolute reference.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Aerial-ground Cross-modal Localization: Dataset, Ground-truth, and Benchmark." pith.science (2026). https://pith.science/paper/VV5UW7HC

@misc{pith2026250907362,
  author       = {Pith},
  title        = {Pith review of: Aerial-ground Cross-modal Localization: Dataset, Ground-truth, and Benchmark},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VV5UW7HC}},
  note         = {Machine review of arXiv:2509.07362}
}
read the original abstract

Accurate visual localization in dense urban environments poses a fundamental task in photogrammetry, geospatial information science, and robotics. While imagery is a low-cost and widely accessible sensing modality, its effectiveness on visual odometry is often limited by textureless surfaces, severe viewpoint changes, and long-term drift. The growing public availability of airborne laser scanning (ALS) data opens new avenues for scalable and precise visual localization by leveraging ALS as a prior map. However, the potential of ALS-based localization remains underexplored due to three key limitations: (1) the lack of platform-diverse datasets, (2) the absence of reliable ground-truth generation methods applicable to large-scale urban environments, and (3) limited validation of existing Image-to-Point Cloud (I2P) algorithms under aerial-ground cross-platform settings. To overcome these challenges, we introduce a new large-scale dataset that integrates ground-level imagery from mobile mapping systems with ALS point clouds collected in Wuhan, Hong Kong, and San Francisco.

Figures

Figures reproduced from arXiv: 2509.07362 by the authors.

Figure 1
Figure 1. Global distribution of our dataset on the map. Despite its potential, ALS-based visual localization in urban environments remains underexplored due to three major challenges. First, most publicly available datasets lack platform diversity, particularly the integration of both ground and aerial sensing platforms. To date, there is no dataset specifically designed to support image-based localization using ALS data in … view at source ↗
Figure 2
Figure 2. Dataset coverage and collection. (a), (b) and (c) illustrate trajectories of the Hong Kong, California and Wuhan datasets, while (d), (e) and (f) are the corresponding data acquisition platforms, with (d) and (e) provided by authors of UrbanNav (Hsu et al., 2023) and UrbanLoco (Wen et al., 2020), respectively. Tsui and Nathan Road, two of Hong Kong’s busiest commer￾cial districts, as well as Whampoa, a large-scale r… view at source ↗
Figure 3
Figure 3. File structure of the dataset (taking Wuhan Loop 1 as an example). along sidewalks in dense urban environments, with a total trajectory length exceeding 11 km. 3.3. Aerial platform The aerial data incorporated into our dataset consists of publicly available ALS point clouds obtained from govern￾ment agencies in Hong Kong (CEDD, Hong Kong, 2020) and San Francisco (USGS, U.S.A., 2023). These 3D point clouds are acquir… view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Definition of coordinate frames. ALS to image projection: To enable the projection of ALS patches onto ground images, the dataset provides the extrinsic calibration between the ground LiDAR and camera, as well as the camera intrinsics. In addition, the file patch_index…
Figure 5
Figure 5. Figure 5: The flowchart of ALS feature extraction. Gray points represent the original ALS point clouds (rendered in grayscale based on elevation values) while red points indicate the projected façade points. 4.1. Formulation Let 𝐗𝑘 denote the state of the 𝑘-th point cloud frame.…
Figure 6
Figure 6. Figure 6: Pose graph formulation for ground-truth generation. Aerial-ground constraints are derived by registering MLS submaps to ALS point clouds using Iterative Closest Point (ICP). The alignment residual is: 𝐫 𝑖 aerial = [ log ( 𝐑̂ −1 aerial𝐑𝑖 ) 𝐭 𝑖 − ̂𝐭 aerial ] . (8) Odomet…
Figure 7
Figure 7. Figure 7: Projection of ALS point clouds to images. (a), (b) and (c) are from Wuhan, Hong Kong and California datasets, respectively. Point clouds are colorized by depth, with colors ranging from blue (near) to red (far), through green and yellow [PITH_FULL_IMAGE:figures/full_f…
Figure 8
Figure 8. Figure 8: Point cloud refinement of the Hong Kong Medium dataset. (a) and (b) illustrate the MLS point clouds before and after optimization. ALS points are rendered in grayscale to represent relative elevation, while green triangles indicate the check points. clouds as the map r…
Figure 9
Figure 9. Figure 9: Point cloud refinement of California Bay Bridge dataset. (a) and (b) illustrate the MLS point clouds before and after optimization. ALS points are rendered in gray to represent relative height, while green triangles indicate the check points. for feature extraction, fo…
Figure 10
Figure 10. Figure 10: Point cloud refinement of the Wuhan Loop 1 dataset. (a) and (b) illustrate the MLS point clouds before and after optimization. ALS points are rendered in grayscale to represent relative elevation, while green triangles indicate the check points [PITH_FULL_IMAGE:figur…
Figure 11
Figure 11. Figure 11: I2P fine localization results. Dark points are from dense matching of images, and ALS point clouds are colorized by height. refines rotation and translation parameters. This design en￾hances robustness under large viewpoint changes and modal￾ity gaps. We retrained the…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

79 extracted references · 71 canonical work pages

  1. [1]

    , author Gronat, P

    author Arandjelovi \'c , R. , author Gronat, P. , author Torii, A. , author Pajdla, T. , author Sivic, J. , year 2018 . title Netvlad: Cnn architecture for weakly supervised place recognition . journal IEEE Transactions on Pattern Analysis and Machine Intelligence volume 40 , pages 1437--1451 . :10.1109/TPAMI.2017.2711011

  2. [2]

    , author Maturana, D

    author Aubry, M. , author Maturana, D. , author Efros, A.A. , author Russell, B.C. , author Sivic, J. , year 2014 . title Seeing 3d chairs: exemplar part-based 2d-3d alignment using a large dataset of cad models , in: booktitle Proceedings of the IEEE conference on computer vision and pattern recognition , pp. pages 3762--3769

  3. [3]

    , author Jiang, G

    author Bai, Z. , author Jiang, G. , author Xu, A. , year 2020 . title Lidar-camera calibration using line correspondences . journal Sensors volume 20 , pages 6319

  4. [4]

    , author Bankiti, V

    author Caesar, H. , author Bankiti, V. , author Lang, A.H. , author Vora, S. , author Liong, V.E. , author Xu, Q. , author Krishnan, A. , author Pan, Y. , author Baldan, G. , author Beijbom, O. , year 2020 . title nuscenes: A multimodal dataset for autonomous driving , in: booktitle Proceedings of the IEEE/CVF conference on computer vision and pattern rec...

  5. [5]

    , author Ushani, A.K

    author Carlevaris-Bianco, N. , author Ushani, A.K. , author Eustice, R.M. , year 2016 . title University of michigan north campus long-term vision and lidar dataset . journal The International Journal of Robotics Research volume 35 , pages 1023--1035

  6. [6]

    title Hong Kong LiDAR Data

    author CEDD, Hong Kong , year 2020 . title Hong Kong LiDAR Data . howpublished https://sdportal.cedd.gov.hk

  7. [7]

    , author Sun, Y

    author Chai, Z. , author Sun, Y. , author Xiong, Z. , year 2018 . title A novel method for lidar camera calibration by plane fitting , in: booktitle 2018 IEEE/ASME International Conference on Advanced Intelligent Mechatronics (AIM) , organization IEEE . pp. pages 286--291

  8. [8]

    , author Cladera, F

    author Chaney, K. , author Cladera, F. , author Wang, Z. , author Bisulco, A. , author Hsieh, M.A. , author Korpela, C. , author Kumar, V. , author Taylor, C.J. , author Daniilidis, K. , year 2023 . title M3ed: Multi-robot, multi-sensor, multi-environment event dataset , in: booktitle 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition Wor...

Show all 79 references
  1. [9]

    , author Mangelson, J

    author Chang, M.F. , author Mangelson, J. , author Kaess, M. , author Lucey, S. , year 2021 . title Hypermap: Compressed 3d map for monocular camera registration , in: booktitle 2021 IEEE International Conference on Robotics and Automation (ICRA) , organization IEEE . pp. page...

  2. [10]

    , year 2018

    author Davis, T.A. , year 2018 . title Graph algorithms via suitesparse: Graphblas: triangle counting and k-truss , in: booktitle 2018 IEEE High Performance extreme Computing Conference (HPEC) , organization IEEE . pp. pages 1--6

  3. [11]

    , author Vallet, B

    author Demantk \'e , J. , author Vallet, B. , author Paparoditis, N. , year 2012 . title Streamed vertical rectangle detection in terrestrial laser scans for facade database production . journal ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Science...

  4. [12]

    , author Eudes, A

    author Dubois, R. , author Eudes, A. , author Frémont, V. , year 2020 . title Airmuseum: a heterogeneous multi-robot dataset for stereo-visual and inertial simultaneous localization and mapping , in: booktitle 2020 IEEE International Conference on Multisensor Fusion and Integr...

  5. [13]

    , author Youssef, A

    author El-Sheimy, N. , author Youssef, A. , year 2020 . title Inertial sensors technologies for navigation applications: State of the art and future trends . journal Satellite navigation volume 1 , pages 2

  6. [14]

    , author Qi, Y

    author Feng, D. , author Qi, Y. , author Zhong, S. , author Chen, Z. , author Chen, Q. , author Chen, H. , author Wu, J. , author Ma, J. , year 2024 . title S3e: A multi-robot multimodal dataset for collaborative slam . journal IEEE Robotics and Automation Letters , pages 1--8...

  7. [15]

    , year 1981

    author FISCHLER AND, M. , year 1981 . title Random sample consensus: a paradigm for model fitting with applications to image analysis and automated cartography . journal Commun. ACM volume 24 , pages 381--395

  8. [16]

    , author Lenz, P

    author Geiger, A. , author Lenz, P. , author Stiller, C. , author Urtasun, R. , year 2013 . title Vision meets robotics: The kitti dataset . journal The international journal of robotics research volume 32 , pages 1231--1237

  9. [17]

    , author Lenz, P

    author Geiger, A. , author Lenz, P. , author Urtasun, R. , year 2012 . title Are we ready for autonomous driving? the kitti vision benchmark suite , in: booktitle 2012 IEEE conference on computer vision and pattern recognition , organization IEEE . pp. pages 3354--3361

  10. [18]

    , author Muthuselvam, A

    author Guan, T. , author Muthuselvam, A. , author Hoover, M. , author Wang, X. , author Liang, J. , author Sathyamoorthy, A.J. , author Conover, D. , author Manocha, D. , year 2023 . title Crossloc3d: Aerial-ground cross-source 3d place recognition , in: booktitle Proceedings ...

  11. [19]

    , author Morin, K

    author Helmberger, M. , author Morin, K. , author Berner, B. , author Kumar, N. , author Cioffi, G. , author Scaramuzza, D. , year 2022 . title The hilti slam challenge dataset . journal IEEE Robotics and Automation Letters volume 7 , pages 7518--7525

  12. [20]

    , author Huang, F

    author Hsu, L.T. , author Huang, F. , author Ng, H.F. , author Zhang, G. , author Zhong, Y. , author Bai, X. , author Wen, W. , year 2023 . title Hong kong urbannav: An open-source multisensory dataset for benchmarking urban navigation algorithms . journal NAVIGATION: Journal ...

  13. [21]

    , author Shen, L

    author Hu, J. , author Shen, L. , author Sun, G. , year 2018 . title Squeeze-and-excitation networks , in: booktitle 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition , pp. pages 7132--7141 . :10.1109/CVPR.2018.00745

  14. [22]

    , author Zhang, X

    author Huang, Z. , author Zhang, X. , author Garcia, A. , author Huang, X. , year 2024 . title A novel, efficient and accurate method for lidar camera calibration , in: booktitle 2024 IEEE International Conference on Robotics and Automation (ICRA) , organization IEEE . pp. pag...

  15. [23]

    , author Nex, F

    author Jende, P. , author Nex, F. , author Gerke, M. , author Vosselman, G. , year 2018 . title A fully automatic approach to register mobile mapping and airborne imagery to support the correction of platform trajectories in gnss-denied urban areas . journal ISPRS journal of p...

  16. [24]

    , author Shin, J

    author Jeong, H. , author Shin, J. , author Rameau, F. , author Kum, D. , year 2024 . title Multi-modal place recognition via vectorized hd maps and images fusion for autonomous driving . journal IEEE Robotics and Automation Letters

  17. [25]

    , author Liao, Y

    author Kang, S. , author Liao, Y. , author Li, J. , author Liang, F. , author Li, Y. , author Zou, X. , author Li, F. , author Chen, X. , author Dong, Z. , author Yang, B. , year 2024 . title Cofii2p: Coarse-to-fine correspondences-based image to point cloud registration . jou...

  18. [26]

    , author Jeong, J

    author Kim, Y. , author Jeong, J. , author Kim, A. , year 2018 . title Stereo camera localization in 3d lidar maps , in: booktitle 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , organization IEEE . pp. pages 1--9

  19. [27]

    , author Zou, Y

    author Li, H. , author Zou, Y. , author Chen, N. , author Lin, J. , author Liu, X. , author Xu, W. , author Zheng, C. , author Li, R. , author He, D. , author Kong, F. , et al., year 2024 a. title Mars-lvig dataset: A multi-sensor aerial robots slam dataset for lidar-visual-in...

  20. [28]

    , author Lee, G.H

    author Li, J. , author Lee, G.H. , year 2021 . title Deepi2p: Image-to-point cloud registration via deep classification , in: booktitle Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pp. pages 15960--15969

  21. [29]

    , author Nguyen, T.M

    author Li, J. , author Nguyen, T.M. , author Cao, M. , author Yuan, S. , author Hung, T.Y. , author Xie, L. , year 2025 a. title Graph optimality-aware stochastic lidar bundle adjustment with progressive spatial smoothing . journal IEEE Transactions on Intelligent Transportati...

  22. [30]

    , author Wu, W

    author Li, J. , author Wu, W. , author Yang, B. , author Zou, X. , author Yang, Y. , author Zhao, X. , author Dong, Z. , year 2023 a. title Whu-helmet: A helmet-based multisensor slam dataset for the evaluation of real-time 3-d mapping in large-scale gnss-denied environments ....

  23. [31]

    , author Yuan, S

    author Li, J. , author Yuan, S. , author Cao, M. , author Nguyen, T.M. , author Cao, K. , author Xie, L. , year 2024 b. title Hcto: Optimality-aware lidar inertial odometry with hybrid continuous time optimization for compact wearable mapping system . journal ISPRS Journal of ...

  24. [32]

    , author Ma, Y

    author Li, L. , author Ma, Y. , author Tang, K. , author Zhao, X. , author Chen, C. , author Huang, J. , author Mei, J. , author Liu, Y. , year 2023 b. title Geo-localization with transformer-based 2d-3d match network . journal IEEE Robotics and Automation Letters volume 8 , p...

  25. [33]

    , author Duan, Y

    author Li, X. , author Duan, Y. , author Wang, B. , author Ren, H. , author You, G. , author Sheng, Y. , author Ji, J. , author Zhang, Y. , year 2024 c. title Edgecalib: Multi-frame weighted edge features for automatic targetless lidar-camera calibration . journal IEEE Robotic...

  26. [34]

    , author Li, J

    author Li, Y. , author Li, J. , author Dong, Z. , author Wang, Y. , author Yang, B. , year 2025 b. title Saliencyi2ploc: Saliency-guided image–point cloud localization using contrastive learning . journal Information Fusion volume 118 , pages 103015 . :10.1016/j.inffus.2025.103015

  27. [35]

    , author Chen, X

    author Liao, Y. , author Chen, X. , author Kang, S. , author Li, J. , author Dong, Z. , author Fan, H. , author Yang, B. , year 2024 a. title Osmloc: Single image-based visual localization in openstreetmap with geometric and semantic guidances . journal arXiv preprint arXiv:2411.08665

  28. [36]

    , author Kang, S

    author Liao, Y. , author Kang, S. , author Li, J. , author Liu, Y. , author Liu, Y. , author Dong, Z. , author Yang, B. , author Chen, X. , year 2024 b. title Mobile-seed: Joint semantic segmentation and boundary detection for mobile robots . journal IEEE Robotics and Automati...

  29. [37]

    , author Wang, C

    author Lin, Y. , author Wang, C. , author Zhai, D. , author Li, W. , author Li, J. , year 2018 . title Toward better boundary preserved supervoxel segmentation for 3d point clouds . journal ISPRS journal of photogrammetry and remote sensing volume 143 , pages 39--47

  30. [38]

    , author Cao, C

    author Liu, H. , author Cao, C. , author Ye, H. , author Cui, H. , author Gao, W. , author Wang, X. , author Shen, S. , year 2024 . title Lightweight structured line map based visual localization . journal IEEE Robotics and Automation Letters volume 9 , pages 5182--5189 . :10....

  31. [39]

    , author Yan, G

    author Luo, Z. , author Yan, G. , author Cai, X. , author Shi, B. , year 2024 . title Zero-training lidar-camera extrinsic calibration method using segment anything model , in: booktitle 2024 IEEE International Conference on Robotics and Automation (ICRA) , pp. pages 14472--14...

  32. [40]

    , author Till, C

    author Majdik, A.L. , author Till, C. , author Scaramuzza, D. , year 2017 . title The zurich urban micro aerial vehicle dataset . journal The International Journal of Robotics Research volume 36 , pages 269--273

  33. [41]

    , author Zhang, L

    author Mao, Q. , author Zhang, L. , author Li, Q. , author Hu, Q. , author Yu, J. , author Feng, S. , author Ochieng, W. , author Gong, H. , year 2015 . title A least squares collocation method for accuracy improvement of mobile lidar systems . journal Remote sensing volume 7 ...

  34. [42]

    , author Cioffi, G

    author Merat, R. , author Cioffi, G. , author Bauersfeld, L. , author Scaramuzza, D. , year 2024 . title Drift-free visual slam using digital twins . journal IEEE Robotics and Automation Letters

  35. [43]

    , year 2006

    author Mor \'e , J.J. , year 2006 . title The levenberg-marquardt algorithm: implementation and theory , in: booktitle Numerical analysis: proceedings of the biennial Conference held at Dundee, June 28--July 1, 1977 , organization Springer . pp. pages 105--116

  36. [44]

    , author El-Sheimy, N

    author Nassar, S. , author El-Sheimy, N. , year 2006 . title A combined algorithm of improving ins error modeling and sensor measurements for accurate ins/gps navigation . journal GPS solutions volume 10 , pages 29--39

  37. [45]

    , author Yuan, S

    author Nguyen, T.M. , author Yuan, S. , author Cao, M. , author Lyu, Y. , author Nguyen, T.H. , author Xie, L. , year 2022 . title Ntu viral: A visual-inertial-ranging-lidar dataset, from an aerial vehicle viewpoint . journal The International Journal of Robotics Research volu...

  38. [46]

    , author Moghadam, P

    author Park, C. , author Moghadam, P. , author Kim, S. , author Sridharan, S. , author Fookes, C. , year 2020 . title Spatiotemporal camera-lidar calibration: A targetless and structureless approach . journal IEEE Robotics and Automation Letters volume 5 , pages 1556--1563

  39. [47]

    , author Zheng, Y

    author Qin, T. , author Zheng, Y. , author Chen, T. , author Chen, Y. , author Su, Q. , year 2021 . title A light-weight semantic map for visual localization towards autonomous driving , in: booktitle 2021 IEEE international conference on robotics and automation (ICRA) , organ...

  40. [48]

    , author Wang, Y

    author Ramezani, M. , author Wang, Y. , author Camurri, M. , author Wisth, D. , author Mattamala, M. , author Fallon, M. , year 2020 . title The newer college dataset: Handheld lidar, inertial and vision with ground truth , in: booktitle 2020 IEEE/RSJ International Conference ...

  41. [49]

    , author Zeng, Y

    author Ren, S. , author Zeng, Y. , author Hou, J. , author Chen, X. , year 2022 . title Corri2p: Deep image-to-point cloud registration via dense correspondence . journal IEEE Transactions on Circuits and Systems for Video Technology volume 33 , pages 1198--1208

  42. [50]

    , author DeTone, D

    author Sarlin, P.E. , author DeTone, D. , author Yang, T.Y. , author Avetisyan, A. , author Straub, J. , author Malisiewicz, T. , author Bulo, S.R. , author Newcombe, R. , author Kontschieder, P. , author Balntas, V. , year 2023 a. title Orienternet: Visual localization in 2d ...

  43. [51]

    , author Trulls, E

    author Sarlin, P.E. , author Trulls, E. , author Pollefeys, M. , author Hosang, J. , author Lynen, S. , year 2023 b. title Snap: Self-supervised neural maps for visual positioning and semantic understanding . journal Advances in Neural Information Processing Systems volume 36 ...

  44. [52]

    , author Vallet, J

    author Schaer, P. , author Vallet, J. , year 2016 . title Trajectory adjustment of mobile laser scan data in gps denied environments . journal The International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences volume 40 , pages 61--64

  45. [53]

    , author Goll, T

    author Schubert, D. , author Goll, T. , author Demmel, N. , author Usenko, V. , author St \"u ckler, J. , author Cremers, D. , year 2018 . title The tum vi benchmark for evaluating visual-inertial odometry , in: booktitle 2018 IEEE/RSJ International Conference on Intelligent R...

  46. [54]

    , author El-Sheimy, N

    author Schwarz, K.P. , author El-Sheimy, N. , year 2004 . title Mobile mapping systems--state of the art and future trends . journal International Archives of Photogrammetry, Remote Sensing and Spatial Information Sciences volume 35 , pages 10

  47. [55]

    , author Omama, M

    author Shubodh, S. , author Omama, M. , author Zaidi, H. , author Parihar, U.S. , author Krishna, M. , year 2024 . title Lip-loc: Lidar image pretraining for cross-modal localization , in: booktitle Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vi...

  48. [56]

    , author De Silva, O

    author Thalagala, R.G. , author De Silva, O. , author Jayasiri, A. , author Gubbels, A. , author Mann, G.K. , author Gosine, R.G. , year 2024 . title Mun-frl: A visual-inertial-lidar dataset for aerial autonomous navigation and mapping . journal The International Journal of Ro...

  49. [57]

    , author Chang, Y

    author Tian, Y. , author Chang, Y. , author Quang, L. , author Schang, A. , author Nieto-Granda, C. , author How, J.P. , author Carlone, L. , year 2023 . title Resilient and distributed multi-robot visual slam: Datasets, experiments, and lessons learned , in: booktitle 2023 IE...

  50. [58]

    , author C ad \' k, M

    author Tome s ek, J. , author C ad \' k, M. , author Brejcha, J. , year 2022 . title Crosslocate: cross-modal large-scale visual geo-localization in natural environments using rendered modalities , in: booktitle Proceedings of the IEEE/CVF Winter Conference on Applications of ...

  51. [59]

    , year 2023

    author USGS, U.S.A. , year 2023 . title San Francisco B23 LiDAR Data . howpublished https://apps.nationalmap.gov/lidar-explorer

  52. [60]

    , author Lee, G.H

    author Uy, M.A. , author Lee, G.H. , year 2018 . title Pointnetvlad: Deep point cloud based retrieval for large-scale place recognition , in: booktitle 2018 IEEE / CVF Conference on Computer Vision and Pattern Recognition , pp. pages 4470--4479 . :10.1109/CVPR.2018.00470

  53. [61]

    , author Chen, C

    author Wang, B. , author Chen, C. , author Cui, Z. , author Qin, J. , author Lu, C.X. , author Yu, Z. , author Zhao, P. , author Dong, Z. , author Zhu, F. , author Trigoni, N. , et al., year 2021 . title P2-net: Joint description and detection of local features for pixel and p...

  54. [62]

    , author Dai, Y

    author Wang, C. , author Dai, Y. , author Elsheimy, N. , author Wen, C. , author Retscher, G. , author Kang, Z. , author Lingua, A. , year 2020 . title Isprs benchmark on multisensory indoor mapping and positioning . journal ISPRS Annals of the Photogrammetry, Remote Sensing a...

  55. [63]

    , author Hauberg, S

    author Warburg, F. , author Hauberg, S. , author Lopez-Antequera, M. , author Gargallo, P. , author Kuang, Y. , author Civera, J. , year 2020 . title Mapillary street-level sequences: A dataset for lifelong place recognition , in: booktitle Proceedings of the IEEE/CVF conferen...

  56. [64]

    , author Jutzi, B

    author Weinmann, M. , author Jutzi, B. , author Hinz, S. , author Mallet, C. , year 2015 . title Semantic point cloud interpretation based on optimal neighborhoods, relevant features and efficient classifiers . journal ISPRS Journal of Photogrammetry and Remote Sensing volume ...

  57. [65]

    , author Zhou, Y

    author Wen, W. , author Zhou, Y. , author Zhang, G. , author Fahandezh-Saadi, S. , author Bai, X. , author Zhan, W. , author Tomizuka, M. , author Hsu, L.T. , year 2020 . title Urbanloco: A full sensor suite dataset for mapping and localization in urban scenes , in: booktitle ...

  58. [66]

    , author Zhang, Z

    author Wu, H. , author Zhang, Z. , author Lin, S. , author Mu, X. , author Zhao, Q. , author Yang, M. , author Qin, T. , year 2024 . title Maplocnet: Coarse-to-fine feature registration for visual re-localization in navigation maps , in: booktitle 2024 IEEE/RSJ International C...

  59. [67]

    , author Shao, R

    author Xie, Y. , author Shao, R. , author Guli, P. , author Li, B. , author Wang, L. , year 2018 . title Infrastructure based calibration of a multi-camera and multi-lidar system using apriltags , in: booktitle 2018 IEEE Intelligent Vehicles Symposium (IV) , organization IEEE ...

  60. [68]

    , author Lv, Z

    author Ye, J. , author Lv, Z. , author Li, W. , author Yu, J. , author Yang, H. , author Zhong, H. , author He, C. , year 2024 . title Cross-view image geo-localization with panorama-bev co-retrieval network , in: booktitle European Conference on Computer Vision , organization...

  61. [69]

    , author Li, A

    author Yin, J. , author Li, A. , author Li, T. , author Yu, W. , author Zou, D. , year 2021 . title M2dgr: A multi-sensor and multi-scenario slam dataset for ground robots . journal IEEE Robotics and Automation Letters volume 7 , pages 2266--2273

  62. [70]

    , author Zhao, H

    author Zhang, C. , author Zhao, H. , author Wang, C. , author Tang, X. , author Yang, M. , year 2023 . title Cross-modal monocular localization in prior lidar maps utilizing semantic consistency , in: booktitle 2023 IEEE International Conference on Robotics and Automation (ICR...

  63. [71]

    , author Helmberger, M

    author Zhang, L. , author Helmberger, M. , author Fu, L.F.T. , author Wisth, D. , author Camurri, M. , author Scaramuzza, D. , author Fallon, M. , year 2022 . title Hilti-oxford dataset: A millimeter-accurate benchmark for simultaneous localization and mapping . journal IEEE R...

  64. [72]

    , author Zhu, S

    author Zhang, X. , author Zhu, S. , author Guo, S. , author Li, J. , author Liu, H. , year 2021 . title Line-based automatic extrinsic calibration of lidar and camera , in: booktitle 2021 IEEE International Conference on Robotics and Automation (ICRA) , pp. pages 9347--9353 . ...

  65. [73]

    , author Gao, Y

    author Zhao, S. , author Gao, Y. , author Wu, T. , author Singh, D. , author Jiang, R. , author Sun, H. , author Sarawata, M. , author Qiu, Y. , author Whittaker, W. , author Higgins, I. , et al., year 2024 . title Subt-mrs dataset: Pushing slam towards all-weather environment...

  66. [74]

    , author Yu, H

    author Zhao, Z. , author Yu, H. , author Lyu, C. , author Yang, W. , author Scherer, S. , year 2023 . title Attention-enhanced cross-modal localization between spherical images and point clouds . journal IEEE Sensors Journal , pages 1--1 :10.1109/JSEN.2023.3306377

  67. [75]

    , author Wen, W

    author Zheng, X. , author Wen, W. , author Hsu, L.T. , year 2024 . title Tightly-coupled visual/inertial/map integration with observability analysis for reliable localization of intelligent vehicles . journal IEEE Transactions on Intelligent Vehicles

  68. [76]

    , author Quang, L

    author Zhou, Y. , author Quang, L. , author Nieto-Granda, C. , author Loianno, G. , year 2024 . title Coped-advancing multi-robot collaborative perception: A comprehensive dataset in real-world environments . journal IEEE Robotics and Automation Letters

  69. [77]

    , author Kong, Y

    author Zhu, Y. , author Kong, Y. , author Jie, Y. , author Xu, S. , author Cheng, H. , year 2023 . title Graco: A multimodal dataset for ground and aerial cooperative localization and mapping . journal IEEE Robotics and Automation Letters volume 8 , pages 966--973 . :10.1109/L...

  70. [78]

    , author Li, J

    author Zou, X. , author Li, J. , author Wu, W. , author Liang, F. , author Yang, B. , author Dong, Z. , year 2025 . title Reliable-loc: Robust sequential lidar global localization in large-scale street scenes based on verifiable cues . journal ISPRS Journal of Photogrammetry a...

  71. [79]

    , author Geneva, P

    author Zuo, X. , author Geneva, P. , author Yang, Y. , author Ye, W. , author Liu, Y. , author Huang, G. , year 2019 . title Visual-inertial localization with prior lidar map constraints . journal IEEE Robotics and Automation Letters volume 4 , pages 3394--3401

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.