REVIEW 4 major objections 6 minor 51 references
Learning Local Feature Descriptor with Motion Attribute for Vision-based Localization
T0 review · 4 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read The paper claims a single fully convolutional network can label each local feature as static, moving, or unstable while computing its descriptor, and that filtering out non-static points before matching sharply improves vision-based…
desk verdict A sensible joint FCN for motion attributes and descriptors whose localization gains rest on an unvalidated semantic-to-motion mapping and a missing comparison to [26]. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is MD-Net, a fully convolutional network with one shared backbone (eight convolutional layers, three max-pooling layers, batch normalization and ReLU) and two light branches sharing that backbone. The motion branch NM outputs per-point probabilities over three classes using a reweighted cross-entropy loss, with class weights inversely proportional to class frequency; the descriptor branch ND regresses 128-dimensional descriptors by matching the output of a HardNet-based teacher under an MSE loss. The combined multi-task loss is $L = \lambda_M L_M + \lambda_D L_D$ with both weights set to 1, and training is done with Adam on Cityscapes images. In the localization pipeline, FAST points whose predicted attribute is not static are discarded, and the remaining points carry MD-Net descriptors into a sliding-window visual-inertial SLAM system.
What would settle it
Run the same visual-inertial localization on a sequence whose only non-static objects are parked cars and traffic lights that remain motionless for the entire traversal; if keeping those points (the 'moving' class) does not increase or even reduces RMSE compared with filtering them out, the semantic-to-motion mapping is not the right model of long-term staticness.
Extended reading notes
Core claim
MD-Net claims that motion attributes and local descriptors can be computed in a single forward pass by a fully convolutional network. The backbone extracts shared features; a light motion branch classifies each point as unstable, moving, or static under a reweighted cross-entropy loss; a light descriptor branch produces 128-dimensional descriptors by imitating a pre-trained HardNet teacher through a mean-square-error loss. Ground truth for motion comes from a hand-specified mapping of Cityscapes semantic classes, with sky, vegetation, and terrain marked unstable, humans and vehicles (including parked cars and traffic lights) marked moving, and buildings, walls, poles, and similar structures marked static. The paper reports that the resulting descriptors roughly match HardNet on patch tasks, reach a mean motion-attribute IoU of 76.2 on Cityscapes validation, raise two-view RANSAC inlier ratios to about 92 percent on average, and, when integrated with FAST detection and a visual-inertial SLAM system, reduce localization RMSE from 2.87 m to 1.10 m on one urban sequence and from 27.06 m to 7.25 m on another.
Load-bearing premise
The load-bearing premise is that the hand-chosen mapping from Cityscapes semantic classes to motion attributes, such as labeling parked cars and traffic lights as moving, matches the true long-term stability of those points in the operating environment; if a parked car stays put for the whole session, filtering it out removes a useful feature.
Editorial extensions
If this is right
- If MD-Net's motion filtering is correct, visual-inertial localization error drops sharply in dynamic street scenes: 1.10 m versus 2.87 m RMSE against FAST+FREAK on series 1, and 7.25 m versus 27.06 m on series 2.
- Because the heavy backbone is shared, the extra cost of labeling motion is small relative to running descriptor extraction alone, which matters for real-time use on robots and cars.
- Filtering before matching raises two-view RANSAC inlier ratios from a mean of 86.2 percent for FAST+FREAK to 92.2 percent, so fewer matches are wasted on moving or unstable points.
- The learned descriptors are distinct enough to beat SuperPoint and hand-crafted features in the tested localization runs, even though they are trained by imitating HardNet rather than with a dedicated descriptor loss.
Reading between the lines
- The semantic-to-motion mapping is the fragile part: one could replace it with a learned classifier trained on multi-session revisits, where 'long-term static' is measured by whether a point is observed at the same 3D location across sessions.
- The same filtering idea could be applied to loop-closure databases, retaining only long-term static points to keep maps from being polluted by objects that moved since the last visit.
- Because the descriptor branch is trained to imitate HardNet, the motion branch could in principle be attached to any FCN descriptor, making motion filtering a drop-in module rather than a bespoke architecture.
- A direct extension would be to predict motion attributes at multiple time horizons, such as minutes versus days, instead of a single three-class label, since short-term and long-term localization need different staticness criteria.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes MD-Net, a fully convolutional network that simultaneously estimates per-point motion attributes (unstable, moving, static) and computes local feature descriptors. The motion attribute branch is trained with a re-weighted cross-entropy loss on labels obtained from a hand-crafted mapping of Cityscapes semantic classes (Table I). The descriptor branch is trained by regression to HardNet feature outputs (Eq. 4), acting as a student to a teacher model. In deployment, features detected by FAST are filtered to keep only static points, and descriptors are extracted from MD-Net; this pipeline is plugged into a visual-inertial SLAM system. Experiments are reported on Cityscapes motion-attribute accuracy (Table II), HPatches descriptor benchmarks (Fig. 5), and three self-collected localization datasets (Table III-V), with claims of improved localization accuracy relative to FAST+FREAK and FAST+SuperPoint baselines.
Significance. If the motion-attribute prediction is reliable, the proposed architecture provides a computationally efficient, single-forward-pass way to discard dynamic and unstable features, which is directly relevant to long-term visual localization. The descriptor branch transfers HardNet's discriminative power to a fully convolutional student, a useful engineering result. However, the central localization claim rests on an unvalidated semantic-to-motion mapping, and the experiments lack statistical support and a comparison to the closest prior method [26]. With those issues addressed, the method could be a practical component for robust visual-inertial localization in dynamic environments.
major comments (4)
- [Section III-A, Table I] The motion-attribute ground-truth labels are derived entirely from the semantic-to-motion mapping in Table I. This mapping assigns the Cityscapes classes 'static' (small static objects such as barriers and trash cans) and 'traffic light' to 'moving', while assigning 'vegetation' and 'terrain' to 'unstable'. These assignments are physically questionable: traffic lights are fixed infrastructure and small static objects are usually stationary over long periods, while tree trunks and terrain can be stable landmarks. Because the IoU scores in Table II are computed against labels generated by the same mapping, they only measure consistency with the mapping, not agreement with true long-term motion. The localization gain in Tables IV-V depends precisely on this filter, so the central claim is only as strong as the mapping. I request an external validation of the mapping (e.g., long-term feature tracking or temporal analysis on data with known dynamics) and an analysis of how mapping errors would affect localization accuracy.
- [Introduction and Section V-C] The abstract and introduction state that the proposed method achieves 'significantly better accuracy, compared to [26]' (the Self-Improving Visual Odometry method). However, no experiment involving [26] is reported in Section V. The comparisons in Tables III-V include FREAK, SuperPoint, and ablations, but not [26]. Either add the [26] baseline to the localization experiments (and descriptor comparisons, if applicable), or remove the claim of superiority over [26].
- [Section V-C, Tables IV-V] The localization experiments are single-run evaluations: one restaurant loop and two street sequences, with no repeated trials, no error bars, and no statistical significance testing. The reported RMSE values (e.g., 1.10 vs 2.87 m on series1) could be influenced by RANSAC randomness, IMU initialization, or dataset-specific conditions. To support the characterization of the improvement as 'significant', the authors should provide repeated runs or otherwise quantify uncertainty, and should clarify what the 'Failed' entries in Table V mean (e.g., tracker divergence, insufficient inliers, or numerical failure).
- [Section V-C, Table III] The inlier-ratio results in Table III are presented as scalar means without the number of image pairs, the variance, or the statistical significance of the differences. Given that the differences between methods (e.g., 92.2 vs 89.1 for SuperPoint on Scene3) are not large, reporting only the mean is insufficient to establish that the proposed filtering consistently improves matching. Please provide per-pair statistics or error bars.
minor comments (6)
- [Section IV] The filtering rule 'only static points are reserved' is not fully specified: is a point kept only when the 'static' class has the highest probability, or is a probability threshold applied? Please state the exact criterion used in the localization experiments.
- [Section V-A] The phrase 'our work park' appears to be a typo; it should likely be 'our workplace' or 'our office park'.
- [Author affiliation] In the author affiliation line, 'Jia Li is with the the National Key Laboratory' contains a duplicated 'the'.
- [Section III-C, Eq. (6)] The learning-rate schedule $l_e = l_0 b^{e/E}$ uses the symbol $e$ both as an epoch index and (if read mathematically) as Euler's number. Consider using a different index, e.g., $t$, for clarity.
- [Fig. 5] The HPatches results are shown as curves without error bars or numeric values. Reporting the exact numbers or adding error bars would help readers assess the claimed 'comparable' performance with HardNet and the margin over SuperPoint.
- [Section V-A, Fig. 4] The generalization evaluation on the Alibaba campus data is qualitative only. If the authors wish to claim generalization beyond Cityscapes, a quantitative evaluation on this data (e.g., motion-attribute IoU or localization results) would be needed.
Circularity Check
No significant circularity: motion labels come from an external semantic dataset, the descriptor is an openly acknowledged HardNet distillation, and localization accuracy is tested on fresh data with external ground truth.
full rationale
The paper's two learned components are (i) motion attribute classification trained on labels derived from Cityscapes semantic annotations via a stated, hand-authored mapping (Table I) and (ii) a descriptor branch trained by regression to HardNet outputs (Eq. 4), disclosed as a teacher-student setup. Neither is a self-derivation: Cityscapes and HardNet are external, and the HPatches and RTK-GPS localization tests are held-out. The localization comparison 'FAST Our' vs 'FAST Our w/o mot. att.' isolates the effect of the motion filter on new street sequences, so the central accuracy claim does not reduce to the training labels. The semantic-to-motion mapping (e.g., traffic lights labeled 'moving') is a correctness and generalization risk, not circularity, because the mapping is fixed before training and the final evaluation is against external trajectories. No load-bearing self-citation chain exists: [22] supplies the baseline visual-inertial implementation, but all feature variants are tested under the same implementation.
Assumptions & free parameters
free parameters (3)
- Multi-task loss weights lambda_M and lambda_D =
1 and 1
- Initial learning rate l0 =
0.01
- Motion attribute mapping table =
Table I
assumptions (3)
- domain assumption Cityscapes semantic labels are a reliable proxy for long-term motion attributes
- domain assumption HardNet descriptor outputs are suitable targets for distillation
- domain assumption Backbone structure of SuperPoint [19] is a good shared feature extractor
Cite this review
Pith. "Pith review of Learning Local Feature Descriptor with Motion Attribute for Vision-based Localization." pith.science (2026). https://pith.science/paper/N6CEPCA2
@misc{pith2026190801180,
author = {Pith},
title = {Pith review of: Learning Local Feature Descriptor with Motion Attribute for Vision-based Localization},
year = {2026},
howpublished = {\url{https://pith.science/paper/N6CEPCA2}},
note = {Machine review of arXiv:1908.01180}
}
read the original abstract
In recent years, camera-based localization has been widely used for robotic applications, and most proposed algorithms rely on local features extracted from recorded images. For better performance, the features used for open-loop localization are required to be short-term globally static, and the ones used for re-localization or loop closure detection need to be long-term static. Therefore, the motion attribute of a local feature point could be exploited to improve localization performance, e.g., the feature points extracted from moving persons or vehicles can be excluded from these systems due to their unsteadiness. In this paper, we design a fully convolutional network (FCN), named MD-Net, to perform motion attribute estimation and feature description simultaneously. MD-Net has a shared backbone network to extract features from the input image and two network branches to complete each sub-task. With MD-Net, we can obtain the motion attribute while avoiding increasing much more computation. Experimental results demonstrate that the proposed method can learn distinct local feature descriptor along with motion attribute only using an FCN, by outperforming competing methods by a wide margin. We also show that the proposed algorithm can be integrated into a vision-based localization algorithm to improve estimation accuracy significantly.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[26]
Self-Improving Visual Odometry,
D. DeTone, T. Malisiewicz, and A. Rabinovich, “Self-Improving Visual Odometry,” CoRR, vol. abs/1812.03245, 2018
arXiv 2018
-
[1]
6-DOF Image Localization from Massive Geo-tagged Reference Images,
Y . Song, X. Chen, X. Wang, Y . Zhang, and J. Li, “6-DOF Image Localization from Massive Geo-tagged Reference Images,” IEEE Transactions on Multimedia , vol. 18, no. 8, pp. 1542–1554, 2016
work page 2016
-
[2]
High-fidelity sensor modeling and self-calibration in vision-aided inertial navigation,
M. Li, H. Yu, X. Zheng, and A. I. Mourikis, “High-fidelity sensor modeling and self-calibration in vision-aided inertial navigation,” in ICRA, May 2014, pp. 409–416
work page 2014
-
[3]
High-precision, consistent ekf-based visual- inertial odometry,
M. Li and A. I. Mourikis, “High-precision, consistent ekf-based visual- inertial odometry,” IJRR, vol. 32, no. 6, pp. 690–711, 2013
work page 2013
-
[4]
Dpptam: Dense piecewise planar tracking and mapping from a monocular sequence,
A. Concha and J. Civera, “Dpptam: Dense piecewise planar tracking and mapping from a monocular sequence,” in IEEE/RSJ International Conference on Intelligent Robots and Systems , 2015, pp. 5686–5693
work page 2015
-
[5]
Active Image-Based Modeling with a Toy Drone,
R. Huang, D. Zou, R. Vaughan, and P. Tan, “Active Image-Based Modeling with a Toy Drone,” in ICRA, 2018
work page 2018
-
[6]
Very Large-Scale Global SfM by Distributed Motion Averaging,
S. Zhu, R. Zhang, L. Zhou, T. Shen, T. Fang, P. Tan, and L. Quan, “Very Large-Scale Global SfM by Distributed Motion Averaging,” in CVPR, 2018
work page 2018
-
[7]
ORB-SLAM2: An Open-Source SLAM System for Monocular, Stereo, and RGB-D Cameras,
R. Mur-Artal and J. D. Tard ´os, “ORB-SLAM2: An Open-Source SLAM System for Monocular, Stereo, and RGB-D Cameras,” IEEE Transactions on Robotics , vol. 33, no. 5, pp. 1255–1262, 2017
work page 2017
Show all 51 references
-
[8]
maplab: An open framework for research in visual-inertial mapping and localization,
T. Schneider, M. Dymczyk, M. Fehr, K. Egger, S. Lynen, I. Gilitschen- ski, and R. Siegwart, “maplab: An open framework for research in visual-inertial mapping and localization,” IEEE Robotics and Automa- tion Letters , 2018
2018
-
[9]
Optimization-based estimator design for vision-aided inertial navigation,
M. Li and A. I. Mourikis, “Optimization-based estimator design for vision-aided inertial navigation,” in Robotics: Science and Systems , 2013, pp. 241–248
2013
-
[10]
Distinctive Image Features from Scale-Invariant Key- points,
D. G. Lowe, “Distinctive Image Features from Scale-Invariant Key- points,” IJCV, vol. 60, no. 2, pp. 91–110, 2004
2004
-
[11]
ORB: An efficient alternative to SIFT or SURF,
E. Rublee, V . Rabaud, K. Konolige, and G. R. Bradski, “ORB: An efficient alternative to SIFT or SURF,” in ICCV, 2011
2011
-
[12]
FREAK: Fast Retina Keypoint,
A. Alahi, R. Ortiz, and P. Vandergheynst, “FREAK: Fast Retina Keypoint,” in CVPR, 2012
2012
-
[13]
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,
S. Ren, K. He, R. B. Girshick, and J. Sun, “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 39, no. 6, pp. 1137–1149, 2017
2017
-
[14]
Fully Convolutional Networks for Semantic Segmentation,
E. Shelhamer, J. Long, and T. Darrell, “Fully Convolutional Networks for Semantic Segmentation,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 39, no. 4, pp. 640–651, 2017
2017
-
[15]
TILDE: A Temporally Invariant Learned DEtector,
Y . Verdie, K. M. Yi, P. Fua, and V . Lepetit, “TILDE: A Temporally Invariant Learned DEtector,” in CVPR, 2015
2015
-
[16]
LIFT: Learned Invariant Feature Transform,
K. M. Yi, E. Trulls, V . Lepetit, and P. Fua, “LIFT: Learned Invariant Feature Transform,” in ECCV, 2016
2016
-
[17]
L2-Net: Deep Learning of Discriminative Patch Descriptor in Euclidean Space,
Y . Tian, B. Fan, and F. Wu, “L2-Net: Deep Learning of Discriminative Patch Descriptor in Euclidean Space,” in CVPR, 2017
2017
-
[18]
Working Hard to Know Your Neighbor’s Margins: Local Descriptor Learning Loss,
A. Mishchuk, D. Mishkin, F. Radenovic, and J. Matas, “Working Hard to Know Your Neighbor’s Margins: Local Descriptor Learning Loss,” in Advances in Neural Information Processing Systems (NeurIPS) , 2017
2017
-
[19]
SuperPoint: Self- Supervised Interest Point Detection and Description,
D. DeTone, T. Malisiewicz, and A. Rabinovich, “SuperPoint: Self- Supervised Interest Point Detection and Description,” in CVPR - Workshops, 2018
2018
-
[20]
R. I. Hartley and A. Zisserman, Multiple View Geometry in Computer Vision, 2nd ed. Cambridge University Press, ISBN: 0521540518, 2004
2004
-
[21]
Get out of my lab: Large-scale, real-time visual-inertial localization
S. Lynen, T. Sattler, M. Bosse, J. A. Hesch, M. Pollefeys, and R. Siegwart, “Get out of my lab: Large-scale, real-time visual-inertial localization.” in Robotics: Science and Systems , 2015
2015
-
[22]
Vision-aided localization for ground robots,
M. Zhang, Y . Chen, and M. Li, “Vision-aided localization for ground robots,” in IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , 2019
2019
-
[23]
Mask- SLAM: Robust Feature-Based Monocular SLAM by Masking Using Semantic Segmentation,
M. Kaneko, K. Iwami, T. Ogawa, T. Yamasaki, and K. Aizawa, “Mask- SLAM: Robust Feature-Based Monocular SLAM by Masking Using Semantic Segmentation,” in CVPR - Workshops , 2018
2018
-
[24]
Semantics- aware Visual Localization Under Challenging Perceptual Conditions,
T. Naseer, G. L. Oliveira, T. Brox, and W. Burgard, “Semantics- aware Visual Localization Under Challenging Perceptual Conditions,” in ICRA, 2017
2017
-
[25]
Rich Feature Hi- erarchies for Accurate Object Detection and Semantic Segmentation,
R. B. Girshick, J. Donahue, T. Darrell, and J. Malik, “Rich Feature Hi- erarchies for Accurate Object Detection and Semantic Segmentation,” in CVPR, 2014
2014
-
[27]
Deep Neural Networks for Object Detection,
C. Szegedy, A. Toshev, and D. Erhan, “Deep Neural Networks for Object Detection,” in Advances in Neural Information Processing Systems (NeurIPS) , 2013
2013
-
[28]
A Combined Corner and Edge Detector,
C. G. Harris and M. Stephens, “A Combined Corner and Edge Detector,” in Alvey Vision Conference (A VC), 1988
1988
-
[29]
Machine Learning for High-Speed Corner Detection,
E. Rosten and T. Drummond, “Machine Learning for High-Speed Corner Detection,” in ECCV, 2006
2006
-
[30]
Edge Detection and Ridge Detection with Automatic Scale Selection,
T. Lindeberg, “Edge Detection and Ridge Detection with Automatic Scale Selection,” IJCV, vol. 30, no. 2, pp. 117–156, 1998
1998
-
[31]
Robust Wide Baseline Stereo from Maximally Stable Extremal Regions,
J. Matas, O. Chum, M. Urban, and T. Pajdla, “Robust Wide Baseline Stereo from Maximally Stable Extremal Regions,” in BMVC, 2002
2002
-
[32]
SURF: Speeded Up Robust Features,
H. Bay, T. Tuytelaars, and L. J. V . Gool, “SURF: Speeded Up Robust Features,” in ECCV, 2006
2006
-
[33]
Fast Human Detection Using a Cascade of Histograms of Oriented Gradients,
Q. Zhu, M. Yeh, K. Cheng, and S. Avidan, “Fast Human Detection Using a Cascade of Histograms of Oriented Gradients,” in CVPR, 2006
2006
-
[34]
BRIEF: Binary Robust Independent Elementary Features,
M. Calonder, V . Lepetit, C. Strecha, and P. Fua, “BRIEF: Binary Robust Independent Elementary Features,” in ECCV, 2010
2010
-
[35]
Quad- networks: unsupervised learning to rank for interest point detection,
N. Savinov, A. Seki, L. Ladicky, T. Sattler, and M. Pollefeys, “Quad- networks: unsupervised learning to rank for interest point detection,” in CVPR, 2017
2017
-
[36]
Learning Dis- criminative and Transformation Covariant Local Feature Detectors,
X. Zhang, F. X. Yu, S. Karaman, and S.-F. Chang, “Learning Dis- criminative and Transformation Covariant Local Feature Detectors,” in CVPR, 2017
2017
-
[37]
Automatic Panoramic Image Stitching using Invariant Features,
M. Brown and D. G. Lowe, “Automatic Panoramic Image Stitching using Invariant Features,” IJCV, vol. 74, no. 1, pp. 59–73, 2007
2007
-
[38]
HPatches: A Benchmark and Evaluation of Handcrafted and Learned Local Descriptors,
V . Balntas, K. Lenc, A. Vedaldi, and K. Mikolajczyk, “HPatches: A Benchmark and Evaluation of Handcrafted and Learned Local Descriptors,” in CVPR, 2017
2017
-
[39]
A Large Dataset for Improving Patch Matching,
R. Mitra, N. Doiphode, U. Gautam, S. Narayan, S. Ahmed, S. Chan- dran, and A. Jain, “A Large Dataset for Improving Patch Matching,” CoRR, vol. abs/1801.01466, 2018
2018 arXiv
-
[40]
GeoDesc: Learning Local Descriptors by Integrating Geometry Constraints,
Z. Luo, T. Shen, L. Zhou, S. Zhu, R. Zhang, Y . Yao, T. Fang, and L. Quan, “GeoDesc: Learning Local Descriptors by Integrating Geometry Constraints,” in ECCV, 2018
2018
-
[41]
LF-Net: Learning Local Features from Images,
Y . Ono, E. Trulls, P. Fua, and K. M. Yi, “LF-Net: Learning Local Features from Images,” in Advances in Neural Information Processing Systems (NeurIPS) , 2018
2018
-
[42]
Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,
S. Ioffe and C. Szegedy, “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” in ICML, 2015
2015
-
[43]
Rectified Linear Units Improve Restricted Boltzmann Machines,
V . Nair and G. E. Hinton, “Rectified Linear Units Improve Restricted Boltzmann Machines,” in ICML, 2010
2010
-
[44]
The Cityscapes Dataset for Semantic Urban Scene Understanding,
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Be- nenson, U. Franke, S. Roth, and B. Schiele, “The Cityscapes Dataset for Semantic Urban Scene Understanding,” in CVPR, 2016
2016
-
[45]
Learning from Large-scale Noisy Web Data with Ubiquitous Reweighting for Image Classification,
J. Li, Y . Song, J. Zhu, L. Cheng, Y . Su, L. Ye, P. Yuan, and S. Han, “Learning from Large-scale Noisy Web Data with Ubiquitous Reweighting for Image Classification,” CoRR, vol. abs/1811.00700, 2018
2018 arXiv
-
[46]
Embedding 3D Geometric Features for Rigid Object Part Segmentation,
Y . Song, X. Chen, J. Li, and Q. Zhao, “Embedding 3D Geometric Features for Rigid Object Part Segmentation,” in ICCV, 2017
2017
-
[47]
Vins-mono: A robust and versatile monoc- ular visual-inertial state estimator,
T. Qin, P. Li, and S. Shen, “Vins-mono: A robust and versatile monoc- ular visual-inertial state estimator,” IEEE Transactions on Robotics , vol. 34, no. 4, pp. 1004–1020, 2018
2018
-
[48]
Keyframe-based visual–inertial odometry using nonlinear optimiza- tion,
S. Leutenegger, S. Lynen, M. Bosse, R. Siegwart, and P. Furgale, “Keyframe-based visual–inertial odometry using nonlinear optimiza- tion,” IJRR, vol. 34, no. 3, pp. 314–334, 2015
2015
-
[49]
Automatic differ- entiation in pytorch,
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer, “Automatic differ- entiation in pytorch,” in Advances in Neural Information Processing Systems (NeurIPS) - Workshop , 2017
2017
-
[50]
Discriminative Learning of Deep Convolutional Feature Point Descriptors,
E. Simo-Serra, E. Trulls, L. Ferraz, I. Kokkinos, P. Fua, and F. Moreno- Noguer, “Discriminative Learning of Deep Convolutional Feature Point Descriptors,” in ICCV, 2015
2015
-
[51]
Least-squares estimation of transformation parameters between two point patterns,
S. Umeyama, “Least-squares estimation of transformation parameters between two point patterns,” IEEE Trans. Pattern Anal. Mach. Intell. , no. 4, pp. 376–380, 1991
1991
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.