REVIEW 3 major objections 4 minor 36 references
Track-leakage-free hold-out self-validation is a fragmentation warning, not an accuracy certificate.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A track-leakage-free hold-out self-consistency score saturates at 1.00 while true reconstruction error swings up to 106 m, proving internal consistency is not absolute accuracy.
T0 review reviewed 2026-08-04 challenge →
load-bearing objection Honest, well-executed negative result on self-validation; the core claim is solid, the 'whole family' generalization is overreach. the 3 major comments →
Track-Leakage-Free Hold-Out Self-Validation for Photogrammetric Reconstruction: Protocol, Sensitivity, and Limits
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that a track-leakage-free hold-out self-validation score—despite being computationally well-posed and able to detect model fragmentation—is not an accuracy measure, and cannot be turned into one by tuning thresholds or restoring the scale axis. The score saturates: on good reconstructions it reports confidence 1.00 with sub-hundredth-degree agreement, but it also reports 1.00 on RTK-referenced models that are 3.4–4.3 m wrong and on injected single-component distortions that are 55–106 m wrong. The blindness is argued to be structural rather than statistical: bundle adjustment has gauge/datum freedom, so any score computed purely from the reconstruction's own geom
What carries the argument
The load-bearing mechanism is the track-level leakage barrier: a deterministic subset of images is withheld, and each withheld view is re-localised by PnP resection against only 3D points observed by at least two retained images, so no view is tested against structure it helped triangulate. Disagreement is measured as a max of geodesic rotation error and scale-free translation-direction error, aggregated as mean average accuracy over 1°, 3°, and 5° thresholds. The paper then invokes bundle adjustment's gauge/datum freedom to argue that this score—and any internal consistency score—is blind to coherent locally near-rigid distortion, with the hold-out fraction and barrier strength shown to be
Load-bearing premise
The load-bearing premise is that no internally computed signal—the hold-out score, bundle-adjustment covariance, reprojection statistics, or track counts—can detect a coherent locally near-rigid distortion, an extrapolation from exact gauge invariance tested on only a few injected warps and a handful of proxy statistics on small samples.
What would settle it
Run the hold-out protocol on a single connected reconstruction that is globally wrong by tens of metres after similarity alignment, with no fragmentation: if the confidence drops, the blindness claim fails. If a purely internal signal—such as disagreement between two independently built sub-models of the same scene—flags that model while the hold-out score stays at 1.00, then the claim that the blind spot covers the whole locally-near-rigid family is over-broad. A practical test: a benchmark of roughly 50 operational captures with survey truth and no shared degradation knob, checking whether a
If this is right
- A high self-consistency score cannot replace a control-point accuracy check; models wrong by tens of metres pass with confidence 1.00.
- A confidence drop is usable only as a qualitative fragmentation warning, not a calibrated detector: benign undersampling can lower confidence as much as fragmentation.
- The blind spot is shared by the wider class of internal consistency signals: bundle-adjustment covariance, reprojection RMSE (which can even invert), and track-length statistics when not riding a degradation-harness confound.
- Escaping the blind spot requires information from outside the reconstruction's own geometry: an external datum, world priors, or cross-consistency between independently built sub-models.
- No validated threshold on the confidence signal currently separates failure from benign undersampling; the negative result rests on the saturation and existence proofs, not on the underpowered correlation analysis.
Where Pith is reading between the lines
- The paper's structural argument implies that any future purely internal 'accuracy' signal must either leave the locally-near-rigid family or add information from outside the single reconstruction; a concrete untested next experiment is cross-submodel agreement between two independently built maps of the same scene.
- The practical value of the protocol hinges on how often coherent distortion arises unprompted in operational captures—the paper explicitly does not measure this; if rare, the fragmentation tripwire combined with registration rate may still be operationally useful.
- A two-signal deployment design follows naturally: hold-out self-consistency to catch fragmentation and outright failure, paired with a coarse independent georeferencing channel (for example, consumer-grade GNSS) to catch coherent scale or datum drift; the paper identifies the remedy but stops short of validating such a combined gate.
- The inversion of reprojection RMSE on coherent distortion implies that routinely reported quality numbers in current photogrammetric practice may actively prefer some catastrophically wrong models in this failure regime, an unstated caution for practitioners who gate on reprojection error.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper formalises a track-leakage-free hold-out protocol for photogrammetric self-validation. A deterministic subset of images is held out; withheld views are re-localised by resection against 3D points supported by at least two retained views; the angular agreement (rotation and translation direction) is aggregated into a ground-truth-free mAA confidence. The protocol is evaluated on five GNSS-referenced captures, 13 ETH3D scenes, EuRoC, 30 IMC scenes, and two KITTI sequences. The main finding is negative: confidence saturates near 1.00 even when true RTK error swings by 14.1x within a capture, and it stays at 1.00 on coherently distorted single-component models wrong by 55–106 m, while a fragmentation failure is caught (1.00 to 0.96). The paper concludes that the protocol measures internal geometric consistency only, is useful as a qualitative fragmentation warning, and is not a substitute for control-point accuracy assessment.
Significance. The empirical core is valuable and unusually clean: existence proofs (Tuniu 0916, 63 m error at confidence 1.00; KITTI scale drift, 19 m at 1.00; ETH3D multi-camera cases at 3–4 m with confidence 1.00) are replicated across independent datasets. The paper ships reproducible scripts and per-reconstruction result tables, gives an honest power analysis, and explicitly lists limitations (underpowered null, single-knob harness confound, no validated detector yet). These strengths make the operational recommendation credible. If the theoretical overreach in Section III-B is tightened as suggested below, this will be a useful cautionary reference for GT-free quality assessment in photogrammetry.
major comments (3)
- [Section III-B, gauge/scale-unobservability mechanism] The sentence 'The blind spot therefore covers the whole family of locally near-rigid transformations' overstates the preceding argument. Exact similarities are pure gauge and leave every internal residual unchanged by construction; non-similarity near-rigid warps are not invariant in that sense. Whether their internal residuals stay below threshold is an empirical magnitude question, not an invariance theorem. The paper's own Section V-H reports wide CIs (e.g., hold-out confidence AUROC on coherent distortion [0.42, 0.93]) and Section VII concedes that the class-level claim is not tested on the sparsity gradient. Please either prove the invariance for the claimed subclass or reformulate as a conjecture supported by the limited experiments, and align the abstract's 'for a structural rather than statistical reason' with that qualification.
- [Section V-D, Table IV and discussion] The scale-aware translation variant is measured only on nominal reconstructions. The text states that 'the scale-aware variant is also saturated' and later uses this to argue that the blindness is 'the unobservable gauge, not the direction-only projection.' But the metric-offset measurement on the coherently distorted models is explicitly left to future work. Please either add that measurement to the Table VI distorted models or state prominently in the conclusion that the scale-aware behaviour on distorted models is predicted by the gauge argument, not empirically demonstrated.
- [Section V-G, last paragraph] The claim 'no informative threshold exists at any scale' is stronger than the evidence. The overlapping benign and distortion-induced bands come from small samples and show that no threshold in these data gives clean separation; they do not prove that no threshold with a useful operating point exists. Please soften to 'no threshold was validated in this study' and use the same hedged wording in the Discussion and Conclusion, which elsewhere appropriately say 'no calibrated threshold.'
minor comments (4)
- [Abstract and Section III-B] Consider harmonising the phrase 'for a structural rather than statistical reason' with the empirical nature of the near-rigid warp claim. The exact-similarity part is structural; the near-rigid part is currently a well-supported conjecture.
- [Section V-E, Table V footnote] The explanation of why the pooled coarse r equals the capture-as-unit continuous r to two decimals is confusing. A one-sentence numerical illustration or a separate column would help the reader distinguish the two estimators.
- [Table VII caption] The instruction 'Read direction, not decimals' is informal. Please state the sample sizes and full bootstrap CIs in the caption, as done for the headline cells in the text.
- [Section V-H] The sentence describing reprojection RMSE as 'fail[ing] worse than chance' is striking and important. Consider adding one explanatory sentence that a coherent warp lowers reprojection residuals, since this mechanism is central and otherwise easy to misread as an artefact.
Circularity Check
No significant circularity: the paper's negative result is an empirical characterization against external ground truth, with its definitional scale-blindness explicitly disclosed and stress-tested.
full rationale
The paper's central claims are empirical, not derived from a fitted input or a self-citation chain. The hold-out confidence is computed by a deterministic protocol (withheld images resected against track-leakage-free 3D anchors) and then compared against RTK, laser-scan, total-station, or benchmark ground truth. No parameter is fitted to the accuracy labels; the hold-out fraction is shown to be rho-invariant, and the paper explicitly declines to propose a validated threshold. The one definitional element—translation error discards scale, so pure-scale error is invisible a priori—is openly stated as 'by construction' and is then probed with a scale-aware variant (Section V-D), which also saturates; the paper uses this ablation to locate the blindness in the gauge, not to disguise a definition as a discovery. The extension to 'the whole family of locally near-rigid transformations' is explicitly labeled a prediction and tested empirically in Section V-H on failure committees, with the paper reporting wide bootstrap CIs and conceding the class-level claim is supported but not proven. There are no load-bearing self-citations: the cited classical reliability/gauge results are independent, standard literature used as context, and the empirical existence proofs (55–106 m at confidence 1.00, KITTI scale drift, ETH3D saturation) stand on released data and external references. The negative meta-analysis is honestly reported as underpowered and explicitly not the basis of the conclusion. No fitted quantity is renamed as a prediction, and no geometrical argument reduces to its own inputs.
Axiom & Free-Parameter Ledger
free parameters (4)
- hold-out fraction ρ =
0.10
- trusted-point barrier τ =
2
- mAA thresholds Θ =
{1°,3°,5°}
- split seed =
42
axioms (5)
- domain assumption Bundle adjustment has intrinsic gauge/datum freedom, so any score computed from the reconstruction's own geometry is invariant to a global similarity.
- domain assumption RTK GNSS per-image positions with RtkFlag=50 are accurate to cm-level and serve as ground truth for camera-position error.
- domain assumption COLMAP's feature matching and verified two-view geometries are reliable enough for the degradation harness to meaningfully simulate corruption.
- domain assumption Seven-parameter (Umeyama) similarity alignment is the appropriate comparison between reconstruction and ground truth, so residual error is shape distortion rather than global scale/rotation/translation.
- standard math P3P resection within a RANSAC loop returns the correct pose for a held-out view when the model is internally self-consistent.
Cite this review
Pith. "Pith review of Track-Leakage-Free Hold-Out Self-Validation for Photogrammetric Reconstruction: Protocol, Sensitivity, and Limits." pith.science (2026). https://pith.science/paper/57JTNVCX
@misc{pith2026260724852,
author = {Pith},
title = {Pith review of: Track-Leakage-Free Hold-Out Self-Validation for Photogrammetric Reconstruction: Protocol, Sensitivity, and Limits},
year = {2026},
howpublished = {\url{https://pith.science/paper/57JTNVCX}},
note = {Machine review of arXiv:2607.24852}
}
read the original abstract
Automated photogrammetric inspection emits metric measurements from reconstructions whose correctness is normally unknown without an external survey. Can a reconstruction estimate its own reliability with no ground truth, and what would such an estimate measure? We formalise a track-leakage-free hold-out protocol: a deterministic image subset is withheld, and each withheld view is re-localised by resection against only those 3D points supported by at least two retained images, so no view is tested against structure it helped create. We evaluate it on five GNSS-referenced captures (four RTK-fixed) across four sites, 13 ETH3D laser-scan scenes, a EuRoC flight, and 30 IMC 2025 scenes. The protocol is computationally well-posed -- good reconstructions score near-perfect self-consistency (median rotation error 0.003 deg) -- but it does not measure accuracy, for a structural rather than statistical reason. It saturates: confidence stays pinned at 1.00 while true error swings 14.1x within a single capture, and holds at 1.00 on survey-grade truth at 3.4 m, 4.3 m, and 1.7 m / 13 deg. It is blind to coherent distortion: corruption that fragments a reconstruction is caught (1.00 -> 0.96), but corruption yielding a single, internally self-consistent, globally distorted model is not -- at three of four captures such models were wrong by 55-106 m at confidence 1.00. On IMC 2025 the same dichotomy appears with no injected degradation: confidence separates failed from successful reconstructions (rho = 0.68) yet ranks nothing among the successful ones (rho = 0.01). A capture-level meta-analysis of the continuous signal is underpowered and sign-unstable (k = 5, 95% CI [-0.45, +0.75]); the negative result does not rest on it. Track-leakage-free hold-out therefore measures internal geometric consistency: a qualitative fragmentation warning, not a substitute for control-point accuracy assessment.
Figures
Reference graph
Works this paper leans on
-
[1]
Parallel tracking and mapping for small AR workspaces,
G. Klein and D. Murray, “Parallel tracking and mapping for small AR workspaces,” inIEEE/ACM Int. Symp. on Mixed and Augmented Reality (ISMAR), 2007
2007
-
[2]
ORB-SLAM: A versatile and accurate monocular SLAM system,
R. Mur-Artal, J. M. M. Montiel, and J. D. Tardós, “ORB-SLAM: A versatile and accurate monocular SLAM system,”IEEE Trans. on Robotics, vol. 31, no. 5, pp. 1147–1163, 2015
2015
-
[3]
Direct sparse odometry,
J. Engel, V . Koltun, and D. Cremers, “Direct sparse odometry,”IEEE Trans. on Pattern Analysis and Machine Intelligence, vol. 40, no. 3, pp. 611–625, 2018
2018
-
[4]
A multi-view stereo benchmark with high- resolution images and multi-camera videos,
T. Schöps, J. L. Schönberger, S. Galliani, T. Sattler, K. Schindler, M. Pollefeys, and A. Geiger, “A multi-view stereo benchmark with high- resolution images and multi-camera videos,” inIEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2017
2017
-
[5]
Tanks and temples: Benchmarking large-scale scene reconstruction,
A. Knapitsch, J. Park, Q.-Y . Zhou, and V . Koltun, “Tanks and temples: Benchmarking large-scale scene reconstruction,”ACM Trans. on Graph- ics, vol. 36, no. 4, 2017
2017
-
[6]
Large- scale data for multiple-view stereopsis,
H. Aanæs, R. R. Jensen, G. V ogiatzis, E. Tola, and A. B. Dahl, “Large- scale data for multiple-view stereopsis,” inInt. Journal of Computer Vision (IJCV), 2016
2016
-
[7]
Image matching across wide baselines: From paper to practice,
Y . Jin, D. Mishkin, A. Mishchuk, J. Matas, P. Fua, K. M. Yi, and E. Trulls, “Image matching across wide baselines: From paper to practice,”Int. Journal of Computer Vision (IJCV), vol. 129, pp. 517–547, 2021
2021
-
[8]
ISPRS benchmark for multi-platform photogrammetry,
F. Nex, M. Gerke, F. Remondino, H.-J. Przybilla, M. Bäumker, and A. Zurhorst, “ISPRS benchmark for multi-platform photogrammetry,” ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. II-3/W4, pp. 135–142, 2015
2015
-
[9]
UseGeo – a UA V-based multi-sensor dataset for geospatial research,
F. Nex, N. Zhang, F. Remondino, E. M. Farella, R. Qin, and I. Toschi, “UseGeo – a UA V-based multi-sensor dataset for geospatial research,” ISPRS Open Journal of Photogrammetry and Remote Sensing, vol. 12, p. 100070, 2024
2024
-
[10]
The Hessigheim 3D (H3D) benchmark on semantic segmentation of high-resolution 3D point clouds and textured meshes from UA V lidar and multi-view-stereo,
M. Kölle, D. Laupheimer, S. Schmohl, N. Haala, F. Rottensteiner, J. D. Wegner, and H. Ledoux, “The Hessigheim 3D (H3D) benchmark on semantic segmentation of high-resolution 3D point clouds and textured meshes from UA V lidar and multi-view-stereo,”ISPRS Open Journal of Photogrammetry and Remote Sensing, vol. 1, p. 100001, 2021
2021
-
[11]
Look ma, no ground truth! ground-truth-free tuning of structure from motion and visual slam,
A. Fontan, J. Civera, T. Fischer, and M. Milford, “Look ma, no ground truth! ground-truth-free tuning of structure from motion and visual slam,”arXiv preprint arXiv:2412.01116, 2024
Pith/arXiv arXiv 2024
-
[12]
RayZer: A self-supervised large view synthesis model,
H. Jiang, H. Tan, K. Sunkavalliet al., “RayZer: A self-supervised large view synthesis model,” inarXiv preprint arXiv:2505.00702, 2025
Pith/arXiv arXiv 2025
-
[13]
E-RayZer: Self-supervised 3D reconstruction as spatial visual pre-training,
Q. Zhao, H. Tan, Q. Wang, S. Bi, K. Zhang, K. Sunkavalli, S. Tulsiani, and H. Jiang, “E-RayZer: Self-supervised 3D reconstruction as spatial visual pre-training,”arXiv preprint arXiv:2512.10950, 2026
arXiv 2026
-
[14]
Benchmarking 6DOF outdoor visual localization in changing condi- tions,
T. Sattler, W. Maddern, C. Toft, A. Torii, L. Hammarstrand, E. Stenborg, D. Safari, M. Okutomi, M. Pollefeys, J. Sivic, F. Kahl, and T. Pajdla, “Benchmarking 6DOF outdoor visual localization in changing condi- tions,” inIEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2018
2018
-
[15]
Are large-scale 3d models really necessary for accurate visual localization?
T. Sattler, A. Torii, J. Sivic, M. Pollefeys, H. Taira, M. Okutomi, and T. Pajdla, “Are large-scale 3d models really necessary for accurate visual localization?” inCVPR, 2017
2017
-
[16]
Camera pose voting for large- scale image-based localization,
B. Zeisl, T. Sattler, and M. Pollefeys, “Camera pose voting for large- scale image-based localization,” inICCV, 2015
2015
-
[17]
PoseNet: A convolutional network for real-time 6-DOF camera relocalization,
A. Kendall, M. Grimes, and R. Cipolla, “PoseNet: A convolutional network for real-time 6-DOF camera relocalization,” inIEEE Int. Conf. on Computer Vision (ICCV), 2015
2015
-
[18]
From coarse to fine: Robust hierarchical localization at large scale,
P.-E. Sarlin, C. Cadena, R. Siegwart, and M. Dymczyk, “From coarse to fine: Robust hierarchical localization at large scale,” inIEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2019
2019
-
[19]
Back to the feature: Learning robust camera localization from pixels to pose,
P.-E. Sarlin, A. Unagar, M. Larsson, H. Germain, C. Toft, V . Larsson, M. Pollefeys, V . Lepetit, L. Hammarstrand, F. Kahl, and T. Sattler, “Back to the feature: Learning robust camera localization from pixels to pose,” inIEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2021
2021
-
[20]
DSAC — differentiable RANSAC for camera local- ization,
E. Brachmann, A. Krull, S. Nowozin, J. Shotton, F. Michel, S. Gumhold, and C. Rother, “DSAC — differentiable RANSAC for camera local- ization,” inIEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2017
2017
-
[21]
What uncertainties do we need in Bayesian deep learning for computer vision?
A. Kendall and Y . Gal, “What uncertainties do we need in Bayesian deep learning for computer vision?” inAdvances in Neural Information Processing Systems (NeurIPS), 2017
2017
-
[22]
Baarda,A Testing Procedure for Use in Geodetic Networks, ser
W. Baarda,A Testing Procedure for Use in Geodetic Networks, ser. Publications on Geodesy, New Series. Delft: Netherlands Geodetic Commission, 1968, vol. 2, no. 5
1968
-
[23]
Reliability analysis of parameter estimation in linear models with applications to mensuration problems in computer vision,
W. Förstner, “Reliability analysis of parameter estimation in linear models with applications to mensuration problems in computer vision,” Computer Vision, Graphics, and Image Processing, vol. 40, no. 3, pp. 273–310, 1987
1987
-
[24]
Bundle adjustment — a modern synthesis,
B. Triggs, P. F. McLauchlan, R. I. Hartley, and A. W. Fitzgibbon, “Bundle adjustment — a modern synthesis,” inVision Algorithms: Theory and Practice (ICCV Workshop), ser. Lecture Notes in Computer Science, vol. 1883. Springer, 2000, pp. 298–372
2000
-
[25]
Scale drift-aware large scale monocular SLAM,
H. Strasdat, J. M. M. Montiel, and A. J. Davison, “Scale drift-aware large scale monocular SLAM,” inRobotics: Science and Systems (RSS), 2010
2010
-
[26]
On-manifold preintegration for real-time visual-inertial odometry,
C. Forster, L. Carlone, F. Dellaert, and D. Scaramuzza, “On-manifold preintegration for real-time visual-inertial odometry,”IEEE Transactions on Robotics, vol. 33, no. 1, pp. 1–21, 2017
2017
-
[27]
An assessment of information criteria for motion model selection,
P. H. S. Torr, “An assessment of information criteria for motion model selection,” inIEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 1997
1997
-
[28]
RANSAC for (quasi-)degenerate data (QDEGSAC),
J.-M. Frahm and M. Pollefeys, “RANSAC for (quasi-)degenerate data (QDEGSAC),” inIEEE Conf. on Computer Vision and Pattern Recog- nition (CVPR), 2006
2006
-
[29]
Simple and scalable predictive uncertainty estimation using deep ensembles,
B. Lakshminarayanan, A. Pritzel, and C. Blundell, “Simple and scalable predictive uncertainty estimation using deep ensembles,” inNeurIPS, 2017
2017
-
[30]
A novel parametrization of the perspective-three-point problem for a direct computation of absolute camera position and orientation,
L. Kneip, D. Scaramuzza, and R. Siegwart, “A novel parametrization of the perspective-three-point problem for a direct computation of absolute camera position and orientation,” inCVPR, 2011
2011
-
[31]
Epnp: An accurate o(n) solution to the pnp problem,
V . Lepetit, F. Moreno-Noguer, and P. Fua, “Epnp: An accurate o(n) solution to the pnp problem,”International Journal of Computer Vision, vol. 81, no. 2, pp. 155–166, 2009
2009
-
[32]
Reconstructing the world* in six days *(as captured by the yahoo 100 million image dataset),
J. Heinly, J. L. Schönberger, E. Dunn, and J.-M. Frahm, “Reconstructing the world* in six days *(as captured by the yahoo 100 million image dataset),” inCVPR, 2015
2015
-
[33]
Least-squares estimation of transformation parameters between two point patterns,
S. Umeyama, “Least-squares estimation of transformation parameters between two point patterns,”IEEE Trans. on Pattern Analysis and Machine Intelligence, vol. 13, no. 4, pp. 376–380, 1991
1991
-
[34]
Structure-from-motion revisited,
J. L. Schönberger and J.-M. Frahm, “Structure-from-motion revisited,” inIEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2016
2016
-
[35]
Are we ready for autonomous driving? The KITTI vision benchmark suite,
A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for autonomous driving? The KITTI vision benchmark suite,” inCVPR, 2012
2012
-
[36]
Vision meets robotics: The KITTI dataset,
A. Geiger, P. Lenz, C. Stiller, and R. Urtasun, “Vision meets robotics: The KITTI dataset,”International Journal of Robotics Research, vol. 32, no. 11, pp. 1231–1237, 2013
2013
This paper was first reviewed by deepseek-v4-flash on August 4, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.