{"id":"89917e8e-34ca-4573-8443-b29be2546c92","arxiv_id":"2607.24852","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A track-leakage-free hold-out self-consistency score saturates at 1.00 while true reconstruction error swings up to 106 m, proving internal consistency is not absolute accuracy.","lead":"A self-check for 3D reconstructions that needs no ground truth was tested across drone, street, and benchmark data: it measures internal consistency, not accuracy. The check stays perfectly confident while reconstructions are wrong by up to 106 metres, so it can flag breakage but can never replace surveyed control points.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section III-B overgeneralizes gauge blindness to all locally near-rigid warps; the paper itself concedes a ground-truth-free cross-submodel escape, so the universal 'whole family' claim is not established.","rationale":"The reader's weakest assumption — that gauge invariance is extrapolated to the whole locally near-rigid family and that only a few internal proxies were tested — is exactly the concern I find most load-bearing. I agree with the reader that the central negative result is well supported by saturation (33/35 rows at confidence 1.00 with up to 14.1x error swing) and by the injected-corruption existence proofs (three of four captures at 55–106 m with confidence 1.00; KITTI scale drift 19 m at 1.00). These are direct observations that no correlation analysis can flip. The conditional element is the generality of the theoretical framing in Section III-B and the empirical reach of Section V-H. The paper is unusually honest about its limitations (Section VII explicitly lists the unrun ROC/PR analysis, the harness confound, the small-sample AUROCs, and the unmeasured frequency of unprompted coherent distortion), which supports a conditional rather than a reject verdict. My recommended verdict is therefore unchanged from the reader's CONDITIONAL: the core claim stands, but the universal 'blind spot covers the whole family' statement should be qualified, and the proposed cross-submodel test would settle whether that qualification is needed. I do not see grounds for rejection: the existence proofs make the 'not a substitute for control points' conclusion robust even if a future internal signal detects some coherent distortions. The concern is about the scope of the limit claim, not about the soundness of the protocol or the honesty of the evaluation.","tokens_in":28526,"tokens_out":9407,"duration_ms":96569,"concrete_test":"Use the released coherent-distortion Tuniu 0916 model (63.65 m RTK RMSE, hold-out confidence 1.00) and the nominal Tuniu 0916 model. Split each model's registered images into two disjoint halves with the same deterministic hash used in the paper, run COLMAP independently on each half (same parameters as the main pipeline), Sim(3)-align the two resulting submodels, and record the median post-alignment camera-centre residual. Repeat over several splits. If the distorted-model residual distribution is cleanly separated from the nominal-model distribution, then a ground-truth-free cross-submodel signal detects the very failure the hold-out missed, and Section III-B's 'whole family' claim must be weakened. If the residuals overlap, the universal blind-spot statement survives this test.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The core negative result — saturation and existence of 55–106 m / confidence 1.00 failures — is solid and does not rest on the universalization. But Section III-B moves from exact gauge invariance (true for similarities) to 'The blind spot therefore covers the whole family of locally near-rigid transformations' and asserts that BA covariance, reprojection statistics, or track length are blind. This is an extrapolation, not a theorem: a smooth non-similarity warp changes internal residuals in principle, and whether the change is sub-threshold is an empirical magnitude question, not a gauge invariance. The paper immediately concedes an escape route — 'cross-consistency between independently-built sub-models' — which is ground-truth-free yet reaches outside a single model's own geometry, undermining the categorical phrasing. Section V-H tests only a handful of proxies on small committees (n=16–19; AUROC CIs span chance), so the class-level claim 'every consistency-based proxy ... saturates or inverts' is not proven. The practical advice ('not a substitute for control-point accuracy') survives; what is at risk is the stronger limit claim that no internal signal can ever detect coherent distortion and therefore that no validated threshold can exist. That stronger claim needs either proof or a narrower formulation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper formalises a track-leakage-free hold-out protocol for photogrammetric self-validation. A deterministic subset of images is held out; withheld views are re-localised by resection against 3D points supported by at least two retained views; the angular agreement (rotation and translation direction) is aggregated into a ground-truth-free mAA confidence. The protocol is evaluated on five GNSS-referenced captures, 13 ETH3D scenes, EuRoC, 30 IMC scenes, and two KITTI sequences. The main finding is negative: confidence saturates near 1.00 even when true RTK error swings by 14.1x within a capture, and it stays at 1.00 on coherently distorted single-component models wrong by 55–106 m, while a fragmentation failure is caught (1.00 to 0.96). The paper concludes that the protocol measures internal geometric consistency only, is useful as a qualitative fragmentation warning, and is not a substitute for control-point accuracy assessment.","tokens_in":28847,"tokens_out":10798,"duration_ms":106125,"significance":"The empirical core is valuable and unusually clean: existence proofs (Tuniu 0916, 63 m error at confidence 1.00; KITTI scale drift, 19 m at 1.00; ETH3D multi-camera cases at 3–4 m with confidence 1.00) are replicated across independent datasets. The paper ships reproducible scripts and per-reconstruction result tables, gives an honest power analysis, and explicitly lists limitations (underpowered null, single-knob harness confound, no validated detector yet). These strengths make the operational recommendation credible. If the theoretical overreach in Section III-B is tightened as suggested below, this will be a useful cautionary reference for GT-free quality assessment in photogrammetry.","major_comments":[{"comment":"The sentence 'The blind spot therefore covers the whole family of locally near-rigid transformations' overstates the preceding argument. Exact similarities are pure gauge and leave every internal residual unchanged by construction; non-similarity near-rigid warps are not invariant in that sense. Whether their internal residuals stay below threshold is an empirical magnitude question, not an invariance theorem. The paper's own Section V-H reports wide CIs (e.g., hold-out confidence AUROC on coherent distortion [0.42, 0.93]) and Section VII concedes that the class-level claim is not tested on the sparsity gradient. Please either prove the invariance for the claimed subclass or reformulate as a conjecture supported by the limited experiments, and align the abstract's 'for a structural rather than statistical reason' with that qualification.","section":"Section III-B, gauge/scale-unobservability mechanism"},{"comment":"The scale-aware translation variant is measured only on nominal reconstructions. The text states that 'the scale-aware variant is also saturated' and later uses this to argue that the blindness is 'the unobservable gauge, not the direction-only projection.' But the metric-offset measurement on the coherently distorted models is explicitly left to future work. Please either add that measurement to the Table VI distorted models or state prominently in the conclusion that the scale-aware behaviour on distorted models is predicted by the gauge argument, not empirically demonstrated.","section":"Section V-D, Table IV and discussion"},{"comment":"The claim 'no informative threshold exists at any scale' is stronger than the evidence. The overlapping benign and distortion-induced bands come from small samples and show that no threshold in these data gives clean separation; they do not prove that no threshold with a useful operating point exists. Please soften to 'no threshold was validated in this study' and use the same hedged wording in the Discussion and Conclusion, which elsewhere appropriately say 'no calibrated threshold.'","section":"Section V-G, last paragraph"}],"minor_comments":[{"comment":"Consider harmonising the phrase 'for a structural rather than statistical reason' with the empirical nature of the near-rigid warp claim. The exact-similarity part is structural; the near-rigid part is currently a well-supported conjecture.","section":"Abstract and Section III-B"},{"comment":"The explanation of why the pooled coarse r equals the capture-as-unit continuous r to two decimals is confusing. A one-sentence numerical illustration or a separate column would help the reader distinguish the two estimators.","section":"Section V-E, Table V footnote"},{"comment":"The instruction 'Read direction, not decimals' is informal. Please state the sample sizes and full bootstrap CIs in the caption, as done for the headline cells in the text.","section":"Table VII caption"},{"comment":"The sentence describing reprojection RMSE as 'fail[ing] worse than chance' is striking and important. Consider adding one explanatory sentence that a coherent warp lowers reprojection residuals, since this mechanism is central and otherwise easy to misread as an artefact.","section":"Section V-H"}],"recommendation":"major_revision","confidential_remarks":"The paper is a strong negative result with reproducible code and unusually transparent handling of limitations. The main risk is that the theoretical framing in Section III-B claims more than the experiments show; the requested revision is to narrow or prove the 'whole family' claim. I expect this to be straightforward. No concerns about scope for The Photogrammetric Record."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, quick read on 2607.24852. The thing to know: the paper's negative result is real. They formalize a track-leakage-free hold-out protocol, show it is well-posed, and then show it saturates—confidence pinned at 1.00 while true RTK error swings 14x, and injected corruption produces single internally-consistent models wrong by 55–106 m at confidence 1.00. That's an existence proof, replicated on three of four captures, plus KITTI scale drift unprompted. This is the kind of result people will want to cite before deploying self-validation as an accuracy substitute.\n\nWhat's genuinely new: the track-level leakage barrier is a precise, needed protocol element; the empirical characterization across COLMAP/RTK, ETH3D, KITTI, and IMC 2025 is the first modern demonstration I know of. The paper is also unusually honest: it reports the inclusion-sensitive pooled correlation (naive −0.08, flips to +0.5 under its own primary criterion), the harness confound, the underpowered null, the wide CIs on the proxy AUROCs. That transparency earns real credit.\n\nSoft spots, in proportion. The core negative claim—internal consistency is not accuracy, and the signal is at best a fragmentation tripwire—is supported. But Section III-B moves from exact gauge invariance (true for similarities) to \"the blind spot covers the whole family of locally near-rigid transformations,\" and that is an extrapolation, not a theorem. The paper itself concedes the escape: cross-consistency between independently-built sub-models is ground-truth-free and can in principle detect coherent warps. So the categorical phrasing is too strong. The empirical test of competing proxies is underpowered (n=16–19 per regime, bootstrap CIs spanning chance), so the class-level claim that every consistency-based proxy saturates or inverts is not proven. And the paper does not measure how often coherent distortion arises unprompted in operational captures—a real limitation for the practical guidance, though not for the theoretical point.\n\nMy take: the load-bearing result is the saturation plus the existence proofs, and those hold up. The \"no internal signal can ever see this\" claim needs either a proof or a narrower formulation. The paper deserves a serious peer review; the referee should push on Section III-B and ask for a more careful statement of what is proven vs. conjectured. I'd cite this in my own work, and I'd bring it to the reading group.","headline":"Honest, well-executed negative result on self-validation; the core claim is solid, the 'whole family' generalization is overreach.","tokens_in":29301,"tokens_out":2228,"would_cite":true,"duration_ms":25014,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Track-leakage-free hold-out self-validation is a fragmentation warning, not an accuracy certificate.","keywords":["photogrammetric reconstruction","self-validation","ground-truth-free confidence","hold-out protocol","internal geometric consistency","gauge freedom","structure-from-motion","failure detection"],"falsifier":"Run the hold-out protocol on a single connected reconstruction that is globally wrong by tens of metres after similarity alignment, with no fragmentation: if the confidence drops, the blindness claim fails. If a purely internal signal—such as disagreement between two independently built sub-models of the same scene—flags that model while the hold-out score stays at 1.00, then the claim that the blind spot covers the whole locally-near-rigid family is over-broad. A practical test: a benchmark of roughly 50 operational captures with survey truth and no shared degradation knob, checking whether a","tokens_in":28410,"feed_emoji":"📐","tokens_out":5498,"duration_ms":71678,"temperature":0.7,"pith_summary":"This paper asks whether a photogrammetric reconstruction can grade its own reliability without ground truth, and answers with a precise protocol and a negative result. The protocol withholds a deterministic subset of images and re-localises each withheld view against only 3D points supported by at least two retained images, so no view is scored against structure it helped create. On operational RTK captures, public benchmarks, and injected corruption, the resulting confidence is internally well-posed but saturates: it stays pinned at 1.00 while true error swings by up to 14.1x within a capture, and reaches 55–106 m on models that are single, self-consistent global distortions. The paper concludes that the score measures internal geometric consistency only—a qualitative fragmentation warning, not a substitute for control-point accuracy assessment. The reason is structural: bundle adjustment has gauge/datum freedom, so any score computed purely from the reconstruction's own geometry is blind to coherent locally near-rigid warps.","feed_headline":"Model check scores 1.00 while true error hits 106 m","feed_subtitle":"A leakage-free hold-out protocol catches fragmentation but is blind to coherent distortion, so it cannot certify accuracy.","key_machinery":"The load-bearing mechanism is the track-level leakage barrier: a deterministic subset of images is withheld, and each withheld view is re-localised by PnP resection against only 3D points observed by at least two retained images, so no view is tested against structure it helped triangulate. Disagreement is measured as a max of geodesic rotation error and scale-free translation-direction error, aggregated as mean average accuracy over 1°, 3°, and 5° thresholds. The paper then invokes bundle adjustment's gauge/datum freedom to argue that this score—and any internal consistency score—is blind to coherent locally near-rigid distortion, with the hold-out fraction and barrier strength shown to be","core_discovery":"The paper's central claim is that a track-leakage-free hold-out self-validation score—despite being computationally well-posed and able to detect model fragmentation—is not an accuracy measure, and cannot be turned into one by tuning thresholds or restoring the scale axis. The score saturates: on good reconstructions it reports confidence 1.00 with sub-hundredth-degree agreement, but it also reports 1.00 on RTK-referenced models that are 3.4–4.3 m wrong and on injected single-component distortions that are 55–106 m wrong. The blindness is argued to be structural rather than statistical: bundle adjustment has gauge/datum freedom, so any score computed purely from the reconstruction's own geom","pith_inferences":["The paper's structural argument implies that any future purely internal 'accuracy' signal must either leave the locally-near-rigid family or add information from outside the single reconstruction; a concrete untested next experiment is cross-submodel agreement between two independently built maps of the same scene.","The practical value of the protocol hinges on how often coherent distortion arises unprompted in operational captures—the paper explicitly does not measure this; if rare, the fragmentation tripwire combined with registration rate may still be operationally useful.","A two-signal deployment design follows naturally: hold-out self-consistency to catch fragmentation and outright failure, paired with a coarse independent georeferencing channel (for example, consumer-grade GNSS) to catch coherent scale or datum drift; the paper identifies the remedy but stops short of validating such a combined gate.","The inversion of reprojection RMSE on coherent distortion implies that routinely reported quality numbers in current photogrammetric practice may actively prefer some catastrophically wrong models in this failure regime, an unstated caution for practitioners who gate on reprojection error."],"forward_implications":["A high self-consistency score cannot replace a control-point accuracy check; models wrong by tens of metres pass with confidence 1.00.","A confidence drop is usable only as a qualitative fragmentation warning, not a calibrated detector: benign undersampling can lower confidence as much as fragmentation.","The blind spot is shared by the wider class of internal consistency signals: bundle-adjustment covariance, reprojection RMSE (which can even invert), and track-length statistics when not riding a degradation-harness confound.","Escaping the blind spot requires information from outside the reconstruction's own geometry: an external datum, world priors, or cross-consistency between independently built sub-models.","No validated threshold on the confidence signal currently separates failure from benign undersampling; the negative result rests on the saturation and existence proofs, not on the underpowered correlation analysis."],"fun_headline_variants":["Self-consistency 1.00, yet 106 m error","Perfect score, 106 m wrong","Hold-out protocol blind to global distortion","Confidence 1.00 but accuracy fails","Track-leak-free check can't catch scale errors"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that no internally computed signal—the hold-out score, bundle-adjustment covariance, reprojection statistics, or track counts—can detect a coherent locally near-rigid distortion, an extrapolation from exact gauge invariance tested on only a few injected warps and a handful of proxy statistics on small samples.","fun_headline_variants_meta":{"raw":{"variants":["Self-consistency 1.00, yet 106 m error","Perfect score, 106 m wrong","Hold-out protocol blind to global distortion","Confidence 1.00 but accuracy fails","Track-leak-free check can't catch scale errors"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000235,"raw_usage":{"total_tokens":1441,"prompt_tokens":954,"completion_tokens":487,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":698,"completion_tokens_details":{"reasoning_tokens":414}},"tokens_in":698,"tokens_out":487,"duration_ms":6274,"temperature":1.0,"reasoning_tokens":414,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T01:29:26.956295+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the hold-out protocol on a single connected reconstruction that is globally wrong by tens of metres after similarity alignment, with no fragmentation: if the confidence drops, the blindness claim fails. If a purely internal signal—such as disagreement between two independently built sub-models of the same scene—flags that model while the hold-out score stays at 1.00, then the claim that the blind spot covers the whole locally-near-rigid family is over-broad. A practical test: a benchmark of roughly 50 operational captures with survey truth and no shared degradation knob, checking whether a","supporting_citations":[],"review_version":2}