{"id":"16b42d7f-aa6c-4e81-91c9-3e57b0cc5c51","arxiv_id":"2608.02404","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"For thermal photogrammetry of heritage buildings, AI super-resolution degrades 3D reconstruction quality; native-resolution thermal images remain the most geometrically accurate, and hardware UltraMax offers only marginal gains.","lead":"A field study of Florence's Loggia dei Lanzi compared six ways to increase the resolution of thermal images before 3D photogrammetry, including hardware microscanning and three AI upscalers. It found that AI-super-resolved images produced 3D models that were less accurate and less complete than the original thermal images, while hardware-based enhancement gave only a small, marginal gain.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"AI models were tested via a 16-bit float32 roundtrip rather than their native 8-bit input; the null result may be an artifact of that integration, so the central claim is only as strong as this assumption.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the AI models were evaluated through a 16-bit float32 roundtrip instead of their native 8-bit training distribution. This is the single most consequential threat to the central claim because it directly undermines the external validity of the negative result. If the models were operating outside their designed input regime, the conclusion that 'AI super-resolution provides no photogrammetric benefit' is not established for the models as intended; it is only established for one particular, possibly unrepresentative integration. The paper's own Discussion acknowledges the tension by stating that obtaining AI output 'requires reducing 16-bit radiometric data to 8-bit imagery,' while the Methods describe a pipeline that does not do this. This internal inconsistency makes the concern concrete rather than speculative. The proposed test—re-running with 8-bit inputs and official preprocessing—would settle whether the null result is a model property or a pipeline artifact. I agree with the reader's conditional verdict: the paper is a useful benchmark, but its broader claims about AI SR in thermal heritage should be tempered until this integration question is resolved. I would not move the verdict; the existing CONDITIONAL already captures the needed qualification.","tokens_in":14211,"tokens_out":9410,"duration_ms":103104,"concrete_test":"Re-run the three AI tiers after converting the 16-bit radiometric PNGs to 8-bit exactly as in the PBVS challenge protocol (per-sequence min-max or the challenge's official normalization), then feed each model through its official preprocessing (including any channel replication and mean/std normalization). Use the identical Metashape pipeline with LiDAR-locked camera positions and fixed scaled calibration, and compare the resolution-independent metrics (RU>100 fraction and ≥3-image consistency) plus normalized reprojection error against the native, UltraMax, and bicubic tiers. If any AI tier reaches native-level quality or clearly outperforms bicubic on these metrics, the central claim is an artifact of the 16-bit float32 integration; if the ranking is unchanged, the conclusion is robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central negative result depends on deploying three AI super-resolution models in a regime they were not designed for. Section 7.2 states the models from the PBVS challenge 'operate on 8-bit data derived from surveillance-class thermal sensors,' yet the paper 'normalizes the 16-bit thermal values to float32, processes them through the model, and rescales the output to the original 16-bit temperature range.' This float32 roundtrip is not the models' native input distribution: learned statistics, batch-normalization moments, and feature detectors are calibrated for 8-bit thermal images with a particular dynamic range, not for 16-bit radiometric values that may occupy only a small fraction of the 16-bit range. The paper never specifies whether the normalization is a global 65535 scale or a per-image min-max, so the actual input distribution is unknown. If the models see an almost-black or abnormally low-contrast input, their internal representations are degraded before any photogrammetric evaluation occurs. The Discussion even argues that AI SR requires 'reducing 16-bit radiometric data to 8-bit imagery'—but the experiments deliberately avoided that reduction, instead using an ad hoc float32 path. Thus the claim 'AI super-resolution methods evaluated here ... yielded no measurable improvement ... and ... actively degraded it' conflates model failure with integration failure. The comparison may be valid for the specific float32 pipeline described, but it does not license the paper's general statements about state-of-the-art AI SR in thermal heritage photogrammetry.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper compares six resolution tiers for thermal photogrammetric reconstruction of the Loggia dei Lanzi rear wall using a FLIR T1020HD: native 1024×768, UltraMax hardware super-resolution, bicubic upscaling, SwinIR, DifIISR, and TongJi/DRCT. All tiers are processed through an identical Agisoft Metashape workflow with LiDAR-locked camera poses, a fixed resolution-scaled calibration, and a common 149-image subset. Metrics include tie-point counts, GSD-normalized reprojection error (mm), reconstruction uncertainty, cross-view consistency, and dense-cloud confidence/completeness. The central finding is that native imagery yields the most accurate and highest-quality reconstruction; UltraMax is marginally useful with the cleanest enhanced sparse cloud; all evaluated AI super-resolution methods fail to improve reconstruction, and the high-factor generative models actively degrade it, concentrating dense points on high-contrast edges.","tokens_in":14487,"tokens_out":4624,"duration_ms":51637,"significance":"If correct, this is a useful negative result for heritage thermography: quantitative photogrammetric documentation should rely on native radiometric imagery, and AI upscaling should not be assumed beneficial. The study's strengths are the controlled comparison (fixed camera network, fixed calibration, common image subset), the analytical GSD normalization tied to LiDAR geometry, the use of two resolution-independent metrics (reconstruction uncertainty and multi-view consistency), and the public release of datasets and processing scripts. The conclusion is, however, only as strong as the premise that the AI models were evaluated in a regime relevant to their design; the float32-input issue is central to that premise.","major_comments":[{"comment":"The AI models are applied to float32-normalized 16-bit thermal values, not to the 8-bit inputs on which they were trained and benchmarked. The normalization scheme is not specified (global 65535 scaling vs. per-image min-max), and the claimed roundtrip-radiometric-fidelity metrics are never reported. If the models see an almost-black or abnormally low-contrast input, the null result could reflect integration failure rather than model capability. The Discussion's argument that AI SR 'requires reducing 16-bit radiometric data to 8-bit imagery' is internally inconsistent with the float32 pipeline actually used. Please add an evaluation with proper 8-bit quantization (or at minimum report input-distribution statistics and roundtrip errors) and restrict the general conclusion accordingly.","section":"§7.2, §9, Table 1"},{"comment":"UltraMax was captured with a deliberately loosened tripod head, and no quantitative check is reported that the induced displacement matches the hand-tremor magnitude and distribution assumed by FLIR's reconstruction. The Discussion appropriately cautions about this, but the paper still calls the UltraMax comparison 'the first such independent evaluation' in §4. If the UltraMax result is to be more than a case study, please provide displacement statistics from the acquired sequences or a justification that the induced motion is representative.","section":"§7.1, §9"},{"comment":"The GSD normalization assumes a constant camera-to-wall standoff derived from LiDAR, but the paper notes local standoff variation. Since the ranking of the 2× tiers is close (17.0–18.7 mm reprojection error), the normalization could affect these small differences. Report per-station standoff values and the resulting GSD range, and confirm that the small inter-tier differences are not within the uncertainty of this assumption. This is not a fatal issue because the two resolution-independent metrics agree with the ranking, but it needs quantification.","section":"§7.4, §8.1"}],"minor_comments":[{"comment":"The column header 'RU>100≥3 img' appears to combine two separate statistics (fraction of tie points with reconstruction uncertainty above 100 and fraction observed in ≥3 images). Please split these into two columns for clarity.","section":"Table 1"},{"comment":"The sentence 'obtaining its output requires reducing 16-bit radiometric data to 8-bit imagery' contradicts §7.2, where the authors deliberately avoid 8-bit conversion and use a float32 path. If this is intended as a general property of available models, it should be stated as such and supported by model documentation, not presented as a consequence of the present experiment.","section":"§9"},{"comment":"The text says 'Carl Frey demonstrated masterfully in 1885 [1]' but reference [1] is listed as 'K. Frey.' Please reconcile the initials.","section":"§1, References"},{"comment":"The caption explains the color mapping well, but adding a color legend directly in the figure would improve readability, especially for readers viewing the figure separately from the caption.","section":"Figure 4"},{"comment":"The paper mentions that 'roundtrip radiometric fidelity (mean absolute error, RMSE, and maximum error against the native input) was recorded at each stage' but these values are not presented. Reporting them would help assess whether the float32 roundtrip itself introduces radiometric distortion that could influence the photogrammetric comparison.","section":"§7.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is well within the scope of the journal and the authors are commendably transparent about data and code. The main concern is that the central negative claim about AI super-resolution depends on a non-native input pipeline; this needs to be resolved experimentally or the claim must be narrowed. The UltraMax 'first independent evaluation' claim is also slightly overstated given the improvised jiggle mechanism, though the authors are appropriately cautious in the Discussion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — this is one of those papers where the negative result is the contribution. The authors compared six resolution tiers (native, FLIR UltraMax, bicubic, SwinIR, DifIISR, TongJi/DRCT) inside a fixed Structure-from-Motion pipeline for thermal images of the Loggia dei Lanzi, with camera positions locked to external LiDAR, calibration held fixed, and a common 149-image subset. They report that native imagery gives the best geometric accuracy by reprojection error, reconstruction uncertainty, and multi-view consistency; UltraMax is the cleanest of the enhanced tiers; and the AI upscalers degrade quality, with the high-factor generative models collapsing onto edges and leaving smooth wall surfaces unreconstructed.\n\nThat is genuinely useful. It is the first independent field test of UltraMax I know of, and the first comparison of hardware microscanning versus AI SR in a thermal SfM pipeline for heritage. The metrics are careful: they normalize reprojection error by analytically derived GSD and back it with two resolution-independent measures. Datasets, code, and per-point exports are public. That makes the study reproducible and citable as a benchmark.\n\nThe big caveat, which I think the reader got right: the AI models were run through a 16-bit-to-float32 roundtrip instead of their native 8-bit input. SwinIR, DRCT, and DifIISR are trained on 8-bit surveillance-class imagery; feeding them float-normalized 16-bit radiometric data is off-distribution. So the null result may say more about that integration than about the models. The paper is upfront about the conversion, but the Discussion goes on to say \"if obtaining its output requires reducing 16-bit radiometric data to 8-bit imagery...\" — which is precisely what they did not do. They used a float path, not an 8-bit reduction. So the general conclusion about AI SR being harmful for quantitative heritage work is stronger than the evidence supports. A cleaner test would run the models on properly converted 8-bit images (and perhaps also on the float path) and see if the ranking changes. Minor issues: the UltraMax acquisition relied on a loosened tripod head to mimic hand tremor, which the authors themselves flag as a first data point; the dataset is one site and one night; and the RU threshold of 100 could use a sensitivity check. There is also an overclaim in the abstract about \"revealing hidden architectural features\" without ground-truth validation of those features.\n\nWho is this for? People building thermal photogrammetry pipelines for heritage will want to read it, and it deserves serious peer review — with the 8-bit/16-bit question addressed or the claims scaled back. I would not cite it in my own work, but I would send it to a referee.","headline":"A careful, reproducible negative result for AI upscaling in heritage thermal photogrammetry — though the AI models were tested via a float32 roundtrip, not their native 8-bit input, so the general 'AI SR degrades reconstruction' claim needs that caveat.","tokens_in":15023,"tokens_out":3056,"would_cite":false,"duration_ms":29460,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI super-resolution of thermal images does not improve photogrammetric 3D reconstruction of heritage surfaces, and aggressive models degrade it; native imagery is the most accurate.","keywords":["thermography","AI super-resolution","photogrammetry","cultural heritage","thermal imaging","Structure-from-Motion","3D reconstruction","radiometric fidelity"],"falsifier":"Re-run the identical six-tier comparison with the same AI models applied to 8-bit-quantized thermal frames (their native training range), keeping the 16-bit native path as control; if a model then matches or beats native reconstruction uncertainty and surface completeness, the null result is an artifact of the float-normalized 16-bit roundtrip, not a property of AI upscaling.","tokens_in":14056,"feed_emoji":"🌡️","tokens_out":7382,"duration_ms":72487,"temperature":0.7,"pith_summary":"This paper asks whether AI upscaling of thermal images helps or hurts 3D photogrammetric reconstruction of heritage buildings, using a winter thermal survey of Florence's Loggia dei Lanzi as the test case. The authors compared six resolution tiers—native sensor captures, hardware pixel-shifted super-resolution, bicubic interpolation, and three state-of-the-art AI super-resolution models—inside an identical structure-from-motion pipeline with camera positions locked to LiDAR reference geometry. Their central finding is that native imagery produces the most accurate reconstruction by every measure, that hardware super-resolution comes close but adds little, and that the AI methods provide no measurable improvement and by quality metrics degrade the result, with the most aggressive models collapsing reconstruction onto edges. This matters because quantitative heritage thermography aims to measure subsurface structures, and the paper indicates that radiometric fidelity at native resolution outweighs any perceptual sharpness AI upscaling adds.","feed_headline":"AI upscaling adds no accuracy to thermal 3D scans","feed_subtitle":"Six resolution tiers on Florence's Loggia: native wins; AI detail fails multi-view checks.","key_machinery":"The load-bearing mechanism is a controlled structure-from-motion comparison: every resolution tier is processed through an identical photogrammetric network with camera positions locked to a LiDAR reference frame and a fixed, resolution-scaled camera calibration, so image content is the only variable. The decisive instrument is multi-view geometric consistency—reprojection error normalized to physical units via an analytically derived ground sample distance (the wall area each pixel covers), plus resolution-independent reconstruction uncertainty and the fraction of tie points observed in three or more images. These quantities expose whether AI-synthesized pixels triangulate to the same 3D lo","core_discovery":"The paper's central claim is that AI super-resolution, whether from a generic transformer, a thermal-challenge winner, or a diffusion model, yields no measurable improvement in photogrammetric reconstruction over native or interpolated baselines—and by quality metrics actively degrades it. Native imagery produced the most geometrically accurate reconstruction by every measure: lowest normalized reprojection error, highest cross-view consistency, and a sparse cloud essentially free of poorly triangulated points. Hardware pixel-shifted super-resolution was the only enhancement that approached native quality, but its benefit was modest. The most aggressive AI upscaling methods increased point c","pith_inferences":["An implication the paper leaves implicit: the same multi-view consistency test could be run on visible-light photogrammetry of low-texture facades; the edge-collapse failure mode may explain mixed results in AI-upscaled architectural scans.","A testable extension: feed the AI models 8-bit quantized versions of the thermal frames while keeping the 16-bit-native branch; if an 8-bit-fed model then matches native quality, the null result is an artifact of the float-normalized roundtrip rather than a general property of AI super-resolution.","The finding that raw point counts invert reconstruction value suggests previous photogrammetry studies that report only point counts may need re-evaluation with uncertainty-filtered metrics.","If hardware microscanning were driven by precise actuator-controlled sub-pixel shifts rather than improvised tripod jiggle, a small genuine benefit might appear; this remains untested and is a natural next experiment."],"forward_implications":["For quantitative heritage thermography, native-resolution 16-bit imagery should be the default; AI upscaling should not be used when measurements are the goal.","Hardware pixel-shifted super-resolution adds little geometric value in the field but remains a low-risk addition; it should not substitute for good acquisition conditions.","Aggressive 4x and 8x AI upscaling actively degrades sparse-cloud quality and collapses dense reconstruction onto edges, leaving the smooth wall surfaces—where hidden features appear—unreconstructed.","Tie-point and dense-point counts are misleading; high-uncertainty or single-depth-map points inflate them, so confidence-filtered quality metrics are essential in reconstruction evaluation.","Because AI super-resolution requires converting 16-bit radiometric data to the 8-bit range most models accept, adopting it discards calibrated temperature information that subsurface analysis depends on."],"fun_headline_variants":["AI upscaling fails to boost thermal 3D accuracy","Native wins: AI super-resolution degrades thermal models","Thermal 3D: AI enhancement loses to plain resolution","AI sharpening adds no value to heritage thermography","Super-resolution hype vs. reality in thermal photogrammetry"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the AI models were fairly tested, but they were run on float-normalized 16-bit radiometric data rather than their native 8-bit training domain; if that mismatch degrades their feature representations, the null result may be an artifact of the integration pipeline rather than a property of the models.","fun_headline_variants_meta":{"raw":{"variants":["AI upscaling fails to boost thermal 3D accuracy","Native wins: AI super-resolution degrades thermal models","Thermal 3D: AI enhancement loses to plain resolution","AI sharpening adds no value to heritage thermography","Super-resolution hype vs. reality in thermal photogrammetry"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000155,"raw_usage":{"total_tokens":1065,"prompt_tokens":769,"completion_tokens":296,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":229}},"tokens_in":513,"tokens_out":296,"duration_ms":3523,"temperature":1.0,"reasoning_tokens":229,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T07:52:17.954962+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the identical six-tier comparison with the same AI models applied to 8-bit-quantized thermal frames (their native training range), keeping the 16-bit native path as control; if a model then matches or beats native reconstruction uncertainty and surface completeness, the null result is an artifact of the float-normalized 16-bit roundtrip, not a property of AI upscaling.","supporting_citations":[],"review_version":1}