{"id":"e5017ebf-9ce1-4890-866e-90e879989f42","arxiv_id":"2512.14200","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"SkyLume contributes 10 real-world UAV urban regions captured at morning, noon, and afternoon with LiDAR-based ground truth, plus the Temporal Consistency Coefficient metric for cross-time albedo stability.","lead":"SkyLume is a new drone dataset that photographs the same ten urban areas at three times of day, adding LiDAR scans and depth/normal maps so 3D reconstruction methods can be tested under changing sunlight. The paper also introduces a metric, TCC, for checking whether inverse-rendering methods separate lighting from true surface color.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Static-scene assumption is unquantified; if urban dynamics are non-negligible, the dataset's 'systematic illumination isolation' and cross-time TCC comparisons are confounded.","rationale":"The reader's weakest assumption points to the static-scene condition, and that is also the most load-bearing concern here. The paper's headline contribution is a real-world dataset that 'systematically isolates' illumination changes; every downstream experiment — TCC-albedo, cross-time geometry, cross-time mesh consistency — relies on the three slots being comparable except for lighting. The paper repeatedly states this (Sec. 3.2, supplementary Sec. 6.1, Fig. 6 caption) but provides no quantification. Urban scenes are dynamic, and the paper itself lists water, glass, and mixed weather conditions, which are precisely where dynamics are worst. This is not an external objection to the consensus; it is an internal mismatch between a strong claim ('solely illumination' / 'predominantly illumination') and the absence of evidence for it. The reader's other conditions (release, large-scale inclusion, TCC degeneracy) are legitimate but secondary: release is a logistics matter, large-scale evaluation is about scope, and the TCC constant-albedo degeneracy would be partially mitigated by the static-scene check since a degenerate baseline would still be detectable with a fidelity term. The proposed test — semantic segmentation of dynamic classes plus a static-pixel TCC recomputation — would settle the concern. Until that is provided, the dataset and its benchmark should be treated as conditional, not as a fully validated illumination-isolation standard.","tokens_in":17995,"tokens_out":5627,"duration_ms":54232,"concrete_test":"Apply a pretrained semantic segmenter (e.g., Mask2Former) to all images from each time slot in all ten scenes; compute the per-slot pixel fraction of dynamic classes (vehicles, pedestrians, vegetation, water), then compute the mean absolute color difference between time-slot pairs on static-only pixels versus all pixels. If any scene has >1% dynamic-class pixels, or if restricting the TCC computation to static pixels changes the TCC scores in Tables 2-3 by more than 0.05 or reorders the baselines, the static-scene assumption fails materially and cross-time differences cannot be attributed to illumination alone.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The load-bearing step is the claim that the three time slots differ only (or predominantly) in illumination. Section 3.2 states the same RTK waypoint trajectory is repeated for each time slot, and the supplement (Sec. 6.1) asserts that differences across slots 'predominantly reflect illumination rather than changes in scene structure'; Fig. 6's caption goes further, saying 'differences arise solely from illumination.' Yet no masking, deghosting, or dynamic-object filtering is described in Sec. 3 or the supplement. The scenes contain vehicles, pedestrians, vegetation, rivers/lakes, and glass facades (Table 6 lists water in six scenes and glass in three). Between morning, noon, and afternoon captures, cars and boats move, tree canopies and water surfaces change, and partly cloudy slots contain moving cloud shadows. Because TCC renders albedo at identical test viewpoints and attributes all disagreement across slots to illumination, any such dynamics inflate the TCC differences and confound method rankings; TCC-Geometry similarly treats mesh disagreement across slots as illumination-induced. The central claim that SkyLume supports city-scale evaluation under varying illumination depends on this isolation being valid. As written, the only evidence is the repeated-trajectory statement, which controls viewpoint but not scene content.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"SkyLume is introduced as a large-scale, real-world UAV aerial dataset for urban scene reconstruction under varying illumination. The dataset comprises 10 urban regions, more than 100K six-megapixel images captured from five directions (four oblique and one nadir) at three times of day (morning, noon, afternoon), with per-scene LiDAR scans, unified 6-DoF poses, LiDAR-derived ground-truth depth and normals, and solar-geometry annotations. The paper also proposes Temporal Consistency Coefficient (TCC), a metric intended to evaluate cross-time albedo stability for inverse rendering, and benchmarks several 3D Gaussian Splatting variants on geometry, novel view synthesis, and albedo-consistency tracks. The central claim is that SkyLume is the first real-world aerial dataset that systematically isolates illumination changes from viewpoint changes at city scale, enabling rigorous evaluation of illumination-robust reconstruction.","tokens_in":18259,"tokens_out":5379,"duration_ms":48701,"significance":"If the central claims hold, SkyLume fills a genuine gap in aerial 3D vision: prior real-world UAV datasets either do not revisit the same area under systematically varying illumination or lack high-precision geometry. The dataset's acquisition is carefully described, and the reported alignment statistics (median reprojection error 0.70 px, millimeter-level relative pose uncertainty, Table 7) suggest a high-quality, unified multi-temporal SfM solution. The per-frame LiDAR depth and solar-geometry annotations add practical value for downstream tasks. The benchmark experiments reveal plausible trends, such as geometry degradation under hard shadows. However, the evaluation methodology has two load-bearing weaknesses: (1) the static-scene assumption under which illumination is isolated is asserted but not quantified or enforced; (2) the TCC metric measures consistency against the temporal mean of the very albedo maps being evaluated, so it cannot validate whether the disentangled albedo is correct, and it rewards trivially constant albedo. These issues do not necessarily invalidate the dataset itself, but they currently weaken the claimed contributions. With substantial revisions—parti","major_comments":[{"comment":"yes","section":"§3.2, §6.1, Table 6, Fig. 6 caption"},{"comment":"yes","section":"§4.1, Eq. (8)-(14)"},{"comment":"yes","section":"§3.3, Fig. 3, §3.2"},{"comment":"yes","section":"General"}],"minor_comments":[{"comment":"The text says 'Period 0 and Period 1' (Sec. 4.3) but Table 3 lists 'Period 1' and 'Period 2'. This inconsistency should be corrected. Also, 'between the the meshes' is a typo.","section":"§4.3, Table 3"},{"comment":"The constants in the saturation functions are not justified. Units of MAE and RMSE should be stated (e.g., 0-1 albedo range), and the choice of 10 should be discussed.","section":"Eq. (10)-(11)"},{"comment":"Typo 'ISPRS-Bencnmark' should be 'ISPRS-Benchmark'. Also, the 'Light' column header could be clarified as 'Varying Illumination'.","section":"Table 1"},{"comment":"The phrase 'differences arise solely from illumination' is too strong given the static-scene issue; it should be rephrased to match the more cautious claim in the supplementary.","section":"Fig. 6 caption"},{"comment":"References [11] and [12] are duplicates of the same 3DGS paper. Please consolidate.","section":"References"},{"comment":"Typo 'RealityScanto' appears in Sec. 3.3; also 'COLMAP' formatting is inconsistent.","section":"§3.3"}],"recommendation":"major_revision","confidential_remarks":"The dataset is potentially valuable, and the acquisition/alignment effort is substantial. However, the current manuscript's evaluation protocol is undermined by the unquantified static-scene assumption and the self-referential TCC metric. These are not mere presentation issues: they affect the validity of the benchmark results and the strength of the dataset's advertised purpose. I believe the authors can address them in a major revision by adding dynamic-object masks or change maps, validating TCC on synthetic data, and providing independent geometry checks. I would not recommend rejection, as the dataset itself, if properly qualified, could make a useful contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague—\n\nSkyLume is worth a look. The dataset is the contribution, and it's a real one: repeated same-trajectory UAV flights at three times of day across ten urban regions, with LiDAR-derived ground truth meshes, per-frame depth and normals, and per-frame solar geometry. That combination is genuinely absent from the cited prior work. The alignment statistics (median reprojection error 0.70 px; relative pose uncertainty around a millimeter) and the acquisition/processing description back the claim that the geometry is usable. The benchmark with 3DGS variants does its job—it shows that lighting shifts hurt albedo stability, geometry, and novel-view quality, which is useful evidence for the community.\n\nThe soft spots are about what the paper promises versus what it shows. The 'systematically isolate illumination changes' claim is load-bearing and unquantified. The supplement (Sec. 6.1) says differences across slots 'predominantly reflect illumination rather than changes in scene structure,' and Figure 6's caption says 'solely.' But the scenes contain vehicles, water, glass facades, and partly cloudy conditions with moving shadows, and no masking or dynamic-object filtering is described. If that's not handled, TCC and cross-time geometry attribute scene dynamics to illumination. The authors need to either show dynamic regions are negligible or mask them.\n\nSecond, TCC is self-referential: each albedo map is compared to the temporal mean of the same three albedo maps (Eq. 8–14). A method that outputs a constant albedo across time would score perfectly on the MAE/RMSE components, and the paper does not discuss this degeneracy. Since there is no external albedo ground truth, TCC is a consistency measure, not an accuracy measure, and should be framed that way.\n\nThird, there is no release URL, license, or checksums in the manuscript. For a dataset paper that's a practical showstopper and an easy fix. Fourth, the large-scale scenes are collected but the benchmark runs on six small/medium scenes; calling the results 'city-scale' is a stretch.\n\nNone of this invalidates the dataset. The paper deserves a serious referee; it needs a revision that quantifies the static-scene assumption, tightens the TCC claims, and actually releases the data. I'd cite it once the data is accessible.","headline":"SkyLume's dataset is the real contribution and worth serious refereeing; the static-scene assumption is unquantified, TCC is self-referential, and there's no release link yet.","tokens_in":18828,"tokens_out":4372,"would_cite":true,"duration_ms":35214,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SkyLume isolates illumination from viewpoint in 100K+ urban aerial images, showing current 3D reconstruction methods are not robust to sunlight changes.","keywords":["UAV dataset","illumination variation","3D reconstruction","Gaussian splatting","inverse rendering","albedo consistency","LiDAR ground truth","novel view synthesis"],"falsifier":"Compare TCC scores and geometry F1 in regions with known moving objects (vehicles, pedestrians, cloud-shadow boundaries) versus static regions; if dynamic regions show systematically higher instability, the attribution to illumination alone fails.","tokens_in":17851,"feed_emoji":"☀️","tokens_out":6451,"duration_ms":56408,"temperature":0.7,"pith_summary":"The paper introduces SkyLume, a large-scale real-world aerial dataset that photographs the same ten urban regions at three times of day—morning, noon, and late afternoon—with LiDAR ground truth, registered camera poses, and per-frame depth and normals. The central claim is that this is the first benchmark that isolates illumination change from viewpoint change at city scale, so differences between captures can be attributed to lighting rather than to camera motion. Using this resource, the authors show that current Gaussian-splatting reconstruction methods (a technique that models scenes as clouds of 3D ellipsoids) degrade noticeably under strong sunlight: shadows get baked into geometry, albedo estimates drift across time, and rendering quality drops on facades, glass, and water. They also introduce a metric, the Temporal Consistency Coefficient (TCC), that scores cross-time albedo stability, turning multi-temporal robustness into a measurable target.","feed_headline":"New aerial dataset shows sunlight wrecks 3D city models","feed_subtitle":"Repeated flights over 10 regions plus LiDAR truth reveal how shadows skew geometry and albedo.","key_machinery":"The load-bearing mechanism is the capture protocol: repeating the same flight trajectory at three times of day and registering all images to a common LiDAR-guided coordinate system, which converts multi-temporal capture into a controlled illumination experiment. The TCC metric is the evaluation mechanism: it renders albedo from K fixed test viewpoints in each of the three time slots, compares each slot against the temporal mean, and combines MAE, RMSE, SSIM, and LPIPS into a single [0,1] consistency score.","core_discovery":"SkyLume consists of more than 100,000 six-megapixel images from five synchronized views over ten urban regions, each region flown in the morning, at noon, and in the late afternoon along identical RTK-guided waypoints. LiDAR scans provide metric ground truth, and a unified structure-from-motion registration locks all time slots into one coordinate system, so the only intended variable is illumination. Benchmarked 3D Gaussian Splatting variants show substantial quality loss under direct sun: cast shadows and moving penumbras are frequently reconstructed as solid geometry, albedo from inverse rendering retains shading and varies across time, and novel view synthesis blurs on weakly textured an","pith_inferences":["The static-scene assumption is the clear risk: urban areas contain moving vehicles, pedestrians, and cloud shadows, so TCC and cross-time geometry scores could conflate illumination effects with object dynamics; masking dynamic regions would sharpen the benchmark.","The TCC design—fixed viewpoints, temporal mean, per-viewpoint aggregation—is general enough to be applied to other repeat-photography setups, such as ground-level cameras or multi-day satellite revisits, as a measure of appearance stability.","The observed result that diffuse illumination yields the most stable geometry suggests a practical capture policy: fly under overcast conditions when geometry is the priority, and use sunlit slots for appearance and albedo studies.","The dataset could double as a real-world domain-shift test for generalizable reconstruction models trained on synthetic data, since it provides true illumination variation with ground truth."],"forward_implications":["Current 3D Gaussian Splatting methods can now be ranked and improved on a controlled illumination-robustness benchmark, rather than on ad hoc captures.","Researchers can isolate shadow-induced geometry bias by comparing reconstructions from sunlit versus overcast slots against the same LiDAR ground truth.","The provided solar geometry (elevation and azimuth per image) enables future de-shadowing, relighting, and illumination-aware training.","The standardized splits allow direct comparison of novel view synthesis methods under identical viewpoint and lighting conditions.","City-scale inverse rendering can be evaluated for material/light disentanglement via TCC, a task previously limited to small objects or synthetic scenes."],"fun_headline_variants":["Aerial dataset reveals sunlight melts 3D city geometry","100k UAV images show why 3D models fail in sun","Sunlight turns shadows into solid 3D errors","SkyLume: benchmarking 3D robustness to daily light shifts"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"Each urban scene is effectively static across the three capture sessions, so measured differences are attributed to illumination rather than to scene change.","fun_headline_variants_meta":{"raw":{"variants":["Aerial dataset reveals sunlight melts 3D city geometry","100k UAV images show why 3D models fail in sun","Sunlight turns shadows into solid 3D errors","SkyLume: benchmarking 3D robustness to daily light shifts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000294,"raw_usage":{"total_tokens":1577,"prompt_tokens":804,"completion_tokens":773,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":548,"completion_tokens_details":{"reasoning_tokens":703}},"tokens_in":548,"tokens_out":773,"duration_ms":7645,"temperature":1.0,"reasoning_tokens":703,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T16:13:00.604043+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare TCC scores and geometry F1 in regions with known moving objects (vehicles, pedestrians, cloud-shadow boundaries) versus static regions; if dynamic regions show systematically higher instability, the attribution to illumination alone fails.","supporting_citations":[],"review_version":1}