{"id":"0248a998-73d3-4d95-b0f3-52172472d2ea","arxiv_id":"2608.07640","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"HeatCast is a new open dataset and evaluation standard for monthly 30 m land surface temperature forecasting across 124 U.S. cities, with baseline RMSEs around 7.7 K.","lead":"HeatCast is a new open benchmark that packages monthly 30 m satellite data for 124 U.S. cities so researchers can compare how well models forecast next-month land surface temperature. The paper's first baselines show a transformer reaching 7.74 K error, but the claim needs checking because the thermal data is natively coarser and only single runs were used.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The '30 m neighborhood-scale' claim rests on a LST product with 100 m native thermal support; if the effective resolution is 100 m, the headline characterization is unsupported, though the benchmark itself remains usable.","rationale":"The reader's weakest assumption is exactly the concern I find most load-bearing: the benchmark's headline claim of 30 m neighborhood-scale LST forecasting depends on the thermal information content of a product whose native resolution is 100 m. The authors are transparent about this in Section 9 and in Table 2, which is a point in their favor, but transparency does not remove the correctness risk. If the effective thermal support is 100 m, then the title and abstract overstate the spatial resolution of the LST target, and any downstream use of the benchmark for block-scale thermal analysis inherits that overstatement. The benchmark's existence, release, fixed split, and reproducible protocol are independently valuable and are not called into question by this concern; the reported RMSE values are internally plausible and the release artifacts are described in unusual detail. The concern is therefore not fatal, but it is real and it justifies a conditional acceptance until either the authors temper the resolution claim or provide evidence of 30 m thermal information content. Since the reader already reached CONDITIONAL and the concrete remedy (either temper the wording or justify the 100 m-to-30 m resampling) is the same, my stress-test does not change the verdict.","tokens_in":13014,"tokens_out":8975,"duration_ms":94314,"concrete_test":"Select one cloud-free monthly scene from the released San Antonio Zarr store and compute the empirical semivariogram of LST along the 30 m grid, estimating range, sill, and nugget. Repeat after aggregating to 100 m cells. If the 30 m semivariogram shows no spatial structure at lags below 100 m (within-TIRS-pixel variance at the noise floor and range near 100 m), the '30 m neighborhood-scale' claim should be tempered. A stronger check is to compare the Landsat LST field against an independent ~70 m ECOSTRESS or airborne thermal observation over the same city and date: if the 30 m Landsat field cannot recover fine-scale thermal variance resolved by the independent sensor, the benchmark target is effectively 100 m.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that HeatCast provides neighborhood-scale (30 m) monthly LST forecasting. The target channel is the USGS Collection 2 single-channel ST product, which the paper itself states is derived from the 100 m native TIRS band and only distributed on the 30 m optical grid (Section 3.1, Table 2, Section 9). The auxiliary channels are truly 30 m, and the LST array is on a 30 m grid, but the thermal information content is resampled from 100 m. If the effective thermal support is 100 m, then HeatCast standardizes forecasting of a 30 m-aligned regridded product rather than 30 m thermal measurements. The advertised advantage over kilometer-scale products, resolving block-scale thermal gradients, is then reduced, because structure below roughly 100 m in the LST field is interpolation artifact rather than measurement. This does not invalidate the benchmark as a reproducible forecasting task, but it weakens the headline 'neighborhood-scale' characterization and the interpretation of LST-only versus auxiliary-channel results, since the target itself is partly derived from coarser thermal observations. The paper acknowledges this only in Section 9 and a table note, while the abstract and title foreground '30 m.'","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"HeatCast introduces a Landsat-based benchmark for monthly land surface temperature forecasting at a nominal 30 m grid resolution across 124 U.S. cities, covering mid-2013 through June 2025. The dataset provides nine aligned channels, Local Climate Zone labels, quality masks, fixed train/validation/test splits, and an evaluation harness reporting LCZ-stratified RMSE in Kelvin. The paper evaluates a CNN+LSTM and Earthformer on a 12-month-to-next-month task, reporting 10.42 K versus 7.74 K aggregate RMSE, and an auxiliary-only Earthformer variant at 7.72 K. The data, code, configurations, and checkpoints are released under an MIT license.","tokens_in":13253,"tokens_out":6855,"duration_ms":67796,"significance":"The contribution is potentially valuable: HeatCast appears to be the first open, reproducible benchmark for urban LST forecasting at a 30 m grid, with a clear task definition, frozen splits, explicit NoData masks, and released artifacts. Its geographic and temporal scope (124 cities, 12 years) is much larger than typical one-to-three-city LST studies, and its LCZ-stratified evaluation is a sensible response to class imbalance. The paper also benefits from transparent limitation statements in Section 9. However, the significance is conditional on a central qualification: the LST target is derived from a 100 m native thermal band and only gridded at 30 m, so the '30 m neighborhood-scale' claim needs to be reframed or empirically supported. The benchmark remains useful as a standardized forecasting task even after this reframing.","major_comments":[{"comment":"The headline characterization of HeatCast as a 30 m neighborhood-scale benchmark is not supported by the paper's own source-product description. Table 2 states that the Landsat TIRS surface temperature product has '100 m→30 m' resolution, and Section 3.1 explains that USGS derives ST from the 100 m TIRS band and distributes it on the 30 m optical grid. Since the target channel carries thermal information at roughly 100 m native support, the 30 m grid contains resampled or interpolated thermal detail rather than 30 m measurements. The abstract, title, and Figures 1 and 5 foreground '30 m' and 'block-scale thermal gradients,' but the caveat appears only in Section 9. This is load-bearing for the central claim. I recommend either qualifying all resolution claims, for example by saying '30 m grid, 100 m native thermal support,' or adding an analysis that quantifies the impact, such as comparing forecasts and RMSE against an evaluation aggregated to 100 m or against a product with genuine 30 m thermal support. The benchmark itself can remain as released, but the characterization needs revision.","section":"3.1, Table 2, Section 9"},{"comment":"The feature-set ablation is presented as a headline result ('Earthformer achieves its lowest RMSE with auxiliary inputs', 7.72 K versus 7.74 K), yet Section 9 states that each baseline uses one training run and that the 0.02 K difference is not a stable ranking. The paper therefore both relies on and disclaims this difference. Because the feature-set ablation is one of the stated contributions, Table 8 and the corresponding text should either provide multiple-seed means and standard deviations or explicitly demote the 7.72 K versus 7.74 K comparison to an observation without a 'best' designation. Without error bars, the claim that non-LST channels are the strongest input is not supported at the reported precision.","section":"6.2, Table 8, Section 9"},{"comment":"The resolution issue also affects the interpretation of the ablation. Because the eight auxiliary channels are natively 30 m while the LST target is natively 100 m, the auxiliary-only Earthformer may be learning to predict the 30 m spatial pattern of the regridded target rather than thermal dynamics; the reported 0.43 K improvement over LST history could reflect the spatial resolution mismatch. At minimum, the paper should compare against a spatial persistence or monthly climatology baseline (which Section 7 lists as future work) and discuss whether the auxiliary advantage would survive an evaluation on 100 m aggregates. This comparison is needed before the auxiliary-only result is presented as evidence about input informativeness.","section":"6.2, Section 7, Section 9"}],"minor_comments":[{"comment":"The monthly lowest-cloud-scene selection does not account for differences in overpass time between Landsat 8 and 9 or across months, which can add non-climatic variance to the target; a sentence acknowledging this and any normalization would help.","section":"3.1, 3.5"},{"comment":"Please clarify how LCZ-stratified RMSE is aggregated: whether pixels are grouped by their own LCZ label before averaging, and whether tile-level weighting applies within each stratum.","section":"4.3"},{"comment":"The LST minimum of -123°C is listed among observed ranges; although the text says residual outliers are masked before evaluation, the table may confuse readers about the valid range, so consider reporting QC-passed ranges instead.","section":"Table 3"},{"comment":"The caption asserts 10–15 K variation within a single neighborhood; if the LST product is resampled from 100 m native support, this statement should be tied to the actual effective resolution or softened.","section":"Figure 1"},{"comment":"The quickstart uses test_years=[2024, 2025], but the test set ends in June 2025; a comment noting the six-month cutoff would avoid ambiguity.","section":"Listing 1"}],"recommendation":"major_revision","confidential_remarks":"This is a solid benchmark contribution with released artifacts and a sensible evaluation protocol, but the 100 m native thermal support issue is central to the '30 m neighborhood-scale' framing and must be addressed before acceptance. The authors' own Section 9 shows awareness, so the fix should be manageable. The single-run issue is also important because the paper makes a small-difference ablation claim. I do not see circularity or data-integrity problems. One editorial observation: the paper cites two methodological precedents that share authors with this manuscript; this is reasonable, but the wording 'Following Stewart et al. and Corley et al.' should be checked to avoid any appearance of self-promotion."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. First, HeatCast is a real service to the field: a 124-city, monthly, 30m-grid LST forecasting benchmark with released data, code, weights, fixed split, and an evaluator, which is genuinely missing from the literature. Second, the title's 'neighborhood-scale' claim is softer than it looks: the LST target is the USGS Collection 2 single-channel product, derived from 100m TIRS and gridded to 30m. The paper admits this in Section 9 but leads with 30m. That does not kill the benchmark, but it changes what the 30m claim should mean.\n\nWhat is good: the gap documentation in Table 1 is honest; prior studies are one-to-three cities, coarse, or unreleased. The protocol is clear — 12-month context, temporal split, LCZ stratification, Kelvin metrics, fixed manifests. Releasing checkpoints and a smoke test is exactly the right kind of reproducibility. The limitation section is unusually candid: it flags the 100m native support, single-run baselines, and the instability of the 0.02K gap between the all-channel and auxiliary-only Earthformer variants.\n\nSoft spots: the baseline results rest on one run per configuration with no error bars. Even the 0.43K auxiliary-vs-LST difference needs variance before it supports a claim. The paper mentions persistence and monthly climatology as future work, but those are reference points a benchmark should ship with; without them, 7.74K is hard to interpret. The CNN+LSTM wins on compact urban (8.62 vs 12.68) but that class is 0.76% of pixels, so aggregate numbers hide the reversal. Also read the auxiliary-only result carefully: those channels come from the same Landsat scenes, so they are not independent of the retrieval. None of this is disqualifying.\n\nThe 30m issue is the one that needs fixing before publication: either temper the wording or show the product supports block-scale structure beyond 100m. The stress-test note is right that sub-100m structure is interpolation artifact. The benchmark itself remains usable and valuable; the forecasting task is well defined on the released product.\n\nWho this is for: urban climate ML and remote sensing researchers who need a shared evaluation standard. It deserves serious peer review; a good reviewer will ask for multi-seed runs, reference baselines, and a more careful resolution claim. Recommendation: send it out, with those questions on the checklist.","headline":"A genuinely useful, released LST forecasting benchmark whose '30m' claim overstates thermal resolution and whose baselines need error bars.","tokens_in":13808,"tokens_out":3442,"would_cite":true,"duration_ms":28860,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces HeatCast, an open, reproducible benchmark for monthly land-surface temperature forecasting at 30 m resolution across 124 U.S. cities, and reports that Earthformer achieves 7.74 K RMSE versus 10.42 K for a CNN+LSTM.","keywords":["land surface temperature","urban heat island","benchmark dataset","spatiotemporal forecasting","remote sensing","climate resilience","Earthformer","local climate zones"],"falsifier":"Run the released reference evaluator on a persistence baseline that predicts the last observed monthly LST, and on a monthly climatology baseline; if either reaches 7.7 K or better, the headline baselines do not demonstrate skill beyond the seasonal cycle, and if the 7.74 K and 7.72 K numbers cannot be regenerated from the released manifests and checkpoints, the benchmark's reproducibility claim fails.","tokens_in":12816,"feed_emoji":"🌡️","tokens_out":9915,"duration_ms":85386,"temperature":0.7,"pith_summary":"This paper introduces HeatCast, an open benchmark for forecasting monthly land-surface temperature at 30 m resolution across 124 U.S. cities from 2013 through June 2025. It defines a fixed task: given twelve months of nine aligned channels (LST, elevation, RGB, three spectral indices, and albedo) on 128x128 tiles, predict the next month's LST, with a temporal split, Local Climate Zone stratified metrics, and a reference evaluator. On this protocol, the paper reports that Earthformer reaches 7.74 K aggregate RMSE versus 10.42 K for a CNN+LSTM, and that an Earthformer using only the eight non-LST channels reaches 7.72 K. If the benchmark is adopted, urban-heat model comparisons become reproducible and comparable across studies. The data, code, and weights are released under MIT.","feed_headline":"Earthformer beats CNN+LSTM on 124-city heat benchmark","feed_subtitle":"Open Landsat dataset sets a common test for next-month surface heat; best model scores 7.74 K RMSE.","key_machinery":"The carrying object is the benchmark protocol itself: monthly 30 m tiles organized as 128x128 patches with nine aligned channels, a fixed one-month-ahead task conditioned on twelve monthly observations, a temporal split (train 2013-2021, validation 2022-2023, test January 2024-June 2025), and an evaluation script that reports mean tile-level RMSE in Kelvin, stratified by Local Climate Zone super-clusters. Earthformer, a space-time transformer using cuboid attention, is the reference model that produces the headline numbers; the released fixed sample manifests and reference evaluator are what let the results be regenerated and compared.","core_discovery":"The paper claims that urban LST forecasting lacks a shared benchmark and that HeatCast fills this gap. The central discovery, stated on the benchmark's own terms, is that Earthformer reaches 7.74 K aggregate test RMSE on the fixed next-month task, compared with 10.42 K for the CNN+LSTM, and that an Earthformer trained on only the eight non-LST channels reaches 7.72 K, against 8.15 K from historical LST alone and 8.68 K from RGB alone. It also reports that the CNN+LSTM performs better on compact urban pixels (8.62 K vs 12.68 K), while Earthformer dominates the open urban, other urban, and natural categories. The claim is that these numbers are reproducible from the released manifests, configurations, and checkpoints, giving the field a common evaluation standard for 30 m urban heat forecasting.","pith_inferences":["Inference: A persistence baseline that predicts last month's LST, or a monthly climatology, would sharpen the claim: if it lands near 7.7 K, the benchmark's baselines would show skill mostly from the seasonal cycle rather than from learned dynamics.","Inference: Because the Landsat thermal band is natively 100 m, users should read '30 m' as grid spacing, not thermal resolution; a planning study that treats 30 m LST hotspots as thermal measurements would over-interpret the product.","Inference: The released artifacts make a natural held-out-city experiment possible, training on a subset of cities and testing on the rest, which the current temporal split does not address."],"forward_implications":["Any future forecasting model can be scored against the same 124 cities, months, and pixels as the published baselines, making cross-study comparisons possible.","The auxiliary-only result (7.72 K) implies that elevation, spectral indices, and albedo carry strong predictive signal for next-month heat, so feature engineering is as important as architecture choice.","LCZ-stratified metrics show compact urban pixels remain the least accurate class for Earthformer (12.68 K vs 8.62 K for CNN+LSTM), pointing to class imbalance as a concrete target.","Because the split is temporal, the benchmark measures how well models forecast future months in known cities; it does not by itself measure transfer to unseen cities."],"supporting_citations":[{"why":"Supplies the USGS Collection 2 single-channel surface-temperature product that HeatCast uses as its target and forecast channel.","marker":"[27]"},{"why":"Defines Earthformer, the spatio-temporal transformer whose cuboid attention yields the benchmark's best reported RMSE.","marker":"[10]"},{"why":"Supplies the LSTM component of the CNN+LSTM reference baseline.","marker":"[17]"},{"why":"Supplies the CNN+LSTM-style architecture used as the weaker reference baseline.","marker":"[54]"},{"why":"Provides the Landsat mosaic, reprojection, and pretraining template the construction pipeline follows.","marker":"[39]"},{"why":"Provides the Landsat benchmark pipeline cited alongside [39] for data packaging choices.","marker":"[7]"},{"why":"Defines the Local Climate Zone classes used for stratified evaluation.","marker":"[40]"},{"why":"Supplies the CONUS LCZ labels that assign every pixel to an LCZ class for stratified metrics.","marker":"[8]"}],"fun_headline_variants":["Earthformer beats CNN+LSTM on 124-city heat forecast benchmark","HeatCast: open 30m LST benchmark for 124 U.S. cities","Non-LST channels alone beat LST history in 124-city heat forecast","HeatCast: 30m heat forecast test for 124 cities, Earthformer top","124-city LST benchmark: Earthformer hits 7.74K RMSE, open data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's 'neighborhood-scale' claim assumes the Landsat Collection 2 LST product, which is derived from the 100 m native TIRS band and only distributed on the 30 m optical grid, actually carries thermal information at 30 m; if the true thermal support is 100 m, the benchmark standardizes a regridded product rather than 30 m thermal measurements.","fun_headline_variants_meta":{"raw":{"variants":["Earthformer beats CNN+LSTM on 124-city heat forecast benchmark","HeatCast: open 30m LST benchmark for 124 U.S. cities","Non-LST channels alone beat LST history in 124-city heat forecast","HeatCast: 30m heat forecast test for 124 cities, Earthformer top","124-city LST benchmark: Earthformer hits 7.74K RMSE, open data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001341,"raw_usage":{"total_tokens":5451,"prompt_tokens":946,"completion_tokens":4505,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":562,"completion_tokens_details":{"reasoning_tokens":4401}},"tokens_in":562,"tokens_out":4505,"duration_ms":29329,"temperature":1.0,"reasoning_tokens":4401,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T00:27:54.483994+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the released reference evaluator on a persistence baseline that predicts the last observed monthly LST, and on a monthly climatology baseline; if either reaches 7.7 K or better, the headline baselines do not demonstrate skill beyond the seasonal cycle, and if the 7.74 K and 7.72 K numbers cannot be regenerated from the released manifests and checkpoints, the benchmark's reproducibility claim fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the CNN+LSTM-style architecture used as the weaker reference baseline."},{"cited_title":"Landsat-Bench: Datasets and Benchmarks for Landsat Foundation Models","cited_arxiv_id":"2506.08780","evidence_quote":"Provides the Landsat benchmark pipeline cited alongside [39] for data packaging choices."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the CONUS LCZ labels that assign every pixel to an LCZ class for stratified metrics."}],"review_version":1}