{"id":"faa7fb6b-80b9-4555-820c-952b57241fad","arxiv_id":"2501.14130","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":8,"one_line_summary":"A seasonal-resolution last-millennium reanalysis, built by assimilating climate proxies into an online ocean-atmosphere-sea-ice linear inverse model, matches modern surface-temperature observations as well as annual products while using far fewer proxies.","lead":"Researchers built a new seasonal-resolution reconstruction of the last 1,000 years of climate by feeding tree rings, corals, ice cores, and other proxies into a coupled ocean-atmosphere-sea-ice forecast model. It is the first such seasonal reanalysis, and it matches modern observations well even in winter, when proxies are scarce.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The pre-instrumental skill claim rests on a stationarity assumption that the paper's holdout verification cannot actually test, because withheld proxies are forward-modeled with the same modern-calibrated PSMs.","rationale":"The reader's weakest assumption—stationarity of the modern-calibrated PSMs—is indeed the load-bearing premise for the pre-instrumental skill claim. My read sharpens this into a concrete circularity: the Section 3.2 bootstrap withholds proxies from assimilation but still evaluates them with the same PSMs used for assimilation, so a multiplicative nonstationarity can cancel out of the verification and leave the holdout correlations artificially high. The reader's conditional verdict already captures the need for caution, so I do not propose changing the verdict; however, the condition should explicitly require a stationarity test that does not share the PSM calibration, such as the proposed OSSE or a split-sample calibration on early versus late instrumental periods. The methodological contributions—seasonal online DA, the seasonal update strategy, and inclusion of sea ice in the LIM—are supported by internal tests and are not called into question by this concern.","tokens_in":22179,"tokens_out":8234,"duration_ms":81747,"concrete_test":"Run an OSSE in which the CCSM4 or MPI-ESM-R last-millennium simulation is treated as truth: prescribe a true proxy forward model whose temperature sensitivity is halved before 1850 CE, generate pseudo-proxies from that true model, then run the full LMR Seasonal DA (Section 2.3) using the modern-calibrated PSM and apply the Section 3.2 20%-holdout bootstrap exactly as in the paper. If the pre-1850 holdout correlations remain comparable to Fig. 11 while the reconstructed temperature field is biased by several tenths of a degree, the verification is blind to PSM nonstationarity and the pre-instrumental skill claim must be regarded as unverified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of robust skill before 1850 depends on the assumption that the linear PSMs calibrated on 20th-century GISTEMP/ERSST data (Section 2.2, Eq. 7) remain valid over 850–1850 CE. The only pre-instrumental verification (Section 3.2) does not actually test this assumption. In the bootstrap test, 20% of proxies are withheld from assimilation, but the withheld proxies are then evaluated by forward-modeling them from the reconstruction using the same modern-calibrated PSM (Eq. 7). A multiplicative drift in the true proxy–climate relationship is invisible to this procedure: if the modern calibration overestimates the true sensitivity by a factor c (a_cal = c·a_true), and assimilated proxies share this bias, the EnKF analysis underestimates the true temperature by roughly 1/c (Eq. 11 in the low-R limit), and the forward model maps that underestimated temperature back to y_pred = a_cal·T_a ≈ a_true·T_true. The holdout correlation, and even CE, can remain near 1 while the reconstructed temperature is substantially biased. Thus the 'robust PSM relationship' conclusion in Section 3.2 is circular with respect to stationarity. The instrumental verification in Section 3.1 also overlaps the PSM calibration period, so it cannot independently establish stationarity either. The headline comparison to other products is additionally supported by global-mean correlation differences of only 0.01–0.07 (Fig. 3) without uncertainty estimates, but the deeper issue is that the pre-instrumental claim lacks a test capable of detecting the assumed failure mode.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces LMR Seasonal, a seasonal-resolution 'online' data assimilation reconstruction of the last millennium. A linear inverse model (LIM) trained on CMIP5 last-millennium simulations provides coupled forecasts of surface temperature, sea surface temperature, upper-ocean heat content, and Northern Hemisphere sea-ice concentration/thickness; an ensemble square-root Kalman filter assimilates PAGES2k v2 proxies with season-specific update timing. The paper claims that this reconstruction achieves the highest correlation skill against instrumental products while using fewer proxies, that verification against held-out proxies demonstrates robust pre-instrumental skill, and that the method captures ENSO evolution and the MCA-LIA contrast better than existing off-line reconstructions.","tokens_in":114,"tokens_out":3263,"duration_ms":85885,"significance":"If the central claims hold, LMR Seasonal would be a significant advance: it is the first seasonal-resolution, coupled ocean-atmosphere-sea-ice reanalysis of the last millennium, and its online update strategy is a clear methodological improvement over static-prior off-line DA for seasonal-to-interannual memory. The manuscript has concrete strengths: it compares against three established paleo-DA products, verifies variables not directly assimilated (OHC300, sea-ice concentration) against independent observational datasets, and demonstrates sensible ENSO composites. However, the headline comparative claim rests on global-mean correlation differences of 0.01-0.07 without uncertainty estimates, and the pre-instrumental verification using withheld proxies is circular with respect to the stationarity of the proxy system models. These issues make the paper a strong candidate for major revision rather than acceptance in its current form.","major_comments":[{"comment":"The abstract and Section 3.1 claim that LMR Seasonal 'achieves the highest correlation skill' compared to other paleo-DA products. The evidence is global-mean correlation differences of 0.01 (vs. LMRv2), 0.04 (vs. PHYDA), and 0.07 (vs. LMR Online) for annual-mean temperature (Fig. 3), and 0.06 (DJF) and 0.01 (JJA) vs. PHYDA (Fig. 4). No confidence intervals, significance tests, or bootstrap estimates are provided for these differences. Given that the verification period 1880-2000 substantially overlaps the PSM calibration period (GISTEMP/ERSST, also 20th-century), these small differences may reflect sampling variability or calibration artifacts. The authors should provide uncertainty bounds on the correlation differences and, where possible, a verification period that does not overlap calibration.","section":"§3.1, Figs. 3-4"},{"comment":"The claim that 'verification against independent proxy records shows that reconstruction skill is robust throughout the last millennium' is not supported by the bootstrap experiment as designed. The 20% withheld proxies are evaluated by forward-modeling them from the reconstructed climate state using the same modern-calibrated PSM (Eq. 7). As the authors correctly note in Section 2.2, the PSMs are linear and calibrated on GISTEMP/ERSST. A multiplicative drift in the true proxy-climate sensitivity is invisible to this procedure: if the calibrated sensitivity is c times the true sensitivity and the assimilated proxies share this bias, the EnKF analysis (Eq. 11 in the low-observation-error limit) underestimates the temperature by roughly 1/c, and the forward model maps this underestimated temperature back to approximately the correct proxy value. Thus high holdout correlation and even high CE do not establish stationarity. The conclusion in Section 3.2 that 'the distribution of correlation values... is very similar, suggesting a robust PSM relationship' is therefore circular. A valid test would require, for example, calibration on one time period and validation on a non-overlapping period, pseudo-proxy experiments in which PSM parameters are perturbed or made time-varying, or comparison against independent non-PSM-based reconstructions.","section":"§3.2, Eq. (7), Fig. 11"},{"comment":"The paper attributes the MCA-LIA difference (0.15°C) to the seasonal-update strategy based on a single experiment in which seasonal proxies update only the annual mean (0.10°C). No uncertainty estimate is given for either value, yet the conclusion 'our seasonal-update strategy appears to be essential to reconstructing the magnitude of the MCA-LIA difference' depends on the 0.05°C difference being statistically distinguishable from ensemble spread and from natural variability. The authors should report confidence intervals on the MCA-LIA difference for both the seasonal and annual-update experiments and test whether the difference between the two experimental outcomes is significant.","section":"§4, Fig. 13"}],"minor_comments":[{"comment":"The word 'shwon' in the caption is a typo for 'shown'.","section":"Fig. 3 caption"},{"comment":"The text states that independent-proxy verification results are 'compare Fig. 11 with Supplementary Fig. S8', but Supplementary Fig. S8 is the Niño3.4 PHYDA comparison, not the proxy-verification figure. The intended cross-reference is likely Supplementary Fig. S9.","section":"§2.2, references to supplementary figures"},{"comment":"The manuscript inconsistently uses 'El Niño', 'Nino3.4', and 'Niño3.4' (e.g., Figs. 8-10 and text). Please standardize the notation.","section":"Throughout"},{"comment":"The EnSRF update equations (Eqs. 10-12) are presented for the serial observation update but the text does not specify how observations are ordered or whether the ensemble is re-orthogonalized after each observation; a sentence clarifying the implementation would help reproducibility.","section":"§2.3"},{"comment":"The data availability statement says the reconstruction data and code 'will be released to the public once this manuscript has been accepted.' For a journal that values reproducibility, releasing code and data at the time of submission (or at least in a preprint repository) would strengthen the paper and allow readers to verify the claims.","section":"Data availability"}],"recommendation":"major_revision","confidential_remarks":"The paper's core methodological contribution is real and likely publishable, but the current framing overstates the comparative skill and the pre-instrumental verification. I would advise the editor to request a revised version that adds uncertainty quantification to the correlation differences and replaces the circular proxy-holdout test with a genuinely independent stationarity test. If the authors can provide those, the paper could be suitable for acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe short version: this is a genuine methodological step forward—the first seasonal-resolution online data-assimilation reanalysis of the last millennium with coupled ocean–atmosphere–sea-ice state variables—but the headline skill claim is thinner than the abstract implies. The pre-instrumental verification cannot actually test the stationarity assumption it depends on.\n\nWhat is new: the season-to-season update strategy (updating specific seasons rather than folding seasonal proxies into an annual mean) and the inclusion of sea ice in the LIM. The implementation is careful: EOF truncation, PSM calibration in truncated space, 800-member EnSRF, and verification against OHC300 and sea-ice concentration, which are not directly assimilated. The citation pattern is appropriate—the paper builds on the LMR program and compares against the right products.\n\nSoft spots, in proportion. The instrumental verification period overlaps the PSM calibration period (both 1880–2000), so the temperature/SST skill partly reflects calibration fit. The holdout proxy verification in Section 3.2 is not fully independent either: withheld proxies are forward-modeled using the same modern-calibrated PSMs. If the true proxy sensitivity drifts by a factor, the analysis is biased in the opposite direction and the forward model maps it back to near-perfect correlation. So the 'robust PSM relationship' conclusion is circular with respect to stationarity. The headline comparison to other DA products rests on global-mean correlation differences of 0.01–0.07 with no uncertainty estimates; that is a weak basis for 'highest skill.' And calling OHC300 and SIC correlations around 0.2 'high' is an overstatement.\n\nNone of this invalidates the central contribution. The seasonal update strategy is a real advance, and the dataset will likely be a community resource once public. But the authors should run a pseudoproxy experiment that varies PSM sensitivity to test the exact failure mode the holdout test misses, and they should report confidence intervals on the skill differences. The language in the abstract and verification sections should be tempered.\n\nRecommendation: worth a serious referee. I'd send it out and ask for those revisions. I'd want code and data public before acceptance.\n\nThis is a good reading-group paper for the methodology and the stationarity discussion.","headline":"First seasonal last-millennium reanalysis with online DA and sea ice, but the pre-instrumental skill claim is thinner than it looks; still a solid contribution worth serious peer review.","tokens_in":23096,"tokens_out":3372,"would_cite":true,"duration_ms":29082,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A season-by-season data-assimilation scheme reconstructs last-millennium climate with higher verification skill than prior paleo-products while using roughly a quarter of the proxies.","keywords":["paleoclimate data assimilation","seasonal resolution","last millennium","linear inverse model","sea ice reconstruction","El Niño diversity","proxy system model","online data assimilation"],"falsifier":"A decisive test would be to run the bootstrap verification while withholding all proxies earlier than 1200 CE; if pre-1200 correlations for non-assimilated proxies collapse toward zero, the claim of consistent skill throughout the last millennium fails. A second, sharper test would compare the DJF temperature field over 1400–1700 against an independent winter-sensitive reconstruction excluded from the proxy database, requiring agreement within the stated ensemble spread.","tokens_in":21941,"feed_emoji":"🌍","tokens_out":9463,"duration_ms":76495,"temperature":0.7,"pith_summary":"This paper introduces a data-assimilation reconstruction of the last millennium that resolves the seasonal cycle (March–May, June–August, September–November, December–February) for coupled atmosphere, ocean, and sea-ice fields. Its central claim is that an \"online\" assimilation scheme, in which a linear inverse model forecasts the climate state from one season to the next, uses proxy information more efficiently than earlier offline reconstructions: it attains the highest correlation skill in surface-temperature verification against instrumental records while assimilating far fewer proxies than competing products do, especially in boreal winter when proxy coverage is sparse. The paper further claims that reconstructed upper-ocean heat content, Arctic sea-ice concentration, and the seasonal evolution of El Niño events verify well against independent instrumental and satellite data, and that the reconstruction shows consistent skill in pre-instrumental epochs when checked against withheld proxies. If these claims hold, the result matters because it offers a seasonal-resolution, physically coupled view of the last millennium, and because the season-to-season update mechanism explains why summer-biased tree-ring proxies no longer dilute winter and annual reconstructions.","feed_headline":"First seasonal reanalysis of the last millennium beats prior skill","feed_subtitle":"Online data assimilation with a linear inverse model carries summer proxy memory into winter using only 521 proxies.","key_machinery":"The engine is a season-to-season Linear Inverse Model (LIM) whose state vector holds principal components of 2-meter temperature, sea surface temperature, upper-300-meter ocean heat content, and Northern-Hemisphere sea-ice concentration and thickness, trained on last-millennium simulations of two CMIP5 climate models. The LIM supplies the forecast prior for an ensemble square-root Kalman filter with 800 members and no localization; linear proxy system models, calibrated on modern instrumental temperatures in the EOF-truncated space with a diagonal error covariance, map proxies into that prior. The defining mechanism is the seasonal update strategy: a proxy with, say, June–August seasonality updates only the June–August ensemble, while annual-mean proxies are assimilated only once an annual mean can be formed from the seasonal forecasts, so the LIM's memory carries information from proxy-rich summer into proxy-poor winter.","core_discovery":"The core discovery is that cycling a linear inverse model forward season to season, and assimilating each proxy in the season it actually represents, produces a last-millennium reconstruction whose verified skill exceeds that of earlier paleo-data-assimilation products even though roughly one quarter as many proxies are used. Global-mean correlation against instrumental surface temperature is about 0.54 for annual means (versus 0.47–0.53 for three earlier reconstructions), and the largest edge comes in boreal winter (about 0.43 versus 0.37), a season with few tree-ring records. Reconstructed upper-ocean heat content correlates near 0.2 with an objective ocean analysis, Arctic sea-ice concentration correlates 0.11–0.21 with satellite data, and the Niño3.4 index correlates about 0.78 with a sea-surface-temperature analysis (coefficient of efficiency around 0.55). The paper attributes the winter advantage to the online scheme: summer proxy information persists through the seasonal forecast and updates the prior for data-poor seasons. It also shows that when seasonal proxies are restricted to updating annual means only, the reconstructed global-mean temperature difference between the Medieval Climate Anomaly and the Little Ice Age shrinks from 0.15 °C to 0.10 °C, arguing that the seasonal update is essential for recovering multicentennial variability.","pith_inferences":["An extension the authors do not pursue is adding atmospheric circulation or precipitation principal components to the LIM state; the paper's exclusion of high-frequency variables to protect forecast skill suggests this would require a careful balance between state dimension and memory.","Because the LIM is trained on two CMIP5 models, a natural test is to train the same system on a CMIP6 last-millennium simulation; the paper's claim that CMIP5-to-CMIP6 seasonal variability statistics are similar predicts verification skill would change little.","The winter-skill advantage is attributed to dynamical memory in the trained LIM, which implies a testable prediction: a model with weaker seasonal memory would show a smaller winter edge, so training the LIM on a model with poorly simulated ENSO persistence should degrade DJF skill more than JJA skill."],"forward_implications":["If the skill claim is correct, future paleoclimate reanalyses can achieve equal or better verification with substantially smaller proxy networks, cutting the data burden for seasonal reconstructions.","Seasonal fields for ocean heat content and Arctic sea ice become available across the whole millennium, enabling direct study of seasonal sea-ice evolution and its links to temperature and orbital forcing.","The millennium-length ENSO reconstruction, verified against four El Niño onset classes in the twentieth century, provides a large sample for studying changes in El Niño diversity and seasonal evolution.","Because the season-to-season update yields a larger Medieval Climate Anomaly-to-Little Ice Age temperature difference than annual-mean updating, earlier annual reconstructions may have underestimated multicentennial variability by letting summer-biased proxies dilute winter and annual signals."],"supporting_citations":[{"why":"Establishes the last-millennium reanalysis framework and the practice of reconstructing variables not directly observed, the methodological foundation this paper modifies toward seasonal online updating.","marker":"Hakim et al. (2016)"},{"why":"Demonstrates online data assimilation with a linear inverse model for the last millennium, the direct precursor whose annual-mean approach this paper extends to four seasons.","marker":"Perkins and Hakim (2021)"},{"why":"Provides the objective proxy-seasonality definition and one of the comparison products; the paper adopts its seasonality approach.","marker":"Tardif et al. (2019)"},{"why":"Supplies the proxy database that provides both the assimilated observations and the withheld-proxy verification.","marker":"PAGES2k Consortium and others (2017)"},{"why":"Introduces linear inverse modeling for seasonal climate prediction, the basis of the forecast model used here.","marker":"Penland and Magorian (1993)"},{"why":"Provides the ensemble Kalman filter theory behind the analysis update.","marker":"Evensen (2003)"},{"why":"Supplies the four-class El Niño onset taxonomy used to verify the seasonal evolution of reconstructed ENSO events.","marker":"Wang et al. (2019)"},{"why":"Gives the orbital-forcing explanation for the seasonal asymmetry in last-millennium temperature trends, which the reconstruction is compared against.","marker":"Lücke et al. (2021)"}],"fun_headline_variants":["Seasonal reanalysis of last millennium tops prior skill","Online DA yields best last-millennium climate skill","Seasonal proxy memory boosts winter reconstructions","Coupling seasonal DA beats earlier paleo-reconstructions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reconstruction's pre-instrumental skill rests on the assumption that the linear proxy system models calibrated against twentieth-century instrumental temperatures stay valid with the same error statistics throughout 850–1850, so that the mapping from climate state to proxy and the assumed error covariances do not drift over time.","fun_headline_variants_meta":{"raw":{"variants":["Seasonal reanalysis of last millennium tops prior skill","Online DA yields best last-millennium climate skill","Seasonal proxy memory boosts winter reconstructions","Coupling seasonal DA beats earlier paleo-reconstructions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000633,"raw_usage":{"total_tokens":2928,"prompt_tokens":960,"completion_tokens":1968,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":576,"completion_tokens_details":{"reasoning_tokens":1906}},"tokens_in":576,"tokens_out":1968,"duration_ms":12316,"temperature":1.0,"reasoning_tokens":1906,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T15:22:09.594967+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test would be to run the bootstrap verification while withholding all proxies earlier than 1200 CE; if pre-1200 correlations for non-assimilated proxies collapse toward zero, the claim of consistent skill throughout the last millennium fails. A second, sharper test would compare the DJF temperature field over 1400–1700 against an independent winter-sensitive reconstruction excluded from the proxy database, requiring agreement within the stated ensemble spread.","supporting_citations":[{"cited_title":"Magorian, 1993: Prediction of Niño 3 sea surface temperatures using linear inverse modeling","cited_arxiv_id":null,"evidence_quote":"Introduces linear inverse modeling for seasonal climate prediction, the basis of the forecast model used here."},{"cited_title":"Luo, Y.-M","cited_arxiv_id":null,"evidence_quote":"Supplies the four-class El Niño onset taxonomy used to verify the seasonal evolution of reconstructed ENSO events."}],"review_version":1}