{"id":"e450b134-e0c1-452a-a7a2-b0b51ff007e8","arxiv_id":"2510.00809","paper_version":3,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Fine-tuning TimesFM sequentially on new synthetic time-series data causes measurable forgetting of earlier tasks, with higher learning rates producing stronger forgetting.","lead":"This paper tests whether a time-series foundation model (TimesFM) forgets previously learned forecasting tasks when fine-tuned on new synthetic datasets. It reports that fine-tuning on a second dataset degrades performance on the first, with the amount of forgetting depending on learning rate and number of epochs.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract claims a comparative study with TimesFM-2.0, Chronos-2, SamFormer, DER, and real-world benchmarks; the body only reports two synthetic TimesFM task pairs, so the central claim is unsupported.","rationale":"The reader's verdict is REJECT, and my analysis supports that. The reader's weakest_assumption focuses on external validity of the synthetic datasets, but the more load-bearing issue is the abstract–body gap: the paper claims a systematic comparative study of multiple models, DER, and real-world benchmarks, none of which appear in the body. This is not an internal inconsistency in a narrow argument; it is a direct failure of the central claim to be supported by the presented evidence. Even if the synthetic datasets were perfect representatives of real-world shifts, the promised comparative results are absent. The reader's rationale does mention this overclaim, but the weakest_assumption field does not. Thus agreement is partial: we share the overall rejection, but I identify the abstract–body mismatch as the primary problem rather than the external validity of the synthetic data. The body alone provides a suggestive preliminary result—TimesFM forgets a synthetic task after fine-tuning—but without multiple seeds, error bars, code, or a comparison baseline, even that narrow claim is weakly supported. Therefore, the rejection stands unchanged.","tokens_in":5468,"tokens_out":3370,"duration_ms":26302,"concrete_test":"Perform a systematic audit: for each claim in the abstract (e.g., 'larger models exhibit greater inherent robustness', 'DER provides disproportionate gains to smaller models', 'match TSFM performance by the end of the continual learning sequence'), find the corresponding experiment in Sections 3–4 or tables. If 'TimesFM-2.0', 'Chronos-2', 'SamFormer', 'DER', or 'energy benchmark' do not appear as experimental results, the central claim is unsubstantiated. Additionally, rerun the D1→D2 experiment with 5 random seeds to check whether BWT = +1.45 is reproducible; if the forgetting magnitude varies widely or vanishes across seeds, even the narrow body claim lacks statistical support.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract promises systematic evaluation of TSFMs (TimesFM-2.0, Chronos-2) versus a specialized SamFormer on synthetic and real-world energy benchmarks, including DER mitigation and conclusions about model-size economics. The body, however, contains only two-stage continual fine-tuning of TimesFM on two pairs of synthetic multi-sinusoidal datasets (D1→D2, D3→D4). There are no experiments with Chronos-2, TimesFM-2.0, SamFormer, DER, or real-world data. No table or figure reports these claimed comparisons. Consequently, the paper's central claim—that fine-tuning consistently triggers forgetting in TSFMs, that larger models exhibit inherent robustness, and that DER levels the playing field—has no direct evidence in the manuscript. The only supported result is that one TSFM forgets one synthetic task after fine-tuning on another, which is a much narrower claim and is, in itself, not surprising given known stability-plasticity trade-offs. This is not a mere external validity concern about synthetic data; it is a fundamental mismatch between the stated contribution and the evidence presented.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper claims to provide the first systematic study of catastrophic forgetting in Time Series Foundation Models (TSFMs), contrasting TimesFM-2.0 and Chronos-2 with a specialized SamFormer model on synthetic and real-world energy benchmarks, and evaluating experience replay (DER) mitigation and model-size economics. The body, however, contains only a two-stage continual fine-tuning protocol applied to a single TSFM (referred to as TimesFM) on two pairs of synthetic multi-sinusoidal datasets (D1→D2 and D3→D4). The results show that fine-tuning on D2 degrades MAE on D1 from 0.15 to 1.60 (BWT reported as +1.45), while D3→D4 shows a smaller degradation from 0.56 to 0.76. A learning-rate and epoch sweep indicates that high learning rates cause more forgetting, while very low rates limit adaptation. The paper concludes that TSFMs suffer from the stability-plasticity dilemma and that mitigation strategies are needed. No experiments with Chronos-2, SamFormer, DER, or real-world data are presented, and no statistical variability is reported.","tokens_in":5690,"tokens_out":3054,"duration_ms":390470,"significance":"If the central claim were properly supported, the paper would fill a gap in the literature on continual fine-tuning of TSFMs. The synthetic-data design that avoids overlap with pretraining data is a sound principle, and the hyperparameter sweep gives a preliminary view of the stability-plasticity trade-off. However, the significance as stated in the abstract is not achieved by the body. The narrow, single-model, two-dataset-pair result is a plausible empirical observation, but it is not a systematic comparative study, and its generality is undermined by the lack of real-world benchmarks, other models, or mitigation techniques. The absence of error bars or multiple seeds further limits the reliability of the quantitative claims. The paper therefore does not deliver on its stated contribution, and its conclusions about model-size economics and the effectiveness of DER are unsupported.","major_comments":[{"comment":"The abstract promises a systematic evaluation of TSFMs (TimesFM-2.0, Chronos-2) versus SamFormer, on synthetic and real-world energy benchmarks, including DER mitigation and conclusions about model-size economics. The body only reports two two-stage fine-tuning experiments with TimesFM on synthetic multi-sinusoidal datasets. No other model, no DER, and no real-world data appear anywhere in the manuscript. The central comparative claim is therefore unsupported by the evidence presented.","section":"Abstract and Section 4 (Table 1)"},{"comment":"All results are reported as single runs without error bars, number of seeds, or statistical significance tests. For example, the D3→D4 forgetting BWT of +0.20 could be within run-to-run noise. The hyperparameter sweep shows considerable sensitivity to learning rate, and without variance estimates, the apparent trends (e.g., that 10^-5 with 5-10 epochs is the best balance) are not established. This undermines the robustness of the paper's main empirical claims.","section":"Tables 1-3"},{"comment":"The metric labeled BWT is not defined in the paper. In Table 1, BWT for D1 is reported as +1.45, which equals 1.60 − 0.15, the increase in MAE. In the continual learning literature, backward transfer is usually negative when performance degrades (e.g., accuracy on old task after new training minus accuracy before). Using a nonstandard definition without explanation makes the quantitative results difficult to interpret and compare with prior work.","section":"Table 1 and Section 3"},{"comment":"The synthetic datasets are sums of a small number of sine waves with hand-selected periods. The paper states that these 'simulate real-world scenarios' where only partial cycles are observed, but no evidence is given that the distribution shift between D1 and D2 (or D3 and D4) is representative of non-stationarity encountered in practice. The conclusion that the findings apply to 'realistic, non-stationary scenarios' is therefore a leap without external validation.","section":"Section 3 and Appendix A"}],"minor_comments":[{"comment":"The abstract and Section 2 mention TimesFM-2.0, but the body refers only to 'TimesFM' without specifying the version. This creates ambiguity about which model was actually fine-tuned.","section":"General"},{"comment":"Figure 1 is referenced in the text ('as shown in Figure 1') but no figure image is present in the manuscript. Either include the figure or remove the reference.","section":"Section 4"},{"comment":"Typographical error: 'InStage One' should be 'In Stage One.'","section":"Section 3"},{"comment":"The claim that this is 'the first empirical study to demonstrate the stability-plasticity trade-off in univariate time series forecasting with foundation models' is too strong given the existing literature on catastrophic forgetting in large models, and no comparison with prior continual learning studies on TSFMs is provided.","section":"Section 2"},{"comment":"Some references are incomplete or inconsistent (e.g., missing page numbers for some conference papers, while others include arXiv version numbers). This should be cleaned up.","section":"References"}],"recommendation":"reject","confidential_remarks":"The manuscript resembles an early workshop-extended abstract rather than a complete research paper. The abstract does not match the body: the promised comparative study with multiple models, DER, and real-world benchmarks is entirely absent. Even as a limited empirical study, the lack of error bars and the nonstandard BWT metric are concerning. The paper's central claim as stated is unsupported, and the experiments that are present are too thin to justify publication in a serious journal. Unless the authors can add the missing experiments and analysis, rejection seems appropriate."},"author_rebuttal":null,"desk_editor":null,"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that TimesFM, a decoder-only time-series foundation model, suffers catastrophic forgetting when fine-tuned sequentially on synthetic forecasting tasks: after learning dataset D2, its error on dataset D1 rises from 0.15 to","keywords":["catastrophic forgetting","continual learning","time series forecasting","foundation models","stability-plasticity dilemma","backward transfer","fine-tuning","TimesFM"],"falsifier":"Run the same two-stage fine-tuning on two real-world series with a genuine distribution shift, such as energy load from different buildings or seasons, and check whether BWT on the first dataset replicates the synthetic 0.15-to-1.60 jump. If the first-task error barely moves, the strong forgetting result is an artifact of the synthetic periodic design.","tokens_in":5329,"feed_emoji":"📉","tokens_out":3709,"duration_ms":34731,"temperature":0.7,"pith_summary":"The paper asks whether time-series foundation models can be fine-tuned on new forecasting tasks without losing what they learned on earlier ones. Using synthetic multi-sinusoidal datasets and a two-stage continual-learning protocol, it shows that TimesFM adapts to the new task but forgets the old one: MAE on the first dataset rises from 0.15 to 1.60 after fine-tuning on the second (D3 to D4: 0.56 to 0.76). The results depend on hyperparameters: high learning rates and more epochs cause more forgetting, while very low learning rates preserve prior knowledge but impair adaptation. The paper concludes that TSFMs exhibit the stability-plasticity dilemma, and that continual learning methods will be needed before such models can be deployed in non-stationary environments.","feed_headline":"Fine-tuning a forecasting model erases its earlier skills","feed_subtitle":"In a two-stage test, adapting to new series raised error on the old series from 0.15 to 1.60.","key_machinery":"The central mechanism is a two-stage sequential fine-tuning protocol on controlled synthetic datasets (D1 to D4) built by summing sine waves with harmonically or non-harmonically aligned periods, designed to avoid overlap with TimesFM's pretraining data. Performance is tracked with Mean Absolute Error (MAE) and Backward Transfer (BWT), a metric that quantifies how much performance on earlier tasks changes after learning a later task.","core_discovery":"Through a two-stage continual learning experiment on synthetic multi-sinusoidal datasets, the paper demonstrates that TimesFM, a pretrained decoder-only transformer for time-series forecasting, shows catastrophic forgetting when fine-tuned sequentially on a new dataset. After fine-tuning on D2, MAE on D1 rises from 0.15 to 1.60, while D2's error falls from 1.27 to 0.08; a second experiment, D3 to D4, shows moderate forgetting (0.56 to 0.76). Hyperparameter sweeps show that higher learning rates and more epochs intensify forgetting, whereas very low learning rates reduce forgetting but also limit new-task adaptation. The paper frames this as direct evidence of the stability-plasticity dilemma","pith_inferences":["The abstract promises comparisons with specialized models and the DER mitigation method, but the reported experiments cover only TimesFM on synthetic data; whether mitigation can let smaller models match large ones is an unverified extension.","If the stability-plasticity tradeoff holds in real settings, deployment strategies that freeze the model after pretraining, or reserve dedicated capacity for new tasks, may be more practical than full fine-tuning.","The design's distinction between harmonic and non-harmonic periods suggests that forgetting may depend on spectral overlap between old and new tasks; a testable extension is to vary period overlap and measure BWT as a function of spectral distance."],"forward_implications":["Sequential fine-tuning of a time-series foundation model on new data without replay or regularization can severely degrade performance on previously learned tasks.","Lower learning rates reduce catastrophic forgetting but at the cost of adaptation, so the operating point must be chosen based on which is more important in deployment.","Backward Transfer is positive in all high-learning-rate settings, showing that forgetting is systematic rather than a one-off artifact.","Pretrained time-series foundation models do not automatically avoid catastrophic forgetting, contradicting the assumption that large-scale pretraining confers continual-learning robustness."],"fun_headline_variants":["Fine-tuning a forecaster on new series spikes old-error by 10x","Forgetting strikes time-series models: old error jumps 10x","Continual fine-tuning erases forecasting skills, but small models recover","Fine-tuning model on new series: old error jumps to 1.60","Adapting a forecaster to new data erases old skill: error up 10x"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The synthetic multi-sinusoidal datasets are assumed to be representative of real-world non-stationary shifts in forecasting, so if actual distribution shifts differ from these hand-crafted periodic signals, the measured forgetting may not transfer beyond this controlled setup.","fun_headline_variants_meta":{"raw":{"variants":["Fine-tuning a forecaster on new series spikes old-error by 10x","Forgetting strikes time-series models: old error jumps 10x","Continual fine-tuning erases forecasting skills, but small models recover","Fine-tuning model on new series: old error jumps to 1.60","Adapting a forecaster to new data erases old skill: error up 10x"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000898,"raw_usage":{"total_tokens":3675,"prompt_tokens":686,"completion_tokens":2989,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":430,"completion_tokens_details":{"reasoning_tokens":2889}},"tokens_in":430,"tokens_out":2989,"duration_ms":21369,"temperature":1.0,"reasoning_tokens":2889,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T13:00:01.569477+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same two-stage fine-tuning on two real-world series with a genuine distribution shift, such as energy load from different buildings or seasons, and check whether BWT on the first dataset replicates the synthetic 0.15-to-1.60 jump. If the first-task error barely moves, the strong forgetting result is an artifact of the synthetic periodic design.","supporting_citations":[],"review_version":1}