{"id":"80c62083-fd0d-482c-a690-77d41f011134","arxiv_id":"1909.02716","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"An autoregressive model with promotion-state dummy variables outperforms judgmentally adjusted forecasts in two FMCG case studies, but with weak statistical validation.","lead":"This paper builds a forecast model that adds sales promotion information to historical demand data using demand uplift states, which group promotions by how much they raise sales. The authors test it on two Australian consumer goods companies and claim it outperforms their manually adjusted forecasts.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported accuracy gains hinge on DUS state labels estimated from as few as two to four promotion weeks per state and never subjected to uncertainty or robustness checks; the central claim of generalizable improvement is not yet established.","rationale":"The paper's central claim is that the FSE model improves forecast accuracy over judgmental forecasts, and the sole evidence is two 20-week hold-outs. The DUS states that define the promotion dummies in Eq. (1) are constructed from in-sample averages of actual-minus-baseline (DUS algorithm Step 3). Company A has only 16 promotional weeks split across five states, implying two to four weeks per state; Company B's state table is internally inconsistent (five stated, six listed) and includes two nearly identical uplift values treated as separate states. These small-n, data-driven state labels are treated as fixed and known in the test period, with no confidence intervals or significance testing for either the state means or the forecast improvements. If the state uplifts are noisy or non-stationary, the state grouping and FSE coefficients overfit the training window and the reported gains will not generalize. The concrete re-split test directly challenges whether the gains survive when the states are estimated from a different training window; if they do not, the abstract's success claim is unsupported. The reader's weakest assumption already identified the small-sample state estimates, so the stress-test agrees with that reading and does not change the conditional verdict.","tokens_in":14163,"tokens_out":10446,"duration_ms":115396,"concrete_test":"Re-run the full DUS+FSE pipeline with a different training/test split: for Company A, estimate states from weeks 1-60 and forecast weeks 61-100 (40-week test); for Company B, estimate states from weeks 1-80 and forecast weeks 81-120. Compare the MSAE of FSE forecasts against the companies' adjusted forecasts in these longer hold-outs. If the improvement over judgmental forecasts is not positive for both companies (or is materially smaller than the reported 38%/59%), the original gains are an artifact of the chosen split and the small number of promotional weeks used to estimate the states, and the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests entirely on two 20-week hold-out comparisons reported in Tables 3 and 6. The DUS algorithm (Section 3, Steps 0-4) constructs the state labels for the FSE regression from in-sample averages of actual-minus-baseline computed over promotional weeks. In Company A, only 16 promotional weeks in the 100-week sample are partitioned into five states (Table 2), so each state's average uplift is built from roughly two to four observations. In Company B, the text says five states while Table 5 lists six, and the In-Store single-buy and multiple-buy states have average uplifts of 16.0 and 15.8, nearly identical yet represented as distinct states. These small-sample, data-driven labels are then treated as fixed and known for the test period with no confidence intervals, bootstrap, or significance test around either the state means or the MSAE improvements. If the mean uplift for a state is noisy or non-stationary, the grouping and the FSE coefficients will overfit the training window, and the 38%/59% improvement over judgmental forecasts will not generalize beyond the reported split. The paper's own limitation notes (Section 5) concede the need for sufficient historical data but do not address this small-n instability.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a regime-switching-inspired approach to demand forecasting in the presence of systematic events, specifically sales promotions. The Demand Uplift States (DUS) algorithm defines discrete states from combinations of promotion factors, using historical actual-minus-baseline averages as state-specific uplift values. These states are then incorporated as binary regressors in an autoregressive model, called the FSE model. The model is evaluated on weekly sales data from two Australian FMCG companies (100 weeks and 120 weeks respectively), using a single 20-week hold-out test period for each, and compared against the companies' final judgmentally adjusted forecasts. The authors report substantial improvements in MSAE, MAE, and MAPE (e.g., 38% and 59% MSAE improvement), concluding that the FSE model can successfully improve forecast accuracy over current industry practice.","tokens_in":14451,"tokens_out":3171,"duration_ms":33471,"significance":"If the reported gains are robust, the FSE approach would be a meaningful, low-complexity contribution to promotional demand forecasting, potentially reducing the cognitive load on forecasters and enabling more structured forecast support systems. The use of real industry data and the explicit comparison to actual judgmental forecasts are strengths, and the DUS algorithm is intuitive and implementable. However, the current empirical evaluation is too thin to establish the central claim: it relies on a single test split per company, lacks any uncertainty quantification or significance tests for the accuracy differences, does not benchmark against standard statistical promotion models, and uses demand states estimated from very small promotional subsamples. The significance of the paper currently rests on an unverified generalization claim.","major_comments":[{"comment":"The central claim of forecast improvement rests entirely on a single 20-week hold-out test split per company, with no confidence intervals, bootstrap, or significance tests on the error differences. The reported MSAE improvements of 38% and 59% could be driven by a few influential weeks; the authors should report uncertainty around these differences (e.g., Diebold-Mariano test, block bootstrap) and, ideally, evaluate multiple rolling-origin test windows to demonstrate stability.","section":"§4.1, Tables 3 and 6; §4.2"},{"comment":"The demand uplift states are computed as average actual-minus-baseline over a very small number of promotional weeks: for Company A, only 16 promotional weeks are partitioned into five states (roughly two to four observations per state). These means are then treated as fixed and known in the test period without any measure of variability or a formal test of whether the states are distinct. The authors should provide bootstrap or cross-validation evidence that the state definitions and their uplift values are stable, and should discuss how small-sample noise in the state means affects forecast accuracy.","section":"§3, DUS algorithm Steps 3–4; Table 2"},{"comment":"There is a direct internal inconsistency in the Company B case study: the text states that the DUS algorithm prescribes five states and that the single-buy/in-store and multiple-buy/in-store combinations are grouped into State 4, but Table 5 lists six states (State 1 through State 6) and assigns these two combinations to separate states (State 4 with uplift 16 and State 5 with uplift 15.8). This inconsistency must be corrected, as it affects the reproducibility of the model and the interpretation of the reported results.","section":"Table 5 versus text in §4.2"},{"comment":"The evaluation benchmarks the FSE model only against the companies' judgmentally adjusted forecasts, not against standard statistical promotional models (e.g., regression with promotion dummies, SCAN*PRO-type models, or the baseline exponential smoothing augmented with promotion indicators). Without such a comparison, it is unclear whether the improvement comes from the DUS/FSE structure specifically or simply from including promotion information in a regression framework. Adding these baselines is necessary to support the paper's positioning of the model as a simple and practical alternative.","section":"§4.1 and §4.2"}],"minor_comments":[{"comment":"There are several garbled table references in the text (e.g., 'Table ,' and 'Table indicates' on page 16), which should be replaced with the correct table numbers.","section":"§4.1"},{"comment":"The word 'Company' is misspelled as 'Comapny' in the captions of Figures 4 and 8.","section":"Figures 4 and 8"},{"comment":"The text says 'There are 60 promotional periods that occur over the 100 observations,' but the dataset for Company B spans 120 weeks; please clarify which number is correct.","section":"§4.2"},{"comment":"The criterion for judging two combinations to have 'distinct' uplifts is not operationalized; the authors should specify a threshold or statistical test for merging or separating states.","section":"§3, Step 4 of the DUS algorithm"}],"recommendation":"major_revision","confidential_remarks":"The core idea is promising and the case study data appear genuine, but the empirical validation is currently not convincing enough for publication. The internal inconsistency in Table 5 (six states minus one grouped state) and the lack of any uncertainty quantification are the most serious problems. I would encourage the authors to strengthen the evaluation with significance tests, baseline comparisons, and robustness checks on the state definitions; with those additions the paper could become a solid applied contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know before you read it. First, the 'regime-switching' language is doing a lot of work: what the FSE model actually is is an autoregressive model with dummy variables for promotion-combination states. Second, the empirical results are honest out-of-sample comparisons against the companies' judgmental forecasts, not against any existing promotional model. That sets the bar low.\\n\\nThe genuinely useful piece is the DUS algorithm: instead of leaving every promotion combination as a separate dummy, it averages historical actual-minus-baseline uplift per combination and groups them into a handful of states. For an FMCG forecaster who currently adjusts everything by hand, that is a practical, explainable step. The two case studies use real data from two different product types, and the model beats the companies' adjusted forecasts in both 20-week test periods. The paper is also honest in its limitations section about needing sufficient history and about the manual effort in running DUS.\\n\\nBut the soft spots are real and they matter. Company A's five states are estimated from only 16 promotional weeks, so each state average is based on two to four observations. Company B's text says five states, the table lists six, and two of those states have almost identical average uplifts (16.0 and 15.8) yet are treated as distinct. There are no confidence intervals around the state means, no significance test on the 38%/59% MSAE improvements, and no comparison against a plain promotional regression with the original promotion-type dummies. Without that, the gains could easily be overfitting a small training window.\\n\\nThe paper is for practitioners who want a simple way to fold promotions into a baseline forecast, and for researchers working on structuring judgmental adjustments. It deserves a serious referee because the research question is practically important and the case data are real, but the current claims overstate both novelty and reliability. I would send it to review with a clear request: add uncertainty quantification, benchmark against existing promotional regression models, and reconcile the state definitions.","headline":"A straightforward promotional regression model with a useful grouping heuristic; the case studies are real but the state definitions are too thinly populated to support the headline accuracy claims.","tokens_in":14967,"tokens_out":2454,"would_cite":false,"duration_ms":24952,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["62M10","62J05"],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that the effects of systematic events such as retail sales promotions can be quantified and embedded directly into a demand forecasting model, replacing much of the judgmental adjustment that forecasters currently apply…","keywords":["Demand Forecasting","Systematic Events","Time Series Regression Models","Sales Promotions","Judgmental Forecasting","Supply Chain","Regime Switching","FMCG"],"falsifier":"A concrete test: take a product with several years of promotion history, split it by promotion combination and by season, and check whether the FSE model with a single fixed uplift per state beats a version that re-estimates each state's uplift on a rolling window. If the rolling version's test-period error is lower, the fixed-state assumption does not carry the claimed improvement.","tokens_in":1607,"feed_emoji":"📈","tokens_out":2412,"duration_ms":77771,"temperature":0.7,"pith_summary":"This paper develops a forecasting model that embeds the effect of retail sales promotions directly into a time series regression, instead of leaving it to human judgment. The model constructs demand uplift states from promotion mechanics, estimates the average uplift for each state, and includes those states as regressors. Tested on weekly sales data from two Australian FMCG companies, the model improves forecast accuracy compared to the companies' judgmentally adjusted forecasts, with MSAE improvements of 38% and 59% in the two cases. The result matters because it offers a simple, systematic way to reduce human intervention in forecasting to only non-systematic events.","feed_headline":"Promotion states in a regression model cut forecast error up to 59%","feed_subtitle":"Two FMCG case studies show a regression model with demand-uplift states beats expert-adjusted forecasts.","key_machinery":"The Demand Uplift States (DUS) algorithm and the Forecasting Systematic Events (FSE) regression model together carry the argument. DUS takes the list of systematic-event factors named by expert forecasters, tests their significance with ANOVA, forms all feasible combinations of their levels, and computes each combination's average demand uplift as the historical difference between actual sales and baseline exponential-smoothing forecasts; combinations with distinct uplifts become states. The FSE model then regresses demand on its own past $p$ values plus one dummy variable per state: $X_t = \\alpha_0 + \\sum_{i=1}^p \\alpha_i X_{t-i} + \\sum_{j=1}^m \\beta_j S_{jt} + \\varepsilon_t$, where $S_{jt}$ is one when the state is active at time $t$. The dummy variables carry the systematic-event effect, so in non-promotion weeks the model reduces to a plain autoregression.","core_discovery":"The central discovery is that the effects of promotion mechanics—promotion type, display type, and advertisement type—can be summarized as discrete demand states, each carrying a fixed uplift estimated from historical weeks, and that adding these states to a low-order autoregressive model produces forecasts that outperform the final judgmentally adjusted forecasts used by the companies. For Company A, the FSE model improved MSAE from 0.18 to 0.11 (38%), MAE from 622.25 to 328.54 (47%), and MAPE by 11%; for Company B, MSAE improved from 0.32 to 0.13 (59%), MAE from 30.88 to 13.62 (55%), and MAPE by 14%. The paper presents this as evidence that systematic events need not be left to unaided human judgment: once the states are identified, the model can serve as a stronger, nearly complete baseline forecast.","pith_inferences":["The comparison in the paper is against each company's own judgmentally adjusted forecasts, not against standard automated promotion models such as ARIMAX or dynamic regression with promotion dummies; a natural test is whether FSE still wins on those benchmarks.","Because uplift states are in-sample means, a promotion combination that later runs at a different discount depth or in a different season could be mis-forecast; an adaptive or shrinkage-based update of the state estimates would be a direct robustness check.","If the state multipliers generalize, the same structure could be applied hierarchically—first estimating states at category or retailer level, then scaling down to SKU level—reducing the per-SKU data requirement.","The DUS algorithm's grouping rule, which puts combinations with similar average uplift into one state, raises a threshold question: how different must two uplifts be to deserve separate states, and does the answer depend on forecast horizon or demand volatility?"],"forward_implications":["Forecasters can restrict judgmental adjustments to non-systematic events—sudden climate change, market shocks, or new campaigns—rather than re-evaluating the effect of every promotion.","The same DUS-plus-FSE machinery can be applied to other systematic events with identifiable levels, such as holidays, catalog drops, or seasonal selling periods.","Once states are fixed, the model can be re-run period after period without re-running the DUS algorithm, so the added cost is mainly the one-time state identification.","In both reported test periods, replacing unstructured promotion adjustments with state dummies improved MSAE by 38% and 59% and reduced MAE by roughly half, implying the practice could cut forecast error and labor cost in FMCG planning.","The approach inherits the limitation of historical methods: it needs enough promotion history, so it is not directly usable for new products with no promotional track record."],"supporting_citations":[{"why":"Supplies the regime-switching concept that the paper adapts to define demand uplift states for non-stationary time series.","marker":"Hamilton (1989)"},{"why":"Provides the empirical evidence on judgmental adjustments to statistical forecasts that motivates structuring those adjustments.","marker":"Fildes et al. (2009)"},{"why":"Documents how promotions drive judgmental adjustments, the practice the FSE model aims to replace.","marker":"Trapero et al. (2013)"},{"why":"Studies identification of sales forecasting models in the presence of promotions and serves as a comparison standard for promotional forecasting.","marker":"Trapero et al. (2015)"},{"why":"Underpins the exponential smoothing baseline that both case companies use for their baseline forecasts.","marker":"Hyndman, Koehler, Ord, & Snyder (2008)"},{"why":"Justifies the choice of MSAE as the scale-independent accuracy measure used to compare forecasts.","marker":"Davydenko & Fildes (2013)"},{"why":"Supports the argument that structuring judgmental information improves forecast accuracy.","marker":"Green & Armstrong (2007)"}],"fun_headline_variants":["Promotion states cut forecast error up to 59%","Systematic events in demand: model beats human judgment","Model quantifies promotions, trims forecast errors by half","Regime-switching model improves forecast accuracy in FMCG"],"cache_read_input_tokens":17152,"weakest_assumption_plain":"The average uplift measured from a handful of past promotions of each type is treated as a fixed multiplier that will apply to every future promotion of that type, even if execution, season, or market conditions change.","fun_headline_variants_meta":{"raw":{"variants":["Promotion states cut forecast error up to 59%","Systematic events in demand: model beats human judgment","Model quantifies promotions, trims forecast errors by half","Regime-switching model improves forecast accuracy in FMCG"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000223,"raw_usage":{"total_tokens":1454,"prompt_tokens":942,"completion_tokens":512,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":558,"completion_tokens_details":{"reasoning_tokens":445}},"tokens_in":558,"tokens_out":512,"duration_ms":5974,"temperature":1.0,"reasoning_tokens":445,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:41:48.411029+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete test: take a product with several years of promotion history, split it by promotion combination and by season, and check whether the FSE model with a single fixed uplift per state beats a version that re-estimates each state's uplift on a rolling window. If the rolling version's test-period error is lower, the fixed-state assumption does not carry the claimed improvement.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Studies identification of sales forecasting models in the presence of promotions and serves as a comparison standard for promotional forecasting."}],"review_version":1}