Pith. sign in

REVIEW 4 major objections 4 minor 9 references

Demand Forecasting in the Presence of Systematic Events: Cases in Capturing Sales Promotions

T0 review · 4 major / 4 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read The paper claims that the effects of systematic events such as retail sales promotions can be quantified and embedded directly into a demand forecasting model, replacing much of the judgmental adjustment that forecasters currently apply…

desk verdict A straightforward promotional regression model with a useful grouping heuristic; the case studies are real but the state definitions are too thinly populated to support the headline accuracy claims. read the letter →

arxiv 1909.02716 v1 pith:HW2YIL6Y submitted 2019-09-06 stat.AP

classification stat.AP MSC 62M1062J05
keywords DemandForecastingSystematicEventsTimeSeriesRegressionModelsSalesPromotionsJudgmentalSupplyChainRegimeSwitchingFMCG
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper develops a forecasting model that embeds the effect of retail sales promotions directly into a time series regression, instead of leaving it to human judgment. The model constructs demand uplift states from promotion mechanics, estimates the average uplift for each state, and includes those states as regressors. Tested on weekly sales data from two Australian FMCG companies, the model improves forecast accuracy compared to the companies' judgmentally adjusted forecasts, with MSAE improvements of 38% and 59% in the two cases. The result matters because it offers a simple, systematic way to reduce human intervention in forecasting to only non-systematic events.

What carries the argument

The Demand Uplift States (DUS) algorithm and the Forecasting Systematic Events (FSE) regression model together carry the argument. DUS takes the list of systematic-event factors named by expert forecasters, tests their significance with ANOVA, forms all feasible combinations of their levels, and computes each combination's average demand uplift as the historical difference between actual sales and baseline exponential-smoothing forecasts; combinations with distinct uplifts become states. The FSE model then regresses demand on its own past $p$ values plus one dummy variable per state: $X_t = \alpha_0 + \sum_{i=1}^p \alpha_i X_{t-i} + \sum_{j=1}^m \beta_j S_{jt} + \varepsilon_t$, where $S_{jt}$ is one when the state is active at time $t$. The dummy variables carry the systematic-event effect, so in non-promotion weeks the model reduces to a plain autoregression.

What would settle it

A concrete test: take a product with several years of promotion history, split it by promotion combination and by season, and check whether the FSE model with a single fixed uplift per state beats a version that re-estimates each state's uplift on a rolling window. If the rolling version's test-period error is lower, the fixed-state assumption does not carry the claimed improvement.

Watch

Extended reading notes

Core claim

The central discovery is that the effects of promotion mechanics—promotion type, display type, and advertisement type—can be summarized as discrete demand states, each carrying a fixed uplift estimated from historical weeks, and that adding these states to a low-order autoregressive model produces forecasts that outperform the final judgmentally adjusted forecasts used by the companies. For Company A, the FSE model improved MSAE from 0.18 to 0.11 (38%), MAE from 622.25 to 328.54 (47%), and MAPE by 11%; for Company B, MSAE improved from 0.32 to 0.13 (59%), MAE from 30.88 to 13.62 (55%), and MAPE by 14%. The paper presents this as evidence that systematic events need not be left to unaided human judgment: once the states are identified, the model can serve as a stronger, nearly complete baseline forecast.

Load-bearing premise

The average uplift measured from a handful of past promotions of each type is treated as a fixed multiplier that will apply to every future promotion of that type, even if execution, season, or market conditions change.

Editorial extensions

If this is right

  • Forecasters can restrict judgmental adjustments to non-systematic events—sudden climate change, market shocks, or new campaigns—rather than re-evaluating the effect of every promotion.
  • The same DUS-plus-FSE machinery can be applied to other systematic events with identifiable levels, such as holidays, catalog drops, or seasonal selling periods.
  • Once states are fixed, the model can be re-run period after period without re-running the DUS algorithm, so the added cost is mainly the one-time state identification.
  • In both reported test periods, replacing unstructured promotion adjustments with state dummies improved MSAE by 38% and 59% and reduced MAE by roughly half, implying the practice could cut forecast error and labor cost in FMCG planning.
  • The approach inherits the limitation of historical methods: it needs enough promotion history, so it is not directly usable for new products with no promotional track record.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The comparison in the paper is against each company's own judgmentally adjusted forecasts, not against standard automated promotion models such as ARIMAX or dynamic regression with promotion dummies; a natural test is whether FSE still wins on those benchmarks.
  • Because uplift states are in-sample means, a promotion combination that later runs at a different discount depth or in a different season could be mis-forecast; an adaptive or shrinkage-based update of the state estimates would be a direct robustness check.
  • If the state multipliers generalize, the same structure could be applied hierarchically—first estimating states at category or retailer level, then scaling down to SKU level—reducing the per-SKU data requirement.
  • The DUS algorithm's grouping rule, which puts combinations with similar average uplift into one state, raises a threshold question: how different must two uplifts be to deserve separate states, and does the answer depend on forecast horizon or demand volatility?
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes a regime-switching-inspired approach to demand forecasting in the presence of systematic events, specifically sales promotions. The Demand Uplift States (DUS) algorithm defines discrete states from combinations of promotion factors, using historical actual-minus-baseline averages as state-specific uplift values. These states are then incorporated as binary regressors in an autoregressive model, called the FSE model. The model is evaluated on weekly sales data from two Australian FMCG companies (100 weeks and 120 weeks respectively), using a single 20-week hold-out test period for each, and compared against the companies' final judgmentally adjusted forecasts. The authors report substantial improvements in MSAE, MAE, and MAPE (e.g., 38% and 59% MSAE improvement), concluding that the FSE model can successfully improve forecast accuracy over current industry practice.

Significance. If the reported gains are robust, the FSE approach would be a meaningful, low-complexity contribution to promotional demand forecasting, potentially reducing the cognitive load on forecasters and enabling more structured forecast support systems. The use of real industry data and the explicit comparison to actual judgmental forecasts are strengths, and the DUS algorithm is intuitive and implementable. However, the current empirical evaluation is too thin to establish the central claim: it relies on a single test split per company, lacks any uncertainty quantification or significance tests for the accuracy differences, does not benchmark against standard statistical promotion models, and uses demand states estimated from very small promotional subsamples. The significance of the paper currently rests on an unverified generalization claim.

major comments (4)
  1. [§4.1, Tables 3 and 6; §4.2] The central claim of forecast improvement rests entirely on a single 20-week hold-out test split per company, with no confidence intervals, bootstrap, or significance tests on the error differences. The reported MSAE improvements of 38% and 59% could be driven by a few influential weeks; the authors should report uncertainty around these differences (e.g., Diebold-Mariano test, block bootstrap) and, ideally, evaluate multiple rolling-origin test windows to demonstrate stability.
  2. [§3, DUS algorithm Steps 3–4; Table 2] The demand uplift states are computed as average actual-minus-baseline over a very small number of promotional weeks: for Company A, only 16 promotional weeks are partitioned into five states (roughly two to four observations per state). These means are then treated as fixed and known in the test period without any measure of variability or a formal test of whether the states are distinct. The authors should provide bootstrap or cross-validation evidence that the state definitions and their uplift values are stable, and should discuss how small-sample noise in the state means affects forecast accuracy.
  3. [Table 5 versus text in §4.2] There is a direct internal inconsistency in the Company B case study: the text states that the DUS algorithm prescribes five states and that the single-buy/in-store and multiple-buy/in-store combinations are grouped into State 4, but Table 5 lists six states (State 1 through State 6) and assigns these two combinations to separate states (State 4 with uplift 16 and State 5 with uplift 15.8). This inconsistency must be corrected, as it affects the reproducibility of the model and the interpretation of the reported results.
  4. [§4.1 and §4.2] The evaluation benchmarks the FSE model only against the companies' judgmentally adjusted forecasts, not against standard statistical promotional models (e.g., regression with promotion dummies, SCAN*PRO-type models, or the baseline exponential smoothing augmented with promotion indicators). Without such a comparison, it is unclear whether the improvement comes from the DUS/FSE structure specifically or simply from including promotion information in a regression framework. Adding these baselines is necessary to support the paper's positioning of the model as a simple and practical alternative.
minor comments (4)
  1. [§4.1] There are several garbled table references in the text (e.g., 'Table ,' and 'Table indicates' on page 16), which should be replaced with the correct table numbers.
  2. [Figures 4 and 8] The word 'Company' is misspelled as 'Comapny' in the captions of Figures 4 and 8.
  3. [§4.2] The text says 'There are 60 promotional periods that occur over the 100 observations,' but the dataset for Company B spans 120 weeks; please clarify which number is correct.
  4. [§3, Step 4 of the DUS algorithm] The criterion for judging two combinations to have 'distinct' uplifts is not operationalized; the authors should specify a threshold or statistical test for merging or separating states.

Circularity Check

1 steps flagged · score 6.0 of 10

DUS state identification uses actuals from the full sample, including the test window, before the train/test split, so the test-period regressors are partly built from the outcome being predicted.

  1. fitted input called prediction [Section 3 (DUS algorithm, Steps 3–4) combined with Section 4.1 train/test split; same pattern in Section 4.2]
    "Demand uplifts in promotional periods are computed by subtracting baseline forecasts from the actual realized sales value in each epoch. ... After running the DUS algorithm, both promotional mechanics prove to be statistically significant at the 5% level. Although, theoretically, there are eight possible combinations of the levels of the two promotional types, based on available data, the DUS algorithm prescribes five demand uplift states, as shown in Table 2. ..."

    The DUS state labels are obtained by averaging actual-minus-baseline uplift over 'available data' (Section 3, Step 3), and the DUS run is reported before the 80/20 split (Section 4.1). Therefore the state assignment of each promotion combination in the 20-week test window is influenced by the actual sales that the FSE model is then asked to forecast. Equation (1)'s test-period regressors S_jt are thus constructed from the test-period outcome variable, so the 38% MSAE improvement is not a clean out-of-sample result; part of the forecast input already encodes the answer. The same ordering appears in Company B (Section 4.2), where DUS is described before the 100/20 split. The paper's limitation remarks (Section 5) ask for sufficient historical data but do not address this leakage.

full rationale

The paper contains no self-citations and no equation-level identity where a defined quantity is its own prediction. The FSE regression coefficients are estimated only on the training portion, and the test-period comparison of forecasts against actual sales is out-of-sample in principle. However, the state-identification step of the DUS algorithm is described as running on 'available data' before the train/test split is introduced. Because those state labels define the S_jt regressors used to forecast the test period, and because they are computed from actual-minus-baseline demand in the same test window, the reported improvements (38% and 59% MSAE) are partially circular: the model inputs for the test period are informed by the outcomes being forecast. Had the DUS states been estimated only on the training set, this would be a standard supervised setup and the circularity score would be low. As written, the out-of-sample claim is weakened by this leakage, so the paper merits a partial circularity score of 6.

Assumptions & free parameters 6 free parameters · 4 assumptions · 1 invented entities

The core modeling choices are the AR order, the state count, the uplift magnitudes, and the selection of promotion factors by experts. These are all estimated or chosen rather than derived from first principles, and the state uplift magnitudes in particular depend on the accuracy of the baseline exponential smoothing forecasts.

free parameters (6)
  • AR order p = 2 for both case studies
    Selected as the order with lowest AICc on the training data (Section 4).
  • Regression coefficients alpha_i and beta_j = Not reported numerically
    Estimated by least squares on 80 or 100 weeks of training data; the paper only states they are significant at the 5% level.
  • Per-state average demand uplifts = Company A: 19816, 14833, 5091, 3466, 4121; Company B: 311, 213.3, 160.48, 16, 15.8 (approximate)
    Computed by DUS algorithm Step 3 as mean of actual sales minus baseline forecast for each promotion combination in the training period.
  • Number of states m = 5 (both companies)
    Determined by DUS algorithm based on distinct average uplifts; some combinations merged.
  • Promotion factors and levels = Company A: promotion type (major/minor), display type (entrance, FGE, other gondola, fixture); Company B: promotion…
    Provided by expert forecasters, not derived from data.
  • Significance threshold = 0.05
    Used for ANOVA variable selection and residual diagnostic tests.
assumptions (4)
  • domain assumption Underlying demand series is stationary after accounting for promotion effects
    Stated in Section 3; KPSS test used but with low power and small sample.
  • domain assumption Promotional plans are known and fixed at forecast time
    Assumed in Section 3 and used in test-set forecasts; if promotions change, states may be misassigned.
  • domain assumption Baseline statistical forecasts (simple exponential smoothing) are an unbiased reference for computing demand uplift
    DUS algorithm Step 3 subtracts baseline forecasts from actuals; if baseline forecasts are biased, the state uplifts inherit that bias.
  • standard math Error term is Gaussian white noise
    Assumed in FSE model (Equation 1) and checked with normality and Ljung-Box tests.
invented entities (1)
  • Demand Uplift State (DUS)
    purpose: Categorical variable representing the expected incremental sales associated with a specific combination of promotion levels; used as regressor in FSE model.
    DUS is a data-derived grouping without external validation; its values are estimated from the same data used for forecasting, and there is no independent test of whether the states capture true uplift.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Demand Forecasting in the Presence of Systematic Events: Cases in Capturing Sales Promotions." pith.science (2026). https://pith.science/paper/HW2YIL6Y

@misc{pith2026190902716,
  author       = {Pith},
  title        = {Pith review of: Demand Forecasting in the Presence of Systematic Events: Cases in Capturing Sales Promotions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/HW2YIL6Y}},
  note         = {Machine review of arXiv:1909.02716}
}
read the original abstract

Reliable demand forecasts are critical for the effective supply chain management. Several endogenous and exogenous variables can influence the dynamics of demand, and hence a single statistical model that only consists of historical sales data is often insufficient to produce accurate forecasts. In practice, the forecasts generated by baseline statistical models are often judgmentally adjusted by forecasters to incorporate factors and information that are not incorporated in the baseline models. There are however systematic events whose effect can be effectively quantified and modeled to help minimize human intervention in adjusting the baseline forecasts. In this paper, we develop and test a novel regime-switching approach to quantify systematic information/events and objectively incorporate them into the baseline statistical model. Our simple yet practical and effective model can help limit forecast adjustments to only focus on the impact of less systematic events such as sudden climate change or dynamic market activities. The proposed model and approach is validated empirically using sales and promotional data from two Australian companies. Discussions focus on a thorough analysis of the forecasting and benchmarking results. Our analysis indicates that the proposed model can successfully improve the forecast accuracy when compared to the current industry practice which heavily relies on human judgment to factor in all types of information/events.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

9 extracted references · 9 canonical work pages

  1. [1]

    2Institute of Transport and Logistics Studies, The University of Sydney, NSW, Australia Abstract Reliable demand forecasts are critical for effective supply chain management

    1 Demand Forecasting in the Presence of Systematic Events: Cases in Capturing Sales Promotions Mahdi Abolghasemi1, Ali Eshragh1, Jason Hurley2, Behnam Fahimnia2 1School of Mathematical and Physical Sciences, The University of Newcastle, NSW, Australia. 2Institute of Transport and Logistics Studies, The University of Sydney, NSW, Australia Abstract Reliabl...

  2. [4]

    is the forecast at time 𝑡, and 𝑥

    4 Model Validation: Empirical Case Studies We use empirical sales and promotional data obtained from two FMCG companies in Australia to investigate the validity and industry application of the FSE model. Both companies are major players in the food and beverage industry. We consider demand forecasting for one major product at each company, a more perishab...

  3. [6]

    episodes across which the dynamic behavior of the series is markedly different

    Find the most significant systematic events by running ANOVA over the data. 3 The terms ‘regime’ and ‘state’ are used interchangeably throughout this paper and are defined as “episodes across which the dynamic behavior of the series is markedly different” (Hamilton 1989, p. 358). 11 Step

  4. [8]

    =|𝑓"−𝑥"| (4) 𝑅𝐸

    The absolute errors and relative errors are calculated from equations (4) and (5), respectively. 𝐴𝐸"=|𝑓"−𝑥"| (4) 𝑅𝐸" = ;<)=<=< (5) Where 𝑓"is the forecast at time𝑡, and 𝑥" is the actual demand at time 𝑡. As illustrated in Figures 2-4 and the summary results reported in Table , the FSE model significantly improves the forecast accuracy when compared to Com...

  5. [9]

    Table 6 clearly indicates the outstanding improvements using different measures such as MSAE (59% improvement), MAE (55% improvement), and MAPE (14% improvement)

    As illustrated in Figures 6-8 and the summary results reported in Table 6, the FSE model significantly improves the forecast accuracy when compared to Company B’s judgmentally adjusted forecasts. Table 6 clearly indicates the outstanding improvements using different measures such as MSAE (59% improvement), MAE (55% improvement), and MAPE (14% improvement)...

  6. [2000]

    statistically sophisticated or complex methods do not necessarily produce more accurate forecasts than simpler ones

    consistently finds that “statistically sophisticated or complex methods do not necessarily produce more accurate forecasts than simpler ones” (Makridakis & Hibon, 2000, p. 452). This notion is also supported by Green and Armstrong (2015) who find that complexity of the forecasting method harms accuracy, and that simpler methods reduce the likelihood of er...

  7. [2008]

    and some of techniques and concepts utilized to tackle those situations could be applicable to a demand forecasting context. In particular, Hamilton (1989) developed a novel approach, the so-called Markov switching model, to more accurately capture and 10 predict changes in the regime or state3 of non-stationary time series. Although, Hamilton initially a...

  8. [2010]

    and dynamic regression involving principal component analysis and transfer functions (Trapero et al., 2015). Despite all those efforts, evidence indicates that lack of resources, expertise, and high costs hinder the widespread implementation of such methods and support systems in practice (Hughes, 2001). We aim to tackle this issue by introducing an easy-...

Show all 9 references
  1. [2014]

    Overall, the debate surrounding the practice of judgmental forecast adjustments has now evolved beyond merely whether they should be utilized or abandoned

    have inevitably resulted in its widespread use in industry (Sanders & Manrodt, 2003b). Overall, the debate surrounding the practice of judgmental forecast adjustments has now evolved beyond merely whether they should be utilized or abandoned. The question is more how to approp...

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.