Pith. sign in

REVIEW 2 cited by

$\spadesuit$ SPADE $\spadesuit$ Split Peak Attention DEcomposition

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.05852 v2 pith:QY45N4JF submitted 2024-11-06 cs.LG stat.ML

classification cs.LGstat.ML
keywords demandpeakattentionforecastingforecastsimprovementperiodsspade
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Demand forecasting faces challenges induced by Peak Events (PEs) corresponding to special periods such as promotions and holidays. Peak events create significant spikes in demand followed by demand ramp down periods. Neural networks like MQCNN and MQT overreact to demand peaks by carrying over the elevated PE demand into subsequent Post-Peak-Event (PPE) periods, resulting in significantly over-biased forecasts. To tackle this challenge, we introduce a neural forecasting model called Split Peak Attention DEcomposition, SPADE. This model reduces the impact of PEs on subsequent forecasts by modeling forecasting as consisting of two separate tasks: one for PEs; and the other for the rest. Its architecture then uses masked convolution filters and a specialized Peak Attention module. We show SPADE's performance on a worldwide retail dataset with hundreds of millions of products. Our results reveal an overall PPE improvement of 4.5%, a 30% improvement for most affected forecasts after promotions and holidays, and an improvement in PE accuracy by 3.9%, relative to current production models.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Investigating Compositional Reasoning in Time Series Foundation Models

    cs.LG 2025-02 conditional novelty 6.0 of 10

    On a benchmark where models train on Fourier components and test on their sums, patch-based Transformers and residual MLP architectures show the strongest compositional generalization, while most standard transformers...

  2. TAT: Temporal-Aligned Transformer for Multi-Horizon Peak Demand Forecasting

    cs.LG 2025-07 conditional novelty 5.0 of 10

    TAT, a transformer with temporal-alignment attention and posterior calibration, improves peak demand forecast accuracy by up to 30% on proprietary e-commerce data.

Pith tools