Pith. sign in

REVIEW 10 cited by

CoST: Contrastive Learning of Disentangled Seasonal-Trend Representations for Time Series Forecasting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.01575 v3 pith:XUQ6RRMX submitted 2022-02-03 cs.LG

classification cs.LG
keywords timecostlearningseriesforecastingrepresentationscontrastivedisentangled
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Deep learning has been actively studied for time series forecasting, and the mainstream paradigm is based on the end-to-end training of neural network architectures, ranging from classical LSTM/RNNs to more recent TCNs and Transformers. Motivated by the recent success of representation learning in computer vision and natural language processing, we argue that a more promising paradigm for time series forecasting, is to first learn disentangled feature representations, followed by a simple regression fine-tuning step -- we justify such a paradigm from a causal perspective. Following this principle, we propose a new time series representation learning framework for time series forecasting named CoST, which applies contrastive learning methods to learn disentangled seasonal-trend representations. CoST comprises both time domain and frequency domain contrastive losses to learn discriminative trend and seasonal representations, respectively. Extensive experiments on real-world datasets show that CoST consistently outperforms the state-of-the-art methods by a considerable margin, achieving a 21.3% improvement in MSE on multivariate benchmarks. It is also robust to various choices of backbone encoders, as well as downstream regressors. Code is available at https://github.com/salesforce/CoST.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation

    cs.LG 2026-08 conditional novelty 7.0 of 10

    A shared cardiac encoder pretrained with delay-aware cross-modal JEPA on ECG, PPG, and PCG beats modality-specific self-supervised baselines on 25 downstream tasks.

  2. Enhancing Contrastive Learning-based Electrocardiogram Pretrained Model with Patient Memory Queue

    eess.SP 2025-05 conditional novelty 6.0 of 10

    A patient memory queue that stores representations with patient IDs improves contrastive ECG pretraining, especially under severe label scarcity.

  3. D-Tracker: Modeling Interest Diffusion in Social Activity Tensor Data Streams

    cs.SI 2025-05 conditional novelty 6.0 of 10

    D-Tracker combines tensor decomposition with a reaction-diffusion system to model and forecast social activity streams while automatically switching models as patterns change.

  4. Pulse-PPG: An Open-Source Field-Trained PPG Foundation Model for Wearable Applications Across Lab and Field Settings

    cs.LG 2025-02 conditional novelty 6.0 of 10

    A field-trained, open-source PPG foundation model outperforms a clinical-data-trained model on 10 of 11 downstream health tasks across wearable and clinical settings.

  5. Memory-enhanced Invariant Prompt Learning for Urban Flow Prediction under Distribution Shifts

    cs.LG 2024-12 reject novelty 6.0 of 10

    MIP uses memory-bank prompts and variance-based invariant training to improve urban flow prediction under distribution shifts on METR-LA and NYCBike1.

  6. BrainStratify: Coarse-to-Fine Disentanglement of Intracranial Neural Dynamics

    eess.SP 2025-05 conditional novelty 5.0 of 10

    BrainStratify's coarse-to-fine disentanglement, electrode clustering plus decoupled product quantization, modestly improves speech decoding over prior methods on sEEG and epidural ECoG datasets.

  7. eMargin: Revisiting Contrastive Learning with Margin-Based Separation

    cs.LG 2025-07 reject novelty 4.0 of 10

    An adaptive margin added to InfoNCE improves time series clustering metrics but hurts linear-probe classification, exposing a disconnect between clustering scores and downstream utility.

  8. Enhancing Masked Time-Series Modeling via Dropping Patches

    stat.ML 2024-12 reject novelty 4.0 of 10

    Randomly dropping 60% of patches before masked pre-training improves PatchTST time-series forecasting accuracy and training speed, though the paper's theoretical explanation is not sound.

  9. Act Now: A Novel Online Forecasting Framework for Large-Scale Streaming Data

    cs.LG 2024-11 conditional novelty 3.0 of 10

    Act-Now combines random subgraph sampling, fast and slow stream buffers, and a label-decomposition forecaster, reporting large MSE reductions on three cellular traffic datasets.

  10. MFF-FTNet: Multi-scale Feature Fusion across Frequency and Temporal Domains for Time Series Forecasting

    cs.LG 2024-11 reject novelty 3.0 of 10

    MFF-FTNet combines CoST-style contrastive learning, TimesNet-style top-k frequency selection, and TCN-style multi-scale convolutions, reporting modest average MSE gains but leaving the forecasting loss undefined.

Pith tools