REVIEW 10 cited by
CoST: Contrastive Learning of Disentangled Seasonal-Trend Representations for Time Series Forecasting
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Deep learning has been actively studied for time series forecasting, and the mainstream paradigm is based on the end-to-end training of neural network architectures, ranging from classical LSTM/RNNs to more recent TCNs and Transformers. Motivated by the recent success of representation learning in computer vision and natural language processing, we argue that a more promising paradigm for time series forecasting, is to first learn disentangled feature representations, followed by a simple regression fine-tuning step -- we justify such a paradigm from a causal perspective. Following this principle, we propose a new time series representation learning framework for time series forecasting named CoST, which applies contrastive learning methods to learn disentangled seasonal-trend representations. CoST comprises both time domain and frequency domain contrastive losses to learn discriminative trend and seasonal representations, respectively. Extensive experiments on real-world datasets show that CoST consistently outperforms the state-of-the-art methods by a considerable margin, achieving a 21.3% improvement in MSE on multivariate benchmarks. It is also robust to various choices of backbone encoders, as well as downstream regressors. Code is available at https://github.com/salesforce/CoST.
Forward citations
Cited by 10 Pith papers
-
CardioState-JEPA: Delay-Aware Cross-Modal Learning of a Shared Cardiac Representation
A shared cardiac encoder pretrained with delay-aware cross-modal JEPA on ECG, PPG, and PCG beats modality-specific self-supervised baselines on 25 downstream tasks.
-
Enhancing Contrastive Learning-based Electrocardiogram Pretrained Model with Patient Memory Queue
A patient memory queue that stores representations with patient IDs improves contrastive ECG pretraining, especially under severe label scarcity.
-
D-Tracker: Modeling Interest Diffusion in Social Activity Tensor Data Streams
D-Tracker combines tensor decomposition with a reaction-diffusion system to model and forecast social activity streams while automatically switching models as patterns change.
-
Pulse-PPG: An Open-Source Field-Trained PPG Foundation Model for Wearable Applications Across Lab and Field Settings
A field-trained, open-source PPG foundation model outperforms a clinical-data-trained model on 10 of 11 downstream health tasks across wearable and clinical settings.
-
Memory-enhanced Invariant Prompt Learning for Urban Flow Prediction under Distribution Shifts
MIP uses memory-bank prompts and variance-based invariant training to improve urban flow prediction under distribution shifts on METR-LA and NYCBike1.
-
BrainStratify: Coarse-to-Fine Disentanglement of Intracranial Neural Dynamics
BrainStratify's coarse-to-fine disentanglement, electrode clustering plus decoupled product quantization, modestly improves speech decoding over prior methods on sEEG and epidural ECoG datasets.
-
eMargin: Revisiting Contrastive Learning with Margin-Based Separation
An adaptive margin added to InfoNCE improves time series clustering metrics but hurts linear-probe classification, exposing a disconnect between clustering scores and downstream utility.
-
Enhancing Masked Time-Series Modeling via Dropping Patches
Randomly dropping 60% of patches before masked pre-training improves PatchTST time-series forecasting accuracy and training speed, though the paper's theoretical explanation is not sound.
-
Act Now: A Novel Online Forecasting Framework for Large-Scale Streaming Data
Act-Now combines random subgraph sampling, fast and slow stream buffers, and a label-decomposition forecaster, reporting large MSE reductions on three cellular traffic datasets.
-
MFF-FTNet: Multi-scale Feature Fusion across Frequency and Temporal Domains for Time Series Forecasting
MFF-FTNet combines CoST-style contrastive learning, TimesNet-style top-k frequency selection, and TCN-style multi-scale convolutions, reporting modest average MSE gains but leaving the forecasting loss undefined.
Discussion (0). Continue with ORCID to comment.