REVIEW 3 major objections 6 minor 61 references
PFML: Self-Supervised Learning of Time-Series Data Without Representation Collapse
T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read PFML, a self-supervised objective that predicts statistical functionals of masked time-series frames instead of reconstructing the signal, learns representations that do not collapse and match data2vec across IMU, speech, and EEG…
desk verdict A simple, practical SSL objective for time-series that mostly delivers on its promises, but the non-collapse proof needs fixing before the claim can stand. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the set of statistical functionals, comprising mean, variance, skewness, kurtosis, minimum and maximum value, zero-crossing rate, and the mean, variance, skewness, and kurtosis of the autocorrelation function. They are precomputed per frame, z-score normalized, and act as prediction targets for masked frames instead of high-dimensional waveforms or learned targets. The mechanism that prevents collapse is the variance argument: if the targets $f_n$ vary across frames, a constant prediction cannot achieve low MSE or L1 loss, so the model is pushed toward output variance; this stands in contrast to methods like data2vec whose targets are produced by the model itself and can collapse.
What would settle it
Monitor the variance of the encoder embeddings $z_n$ during PFML pre-training on a dataset whose frames and functionals are near-constant; the invariant claim predicts collapse has no route, while a run where $z_n$ variance drops below the 0.01 threshold while loss keeps decreasing would refute it.
Extended reading notes
Core claim
The paper's claim is that predicting statistical functionals from masked latent embeddings is sufficient to learn useful, non-collapsed time-series representations without contrastive sampling, clustering, or target networks. Concretely, PFML frames the signal, computes 11 functionals per frame, randomly masks a block of encoder embeddings, and trains a Transformer to predict the functionals of the masked frames from the unmasked context. Under the stated assumptions that the signal frames vary over time and so do their functionals, low prediction loss forces the model's outputs to vary; the paper argues this makes collapsed representations impossible and confirms empirically that PFML never triggers its collapse criterion. In downstream evaluations PFML outperforms MAE and TS2Vec on the five tested tasks and is on par with data2vec, with the largest gains on sleep-stage classification from EEG.
Load-bearing premise
The collapse guarantee rests on the assumption that the chosen functionals of real signal frames vary over time; if a dataset's frames or functionals are nearly constant, the target variance disappears and the provided derivation no longer rules out collapse.
Editorial extensions
If this is right
- PFML should transfer to new sensor modalities with little hyperparameter search, since its only algorithm-specific knobs are masking probability, mask length, and the choice of functionals.
- Because the prediction targets are pre-computable and require no teacher network or negative sampling, pre-training is cheap and runs on a single 16 GB GPU.
- The ablation results imply that using a richer functional set improves downstream performance, so the method's ceiling is tied to how well the chosen functionals describe the signal frames.
- Masking embeddings rather than raw inputs helps downstream tasks, meaning the encoder is protected from the hardest part of reconstruction.
- Since PFML never collapsed in 10 runs per modality, it removes the need for the collapse-detection restarts that data2vec required in these experiments.
Reading between the lines
- The variance proof covers the model outputs $y_n$, not the encoder embeddings $z_n$ used in downstream classifiers; the claim that embeddings do not collapse is therefore an empirical one resting on the optimization and on the chosen 0.01 variance threshold.
- A dataset whose frames are nearly constant, or a functional set that is nearly constant on that data, would void Assumptions 1 and 2; testing on such data would delineate when the collapse guarantee actually holds.
- The same functional-prediction idea could be applied to images by computing functionals over image patches, which the authors mention as a possible extension; the main unknown is which local statistics would carry enough information.
- The comparison with TS2Vec is partly confounded by architecture flexibility: TS2Vec is locked into its built-in encoder and cannot pre-train the Transformer, which likely explains part of its performance gap.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes PFML (Prediction of Functionals from Masked Latents), a self-supervised objective for time-series data in which an encoder produces frame-level embeddings, a random subset of embeddings is masked, and a Transformer predicts statistical functionals (mean, variance, skewness, kurtosis, ZCR, ACF statistics) of the input frames corresponding to the masked embeddings. The authors argue that because the target functionals have variance across frames (Assumptions 1 and 2), the model cannot converge to a collapsed representation; Appendix A gives a formal proof that the predictions y_n must have nonzero variance. The method is evaluated on infant IMU posture/movement classification, speech emotion recognition, and EEG sleep staging, comparing against MAE, data2vec, TS2Vec, and no pre-training. Results show PFML approximately matches data2vec, outperforms MAE and TS2Vec, and, in 10 runs, never triggers the authors' collapse-detection heuristic, whereas data2vec collapses in 8–9 of 10 runs depending on modality.
Significance. If the central claims are correct, PFML is a conceptually simple and architecture-flexible SSL objective that avoids representation collapse by construction and achieves state-of-the-art-level downstream performance across multiple sensor modalities. The paper's strengths include a clear and simple formulation, publicly released code, evaluation on three real-world clinical/affective datasets, and a direct comparison with a strong modality-agnostic baseline (data2vec). However, the theoretical guarantee of non-collapse is not actually established by the provided proof, because the proof concerns the predictions y_n rather than the encoder embeddings z_n that are used for downstream tasks. This gap weakens the paper's main selling point. Additionally, the empirical results are reported without error bars or significance tests, making the performance comparisons difficult to interpret. The paper is therefore promising but needs substantial revision to support its headline claims.
major comments (3)
- [Section III-B and Appendix A] The proof in Appendix A shows only that the prediction outputs y_n must have nonzero variance if the loss is low, because the targets f_n have variance. It does not show that the encoder embeddings z_n have nonzero variance or that they depend on the input content. The paper's own definition of representation collapse in Section III-A is 'a constant, input-invariant feature representation' — i.e., a property of the embeddings. Since the architecture uses relative positional encoding (Section IV-A) and masked embeddings are replaced by a vector of ones, the Transformer can produce position- and mask-dependent y_n even when every z_n is identical. Thus the claim in the abstract, Section III-B ('does not converge to collapsed feature representations, as long as Assumptions 1 and 2 hold true'), and the conclusion is not supported by the derivation. The proof must be extended to bound the variance or content-dependence of z_n under suitable assumptions, or the theoretical claim must be weakened to non-degeneracy of predictions, with embedding non-collapse presented as an empirical observation.
- [Section IV-A and Table 3] The collapse-detection criterion used for Table 3 — variance of embeddings or outputs falling below 0.01 for 10 consecutive epochs — cannot detect embeddings that are constant with respect to input content but vary with position or mask pattern, because such embeddings have high variance. Therefore the reported zero collapses for PFML do not establish that the learned representations are input-invariant-free in the sense of Section III-A. The authors should report an additional diagnostic that measures content-dependence, for example the variance of z_n after conditioning on position and mask configuration, or the mutual information between z_n and x_n, to support the non-collapse claim empirically.
- [Tables 1 and 2] The fine-tuning and linear evaluation results appear to be from single runs, with no standard deviations, confidence intervals, or significance tests. Several differences between PFML and data2vec are very small (e.g., 81.8 vs 81.9 UAF1 for IMU movement; 70.7 vs 70.7 UAR for speech valence), so the statements that PFML is 'superior to MAE' and 'on par with the current state-of-the-art' are not statistically supported by the reported evidence. Reporting mean ± std over at least three seeds and, where appropriate, a paired significance test would be necessary to substantiate the comparative claims.
minor comments (6)
- [Section III-B, Equations (1)–(2)] The description of the autocorrelation function (ACF) is ambiguous: Equation (2) returns a vector over lags, and the paper later refers to the mean, variance, skewness, and kurtosis of the ACF. Please clarify that these four functionals are computed on the ACF vector, and define the set of lags used.
- [Appendix A] The notation σ²(x_n) denotes the variance of frames across the sequence, but x_n is also used for a single frame vector; this should be clarified, for instance by writing σ²_n(x_n) or defining the variance over the frame index n explicitly. The proof would also benefit from explicitly modeling the dependence of y_n on the mask pattern and positional encoding, which is where the current gap arises.
- [Section IV-A] The collapse-detection threshold of 0.01 for the variance of embeddings or outputs is introduced heuristically. Please provide a justification or a reference for this threshold, and report how sensitive the conclusions in Table 3 are to its value.
- [Appendix F, Table 15] The TS2Vec citation in the table caption appears as '[?]' and should be fixed to the proper reference [35].
- [Table 4 and Appendix D] In Table 4, MAE is marked as robust to representation collapse despite one collapse in Table 3. The text attributes this to bad weight initialization, but the table and its caption should explicitly state this qualification so readers are not misled.
- [Section V] The statement that PFML 'outperformed MAE and TS2Vec' is too strong given the small absolute differences in several tasks (e.g., IMU posture 95.7 vs 95.6; IMU movement 81.8 vs 81.0). Consider softening the language or adding statistical support.
Circularity Check
No material circularity: PFML's targets are deterministic functionals and the downstream comparisons are external, though Appendix A's non-collapse proof covers prediction variance rather than encoder embedding variance.
full rationale
PFML's pre-training objective is a regression from masked-context latent representations to deterministic statistical functionals f_n of the input frames; the functionals are precomputed and z-score normalized before training, and the model's predictions y_n are compared with them under MSE/L1 losses (Section III-B, Appendix A, Equations 3-4). The targets are therefore not produced by the model being trained, and no fitted parameter is later relabeled as a prediction. The non-collapse argument is a direct variance-transfer argument: under Assumptions 1 and 2, f_n has variance across frames, so a low-loss solution cannot have constant y_n. That is a mathematical consequence of the loss, not a restatement of the result, and it does not rely on a self-citation or on the data2vec/MAE/TS2Vec comparisons. The central empirical claims are benchmarked against external baselines on held-out classification tasks (Tables 1-3), so the derivation is self-contained with respect to those benchmarks. The main weakness is a proof gap rather than circularity: Appendix A establishes variance of the prediction outputs y_n, whereas the paper's collapse definition concerns feature representations and its operational collapse check tracks both embeddings and outputs (Section IV-A); a model with constant encoder embeddings z_n could in principle still produce position- or mask-dependent y_n through relative positional encoding and the ones-vector mask pattern. This unsupported inference is a correctness risk, not a circular reduction. Self-citations [53], [54], [55], [56] are used for datasets, encoder architecture, and prior subset-selection results; they are ancillary empirical resources and are not load-bearing for the non-collapse theorem or for the comparisons against external methods. No uniqueness theorem is imported from the authors' prior work, and no known result is merely renamed. Accordingly, the paper receives a low score reflecting only minor, non-load-bearing self-citations.
Assumptions & free parameters
free parameters (3)
- Set of 11 statistical functionals
- Masking hyperparameters (pm, lm) =
IMU: 0.15, 3; Speech: 0.065, 10; EEG: 0.1, 3
- Collapse detection threshold =
0.01 (variance) for 10 epochs
assumptions (4)
- domain assumption Assumption 1: temporal variability across frames
- domain assumption Assumption 2: non-trivial functionals computed from frames have variance
- domain assumption Low prediction loss on functionals transfers to useful downstream representations
- ad hoc to paper Optimization does not exploit positional shortcuts to keep embeddings collapsed
Cite this review
Pith. "Pith review of PFML: Self-Supervised Learning of Time-Series Data Without Representation Collapse." pith.science (2026). https://pith.science/paper/BMZJLKBE
@misc{pith2026241110087,
author = {Pith},
title = {Pith review of: PFML: Self-Supervised Learning of Time-Series Data Without Representation Collapse},
year = {2026},
howpublished = {\url{https://pith.science/paper/BMZJLKBE}},
note = {Machine review of arXiv:2411.10087}
}
read the original abstract
Self-supervised learning (SSL) is a data-driven learning approach that utilizes the innate structure of the data to guide the learning process. In contrast to supervised learning, which depends on external labels, SSL utilizes the inherent characteristics of the data to produce its own supervisory signal. However, one frequent issue with SSL methods is representation collapse, where the model outputs a constant input-invariant feature representation. This issue hinders the potential application of SSL methods to new data modalities, as trying to avoid representation collapse wastes researchers' time and effort. This paper introduces a novel SSL algorithm for time-series data called Prediction of Functionals from Masked Latents (PFML). Instead of predicting masked input signals or their latent representations directly, PFML operates by predicting statistical functionals of the input signal corresponding to masked embeddings, given a sequence of unmasked embeddings. The algorithm is designed to avoid representation collapse, rendering it straightforwardly applicable to different time-series data domains, such as novel sensor modalities in clinical data. We demonstrate the effectiveness of PFML through complex, real-life classification tasks across three different data modalities: infant posture and movement classification from multi-sensor inertial measurement unit data, emotion recognition from speech data, and sleep stage classification from EEG data. The results show that PFML is superior to a conceptually similar SSL method and a contrastive learning-based SSL method. Additionally, PFML is on par with the current state-of-the-art SSL method, while also being conceptually simpler and without suffering from representation collapse.
Figures
Reference graph
Works this paper leans on
-
[1]
R. Balestriero, M. Ibrahim, V . Sobal, A. Morcos, S. Shekhar, T. Goldstein, F. Bordes, A. Bardes, G. Mialon, Y . Tian, A. Schwarzschild, A. G. Wilson, J. Geiping, Q. Garrido, P . Fernandez, A. Bar, H. Pirsiavash, Y . LeCun, and M. Goldblum, ‘‘A Cookbook of Self-Supervised Learning,’’ arXiv preprint arXiv: 2304.12210, 2023
arXiv 2023
-
[2]
J. Gui, T. Chen, J. Zhang, Q. Cao, Z. Sun, H. Luo, and D. Tao, ‘‘A Survey on Self-Supervised Learning: Algorithms, Applications, and Future Trends,’’ IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 46, no. 12, pp. 9052–9071, 2024
work page 2024
- [3]
-
[4]
A. van den Oord, Y . Li, and O. Vinyals, ‘‘Representation Learning with Contrastive Predictive Coding,’’ arXiv preprint arXiv: 1807.03748 , 2018
arXiv 2018
-
[5]
A. Baevski, Y . Zhou, A. Mohamed, and M. Auli, ‘‘wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,’’ in Proc. NeurIPS, 2020, pp. 12 449–12 460
work page 2020
-
[6]
T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, ‘‘A simple framework for contrastive learning of visual representations,’’ in Proc. ICML, 2020, pp. 1597–1607
work page 2020
- [7]
-
[8]
A. Baevski, W. Hsu, Q. Xu, A. Babu, J. Gu, and M. Auli, ‘‘data2vec: A General Framework for Self-supervised Learning in Speech, Vision and Language,’’ in Proc. ICML, 2022, pp. 1298–1312
work page 2022
Show all 61 references
-
[9]
W. Wang, H. Bao, L. Dong, J. Bjorck, Z. Peng, Q. Liu, K. Aggarwal, O. K. Mohammed, S. Singhal, S. Som, and F. Wei, ‘‘Image as a Foreign Language: BEIT Pretraining for Vision and Vision-Language Tasks,’’ in Proc. IEEE CVPR, 2023, pp. 19 175–19 186
2023
-
[10]
O. J. Hénaff, A. Srinivas, J. De Fauw, A. Razavi, C. Doersch, S. M. A. Eslami, and A. V an Den Oord, ‘‘Data-efficient image recognition with contrastive predictive coding,’’ in Proc. ICML, 2020, pp. 4182–4192
2020
-
[11]
Baevski, A
A. Baevski, A. Babu, W.-N. Hsu, and M. Auli, ‘‘Efficient self-supervised learning with contextualized target representations for vision, speech and language,’’ in Proc. ICML, 2023, pp. 1416–1429
2023
-
[12]
J. W. Y oon, S. M. Kim, and N. S. Kim, ‘‘MCR-Data2vec 2.0: Improving Self-supervised Speech Pre-training via Model-level Consistency Regular- ization,’’ in Proc. INTERSPEECH, 2023, pp. 2833–2837
2023
-
[13]
Q.-S. Zhu, L. Zhou, J. Zhang, S.-J. Liu, Y .-C. Hu, and L.-R. Dai, ‘‘Robust Data2VEC: Noise-Robust Speech Representation Learning for ASR by Combining Regression and Improved Contrastive Learning,’’ in Proc. IEEE ICASSP, 2023, pp. 1–5
2023
-
[14]
J. Lian, A. Baevski, W.-N. Hsu, and M. Auli, ‘‘Av-Data2V ec: Self- Supervised Learning of Audio-Visual Speech Representations with Contex- tualized Target Representations,’’ in IEEE ASRU, 2023, pp. 1–8
2023
-
[15]
Kalantidis, M
Y . Kalantidis, M. B. Sariyildiz, N. Pion, P . Weinzaepfel, and D. Larlus, ‘‘Hard negative mixing for contrastive learning,’’ in Proc. NeurIPS, 2020, pp. 21 798—-21 809
2020
-
[16]
Robinson, C.-Y
J. Robinson, C.-Y . Chuang, S. Sra, and S. Jegelka, ‘‘Contrastive Learning with Hard Negative Samples,’’ in Proc. ICLR, 2021
2021
-
[17]
Caron, I
M. Caron, I. Misra, J. Mairal, P . Goyal, P . Bojanowski, and A. Joulin, ‘‘Un- supervised learning of visual features by contrasting cluster assignments,’’ in Proc. NeurIPS, 2020, pp. 9912—-9924
2020
-
[18]
W.-N. Hsu, B. Bolte, Y .-H. H. Tsai, K. Lakhotia, R. Salakhutdinov, and A. Mohamed, ‘‘HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,’’ IEEE/ACM Trans. Audio, Speech and Lang. Proc., vol. 29, p. 3451–3460, 2021
2021
-
[19]
T. Hua, W. Wang, Z. Xue, S. Ren, Y . Wang, and H. Zhao, ‘‘On Feature Decorrelation in Self-Supervised Learning,’’ in Proc. IEEE ICCV, 2021, pp. 9578–9588
2021
-
[20]
L. Jing, P . Vincent, Y . LeCun, and Y . Tian, ‘‘Understanding Dimensional Collapse in Contrastive Self-supervised Learning,’’ in Proc. ICLR, 2022
2022
-
[21]
Garrido, R
Q. Garrido, R. Balestriero, L. Najman, and Y . LeCun, ‘‘RankMe: Assessing the Downstream Performance of Pretrained Self-Supervised Representations by Their Rank,’’ in Proc. ICML, 2023, p. 10929–10974
2023
-
[22]
S. Chen, C. Wang, Z. Chen, Y . Wu, S. Liu, Z. Chen, J. Li, N. Kanda, T. Y oshioka, X. Xiao, J. Wu, L. Zhou, S. Ren, Y . Qian, Y . Qian, J. Wu, M. Zeng, X. Y u, and F. Wei, ‘‘WavLM: Large-Scale Self-Supervised Pre- Training for Full Stack Speech Processing,’’IEEE Journal of Sel...
2022
-
[23]
Lee, J.-B
H.-Y . Lee, J.-B. Huang, M. Singh, and M.-H. Y ang, ‘‘Unsupervised Representation Learning by Sorting Sequences,’’ in Proc. ICCV , 2017, pp. 667–676
2017
-
[24]
Gidaris, P
S. Gidaris, P . Singh, and N. Komodakis, ‘‘Unsupervised Representation Learning by Predicting Image Rotations,’’ in Proc. ICLR, 2018
2018
-
[25]
Caron, P
M. Caron, P . Bojanowski, A. Joulin, and M. Douze, ‘‘Deep Clustering for Unsupervised Learning of Visual Features,’’ in Proc. ECCV, 2018, pp. 132—-149
2018
-
[26]
Grill, F
J.-B. Grill, F. Strub, F. Altché, C. Tallec, P . H. Richemond, E. Buchatskaya, C. Doersch, B. A. Pires, Z. D. Guo, M. G. Azar, B. Piot, K. Kavukcuoglu, R. Munos, and M. V alko, ‘‘Bootstrap Y our Own Latent a New Approach to Self-Supervised Learning,’’ in Proc. NeurIPS, 2020, p...
2020
-
[27]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sas- try, A. Askell, P . Mishkin, J. Clark, G. Krueger, and I. Sutskever, ‘‘Learning Transferable Visual Models From Natural Language Supervision,’’ in Proc. ICML, 2021, pp. 8748–8763
2021
-
[28]
K. He, X. Chen, S. Xie, Y . Li, P . Dollár, and R. B. Girshick, ‘‘Masked Autoencoders Are Scalable Vision Learners,’’ in Proc. IEEE CVPR, 2022, pp. 15 979–15 988
2022
-
[29]
H. Bao, L. Dong, S. Piao, and F. Wei, ‘‘BEiT: BERT Pre-Training of Image Transformers,’’ in Proc. ICLR, 2022
2022
-
[30]
Oquab, T
M. Oquab, T. Darcet, T. Moutakanni, H. V . V o, M. Szafraniec, V . Khalidov, P . Fernandez, D. HAZIZA, F. Massa, A. El-Nouby, M. Assran, N. Bal- las, W. Galuba, R. Howes, P .-Y . Huang, S.-W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P . Lab...
2024
-
[31]
Devlin, M
J. Devlin, M. Chang, K. Lee, and K. Toutanova, ‘‘BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,’’ in Proc. NAACL-HLT, 2019, pp. 4171–4186. 10 VOLUME 13, 2025 E. Vaaras et al.: PFML: Self-Supervised Learning of Time-Series Data Without Represe...
2019
-
[32]
Brown, B
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P . Dhariwal, A. Neelakantan, P . Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert- V oss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B....
2020
-
[33]
Y . Tay, M. Dehghani, V . Q. Tran, X. Garcia, J. Wei, X. Wang, H. W. Chung, D. Bahri, T. Schuster, S. Zheng, D. Zhou, N. Houlsby, and D. Metzler, ‘‘UL2: Unifying Language Learning Paradigms,’’ in Proc. ICLR, 2023
2023
-
[34]
OpenAI, ‘‘GPT-4 Technical Report,’’ arXiv preprint arXiv: 2303.08774 , 2023
2023 arXiv
-
[35]
Z. Y ue, Y . Wang, J. Duan, T. Y ang, C. Huang, Y . Tong, and B. Xu, ‘‘TS2V ec: Towards Universal Representation of Time Series,’’ in Proc. AAAI, 2022, pp. 8980–8987
2022
-
[36]
K. He, H. Fan, Y . Wu, S. Xie, and R. Girshick, ‘‘Momentum Contrast for Unsupervised Visual Representation Learning,’’ in Proc. IEEE CVPR, 2020, pp. 9726–9735
2020
-
[37]
Pizzi, S
E. Pizzi, S. D. Roy, S. N. Ravindra, P . Goyal, and M. Douze, ‘‘A Self- Supervised Descriptor for Image Copy Detection,’’ in Proc. IEEE CVPR, 2022, pp. 14 512–14 522
2022
-
[38]
Y . M. Asano, C. Rupprecht, and A. V edaldi, ‘‘Self-labelling via simultaneous clustering and representation learning,’’ in Proc. ICLR, 2020
2020
-
[39]
W. Wang, Q. Tang, and K. Livescu, ‘‘Unsupervised Pre-Training of Bidirectional Speech Encoders via Masked Reconstruction,’’ in Proc. IEEE ICASSP, 2020, pp. 6889–6893
2020
-
[40]
Z. Xie, Z. Zhang, Y . Cao, Y . Lin, J. Bao, Z. Y ao, Q. Dai, and H. Hu, ‘‘SimMIM: a Simple Framework for Masked Image Modeling,’’ in Proc. IEEE CVPR, 2022, pp. 9643–9653
2022
-
[41]
Bardes, J
A. Bardes, J. Ponce, and Y . LeCun, ‘‘VICReg: V ariance-Invariance- Covariance Regularization for Self-Supervised Learning,’’ in Proc. ICLR, 2022
2022
-
[42]
J. H. McDermott and E. P . Simoncelli, ‘‘Sound Texture Perception via Statistics of the Auditory Periphery: Evidence from Sound Synthesis,’’ Neuron, vol. 71, no. 5, pp. 926–940, 2011
2011
-
[43]
L. R. Rabiner and R. W. Schafer, ‘‘Introduction to Digital Speech Process- ing,’’ F oundations and Trends Signal Processing, vol. 1, no. 1-2, pp. 1–194, 2007
2007
-
[44]
Gulati, C.-C
A. Gulati, C.-C. Chiu, J. Qin, J. Y u, N. Parmar, R. Pang, S. Wang, W. Han, Y . Wu, Y . Zhang, and Z. Zhang, ‘‘Conformer: Convolution-augmented Transformer for Speech Recognition,’’ in Proc. INTERSPEECH, 2020, pp. 5036–5040
2020
-
[45]
Hochreiter and J
S. Hochreiter and J. Schmidhuber, ‘‘Long Short-Term Memory,’’ Neural Computation, vol. 9, no. 8, pp. 1735–1780, 1997
1997
-
[46]
K. Cho, B. Merrienboer, C. Gulcehre, F. Bougares, H. Schwenk, and Y . Bengio, ‘‘Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation,’’ in Proc. EMNLP, 2014, pp. 1724– 1734
2014
-
[47]
Hendrycks and K
D. Hendrycks and K. Gimpel, ‘‘Gaussian Error Linear Units (GELUs),’’ arXiv preprint arXiv: 1606.08415 , 2016
2016 arXiv
-
[48]
J. L. Ba, J. R. Kiros, and G. E. Hinton, ‘‘Layer Normalization,’’ arXiv preprint arXiv: 1607.06450, 2016
2016 arXiv
-
[49]
Ulyanov, A
D. Ulyanov, A. V edaldi, and V . Lempitsky, ‘‘Instance Normalization: The Missing Ingredient for Fast Stylization,’’ arXiv preprint arXiv: 1607.08022 , 2016
2016 arXiv
-
[50]
L. Liu, H. Jiang, P . He, W. Chen, X. Liu, J. Gao, and J. Han, ‘‘On the V ariance of the Adaptive Learning Rate and Beyond,’’ in Proc. ICLR, 2020
2020
-
[51]
Loshchilov and F
I. Loshchilov and F. Hutter, ‘‘Decoupled Weight Decay Regularization,’’ in Proc. ICLR, 2019
2019
-
[52]
D. P . Kingma and J. Ba, ‘‘Adam: A Method for Stochastic Optimization,’’ in Proc. ICLR, 2015
2015
-
[53]
Airaksinen, A
M. Airaksinen, A. Gallen, A. Kivi, P . Vijayakrishnan, T. Häyrinen, E. Ilen, O. Räsänen, L. M. Haataja, and S. V anhatalo, ‘‘Intelligent wearable allows out-of-the-lab tracking of developing motor abilities in infants,’’ Communications Medicine, vol. 2, no. 69, 2022
2022
-
[54]
V aaras, M
E. V aaras, M. Airaksinen, S. V anhatalo, and O. Räsänen, ‘‘Evaluation of self-supervised pre-training for automatic infant movement classification using wearable movement sensors,’’ in Proc. IEEE EMBC, 2023, pp. 1–6
2023
-
[55]
Airaksinen, O
M. Airaksinen, O. Räsänen, E. Ilén, T. Häyrinen, A. Kivi, V . Marchi, A. Gallen, S. Blom, A. V arhe, N. Kaartinen, L. Haataja, and S. V anhatalo, ‘‘Automatic Posture and Movement Tracking of Infants with Wearable Movement Sensors,’’ Scientific Reports, vol. 10, no. 169, 2020
2020
-
[56]
V aaras, S
E. V aaras, S. Ahlqvist-Björkroth, K. Drossos, L. Lehtonen, and O. Räsänen, ‘‘Development of a speech emotion recognizer for large-scale child-centered audio recordings from a hospital environment,’’ Speech Communication, vol. 148, pp. 9–22, 2023
2023
-
[57]
B. Kemp, A. Zwinderman, B. Tuk, H. Kamphuisen, and J. Oberye, ‘‘Analysis of a sleep-dependent neuronal feedback loop: the slow-wave microcontinuity of the EEG,’’ IEEE Transactions on Biomedical Engineering , vol. 47, no. 9, pp. 1185–1194, 2000
2000
-
[58]
A. L. Goldberger, L. A. N. Amaral, L. Glass, J. M. Hausdorff, P . C. Ivanov, R. G. Mark, J. E. Mietus, G. B. Moody, C.-K. Peng, and H. E. Stanley, ‘‘PhysioBank, PhysioToolkit, and PhysioNet: Components of a New Research Resource for Complex Physiologic Signals,’’ Circulation, ...
2000
-
[59]
Eldele, Z
E. Eldele, Z. Chen, C. Liu, M. Wu, C.-K. Kwoh, X. Li, and C. Guan, ‘‘An Attention-Based Deep Learning Approach for Sleep Stage Classification With Single-Channel EEG,’’ IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol. 29, pp. 809–818, 2021
2021
-
[60]
Watson, J
D. Watson, J. Krutzinna, I. Bruce, C. Griffiths, I. McInnes, M. Barnes, and L. Floridi, ‘‘Clinical Applications of Machine Learning Algorithms: Beyond the Black Box,’’ BMJ, vol. 364, no. l886, 2019. EINARI VAARAS was born in Finland, in 1996. He received the B.Sc. (Tech.) and ...
2019
-
[1986]
PFML: Self-Supervised Learning of Time-Series Data Without Representation Collapse
He received the M.Sc. (Tech.) and D.Sc. (Tech.) degrees in speech and language technology from Aalto University, Espoo, Finland, in 2012 and 2018, respectively. He is currently a Senior Research Engineer with the BABA Center, Univer- sity of Helsinki, and Helsinki University H...
2012 arXiv
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.