A 22M-parameter encoder pre-trained on audio and music achieves state-of-the-art zero-shot KNN performance across 19 industrial signal datasets spanning four modalities by modeling spectrogram sub-bands rather than fixed-rate inputs.
wav2vec 2.0: A framework for self-supervised learning of speech representations,
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
citation-role summary
background 1
citation-polarity summary
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
background 1representative citing papers
citing papers explorer
-
FISHER: A Foundation Model for Multi-Modal Industrial Signal Comprehensive Representation
A 22M-parameter encoder pre-trained on audio and music achieves state-of-the-art zero-shot KNN performance across 19 industrial signal datasets spanning four modalities by modeling spectrogram sub-bands rather than fixed-rate inputs.