Pith. sign in

Paper Citation Record · LEDGER

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study

As of 10 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2502.02366.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02366 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T12:27:19.487461Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact8
  • verified fuzzy8
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f22b93c6-b266-49fe-8059-01aa1f00697d · outbound

This paper cites PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.256603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.256603Z digest=sha256:67f209325cd87b3d2b6171ea48310b2bd2f0cf62542e6d6530418a7154e71e13

Observation 42358cba-21a7-4453-9d26-22cddc5d4bff · outbound

This paper cites Embeddings for up to 2000 random samples from the validation partition of select datasets representing speech, non-speech and VAD audio domains.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Embeddings for up to 2000 random samples from the validation partition of select datasets representing speech, non-speech and VAD audio domains

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:27:21.159538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.250437Z digest=sha256:97af2c4b966ca6e0feba98f798acb239ffcece7f670f140666179fe88f2f7a1d

Observation e57ecd9a-7554-4c9d-bcd5-aafea976dbb3 · outbound

This paper cites Using State of the Art Speaker Recognition and Natural Language Processing Technologies to Detect Alzheimer’s Disease and Assess its Severity,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Using State of the Art Speaker Recognition and Natural Language Processing Technologies to Detect Alzheimer’s Disease and Assess its Severity,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.271776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.271776Z digest=sha256:9b45e328dbabb27882eee6771b8687164b3c00ee08e6b3ac98681dff0deddedb

Observation dc4165f4-1aa9-41f7-bc33-719a3bbcde75 · outbound

This paper cites wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:27:21.143137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.261659Z digest=sha256:d852a29b28eb461e61ad959aa98b9f485e16be2fdef6acaead637d76478e7d55

Observation d362b1a4-41e5-4904-8248-60c7a41cf417 · outbound

This paper cites Characterizing soundscapes across diverse ecosystems using a universal acoustic feature set,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Characterizing soundscapes across diverse ecosystems using a universal acoustic feature set,

Reference 5

Resolution
verified exact
doi, observed 2026-08-09T12:27:20.051458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.282694Z digest=sha256:9b2ac129e477871f8ea5551ff942676d0079839b8f2e567cfbea065fea1e269d

Observation 26e6edd9-68e2-4b4d-b8e1-86c32ed79e01 · outbound

This paper cites Soundscapes and deep learning enable tracking biodiversity recovery in tropical forests,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Soundscapes and deep learning enable tracking biodiversity recovery in tropical forests,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.287723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.287723Z digest=sha256:9227fe256f329434d0282d8f61ffea92367d371b380a4a7e49fc0beeada31f3f

Observation 0ed909f8-a55a-4391-8c4e-b41ddafeec3d · outbound

This paper cites Using X-Vectors to Automatically Detect Parkinson’s Disease from Speech,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Using X-Vectors to Automatically Detect Parkinson’s Disease from Speech,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.277150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.277150Z digest=sha256:cee8057ee58c9905012834ce2a84bd8ef036154d9a011a949e95fcf7d3676993

Observation 249558b5-4a6a-4e15-af5e-751f385b1726 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Representation Learning with Contrastive Predictive Coding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.297365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.297365Z digest=sha256:308b9ddcf7942ec229e08e6be48982d5836c4aa0aa279966b886bedc2ff7629a

Observation 53867186-9827-4bb7-a66c-f5990b2fb264 · outbound

This paper cites Unsupervised Cross-lingual Representation Learning for Speech Recognition.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Unsupervised Cross-lingual Representation Learning for Speech Recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.302520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.302520Z digest=sha256:8958ebdf41c37585dde72b7f28c13217f607f8b3509fc1663a0fa28d7b422c7f

Observation 6293db8b-5677-47ad-92ae-2d5f5b9d0681 · outbound

This paper cites Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Robust wav2vec 2.0: Analyzing Domain Shift in Self-Supervised Pre-Training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.292513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.292513Z digest=sha256:8b15c4cc61f4d62e2756e932daa93110c9522e6d55b1b772102f7f5c7e5a330f

Observation 7cf3e73f-78d9-450d-b054-e0a62a91c99a · outbound

This paper cites BYOL for Audio: Exploring Pre-Trained General-Purpose Audio Representations,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study BYOL for Audio: Exploring Pre-Trained General-Purpose Audio Representations,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.312022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.312022Z digest=sha256:dcabce0216351b56ec51dc1d05350c8a3b11510d24e1b0f996ac54f2c14059d3

Observation fd418345-df05-42d9-9378-9c66edf65639 · outbound

This paper cites BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study BYOL for Audio: Self-Supervised Learning for General-Purpose Audio Representation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.316685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.316685Z digest=sha256:4a16afbbba1602c7bad97072e9f91d8a2e6146b0ba8458dccc8da46b557bcb37

Observation ff44cd45-6874-4b15-ba76-f18e5df4f0e0 · outbound

This paper cites WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,

Reference 13

Resolution
verified exact
doi, observed 2026-08-09T12:27:19.973418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.307198Z digest=sha256:ddc9eba021b23446db2b4abada62b9424796c1cbba14b401df04777a2d361cd2

Observation 1f9be9d3-9ac4-4bfe-8c11-f17534111143 · outbound

This paper cites The fifth 'CHiME' Speech Separation and Recognition Challenge: Dataset, task and baselines.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study The fifth 'CHiME' Speech Separation and Recognition Challenge: Dataset, task and baselines

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.333240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.333240Z digest=sha256:71e4e848d435ba640059089863882a0e129a708dc372fd2f09d654c8e986f472

Observation e7bbcfef-e100-4020-833b-67ce9bf0711c · outbound

This paper cites Unleashing the killer corpus: experiences in creating the multi-everything AMI Meeting Corpus,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Unleashing the killer corpus: experiences in creating the multi-everything AMI Meeting Corpus,

Reference 15

Resolution
verified exact
doi, observed 2026-08-09T12:27:19.857230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.339153Z digest=sha256:56e507059e868143ea25a1fa4dc8d8e38ee40fdba061c4df8fd119ac1f245dac

Observation ff13feae-5940-4199-94fc-d8c14ac8c2c5 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Common Voice: A Massively-Multilingual Speech Corpus,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:27:21.108393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.322091Z digest=sha256:a74dc3e89e04ca3152523cea9e04da67bdbd1d895fdacf32c33187c974ee9db4

Observation 0b8535c4-248a-4dd8-b80e-e720f53b3633 · outbound

This paper cites Available: https://aclanthology.org/2020.lrec-1.520.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Available: https://aclanthology.org/2020.lrec-1.520

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:27:21.092504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.326958Z digest=sha256:96d7f355160ef7ae7c2ab90e1d5acc0a48f7c0adf3b3d77b63489f3d00454894

Observation 520a4489-9a03-4f98-a9f0-79822b649030 · outbound

This paper cites Recognition and understanding of meetings the AMI and AMIDA projects,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Recognition and understanding of meetings the AMI and AMIDA projects,

Reference 18

Resolution
metadata mismatch
raw_fallback, observed 2026-08-09T12:27:20.657531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.353246Z digest=sha256:4ec9543f6381772cc90b3aae592512158a8053cf268a4185cb83198662da3968

Observation 74dc6609-6f2d-4697-a419-2495e7c30850 · outbound

This paper cites Enhancing the TED-LIUM Corpus with Selected Data for Language Modeling and More TED Talks,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Enhancing the TED-LIUM Corpus with Selected Data for Language Modeling and More TED Talks,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:27:21.075801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.358039Z digest=sha256:6a4bb6788a1a993f267868149780b581f1fc7ed338037e101403be7368eafb14

Observation 314f1b96-be1f-498b-bb11-68e5e25b1745 · outbound

This paper cites VoxCeleb: A Large-Scale Speaker Identification Dataset,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study VoxCeleb: A Large-Scale Speaker Identification Dataset,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.343966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.343966Z digest=sha256:03023d6360f3e7afeb56c385b0591d3030019936d9cb43bf6990e6d8dcd5ddda

Observation b2a82536-1475-4947-9d6a-5802f59c35bd · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Librispeech: An ASR corpus based on public domain audio books,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.348655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.348655Z digest=sha256:b8df70dc2128da941b29161157a9ac50ec8de9e2894fe5af2ff27080a5cd8895

Observation c0bbf53d-9141-497d-9450-cb929aabf208 · outbound

This paper cites General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study General-purpose Tagging of Freesound Audio with AudioSet Labels: Task Description, Dataset, and Baseline

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.373247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.373247Z digest=sha256:af059f8fc8dbec5f3a5ae411587ce71ff904d87bf9885d30e628f8dc42c2e02e

Observation 86254a16-ab1c-4f01-99ba-f180d0b9b32a · outbound

This paper cites Audio tagging with noisy labels and minimal supervision.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Audio tagging with noisy labels and minimal supervision

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-09T12:27:19.765128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.378879Z digest=sha256:1cf24a1e13bd794898d948cd58ab6fb0ba1410038e751c607983958f5d6d1816

Observation 36c1e223-ea52-4c08-8b1e-ae6e318c66bd · outbound

This paper cites The Multilingual TEDx Corpus for Speech Recognition and Translation.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study The Multilingual TEDx Corpus for Speech Recognition and Translation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.362872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.362872Z digest=sha256:62fa4022913207931dee1551b9fbc6f964b55887e58b90060d285cbe86afe858

Observation 2d3bd654-f1e8-4874-a971-b446dc5715f2 · outbound

This paper cites SONYC-UST-V2: An Urban Sound Tagging Dataset with Spatiotemporal Context.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study SONYC-UST-V2: An Urban Sound Tagging Dataset with Spatiotemporal Context

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.367959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.367959Z digest=sha256:5385881788baa77cecab2b13b368280c8d1aebb2d03b8505490c4c892e8a7089

Observation c54c481e-8447-4f99-aa94-9734d26c196a · outbound

This paper cites MUSAN: A Music, Speech, and Noise Corpus.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study MUSAN: A Music, Speech, and Noise Corpus

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.393622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.393622Z digest=sha256:3578d50c841c13e64c12e8bf3b230f2789785f12c14ed8e44e30d952a445d7e9

Observation 14668544-0e06-48e2-a15d-8450110aadbb · outbound

This paper cites An open dataset for research on audio field recording archives: freefield1010.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study An open dataset for research on audio field recording archives: freefield1010

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-09T12:27:19.720639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.399285Z digest=sha256:b34cb5d648b9ac33307908a48dd3d33fd00f6679bb677fee5762d2935d2c5187

Observation 110d69d4-f231-4a14-aa5d-bfa7b7d19866 · outbound

This paper cites FSD50K: An Open Dataset of Human-Labeled Sound Events,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study FSD50K: An Open Dataset of Human-Labeled Sound Events,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.384159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.384159Z digest=sha256:c34b829585b593034649ad3b48dab8a985ba552c20cc4c4897043ee5d13de3e3

Observation 1b64d332-ca61-423a-9c36-6ce02c484be4 · outbound

This paper cites an unresolved cited work.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-09T12:27:21.176477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.244016Z digest=sha256:af687e195e6b334989d78e5ee9df637982aa84423906f550c22eee22b50b05c9

Observation 9b592abc-a6af-4bd6-a354-5a2856c81cb5 · outbound

This paper cites Audio Set: An ontology and human-labeled dataset for audio events,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Audio Set: An ontology and human-labeled dataset for audio events,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.388827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.388827Z digest=sha256:23096330a743d7f29f06f4230052ac784d13cbfddd411024a1f464ed4601c4a4

Observation 20a62922-1eda-4a1f-ae5d-c67e70a5b59f · outbound

This paper cites CNN architectures for large-scale audio classification,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study CNN architectures for large-scale audio classification,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.418526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.418526Z digest=sha256:bb473df1821a61f8b55826cee3d74d49deb3735a18eb1cf776b6e00b45b339e1

Observation fb81c49d-4cbf-4b1e-9247-254b030e661e · outbound

This paper cites WHAM!: Extending Speech Separation to Noisy Environments.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study WHAM!: Extending Speech Separation to Noisy Environments

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.404632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.404632Z digest=sha256:0cca9f39c78b95001f1e9df514544e5f34cc430ffff9435b248509e00e3f7110

Observation 8433630e-532e-4569-8d87-47c260851edb · outbound

This paper cites Bootstrap your own latent a new approach to self-supervised learning,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Bootstrap your own latent a new approach to self-supervised learning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:27:21.057812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.409636Z digest=sha256:36e6dde41435a69d84afb0dede95fc5beb7737961df0bac0251698b74ecda848

Observation ec4dd20b-a7e9-4b7e-95c4-6c7677caec9e · outbound

This paper cites MatchboxNet: 1D Time-Channel Separable Convolutional Neural Network Architecture for Speech Commands Recognition,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study MatchboxNet: 1D Time-Channel Separable Convolutional Neural Network Architecture for Speech Commands Recognition,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.413855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.413855Z digest=sha256:f11880979854f3ede0bb1738f460951d704e6b1a3d7794f9a2810eee95474a23

Observation 074a7df4-19ee-43f9-94eb-87bcdb907136 · outbound

This paper cites Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Speech Commands: A Dataset for Limited-Vocabulary Speech Recognition

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.436477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.436477Z digest=sha256:8d04d0efe3cedfb32f486c41a348e970bd6987438b796c7fbaa48acd9d0e3db7

Observation acd3b149-c1c7-4a2a-ab38-61d016f7b473 · outbound

This paper cites Environmental sound classification with convolutional neural networks,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Environmental sound classification with convolutional neural networks,

Reference 36

Resolution
metadata mismatch
raw_fallback, observed 2026-08-09T12:27:20.321447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.422945Z digest=sha256:1107783db6b2285ae0ecc5fbfd5621e57559621ff87dc81c88a1043dbf0f1f72

Observation fe80b230-89ce-445f-9f8f-7be5761927af · outbound

This paper cites A Dataset and Taxonomy for Urban Sound Research,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study A Dataset and Taxonomy for Urban Sound Research,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.427052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.427052Z digest=sha256:67815e967ca40bb1d172e42b2725105192a9284d5f0c4d2b3c03a11a34103924

Observation 0ca4c5ec-044d-4176-8293-d70414a0651d · outbound

This paper cites Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Neural Audio Synthesis of Musical Notes with WaveNet Autoencoders,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:27:21.040629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.431944Z digest=sha256:ab8d0a540eb739ddca5b5cbe547e6022707edf6bac26dfe01703af1fe25ba8dc

Observation 53af1c20-6702-4b24-af27-330b3785cc1b · outbound

This paper cites CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit (version 0.92),.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR Voice Cloning Toolkit (version 0.92),

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.441197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.441197Z digest=sha256:16be64e56cae4fa58f6ba1555d96d46ca2396c1208191a1656501778f16246ed

Observation 4c5150ae-374a-4d78-a102-d8f1a92e5417 · outbound

This paper cites AVA-Speech: A Densely Labeled Dataset of Speech Activity in Movies.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study AVA-Speech: A Densely Labeled Dataset of Speech Activity in Movies

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-09T12:27:19.631797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.446594Z digest=sha256:c22be643196963196d5bbaac2aa7364055241787e856c2c30731fd58cb1c8234

Observation 956c6d5a-f35a-423e-912e-b1bfa9ea51ad · outbound

This paper cites Representational geometry: integrating cognition, computation, and the brain,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Representational geometry: integrating cognition, computation, and the brain,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.451806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.451806Z digest=sha256:519f0fffb83eba4140cc4eb0ffc0e5ce116005e988bfcf06b7f38dfe8ea5b491

Observation 9a4ce4d0-467c-4015-a675-305a16683bf3 · outbound

This paper cites The Timbre Toolbox: Extracting audio descriptors from musical signals,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study The Timbre Toolbox: Extracting audio descriptors from musical signals,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.456940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.456940Z digest=sha256:b4eea6bb19bbb6eeceee1d195b2a500e09b48a1db0886dae1680c4c09f77302d

Observation d43fedd3-434b-4923-a10c-709ec2a3479b · outbound

This paper cites The Modulation Transfer Function for Speech Intelligibility,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study The Modulation Transfer Function for Speech Intelligibility,

Reference 44

Resolution
verified exact
doi, observed 2026-08-09T12:27:19.586295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.462136Z digest=sha256:d494db74886f16cd1ea2cb455bf21f8c0c0229a8fe2e5164e6eb35c746cac397

Observation 8992706d-8d5d-49cc-bf50-e2ddb1370209 · outbound

This paper cites YIN, a fundamental frequency estimator for speech and music,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study YIN, a fundamental frequency estimator for speech and music,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.467070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.467070Z digest=sha256:0b7c27defccc9bffa12b5ea497be1ae771edd756e9f81b9df27f9f8af30f5ab1

Observation 6fdc6041-464b-43b8-8394-9cc7d6ab7f27 · outbound

This paper cites Acoustic Event Detection Using Speaker Recognition Techniques: Model Optimization and Explainable Features,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Acoustic Event Detection Using Speaker Recognition Techniques: Model Optimization and Explainable Features,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.472186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.472186Z digest=sha256:8e0a15a2c5fbbc74d7b4e036fbd8b20c64bae54ee53f8e52c61fca4597934a16

Observation 4323c434-8026-43df-94e6-1aba819f9913 · outbound

This paper cites Acoustic Correlates of Auditory Object and Event Perception: Speakers, Musical Timbres, and Environmental Sounds,.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Acoustic Correlates of Auditory Object and Event Perception: Speakers, Musical Timbres, and Environmental Sounds,

Reference 47

Resolution
verified exact
raw_fallback, observed 2026-08-09T12:27:20.150848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.477019Z digest=sha256:4457b0925f35819d12d5fe673e39da25d04d7644c06eb3e93a2f0ab6ee2017d2

Observation aca9b87a-2b79-439c-849e-90e5273d74cb · outbound

This paper cites AVES: Animal Vocalization Encoder based on Self-Supervision.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study AVES: Animal Vocalization Encoder based on Self-Supervision

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.482197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.482197Z digest=sha256:56e8eb4512e79871abce6c2308e7f53f6584bf7651cb09cce99ca270a4a243d0

Observation 1d76a4ec-b7cf-4bbe-82ce-6e655527bbe9 · outbound

This paper cites SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T12:27:19.487461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T12:27:19.487461Z digest=sha256:dce9297ad0bd91039db86137f411d689291c658e30f71963d5e25085d22924f1

Observation a6dbf56d-3d40-4535-a098-7f095a33addc · outbound

This paper cites Available: https://proceedings.neurips.cc/paper/2020/hash/92d1e1eb1cd6f9fba3227870bb6d7f07-Abstract.html.

Self-Supervised Convolutional Audio Models are Flexible Acoustic Feature Learners: A Domain Specificity and Transfer-Learning Study Available: https://proceedings.neurips.cc/paper/2020/hash/92d1e1eb1cd6f9fba3227870bb6d7f07-Abstract.html

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T12:27:21.126288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T12:27:19.266785Z digest=sha256:97428c705b307639afd87ce124a806d0d260caecd17889d7a2742ebe715d20ee

Pith citing papers

No inbound Pith citation observations are available.