REVIEW 2 major objections 2 minor 43 references
Cross Domain Few-Shot Class-Incremental Audio Classification Via Adversarial Contrastive Learning
T0 review · 2 major / 2 minor · reviewed 2026-07-03 · grok-4.3
Pith's one-line read Adversarial contrastive training lets a frozen encoder and updated classifier classify new audio classes from shifted domains.
desk verdict The paper introduces cross-domain few-shot class-incremental audio classification and claims gains from adversarial contrastive training on a frozen encoder, but the lack of domain adaptation at the feature level raises questions about whether it truly handles the shift. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Adversarial contrastive training applied to an encoder-classifier model that freezes the encoder after base-session training.
What would settle it
If the method recorded lower average accuracy than state-of-the-art approaches across the six pairs of cross-domain datasets, the central performance claim would be falsified.
Extended reading notes
Core claim
The paper claims that a strategy of adversarial contrastive training enables the model to effectively classify samples of different classes from unseen domains in cross-domain few-shot class-incremental audio classification, where the encoder is trained in the base session but frozen in incremental sessions and the classifier is trained in all sessions, exceeding state-of-the-art methods in average accuracy on six pairs of cross-domain datasets.
Load-bearing premise
The domain shift between base and incremental class samples can be adequately addressed by adversarial contrastive training with a frozen encoder after the base session and an updated classifier, without needing explicit domain adaptation modules.
Editorial extensions
If this is right
- The classifier can adapt to new classes from different domains while the encoder remains fixed after base training.
- Average accuracy improves over prior methods on multiple source-to-target domain pairs without extra adaptation components.
- The approach applies directly to incremental audio tasks where base and new samples follow different distributions.
- Training the classifier in every session while freezing the encoder reduces the risk of overwriting earlier knowledge.
Reading between the lines
- The method may lower deployment cost by avoiding full retraining when new domains appear.
- Success across six dataset pairs suggests the training objectives produce features that transfer across common audio domain shifts.
- The frozen-encoder design could be tested on incremental tasks in other sensor modalities that exhibit similar distribution changes.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the Cross Domain Few-Shot Class-Incremental Audio Classification (FCAC) problem, where base and incremental classes exhibit domain shift. It proposes an adversarial contrastive learning strategy consisting of an encoder trained only on base-session data (then frozen) and a classifier updated across all sessions. The method is evaluated on six pairs of cross-domain audio datasets and reported to exceed state-of-the-art average accuracy; code is released at https://github.com/YongjieSi/ACL.
Significance. If the empirical superiority holds under detailed scrutiny, the work would address a practically relevant gap in audio classification by relaxing the same-domain assumption common in FCAC. Releasing code supports reproducibility and is a positive contribution.
major comments (2)
- [Method] Method section (architecture description): the encoder is trained solely on base data and frozen thereafter, with only the classifier updated in incremental sessions. No explicit domain-adversarial loss, feature alignment term, or encoder fine-tuning is described to mitigate distribution mismatch; the contrastive component therefore operates on potentially misaligned frozen representations. This assumption is load-bearing for the central claim that the strategy handles cross-domain shift.
- [Experiments] Experiments section: the abstract and results claim outperformance on six dataset pairs, yet no specific baselines, metrics (e.g., per-session accuracy with standard deviation), number of runs, statistical significance tests, or quantification of domain shift (e.g., via MMD or classifier accuracy on domain labels) are referenced. Without these, it is impossible to verify whether the reported average accuracy improvement is robust or merely reflects weak baselines.
minor comments (2)
- [Method] Notation for the adversarial contrastive loss should be formalized with an equation; the current prose description leaves unclear whether an adversarial objective is present or whether the term is used descriptively.
- [Experiments] The six dataset pairs should be explicitly listed with their domain characteristics (e.g., recording conditions, sampling rates) in a table to allow readers to assess the severity of the shifts.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback on our manuscript. We address each major comment below and will make revisions to improve clarity and completeness.
read point-by-point responses
-
Referee: [Method] Method section (architecture description): the encoder is trained solely on base data and frozen thereafter, with only the classifier updated in incremental sessions. No explicit domain-adversarial loss, feature alignment term, or encoder fine-tuning is described to mitigate distribution mismatch; the contrastive component therefore operates on potentially misaligned frozen representations. This assumption is load-bearing for the central claim that the strategy handles cross-domain shift.
Authors: The proposed adversarial contrastive training is the core mechanism intended to promote robustness to domain shift even with a frozen encoder, by structuring the contrastive objective during base-session training to support generalization to incremental domains. We acknowledge that the current description does not explicitly detail an additional domain-adversarial loss term. To address this, we will revise the method section to more clearly articulate how the adversarial contrastive strategy mitigates the distribution mismatch without requiring encoder updates or explicit alignment losses in incremental sessions. revision: yes
-
Referee: [Experiments] Experiments section: the abstract and results claim outperformance on six dataset pairs, yet no specific baselines, metrics (e.g., per-session accuracy with standard deviation), number of runs, statistical significance tests, or quantification of domain shift (e.g., via MMD or classifier accuracy on domain labels) are referenced. Without these, it is impossible to verify whether the reported average accuracy improvement is robust or merely reflects weak baselines.
Authors: We agree that the experimental reporting requires additional detail to substantiate the claims. In the revised manuscript, we will expand the experiments section to specify the baselines, report per-session accuracies with standard deviations over the number of runs performed, include statistical significance tests, and add quantification of domain shift (e.g., via MMD). revision: yes
Circularity Check
No circularity: empirical method proposal without derivations or self-referential predictions
full rationale
The paper presents an empirical ML method for cross-domain few-shot class-incremental audio classification using adversarial contrastive training. The architecture (encoder trained then frozen after base session, classifier updated across sessions) is described and evaluated directly on six pairs of cross-domain datasets, with results compared to SOTA. No equations, first-principles derivations, fitted parameters renamed as predictions, or load-bearing self-citations appear in the provided text. The central claim rests on experimental accuracy improvements rather than any chain that reduces to its own inputs by construction, rendering the work self-contained.
Assumptions & free parameters
Cite this review
Pith. "Pith review of Cross Domain Few-Shot Class-Incremental Audio Classification Via Adversarial Contrastive Learning." pith.science (2026). https://pith.science/paper/YGH5B2VI
@misc{pith2026260702254,
author = {Pith},
title = {Pith review of: Cross Domain Few-Shot Class-Incremental Audio Classification Via Adversarial Contrastive Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/YGH5B2VI}},
note = {Machine review of arXiv:2607.02254}
}
read the original abstract
Current Few-shot Class-incremental Audio Classification (FCAC) methods assume that samples of base and incremental classes are in the same domain (following the same distribution). However, there is generally a domain shift between the above two types of samples. In this paper, we explore the problem of Cross Domain FCAC where samples of base and incremental classes have domain shift. We propose a strategy of adversarial contrastive training which enables the model to effectively classify samples of different classes from unseen domains. The model consists of an encoder and a classifier. The encoder is trained in base session but frozen in incremental sessions, whereas the classifier is trained in all sessions. Experiments are done on six pairs of cross-domain datasets. Results show that our method exceeds state-of-the-art methods in average accuracy. The code is at https://github.com/YongjieSi/ACL.
Figures
Reference graph
Works this paper leans on
-
[1]
Introduction Audio classification (AC) is a task to identify different audio classes. It has widespread app lications, such as wildlife protection [1], road surveillance [2], classification of acoust ic scene [3], video analysis [4] and speaker analysis [5]. Many AC methods [6]-[9] require abundant samples of predefined classes for model training. The tra...
-
[2]
We explore a new problem of CD-FCAC where the domain shift and class increment coexist
-
[3]
We propose an adversarial contrastive training strategy which is a fusion of adversarial training and contrastive learning
-
[4]
We propose a CD-FCAC me thod which exceeds state- of-the-art methods in average accuracy
-
[5]
Method 2.1. Problem Definition The CD-FCAC consists of M sessions, including one base session (session 0) and ( M-1) incremental sessions (sessions 1 to (M-1)). we consider a user-friendly setting where the samples * Corresponding author: Yanxiong Li (eeyxli@scut.edu.cn). in base session are from a single source domain, while the samples in incremental se...
-
[6]
Select audio samples from 𝑫 ௧and extract its log-Mel spectrogram (𝑿௦,𝒀 ௦)
-
[7]
While maximizing training do 4 . 𝑿௧←Spectral disturbance on 𝑿௦
-
[8]
𝑿௧←𝑿 ௧+ ζ𝑚𝑎𝑥∇𝑿(ℓ((𝑿௧,𝒀 ௦); 𝜃) − 𝛾𝑑((𝑿௧,𝒀 ௦), (𝑿௦,𝒀 ௦)))
F o r R epochs do 6 . 𝑿௧←𝑿 ௧+ ζ𝑚𝑎𝑥∇𝑿(ℓ((𝑿௧,𝒀 ௦); 𝜃) − 𝛾𝑑((𝑿௧,𝒀 ௦), (𝑿௦,𝒀 ௦)))
Show all 43 references
-
[9]
𝑿= [𝑿௦,𝑿 ௧], 𝒀= [𝒀௦,𝒀 ௦]
-
[10]
𝜃 ← 𝜃 − ζ ∇ఏℓ(𝑿,𝒀 ;𝜃 )
-
[11]
Experimental Datasets Table 2: Detailed information of the LS-100/NSynth -100/FSC-89
Experiments 3.1. Experimental Datasets Table 2: Detailed information of the LS-100/NSynth -100/FSC-89. Param D0 Dm (1 ≤ m ≤ (M-1)) 𝑫 ௧ 𝑫 ௧ 𝑫 ௧ 𝑫 ௧ #C 60/55/59 60/55/59 40/45/30 40/45/30 L(h) 16.66/12.23 /13.11 3.33/1.52 /3.28 11.11/5.00 /4.17 2.22/5.00 /1.67 #S/C 500/2...
-
[12]
Based on the description of our method and results, we can draw two conclusions
Conclusions We solved the CD-FCAC problem via the adversarial training and supervised contrastive loss. Based on the description of our method and results, we can draw two conclusions. First, our method exceeds state-of-the-art methods in average accuracy. Second, the proposed...
-
[13]
Acknowledgements This work was supported by the national natural science foundation of China (62371195, 62111530145, 61771200), the exchange project of the 10th Meeting of the China-Croatia Science and Technology Cooperation Committee (No. 10-34)
-
[14]
The authors take full responsibility for the content
Generative AI Use Disclosure No generative AI tools (e.g., LLM/ChatGPT) were used to produce any part of this work, except that a limited amount of grammar/spelling checking may have been performed on the final text using standard to ols. The authors take full responsibility f...
-
[15]
Comparison of feature extraction methods for sound-based classification of honey bee activity,
A. Terni, N. Ortolani, I. Nolasco, E. Bonitos, and S. Cecchi, “Comparison of feature extraction methods for sound-based classification of honey bee activity,” IEEE/ACM TASLP, vol. 30, pp. 112-122, 2022
2022
-
[16]
Anomalous sou nd detection using deep audio repr esentation and a BLSTM network for audio surveillance of roads,
Y. Li, X. Li, Y. Zhang, M. Liu, and W. Wang, “Anomalous sou nd detection using deep audio repr esentation and a BLSTM network for audio surveillance of roads,” IEEE Access, vol. 6, pp. 58043- 58055, 2018
2018
-
[17]
Low-compl exity acoustic scene classification us ing parallel attention-convolut ion network,
Y. Li, J. Tan, G. Chen, J. Li, Y. Si, and Q. He, “Low-compl exity acoustic scene classification us ing parallel attention-convolut ion network,” in Proc. of Interspeech, 2024, pp. 1-5
2024
-
[18]
Detecting video anomalies by jointly util izing appearance and skeleton information,
W. Pang, et al., “Detecting video anomalies by jointly util izing appearance and skeleton information,” Expert Syst. Appl., vol. 246, pp. 1-12, 2024
2024
-
[19]
Speaker clust ering by co-optimizing deep representation learning and cluster estimation,
Y. Li, W. Wang, M. Liu, Z. Jiang, and Q. He, “Speaker clust ering by co-optimizing deep representation learning and cluster estimation,” IEEE TMM, vol. 23, pp. 3377-3387, 2021
2021
-
[20]
Learnable counterfactual attention for music classification,
Y.-X. Lin, J.-C. Lin, W.-L. Wei and J.-C. Wang, “Learnable counterfactual attention for music classification,” IEEE TASLP, vol. 33, pp. 570-585, 2025
2025
-
[21]
Classification of urban sound using sequential convolutional neural network model and its visualisation,
M. Agarwal, K. S. Gill, S. Chattopadhyay and M. Singh, “Classification of urban sound using sequential convolutional neural network model and its visualisation,” in Proc. of ICITEICSI, 2024, pp. 1-5. [ 8 ] K . M . L i m , C . P . L e e , Z . Y . L e e , J . Y . L i m a n d J ....
2024
-
[22]
Enhancing conformer-based sound event detection using frequency dynamic convolutions a nd beats audio embeddings,
S. Barahona, D. de Benito-Gorrón, D. T. Toledano and D. Ram os, “Enhancing conformer-based sound event detection using frequency dynamic convolutions a nd beats audio embeddings,” IEEE/ACM TASLP, vol. 32, pp. 3896-3907, 2024
2024
-
[23]
Few shot continual learning for audio classification,
Y. Wang, N. J. Bryan, M. Cartwright, J. P. Bello, and J. Salomon, “Few shot continual learning for audio classification,” in Proc. of IEEE ICASSP, 2021, pp. 321-325
2021
-
[24]
Few-shot class- incremental audio classifica tion using dynamically expanded classifier with self-attention modified prototypes,
Y. Li, W. Cao, W. Xian, J. Li, and E. Benito’s, “Few-shot class- incremental audio classifica tion using dynamically expanded classifier with self-attention modified prototypes,” IEEE TMM , vol.26, pp.1346-1360, 2024
2024
-
[25]
Few-shot class-incremental audio classification via discriminative prototype learning,
W. Xie, et al., “Few-shot class-incremental audio classification via discriminative prototype learning,” Expert Syst. Appl. , vol. 225, 2023, Art. no. 120044
2023
-
[26]
Few-shot c lass-incremental audio classifi cation using adaptively-refined prototypes,
W. Xie, et al., “Few-shot c lass-incremental audio classifi cation using adaptively-refined prototypes,” in Proc. of Interspeech, 2023, pp. 301-305. [ 1 4 ] Y . L i , W . C a o , J . L i , W . X i e , a n d Q . H e , “ F e w - s h o t c l a s s - incremental audio classificati...
2023
-
[27]
Few-shot class-incremental audio classification with adap tive mitigation of forgetting and overfitting,
Y. Li, J. Li, Y. Si, J. Tan, and Q. He, “Few-shot class-incremental audio classification with adap tive mitigation of forgetting and overfitting,” IEEE/ACM TASLP, vol. 32, pp. 2297-2311, 2024
2024
-
[28]
Fully few-shot class-incremental audio classifi cation with adaptive improvemen t of stability and plasticity,
Y. Si, Y. Li, J. Tan, G. Chen, Q. Li, and M. Russo, “Fully few-shot class-incremental audio classifi cation with adaptive improvemen t of stability and plasticity,” IEEE TASLP, vol. 33, pp. 418-433, 2025
2025
-
[29]
Fully few-shot class- incremental audio classifica tion using multi-level embedding extractor and ridge regression classifier,
Y. Si, Y. Li, J. Tan, Q. He, and I. Kwak, “Fully few-shot class- incremental audio classifica tion using multi-level embedding extractor and ridge regression classifier,” in Proc. of Interspeech , 2025, pp. 1318-1322
2025
-
[30]
Supervised contrastive learning,
P. Khosla, P. Teterwak, C. Wang, A. Sarna, Y. Tian, P. Iso la, A. Maschinot, C. Liu, and D. Krishnan, “Supervised contrastive learning,” in Proc. of NIPS, 2020. pp. 18661-18673
2020
-
[31]
Certifying some distributional robustness with p rincipled adversarial training,
A. Sinha, H. Namkoong, and J. Duchi, “Certifying some distributional robustness with p rincipled adversarial training, ” in Proc. of International Conference on Learning Representations , 2018, pp. 1-34
2018
-
[32]
Generalizing to unseen domains via adversarial da ta augmentation,
R. Volpi, H. Namkoong, O. Sener, and J.C. Duchi, V. Murino, and S. Savarese, “Generalizing to unseen domains via adversarial da ta augmentation,” in Proc. of Advances in Neural Information Processing Systems, 2018, pp. 1-11
2018
-
[33]
Adversaria l feature augmentation for cros s- domain few-shot classification,
Y. Hu, A.J. Ma, “Adversaria l feature augmentation for cros s- domain few-shot classification,” in Proc. of ECCV, 2022, pp. 20-37
2022
-
[34]
An adversarial meta-training framewo rk for cross-domain few-shot learning,
P. Tian, and S. Xie, “An adversarial meta-training framewo rk for cross-domain few-shot learning," IEEE Transactions on Multimedia, vol. 25, pp. 6881-6891, 2023
2023
-
[35]
Network randomizatio n: A simple technique for generalization in deep reinforcement learning,
K. Lee, K. Lee, J. Shin, and H. Lee, “Network randomizatio n: A simple technique for generalization in deep reinforcement learning,” in Proc. of ICLR, 2020, pp. 1-22
2020
-
[36]
Understanding the difficulty of training deep feedforward neur al networks,
X. Glorot, and Y. Bengio, “Understanding the difficulty of training deep feedforward neur al networks,” in Proc. of the thirteenth international conference on artific ial intelligence and statistics , 2010, pp. 249-256
2010
-
[37]
Cross-domain Few-shot Learnin g with Task-specific Adapters,
W. Li, X. Liu and H. Bilen, "Cross-domain Few-shot Learnin g with Task-specific Adapters," in Proc. of CVPR , 2022, pp. 7151- 7160,
2022
-
[38]
Revisiti ng Prototypical Network for Cross Domain Few-Shot Learning,
F. Zhou, P. Wang, L. Zhang, W. Wei and Y. Zhang, "Revisiti ng Prototypical Network for Cross Domain Few-Shot Learning," in Proc. Of CVPR, 2023, pp. 20061-20070
2023
-
[39]
F ew- shot incremental learning with continually evolved classifiers,
C. Zhang, N. Song, G. Lin, Y. Zheng, P. Pan, and Y. Xu, “F ew- shot incremental learning with continually evolved classifiers, ” in Proc. of CVPR, 2021, pp. 12450 -1245
2021
-
[40]
Cross-domain few-shot incremental learning for point-cloud recognition,
Y. Tan, and X. Xiang, “Cross-domain few-shot incremental learning for point-cloud recognition,” in Proc. of IEEE/CVF Winter Conference on Applications of Computer Vision , 2024, pp. 2296- 2305
2024
-
[41]
Visualizing data using T- SNE,
L. V. D. Maaten, and G. Hinton, “Visualizing data using T- SNE,” J. Mach. Learn. Res., vol. 9, no. 96, pp. 2579-2605, 2008
2008
-
[42]
Deep Residual Learnin g for Image Recognition,
K. He, X. Zhang, S. Ren and J. Sun, "Deep Residual Learnin g for Image Recognition," in Proc. of IEEE CVPR, 2016, 1022 pp. 770- 778
2016
-
[43]
Networ k randomization: A simple technique for generalization in deep reinforcement learning
Kimin Lee, Kibok Lee, Jinw oo Shin, and Honglak Lee. Networ k randomization: A simple technique for generalization in deep reinforcement learning. in Proc. of IEEE ICLR, 2020
2020
Reviewed July 3, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.