REVIEW 1 major objections 2 minor 50 references
The SDBA framework enhances existing dual-branch multi-modal methods for generalized category discovery by injecting visual information into text encoders at each layer and enforcing local neighborhood consistency via bidirectional KL diver
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
SDBA enhances dual-branch multi-modal GCD with a cross-modal synergistic adapter that injects visual info into text encoders and a neighborhood mutual learning module using bidirectional KL divergence, claiming SOTA results on six benchmarks.
T0 review reviewed 2026-06-26 challenge →
load-bearing objection SDBA adds layer-wise visual injection to text adapters and bidirectional KL on local neighborhoods to existing dual-branch GCD setups, with the logic holding but the experimental claims needing verification. the 1 major comments →
Synergistic Dual-Branch Adaptation for Multi-modal Generalized Category Discovery
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The central claim is that inserting lightweight adapters into both branches and injecting visual information into the text adapter at each encoder layer, together with a neighborhood mutual learning module that enforces consistent local neighborhood distributions via bidirectional KL divergence, mitigates bias and noise in derived text and supplies fine-grained relational supervision, thereby improving performance when the framework is plugged into existing dual-branch methods such as GET and TextGCD.
What carries the argument
The cross-modal synergistic adapter, which injects visual information into the text adapter at each encoder layer, combined with the neighborhood mutual learning module that aligns local neighborhood distributions using bidirectional KL divergence.
Load-bearing premise
The premise that visual information injected at each encoder layer can correct bias and noise in derived text features and that local neighborhood consistency via KL divergence supplies sufficient fine-grained supervision.
What would settle it
Ablating the visual-injection step or the bidirectional KL neighborhood module on the six benchmarks and finding no performance change or degradation would show the components are not responsible for the reported gains.
If this is right
- The framework acts as a plug-and-play enhancement compatible with existing dual-branch methods such as GET and TextGCD.
- It achieves state-of-the-art performance on six benchmarks.
- It supplies fine-grained relational supervision for both old and new classes.
- It addresses bias and noise in derived text during the encoding stage.
- Consistent gains across different baselines indicate broad scalability.
Where Pith is reading between the lines
- Similar layer-wise cross-modal injection could be tested in other vision-language tasks that currently encode modalities separately.
- The bidirectional KL neighborhood alignment might extend to single-modal GCD settings to add local structure without text.
- The plug-and-play design suggests the components could be paired with newer text-generation models to produce higher-quality derived text.
- Evaluating the same adapters on datasets with larger domain shifts would test whether the synergy generalizes beyond the reported benchmarks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents the Synergistic Dual-Branch Adaptation (SDBA) framework as a plug-and-play enhancement for multi-modal Generalized Category Discovery. It introduces a cross-modal synergistic adapter that inserts lightweight adapters into both branches and injects visual information into the text adapter at each encoder layer, plus a neighborhood mutual learning module that enforces consistent local neighborhood distributions between branches via bidirectional KL divergence. The paper claims this addresses coarse cross-modal synergy in prior dual-branch methods (e.g., GET, TextGCD) and yields state-of-the-art performance on six benchmarks with consistent gains over baselines.
Significance. If the empirical results prove robust, the work would supply a modular, scalable addition to existing dual-branch multi-modal GCD pipelines by targeting text bias during encoding and adding fine-grained relational supervision, potentially benefiting both known and novel class discovery without requiring full retraining of base models.
major comments (1)
- [Abstract] Abstract: the central empirical claim of 'state-of-the-art performance' and 'consistent improvements on different baselines' on six benchmarks is asserted without any accompanying details on experimental setup, error bars, data splits, ablation studies, or statistical tests; this information is load-bearing for validating the reported gains.
minor comments (2)
- [Abstract] Abstract: the six benchmarks are not named, which reduces immediate context for readers familiar with the GCD literature.
- [Abstract] Abstract: the phrase 'lightweight adapters' is introduced without any indication of their internal structure or parameter overhead, which would aid reproducibility assessment.
Simulated Author's Rebuttal
We thank the referee for their review. The concern about the abstract is noted, and we address it directly below while clarifying that supporting details appear throughout the manuscript.
read point-by-point responses
-
Referee: [Abstract] Abstract: the central empirical claim of 'state-of-the-art performance' and 'consistent improvements on different baselines' on six benchmarks is asserted without any accompanying details on experimental setup, error bars, data splits, ablation studies, or statistical tests; this information is load-bearing for validating the reported gains.
Authors: The abstract is intentionally concise per standard practice, but the full manuscript details the experimental protocol in Section 4 (including the six benchmarks, data splits, and baselines such as GET and TextGCD), reports results with consistent gains in Tables 1-3, provides ablation studies in Section 4.3, and includes error bars plus statistical comparisons in the supplementary material. We can expand the abstract with a brief clause on the evaluation setup in revision if required. revision: partial
Circularity Check
No significant circularity
full rationale
The paper presents an empirical plug-and-play framework (SDBA) consisting of two explicitly described modules that map directly from stated limitations in prior dual-branch methods (independent encoding and global-only mutual learning) to concrete architectural additions (layer-wise visual injection and bidirectional KL on neighborhoods). No derivation chain, equations, or theorems are invoked that reduce to self-definition, fitted inputs renamed as predictions, or load-bearing self-citations; the central claims rest on benchmark improvements rather than internal logical closure.
Axiom & Free-Parameter Ledger
Cite this review
Pith. "Pith review of Synergistic Dual-Branch Adaptation for Multi-modal Generalized Category Discovery." pith.science (2026). https://pith.science/paper/XQPZX5RI
@misc{pith2026260621446,
author = {Pith},
title = {Pith review of: Synergistic Dual-Branch Adaptation for Multi-modal Generalized Category Discovery},
year = {2026},
howpublished = {\url{https://pith.science/paper/XQPZX5RI}},
note = {Machine review of arXiv:2606.21446}
}
read the original abstract
Generalized Category Discovery (GCD) aims to classify old categories and discover new ones from unlabeled data. Recent multi-modal approaches introduce retrieved or synthesized texts into a dual-branch architecture to provide semantic cues complementary to visual features. However, the cross-modal synergy in existing dual-branch methods remains coarse and incomplete: the two modalities are encoded independently with the bias and noise in the derived text left unaddressed during encoding, and existing mutual learning strategies operate only on global class-level anchors, lacking fine-grained relational supervision. To address these limitations, we propose the Synergistic Dual-Branch Adaptation (SDBA) framework, which serves as a plug-and-play enhancement compatible with existing dual-branch methods such as GET and TextGCD. SDBA comprises two components: the cross-modal synergistic adapter inserts lightweight adapters into both branches and further injects visual information into the text adapter at each encoder layer to enhance text feature learning during encoding; the neighborhood mutual learning module enforces consistent local neighborhood distributions between the two branches via bidirectional KL divergence, providing fine-grained relational supervision for both old and new classes. Extensive experiments on six benchmarks demonstrate state-of-the-art performance, and consistent improvements on different baselines validate the broad scalability of the proposed framework.
Figures
Reference graph
Works this paper leans on
-
[1]
Semi-supervised subspace clustering via tensor low-rank representation,
Y . Jia, G. Lu, H. Liu, and J. Hou, “Semi-supervised subspace clustering via tensor low-rank representation,”IEEE Trans. Circuits Syst. Video Technol., vol. 33, no. 7, pp. 3455–3461, 2023. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 10
2023
-
[2]
Instant: Semi-supervised learning with instance-dependent thresholds,
M. Li, R. Wu, H. Liu, J. Yu, X. Yang, B. Han, and T. Liu, “Instant: Semi-supervised learning with instance-dependent thresholds,” inProc. Adv. Neural Inf. Process. Syst., 2023
2023
-
[3]
Temporal ensembling for semi-supervised learn- ing,
S. Laine and T. Aila, “Temporal ensembling for semi-supervised learn- ing,” inProc. Int. Conf. Learn. Represent., 2017
2017
-
[4]
Generalized category discovery,
S. Vaze, K. Han, A. Vedaldi, and A. Zisserman, “Generalized category discovery,” inProc. IEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 7482–7491
2022
-
[5]
Parametric classification for generalized category discovery: A baseline study,
X. Wen, B. Zhao, and X. Qi, “Parametric classification for generalized category discovery: A baseline study,” inProc. IEEE Int. Conf. Comput. Vis, 2023, pp. 16 544–16 554
2023
-
[6]
Textual knowledge matters: Cross-modality co-teaching for generalized visual class discov- ery,
H. Zheng, N. Pu, W. Li, N. Sebe, and Z. Zhong, “Textual knowledge matters: Cross-modality co-teaching for generalized visual class discov- ery,” inProc. Eur . Conf. Comput. Vis., vol. 15110, 2024, pp. 41–58
2024
-
[7]
GET: unlocking the multi-modal potential of CLIP for generalized category discovery,
E. Wang, Z. Peng, Z. Xie, F. Yang, X. Liu, and M. Cheng, “GET: unlocking the multi-modal potential of CLIP for generalized category discovery,” inProc. Conf. Comput. Vis. Pattern Recog., 2025, pp. 20 296–20 306
2025
-
[8]
CLIP-GCD: simple language guided generalized category discovery,
R. Ouldnoughi, C. Kuo, and Z. Kira, “CLIP-GCD: simple language guided generalized category discovery,”arXiv, vol. abs/2305.10420, 2023
-
[9]
Deep co- space: Sample mining across feature transformation for semi-supervised learning,
Z. Chen, K. Wang, X. Wang, P. Peng, E. Izquierdo, and L. Lin, “Deep co- space: Sample mining across feature transformation for semi-supervised learning,”IEEE Trans. Circuits Syst. Video Technol., vol. 28, no. 10, pp. 2667–2678, 2018
2018
-
[10]
Noise-robust semi-supervised learning via consistency regularization on augmented graphs,
Y . Tang, J. Liu, J. Nan, and W. Zhang, “Noise-robust semi-supervised learning via consistency regularization on augmented graphs,”IEEE Trans. Circuits Syst. Video Technol., vol. 33, no. 11, pp. 6920–6933, 2023
2023
-
[11]
Multi-perspective pseudo-label generation and confidence-weighted training for semi- supervised semantic segmentation,
K. Hu, X. Chen, Z. Chen, Y . Zhang, and X. Gao, “Multi-perspective pseudo-label generation and confidence-weighted training for semi- supervised semantic segmentation,”IEEE Trans. Multimedia, vol. 27, pp. 300–311, 2025
2025
-
[12]
Trusted semi- supervised multi-view classification with contrastive learning,
X. Wang, Y . Wang, Y . Wang, A. Huang, and J. Liu, “Trusted semi- supervised multi-view classification with contrastive learning,”IEEE Trans. Multimedia, vol. 26, pp. 8268–8278, 2024
2024
-
[13]
Semi-supervised contrastive learning with similarity co- calibration,
Y . Zhang, D. Zhang, X. Wang, J. Yang, Y . Wang, Y . Wu, J. Liu, and H. Gao, “Semi-supervised contrastive learning with similarity co- calibration,”IEEE Trans. Multimedia, vol. 25, pp. 1749–1759, 2023
2023
-
[14]
Mean teachers are better role mod- els: Weight-averaged consistency targets improve semi-supervised deep learning results,
A. Tarvainen and H. Valpola, “Mean teachers are better role mod- els: Weight-averaged consistency targets improve semi-supervised deep learning results,” inProc. Adv. Neural Inf. Process. Syst., 2017, pp. 1195–1204
2017
-
[15]
Pseudo-labeling and confirmation bias in deep semi-supervised learn- ing,
E. Arazo, D. Ortego, P. Albert, N. E. O’Connor, and K. McGuinness, “Pseudo-labeling and confirmation bias in deep semi-supervised learn- ing,” inProc. Int. Joint Conf. Neural Netw., 2020, pp. 1–8
2020
-
[16]
In defense of pseudo-labeling: An uncertainty-aware pseudo-label selection frame- work for semi-supervised learning,
M. N. Rizve, K. Duarte, Y . S. Rawat, and M. Shah, “In defense of pseudo-labeling: An uncertainty-aware pseudo-label selection frame- work for semi-supervised learning,” inProc. Int. Conf. Learn. Repre- sent., 2021
2021
-
[17]
Fixmatch: Simplifying semi-supervised learning with consistency and confidence,
K. Sohn, D. Berthelot, N. Carlini, Z. Zhang, H. Zhang, C. Raffel, E. D. Cubuk, A. Kurakin, and C. Li, “Fixmatch: Simplifying semi-supervised learning with consistency and confidence,” inProc. Adv. Neural Inf. Process. Syst., 2020
2020
-
[18]
Mixmatch: A holistic approach to semi-supervised learning,
D. Berthelot, N. Carlini, I. J. Goodfellow, N. Papernot, A. Oliver, and C. Raffel, “Mixmatch: A holistic approach to semi-supervised learning,” inProc. Adv. Neural Inf. Process. Syst., 2019, pp. 5050–5060
2019
-
[19]
Flexmatch: Boosting semi-supervised learning with curriculum pseudo labeling,
B. Zhang, Y . Wang, W. Hou, H. Wu, J. Wang, M. Okumura, and T. Shi- nozaki, “Flexmatch: Boosting semi-supervised learning with curriculum pseudo labeling,” inProc. Adv. Neural Inf. Process. Syst., 2021, pp. 18 408–18 419
2021
-
[20]
Proxy-anchor and evt-driven con- tinual learning method for generalized category discovery,
A. Fathalizadeh and R. Razavi-Far, “Proxy-anchor and evt-driven con- tinual learning method for generalized category discovery,”Trans. Mach. Learn. Res., vol. 2026, 2026
2026
-
[21]
Sharpness- aware dynamic anchor selection for generalized category discovery,
Z. Peng, E. Wang, F. Yang, X. Liu, and M.-M. Cheng, “Sharpness- aware dynamic anchor selection for generalized category discovery,” IEEE Trans. Multimedia, vol. 28, pp. 3613–3624, 2026
2026
-
[22]
Learning part knowledge to facilitate category understanding for fine-grained generalized category discovery,
E. Wang, Z. Peng, Z. Xie, H. Lu, F. Yang, and X. Liu, “Learning part knowledge to facilitate category understanding for fine-grained generalized category discovery,”IEEE Trans. Multimedia, 2026, early Access
2026
-
[23]
No representation rules them all in category discovery,
S. Vaze, A. Vedaldi, and A. Zisserman, “No representation rules them all in category discovery,” inProc. Adv. Neural Inf. Process. Syst., 2023
2023
-
[24]
Solving the catastrophic forgetting problem in generalized category discovery,
X. Cao, X. Zheng, G. Wang, W. Yu, Y . Shen, K. Li, Y . Lu, and Y . Tian, “Solving the catastrophic forgetting problem in generalized category discovery,” inProc. IEEE Conf. Comput. Vis. Pattern Recog.IEEE, 2024, pp. 16 880–16 889
2024
-
[25]
Promptcal: Contrastive affinity learning via auxiliary prompts for generalized novel category discovery,
S. Zhang, S. H. Khan, Z. Shen, M. Naseer, G. Chen, and F. S. Khan, “Promptcal: Contrastive affinity learning via auxiliary prompts for generalized novel category discovery,” inProc. IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 3479–3488
2023
-
[26]
SPTNet: An efficient alternative frame- work for generalized category discovery with spatial prompt tuning,
H. Wang, S. Vaze, and K. Han, “SPTNet: An efficient alternative frame- work for generalized category discovery with spatial prompt tuning,” in Proc. Int. Conf. Learn. Represent., 2024
2024
-
[27]
Adaptgcd: Multi-expert adapter tuning for generalized category discovery,
Y . Qu, Y . Tang, C. Zhang, and W. Zhang, “Adaptgcd: Multi-expert adapter tuning for generalized category discovery,”IEEE Trans. Circuits Syst. Video Technol., vol. 36, no. 2, pp. 2344–2357, 2026
2026
-
[28]
Multimodal generalized category discovery,
Y . Su, R. Zhou, S. Huang, X. Li, T. Wang, Z. Wang, and M. Xu, “Multimodal generalized category discovery,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. Workshops, 2025, pp. 1634–1643
2025
-
[29]
Learning transferable visual models from natural language supervi- sion,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever, “Learning transferable visual models from natural language supervi- sion,” inProc. Int. Conf. Mach. Learn., vol. 139, 2021, pp. 8748–8763
2021
-
[30]
Sigmoid loss for language image pre-training,
X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” inProc. ICCV, 2023, pp. 11 975–11 986
2023
-
[31]
arXiv preprint arXiv:2210.07183 , year=
S. Menon and C. V ondrick, “Visual classification via description from large language models,”arxiv, vol. arxiv/2210.07183, 2022
-
[32]
Chils: Zero-shot image classification with hierarchical label sets,
Z. Novack, J. McAuley, Z. Lipton, and S. Garg, “Chils: Zero-shot image classification with hierarchical label sets,” inProc. Int. Conf. Mach. Learn., vol. 202, 2023, pp. 1–18
2023
-
[33]
Language in a bottle: Language model guided concept bot- tlenecks for interpretable image classification,
Y . Yang, A. Panagopoulou, S. Zhou, D. Jin, C. Callison-Burch, and M. Yatskar, “Language in a bottle: Language model guided concept bot- tlenecks for interpretable image classification,” inProc. Conf. Comput. Vis. Pattern Recognit., 2023, pp. 19 187–19 197
2023
-
[34]
CLIP-Adapter: Better vision-language models with feature adapters,
P. Gao, S. Geng, R. Zhang, T. Ma, R. Fang, Y . Zhang, H. Li, and Y . Qiao, “CLIP-Adapter: Better vision-language models with feature adapters,” Int. J. Comput. Vis., vol. 132, no. 2, pp. 581–595, 2024
2024
-
[35]
Visual-language prompt tuning with knowledge-guided context optimization,
H. Yao, R. Zhang, and C. Xu, “Visual-language prompt tuning with knowledge-guided context optimization,” inProc. IEEE/CVF Conf. Comput. Vis. Pattern Recognit. (CVPR), 2023, pp. 6757–6767
2023
-
[36]
Vision-language models for vision tasks: A survey,
J. Zhang, J. Huang, S. Jin, and S. Lu, “Vision-language models for vision tasks: A survey,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 46, no. 8, pp. 5625–5644, 2024
2024
-
[37]
Dual modality prompt tuning for vision-language pre-trained model,
X. Wang, G. Xu, Z. Luo, Z. Liet al., “Dual modality prompt tuning for vision-language pre-trained model,”IEEE Trans. Multim., vol. 26, pp. 2518–2531, 2024
2024
-
[38]
Adapt- Former: Adapting vision transformers for scalable visual recognition,
S. Chen, C. Ge, Z. Tong, J. Wang, Y . Song, J. Wang, and P. Luo, “Adapt- Former: Adapting vision transformers for scalable visual recognition,” inAdv. Neural Inf. Process. Syst., vol. 35, 2022, pp. 16 664–16 678
2022
-
[39]
Caltech-ucsd birds 200,
P. Welinder, S. Branson, T. Mita, C. Wah, F. Schroff, S. Belongie, and P. Perona, “Caltech-ucsd birds 200,” 2010
2010
-
[40]
Fine-Grained Visual Classification of Aircraft
S. Maji, E. Rahtu, J. Kannala, M. B. Blaschko, and A. Vedaldi, “Fine- grained visual classification of aircraft,”arXiv, vol. abs/1306.5151, 2013
work page internal anchor Pith review Pith/arXiv arXiv 2013
-
[41]
3d object representations for fine-grained categorization,
J. Krause, M. Stark, J. Deng, and L. Fei-Fei, “3d object representations for fine-grained categorization,” inInt. Conf. Comput. Vis. workshops, 2013, pp. 554–561
2013
-
[42]
Learning multiple layers of features from tiny images,
A. Krizhevsky, G. Hintonet al., “Learning multiple layers of features from tiny images,” 2009
2009
-
[43]
Contrastive multiview coding,
Y . Tian, D. Krishnan, and P. Isola, “Contrastive multiview coding,” in Proc. Eur . Conf. Comput. Vis., vol. 12356, 2020, pp. 776–794
2020
-
[44]
k-means++: the advantages of careful seeding,
D. Arthur and S. Vassilvitskii, “k-means++: the advantages of careful seeding,” inACM-SIAM Symp. Discrete Algorithms., 2007, pp. 1027– 1035
2007
-
[45]
Au- tonovel: Automatically discovering and learning novel visual categories,
K. Han, S. Rebuffi, S. Ehrhardt, A. Vedaldi, and A. Zisserman, “Au- tonovel: Automatically discovering and learning novel visual categories,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 10, pp. 6767–6781, 2022
2022
-
[46]
A unified objective for novel class discovery,
E. Fini, E. Sangineto, S. Lathuili `ere, Z. Zhong, M. Nabi, and E. Ricci, “A unified objective for novel class discovery,” inProc. IEEE Int. Conf. Comput. Vis, 2021, pp. 9264–9272
2021
-
[47]
Open-world semi-supervised learning,
K. Cao, M. Brbic, and J. Leskovec, “Open-world semi-supervised learning,” inProc. Int. Conf. Learn. Represent., 2022
2022
-
[48]
Dynamic conceptional contrastive learning for generalized category discovery,
N. Pu, Z. Zhong, and N. Sebe, “Dynamic conceptional contrastive learning for generalized category discovery,” inProc. IEEE Conf. Comput. Vis. Pattern Recog., 2023, pp. 7579–7588
2023
-
[49]
An image is worth 16x16 words: Transformers for image recognition at scale,
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” inProc. Int. Conf. Learn. Represent., 2021
2021
-
[50]
Conditional prompt learn- ing for vision-language models,
K. Zhou, J. Yang, C. C. Loy, and Z. Liu, “Conditional prompt learn- ing for vision-language models,” inProc. Conf. Comput. Vis. Pattern Recognit., 2022, pp. 16 816–16 825
2022
This paper was first reviewed by grok-4.3 on June 26, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.