REVIEW 2 cited by
Parametric Classification for Generalized Category Discovery: A Baseline Study
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Generalized Category Discovery (GCD) aims to discover novel categories in unlabelled datasets using knowledge learned from labelled samples. Previous studies argued that parametric classifiers are prone to overfitting to seen categories, and endorsed using a non-parametric classifier formed with semi-supervised k-means. However, in this study, we investigate the failure of parametric classifiers, verify the effectiveness of previous design choices when high-quality supervision is available, and identify unreliable pseudo-labels as a key problem. We demonstrate that two prediction biases exist: the classifier tends to predict seen classes more often, and produces an imbalanced distribution across seen and novel categories. Based on these findings, we propose a simple yet effective parametric classification method that benefits from entropy regularisation, achieves state-of-the-art performance on multiple GCD benchmarks and shows strong robustness to unknown class numbers. We hope the investigation and proposed simple framework can serve as a strong baseline to facilitate future studies in this field. Our code is available at: https://github.com/CVMI-Lab/SimGCD.
Forward citations
Cited by 2 Pith papers
-
Generalized Category Discovery under the Long-Tailed Distribution
A confidence-plus-density sample selection framework improves generalized category discovery accuracy on long-tailed image benchmarks and provides a faster density-peak class number estimator.
-
Unleashing the Potential of Model Bias for Generalized Category Discovery
SDC reuses the biased outputs of a pre-trained model to adjust logits and generate better pseudo-labels, improving novel category discovery in text classification.
Discussion (0). Continue with ORCID to comment.