Pith. sign in

REVIEW 4 major objections 7 minor 76 references

PI-H2T: Enhancing Long-Tailed Visual Recognition with Permutation-Invariant and Head-to-Tail Feature Fusion

T0 review · 4 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read PI-H2T fuses permutation-invariant features into the representation and interpolates head-class features into tail-class features during classifier retraining, improving long-tailed recognition accuracy on four benchmarks.

desk verdict A useful plug-in extension of the authors' H2T line, with broad benchmarks and real incremental novelty, but the theory is shaky and the headline tail regressions need explanation before the numbers can be trusted. read the letter →

arxiv 2506.00625 v2 pith:O3DXTSPF submitted 2025-05-31 cs.CV

classification cs.CV
keywords pi-h2tclassesfusionrepresentationtailhead-to-taillong-tailedpermutation-invariant
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Long-tailed recognition is the task of classifying images when some classes have many examples and others have very few. Models trained on such data tend to classify the frequent head classes well and the rare tail classes poorly. PI-H2T is a two-stage add-on that tries to fix both the feature representation and the final classifier.

In the first stage, it takes the feature map produced by a backbone network and also computes an order-agnostic version of it by averaging across the channel dimension. Because averaging is commutative, this permutation-invariant version does not depend on the order in which channels appear. The method learns two small weights to merge the original feature map with this averaged version. The authors argue this fusion makes each class's features more compact and leaves more room between classes, although the proof is largely a restatement of the classifier's own margin.

In the second stage, the feature extractor is frozen and only the classifier is retrained. The method samples two streams of data: one balanced across classes and one following the original imbalanced distribution, so it frequently draws head-class samples. For each balanced-sample feature, it picks a head-class feature and interpolates between them, keeping the balanced sample's tail-class label. The mixing ratio is set by how far the sample is from its class center. The idea is that head-class semantics can fill in the sparse tail-class regions and push the decision boundary into a better position.

The authors report gains of about 1 to 3 accuracy points when PI-H2T is added to existing single-model, multi-expert, and prompt-tuned vision transformer baselines on CIFAR100-LT, ImageNet-LT, iNaturalist 2018, and Places-LT.

Extended reading notes

Core claim

The paper's central claim is that PI-H2T 'enhances the representation space through permutation-invariant representation fusion (PIF), yielding more clustered features and automatic class margins,' and that it 'adjusts the biased classifier by transferring semantic information from head to tail classes via head-to-tail fusion (H2TF).' If correct, adding PI-H2T to existing long-tailed methods yields consistent top-1 gains, for example GCL+PI-H2T reaches 55.08% on ImageNet-LT, RIDE+PI-H2T reaches 76.03% on iNaturalist 2018, and LPT+PI-H2T reaches 50.12% on Places-LT.

Load-bearing premise

Equations (7) and (8) fuse features as \tilde f = r f_t + (1-r) f_h and label the result with the tail class. The method assumes that linear interpolation between a tail feature and a head feature, with the tail label retained, transfers head semantics into the tail region without creating mislabeled synthetic features. Section 3.4 states this is based on the assumption that tail classes are easily affected by head classes, but no proof or separate validation is provided that head-to-tail interpolation stays within the tail decision region, especially for samples far from the class center. The paper's own Places-LT results show the improvement can be small or inconsistent for some baselines.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper proposes PI-H2T, a two-stage plug-and-play module for long-tailed visual recognition. In stage I, PIF fuses permutation-invariant feature maps (channel mean) with the original feature maps through a learned 1D convolution, with the goal of producing more clustered features and automatic class margins. In stage II, H2TF retrains the classifier on features interpolated between class-balanced (tail-emphasized) samples and instance-wise (head-emphasized) samples, retaining the tail label and setting the fusion ratio via cosine distance to the class center. The authors provide two theoretical remarks about margins and about two opposing "forces" in H2TF, and evaluate the method on CIFAR100-LT, ImageNet-LT, iNaturalist 2018, and Places-LT with single and multi-expert CNN backbones as well as a ViT backbone, reporting top-1 improvements over several baselines. Source code is provided.

Significance. If the reported gains are reproducible, PI-H2T would be a useful low-cost add-on to existing long-tailed learning methods: it introduces very few parameters, leaves the backbone architecture intact, and is demonstrated on single-model, multi-expert, and prompt-tuning settings. The release of source code and the breadth of benchmarks are positive features. However, the theoretical support is incomplete, and the headline experiments do not consistently support the central tail-improvement mechanism; several large discrepancies and the absence of error bars make it impossible to assess the method's true gains at face value.

major comments (4)
  1. [Section 3.5, Eq. (10)] The proof in Remark 1 is internally inconsistent. Equation (10) states f_fuse = a(f - f_PI), which omits the b·f term present in Eq. (3), yet Eqs. (12)-(13) reintroduce b·w_y^T f and b·w_i^T f. The derivation of Eq. (14) only goes through if Eq. (10) is corrected to f_fuse = a(f - f_PI) + b·f. In addition, the claim that the margin in Eq. (14) is positive relies on the unproven alignment condition w_y^T f_PI > w_i^T f_PI; the sentence "should be closer to w_y" is an assertion, not a proof. Since this margin is the basis for the "automatic class margins" claim, the proof needs to be fixed or the claim weakened.
  2. [Table 2 and Table 4] The headline results contradict the proposed mechanism on the metric the method is designed to improve. On ImageNet-LT, GCL+PI-H2T has a tail accuracy of 37.25%, down from 52.12% for GCL and 52.15% for GCL+H2T, while the overall gain to 55.08% comes from head and medium classes. On Places-LT, GCL+PI-H2T has tail accuracy 33.74% versus 38.44% for GCL, and its overall accuracy 40.56% is below GCL+H2T's 40.73%. Section 4.3 acknowledges only that gains are "less obvious" on Places-LT, but does not mention or explain these tail regressions. If these numbers are correct, the central claim that H2TF transfers head semantics to tail classes and expands tail space is unsupported on the largest benchmarks; if they are typos, the tables must be corrected.
  3. [Section 3.5, Remark 2] The "two forces" derivation is not a valid proof as written. Equations (15) and (16) concern different objects: Eq. (15) states an inequality for a correctly classified tail sample, while Eq. (16) states an inequality for a different "H2TF head" quantity. Adding these inequalities requires that both hold for the same pair (f_t, f_h), but no coupling or joint distribution is supplied. The notation \tilde f_i in the sentence preceding Eq. (15) is undefined. Furthermore, the claim that correctly classified training samples outnumber misclassified ones is an empirical property that is not established, and Remark 2 provides no formal connection from these sample-level inequalities to the geometry of the learned decision boundary.
  4. [Section 4, empirical methodology] All reported experimental numbers are single runs with no error bars, standard deviations, or significance tests. Given the large and inconsistent differences in the tables—for example, GCL+PI-H2T improves by 3.40 points on iNaturalist 2018 (75.02 vs. 71.62) but loses 14.87 points on ImageNet-LT tail accuracy (37.25 vs. 52.12)—the results as printed cannot be distinguished from seed artifacts, implementation bugs, or real effects. The paper should report mean and standard deviation over multiple seeds for the main comparisons, at least for the core claims of tail improvement and overall accuracy gains.
minor comments (7)
  1. [Section 3.3, Eq. (6)] Equation (6) defines F_fuse = Conv1D(Concat{F - F_PI, F}) but does not specify the convolution dimension or how the learnable parameters a and b in Eq. (3) map to the 1D convolution kernel; this should be clarified.
  2. [Section 3.2 and Abstract] There is a duplicated phrase "we propose the fusion of the fusion of permutation-invariant features" in Section 3.2; the abstract also has a minor wording issue in "yielding more clustered features and automatic class margins." Please proofread.
  3. [Section 4.2, Basic Settings] The experimental settings do not report the hyperparameters introduced by PI-H2T, such as the initialization of the 1D convolution in Eq. (6), the number of fused channels, and any regularization; these details are needed for reproducibility.
  4. [Eq. (8)] The notation Norm[d_x]_0^1 is nonstandard; please define explicitly how the cosine distance is normalized to [0,1].
  5. [Algorithm 1, lines 12-14] The variable names f^I and f^B in Algorithm 1 are not clearly mapped to f_t (fused sample) and f_h (fusing sample) from Eq. (7); the text should explicitly state which branch provides the fused and fusing features.
  6. [Figures 5 and 6] The claims that PIF "yields more clustered features" and "automatic class margins" are supported only by qualitative t-SNE plots; quantitative measures such as intra-class variance, inter-class distance, or nearest-class accuracy would strengthen the evaluation.
  7. [References] Reference [56] duplicates [27] (both are Radford et al., CLIP), and reference [73] in the bibliography lists an incomplete author string; these should be cleaned up.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the empirical evaluation is against external benchmarks and the method components are defined independently of the reported results.

full rationale

The derivation chain is not circular in any of the enumerated senses. The method's two components are defined directly from the input features: PIF fuses the backbone feature map with a channel-wise mean (Eqs. 2-6), and H2TF linearly interpolates a balanced-branch feature with an instance-wise-branch feature using a ratio set by cosine distance to the class center (Eqs. 7-9). These are algorithmic constructions, not parameters fitted to the benchmark accuracies, and no reported 'prediction' is a renamed fit. The theoretical Remarks 1 and 2 are largely tautological rearrangements of the classifier inequalities and rely on the assumption that training-set classification already succeeds, but that is a derivation weakness rather than input-output circularity; the central empirical claim is supported by comparisons on CIFAR100-LT, ImageNet-LT, iNaturalist, and Places-LT against external baselines. The paper cites its own prior work H2T [38], but only as the direct predecessor that PI-H2T extends and as a comparison row, not as the evidence for PI-H2T's gains; the self-citation is therefore not load-bearing. The suspicious tail-accuracy drops in Tables 2 and 4 are substantive correctness or reproducibility concerns, but they do not show that the method's outputs are equivalent to its inputs.

Assumptions & free parameters 2 free parameters · 7 assumptions · 0 invented entities

The central empirical claim rests on two learned operations: PIF, which adds mean-pooled permutation-invariant features to the backbone output, and H2TF, which interpolates head-class features into tail-class features during classifier retraining. The theoretical justifications in Section 3.5 rely on assumptions about training-set correctness, alignment of PI features with class centers, and semantic preservation under linear interpolation. No external fixed constants are derived; both modules are trained on the target datasets.

free parameters (2)
  • PIF fusion weights a and b (or 1D convolution kernel in Eq. 6) = learned during training; values not reported
    These weights determine how much permutation-invariant signal is mixed into the representation; they are fitted to the target dataset by gradient descent and are part of the proposed module.
  • Fusion ratio r (defined by Eq. 8) = per-sample normalized distance; no single reported value
    The ratio controls the balance between tail and head features in H2TF and is derived from cosine distance to the classifier weight, which is itself being trained in stage II. Its normalization statistics are data-dependent.
assumptions (7)
  • standard math Mean operation across feature channels is permutation-invariant.
    Theorem 1 in Section 3.3; follows from commutativity of addition.
  • domain assumption Permutation-invariant features provide useful complementary information to the original ordered feature maps.
    Section 3.2 motivates PIF by claiming PI features reduce overfitting to feature order and enable inner-set interactions; this is not proven and is the basis for Eq. 3.
  • domain assumption After stage I, target-class logits exceed non-target logits for training samples, and specifically (w_y - w_i)^T f_PI > 0.
    Remark 1, Eq. (14); the automatic margin is positive only if the PI feature of a sample is better aligned with its own class center than with other class centers. The paper asserts this should hold but provides no proof.
  • domain assumption Linear interpolation between a tail-class feature and a head-class feature preserves tail-class semantics under the tail label.
    Section 3.4 and Eq. (7); H2TF labels the fused feature with the tail class and assumes head semantics fill the tail region rather than creating mislabeled samples.
  • domain assumption Classifier weights w_y can serve as class centers for the distance-based fusion ratio.
    Section 3.4, Eq. (8); cosine distance to w_y is used to set r, but w_y in stage I is biased toward head classes.
  • domain assumption In stage II, correctly classified training samples outnumber misclassified ones, so the force in Eq. (18) dominates Eq. (21).
    Section 3.5, Remark 2; the calibration conclusion depends on this imbalance in training-set correctness, which is not guaranteed for long-tailed data.
  • domain assumption Freezing the feature extractor during classifier fine-tuning keeps the representation stable.
    Algorithm 1, line 15; standard two-stage protocol from Kang et al., but it is an assumption that classifier retraining alone can correct boundary bias.

how reviews work

0 comments
Cite this review

Pith. "Pith review of PI-H2T: Enhancing Long-Tailed Visual Recognition with Permutation-Invariant and Head-to-Tail Feature Fusion." pith.science (2026). https://pith.science/paper/O3DXTSPF

@misc{pith2026250600625,
  author       = {Pith},
  title        = {Pith review of: PI-H2T: Enhancing Long-Tailed Visual Recognition with Permutation-Invariant and Head-to-Tail Feature Fusion},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/O3DXTSPF}},
  note         = {Machine review of arXiv:2506.00625}
}
read the original abstract

The imbalanced distribution of long-tailed data presents a significant challenge for deep learning models, causing them to prioritize head classes while neglecting tail classes. Two key factors contributing to low recognition accuracy are the deformed representation space and a biased classifier, stemming from insufficient semantic information in tail classes. To address these issues, we propose permutation-invariant and head-to-tail feature fusion (PI-H2T), a highly adaptable method. PI-H2T enhances the representation space through permutation-invariant representation fusion (PIF), yielding more clustered features and automatic class margins. Additionally, it adjusts the biased classifier by transferring semantic information from head to tail classes via head-to-tail fusion (H2TF), improving tail class diversity. Theoretical analysis and experiments show that PI-H2T optimizes both the representation space and decision boundaries. Its plug-and-play design ensures seamless integration into existing methods, providing a straightforward path to further performance improvements. Extensive experiments on long-tailed benchmarks confirm the effectiveness of PI-H2T.

Figures

Figures reproduced from arXiv: 2506.00625 by the authors.

Figure 1
Figure 1. Comparison between decision boundaries produced by (a) exist [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Illustration of the relationship between the fusion ratio and the [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Framework of permutation-invariant and head-to-tail feature fusion (PI-H2T). [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Rationale analysis of H2TF. Forces ⃝1 and ⃝2 are generated by Eqs. (18) and (21), respectively. ⃝1 > ⃝2 makes the tail sample to “pull" closer to wt and “push" further away from wh, leading to the adjustment of decision boundary and enlargement of the tail class space.…
Figure 5
Figure 5. Figure 5: t-SNE visualization of feature distribution. The dataset is [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Decision boundary between Class 0 and Class 9 on CIFAR10-LT [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

76 extracted references · 74 canonical work pages

  1. [1]

    Long-tail learning with foundation model: Heavy fine-tuning hurts,

    Jiang-Xin Shi, Tong Wei, Zhi Zhou, Jie-Jing Shao, Xin-Yan Han, and Yu-Feng Li, “Long-tail learning with foundation model: Heavy fine-tuning hurts,” inICML, 2024

  2. [2]

    The pareto, zipf and other power laws,

    William J Reed, “The pareto, zipf and other power laws,”Eco- nomics letters, vol. 74, no. 1, pp. 15–19, 2001

  3. [3]

    Large-scale long-tailed recognition in an open world,

    Ziwei Liu, Zhongqi Miao, Xiaohang Zhan, Jiayun Wang, Boqing Gong, and Stella X. Yu, “Large-scale long-tailed recognition in an open world,” inIEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 2537–2546

  4. [4]

    Deep long-tailed learning: A survey,

    Yifan Zhang, Bingyi Kang, Bryan Hooi, Shuicheng Yan, and Jiashi Feng, “Deep long-tailed learning: A survey,”IEEE Trans. Pattern Anal. Mach. Intell., pp. 1–20, 2023

  5. [5]

    thesis, Hong Kong Baptist University, 2022

    Mengke Li,Advances in Long-Tailed Visual Recognition, Ph.D. thesis, Hong Kong Baptist University, 2022

  6. [6]

    SMOTE: synthetic minority over-sampling technique,

    Nitesh V . Chawla, Kevin W. Bowyer, Lawrence O. Hall, and W. Philip Kegelmeyer, “SMOTE: synthetic minority over-sampling technique,”Journal of artificial intelligence research, vol. 16, pp. 321– 357, 2002

  7. [7]

    Learning deep representation for imbalanced classification,

    Chen Huang, Yining Li, Chen Change Loy, and Xiaoou Tang, “Learning deep representation for imbalanced classification,” in IEEE Conf. Comput. Vis. Pattern Recog., June 2016

  8. [8]

    M2m: Imbal- anced classification via major-to-minor translation,

    Jaehyung Kim, Jongheon Jeong, and Jinwoo Shin, “M2m: Imbal- anced classification via major-to-minor translation,” inIEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 13896–13905

Show all 76 references
  1. [9]

    RSG: A simple but effective module for learning imbalanced datasets,

    Jianfeng Wang, Thomas Lukasiewicz, Xiaolin Hu, Jianfei Cai, and Xu Zhenghua, “RSG: A simple but effective module for learning imbalanced datasets,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 3784–3793

  2. [10]

    MetaSAug: Meta semantic augmenta- tion for long-tailed visual recognition,

    Shuang Li, Kaixiong Gong, Chi Harold Liu, Yulin Wang, Feng Qiao, and Xinjing Cheng, “MetaSAug: Meta semantic augmenta- tion for long-tailed visual recognition,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 5212–5221

  3. [11]

    The majority can help the minority: Context- rich minority oversampling for long-tailed classification,

    Seulki Park, Youngkyu Hong, Byeongho Heo, Sangdoo Yun, and Jin Young Choi, “The majority can help the minority: Context- rich minority oversampling for long-tailed classification,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 6887–6896

  4. [12]

    BBN: Bilateral-branch network with cumulative learning for long-tailed visual recognition,

    Boyan Zhou, Quan Cui, Xiu-Shen Wei, and Zhao-Min Chen, “BBN: Bilateral-branch network with cumulative learning for long-tailed visual recognition,” inCVPR, 2020, pp. 9719–9728

  5. [13]

    Decoupling represen- tation and classifier for long-tailed recognition,

    Bingyi Kang, Saining Xie, Marcus Rohrbach, Zhicheng Yan, Albert Gordo, Jiashi Feng, and Yannis Kalantidis, “Decoupling represen- tation and classifier for long-tailed recognition,” inInt. Conf. Learn. Represent., 2020

  6. [14]

    Long-tailed recognition by routing diverse distribution- aware experts,

    Xudong Wang, Long Lian, Zhongqi Miao, Ziwei Liu, and Stella X. Yu, “Long-tailed recognition by routing diverse distribution- aware experts,” inInt. Conf. Learn. Represent., 2021

  7. [15]

    The devil is in classifi- cation: A simple framework for long-tail instance segmentation,

    Tao Wang, Yu Li, Bingyi Kang, Junnan Li, Jun Hao Liew, Sheng Tang, Steven C. H. Hoi, and Jiashi Feng, “The devil is in classifi- cation: A simple framework for long-tail instance segmentation,” inEur. Conf. Comput. Vis., 2020, vol. 12359, pp. 728–744

  8. [16]

    Trustworthy long-tailed classification,

    Bolian Li, Zongbo Han, Haining Li, Huazhu Fu, and Changqing Zhang, “Trustworthy long-tailed classification,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 6970–6979

  9. [17]

    Long-tailed visual recognition via self-heterogeneous integration with knowledge excavation,

    Yan Jin, Mengke Li, Yang Lu, Yiu-ming Cheung, and Hanzi Wang, “Long-tailed visual recognition via self-heterogeneous integration with knowledge excavation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), June 2023, pp. 23695–23704

  10. [18]

    Focal loss for dense object detection,

    Tsung-Yi Lin, Priya Goyal, Ross B. Girshick, Kaiming He, and Piotr Dollár, “Focal loss for dense object detection,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 42, no. 2, pp. 318–327, 2020

  11. [19]

    Class-balanced loss based on effective number of samples,

    Yin Cui, Menglin Jia, Tsung-Yi Lin, Yang Song, and Serge Belongie, “Class-balanced loss based on effective number of samples,” in IEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 9268–9277

  12. [20]

    Balanced meta-softmax for long-tailed visual recognition,

    Jiawei Ren, Cunjun Yu, shunan sheng, Xiao Ma, Haiyu Zhao, Shuai Yi, and hongsheng Li, “Balanced meta-softmax for long-tailed visual recognition,” inAdv. Neural Inform. Process. Syst., 2020, vol. 33, pp. 4175–4186

  13. [21]

    Long-tail learning via logit adjustment,

    Aditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain, Andreas Veit, and Sanjiv Kumar, “Long-tail learning via logit adjustment,” inInt. Conf. Learn. Represent., 2021

  14. [22]

    Disentangling label distribution for long-tailed visual recognition,

    Youngkyu Hong, Seungju Han, Kwanghee Choi, Seokjun Seo, Beomsu Kim, and Buru Chang, “Disentangling label distribution for long-tailed visual recognition,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2021, pp. 6626–6636

  15. [23]

    Long-tailed visual recognition via gaussian clouded logit adjustment,

    Mengke Li, Yiu-ming Cheung, and Yang Lu, “Long-tailed visual recognition via gaussian clouded logit adjustment,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2022, pp. 6929–6938

  16. [24]

    A simple long-tailed recognition baseline via vision-language model,

    Teli Ma, Shijie Geng, Mengmeng Wang, Jing Shao, Jiasen Lu, Hongsheng Li, Peng Gao, and Yu Qiao, “A simple long-tailed recognition baseline via vision-language model,”arXiv preprint arXiv:2111.14745, 2021

  17. [25]

    VL-LTR: Learning class-wise visual-linguistic represen- tation for long-tailed visual recognition,

    Changyao Tian, Wenhai Wang, Xizhou Zhu, Jifeng Dai, and Yu Qiao, “VL-LTR: Learning class-wise visual-linguistic represen- tation for long-tailed visual recognition,” inEur. Conf. Comput. Vis. Springer, 2022, pp. 73–91

  18. [26]

    LPT: long-tailed prompt tuning for image classification,

    Bowen Dong, Pan Zhou, Shuicheng Yan, and Wangmeng Zuo, “LPT: long-tailed prompt tuning for image classification,” inInt. Conf. Learn. Represent., 2023

  19. [28]

    ImageNet-21K pretraining for the masses,

    Tal Ridnik, Emanuel Ben Baruch, Asaf Noy, and Lihi Zelnik, “ImageNet-21K pretraining for the masses,” inAdv. Neural Inform. Process. Syst., 2021

  20. [29]

    Learning imbalanced datasets with label-distribution-aware margin loss,

    Kaidi Cao, Colin Wei, Adrien Gaidon, Nikos Arechiga, and Tengyu Ma, “Learning imbalanced datasets with label-distribution-aware margin loss,” inAdv. Neural Inform. Process. Syst., 2019, pp. 1567– 1578

  21. [30]

    Arcface: Additive angular margin loss for deep face recognition,

    Jiankang Deng, Jia Guo, Niannan Xue, and Stefanos Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” inIEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 4690–4699. JOURNAL OF LATEX CLASS FILES 12

  22. [31]

    Cosface: Large margin cosine loss for deep face recognition,

    Hao Wang, Yitong Wang, Zheng Zhou, Xing Ji, Dihong Gong, Jingchao Zhou, Zhifeng Li, and Wei Liu, “Cosface: Large margin cosine loss for deep face recognition,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 5265–5274

  23. [32]

    Normface: L2 hypersphere embedding for face verification,

    Feng Wang, Xiang Xiang, Jian Cheng, and Alan Loddon Yuille, “Normface: L2 hypersphere embedding for face verification,” in ACM Int. Conf. Multimedia, 2017, pp. 1041–1049

  24. [33]

    Mtfh: A matrix tri-factorization hashing framework for efficient cross- modal retrieval,

    Xin Liu, Zhikai Hu, Haibin Ling, and Yiu-Ming Cheung, “Mtfh: A matrix tri-factorization hashing framework for efficient cross- modal retrieval,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 43, no. 3, pp. 964–981, 2021

  25. [34]

    Joint semantic preserving sparse hashing for cross-modal retrieval,

    Zhikai Hu, Yiu-ming Cheung, Mengke Li, Weichao Lan, Donglin Zhang, and Qiang Liu, “Joint semantic preserving sparse hashing for cross-modal retrieval,”IEEE Trans. Circuit Syst. Video Technol., vol. 34, no. 4, pp. 2989–3002, 2024

  26. [35]

    Revisiting unreasonable effectiveness of data in deep learning era,

    Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta, “Revisiting unreasonable effectiveness of data in deep learning era,” inInt. Conf. Comput. Vis., 2017, pp. 843–852

  27. [36]

    Does head label help for long-tailed multi-label text classification,

    Lin Xiao, Xiangliang Zhang, Liping Jing, Chi Huang, and Mingyang Song, “Does head label help for long-tailed multi-label text classification,”AAAI, vol. 35, no. 16, pp. 14103–14111, 2021

  28. [37]

    Permutation-invariant fea- ture restructuring for correlation-aware image set-based recogni- tion,

    Xiaofeng Liu, Zhenhua Guo, Site Li, Lingsheng Kong, Ping Jia, Jane You, and B.V .K. Vijaya Kumar, “Permutation-invariant fea- ture restructuring for correlation-aware image set-based recogni- tion,” inInt. Conf. Comput. Vis., October 2019

  29. [38]

    Feature fusion from head to tail for long-tailed visual recognition,

    Mengke Li, Zhikai HU, Yang Lu, Weichao Lan, Yiu-ming Cheung, and Hui Huang, “Feature fusion from head to tail for long-tailed visual recognition,”AAAI, vol. 38, no. 12, pp. 13581–13589, 2024

  30. [39]

    Autoaugment: Learning augmentation strategies from data,

    Ekin D. Cubuk, Barret Zoph, Dandelion Mané, Vijay Vasudevan, and Quoc V . Le, “Autoaugment: Learning augmentation strategies from data,” inIEEE Conf. Comput. Vis. Pattern Recog., 2019, pp. 113–123

  31. [40]

    Fast autoaugment,

    Sungbin Lim, Ildoo Kim, Taesup Kim, Chiheon Kim, and Sung- woong Kim, “Fast autoaugment,” inAdv. Neural Inform. Process. Syst., 2019, pp. 6662–6672

  32. [41]

    Pop- ulation based augmentation: Efficient learning of augmentation policy schedules,

    Daniel Ho, Eric Liang, Xi Chen, Ion Stoica, and Pieter , “Pop- ulation based augmentation: Efficient learning of augmentation policy schedules,” inICML, 2019, vol. 97, pp. 2731–2741

  33. [42]

    Randaugment: Practical automated data augmentation with a reduced search space,

    Ekin D. Cubuk, Barret Zoph, Jonathon Shlens, and Quoc V . Le, “Randaugment: Practical automated data augmentation with a reduced search space,” inIEEE Conf. Comput. Vis. Pattern Recog. Worksh., 2020, pp. 3008–3017

  34. [43]

    Parametric contrastive learning,

    Jiequan Cui, Zhisheng Zhong, Shu Liu, Bei Yu, and Jiaya Jia, “Parametric contrastive learning,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 715–724

  35. [44]

    Nested collaborative learning for long-tailed visual recognition,

    Jun Li, Zichang Tan, Jun Wan, Zhen Lei, and Guodong Guo, “Nested collaborative learning for long-tailed visual recognition,” inIEEE Conf. Comput. Vis. Pattern Recog., 2022, pp. 6949–6958

  36. [45]

    Improving calibration for long-tailed recognition,

    Zhisheng Zhong, Jiequan Cui, Shu Liu, and Jiaya Jia, “Improving calibration for long-tailed recognition,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 16489–16498

  37. [46]

    Distribution alignment: A unified framework for long-tail visual recognition,

    Songyang Zhang, Zeming Li, Shipeng Yan, Xuming He, and Jian Sun, “Distribution alignment: A unified framework for long-tail visual recognition,” inIEEE Conf. Comput. Vis. Pattern Recog., 2021, pp. 2361–2370

  38. [47]

    Ace: Ally complementary experts for solving long-tailed recognition in one- shot,

    Jiarui Cai, Yizhou Wang, and Jenq-Neng Hwang, “Ace: Ally complementary experts for solving long-tailed recognition in one- shot,” inInt. Conf. Comput. Vis., 2021, pp. 112–121

  39. [48]

    Reslt: Residual learning for long-tailed recognition,

    Jiequan Cui, Shu Liu, Zhuotao Tian, Zhisheng Zhong, and Jiaya Jia, “Reslt: Residual learning for long-tailed recognition,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 3, pp. 3695–3706, 2023

  40. [49]

    Learning from multiple experts: Self-paced knowledge distillation for long-tailed classification,

    Liuyu Xiang, Guiguang Ding, and Jungong Han, “Learning from multiple experts: Self-paced knowledge distillation for long-tailed classification,” inEur. Conf. Comput. Vis., 2020, pp. 247–263

  41. [50]

    Overcoming classifier imbalance for long- tail object detection with balanced group softmax,

    Yu Li, Tao Wang, Bingyi Kang, Sheng Tang, Chunfeng Wang, Jintao Li, and Jiashi Feng, “Overcoming classifier imbalance for long- tail object detection with balanced group softmax,” inIEEE Conf. Comput. Vis. Pattern Recog., 2020, pp. 10988–10997

  42. [51]

    Key point sen- sitive loss for long-tailed visual recognition,

    Mengke Li, Yiu-ming Cheung, and Zhikai Hu, “Key point sen- sitive loss for long-tailed visual recognition,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 45, no. 4, pp. 4812–4825, 2023

  43. [52]

    Seesaw loss for long-tailed instance segmentation,

    Jiaqi Wang, Wenwei Zhang, Yuhang Zang, Yuhang Cao, Jiangmiao Pang, Tao Gong, Kai Chen, Ziwei Liu, Chen Change Loy, and Dahua Lin, “Seesaw loss for long-tailed instance segmentation,” inProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Ju...

  44. [53]

    An image is worth 16x16 words: Transformers for image recog- nition at scale,

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa De- hghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al., “An image is worth 16x16 words: Transformers for image recog- nition at scale,” inInt. Conf. ...

  45. [54]

    Attention is all you need,

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin, “Attention is all you need,”Adv. Neural Inform. Process. Syst., vol. 30, 2017

  46. [55]

    Learning imbalanced data with vision transformers,

    Zhengzhuo Xu, Ruikang Liu, Shuo Yang, Zenghao Chai, and Chun Yuan, “Learning imbalanced data with vision transformers,” in IEEE Conf. Comput. Vis. Pattern Recog., June 2023, pp. 15793–15803

  47. [56]

    Learning transferable visual models from natural language supervision,

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al., “Learning transferable visual models from natural language supervision,” inICML, 2021, pp. 8748–8763

  48. [57]

    Visual prompt tuning,

    Menglin Jia, Luming Tang, Bor-Chun Chen, Claire Cardie, Serge Belongie, Bharath Hariharan, and Ser-Nam Lim, “Visual prompt tuning,” inEur. Conf. Comput. Vis.Springer, 2022, pp. 709–727

  49. [58]

    Adaptformer: Adapting vision transformers for scalable visual recognition,

    Shoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang, Yibing Song, Jue Wang, and Ping Luo, “Adaptformer: Adapting vision transformers for scalable visual recognition,”Adv. Neural Inform. Process. Syst., vol. 35, pp. 16664–16678, 2022

  50. [59]

    Lora: Low- rank adaptation of large language models,

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen, “Lora: Low- rank adaptation of large language models,” inInt. Conf. Learn. Represent., 2022

  51. [60]

    Retrieval augmented classification for long-tail visual recognition,

    Alexander Long, Wei Yin, Thalaiyasingam Ajanthan, Vu Nguyen, Pulak Purkait, Ravi Garg, Alan Blair, Chunhua Shen, and Anton van den Hengel, “Retrieval augmented classification for long-tail visual recognition,” inIEEE Conf. Comput. Vis. Pattern Recog., June 2022, pp. 6959–6969

  52. [61]

    Vl-ltr: Learning class-wise visual-linguistic represen- tation for long-tailed visual recognition,

    Changyao Tian, Wenhai Wang, Xizhou Zhu, Jifeng Dai, and Yu Qiao, “Vl-ltr: Learning class-wise visual-linguistic represen- tation for long-tailed visual recognition,” inEur. Conf. Comput. Vis.Springer, 2022, pp. 73–91

  53. [62]

    Deep over-sampling frame- work for classifying imbalanced data,

    Shin Ando and Chun-Yuan Huang, “Deep over-sampling frame- work for classifying imbalanced data,” inECML/PKDD, 2017, pp. 770–785

  54. [63]

    A survey on long- tailed visual recognition,

    Lu Yang, He Jiang, Qing Song, and Jun Guo, “A survey on long- tailed visual recognition,”Int. J. Comput. Vis., vol. 130, no. 7, pp. 1837–1872, 2022

  55. [64]

    Deep sets,

    Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Russ R Salakhutdinov, and Alexander J Smola, “Deep sets,” inAdv. Neural Inform. Process. Syst., 2017, vol. 30

  56. [65]

    Deep set prediction networks,

    Yan Zhang, Jonathon Hare, and Adam Prugel-Bennett, “Deep set prediction networks,” inAdv. Neural Inform. Process. Syst., 2019, vol. 32

  57. [66]

    Deep residual learning for image recognition,

    Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, “Deep residual learning for image recognition,” inIEEE Conf. Comput. Vis. Pattern Recog., 2016, pp. 770–778

  58. [67]

    The inaturalist species classification and detection dataset,

    Grant Van Horn, Oisin Mac Aodha, Yang Song, Yin Cui, Chen Sun, Alex Shepard, Hartwig Adam, Pietro Perona, and Serge J. Belongie, “The inaturalist species classification and detection dataset,” inIEEE Conf. Comput. Vis. Pattern Recog., 2018, pp. 8769– 8778

  59. [68]

    Open long-tailed recognition in a dynamic world,

    Ziwei Liu, Zhongqi Miao, Xiaohang Zhan, Jiayun Wang, Boqing Gong, and Stella X. Yu, “Open long-tailed recognition in a dynamic world,”IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 3, pp. 1836–1851, 2024

  60. [69]

    Learning multiple layers of features from tiny images,

    Alex Krizhevsky, Geoffrey Hinton, et al., “Learning multiple layers of features from tiny images,”Tech Report, 2009

  61. [70]

    Feature space augmentation for long-tailed data,

    Peng Chu, Xiao Bian, Shaopeng Liu, and Haibin Ling, “Feature space augmentation for long-tailed data,” inEur. Conf. Comput. Vis., 2020, vol. 12374, pp. 694–710

  62. [71]

    Imagenet large scale visual recognition challenge,

    Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei, “Imagenet large scale visual recognition challenge,”Int. J. Comput. Vis., vol. 115, no. 3, pp. 2...

  63. [72]

    Places: A 10 million image database for scene recognition,

    Bolei Zhou, Agata Lapedriza, Aditya Khosla, Aude Oliva, and Antonio Torralba, “Places: A 10 million image database for scene recognition,”IEEE Trans. Pattern Anal. Mach. Intell., vol. 40, no. 6, pp. 1452–1464, 2017

  64. [73]

    Memory-based jitter: JOURNAL OF LATEX CLASS FILES 13 Improving visual recognition on long-tailed data with diversity in memory,

    Jialun Liu, Wenhui Li, and Yifan Sun, “Memory-based jitter: JOURNAL OF LATEX CLASS FILES 13 Improving visual recognition on long-tailed data with diversity in memory,” inAAAI. 2022, pp. 1720–1728, AAAI Press

  65. [74]

    An optimal transport view of class-imbalanced visual recognition,

    Lianbao Jin, Dayu Lang, and Na Lei, “An optimal transport view of class-imbalanced visual recognition,”International Journal of Computer Vision, pp. 1–19, 2023

  66. [75]

    mixup: Beyond empirical risk minimization,

    Hongyi Zhang, Moustapha Cisse, Yann N Dauphin, and David Lopez-Paz, “mixup: Beyond empirical risk minimization,”ICLR, 2018

  67. [76]

    Exploring vision-language models for imbalanced learning,

    Yidong Wang, Zhuohao Yu, Jindong Wang, Qiang Heng, Hao Chen, Wei Ye, Rui Xie, Xing Xie, and Shikun Zhang, “Exploring vision-language models for imbalanced learning,”Int. J. Comput. Vis., vol. 132, no. 1, pp. 224–237, 2024

  68. [77]

    Bag of tricks for long-tailed visual recognition with deep convo- lutional neural networks,

    Yongshun Zhang, Xiu-Shen Wei, Boyan Zhou, and Jianxin Wu, “Bag of tricks for long-tailed visual recognition with deep convo- lutional neural networks,” inAAAI, 2021, pp. 3447–3455. Mengke Lireceived the B.S. degree in Com- munication Engineering from Southwest Univer- sity, Ch...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.