Pith. sign in

REVIEW 4 minor 1 cited by

Optimized Deferral for Imbalanced Settings

T0 review · 0 major / 4 minor · reviewed 2026-05-07 · grok-4.3

Pith's one-line read Casting deferral optimization as cost-sensitive learning over input-expert pairs yields algorithms that handle expert imbalance better.

desk verdict This paper recasts imbalanced deferral as cost-sensitive learning over input-expert pairs, derives margin losses for it, and introduces MILD, which shows gains on image and LLM tasks but keeps the new guarantees light on detail. read the letter →

arxiv 2604.27723 v1 submitted 2026-04-30 cs.LG stat.ML

classification cs.LGstat.ML
keywords learningtodeferexpertimbalancecost-sensitivemargin-basedlossesMILDalgorithmLLMroutingtwo-stagedeferralimbalancedclassification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tackles the expert imbalance problem in two-stage learning to defer, where standard methods tend to route everything to the most common expert and produce suboptimal accuracy. It reframes the entire deferral loss as a cost-sensitive classification task defined on the joint domain of inputs and experts rather than on inputs alone. From this view the authors derive fresh margin-based surrogate losses together with generalization guarantees, then build the MILD algorithm that puts these losses to work. The resulting deferral decisions improve both error rates and resource use in settings such as routing queries among several LLMs or classifiers. The work matters because many practical deferral pipelines already possess a fixed collection of experts whose usage frequencies are naturally skewed.

What carries the argument

The cost-sensitive learning formulation over the input-expert domain, which turns deferral into a weighted classification problem whose weights encode expert imbalance and enables margin-based losses with provable guarantees.

What would settle it

Train MILD and standard two-stage deferral baselines on a controlled dataset with known expert imbalance, then measure whether MILD produces a statistically significant drop in overall error rate or deferral cost; failure to do so would refute the central claim.

Watch

Extended reading notes

Core claim

We cast the deferral loss optimization as a novel cost-sensitive learning problem over the input-expert domain. We derive new margin-based loss functions and guarantees tailored to this setting, and develop novel algorithms for cost-sensitive learning. Leveraging these results, we design principled deferral algorithms, MILD (Margin-based Imbalanced Learning to Defer), specifically suited for expert imbalance settings.

Load-bearing premise

That treating deferral as cost-sensitive classification in the input-expert product space produces loss functions and algorithms whose theoretical guarantees translate into measurable gains on real imbalanced data.

Editorial extensions

If this is right

  • MILD outperforms existing deferral baselines on image classification tasks that exhibit expert imbalance.
  • MILD improves routing accuracy and efficiency when directing queries to collections of LLMs.
  • The new margin-based losses supply generalization bounds specific to the imbalanced deferral setting.
  • The cost-sensitive algorithms developed for the input-expert domain can be reused for other routing or selection tasks with skewed expert usage.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same input-expert cost-sensitive view could be applied to deferral pipelines that involve more than one deferral stage.
  • Similar cost-sensitive reductions might help balance load across heterogeneous models in distributed inference systems.
  • Empirical tests that vary the degree of imbalance while holding other factors fixed would clarify how MILD scales with skew severity.
  • The approach suggests examining whether cost-sensitive losses can also mitigate imbalance when the experts themselves are being trained rather than fixed.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

0 major / 4 minor

Summary. The manuscript examines two-stage learning to defer under expert imbalance, where policies tend to favor majority experts. It reformulates deferral optimization as cost-sensitive learning over the joint input-expert domain, derives new margin-based surrogate losses together with generalization guarantees, develops supporting algorithms for cost-sensitive learning, and introduces the MILD algorithm. Experiments on image classification and LLM routing tasks are reported to show improvements over existing baselines.

Significance. If the derived margin-based losses and associated guarantees are valid, the work supplies a principled, cost-sensitive treatment of expert imbalance that is directly relevant to practical deferral settings such as LLM routing. The modeling choice of operating in the input-expert space is coherent with existing cost-sensitive techniques and the empirical evaluation on both vision and language tasks provides concrete evidence of utility. The derivation of tailored losses and the focus on imbalance constitute the primary contributions.

minor comments (4)
  1. [§3.1] §3.1, Definition 1: the cost matrix C(x, e) is introduced without an explicit statement of how the imbalance ratios are encoded; a short paragraph clarifying the mapping from observed expert frequencies to the cost entries would improve readability.
  2. [§4.2] §4.2, Theorem 2: the generalization bound is stated in terms of the Rademacher complexity of the joint hypothesis class; it would be helpful to include a brief comparison (one sentence) to the corresponding bound for standard cost-sensitive classification to highlight the novelty of the input-expert formulation.
  3. [Figure 3] Figure 3 and Table 2: axis labels and legend entries use inconsistent abbreviations (e.g., “MILD” vs. “Mild”); uniform notation across all figures and tables is needed.
  4. [§5.3] §5.3: the LLM routing experiments report accuracy and deferral rate but do not include a statistical significance test across the five random seeds; adding p-values or confidence intervals would strengthen the empirical claims.

Simulated Author's Rebuttal

0 responses · 0 unresolved

We thank the referee for their careful reading of the manuscript and for the positive evaluation, including the accurate summary of our contributions and the recommendation for minor revision. The referee correctly identifies the core technical approach: reformulating two-stage deferral under expert imbalance as cost-sensitive learning over the input-expert domain, deriving margin-based surrogate losses with generalization guarantees, and introducing the MILD algorithm. We will incorporate any minor suggestions in the revised version.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; derivation self-contained via standard cost-sensitive modeling

full rationale

The paper casts deferral loss optimization as a cost-sensitive learning problem over the input-expert domain, then derives new margin-based losses, guarantees, and the MILD algorithm from that formulation. This modeling step is a coherent extension of existing imbalance-handling techniques rather than a self-definitional loop or a fitted parameter renamed as a prediction. The abstract explicitly presents the losses and algorithms as derived results, with no indication that they reduce by construction to the input data or to prior self-citations that bear the central claim. Experiments on image classification and LLM routing tasks supply external validation outside the derivation. No load-bearing equation or uniqueness theorem is shown to collapse to its own inputs.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

Abstract provides no details on parameters axioms or new entities.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Optimized Deferral for Imbalanced Settings." pith.science (2026). https://pith.science/paper/2604.27723

@misc{pith2026260427723,
  author       = {Pith},
  title        = {Pith review of: Optimized Deferral for Imbalanced Settings},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2604.27723}},
  note         = {Machine review of arXiv:2604.27723}
}
read the original abstract

Learning algorithms can be significantly improved by routing complex or uncertain inputs to specialized experts, balancing accuracy with computational cost. This approach, known as learning to defer, is essential in domains like natural language generation, medical diagnosis, and computer vision, where an effective deferral can reduce errors at low extra resource consumption. However, the two-stage learning to defer setting, which leverages existing predictors such as a collection of LLMs or other classifiers, often faces challenges due to an expert imbalance problem. This imbalance can lead to suboptimal performance, with deferral algorithms favoring the majority expert. We present a comprehensive study of two-stage learning to defer in expert imbalance settings. We cast the deferral loss optimization as a novel cost-sensitive learning problem over the input-expert domain. We derive new margin-based loss functions and guarantees tailored to this setting, and develop novel algorithms for cost-sensitive learning. Leveraging these results, we design principled deferral algorithms, MILD (Margin-based Imbalanced Learning to Defer), specifically suited for expert imbalance settings. Extensive experiments demonstrate the effectiveness of our approach, showing clear improvements over existing baselines on both image classification and real-world Large Language Model (LLM) routing tasks.

Figures

Figures reproduced from arXiv: 2604.27723 by the authors.

Figure 1
Figure 1. 1-D example with 4 classes (colored densities) and 3 experts (gray lines indicating accuracies). Experts have increasing costs indicated by darker shades of gray. not class labels. A 1-D example is provided in view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Online Learning-to-Defer with Varying Experts

    stat.ML 2026-05 unverdicted novelty 8.0 of 10

    Presents the first online learning-to-defer algorithm with regret bounds O((n + n_e) T^{2/3}) generally and O((n + n_e) sqrt(T)) under low noise for multiclass classification with varying experts.

Reference graph

Works this paper leans on

143 extracted references · 143 canonical work pages · cited by 1 Pith paper

  1. [1]

    H -consistency bounds for surrogate loss minimizers

    Awasthi, P., Mao, A., Mohri, M., and Zhong, Y. H -consistency bounds for surrogate loss minimizers. In International Conference on Machine Learning, pp.\ 1117--1174, 2022 a

  2. [2]

    Multi-class H -consistency bounds

    Awasthi, P., Mao, A., Mohri, M., and Zhong, Y. Multi-class H -consistency bounds. In Advances in Neural Information Processing Systems, pp.\ 782--795, 2022 b

  3. [3]

    L., Jordan, M

    Bartlett, P. L., Jordan, M. I., and McAuliffe, J. D. Convexity, classification, and risk bounds. Journal of the American Statistical Association, 101 0 (473): 0 138--156, 2006

  4. [4]

    Spectrally-normalized margin bounds for neural networks

    Bartlett, P. L., Foster, D. J., and Telgarsky, M. Spectrally-normalized margin bounds for neural networks. CoRR, abs/1706.08498, 2017

  5. [5]

    Sparks of Artificial General Intelligence: Early experiments with GPT-4

    Bubeck, S., Chandrasekaran, V., Eldan, R., Gehrke, J., Horvitz, E., Kamar, E., Lee, P., Lee, Y. T., Li, Y., Lundberg, S., et al. Sparks of artificial general intelligence: Early experiments with gpt-4. arXiv preprint arXiv:2303.12712, 2023

  6. [6]

    Learning imbalanced datasets with label-distribution-aware margin loss

    Cao, K., Wei, C., Gaidon, A., Arechiga, N., and Ma, T. Learning imbalanced datasets with label-distribution-aware margin loss. In Advances in Neural Information Processing Systems, 2019

  7. [7]

    Generalizing consistent multi-class classification with rejection to be compatible with arbitrary losses

    Cao, Y., Cai, T., Feng, L., Gu, L., Gu, J., An, B., Niu, G., and Sugiyama, M. Generalizing consistent multi-class classification with rejection to be compatible with arbitrary losses. In Advances in Neural Information Processing Systems, 2022

  8. [8]

    Sample efficient learning of predictors that complement humans

    Charusaie, M.-A., Mozannar, H., Sontag, D., and Samadi, S. Sample efficient learning of predictors that complement humans. In International Conference on Machine Learning, pp.\ 2972--3005, 2022

Show all 143 references
  1. [9]

    V., Bowyer, K

    Chawla, N. V., Bowyer, K. W., Hall, L. O., and Kegelmeyer, W. P. SMOTE : synthetic minority over-sampling technique. Journal of Artificial Intelligence Research, 16: 0 321--357, 2002

  2. [10]

    Regression with cost-based rejection

    Cheng, X., Cao, Y., Wang, H., Wei, H., An, B., and Feng, L. Regression with cost-based rejection. In Advances in Neural Information Processing Systems, 2023

  3. [11]

    Learning with rejection

    Cortes, C., DeSalvo, G., and Mohri, M. Learning with rejection. In International Conference on Algorithmic Learning Theory, pp.\ 67--82, 2016 a

  4. [12]

    Boosting with abstention

    Cortes, C., DeSalvo, G., and Mohri, M. Boosting with abstention. In Advances in Neural Information Processing Systems, pp.\ 1660--1668, 2016 b

  5. [13]

    Structured prediction theory based on factor graph complexity

    Cortes, C., Kuznetsov, V., Mohri, M., and Yang, S. Structured prediction theory based on factor graph complexity. In Advances in Neural Information Processing Systems, 2016 c

  6. [14]

    Adanet: Adaptive structural learning of artificial neural networks

    Cortes, C., Gonzalvo, X., Kuznetsov, V., Mohri, M., and Yang, S. Adanet: Adaptive structural learning of artificial neural networks. In International Conference on Machine Learning, pp.\ 874--883, 2017

  7. [15]

    Theory and algorithms for learning with rejection in binary classification

    Cortes, C., DeSalvo, G., and Mohri, M. Theory and algorithms for learning with rejection in binary classification. Annals of Mathematics and Artificial Intelligence, 92 0 (2): 0 277--315, 2024 a

  8. [16]

    Cardinality-aware set prediction and top- k classification

    Cortes, C., Mao, A., Mohri, C., Mohri, M., and Zhong, Y. Cardinality-aware set prediction and top- k classification. In Advances in Neural Information Processing Systems, 2024 b

  9. [17]

    Balancing the scales: A theoretical and algorithmic framework for learning from imbalanced data

    Cortes, C., Mao, A., Mohri, M., and Zhong, Y. Balancing the scales: A theoretical and algorithmic framework for learning from imbalanced data. In International Conference on Machine Learning, 2025

  10. [18]

    A theoretical framework for modular learning of robust generative models

    Cortes, C., Mohri, M., and Zhong, Y. A theoretical framework for modular learning of robust generative models. In International Conference on Machine Learning, 2026

  11. [19]

    Parametric contrastive learning

    Cui, J., Zhong, Z., Liu, S., Yu, B., and Jia, J. Parametric contrastive learning. In International Conference on Computer Vision, 2021

  12. [20]

    Reslt: Residual learning for long-tailed recognition

    Cui, J., Liu, S., Tian, Z., Zhong, Z., and Jia, J. Reslt: Residual learning for long-tailed recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2022

  13. [21]

    Class-balanced loss based on effective number of samples

    Cui, Y., Jia, M., Lin, T.-Y., Song, Y., and Belongie, S. Class-balanced loss based on effective number of samples. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.\ 9268--9277, 2019

  14. [22]

    Regression under human assistance

    De, A., Koley, P., Ganguly, N., and Gomez-Rodriguez, M. Regression under human assistance. In Proceedings of the AAAI Conference on Artificial Intelligence, pp.\ 2611--2620, 2020

  15. [23]

    Budgeted multiple-expert deferral

    DeSalvo, G., Mohri, C., Mohri, M., and Zhong, Y. Budgeted multiple-expert deferral. arXiv preprint arXiv:2510.26706, 2025

  16. [24]

    Simpro: A simple probabilistic framework towards realistic long-tailed semi-supervised learning

    Du, C., Han, Y., and Huang, G. Simpro: A simple probabilistic framework towards realistic long-tailed semi-supervised learning. In International Conference on Machine Learning, 2024

  17. [25]

    and Wiener, Y

    El-Yaniv, R. and Wiener, Y. Active learning via perfect selective classification. Journal of Machine Learning Research, 13 0 (2), 2012

  18. [26]

    El-Yaniv, R. et al. On the foundations of noise-free selective classification. Journal of Machine Learning Research, 11 0 (5), 2010

  19. [27]

    The foundations of cost-sensitive learning

    Elkan, C. The foundations of cost-sensitive learning. In International Joint Conference on Artificial Intelligence, 2001

  20. [28]

    A multiple resampling method for learning from imbalanced data sets

    Estabrooks, A., Jo, T., and Japkowicz, N. A multiple resampling method for learning from imbalanced data sets. Computational Intelligence, 20 0 (1): 0 18--36, 2004

  21. [29]

    Learning with average top-k loss

    Fan, Y., Lyu, S., Ying, Y., and Hu, B. Learning with average top-k loss. In Advances in Neural Information Processing Systems, pp.\ 497--505, 2017

  22. [30]

    Gabidolla, M., Zharmagambetov, A., and Carreira - Perpi \ n \' a n, M. \' A . Beyond the ROC curve: Classification trees using cost-optimal curves, with application to imbalanced datasets. In International Conference on Machine Learning, 2024

  23. [31]

    Enhancing minority classes by mixing: an adaptative optimal transport approach for long-tailed classification

    Gao, J., Zhao, H., Li, Z., and Guo, D. Enhancing minority classes by mixing: an adaptative optimal transport approach for long-tailed classification. In Advances in Neural Information Processing Systems, 2023

  24. [32]

    Distribution alignment optimization through neural collapse for long-tailed classification

    Gao, J., Zhao, H., dan Guo, D., and Zha, H. Distribution alignment optimization through neural collapse for long-tailed classification. In International Conference on Machine Learning, 2024

  25. [33]

    and El-Yaniv, R

    Geifman, Y. and El-Yaniv, R. Selective classification for deep neural networks. In Advances in Neural Information Processing Systems, 2017

  26. [34]

    and El-Yaniv, R

    Geifman, Y. and El-Yaniv, R. Selectivenet: A deep neural network with an integrated reject option. In International Conference on Machine Learning, pp.\ 2151--2159, 2019

  27. [35]

    Wrapped cauchy distributed angular softmax for long-tailed visual recognition

    Han, B. Wrapped cauchy distributed angular softmax for long-tailed visual recognition. In International Conference on Machine Learning, pp.\ 12368--12388, 2023

  28. [36]

    Borderline-smote: a new over-sampling method in imbalanced data sets learning

    Han, H., Wang, W.-Y., and Mao, B.-H. Borderline-smote: a new over-sampling method in imbalanced data sets learning. In International Conference on Intelligent Computing, pp.\ 878--887, 2005

  29. [37]

    Deep residual learning for image recognition

    He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.\ 770--778, 2016

  30. [38]

    Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing

    He, P., Gao, J., and Chen, W. Debertav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing. arXiv preprint arXiv:2111.09543, 2021

  31. [39]

    Measuring massive multitask language understanding

    Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020

  32. [40]

    Disentangling label distribution for long-tailed visual recognition

    Hong, Y., Han, S., Choi, K., Seo, S., Kim, B., and Chang, B. Disentangling label distribution for long-tailed visual recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021

  33. [41]

    Cost-sensitive support vector machines

    Iranmehr, A., Masnadi - Shirazi, H., and Vasconcelos, N. Cost-sensitive support vector machines. Neurocomputing, 343: 0 50--64, 2019

  34. [42]

    A., Brown, M., Yang, M.-H., Wang, L., and Gong, B

    Jamal, M. A., Brown, M., Yang, M.-H., Wang, L., and Gong, B. Rethinking class-balanced methods for long-tailed visual recognition from a domain adaptation perspective. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.\ 7610--7619, 2020

  35. [43]

    Risk-controlled selective prediction for regression deep neural network models

    Jiang, W., Zhao, Y., and Wang, Z. Risk-controlled selective prediction for regression deep neural network models. In International Joint Conference on Neural Networks, pp.\ 1--8, 2020

  36. [44]

    Balanced meta-softmax for long-tailed visual recognition

    Jiawei, R., Yu, C., Ma, X., Zhao, H., Yi, S., et al. Balanced meta-softmax for long-tailed visual recognition. In Advances in Neural Information Processing Systems, 2020

  37. [45]

    Decoupling representation and classifier for long-tailed recognition

    Kang, B., Xie, S., Rohrbach, M., Yan, Z., Gordo, A., Feng, J., and Kalantidis, Y. Decoupling representation and classifier for long-tailed recognition. In International Conference on Learning Representations, 2020

  38. [46]

    Maximum class separation as inductive bias in one matrix

    Kasarla, T., Burghouts, G., Van Spengler, M., Van Der Pol, E., Cucchiara, R., and Mettes, P. Maximum class separation as inductive bias in one matrix. In Advances in Neural Information Processing Systems, pp.\ 19553--19566, 2022

  39. [47]

    Towards unbiased and accurate deferral to multiple experts

    Keswani, V., Lease, M., and Kenthapadi, K. Towards unbiased and accurate deferral to multiple experts. In Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, pp.\ 154--165, 2021

  40. [48]

    W., Shen, J., and Shao, L

    Khan, S., Hayat, M., Zamir, S. W., Shen, J., and Shao, L. Striking the right balance with uncertainty. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.\ 103--112, 2019

  41. [49]

    Kingma, D. P. and Ba, J. Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980, 2014

  42. [50]

    R., Paraskevas, O., Oymak, S., and Thrampoulidis, C

    Kini, G. R., Paraskevas, O., Oymak, S., and Thrampoulidis, C. Label-imbalanced and group-sensitive classification under overparameterization. In Advances in Neural Information Processing Systems, pp.\ 18970--18983, 2021

  43. [51]

    Learning multiple layers of features from tiny images

    Krizhevsky, A. Learning multiple layers of features from tiny images. Technical report, Toronto University, 2009

  44. [52]

    and Matwin, S

    Kubat, M. and Matwin, S. Addressing the curse of imbalanced training sets: One-sided selection. In International Conference on Machine Learning, 1997

  45. [53]

    and Yang, X

    Le, Y. and Yang, X. Tiny imagenet visual recognition challenge. CS 231N, 7 0 (7): 0 3, 2015

  46. [54]

    Size-invariance matters: Rethinking metrics and losses for imbalanced multi-object salient object detection

    Li, F., Xu, Q., Bao, S., Yang, Z., Cong, R., Cao, X., and Huang, Q. Size-invariance matters: Rethinking metrics and losses for imbalanced multi-object salient object detection. In International Conference on Machine Learning, 2024

  47. [55]

    When no-rejection learning is optimal for regression with rejection

    Li, X., Liu, S., Sun, C., and Wang, H. When no-rejection learning is optimal for regression with rejection. arXiv preprint arXiv:2307.02932, 2023

  48. [56]

    Focal loss for dense object detection

    Lin, T.-Y., Goyal, P., Girshick, R., He, K., and Doll \'a r, P. Focal loss for dense object detection. In International Conference on Computer Vision, pp.\ 2980--2988, 2017

  49. [57]

    Incorporating uncertainty in learning to defer algorithms for safe computer-aided diagnosis

    Liu, J., Gallego, B., and Barbieri, S. Incorporating uncertainty in learning to defer algorithms for safe computer-aided diagnosis. Scientific Reports, 12 0 (1): 0 1762, 2022

  50. [58]

    Elta: An enhancer against long-tail for aesthetics-oriented models

    Liu, L., He, S., Ming, A., Xie, R., and Ma, H. Elta: An enhancer against long-tail for aesthetics-oriented models. In International Conference on Machine Learning, 2024

  51. [59]

    Exploratory undersampling for class-imbalance learning

    Liu, X.-Y., Wu, J., and Zhou, Z.-H. Exploratory undersampling for class-imbalance learning. IEEE Transactions on Systems, Man, and Cybernetics, 39 0 (2): 0 539--550, 2008

  52. [60]

    Liu, Z., Miao, Z., Zhan, X., Wang, J., Gong, B., and Yu, S. X. Large-scale long-tailed recognition in an open world. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.\ 2537--2546, 2019

  53. [61]

    and Hutter, F

    Loshchilov, I. and Hutter, F. Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017

  54. [62]

    Learning adversarially fair and transferable representations

    Madras, D., Creager, E., Pitassi, T., and Zemel, R. Learning adversarially fair and transferable representations. arXiv preprint arXiv:1802.06309, 2018

  55. [63]

    Theory and Algorithms for Learning with Multi-Class Abstention and Multi-Expert Deferral

    Mao, A. Theory and Algorithms for Learning with Multi-Class Abstention and Multi-Expert Deferral. PhD thesis, New York University, 2025

  56. [64]

    Two-stage learning to defer with multiple experts

    Mao, A., Mohri, C., Mohri, M., and Zhong, Y. Two-stage learning to defer with multiple experts. In Advances in Neural Information Processing Systems, 2023 a

  57. [65]

    H -consistency bounds: Characterization and extensions

    Mao, A., Mohri, M., and Zhong, Y. H -consistency bounds: Characterization and extensions. In Advances in Neural Information Processing Systems, 2023 b

  58. [66]

    H -consistency bounds for pairwise misranking loss surrogates

    Mao, A., Mohri, M., and Zhong, Y. H -consistency bounds for pairwise misranking loss surrogates. In International Conference on Machine learning, 2023 c

  59. [67]

    Ranking with abstention

    Mao, A., Mohri, M., and Zhong, Y. Ranking with abstention. In ICML 2023 Workshop The Many Facets of Preference-Based Learning, 2023 d

  60. [68]

    Cross-entropy loss functions: Theoretical analysis and applications

    Mao, A., Mohri, M., and Zhong, Y. Cross-entropy loss functions: Theoretical analysis and applications. In International Conference on Machine Learning, 2023 e

  61. [69]

    Structured prediction with stronger consistency guarantees

    Mao, A., Mohri, M., and Zhong, Y. Structured prediction with stronger consistency guarantees. In Advances in Neural Information Processing Systems, pp.\ 46903--46937, 2023 f

  62. [70]

    Principled approaches for learning to defer with multiple experts

    Mao, A., Mohri, M., and Zhong, Y. Principled approaches for learning to defer with multiple experts. In International Symposium on Artificial Intelligence and Mathematics, 2024 a

  63. [71]

    Predictor-rejector multi-class abstention: Theoretical analysis and algorithms

    Mao, A., Mohri, M., and Zhong, Y. Predictor-rejector multi-class abstention: Theoretical analysis and algorithms. In International Conference on Algorithmic Learning Theory, pp.\ 822--867, 2024 b

  64. [72]

    Theoretically grounded loss functions and algorithms for score-based multi-class abstention

    Mao, A., Mohri, M., and Zhong, Y. Theoretically grounded loss functions and algorithms for score-based multi-class abstention. In International Conference on Artificial Intelligence and Statistics, pp.\ 4753--4761, 2024 c

  65. [73]

    H -consistency guarantees for regression

    Mao, A., Mohri, M., and Zhong, Y. H -consistency guarantees for regression. In International Conference on Machine Learning, pp.\ 34712--34737, 2024 d

  66. [74]

    Multi-label learning with stronger consistency guarantees

    Mao, A., Mohri, M., and Zhong, Y. Multi-label learning with stronger consistency guarantees. In Advances in Neural Information Processing Systems, 2024 e

  67. [75]

    Regression with multi-expert deferral

    Mao, A., Mohri, M., and Zhong, Y. Regression with multi-expert deferral. In International Conference on Machine Learning, pp.\ 34738--34759, 2024 f

  68. [76]

    A universal growth rate for learning with smooth surrogate losses

    Mao, A., Mohri, M., and Zhong, Y. A universal growth rate for learning with smooth surrogate losses. In Advances in Neural Information Processing Systems, 2024 g

  69. [77]

    Realizable H -consistent and B ayes-consistent loss functions for learning to defer

    Mao, A., Mohri, M., and Zhong, Y. Realizable H -consistent and B ayes-consistent loss functions for learning to defer. In Advances in Neural Information Processing Systems, 2024 h

  70. [78]

    Mastering multiple-expert routing: Realizable H -consistency and strong guarantees for learning to defer

    Mao, A., Mohri, M., and Zhong, Y. Mastering multiple-expert routing: Realizable H -consistency and strong guarantees for learning to defer. In International Conference on Machine Learning, 2025 a

  71. [79]

    Principled algorithms for optimizing generalized metrics in binary classification

    Mao, A., Mohri, M., and Zhong, Y. Principled algorithms for optimizing generalized metrics in binary classification. In International Conference on Machine Learning, 2025 b

  72. [80]

    Enhanced -consistency bounds

    Mao, A., Mohri, M., and Zhong, Y. Enhanced -consistency bounds. In International Conference on Algorithmic Learning Theory, 2025 c

  73. [81]

    and Vasconcelos, N

    Masnadi-Shirazi, H. and Vasconcelos, N. Risk minimization, probability elicitation, and cost-sensitive SVM s. In International Conference on Machine Learning, 2010

  74. [82]

    A vector-contraction inequality for rademacher complexities

    Maurer, A. A vector-contraction inequality for rademacher complexities. In International Conference on Algorithmic Learning Theory, 2016

  75. [83]

    Learning from rich semantics and coarse locations for long-tailed object detection

    Meng, L., Dai, X., Yang, J., Chen, D., Chen, Y., Liu, M., Chen, Y.-L., Wu, Z., Yuan, L., and Jiang, Y.-G. Learning from rich semantics and coarse locations for long-tailed object detection. In Advances in Neural Information Processing Systems, 2023

  76. [84]

    K., Jayasumana, S., Rawat, A

    Menon, A. K., Jayasumana, S., Rawat, A. S., Jain, H., Veit, A., and Kumar, S. Long-tail learning via logit adjustment. In International Conference on Learning Representations, 2021

  77. [85]

    Learning to reject with a fixed predictor: Application to decontextualization

    Mohri, C., Andor, D., Choi, E., Collins, M., Mao, A., and Zhong, Y. Learning to reject with a fixed predictor: Application to decontextualization. In International Conference on Learning Representations, 2024

  78. [86]

    and Zhong, Y

    Mohri, M. and Zhong, Y. Mind the gap: Structure-aware consistency in preference learning. In International Conference on Machine Learning, 2026 a

  79. [87]

    and Zhong, Y

    Mohri, M. and Zhong, Y. Linear-core surrogates: Smooth loss functions with linear rates for classification and structured prediction. In International Conference on Machine Learning, 2026 b

  80. [88]

    and Zhong, Y

    Mohri, M. and Zhong, Y. Beyond tsybakov: Model margin noise and H -consistency bounds. In International Symposium on Artificial Intelligence and Mathematics, 2026 c

  81. [89]

    Foundations of Machine Learning

    Mohri, M., Rostamizadeh, A., and Talwalkar, A. Foundations of Machine Learning. MIT Press, second edition, 2018

  82. [90]

    H., Carlier, A., Ng, L

    Montreuil, Y., Yeo, S. H., Carlier, A., Ng, L. X., and Ooi, W. T. Optimal query allocation in extractive QA with LLMs : A learning-to-defer framework with theoretical guarantees. arXiv preprint arXiv:2410.15761, 2024

  83. [91]

    X., and Ooi, W

    Montreuil, Y., Carlier, A., Ng, L. X., and Ooi, W. T. Adversarial robustness in two-stage learning-to-defer: Algorithms and guarantees. In International Conference on Machine Learning, 2025 a

  84. [92]

    X., and Ooi, W

    Montreuil, Y., Carlier, A., Ng, L. X., and Ooi, W. T. Why ask one when you can ask k ? learning-to-defer to the top- k experts. arXiv preprint arXiv:2504.12988, 2025 b

  85. [93]

    H., Carlier, A., Ng, L

    Montreuil, Y., Yeo, S. H., Carlier, A., Ng, L. X., and Ooi, W. T. A two-stage learning-to-defer approach for multi-task learning. In International Conference on Machine Learning, 2025 c

  86. [94]

    and Sontag, D

    Mozannar, H. and Sontag, D. Consistent estimators for learning to defer to an expert. In International Conference on Machine Learning, pp.\ 7076--7087, 2020

  87. [95]

    Who should predict? exact algorithms for learning to defer to humans

    Mozannar, H., Lang, H., Wei, D., Sattigeri, P., Das, S., and Sontag, D. Who should predict? exact algorithms for learning to defer to humans. In International Conference on Artificial Intelligence and Statistics, pp.\ 10520--10545, 2023

  88. [96]

    K., Rawat, A

    Narasimhan, H., Jitkrittum, W., Menon, A. K., Rawat, A. S., and Kumar, S. Post-hoc estimators for learning to defer to an expert. In Advances in Neural Information Processing Systems, pp.\ 29292--29304, 2022

  89. [97]

    Netzer, Y., Wang, T., Coates, A., Bissacco, A., Wu, B., and Ng, A. Y. Reading digits in natural images with unsupervised feature learning. In Advances in Neural Information Processing Systems, 2011

  90. [98]

    Norm-based capacity control in neural networks

    Neyshabur, B., Tomioka, R., and Srebro, N. Norm-based capacity control in neural networks. CoRR, abs/1503.00036, 2015

  91. [99]

    Differentiable learning under triage

    Okati, N., De, A., and Rodriguez, M. Differentiable learning under triage. In Advances in Neural Information Processing Systems, pp.\ 9140--9151, 2021

  92. [100]

    F., Zazo, J., Parbhoo, S., Perlis, R

    Pradier, M. F., Zazo, J., Parbhoo, S., Perlis, R. H., Zazzi, M., and Doshi-Velez, F. Preferential mixture-of-experts: Interpretable models that rely on human expertise as much as possible. AMIA Summits on Translational Science Proceedings, 2021: 0 525, 2021

  93. [101]

    and Liu, Y

    Qiao, X. and Liu, Y. Adaptive weighted learning for unbalanced multicategory classification. Biometrics, 65: 0 159--68, 2008

  94. [102]

    The algorithmic automation problem: Prediction, triage, and human effort

    Raghu, M., Blumer, K., Corrado, G., Kleinberg, J., Obermeyer, Z., and Mullainathan, S. The algorithmic automation problem: Prediction, triage, and human effort. arXiv preprint arXiv:1903.12220, 2019

  95. [103]

    and Yee, M

    Raman, N. and Yee, M. Improving learning-to-defer algorithms through fine-tuning. arXiv preprint arXiv:2112.10768, 2021

  96. [104]

    K., Das, S., Panda, R., Sattigeri, P., and Wornell, G

    Shah, A., Bu, Y., Lee, J. K., Das, S., Panda, R., Sattigeri, P., and Wornell, G. W. Selective regression under fairness criteria. In International Conference on Machine Learning, pp.\ 19598--19615, 2022

  97. [105]

    Long-tail learning with foundation model: Heavy fine-tuning hurts

    Shi, J.-X., Wei, T., Zhou, Z., Shao, J.-J., Han, X.-Y., and Li, Y.-F. Long-tail learning with foundation model: Heavy fine-tuning hurts. In International Conference on Machine Learning, 2024

  98. [106]

    How to compare different loss functions and their risks

    Steinwart, I. How to compare different loss functions and their risks. Constructive Approximation, 26 0 (2): 0 225--287, 2007

  99. [107]

    S., Wong, A

    Sun, Y., Kamel, M. S., Wong, A. K., and Wang, Y. Cost-sensitive boosting for classification of imbalanced data. Pattern Recognition, 40 0 (12): 0 3358--3378, 2007

  100. [108]

    Learning to defer to a population: A meta-learning approach

    Tailor, D., Patra, A., Verma, R., Manggala, P., and Nalisnick, E. Learning to defer to a population: A meta-learning approach. In International Conference on Artificial Intelligence and Statistics, pp.\ 3475--3483, 2024

  101. [109]

    Equalization loss for long-tailed object recognition

    Tan, J., Wang, C., Li, B., Li, Q., Ouyang, W., Yin, C., and Yan, J. Equalization loss for long-tailed object recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.\ 11662--11671, 2020

  102. [110]

    Long-tailed classification by keeping the good and removing the bad momentum causal effect

    Tang, K., Huang, J., and Zhang, H. Long-tailed classification by keeping the good and removing the bad momentum causal effect. In Advances in Neural Information Processing Systems, 2020

  103. [111]

    Posterior re-calibration for imbalanced datasets

    Tian, J., Liu, Y.-C., Glaser, N., Hsu, Y.-C., and Kira, Z. Posterior re-calibration for imbalanced datasets. In Advances in Neural Information Processing Systems, 2020

  104. [112]

    M., and Napolitano, A

    Van Hulse, J., Khoshgoftaar, T. M., and Napolitano, A. Experimental perspectives on learning from imbalanced data. In International Conference on Machine Learning, 2007

  105. [113]

    and Nalisnick, E

    Verma, R. and Nalisnick, E. Calibrated learning to defer with one-vs-all classifiers. In International Conference on Machine Learning, pp.\ 22184--22202, 2022

  106. [114]

    Learning to defer to multiple experts: Consistent surrogate losses, confidence calibration, and conformal ensembles

    Verma, R., Barrej \'o n, D., and Nalisnick, E. Learning to defer to multiple experts: Consistent surrogate losses, confidence calibration, and conformal ensembles. In International Conference on Artificial Intelligence and Statistics, pp.\ 11415--11434, 2023

  107. [115]

    C., Small, K., Brodley, C

    Wallace, B. C., Small, K., Brodley, C. E., and Trikalinos, T. A. Class imbalance, redux. In International Conference on Data Mining, pp.\ 754--763, 2011

  108. [116]

    Rsg: A simple but effective module for learning imbalanced datasets

    Wang, J., Lukasiewicz, T., Hu, X., Cai, J., and Xu, Z. Rsg: A simple but effective module for learning imbalanced datasets. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.\ 3784--3793, 2021 a

  109. [117]

    Wang, X., Lian, L., Miao, Z., Liu, Z., and Yu, S. X. Long-tailed recognition by routing diverse distribution-aware experts. In International Conference on Learning Representations, 2021 b

  110. [118]

    H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W

    Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., Chi, E. H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W. Emergent abilities of large language models. CoRR, abs/2206.07682, 2022

  111. [119]

    Learning label shift correction for test-agnostic long-tailed recognition

    Wei, T., Mao, Z., Zhou, Z.-H., Wan, Y., and Zhang, M.-L. Learning label shift correction for test-agnostic long-tailed recognition. In International Conference on Machine Learning, 2024

  112. [120]

    and El-Yaniv, R

    Wiener, Y. and El-Yaniv, R. Agnostic selective classification. In Advances in Neural Information Processing Systems, 2011

  113. [121]

    and El-Yaniv, R

    Wiener, Y. and El-Yaniv, R. Pointwise tracking the optimal regression function. In Advances in Neural Information Processing Systems, 2012

  114. [122]

    and El-Yaniv, R

    Wiener, Y. and El-Yaniv, R. Agnostic pointwise-competitive selective classification. Journal of Artificial Intelligence Research, 52: 0 171--201, 2015

  115. [123]

    Learning to complement humans

    Wilder, B., Horvitz, E., and Kamar, E. Learning to complement humans. In International Joint Conferences on Artificial Intelligence, pp.\ 1526--1533, 2021

  116. [124]

    Learning from multiple experts: Self-paced knowledge distillation for long-tailed classification

    Xiang, L., Ding, G., and Han, J. Learning from multiple experts: Self-paced knowledge distillation for long-tailed classification. In European Conference on Computer Vision, pp.\ 247--263, 2020

  117. [125]

    Yang, A., Yu, B., Li, C., Liu, D., Huang, F., Huang, H., Jiang, J., Tu, J., Zhang, J., Zhou, J., et al. Qwen2. 5-1m technical report. arXiv preprint arXiv:2501.15383, 2025

  118. [126]

    Yang, Y., Chen, S., Li, X., Xie, L., Lin, Z., and Tao, D. Inducing neural collapse in imbalanced learning: Do we really need a learnable classifier at the end of deep neural network? In Advances in Neural Information Processing Systems, pp.\ 37991--38002, 2022

  119. [127]

    Harnessing hierarchical label distribution variations in test agnostic long-tail recognition

    Yang, Z., Xu, Q., Wang, Z., Li, S., Han, B., Bao, S., Cao, X., and Huang, Q. Harnessing hierarchical label distribution variations in test agnostic long-tail recognition. In International Conference on Machine Learning, 2024

  120. [128]

    Identifying and compensating for feature deviation in imbalanced deep learning, 2020

    Ye, H.-J., Chen, H.-Y., Zhan, D.-C., and Chao, W.-L. Identifying and compensating for feature deviation in imbalanced deep learning, 2020

  121. [129]

    Regression with reject option and application to knn

    Zaoui, A., Denis, C., and Hebiri, M. Regression with reject option and application to knn. In Advances in Neural Information Processing Systems, pp.\ 20073--20082, 2020

  122. [130]

    Statistical behavior and consistency of classification methods based on convex risk minimization

    Zhang, T. Statistical behavior and consistency of classification methods based on convex risk minimization. The Annals of Statistics, 32 0 (1): 0 56--85, 2004

  123. [131]

    Online adaptive asymmetric active learning for budgeted imbalanced data

    Zhang, Y., Zhao, P., Cao, J., Ma, W., Huang, J., Wu, Q., and Tan, M. Online adaptive asymmetric active learning for budgeted imbalanced data. In SIGKDD International Conference on Knowledge Discovery & Data Mining, pp.\ 2768--2777, 2018

  124. [132]

    Online adaptive asymmetric active learning with limited budgets

    Zhang, Y., Zhao, P., Niu, S., Wu, Q., Cao, J., Huang, J., and Tan, M. Online adaptive asymmetric active learning with limited budgets. IEEE Transactions on Knowledge and Data Engineering, 2019

  125. [133]

    Self-supervised aggregation of diverse experts for test-agnostic long-tailed recognition

    Zhang, Y., Hooi, B., Hong, L., and Feng, J. Self-supervised aggregation of diverse experts for test-agnostic long-tailed recognition. In Advances in Neural Information Processing Systems, 2022

  126. [134]

    Deep long-tailed learning: A survey

    Zhang, Y., Kang, B., Hooi, B., Yan, S., and Feng, J. Deep long-tailed learning: A survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45 0 (9): 0 10795--10816, 2023

  127. [135]

    and Pfister, T

    Zhang, Z. and Pfister, T. Learning fast sample re-weighting without reward data. In International Conference on Computer Vision, 2021

  128. [136]

    C., Tan, M., and Huang, J

    Zhao, P., Zhang, Y., Wu, M., Hoi, S. C., Tan, M., and Huang, J. Adaptive cost-sensitive online classification. IEEE Transactions on Knowledge and Data Engineering, 31 0 (2): 0 214--228, 2018

  129. [137]

    Fundamental Novel Consistency Theory: H-Consistency Bounds

    Zhong, Y. Fundamental Novel Consistency Theory: H-Consistency Bounds. PhD thesis, New York University, 2025

  130. [138]

    Improving calibration for long-tailed recognition

    Zhong, Z., Cui, J., Liu, S., and Jia, J. Improving calibration for long-tailed recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021

  131. [139]

    Bbn: Bilateral-branch network with cumulative learning for long-tailed visual recognition

    Zhou, B., Cui, Q., Wei, X.-S., and Chen, Z.-M. Bbn: Bilateral-branch network with cumulative learning for long-tailed visual recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp.\ 9719--9728, 2020

  132. [140]

    and Liu, X.-Y

    Zhou, Z.-H. and Liu, X.-Y. Training cost-sensitive neural networks with methods addressing the class imbalance problem. IEEE Transactions on Knowledge and Data Engineering, 18 0 (1): 0 63--77, 2005

  133. [141]

    Generalized logit adjustment: Calibrating fine-tuned models by removing label bias in foundation models

    Zhu, B., Tang, K., Sun, Q., and Zhang, H. Generalized logit adjustment: Calibrating fine-tuned models by removing label bias in foundation models. In Advances in Neural Information Processing Systems, pp.\ 64663--64680, 2023

  134. [142]

    Generative active learning for long-tailed instance segmentation

    Zhu, M., Fan, C., Chen, H., Liu, Y., Mao, W., Xu, X., and Shen, C. Generative active learning for long-tailed instance segmentation. In International Conference on Machine Learning, 2024

  135. [143]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed May 7, 2026 · model on record in the stance chip above.