Pith. sign in

REVIEW 2 major objections 5 minor 42 references

Improving Task Diversity in Label Efficient Supervised Finetuning of LLMs

T0 review · 2 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read This paper argues that a simple inverse-confidence task allocation can beat full-dataset LLM finetuning with up to 80% fewer labels.

desk verdict A simple task-level inverse-confidence selection method with a promising MMLU result, but the confidence metric is confounded with output length and the abstract overclaims consistency. read the letter →

arxiv 2507.21482 v1 pith:VWCELGOT submitted 2025-07-29 cs.CL cs.AI

classification cs.CLcs.AI
keywords label-efficientsupervisedfinetuningtaskdiversityinverseconfidenceweightinground-robinsamplinginstructiontuningdataselectionone-batchactivelearninguncertaintyquantification
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that task-level diversity, rather than prompt-level diversity in an embedding space, is the right organizing principle for label-efficient supervised finetuning of LLMs. Its Weighted Task Diversity rule gives every task a small base budget and then spends the remaining annotation budget on tasks in proportion to the inverse of the base model's average confidence, so low-confidence tasks receive more labels. On the Dolly dataset, models finetuned on 3K labels selected this way score 39.74 on MMLU versus 35.33 for models trained on the full 13.5K pool, a gain of more than 4 points with roughly 80% fewer labels. On FLAN V2, the same strategy matches full-data training at half the budget and achieves the highest AlpacaEval win rates. If correct, this means a cheap, reusable confidence signal can replace elaborate embedding-based and uncertainty-based selectors.

What carries the argument

The load-bearing object is the task-level inverse-confidence allocation $\alpha_t \propto 1/\mathrm{conf}_t$, realized through a round-robin sampler. Task confidence is defined as $\mathrm{conf}_t = \frac{1}{|X_t|}\sum_{x \in X_t} \prod_{j=1}^m g(y_j \mid y_{<j}, x)$, the average over a task's prompts of the product of token probabilities of the base model's generated response. The allocation first gives every task a small base budget, then distributes the remainder in proportion to $1/\mathrm{conf}_t$, clamped to each task's available prompts. The round-robin pass converts fractional allocations into an integer selection and enforces that small tasks are covered before the budget moves on. This machinery does two jobs at once: it guarantees task coverage, and it shifts labels toward tasks where the base model is most uncertain.

What would settle it

Run Weighted Task Diversity with confidence normalized by response length or computed as a per-token geometric mean, holding the budget, pool, and training recipe fixed; if the 3K-label MMLU advantage over full-data training disappears, the reported gain rests on the length confound rather than on task difficulty.

Watch

Extended reading notes

Core claim

The central claim is that a pretrained model's own confidence, averaged per task, is a usable signal for deciding which instruction-tuning prompts to pay humans to label. Under Weighted Task Diversity, each task first receives a small coverage floor (the paper uses 5 examples), and the remaining budget is split across tasks proportionally to $1/\mathrm{conf}_t$, where $\mathrm{conf}_t$ is the mean over prompts in task $t$ of the product of token probabilities of the base model's generated response; a round-robin procedure then selects examples uniformly within each task. The paper reports that on Dolly, this selection at a 3K budget reaches 39.74 MMLU versus 35.33 for training on the full 13.5K pool, that on FLAN V2 a 45K budget reaches the same MMLU as all 90K examples, and that the selected subsets win more often than full-data-trained models in GPT-4-judged AlpacaEval comparisons. These results are presented as evidence that task diversity plus uncertainty weighting is at least as powerful as, and much simpler than, embedding-based diversity methods.

Load-bearing premise

The method rests on treating the base model's average per-task confidence as a trustworthy signal of which tasks most need labels; if confidence mostly reflects pretraining exposure, output length, or decoding style rather than labeling value, the inverse-confidence allocation points at the wrong tasks.

Editorial extensions

If this is right

  • On the Dolly pool, a 3K-label Weighted Task Diversity selection produces an MMLU score of 39.74, higher than the 35.33 from training on all 13.5K labels, an annotation saving of roughly 80% with an accuracy gain.
  • On FLAN V2, the same strategy reaches full-data MMLU performance at a 45K budget, cutting the annotation load in half at matched accuracy.
  • Weighted Task Diversity and plain Task Diversity match or beat all tested baselines—random, entropy, confidence, margin, k-center, facility location, DPP, and ActiveIT—on AlpacaEval win rates at matched budgets.
  • The method needs no precomputed embeddings, no clustering, no ground-truth-based quality scores, and no careful kernel hyperparameters; task labels already present in instruction datasets suffice.
  • Because allocation happens before any annotation is collected, the rule is immediately usable in the one-batch active learning setting where ground-truth responses are unknown.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: an unnormalized product-of-token-probabilities confidence conflates task difficulty with response length and decoding choices; a length-normalized or per-token-geometric-mean version is a natural ablation the paper does not run.
  • Inference: the same allocation rule could be recycled across active-learning rounds, recomputing task confidence after each finetune; the paper studies only the one-batch setting.
  • Inference: the over-4-point gain over full-data training on Dolly might be partly a curation effect, since the 3K subset drops many redundant or lower-quality examples; a deduplicated pool would separate allocation value from data cleaning.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper studies one-batch label-efficient supervised finetuning: given an unlabeled prompt pool with task labels, select k prompts to annotate. It proposes Task Diversity (round-robin equal allocation across tasks) and Weighted Task Diversity (allocation proportional to the inverse of the base model's average task-level confidence, after a base allocation of five examples per task). Confidence for a prompt is the product of token probabilities of the base model's generated response. Experiments on LLaMA-2 7B use a 90K FLAN V2 subset and Dolly, with budgets 20K/30K/45K/90K and 3K/6K/13.5K, evaluating MMLU, BBH, and AlpacaEval. The main claims are that Weighted Task Diversity beats full-data training by about 4% MMLU on Dolly at 3K labels and that it generally matches or exceeds more complex baselines while saving up to 80% of labels.

Significance. If the empirical results hold, the contribution is practically useful: a simple, transparent task-level allocation rule using readily available task labels and one forward pass over the pool, with no embedding optimization or per-example score tuning, can produce strong SFT models under small annotation budgets. The paper reports means and standard errors over three seeds, includes multiple diversity and uncertainty baselines, and provides length-controlled AlpacaEval win rates, which are strengths. The main caveat is that the uncertainty mechanism underlying the method is not identified: the confidence score conflates output length with uncertainty, so the headline gains may be explained by a length bias rather than by a calibrated notion of task difficulty. With an added length-normalized ablation and a qualified abstract, the practical claim would be credible.

major comments (2)
  1. [Section 4.3] The definition of conf(x) as the product of token probabilities makes the measure exponentially sensitive to output length: for a response of length m with geometric mean per-token probability p, conf(x) = p^m. Task-level averages therefore rank tasks primarily by how long the base model's generations are, rather than by calibrated uncertainty or task difficulty. This is visible in the allocations: in Figure 4, Dolly's long-generation tasks receive 2288 (open QA), 1400 (brainstorming), and 1275 (general QA) samples, while classification, closed QA, information extraction, and summarization receive 196, 75, 86, and 60; Figure 6 shows the same pattern for FLAN. Consequently, Table 2's headline result (Weighted Task Diversity at k=3K, MMLU 39.74 vs. full-pool 35.33) does not separate 'prioritize low-confidence tasks' from 'prioritize long-output tasks.' The Limitations paragraph acknowledges that confidence may reflect pretraining exposure, but it does not address the length dependence that is built into the product definition, and no length-normalized confidence ablation is reported. Please report results with a per-token average log-probability or another length-normalized uncertainty score, or otherwise explicitly control for output length.
  2. [Abstract / Section 5.2] The abstract's claim that the algorithm 'consistently performs at or above the level of the best existing methods' is contradicted by the BBH columns of Table 1. At k=20K, Weighted Task Diversity scores 39.96±0.52, below Min Margin (40.44±0.48) and DPP (40.68±0.50); at k=45K, FL(γ=0.002) reaches 41.68±0.34 and Task Diversity 41.10±0.53, both above Weighted Task Diversity's 40.86±0.15. Section 5.2 itself concedes that on BBH 'no single method consistently dominates.' The abstract should be qualified to the datasets and metrics where the claim actually holds (e.g., MMLU and AlpacaEval, or 'on most budgets'), rather than stated as a universal property.
minor comments (5)
  1. [Section 4.1 / Tables 2-3] The text describes Dolly as a 15K dataset but the experiments use a 13.5K pool; please explain how the 13.5K subset was constructed and why it differs from the full dataset.
  2. [Section 4.3] The clamping notation in the allocation formula is typeset ambiguously, with the bounds appearing as |X_t| and 5 without a clear order. Please define [·]_a^b explicitly and state which bound is lower and which is upper.
  3. [Figures 3 and 5] The bar labels 'T ask' and 'WTD T ask Div' contain spacing artifacts; please correct these to 'Task' and 'WTD Task Div.'
  4. [Section 5.2] The sentence about the 45K budget 'saving 50% annotation budget when compared to random sampling' is imprecise; the saving is relative to the 90K full-data model, not to random sampling per se.
  5. [Reproducibility] The paper does not mention a code or data-release plan; releasing the selected prompt indices and the finetuning code would substantially aid reproducibility and comparison.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the inverse-confidence task allocation is computed directly from base-model confidences and evaluated on external benchmarks; self-citations to Bhatt et al. (2024) are not load-bearing.

full rationale

The paper's central claim is empirical: selecting prompts by task-level inverse confidence yields better downstream MMLU scores than full-data training. The selection rule in Section 4.3, alpha_t = clamp(C / conft, 5, |X_t|), uses only pretrained-model confidence averaged over unlabeled prompts. No parameter is fit to MMLU, BBH, or AlpacaEval, and the base budget of five examples is hand-set rather than tuned on test performance. The confidence score is a standard product of token probabilities, and although it may confound output length with uncertainty, that is a correctness risk acknowledged in the Limitations, not a circular reduction. The only self-citations to Bhatt et al. (2024) are used to frame the one-batch active-learning problem and to name a standard confidence score; they do not supply a theorem or fitted value on which the reported result depends. The 4% MMLU gain in Table 2 is therefore an independent experimental outcome, not an artifact of the method's definition or of the cited prior work.

Assumptions & free parameters 1 free parameters · 5 assumptions · 0 invented entities

The method introduces no new entities. Its only hand-set free parameter is the base per-task budget of 5 examples. The load-bearing assumptions concern the informativeness of task-level confidence and the availability of task labels. The core heuristic, low-confidence tasks deserve more labels, is a domain assumption rather than something derived.

free parameters (1)
  • Base per-task allocation = 5 examples
    Section 4.3 states each task is initially allocated a small base budget (e.g., 5 examples) to ensure minimum coverage. The value is chosen by hand and no sensitivity analysis is provided; it affects both low-resource and high-resource tasks.
assumptions (5)
  • domain assumption Task labels are available and partition the prompt pool into meaningful groups.
    The method relies on FLAN's 1,691 tasks and Dolly's 8 categories (Section 4.1). The Limitations section concedes that such categorizations may not exist in specialized domains.
  • domain assumption Average generated-sequence confidence is a valid proxy for annotation informativeness.
    Section 4.3 defines conft as the mean over prompts of the product of token probabilities; the whole allocation rule depends on this proxy. The Limitations note this may reflect pretraining exposure rather than task difficulty.
  • domain assumption The unnormalized product of token probabilities is comparable across tasks of different output lengths.
    No length normalization is applied, so longer generations receive systematically lower confidence and therefore more allocation (e.g., open QA in Figure 4). This modeling choice biases the allocations.
  • domain assumption Annotating more examples from low-confidence tasks improves finetuning more than annotating high-confidence tasks.
    The inverse-confidence weighting in Section 4.3 assumes a monotone relationship between model uncertainty and labeling value; this is the core heuristic and is not derived.
  • domain assumption The one-batch active learning setting, where annotations are unknown at selection time and become available for the selected subset, is the correct problem framing.
    Section 2.2 distinguishes label-efficient learning from data selection; all promoted gains are conditional on this framing.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Improving Task Diversity in Label Efficient Supervised Finetuning of LLMs." pith.science (2026). https://pith.science/paper/VWCELGOT

@misc{pith2026250721482,
  author       = {Pith},
  title        = {Pith review of: Improving Task Diversity in Label Efficient Supervised Finetuning of LLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VWCELGOT}},
  note         = {Machine review of arXiv:2507.21482}
}
read the original abstract

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse domains, but developing high-performing models for specialized applications often requires substantial human annotation -- a process that is time-consuming, labor-intensive, and expensive. In this paper, we address the label-efficient learning problem for supervised finetuning (SFT) by leveraging task-diversity as a fundamental principle for effective data selection. This is markedly different from existing methods based on the prompt-diversity. Our approach is based on two key observations: 1) task labels for different prompts are often readily available; 2) pre-trained models have significantly varying levels of confidence across tasks. We combine these facts to devise a simple yet effective sampling strategy: we select examples across tasks using an inverse confidence weighting strategy. This produces models comparable to or better than those trained with more complex sampling procedures, while being significantly easier to implement and less computationally intensive. Notably, our experimental results demonstrate that this method can achieve better accuracy than training on the complete dataset (a 4\% increase in MMLU score). Across various annotation budgets and two instruction finetuning datasets, our algorithm consistently performs at or above the level of the best existing methods, while reducing annotation costs by up to 80\%.

Figures

Figures reproduced from arXiv: 2507.21482 by the authors.

Figure 1
Figure 1. Comparison of three diversity-based data selection baseline algorithms. In contrast to task diversity in [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. (a) Task diversity: Fixed number of examples are selected in a round-robin manner across tasks, without [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Evaluation of 30K prompt selection strategies on the FLAN [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Number of allocated samples per task in the Dolly dataset under the Weighted Task Diversity strategy. “BRNSTRM" indicates Brainstorming Tasks, “CLS" indicates Classification Tasks, “Cld qa" indicates Closed QA and “Summ" indicates Summarization. In [PITH_FULL_IMAGE:fi…
Figure 5
Figure 5. Figure 5: Evaluation of 6K prompt selection strategies on the Dolly dataset using AlpacaEval. The win rate [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: Task-wise sample allocation in the FLAN V2 dataset using the Weighted Task Diversity strategy. Tasks where the model exhibits greater uncertainty receive a larger portion of the annotation budget. Nevertheless, such tasks still receive limited annotations, helping to p…

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

42 extracted references · 13 canonical work pages

  1. [1]

    Jordan Ash, Surbhi Goel, Akshay Krishnamurthy, and Sham Kakade. 2021. Gone fishing: Neural active learning with fisher embeddings. Advances in Neural Information Processing Systems, 34:8927--8939

  2. [2]

    Jordan T Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford, and Alekh Agarwal. 2019. Deep batch active learning by diverse, uncertain gradient lower bounds. arXiv preprint arXiv:1906.03671

  3. [3]

    Gantavya Bhatt, Yifang Chen, Arnav M Das, Jifan Zhang, Sang T Truong, Stephen Mussmann, Yinglun Zhu, Jeffrey Bilmes, Simon S Du, Kevin Jamieson, Jordan T Ash, and Robert D Nowak. 2024. An experimental design framework for label-efficient supervised finetuning of large language models. arXiv preprint arXiv:2401.06692

  4. [4]

    Jeff Bilmes. 2022. Submodularity in machine learning and artificial intelligence. arXiv preprint arXiv:2202.00132

  5. [5]

    Alexander Bukharin and Tuo Zhao. 2023. https://arxiv.org/abs/2311.14736 Data diversity matters for robust instruction tuning . Preprint, arXiv:2311.14736

  6. [6]

    Hardy Chen, Haoqin Tu, Fali Wang, Hui Liu, Xianfeng Tang, Xinya Du, Yuyin Zhou, and Cihang Xie. 2025. Sft or rl? an early investigation into training r1-like reasoning large vision-language models. arXiv preprint arXiv:2504.11468

  7. [7]

    Lichang Chen, Shiyang Li, Jun Yan, Hai Wang, Kalpa Gunaratna, Vikas Yadav, Zheng Tang, Vijay Srinivasan, Tianyi Zhou, Heng Huang, and Hongxia Jin. 2023. https://arxiv.org/abs/2307.08701 Alpagasus: Training a better alpaca with fewer data . Preprint, arXiv:2307.08701

  8. [8]

    Gui Citovsky, Giulia DeSalvo, Claudio Gentile, Lazaros Karydas, Anand Rajagopalan, Afshin Rostamizadeh, and Sanjiv Kumar. 2021. Batch active learning at scale. Advances in Neural Information Processing Systems, 34:11933--11944

Show all 42 references
  1. [9]

    Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023. Free dolly: Introducing the world’s first truly open instruction-tuned llm

  2. [10]

    Qianlong Du, Chengqing Zong, and Jiajun Zhang. 2023. https://arxiv.org/abs/2311.15653 Mods: Model-oriented data selection for instruction tuning . Preprint, arXiv:2311.15653

  3. [11]

    Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2020. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300

  4. [12]

    Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2021. Lora: Low-rank adaptation of large language models. arXiv preprint arXiv:2106.09685

  5. [13]

    Smith, Hannaneh Hajishirzi, and Pradeep Dasigi

    Hamish Ivison, Noah A. Smith, Hannaneh Hajishirzi, and Pradeep Dasigi. 2022. Data-efficient fine-tuning using cross-task nearest neighbors. arXiv preprint arXiv:2212.00196

  6. [14]

    Jan Kremer, Kim Steenstrup Pedersen, and Christian Igel. 2014. Active learning with support vector machines. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 4(4):313--326

  7. [15]

    Alex Kulesza, Ben Taskar, and 1 others. 2012. Determinantal point processes for machine learning. Foundations and Trends in Machine Learning , 5(2--3):123--286

  8. [16]

    Po-Nien Kung, Fan Yin, Di Wu, Kai-Wei Chang, and Nanyun Peng. 2023. Active instruction tuning: Improving cross-task generalization by training on prompt sensitive tasks. arXiv preprint arXiv:2311.00288

  9. [17]

    David D Lewis. 1995. A sequential algorithm for training text classifiers: Corrigendum and additional data. In Acm Sigir Forum, volume 29, pages 13--19. ACM New York, NY, USA

  10. [18]

    Hashimoto

    Xuechen Li, Tianyi Zhang, Yann Dubois, Rohan Taori, Ishaan Gulrajani, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto. 2023. Alpacaeval: An automatic evaluator of instruction-following models. https://github.com/tatsu-lab/alpaca_eval

  11. [19]

    Shayne Longpre, Le Hou, Tu Vu, Albert Webson, Hyung Won Chung, Yi Tay, Denny Zhou, Quoc V Le, Barret Zoph, Jason Wei, and 1 others. 2023. The flan collection: Designing data and methods for effective instruction tuning. arXiv preprint arXiv:2301.13688

  12. [20]

    Pitu B Mirchandani and Richard L Francis. 1990. Discrete location theory

  13. [21]

    Baharan Mirzasoleiman, Jeff Bilmes, and Jure Leskovec. 2020. Coresets for data-efficient training of machine learning models. In International Conference on Machine Learning, pages 6950--6960. PMLR

  14. [22]

    Mohamad Amin Mohamadi, Wonho Bae, and Danica J Sutherland. 2022. Making look-ahead active learning strategies feasible with neural tangent kernels. Advances in Neural Information Processing Systems, 35:12542--12553

  15. [23]

    Shyam Nuggehalli, Jifan Zhang, Lalit Jain, and Robert Nowak. 2023. Direct: Deep active learning under imbalance and label noise. arXiv preprint arXiv:2312.09196

  16. [24]

    Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and...

  17. [25]

    Yulei Qin, Yuncheng Yang, Pengcheng Guo, Gang Li, Hang Shao, Yuchen Shi, Zihan Xu, Yun Gu, Ke Li, and Xing Sun. 2024. Unleashing the power of data tsunami: A comprehensive survey on data assessment and selection for instruction tuning of language models. arXiv preprint arXiv:2...

  18. [26]

    Ozan Sener and Silvio Savarese. 2018. https://openreview.net/forum?id=H1aIuk-RW Active learning for convolutional neural networks: A core-set approach . In International Conference on Learning Representations (ICLR)

  19. [27]

    Burr Settles. 2009. Active learning literature survey

  20. [28]

    Mirac Suzgun, Nathan Scales, Nathanael Sch \"a rli, Sebastian Gehrmann, Yi Tay, Hyung Won Chung, Aakanksha Chowdhery, Quoc V Le, Ed H Chi, Denny Zhou, and 1 others. 2022. Challenging big-bench tasks and whether chain-of-thought can solve them. arXiv preprint arXiv:2210.09261

  21. [29]

    Simon Tong and Daphne Koller. 2001. Support vector machine active learning with applications to text classification. Journal of machine learning research, 2(Nov):45--66

  22. [30]

    Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, and 1 others. 2023. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288

  23. [31]

    Peiqi Wang, Yikang Shen, Zhen Guo, Matthew Stallone, Yoon Kim, Polina Golland, and Rameswar Panda. 2024. Diversity measurement and subset selection for instruction tuning datasets. arXiv preprint arXiv:2402.02318

  24. [32]

    Yizhong Wang, Hamish Ivison, Pradeep Dasigi, Jack Hessel, Tushar Khot, Khyathi Raghavi Chandu, David Wadden, Kelsey MacMillan, Noah A Smith, Iz Beltagy, and 1 others. 2023. How far can camels go? exploring the state of instruction tuning on open resources. arXiv preprint arXiv...

  25. [33]

    Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Anjana Arunkumar, Arjun Ashok, Arut Selvan Dhanasekaran, Atharva Naik, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Gary Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby ...

  26. [34]

    Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M

    Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022. https://arxiv.org/abs/2109.01652 Finetuned language models are zero-shot learners . Preprint, arXiv:2109.01652

  27. [35]

    Kai Wei, Rishabh Iyer, and Jeff Bilmes. 2015. Submodularity in data subset selection and active learning. In International conference on machine learning, pages 1954--1963. PMLR

  28. [36]

    Tian Xie, Jifan Zhang, Haoyue Bai, and Robert Nowak. 2024. Deep active learning in the open world. arXiv preprint arXiv:2411.06353

  29. [37]

    Yu Yang, Siddhartha Mishra, Jeffrey Chiang, and Baharan Mirzasoleiman. 2024. Smalltolarge (s2l): Scalable data selection for fine-tuning large language models by summarizing training trajectories of small models. Advances in Neural Information Processing Systems, 37:83465--83496

  30. [38]

    Jifan Zhang, Gregory Canal, Yinglun Zhu, Robert D Nowak, Yifang Chen, Arnav M Das, Gantavya Bhatt, Stephen Mussmann, Jeffrey Bilmes, Simon S Du, and 1 others. 2023. Labelbench: A comprehensive framework for benchmarking adaptive label-efficient learning. arXiv preprint arXiv:2...

  31. [39]

    Jifan Zhang, Julian Katz-Samuels, and Robert Nowak. 2022. Galaxy: Graph-based active learning at the extreme. In International Conference on Machine Learning, pages 26223--26238. PMLR

  32. [40]

    Kuan Lok Zhou, Jiayi Chen, Siddharth Suresh, Reuben Narad, Timothy T Rogers, Lalit K Jain, Robert D Nowak, Bob Mankoff, and Jifan Zhang. 2025. Bridging the creativity understanding gap: Small-scale human alignment enables expert-level humor ranking in llms. arXiv preprint arXi...

  33. [41]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  34. [42]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.