REVIEW 3 major objections 5 minor 60 references
CorrSynth -- A Correlated Sampling Method for Diverse Dataset Generation from LLMs
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read CorrSynth claims that generating synthetic training examples in lockstep pairs that contrast their logits at every token produces datasets with both higher diversity and better downstream student accuracy than independent few-shot…
desk verdict The parallel-contrast decoding idea is genuinely new and the student-accuracy gains look real, but the paper overclaims its intrinsic metrics: Table 2 shows MAUVE gets worse with the Phi-3-mini teacher. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing identity is the lockstep contrastive sampling rule, equations (1)–(2) for two labels and (3) for $M$ parallel sequences: at token $i$, each sequence samples from its own label-conditioned distribution raised to power $\gamma$, divided by the geometric mean of the other active sequences' distributions raised to $\gamma-\delta$. The paper implements this in logits space by subtracting weighted contrast logits from the numerator logits (equation (13)) and adds a plausibility constraint (equations (14)–(15)) that zeroes out any token below $\alpha$ fraction of the numerator mode. This mechanism does the work because it makes the contrast depend on the partner's evolving text, not on a static prompt, so the repulsion does not decay with generation length.
What would settle it
A direct test is to run the method on a held-out domain not in the paper and compare the average pairwise embedding distance between parallel-generated pairs with pairs of independent few-shot generations: if the parallel pairs are not consistently farther apart, the claimed anti-correlation has not generalized. A second check is to track the $\infty$-norm of the logit difference between the numerator and denominator across token positions on long generations; if it decays to the level seen with classifier-free guidance, the persistence advantage that the method claims disappears.
Extended reading notes
Core claim
For two classes with verbalized labels $y$ and $\bar y$, the $i$-th token of $x$ is drawn from $\tilde P_i(\cdot)\propto P(\cdot\mid \mathrm{prompt}(y),x_{<i})^\gamma\,P(\cdot\mid \mathrm{prompt}(\bar y),\bar x_{<i})^{\gamma-\delta}$, while $\bar x_i$ is drawn from the symmetric distribution with the two labels swapped. The paper's central claim is that the two sequences generated this way are naturally anti-correlated—they move apart in the LLM's embedding space—because each sequence's own partial text is fed into the other sequence's contrast term, keeping the guidance signal alive where classifier-free guidance would let it fade. The paper supports the claim with intrinsic metrics (Self-BLEU, entity entropy, MAUVE) and student accuracy on AG News, TOI Headlines, Humor, and IMDb, reporting that CorrSynth variants outperform few-shot generation on all four and outperform several published baselines where those baselines released data.
Load-bearing premise
The whole method rests on the assumption that feeding each sequence's own partial text into the other sequence's contrast term keeps the two generations actively repelling each other all the way to the end; if that anti-correlation fades or fails to transfer to a new task or domain, the accuracy and diversity gains reported here would not generalize.
Editorial extensions
If this is right
- Synthetic classification datasets can be made more diverse without a larger teacher or an external retrieval corpus; the gains come from the sampling procedure itself.
- Because the $K$ parallel sequences are themselves the outputs, the method costs $N \times L$ forward passes for $N$ generations, whereas an equivalent $K$-way classifier-free-guidance formulation costs a factor of $K$ or $R$ more, so the saved compute can be spent on more examples.
- The parameter $\delta$ provides a practical dial between label separation and hard negatives: high $\delta$ gives well-separated clusters, low $\delta$ generates overlapping examples that look like mislabels but can still help the student.
- The same contrast mechanism is portable to other decoding-time synthesis pipelines, such as retrieval-augmented generation, since the correlated sampling operates at the token-sampling stage rather than at the prompt level.
Reading between the lines
- Beyond the paper's four English classification tasks, the most direct untested promise is whether lockstep contrast survives in long-form or open-ended generation; the paper's own argument that guidance persistence matters more for longer outputs makes long-form the natural next experiment.
- Because the method needs white-box logits, it cannot be used with API-only teacher models; an approximate version that estimates contrast from sampled text rather than logits would test whether the mechanism survives closed models, but the paper does not explore this.
- The paper reports that higher entity entropy and lower self-BLEU accompany higher student accuracy, but it does not establish that the diversity gains cause the accuracy gains; an experiment that fixes the number of training examples and varies only entity overlap could separate cause from correlation.
- Contrasting against arbitrary unrelated prompts rather than opposite labels would test whether the benefit comes from semantic label opposition or from any parallel repulsion between sequences; such an ablation could simplify or broaden the method.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CorrSynth, a decoding-time sampling method for LLM-based synthetic classification dataset generation. Instead of sampling each class-conditioned sequence independently, CorrSynth generates several sequences in parallel and uses the logits of the other parallel sequences as contrast terms, generalizing classifier-free guidance. The authors define cross-label, intra-label, and hybrid variants, and evaluate them on AG News, TOI Headlines, Humor, and IMDB using Mixtral and Phi-3-mini teachers, training DistilBERT students. They report improvements over FewGen and selected prior works in student accuracy, Self-BLEU, and entity entropy, and additionally claim better MAUVE.
Significance. If the empirical findings hold, this is a useful and simple contribution to synthetic data generation: it extends classifier-free guidance with a parallel contrast formulation, reduces the number of forward passes compared to CFG by a factor related to the number of classes, and works with open-weight local models. The paper's strengths are its breadth of experiments (four datasets, two teachers, zero-shot and few-shot settings), the computational-complexity analysis of CFG, the visual analyses with UMAP and dataset cartography, and an unusually honest limitations section. The main weakness is that the abstract and Section 5.2 overstate the intrinsic-metric results: the MAUVE claim is not robust for the Phi-3-mini teacher, and the absence of error bars or significance tests makes it difficult to assess how consistent the smaller accuracy gains are.
major comments (3)
- [Abstract; §5.2; Table 2] The abstract's central claim that CorrSynth 'improves both student metrics and intrinsic metrics' is contradicted by Table 2 for MAUVE with the Phi-3-mini teacher. FEWGEN has average MAUVE 82.2, Corr-Intra 81.3, and Corr-Hybrid 78.1; per dataset, Corr-Intra is worse than FEWGEN on AG News (82.3 vs 91.0), TOI Headlines (83.2 vs 86.3), and Humor (82.3 vs 83.7), improving only on IMDB. The sentence in Section 5.2 claiming a 'better match with human-written text (better MAUVE)' is therefore false for this teacher. The robust gains are in diversity (Self-BLEU, entity entropy) and student accuracy; the paper should qualify the intrinsic-metric claim, report variance across seeds and teachers, or remove MAUVE from the blanket statement.
- [Table 2; §4 Evaluation criteria] No error bars, confidence intervals, or significance tests are reported for the student accuracy means (described as averaged over 5 runs) or for MAUVE. Several of the reported accuracy improvements are small, for example 83.8 to 84.8 on AG News for Corr-Intra with Phi-3-mini, so without variance estimates the claim that improvements are 'consistently reported' across all four datasets is not fully supported. Please add standard deviations or confidence intervals, and ideally a paired significance test for the main FEWGEN vs CorrSynth comparisons.
- [§3.1; Figure 2; Figure 3] The method is motivated by the hypothesis that sequences generated by equations (1) and (2) are 'naturally anti-correlated' and that the contrast signal remains alive throughout generation. This hypothesis is demonstrated only with five prompts on IMDB in Figure 2 and with heatmaps on TOI Headlines in Figure 3. There is no formal argument and no evidence that the anti-correlation generalizes across the two teachers and the remaining datasets. Because this mechanism is the paper's stated foundation, the authors should either provide additional evidence (for example, the same contrast-persistence plot for another dataset and for Mixtral) or explicitly state that the method is empirical and that the mechanism is only qualitatively supported.
minor comments (5)
- [§3.1, equations (1)-(2)] The exact sampling order is not fully specified: it should be stated explicitly whether x_i and bar-x_i are both sampled from distributions conditioned on the other sequence's prefix at position i-1, or whether one current token is sampled first and then used in the other distribution. The equations suggest simultaneous sampling, but the text should say so.
- [§5.3, Table 3] The comparison with REGEN, SynthesizRR, SunGen, S3, and AttrPrompt uses numbers quoted from Divekar and Durrett (2024), where the teacher models differ (BERT, Llama2, GPT2-XL, GPT3.5-T, vs Phi-3-mini here) and no error bars are available for the quoted cells. This is not a controlled head-to-head, so the text should avoid suggesting a direct win over these prior methods and instead frame the table as an indicative comparison with published numbers.
- [Appendix D] There is a typo: 'a titled distribution' should be 'a tilted distribution' in the sentence introducing equation (5).
- [§4, Entity entropy] The entity entropy metric is described only vaguely as using the distribution of 16 entity types from a pre-trained NER model. Please provide the exact formula and the NER model used, since this metric is one of the two 'diversity' metrics that carry the paper's central claim.
- [Appendix F.2; Table 5] The CFG-vs-CorrSynth comparison is performed only for the intra-label variant on TOI Headlines, and the paper correctly notes the comparison is not fair because CFG receives twice the compute budget. This limitation is acknowledged; still, the main-text sentence in Section 5.1 that CorrSynth is 'better suited for longer generations' is based on only one dataset and should be softened accordingly.
Circularity Check
No significant circularity: CorrSynth's gains are evaluated on held-out test sets and its sampling equations are not derived from the metrics they predict.
full rationale
The paper is an empirical method paper; it does not claim to derive a quantitative law from first principles. The central results (student accuracy on held-out test sets, Self-BLEU, entity entropy, MAUVE) are measured on outputs of a sampling procedure whose parameters (γ, δ, α, R) are explicitly tuned. This is standard practice, not circular. The anti-correlation property is introduced as a hypothesis: 'We hypothesize that the sequences x, x̄ generated auto-regressively using equations (1) and (2) are naturally anti-correlated,' with the paper openly stating that the reasoning 'is based on intuition that requires careful experiments to prove.' The subsequent heatmaps and UMAP plots are sanity checks of the mechanism, not predictions derived from the method. The only self-citation of note is Divekar & Durrett (2024), used to justify ICL example counts, evaluation metrics, and to quote baseline numbers in Table 3; one of the present authors is a co-author of that work. However, these citations are not load-bearing for the central claim: the primary FEWGEN comparison is run in-house in Table 2, and all student accuracies are evaluated on held-out test sets. The paper therefore does not reduce to its own inputs by construction. A separate correctness concern is that the abstract's claim that CorrSynth improves intrinsic metrics 'across four datasets' is contradicted by Table 2 for MAUVE with Phi-3-mini (e.g., AG News 82.3 vs FEWGEN 91.0), but this is an overstatement, not circularity.
Assumptions & free parameters
free parameters (4)
- gamma =
1.0
- delta =
0.9*gamma (cross-label), 0.5*gamma (intra-label); gamma_intra=gamma/2, gamma_cross=gamma/10 (hybrid)
- plausibility threshold alpha =
1e-3
- repeat factor R =
1 (cross-label), 2 (intra/hybrid)
assumptions (3)
- domain assumption Next-token distributions from the LLM under different prompts have comparable logits so they can be combined by weighted product.
- domain assumption The plausibility set with threshold alpha retains enough of the true next-token support to keep generations coherent.
- domain assumption The seed set of 50-100 examples per class is sufficient to provide in-context demonstrations without dominating the generated distribution.
Cite this review
Pith. "Pith review of CorrSynth -- A Correlated Sampling Method for Diverse Dataset Generation from LLMs." pith.science (2026). https://pith.science/paper/D6LBGRRG
@misc{pith2026241108553,
author = {Pith},
title = {Pith review of: CorrSynth -- A Correlated Sampling Method for Diverse Dataset Generation from LLMs},
year = {2026},
howpublished = {\url{https://pith.science/paper/D6LBGRRG}},
note = {Machine review of arXiv:2411.08553}
}
read the original abstract
Large language models (LLMs) have demonstrated remarkable performance in diverse tasks using zero-shot and few-shot prompting. Even though their capabilities of data synthesis have been studied well in recent years, the generated data suffers from a lack of diversity, less adherence to the prompt, and potential biases that creep into the data from the generator model. In this work, we tackle the challenge of generating datasets with high diversity, upon which a student model is trained for downstream tasks. Taking the route of decoding-time guidance-based approaches, we propose CorrSynth, which generates data that is more diverse and faithful to the input prompt using a correlated sampling strategy. Further, our method overcomes the complexity drawbacks of some other guidance-based techniques like classifier-based guidance. With extensive experiments, we show the effectiveness of our approach and substantiate our claims. In particular, we perform intrinsic evaluation to show the improvements in diversity. Our experiments show that CorrSynth improves both student metrics and intrinsic metrics upon competitive baselines across four datasets, showing the innate advantage of our method.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
online" 'onlinestring :=
ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence after.sentence after.block STRING...
-
[2]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 global.max substring 't := if while FUNCTION word.in bbl.in capitalize " " * FUNCT...
-
[3]
Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan, Jyoti Aneja, Ahmed Awadallah, Hany Awadalla, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Harkirat Behl, et al. 2024. Phi-3 technical report: A highly capable language model locally on your phone. arXiv preprint arXiv:2404.14219
arXiv 2024
-
[4]
OpenAI Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mo Bavarian, Jeff Belgum, Irwan Bello, Jake Berdine, Gabriel Bernadett-Shapiro, Christopher Berner, Lenny...
2023
-
[5]
Ateret Anaby-Tavor, Boaz Carmeli, Esther Goldbraich, Amir Kantor, George Kour, Segev Shlomov, N. Tepper, and Naama Zwerdling. 2019. https://api.semanticscholar.org/CorpusID:212821571 Do not have enough data? deep learning to the rescue! In AAAI Conference on Artificial Intelligence
work page 2019
-
[6]
Brown, Jack Clark, Sam McCandlish, Christopher Olah, Benjamin Mann, and Jared Kaplan
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, John Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dar...
arXiv 2022
-
[7]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...
2020
-
[8]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...
2020
Show all 60 references
-
[9]
Yu, Qiang Yang, and Xing Xie
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Linyi Yang, Kaijie Zhu, Hao Chen, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, Wei Ye, Yue Zhang, Yi Chang, Philip S. Yu, Qiang Yang, and Xing Xie. 2024. https://doi.org/10.1145/3641289 A survey on evaluation of large language models ....
2024 doi
-
[10]
Yung-Sung Chuang, Yujia Xie, Hongyin Luo, Yoon Kim, James Glass, and Pengcheng He. 2023. Dola: Decoding by contrasting layers improves factuality in large language models. arXiv preprint arXiv:2309.03883
2023 arXiv
-
[11]
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021. https://api.semanticscholar.org/CorpusID:239998651 Training verifiers to solve mat...
2021 arXiv
-
[12]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...
2019 doi
-
[13]
Abhishek Divekar and Greg Durrett. 2024. Synthesizrr: Generating diverse datasets with retrieval augmentation. arXiv preprint arXiv:2405.10040
2024 arXiv
-
[14]
Jiahui Gao, Renjie Pi, LIN Yong, Hang Xu, Jiacheng Ye, Zhiyong Wu, WEIZHONG ZHANG, Xiaodan Liang, Zhenguo Li, and Lingpeng Kong. 2022. Self-Guided Noise-Free Data Generation for Efficient Zero-Shot Learning . In The Eleventh International Conference on Learning Representations
2022
-
[15]
Jiahui Gao, Renjie Pi, LIN Yong, Hang Xu, Jiacheng Ye, Zhiyong Wu, Weizhong Zhang, Xiaodan Liang, Zhenguo Li, and Lingpeng Kong. 2023. https://openreview.net/forum?id=h5OpjGd_lo6 Self-guided noise-free data generation for efficient zero-shot learning . In The Eleventh Internat...
2023
-
[16]
Ariel Gera, Roni Friedman, Ofir Arviv, Chulaka Gunasekara, Benjamin Sznajder, Noam Slonim, and Eyal Shnarch. 2023. https://doi.org/10.18653/v1/2023.acl-long.580 The benefits of bad advice: Autocontrastive decoding across model layers . In Proceedings of the 61st Annual Meeting...
2023 doi
-
[17]
Xu Guo and Yiqiang Chen. 2024. Generative AI for Synthetic Data Generation: Methods, Challenges and the Future . arXiv preprint arXiv:2403.04190
2024 arXiv
-
[18]
Jonathan Ho and Tim Salimans. 2021. https://openreview.net/forum?id=qw8AKxfYbI Classifier-Free Diffusion Guidance . In NeurIPS 2021 Workshop on Deep Generative Models and Downstream Applications
2021
-
[19]
Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2023. https://doi.org/10.18653/v1/2023.acl-long.806 Unnatural instructions: Tuning language models with (almost) no human labor . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics...
2023 doi
-
[20]
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al. 2023. Mistral 7b. arXiv preprint arXiv:2310.06825
2023 arXiv
-
[21]
Albert Q Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al. 2024. Mixtral of experts. arXiv preprint arXiv:2401.04088
2024 arXiv
-
[22]
Xu, Jun Araki, and Graham Neubig
Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020. https://doi.org/10.1162/tacl_a_00324 How can we know what language models know? Transactions of the Association for Computational Linguistics, 8:423--438
2020 doi
-
[23]
Rohit Kulkarni. 2020. https://doi.org/10.7910/DVN/DPQMQH Times of India News Headlines
2020 doi
-
[24]
Varun Kumar, Ashutosh Choudhary, and Eunah Cho. 2020. https://aclanthology.org/2020.lifelongnlp-1.3 Data augmentation using pre-trained transformer models . In Proceedings of the 2nd Workshop on Life-long Learning for Spoken Language Systems, pages 18--26, Suzhou, China. Assoc...
2020
-
[25]
Kenton Lee, Kelvin Guu, Luheng He, Timothy Dozat, and Hyung Won Chung. 2021. https://api.semanticscholar.org/CorpusID:231749880 Neural Data Augmentation via Example Extrapolation . ArXiv, abs/2102.01335
2021 arXiv
-
[26]
u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K \"u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \"a schel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Proc...
2020
-
[27]
Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori B Hashimoto, Luke Zettlemoyer, and Mike Lewis. 2023. Contrastive Decoding: Open-ended Text Generation as Optimization . In Proceedings of the 61st Annual Meeting of the Association for Computationa...
2023
-
[28]
Lang Liu, Krishna Pillutla, Sean Welleck, Sewoong Oh, Yejin Choi, and Zaid Harchaoui. 2021. Divergence Frontiers for Generative Models: Sample Complexity, Quantization Effects, and Frontier Integrals . In Advances in Neural Information Processing Systems
2021
-
[29]
Maas, Raymond E
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. P11-1015 Learning word vectors for sentiment analysis . pages 142--150, Portland, Oregon, USA
2011
-
[30]
Leland McInnes, John Healy, and James Melville. 2020. https://arxiv.org/abs/1802.03426 Umap: Uniform manifold approximation and projection for dimension reduction . Preprint, arXiv:1802.03426
2020 arXiv
-
[31]
Yu Meng, Jiaxin Huang, Yu Zhang, and Jiawei Han. 2022 a . https://api.semanticscholar.org/CorpusID:246680398 Generating training data with language models: Towards zero-shot language understanding . ArXiv, abs/2202.04538
2022 arXiv
-
[32]
Yu Meng, Jiaxin Huang, Yu Zhang, and Jiawei Han. 2022 b . https://proceedings.neurips.cc/paper_files/paper/2022/file/0346c148ba1c21c6b4780a961ea141dc-Paper-Conference.pdf Generating training data with language models: Towards zero-shot language understanding . In Advances in N...
2022
-
[33]
Yu Meng, Martin Michalski, Jiaxin Huang, Yu Zhang, Tarek Abdelzaher, and Jiawei Han. 2023 a . Tuning language models as training data generators for augmentation-enhanced few-shot learning. In International Conference on Machine Learning, pages 24457--24477. PMLR
2023
-
[34]
Yu Meng, Martin Michalski, Jiaxin Huang, Yu Zhang, Tarek Abdelzaher, and Jiawei Han. 2023 b . Tuning language models as training data generators for augmentation-enhanced few-shot learning. In International Conference on Machine Learning
2023
-
[35]
Sean O'Brien and Mike Lewis. 2023. Contrastive decoding improves reasoning in large language models. arXiv preprint arXiv:2309.09117
2023 arXiv
-
[36]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. https://doi.org/10.3115/1073083.1073135 BLEU: A method for automatic evaluation of machine translation . In Proceedings of the 40th Annual Meeting on Association for Computational Linguistics, ACL '02, page 311...
2002
-
[37]
Raul Puri, Ryan Spring, Mohammad Shoeybi, Mostofa Patwary, and Bryan Catanzaro. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.468 Training question answering models from synthetic data . In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Process...
2020 doi
-
[38]
Laria Reynolds and Kyle McDonell. 2021. https://doi.org/10.1145/3411763.3451760 Prompt programming for large language models: Beyond the few-shot paradigm . In Extended Abstracts of the 2021 CHI Conference on Human Factors in Computing Systems, CHI EA '21, New York, NY, USA. A...
2021
-
[39]
Guillaume Sanchez, Honglu Fan, Alexander Spangher, Elad Levi, Pawan Sasanka Ammanamanchi, and Stella Biderman. 2023. Stay on topic with classifier-free guidance . arXiv preprint arXiv:2306.17806
2023 arXiv
-
[40]
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019. https://arxiv.org/abs/1910.01108 DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter . In 5th Workshop on Energy Efficient Machine Learning and Cognitive Computing @ NeurIPS 2019
2019 arXiv
-
[41]
Timo Schick and Hinrich Sch \"u tze. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.555 Generating datasets with pretrained language models . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6943--6951, Online and Punta Cana, ...
2021 doi
-
[42]
Weijia Shi, Xiaochuang Han, Mike Lewis, Yulia Tsvetkov, Luke Zettlemoyer, and Scott Yih. 2023. https://api.semanticscholar.org/CorpusID:258866080 Trusting your evidence: Hallucinate less with context-aware decoding . ArXiv, abs/2305.14739
2023 arXiv
-
[43]
Logan IV, Eric Wallace, and Sameer Singh
Taylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace, and Sameer Singh. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.346 A uto P rompt: E liciting K nowledge from L anguage M odels with A utomatically G enerated P rompts . In Proceedings of the 2020 Conference o...
2020 doi
-
[44]
Smith, and Yejin Choi
Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A. Smith, and Yejin Choi. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.746 Dataset cartography: Mapping and diagnosing datasets with training dynamics . In Proceedings of the 2020 C...
2020 doi
-
[45]
Ruida Wang, Wangchunshu Zhou, and Mrinmaya Sachan. 2023 a . https://doi.org/10.18653/v1/2023.findings-emnlp.791 Let ' s synthesize step by step: Iterative dataset synthesis with large language models by extrapolating errors from small models . In Findings of the Association fo...
2023 doi
-
[46]
Smith, Daniel Khashabi, and Hannaneh Hajishirzi
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023 b . https://doi.org/10.18653/v1/2023.acl-long.754 Self-instruct: Aligning language models with self-generated instructions . In Proceedings of the 61st Annual ...
2023 doi
-
[47]
Zirui Wang, Adams Wei Yu, Orhan Firat, and Yuan Cao. 2021. https://api.semanticscholar.org/CorpusID:237572306 Towards zero-label language learning . ArXiv, abs/2109.09193
2021 arXiv
-
[48]
Peter West, Chandra Bhagavatula, Jack Hessel, Jena Hwang, Liwei Jiang, Ronan Le Bras, Ximing Lu, Sean Welleck, and Yejin Choi. 2022. https://doi.org/10.18653/v1/2022.naacl-main.341 Symbolic knowledge distillation: from general language models to commonsense models . In Proceed...
2022 doi
-
[49]
Jiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu, Jiangtao Feng, Zhiyong Wu, Tao Yu, and Lingpeng Kong. 2022 a . ZeroGen: Efficient Zero-shot Learning via Dataset Generation . In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11653--11669
2022
-
[50]
Jiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu, Jiangtao Feng, Zhiyong Wu, Tao Yu, and Lingpeng Kong. 2022 b . https://api.semanticscholar.org/CorpusID:246867045 Zerogen: Efficient zero-shot learning via dataset generation . ArXiv, abs/2202.07922
2022 arXiv
-
[51]
Jiacheng Ye, Jiahui Gao, Zhiyong Wu, Jiangtao Feng, Tao Yu, and Lingpeng Kong. 2022 c . ProGen: Progressive Zero-shot Dataset Generation via In-context Feedback . In Findings of the Association for Computational Linguistics: EMNLP 2022, pages 3671--3683
2022
-
[52]
Jiacheng Ye, Jiahui Gao, Zhiyong Wu, Jiangtao Feng, Tao Yu, and Lingpeng Kong. 2022 d . https://doi.org/10.18653/v1/2022.findings-emnlp.269 P ro G en: Progressive zero-shot dataset generation via in-context feedback . In Findings of the Association for Computational Linguistic...
2022 doi
-
[53]
Asaf Yehudai, Boaz Carmeli, Yosi Mass, Ofir Arviv, Nathaniel Mills, Assaf Toledo, Eyal Shnarch, and Leshem Choshen. 2024. https://api.semanticscholar.org/CorpusID:267211959 Genie: Achieving human parity in content-grounded datasets generation . ArXiv, abs/2401.14367
2024 arXiv
-
[54]
Yue Yu, Yuchen Zhuang, Jieyu Zhang, Yu Meng, Alexander Ratner, Ranjay Krishna, Jiaming Shen, and Chao Zhang. 2023 a . https://openreview.net/forum?id=6hZIfAY9GD Large language model as attributed training data generator: A tale of diversity and bias . In Thirty-seventh Confere...
2023
-
[55]
Yue Yu, Yuchen Zhuang, Jieyu Zhang, Yu Meng, Alexander J Ratner, Ranjay Krishna, Jiaming Shen, and Chao Zhang. 2024. Large language model as attributed training data generator: A tale of diversity and bias. Advances in Neural Information Processing Systems, 36
2024
-
[56]
Yue Yu, Yuchen Zhuang, Rongzhi Zhang, Yu Meng, Jiaming Shen, and Chao Zhang. 2023 b . Regen: Zero-shot text classification via training data generation with progressive dense retrieval. arXiv preprint arXiv:2305.10703
2023 arXiv
-
[57]
Yue Yu, Yuchen Zhuang, Rongzhi Zhang, Yu Meng, Jiaming Shen, and Chao Zhang. 2023 c . https://doi.org/10.18653/v1/2023.findings-acl.748 R e G en: Zero-shot text classification via training data generation with progressive dense retrieval . In Findings of the Association for Co...
2023 doi
-
[58]
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015. Character-level convolutional networks for text classification. In Proceedings of the 28th International Conference on Neural Information Processing Systems - Volume 1, NIPS'15, page 649–657, Cambridge, MA, USA. MIT Press
2015
-
[59]
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018. Texygen: A benchmarking platform for text generation models. SIGIR
2018
-
[60]
Yftah Ziser, Elad Kravi, and David Carmel. 2020. https://doi.org/10.1145/3397271.3401077 Humor detection in product question answering systems . In Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, SIGIR '20, page ...
2020
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.