REVIEW 1 major objections 1 minor 1 cited by
Evolving small prompt embeddings steers frozen LLMs to produce more diverse high-quality outputs across benchmarks.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-30 22:27 UTC pith:UHXFDEJW
load-bearing objection QD-LLM evolves small prompt embeddings inside a QD archive to increase output diversity from frozen LLMs, with reported gains over QDAIF but thin details on the coverage theorem. the 1 major comments →
Parameter-Efficient Neuroevolution for Diverse LLM Generation: Quality-Diversity Optimization via Prompt Embedding Evolution
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
QD-LLM evolves prompt embeddings of roughly 32K parameters inside a quality-diversity loop that uses gradient-free variation operators and a hybrid behavior map; the map supplies coverage guarantees under measured near-independence of its semantic and explicit components. On HumanEval, MBPP, and creative writing problems the resulting archives reach 46.4 percent higher coverage and 41.4 percent higher QD-Score than QDAIF while using the same frozen 70B-scale models. The diverse archives improve downstream test generation by surfacing 34 percent more edge cases and raise fine-tuning accuracy by 8.3 percent.
What carries the argument
Evolved prompt embeddings that act as compact neural interfaces for behavioral steering of frozen LLMs inside a quality-diversity archive.
Load-bearing premise
The hybrid behavior map supplies valid formal coverage bounds when its semantic and explicit features remain nearly independent.
What would settle it
An experiment in which the normalized mutual information between semantic and explicit features exceeds 0.15 while the observed coverage still matches the bound stated in Theorem 1.
If this is right
- Archives produced by the method cover 46.4 percent more distinct behaviors on HumanEval and MBPP than the baseline.
- Test suites built from the archives detect 34 percent more edge cases than those built from baseline outputs.
- Fine-tuning data drawn from the archives raises downstream accuracy by 8.3 percent.
- The same gains appear on both Llama-3-70B and Mistral-Large when full embedding access is available.
Where Pith is reading between the lines
- The approach could be tested on tasks whose behavior spaces are harder to factor into independent features.
- Because the evolved interface is only 32K parameters, the method may extend to models whose internal representations are inaccessible.
- If the near-independence condition holds more broadly, hybrid maps of this form could replace hand-designed behavior descriptors in other QD applications.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript presents QD-LLM, a parameter-efficient neuroevolution framework that evolves compact prompt embeddings (~32K parameters) to steer generation in frozen LLMs (70B+) inside a Quality-Diversity optimization loop. It claims 46.4% higher coverage and 41.4% higher QD-Score than QDAIF on HumanEval, MBPP and creative-writing benchmarks (p<0.001, 30 runs, A=0.94), supported by a hybrid semantic/explicit behavior characterization that supplies formal coverage bounds (Theorem 1) under measured near-independence (NMI=0.08±0.02). Downstream utility for test-case generation and fine-tuning data is also reported.
Significance. If the formal coverage bounds and the reported performance deltas hold, the work supplies a practical bridge between neuroevolution and contemporary LLMs by achieving behavioral diversity without weight updates. The parameter count, cross-model validation, and downstream task improvements would be genuine strengths.
major comments (1)
- [Theorem 1] Theorem 1 (hybrid behavior characterization): the coverage bounds are stated to hold under the observed NMI=0.08±0.02. The manuscript must explicitly show whether the proof includes an additive error term linear in mutual information or instead assumes exact independence (joint coverage = product of marginals). If the latter, the residual MI introduces an unquantified slack that directly affects the validity of the hybrid archive's coverage claim, which is load-bearing for the headline performance comparison.
minor comments (1)
- [Experimental results] The abstract reports statistical significance and effect sizes but supplies no derivation details, error-bar methodology, or full experimental protocol; these should be added to the main text or supplementary material so that the 46.4%/41.4% gains can be reproduced.
Simulated Author's Rebuttal
We thank the referee for the constructive comment on Theorem 1. We address the concern directly below and will revise the manuscript to make the treatment of residual mutual information explicit.
read point-by-point responses
-
Referee: [Theorem 1] Theorem 1 (hybrid behavior characterization): the coverage bounds are stated to hold under the observed NMI=0.08±0.02. The manuscript must explicitly show whether the proof includes an additive error term linear in mutual information or instead assumes exact independence (joint coverage = product of marginals). If the latter, the residual MI introduces an unquantified slack that directly affects the validity of the hybrid archive's coverage claim, which is load-bearing for the headline performance comparison.
Authors: We agree that the current presentation of Theorem 1 requires clarification on this point. The proof derives the joint coverage bound under the assumption of exact independence (product of marginals) and cites the measured NMI=0.08±0.02 only as empirical justification that the assumption is approximately satisfied. We will revise the manuscript to (i) state this assumption explicitly in the theorem statement, (ii) add a corollary that supplies an additive error term linear in mutual information (derived via standard information-theoretic inequalities such as the data processing inequality), and (iii) quantify the resulting slack for the observed NMI value, showing it remains below 4% relative to the reported coverage figures. These additions will be placed immediately after the theorem and will not alter the experimental results or conclusions. revision: yes
Circularity Check
No circularity: empirical gains rest on external baseline comparison; Theorem 1 presented as independent formal bound
full rationale
The reported performance improvements (46.4% coverage, 41.4% QD-Score vs. QDAIF) are measured against an external baseline with statistical validation (p<0.001, 30 runs). No equations or derivations in the provided text reduce these quantities to fitted parameters or self-defined constructs. The hybrid characterization invokes Theorem 1 under an empirically measured NMI value; this is an external validation step rather than a self-referential definition or fitted-input prediction. No self-citations, uniqueness theorems, or ansatzes from prior author work are load-bearing in the abstract or described contributions. The derivation chain is therefore self-contained against external benchmarks.
Axiom & Free-Parameter Ledger
read the original abstract
Large Language Models exhibit mode collapse, producing homogeneous outputs that fail to explore valid solution spaces. We present QD-LLM, a framework for parameter-efficient neuroevolution that evolves prompt embeddings, compact neural interfaces (~32K parameters) that steer generation in frozen LLMs (70B+ parameters), within a Quality-Diversity (QD) optimization framework. Our contributions: (1) evolved prompt embeddings via gradient-free optimization enabling behavioral steering without model fine-tuning; (2) hybrid behavior characterization combining semantic and explicit features with formal coverage bounds (Theorem 1) under validated near-independence (NMI $= 0.08 \pm 0.02$); (3) co-evolutionary variation operators including targeted behavioral mutation via finite-difference gradient estimation. On HumanEval (164 problems), MBPP, and creative writing benchmarks, QD-LLM achieves 46.4% higher coverage and 41.4% higher QD-Score than QDAIF ($p<0.001$, 30 runs, Vargha-Delaney $A=0.94$). We demonstrate downstream utility: diverse archives improve test generation (34% more edge cases) and fine-tuning data quality (8.3% accuracy gain). We validate across open-source LLMs (Llama-3-70B, Mistral-Large) with full embedding access, establishing prompt embedding evolution as an effective paradigm bridging neuroevolution and modern LLMs.
Figures
Forward citations
Cited by 1 Pith paper
-
Self-Evolving Agent Harnesses via Gated Semantic Quality-Diversity
A gated, statistically-checked self-evolution loop improves frozen agents' harnesses by +9 to +15.5 points on sealed tests across six benchmarks, retaining 86-147% of the training gain.
Reference graph
Works this paper leans on
-
[1]
Ekin Akyürek, Dale Schuurmans, Jacob Andreas, Tengyu Ma, and Denny Zhou
-
[2]
What learning algorithm is in-context learning? Investigations with lin- ear models. InThe Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net
work page 2023
-
[3]
Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Parameter-Efficient Neuroevolution for Diverse LLM Generation GECCO ’26, July 13–17, 2026, San Jose, Costa Rica Charles Sutton. 2021. Program Synthesis with Large Language Models.arXiv preprintarXiv.2108.07732 (20...
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[4]
Constitutional AI: Harmlessness from AI Feedback
Yuntao Bai, Saurav Kadavath, Sandipan Kundu, Amanda Askell, Jackson Kernion, Andy Jones, Anna Chen, Anna Goldie, Azalia Mirhoseini, Cameron McKinnon, Carol Chen, Catherine Olsson, Christopher Olah, Danny Hernan- dez, Dawn Drain, Deep Ganguli, Dustin Li, Eli Tran-Johnson, Ethan Perez, Jamie Kerr, Jared Mueller, Jeffrey Ladish, Joshua Landau, Kamal Ndousse,...
work page internal anchor Pith review Pith/arXiv arXiv 2022
-
[5]
Stanley, Grégory Schott, and Joel Lehman
Herbie Bradley, Andrew Dai, Hannah Benita Teufel, Jenny Zhang, Koen Oost- ermeijer, Marco Bellagente, Jeff Clune, Kenneth O. Stanley, Grégory Schott, and Joel Lehman. 2024. Quality-Diversity through AI Feedback. InThe Twelfth In- ternational Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net
work page 2024
-
[6]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin...
work page 2020
-
[7]
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian...
work page internal anchor Pith review Pith/arXiv arXiv 2021
-
[8]
Cédric Colas, Vashisht Madhavan, Joost Huizinga, and Jeff Clune. 2020. Scal- ing MAP-Elites to deep neuroevolution. InGECCO ’20: Genetic and Evolutionary Computation Conference, Cancún Mexico, July 8-12, 2020, Carlos Artemio Coello Coello (Ed.). ACM, 67–75. doi:10.1145/3377930.3390217
-
[9]
Antoine Cully. 2019. Autonomous skill discovery with quality-diversity and unsupervised descriptors. InProceedings of the Genetic and Evolutionary Compu- tation Conference, GECCO 2019, Prague, Czech Republic, July 13-17, 2019, Anne Auger and Thomas Stützle (Eds.). ACM, 81–89. doi:10.1145/3321707.3321804
-
[10]
Antoine Cully, Jeff Clune, Danesh Tarapore, and Jean-Baptiste Mouret. 2015. Robots that can adapt like animals.Nat.521, 7553 (2015), 503–507. doi:10.1038/ NATURE14422
work page 2015
-
[11]
Antoine Cully and Yiannis Demiris. 2018. Quality and Diversity Optimization: A Unifying Modular Framework.IEEE Trans. Evol. Comput.22, 2 (2018), 245–259. doi:10.1109/TEVC.2017.2704781
-
[12]
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. QLoRA: Efficient Finetuning of Quantized LLMs. InAdvances in Neural Informa- tion Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenk...
work page 2023
-
[13]
Li Ding, Jenny Zhang, Jeff Clune, Lee Spector, and Joel Lehman. 2024. Quality Di- versity through Human Feedback: Towards Open-Ended Diversity-Driven Opti- mization. InForty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024 (Proceedings of Machine Learning Research), Rus- lan Salakhutdinov, Zico Kolter, Kathe...
work page 2024
-
[14]
Maxence Faldor, Félix Chalumeau, Manon Flageat, and Antoine Cully. 2023. MAP-Elites with Descriptor-Conditioned Gradients and Archive Distillation into a Single Policy. InProceedings of the Genetic and Evolutionary Computation Conference, GECCO 2023, Lisbon, Portugal, July 15-19, 2023, Sara Silva and Luís Paquete (Eds.). ACM, 138–146. doi:10.1145/3583131.3590503
-
[15]
Zhangyin Feng, Daya Guo, Duyu Tang, Nan Duan, Xiaocheng Feng, Ming Gong, Linjun Shou, Bing Qin, Ting Liu, Daxin Jiang, and Ming Zhou. 2020. CodeBERT: A Pre-Trained Model for Programming and Natural Languages. InFindings of the Association for Computational Linguistics: EMNLP 2020, Online Event, 16-20 November 2020 (Findings of ACL), Trevor Cohn, Yulan He,...
-
[16]
Chrisantha Fernando, Dylan Banarse, Henryk Michalewski, Simon Osindero, and Tim Rocktäschel. 2024. Promptbreeder: Self-Referential Self-Improvement via Prompt Evolution. InForty-first International Conference on Machine Learn- ing, ICML 2024, Vienna, Austria, July 21-27, 2024 (Proceedings of Machine Learning Research), Ruslan Salakhutdinov, Zico Kolter, K...
work page 2024
-
[17]
Manon Flageat and Antoine Cully. 2024. Uncertain Quality-Diversity: Evalu- ation Methodology and New Methods for Quality-Diversity in Uncertain Do- mains.IEEE Trans. Evol. Comput.28, 4 (2024), 891–902. doi:10.1109/TEVC.2023. 3273560
-
[18]
Manon Flageat, Johann Huber, François Hélénon, Stéphane Doncieux, and An- toine Cully. 2025. Extract-QD Framework: A Generic Approach for Quality- Diversity in Noisy, Stochastic or Uncertain Domains. InProceedings of the Ge- netic and Evolutionary Computation Conference, GECCO 2025, NH Malaga Hotel, Malaga, Spain, July 14-18, 2025, Bogdan Filipic (Ed.). A...
-
[19]
Fontaine and Stefanos Nikolaidis
Matthew C. Fontaine and Stefanos Nikolaidis. 2021. Differentiable Quality Di- versity. InAdvances in Neural Information Processing Systems 34: Annual Confer- ence on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, Marc’Aurelio Ranzato, Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan (Ed...
work page 2021
-
[20]
Fontaine and Stefanos Nikolaidis
Matthew C. Fontaine and Stefanos Nikolaidis. 2023. Covariance Matrix Adapta- tion MAP-Annealing. InProceedings of the Genetic and Evolutionary Computation Conference, GECCO 2023, Lisbon, Portugal, July 15-19, 2023, Sara Silva and Luís Paquete (Eds.). ACM, 456–465. doi:10.1145/3583131.3590389
-
[21]
Fontaine, Julian Togelius, Stefanos Nikolaidis, and Amy K
Matthew C. Fontaine, Julian Togelius, Stefanos Nikolaidis, and Amy K. Hoover
-
[22]
Covariance matrix adaptation for the rapid illumination of behavior space. InGECCO ’20: Genetic and Evolutionary Computation Conference, Cancún Mexico, July 8-12, 2020, Carlos Artemio Coello Coello (Ed.). ACM, 94–102. doi:10.1145/ 3377930.3390232
-
[23]
Adam Gaier and David Ha. 2019. Weight Agnostic Neural Networks. InAdvances in Neural Information Processing Systems 32: Annual Conference on Neural Infor- mation Processing Systems 2019, NeurIPS 2019, December 8-14, 2019, Vancouver, BC, Canada, Hanna M. Wallach, Hugo Larochelle, Alina Beygelzimer, Florence d’Alché-Buc, Emily B. Fox, and Roman Garnett (Eds...
work page 2019
-
[25]
Daniele Gravina, Ahmed Khalifa, Antonios Liapis, Julian Togelius, and Geor- gios N. Yannakakis. 2019. Procedural Content Generation through Quality Di- versity. InIEEE Conference on Games, CoG 2019, London, United Kingdom, August 20-23, 2019. IEEE, 1–8. doi:10.1109/CIG.2019.8848053
-
[26]
Luca Grillotti and Antoine Cully. 2022. Unsupervised Behavior Discovery With Quality-Diversity Optimization.IEEE Trans. Evol. Comput.26, 6 (2022), 1539–
work page 2022
-
[27]
doi:10.1109/TEVC.2022.3159855
-
[28]
Qingyan Guo, Rui Wang, Junliang Guo, Bei Li, Kaitao Song, Xu Tan, Guoqing Liu, Jiang Bian, and Yujiu Yang. 2024. Connecting Large Language Models with Evolutionary Algorithms Yields Powerful Prompt Optimizers. InThe Twelfth In- ternational Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net
work page 2024
-
[29]
Nikolaus Hansen. 2023. The CMA Evolution Strategy: A Tutorial.arXiv preprint arXiv.1604.00772 (2023). https://arxiv.org/abs/1604.00772
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[30]
Francis Heylighen and Jean-Marc Dewaele. 1999. Formality of language: defi- nition, measurement and behavioral determinants.Interner Bericht, Center “Leo Apostel”, Vrije Universiteit Brüssel4, 1 (1999)
work page 1999
-
[31]
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The Curi- ous Case of Neural Text Degeneration. In8th International Conference on Learn- ing Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenRe- view.net
work page 2020
-
[32]
Neil Houlsby, Andrei Giurgiu, Stanislaw Jastrzebski, Bruna Morrone, Quentin de Laroussilhe, Andrea Gesmundo, Mona Attariyan, and Sylvain Gelly. 2019. Parameter-Efficient Transfer Learning for NLP. InProceedings of the 36th Inter- national Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA (Proceedings of Machine Learnin...
work page 2019
-
[33]
Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022. LoRA: Low-Rank Adaptation of Large GECCO ’26, July 13–17, 2026, San Jose, Costa Rica D. Guo, J. Wu, and S. M. Yiu Language Models. InThe Tenth International Conference on Learning Representa- tions, ICLR 2022, Virtual Event, April 25-29, 20...
work page 2022
-
[34]
2019.𝜀-Entropy and𝜀-Capacity of Sets in Functional Spaces (Excerpt)
AN Kolmogorov and VM Tihomirov. 2019.𝜀-Entropy and𝜀-Capacity of Sets in Functional Spaces (Excerpt). InClassics On Fractals. CRC Press, 298–339
work page 2019
-
[35]
Alexander Kraskov, Harald Stögbauer, and Peter Grassberger. 2004. Estimating mutual information.Physical Review E69, 6 (jun 2004). doi:10.1103/physreve.69. 066138
-
[36]
Brian Lester, Rami Al-Rfou, and Noah Constant. 2021. The Power of Scale for Parameter-Efficient Prompt Tuning. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, Marie-Francine Moens, Xuanjing Huang, Lucia Specia, and Scott Wen-tau Yih (...
-
[37]
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein. 2018. Visualizing the Loss Landscape of Neural Nets. InAdvances in Neural Informa- tion Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, December 3-8, 2018, Montréal, Canada, Samy Ben- gio, Hanna M. Wallach, Hugo Larochelle, Kristen Gr...
work page 2018
-
[38]
Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Ko- cetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier De- haene, Mishig Davaadorj, Joel Lamy-Poirier, João Monteiro, Oleh Shliazhko, Nicolas Gontier, Nicholas Meade, Armel Zebaze, Ming-Ho Yee, ...
work page 2023
-
[39]
Xiang Lisa Li, Ari Holtzman, Daniel Fried, Percy Liang, Jason Eisner, Tatsunori Hashimoto, Luke Zettlemoyer, and Mike Lewis. 2023. Contrastive Decoding: Open-ended Text Generation as Optimization. InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, A...
-
[40]
Xiang Lisa Li and Percy Liang. 2021. Prefix-Tuning: Optimizing Continuous Prompts for Generation. InProceedings of the 59th Annual Meeting of the As- sociation for Computational Linguistics and the 11th International Joint Confer- ence on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Pa- pers), Virtual Event, August 1-6, 2021, Chengqing Zo...
-
[41]
Yujia Li, David Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, Rémi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d’Autume, Igor Babuschkin, Xinyun Chen, Po-Sen Huang, Johannes Welbl, Sven Gowal, Alexey Cherepanov, James Molloy, Daniel J. Mankowitz, Esme Sutherland Robson, Pushm...
work page 2022
-
[42]
arXiv:https://www.science.org/doi/pdf/10.1126/science.abq1158 doi:10. 1126/science.abq1158
- [43]
-
[44]
Hong Seo Lim and Peng Qiu. 2023. Quantifying Cell-Type-Specific Differences of Single-Cell Datasets Using Uniform Manifold Approximation and Projection for Dimension Reduction and Shapley Additive exPlanations.J. Comput. Biol. 30, 7 (2023), 738–750. doi:10.1089/CMB.2022.0366
-
[45]
Xiao Liu, Kaixuan Ji, Yicheng Fu, Weng Tam, Zhengxiao Du, Zhilin Yang, and Jie Tang. 2022. P-Tuning: Prompt Tuning Can Be Comparable to Fine-tuning Across Scales and Tasks. InProceedings of the 60th Annual Meeting of the Asso- ciation for Computational Linguistics (Volume 2: Short Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, Smaranda Muresan, Pres...
work page 2022
-
[46]
Shuai Lu, Daya Guo, Shuo Ren, Junjie Huang, Alexey Svyatkovskiy, Ambro- sio Blanco, Colin B. Clement, Dawn Drain, Daxin Jiang, Duyu Tang, Ge Li, Li- dong Zhou, Linjun Shou, Long Zhou, Michele Tufano, Ming Gong, Ming Zhou, Nan Duan, Neel Sundaresan, Shao Kun Deng, Shengyu Fu, and Shujie Liu. 2021. CodeXGLUE: A Machine Learning Benchmark Dataset for Code Un...
work page 2021
-
[47]
Nelson, Herbie Bradley, Adam Gaier, Arash Moradi Karkaj, Amy K
Elliot Meyerson, Mark J. Nelson, Herbie Bradley, Adam Gaier, Arash Moradi Karkaj, Amy K. Hoover, and Joel Lehman. 2024. Language Model Crossover: Variation through Few-Shot Prompting.ACM Trans. Evol. Learn. Optim.4, 4 (2024), 27:1–27:40. doi:10.1145/3694791
-
[48]
Nasrin Mostafazadeh, Nathanael Chambers, Xiaodong He, Devi Parikh, Dhruv Batra, Lucy Vanderwende, Pushmeet Kohli, and James F. Allen. 2016. A Cor- pus and Cloze Evaluation for Deeper Understanding of Commonsense Stories. InNAACL HLT 2016, The 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tec...
-
[49]
Jean-Baptiste Mouret and Jeff Clune. 2015. Illuminating search spaces by map- ping elites.arXiv preprintarXiv.1504.04909 (2015). https://arxiv.org/abs/1504. 04909
work page internal anchor Pith review Pith/arXiv arXiv 2015
-
[50]
Olle Nilsson and Antoine Cully. 2021. Policy gradient assisted MAP-Elites. In GECCO ’21: Genetic and Evolutionary Computation Conference, Lille, France, July 10-14, 2021, Francisco Chicano and Krzysztof Krawiec (Eds.). ACM, 866–875. doi:10.1145/3449639.3459304
-
[51]
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leike, and Ryan Lowe
-
[52]
Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems 35: Annual Conference on Neu- ral Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, No- vember 28 - December 9, 2022, Sanmi Koyejo, S. Mohamed, A. Agarwal, Danielle Belgrave, K. Cho, and A. Oh (Eds.)
work page 2022
-
[53]
Thomas Pierrot, Guillaume Richard, Karim Beguir, and Antoine Cully. 2022. Multi-objective quality diversity optimization. InGECCO ’22: Genetic and Evo- lutionary Computation Conference, Boston, Massachusetts, USA, July 9 - 13, 2022, Jonathan E. Fieldsend and Markus Wagner (Eds.). ACM, 139–147. doi:10.1145/ 3512290.3528823
-
[54]
Justin K. Pugh, Lisa B. Soros, and Kenneth O. Stanley. 2016. Quality Diversity: A New Frontier for Evolutionary Computation.Frontiers Robotics AI3 (2016), 40. doi:10.3389/FROBT.2016.00040
-
[55]
Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. InProceedings of the 2019 Conference on Empir- ical Methods in Natural Language Processing and the 9th International Joint Con- ference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, Kentaro Inui, Jing Jiang, Vin...
-
[56]
Pawan Kumar, Emilien Dupont, Francisco J
Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M. Pawan Kumar, Emilien Dupont, Francisco J. R. Ruiz, Jordan S. Ellenberg, Pengming Wang, Omar Fawzi, Pushmeet Kohli, and Alhussein Fawzi
-
[57]
Pawan Kumar, Emilien Dupont, Francisco J
Mathematical discoveries from program search with large language mod- els.Nat.625, 7995 (2024), 468–475. doi:10.1038/S41586-023-06924-6
-
[58]
Tim Salimans, Jonathan Ho, Xi Chen, Szymon Sidor, and Ilya Sutskever. 2017. Evolution Strategies as a Scalable Alternative to Reinforcement Learning.arXiv preprintarXiv.1703.03864 (2017). https://arxiv.org/abs/1703.03864
work page internal anchor Pith review Pith/arXiv arXiv 2017
-
[59]
Kenneth O. Stanley, David B. D’Ambrosio, and Jason Gauci. 2009. A Hypercube- Based Encoding for Evolving Large-Scale Neural Networks.Artif. Life15, 2 (2009), 185–212. doi:10.1162/ARTL.2009.15.2.15202
-
[60]
Stanley and Risto Miikkulainen
Kenneth O. Stanley and Risto Miikkulainen. 2002. Evolving Neural Networks through Augmenting Topologies.Evolutionary Computation10, 2 (jun 2002), 99–127. doi:10.1162/106365602320169811
-
[61]
Felipe Petroski Such, Vashisht Madhavan, Edoardo Conti, Joel Lehman, Ken- neth O. Stanley, and Jeff Clune. 2018. Deep Neuroevolution: Genetic Algorithms Are a Competitive Alternative for Training Deep Neural Networks for Reinforce- ment Learning.arXiv preprintarXiv.1712.06567 (2018). https://arxiv.org/abs/ 1712.06567
work page internal anchor Pith review Pith/arXiv arXiv 2018
-
[62]
András Vargha, Harold D. Delaney, and Andras Vargha. 2000. A Critique and Improvement of the “CL” Common Language Effect Size Statistics of McGraw and Wong.Journal of Educational and Behavioral Statistics25, 2 (2000), 101. doi:10.2307/1165329
-
[63]
Chatzilygeroudis, and Jean-Baptiste Mouret
Vassilis Vassiliades, Konstantinos I. Chatzilygeroudis, and Jean-Baptiste Mouret
-
[64]
Using Centroidal Voronoi Tessellations to Scale Up the Multidimensional Archive of Phenotypic Elites Algorithm.IEEE Trans. Evol. Comput.22, 4 (2018), 623–630. doi:10.1109/TEVC.2017.2735550
-
[65]
Diverse Beam Search: Decoding Diverse Solutions from Neural Sequence Models
Ashwin K Vijayakumar, Michael Cogswell, Ramprasath R. Selvaraju, Qing Sun, Stefan Lee, David Crandall, and Dhruv Batra. 2018. Diverse Beam Search: Decoding Diverse Solutions from Neural Sequence Models.arXiv preprint arXiv.1610.02424 (2018). https://arxiv.org/abs/1610.02424 Parameter-Efficient Neuroevolution for Diverse LLM Generation GECCO ’26, July 13–1...
work page internal anchor Pith review Pith/arXiv arXiv 2018
-
[66]
Ren-Jian Wang, Ke Xue, Haopu Shang, Chao Qian, Haobo Fu, and Qiang Fu
-
[67]
Multi-objective Optimization-based Selection for Quality-Diversity by Non-surrounded-dominated Sorting. InProceedings of the Thirty-Second Inter- national Joint Conference on Artificial Intelligence, IJCAI 2023, 19th-25th August 2023, Macao, SAR, China. ijcai.org, 4335–4343. doi:10.24963/IJCAI.2023/482
-
[68]
Tongzhou Wang and Phillip Isola. 2020. Understanding Contrastive Representa- tion Learning through Alignment and Uniformity on the Hypersphere. InPro- ceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event (Proceedings of Machine Learning Research). PMLR, 9929–9939
work page 2020
-
[69]
Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi. 2021. CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Under- standing and Generation. InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, Marie-Fran...
-
[70]
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-Thought Prompting Elicits Reasoning in Large Language Models. InAdvances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Sys- tems 2022, NeurIPS 2022, New Orleans, LA, USA, Nov...
work page 2022
-
[71]
Daan Wierstra, Tom Schaul, Jan Peters, and Jürgen Schmidhuber. 2008. Nat- ural Evolution Strategies. InProceedings of the IEEE Congress on Evolutionary Computation, CEC 2008, June 1-6, 2008, Hong Kong, China. IEEE, 3381–3387. doi:10.1109/CEC.2008.4631255
-
[72]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-Judge with MT- Bench and Chatbot Arena. InAdvances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems...
work page 2023
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.