Pith. sign in

REVIEW 4 major objections 5 minor 57 references

WarriorCoder: Learning from Expert Battles to Augment Code Large Language Models

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Battles among open-source code LLMs train a 6.7B model to 80.5% on HumanEval without proprietary data.

desk verdict WarriorCoder is a solid code-data flywheel paper with strong empirical results, undermined by a judge-based selection signal that is never validated against execution. read the letter →

arxiv 2412.17395 v3 pith:SM2WWDC5 submitted 2024-12-23 cs.CL

classification cs.CL
keywords codegenerationLLM-as-a-judgedataflywheelEloratinginstructionminingopen-sourceLLMsexpertbattles
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that high-quality code training data can be generated from scratch by having open-source code LLMs battle each other, with no proprietary LLM annotations and no human-written seed prompts. It builds an arena where one expert attacks another with instructions sampled from its own chat-template distribution, judges vote on the correctness and helpfulness of the two responses, and a blend of local vote share and global Elo rating selects the winner. Fine-tuning a 6.7B base model on these winning responses yields 80.5% pass@1 on HumanEval and competitive results on CRUXEval and DS-1000, surpassing same-size fine-tuned baselines. If true, this offers a low-cost, scalable way to build diverse code instruction data and absorb the strengths of multiple experts.

What carries the argument

The load-bearing object is the battle arena itself: one attacker expert mines an instruction by completing its own chat-template prefix, a defender expert answers it, and the remaining experts vote as judges. The selection rule combines the local vote share $x^i_{A>B}$ with an Elo rating expectation $X^{Elo}_{A>B}$ through Equation (5), with $\alpha=0.7$, so that global consistency tempers noisy local votes. Instruction quality is controlled by four-level difficulty filtering and KCenterGreedy embedding-based compression, where KCenterGreedy is a coreset selection algorithm that chooses a diverse subset from embeddings. The mechanism's work is to convert unlabeled expert knowledge into paired responses with a winner label, which then becomes supervised fine-tuning data.

What would settle it

Retrain the same 6.7B base on the same mined instructions with responses chosen randomly instead of by judge-plus-Elo selection; if the random-response model scores within a few points of WarriorCoder on HumanEval, then the battle-selection step is not what produces the gain.

Watch

Extended reading notes

Core claim

The central claim is that a code LLM can be improved to state-of-the-art same-size performance by learning from pairwise battles among open-source expert code LLMs, rather than from data expanded by proprietary models. The authors construct an arena in which each expert alternately attacks and defends, generating instructions by completing the prefix of its own chat template, and the other experts act as judges. The winning response for each instruction is selected by combining the local judge-vote share with an Elo-rating-based global expectation, then used as supervised fine-tuning data for a 6.7B DeepSeekCoder base. The paper reports 80.5% pass@1 on HumanEval, 75.6% on HumanEval+, 76.2% on MBPP, and 64.8% on MBPP+, all without relying on proprietary LLMs, and interprets this as direct evidence that competitive data generation can absorb the strengths of multiple experts.

Load-bearing premise

Everything rests on the open-source judge LLMs being able to tell which of two code answers is more correct and helpful; if their votes are noisy or systematically biased, the winner labels that form the training data may point the model at the wrong answers.

Editorial extensions

If this is right

  • If the central claim holds, a 6.7B code model trained on battle-selected data reaches 80.5% pass@1 on HumanEval and 76.2% on MBPP, surpassing all same-size fine-tuned baselines in Table 1.
  • The same model reaches 42.9% and 45.4% pass@1 on CRUXEval input/output and 38.1% overall on DS-1000, indicating gains extend beyond simple generation into code reasoning and library usage.
  • Learning from more experts monotonically improves all four main benchmarks (Table 5), so the data flywheel should keep benefiting as the competitor pool grows.
  • Because the pipeline needs no seed dataset, no human prompts, and no proprietary LLM annotations, it lowers the cost and widens access to building code instruction data.
  • The paper also argues the mined instructions are largely novel, with ROUGE scores below 0.6 against existing datasets, so the approach adds independent training examples.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • An implication the paper leaves implicit: the same battle framework could generate fine-tuning data for other domains, but only in domains where a panel of open-source judges can reliably rank answers; judge quality would need to be validated first.
  • The Elo-plus-vote blend suggests a general recipe for aggregating pairwise preferences under noisy judges; a natural next test is whether simpler aggregators, such as majority vote alone, lose the gains on harder problems.
  • Because the instruction pool is mined from the chat templates of the five chosen experts, the diversity ceiling is set by those models; adding more or more diverse open-source experts should push the benchmarks higher, a trend the paper's Table 5 already hints at.
  • A cautious reader would want a contamination check: the ROUGE comparisons were done against CodeAlpaca and CodeUltraFeedback, not against HumanEval or MBPP; testing overlap with the evaluation sets would clarify how much of the gain is genuine generalization.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. WarriorCoder proposes a data flywheel for code LLMs in which open-source expert models compete in pairwise battles, LLM judges vote on the responses, and the target model is fine-tuned on the responses that win under a combination of local vote share and global Elo ratings. The pipeline mines instructions by prompting chat models with only the prefix of their chat templates, filters by judged difficulty, compresses via KCenterGreedy, and uses the battle winners as SFT targets. The paper reports pass@1 results on HumanEval, HumanEval+, MBPP, MBPP+, CRUXEval, and DS-1000, claiming state-of-the-art performance for a 6.7B model without proprietary LLM annotation.

Significance. If the central claims hold, the paper makes a useful contribution: it provides a fully open-source data construction pipeline for code instruction tuning, removes dependence on proprietary annotators, and reports large gains over same-scale baselines. The strengths are the concrete pipeline description, the analysis of instruction diversity and difficulty, and the multi-benchmark evaluation. The main limitation is that the selection signal (judge votes and Elo) is never validated against executable correctness, so the causal mechanism behind the gains remains unproven.

major comments (4)
  1. [Section 3.3, Eq. (5)] The summation \sum_{B\in Com\A} in Eq. (5) is not operationalized by the described arena. In each battle round only one attacker and one defender compete (Section 3.1), and Eq. (2) defines x_i^{A>B} for a single opponent B. If the final score is instead computed against all other competitors, the paper does not specify how responses from every model are obtained for the same instruction or how the Elo ratings used in the sum are updated. Since e_i^A is the selection criterion for the training response, the reader cannot reproduce or interpret Eq. (5) without clarification.
  2. [Section 3.3, Eqs. (3)-(5)] The Elo term is not an independent source of global consistency: R_A and R_B are updated in Eq. (4) from the same judge votes t_A and t_B that define the local score in Eq. (2). The combined score therefore reweights the same judge signal rather than adding a separate measurement of model strength. The paper should either justify analytically why this reweighting corrects judge noise or validate it empirically, for example by comparing selections made with and without the Elo term against a held-out execution-based correctness signal.
  3. [Section 2.3 and Appendix C] Judge reliability is the core assumption of the method, but it is never validated against executable ground truth. The paper itself notes that judge models struggle with complex problems and exhibit position, verbosity, and self-enhancement biases, and the only reported mitigation is order shuffling. Without reporting inter-judge agreement, judge agreement with test-case execution, or a human-annotated sample, the claim that WarriorCoder learns from the winner (Section 3.4) is not established; the selected responses could merely be the more verbose or stylistically preferred ones.
  4. [Table 5] The ablation varies the number of experts but does not isolate the contribution of the battle-based selection mechanism. A reader cannot tell whether the gains come from selecting the judge-preferred response, from the diversity of multiple expert responses, or simply from fine-tuning on additional open-source generated data. Adding comparisons against random selection among experts, local-vote-only selection, Elo-only selection, or a pooled dataset without any selection would be needed to support the causal claim implicit in the title and in Section 3.4.
minor comments (5)
  1. [Section 3.2] The paper states that 'we deduplicate the data and adopt judges to assess their difficulty' but does not specify which models serve as difficulty judges or how the 1-10 score is elicited; please provide the prompt and the judge model.
  2. [Section 4.1] The explanation for setting alpha to 0.7 ('because we need the Elo Rating only when judges' opinions are divided') is unclear, since alpha=0.7 makes the Elo term dominate rather than apply only in tied cases; please clarify the design.
  3. [Table 1] The 'Rely on proprietary LLMs?' column uses the symbols '%' and '!' without a legend in the table or caption, so the reader cannot interpret the last column.
  4. [Equations (2)-(4)] The paper should define what counts as a win and a draw when judges vote; currently only the vote counts t_A and t_B are given, and the mapping from votes to the actual scores s_i^{A>B} and s_i^{B>A} in Eq. (4) is not fully specified.
  5. [Section 4.4.1 and Figure 3] The ROUGE overlap analysis covers CodeAlpaca and CodeUltraFeedback but not the evaluation benchmarks used in Section 4; an overlap check against HumanEval, MBPP, CRUXEval, and DS-1000 would strengthen the contamination discussion.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: benchmark evaluation is external; Elo and local votes are two views of the same judge signal by design, not a hidden reduction.

full rationale

WarriorCoder's claimed output—a 6.7B model scoring 80.5/75.6/76.2/64.8 on HumanEval/HumanEval+/MBPP/MBPP+—is verified against external execution-based benchmarks, not against the judge votes used to construct the training set. Within the data-generation pipeline, the local vote fraction (Eq. 2) and the Elo rating (Eqs. 3–4) are indeed derived from the same LLM-judge outcomes, and Eq. 5 combines them; however, the paper explicitly frames the Elo term as a temporally aggregated reweighting of the same preference signal for 'global consistency,' not as an independent correctness measurement. The combination is therefore a disclosed design choice rather than a hidden equivalence between a prediction and its input. The completion-based instruction mining, difficulty filtering, and KCenterGreedy compression operate on instructions sampled from the expert models before any external evaluation and do not import the benchmark outcomes into the selection. Self-citations (e.g., Luo et al. 2024a Arena Learning, Luo et al. 2024b WizardCoder, Xu et al. 2024a WizardLM) appear in related-work or baseline contexts and do not carry the load-bearing derivation; no uniqueness theorem or ansatz is imported from them. The lack of judge-vs-execution agreement analysis and the absence of a random-selection ablation are genuine validity and correctness limitations, but under the stated criteria they are not circularity, because the final state-of-the-art claim does not reduce by construction to the judge-vote input.

Assumptions & free parameters 7 free parameters · 5 assumptions · 0 invented entities

The paper introduces no new physical or conceptual entities; it repurposes existing models and methods. The free parameters are the manual settings for Elo mixing, difficulty filtering, sampling, and arena scale.

free parameters (7)
  • alpha = 0.7
    Balances Elo (global) and local vote share in Eq 5; chosen by hand.
  • K = 40
    Sensitivity of Elo updates in Eq 4; chosen by hand.
  • difficulty_threshold = 6
    Instructions with d(i) >= 6 are kept; threshold chosen by hand in Eq 1.
  • temperature_settings = 1.0, 1.1, 1.2
    Nine generation configs used in instruction mining; chosen without stated tuning.
  • top_p_settings = 0.99, 0.995, 1.0
    Nine generation configs used in instruction mining; chosen without stated tuning.
  • battle_rounds = 70000
    Number of arena rounds; chosen.
  • initial_elo = not specified
    Initial Elo rating used in Eq 4 is not reported.
assumptions (5)
  • domain assumption Judge votes from open-source LLMs accurately reflect the relative quality of code responses.
    The entire winner-selection pipeline relies on this; the paper acknowledges biases and complexity limits of LLM judges in Sections 2.3 and Appendix C.
  • domain assumption Completion-based mining from chat template prefixes produces useful, diverse instructions drawn from the expert's mastered distribution.
    Section 3.2, inspired by Magpie; no direct validation that sampled instructions are high-quality beyond judge difficulty filtering.
  • domain assumption KCenterGreedy on RoBERTa embeddings preserves diversity and representativeness of instructions.
    Section 3.2; this is a standard coreset assumption, unvalidated here.
  • domain assumption Public benchmark pass@1 and pass@5 scores are reliable indicators of code ability.
    Used for all main results in Tables 1 and 2 and Table 3.
  • domain assumption Elo ratings provide a valid global measure of model skill in this arena.
    Borrowed from game rating systems and applied to LLMs in Eq 3 and Eq 4.

how reviews work

0 comments
Cite this review

Pith. "Pith review of WarriorCoder: Learning from Expert Battles to Augment Code Large Language Models." pith.science (2026). https://pith.science/paper/SM2WWDC5

@misc{pith2026241217395,
  author       = {Pith},
  title        = {Pith review of: WarriorCoder: Learning from Expert Battles to Augment Code Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/SM2WWDC5}},
  note         = {Machine review of arXiv:2412.17395}
}
read the original abstract

Despite recent progress achieved by code large language models (LLMs), their remarkable abilities are largely dependent on fine-tuning on the high-quality data, posing challenges for data collection and annotation. To address this, current methods often design various data flywheels to collect complex code instructions, enabling models to handle more intricate tasks. However, these approaches typically rely on off-the-shelf datasets and data augmentation from a limited set of proprietary LLMs (e.g., Claude, GPT4, and so on), which restricts the diversity of the constructed data and makes it prone to systemic biases. In this paper, we propose WarriorCoder, a novel paradigm learns from expert battles to address these limitations. Specifically, we create an arena where leading expert code LLMs challenge each other, with evaluations conducted by impartial judges. This competitive framework generates novel training data from scratch, leveraging the strengths of all participants. Experimental results show that WarriorCoder achieves state-of-the-art performance compared to previous models of the same size, even without relying on proprietary LLMs.

Figures

Figures reproduced from arXiv: 2412.17395 by the authors.

Figure 1
Figure 1. The comparisons between our method and traditional data flywheels. Unlike previous work, we guides the target model to learn from pairwise competi￾tions. No demand for seed datasets, human-generated prompts, or annotations from proprietary models, the target model integrates the strengths of its competitors. training is heavily dependent on the availability of high-quality data (Xu et al., 2024a), and challenges of … view at source ↗
Figure 2
Figure 2. The diagram of learning from expert battles. In each round of the arena, the attacker challenges the [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 5
Figure 5. The proportion of difficulties of mined in [PITH_FULL_IMAGE:figures/full_fig_p008_5.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: The heatmap of win rates of the selected code [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

57 extracted references · 10 canonical work pages

  1. [1]

    Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J

    Jacob Austin, Augustus Odena, Maxwell I. Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie J. Cai, Michael Terry, Quoc V. Le, and Charles Sutton. 2021. https://arxiv.org/abs/2108.07732 Program synthesis with large language models . CoRR, abs/2108.07732

  2. [2]

    Brown, Jack Clark, Sam McCandlish, Chris Olah, Benjamin Mann, and Jared Kaplan

    Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El Showk, Nelson Elhage, Zac Hatfield - Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson...

  3. [3]

    Egor Bogomolov, Aleksandra Eliseeva, Timur Galimzyanov, Evgeniy Glukhov, Anton Shapkin, Maria Tigina, Yaroslav Golubev, Alexander Kovrigin, Arie van Deursen, Maliheh Izadi, and Timofey Bryksin. 2024. https://doi.org/10.48550/ARXIV.2406.11612 Long code arena: a set of benchmarks for long-context code models . CoRR, abs/2406.11612

  4. [4]

    Sahil Chaudhary. 2023. Code alpaca: An instruction-following llama model for code generation. https://github.com/sahil280114/codealpaca

  5. [5]

    Guiming Chen, Shunian Chen, Ziche Liu, Feng Jiang, and Benyou Wang. 2024 a . https://aclanthology.org/2024.emnlp-main.474 Humans or llms as the judge? A study on judgement bias . In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, FL, USA, November 12-16, 2024 , pages 8301--8327. Association for Co...

  6. [6]

    Jie Chen, Yupeng Zhang, Bingning Wang, Xin Zhao, Ji - Rong Wen, and Weipeng Chen. 2024 b . https://aclanthology.org/2024.findings-emnlp.873 Unveiling the flaws: Exploring imperfections in synthetic data and mitigation strategies for large language models . In Findings of the Association for Computational Linguistics: EMNLP 2024, Miami, Florida, USA, Novem...

  7. [7]

    Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pond \' e de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bava...

  8. [8]

    David Cheng - Han Chiang and Hung - yi Lee. 2023. https://doi.org/10.18653/V1/2023.ACL-LONG.870 Can large language models be an alternative to human evaluations? In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023 , pages 15607--15631. Association fo...

Show all 57 references
  1. [9]

    Gonzalez, Ion Stoica, and Eric P

    Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023. https://lmsys.org/blog/2023-03-30-vicuna/ Vicuna: An open-source chatbot impressing gpt-4 with 90\

  2. [10]

    Jordan, Joseph E

    Wei - Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Banghua Zhu, Hao Zhang, Michael I. Jordan, Joseph E. Gonzalez, and Ion Stoica. 2024. https://openreview.net/forum?id=3MW8GKNyzI Chatbot arena: An open platform for evaluating l...

  3. [11]

    DeepSeek - AI, Qihao Zhu, Daya Guo, Zhihong Shao, Dejian Yang, Peiyi Wang, Runxin Xu, Y. Wu, Yukun Li, Huazuo Gao, Shirong Ma, Wangding Zeng, Xiao Bi, Zihui Gu, Hanwei Xu, Damai Dai, Kai Dong, Liyue Zhang, Yishi Piao, Zhibin Gou, Zhenda Xie, Zhewen Hao, Bingxuan Wang, Junxiao ...

  4. [12]

    Ning Ding, Yulin Chen, Bokai Xu, Yujia Qin, Shengding Hu, Zhiyuan Liu, Maosong Sun, and Bowen Zhou. 2023. https://doi.org/10.18653/V1/2023.EMNLP-MAIN.183 Enhancing chat language models by scaling high-quality instructional conversations . In Proceedings of the 2023 Conference ...

  5. [13]

    Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al - Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, Arun Rao, Aston Z...

  6. [14]

    Daniel Fried, Armen Aghajanyan, Jessy Lin, Sida Wang, Eric Wallace, Freda Shi, Ruiqi Zhong, Scott Yih, Luke Zettlemoyer, and Mike Lewis. 2023. https://openreview.net/forum?id=hQwb-lbM6EL Incoder: A generative model for code infilling and synthesis . In The Eleventh Internation...

  7. [15]

    Alex Gu, Baptiste Rozière, Hugh Leather, Armando Solar-Lezama, Gabriel Synnaeve, and Sida I. Wang. 2024. Cruxeval: A benchmark for code reasoning, understanding and execution. arXiv preprint arXiv:2401.03065

  8. [16]

    Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y. Wu, Y. K. Li, Fuli Luo, Yingfei Xiong, and Wenfeng Liang. 2024. https://doi.org/10.48550/ARXIV.2401.14196 Deepseek-coder: When the large language model meets programming - the rise...

  9. [17]

    Grundy, and Haoyu Wang

    Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John C. Grundy, and Haoyu Wang. 2023. https://doi.org/10.48550/ARXIV.2308.10620 Large language models for software engineering: A systematic literature review . CoRR, abs/2308.10620

  10. [18]

    Binyuan Hui, Jian Yang, Zeyu Cui, Jiaxi Yang, Dayiheng Liu, Lei Zhang, Tianyu Liu, Jiajun Zhang, Bowen Yu, Kai Dang, An Yang, Rui Men, Fei Huang, Xingzhang Ren, Xuancheng Ren, Jingren Zhou, and Junyang Lin. 2024. https://doi.org/10.48550/ARXIV.2409.12186 Qwen2.5-coder technica...

  11. [19]

    Naman Jain, King Han, Alex Gu, Wen - Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar - Lezama, Koushik Sen, and Ion Stoica. 2024. https://doi.org/10.48550/ARXIV.2403.07974 Livecodebench: Holistic and contamination free evaluation of large language models for code ...

  12. [20]

    Yuhang Lai, Chengxi Li, Yiming Wang, Tianyi Zhang, Ruiqi Zhong, Luke Zettlemoyer, Scott Wen tau Yih, Daniel Fried, Sida Wang, and Tao Yu. 2022. Ds-1000: A natural and reliable benchmark for data science code generation. ArXiv, abs/2211.11501

  13. [21]

    Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, Qian Liu, Evgenii Zheltonozhskii, Terry Yue Zhuo, Thomas Wang, Olivier Dehaene, Mishig Davaadorj, Joel Lamy - Poirier, Jo \ a o Mont...

  14. [22]

    Gonzalez, and Ion Stoica

    Tianle Li, Wei - Lin Chiang, Evan Frick, Lisa Dunlap, Tianhao Wu, Banghua Zhu, Joseph E. Gonzalez, and Ion Stoica. 2024. https://doi.org/10.48550/ARXIV.2406.11939 From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline . CoRR, abs/2406.11939

  15. [23]

    Yujia Li, David H. Choi, Junyoung Chung, Nate Kushman, Julian Schrittwieser, R \' e mi Leblond, Tom Eccles, James Keeling, Felix Gimeno, Agustin Dal Lago, Thomas Hubert, Peter Choy, Cyprien de Masson d'Autume, Igor Babuschkin, Xinyun Chen, Po - Sen Huang, Johannes Welbl, Sven ...

  16. [24]

    Jiawei Liu, Chunqiu Steven Xia, Yuyao Wang, and Lingming Zhang. 2023. https://openreview.net/forum?id=1qvx610Cu7 Is your code generated by chat GPT really correct? rigorous evaluation of large language models for code generation . In Thirty-seventh Conference on Neural Informa...

  17. [25]

    Jiawei Liu, Songrun Xie, Junhao Wang, Yuxiang Wei, Yifeng Ding, and Lingming Zhang. 2024. https://openreview.net/forum?id=IBCBMeAhmC Evaluating language models for efficient code generation . In First Conference on Language Modeling

  18. [26]

    Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. https://arxiv.org/abs/1907.11692 Roberta: A robustly optimized BERT pretraining approach . CoRR, abs/1907.11692

  19. [27]

    Haipeng Luo, Qingfeng Sun, Can Xu, Pu Zhao, Qingwei Lin, Jianguang Lou, Shifeng Chen, Yansong Tang, and Weizhu Chen. 2024 a . https://doi.org/10.48550/ARXIV.2407.10627 Arena learning: Build data flywheel for llms post-training via simulated chatbot arena . CoRR, abs/2407.10627

  20. [28]

    Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. 2024 b . https://openreview.net/forum?id=UnUwSIgK5W Wizardcoder: Empowering code large language models with evol-instruct . In The Twelfth International Co...

  21. [29]

    Lyu, Baishakhi Ray, Abhik Roychoudhury, Shin Hwei Tan, and Patanamon Thongtanunam

    Michael R. Lyu, Baishakhi Ray, Abhik Roychoudhury, Shin Hwei Tan, and Patanamon Thongtanunam. 2024. https://doi.org/10.48550/ARXIV.2405.02213 Automatic programming: Large language models and beyond . CoRR, abs/2405.02213

  22. [30]

    Niklas Muennighoff, Qian Liu, Armel Randy Zebaze, Qinkai Zheng, Binyuan Hui, Terry Yue Zhuo, Swayam Singh, Xiangru Tang, Leandro von Werra, and Shayne Longpre. 2024. https://openreview.net/forum?id=mw1PWNSWZP Octopack: Instruction tuning code large language models . In The Twe...

  23. [31]

    Aidar Myrzakhan, Sondos Mahmoud Bsharat, and Zhiqiang Shen. 2024. https://doi.org/10.48550/ARXIV.2406.07545 Open-llm-leaderboard: From multi-choice to open-style questions for llms evaluation, benchmark, and arena . CoRR, abs/2406.07545

  24. [32]

    Erik Nijkamp, Bo Pang, Hiroaki Hayashi, Lifu Tu, Huan Wang, Yingbo Zhou, Silvio Savarese, and Caiming Xiong. 2023. https://openreview.net/forum?id=iaYcJKpY2B\_ Codegen: An open large language model for code with multi-turn program synthesis . In The Eleventh International Conf...

  25. [33]

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, Red Avila, Igor Babuschkin, Suchir Balaji, Valerie Balcom, Paul Baltescu, Haiming Bao, Mohammad Bavarian, Jeff ...

  26. [34]

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F. Christiano, Jan Leik...

  27. [35]

    Baptiste Rozi \` e re, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, J \' e r \' e my Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton - Ferrer, Aaron Grattafiori, Wenhan Xiong,...

  28. [36]

    Ozan Sener and Silvio Savarese. 2018. https://openreview.net/forum?id=H1aIuk-RW Active learning for convolutional neural networks: A core-set approach . In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Confe...

  29. [37]

    DiJia Su, Hanlin Zhu, Yingchen Xu, Jiantao Jiao, Yuandong Tian, and Qinqing Zheng. 2025. https://arxiv.org/abs/2502.03275 Token assorted: Mixing latent and text tokens for improved language model reasoning . Preprint, arXiv:2502.03275

  30. [38]

    Qwen Team. 2024. https://qwenlm.github.io/blog/qwq-32b-preview/ Qwq: Reflect deeply on the boundaries of the unknown

  31. [39]

    Aman Singh Thakur, Kartik Choudhary, Venkat Srinik Ramayapally, Sankaran Vaidyanathan, and Dieuwke Hupkes. 2024. https://doi.org/10.48550/ARXIV.2406.12624 Judging the judges: Evaluating alignment and vulnerabilities in llms-as-judges . CoRR, abs/2406.12624

  32. [40]

    Smith, Daniel Khashabi, and Hannaneh Hajishirzi

    Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. https://doi.org/10.18653/V1/2023.ACL-LONG.754 Self-instruct: Aligning language models with self-generated instructions . In Proceedings of the 61st Annual Mee...

  33. [41]

    Joty, and Steven C

    Yue Wang, Weishi Wang, Shafiq R. Joty, and Steven C. H. Hoi. 2021. https://doi.org/10.18653/V1/2021.EMNLP-MAIN.685 Codet5: Identifier-aware unified pre-trained encoder-decoder models for code understanding and generation . In Proceedings of the 2021 Conference on Empirical Met...

  34. [42]

    Yuxiang Wei, Zhe Wang, Jiawei Liu, Yifeng Ding, and Lingming Zhang. 2024. https://openreview.net/forum?id=XUeoOBid3x Magicoder: Empowering code generation with oss-instruct . In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2...

  35. [43]

    Sahraoui

    Martin Weyssow, Aton Kamanda, and Houari A. Sahraoui. 2024. https://doi.org/10.48550/ARXIV.2403.09032 Codeultrafeedback: An llm-as-a-judge dataset for aligning large language models to coding preferences . CoRR, abs/2403.09032

  36. [44]

    Tianhao Wu, Weizhe Yuan, Olga Golovneva, Jing Xu, Yuandong Tian, Jiantao Jiao, Jason Weston, and Sainbayar Sukhbaatar. 2024 a . https://doi.org/10.48550/ARXIV.2407.19594 Meta-rewarding language models: Self-improving alignment with llm-as-a-meta-judge . CoRR, abs/2407.19594

  37. [45]

    Yutong Wu, Di Huang, Wenxuan Shi, Wei Wang, Lingzhe Gao, Shihao Liu, Ziyuan Nan, Kaizhao Yuan, Rui Zhang, Xishan Zhang, Zidong Du, Qi Guo, Yewen Pu, Dawei Yin, Xing Hu, and Yunji Chen. 2024 b . https://doi.org/10.48550/ARXIV.2407.05700 Inversecoder: Unleashing the power of ins...

  38. [46]

    Can Xu, Qingfeng Sun, Kai Zheng, Xiubo Geng, Pu Zhao, Jiazhan Feng, Chongyang Tao, Qingwei Lin, and Daxin Jiang. 2024 a . https://openreview.net/forum?id=CfXh93NDgH Wizardlm: Empowering large pre-trained language models to follow complex instructions . In The Twelfth Internati...

  39. [47]

    Zhangchen Xu, Fengqing Jiang, Luyao Niu, Yuntian Deng, Radha Poovendran, Yejin Choi, and Bill Yuchen Lin. 2024 b . https://doi.org/10.48550/ARXIV.2406.08464 Magpie: Alignment data synthesis from scratch by prompting aligned llms with nothing . CoRR, abs/2406.08464

  40. [48]

    Ratner, Ranjay Krishna, Jiaming Shen, and Chao Zhang

    Yue Yu, Yuchen Zhuang, Jieyu Zhang, Yu Meng, Alexander J. Ratner, Ranjay Krishna, Jiaming Shen, and Chao Zhang. 2023. http://papers.nips.cc/paper\_files/paper/2023/hash/ae9500c4f5607caf2eff033c67daa9d7-Abstract-Datasets\_and\_Benchmarks.html Large language model as attributed ...

  41. [49]

    Zhaojian Yu, Xin Zhang, Ning Shang, Yangyu Huang, Can Xu, Yishujie Zhao, Wenxiang Hu, and Qiufeng Yin. 2024. https://doi.org/10.18653/V1/2024.ACL-LONG.280 Wavecoder: Widespread and versatile enhancement for code large language models by instruction tuning . In Proceedings of t...

  42. [50]

    Daoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu, Bingchao Wu, Bei Guan, Yongji Wang, and Jian - Guang Lou. 2023. https://doi.org/10.18653/V1/2023.ACL-LONG.411 Large language models meet nl2code: A survey . In Proceedings of the 61st Annual Meeting of the Association for Compu...

  43. [51]

    Ruochen Zhao, Wenxuan Zhang, Yew Ken Chia, Deli Zhao, and Lidong Bing. 2024. https://doi.org/10.48550/ARXIV.2405.20267 Auto arena of llms: Automating LLM evaluations with agent peer-battles and committee discussions . CoRR, abs/2405.20267

  44. [52]

    Xing, Joseph E

    Lianmin Zheng, Wei - Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric P. Xing, Joseph E. Gonzalez, Ion Stoica, and Hao Zhang. 2024. https://openreview.net/forum?id=BOfDKxfwt0 Lmsys-chat-1m: A large-scale real-world LLM con...

  45. [53]

    Xing, Hao Zhang, Joseph E

    Lianmin Zheng, Wei - Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023 a . http://papers.nips.cc/paper\_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Ab...

  46. [54]

    Qinkai Zheng, Xiao Xia, Xu Zou, Yuxiao Dong, Shan Wang, Yufei Xue, Zihan Wang, Lei Shen, Andi Wang, Yang Li, Teng Su, Zhilin Yang, and Jie Tang. 2023 b . https://doi.org/10.48550/ARXIV.2303.17568 Codegeex: A pre-trained model for code generation with multilingual evaluations o...

  47. [55]

    Terry Yue Zhuo, Minh Chien Vu, Jenny Chim, Han Hu, Wenhao Yu, Ratnadira Widyasari, Imam Nur Bani Yusuf, Haolan Zhan, Junda He, Indraneil Paul, Simon Brunner, Chen Gong, Thong Hoang, Armel Randy Zebaze, Xiaoheng Hong, Wen - Ding Li, Jean Kaddour, Ming Xu, Zhihan Zhang, Prateek ...

  48. [56]

    online" 'onlinestring :=

    ENTRY address archivePrefix author booktitle chapter edition editor eid eprint eprinttype howpublished institution journal key month note number organization pages publisher school series title type volume year doi pubmed url lastchecked label extra.label sort.label short.list...

  49. [57]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.