Pith. sign in

REVIEW 3 cited by

Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.10627 v1 pith:6TEATSL6 submitted 2024-07-15 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords arenalearningdatamodelbattlecontinuousllmsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Assessing the effectiveness of large language models (LLMs) presents substantial challenges. The method of conducting human-annotated battles in an online Chatbot Arena is a highly effective evaluative technique. However, this approach is limited by the costs and time required for human annotation. In this paper, we introduce Arena Learning, an innovative offline strategy designed to simulate these arena battles using AI-driven annotations to evaluate battle outcomes, thus facilitating the continuous improvement of the target model through both supervised fine-tuning and reinforcement learning. Arena Learning comprises two key elements. First, it ensures precise evaluations and maintains consistency between offline simulations and online competitions via WizardArena, a pipeline developed to accurately predict the Elo rankings of various models using a meticulously designed offline test set. Our results demonstrate that WizardArena's predictions closely align with those from the online Arena. Second, it involves the continuous improvement of training data based on the battle results and the refined model. We establish a data flywheel to iteratively update the training data by highlighting the weaknesses of the target model based on its battle results, enabling it to learn from the strengths of multiple different models. We apply Arena Learning to train our target model, WizardLM-$\beta$, and demonstrate significant performance enhancements across various metrics. This fully automated training and evaluation pipeline sets the stage for continuous advancements in various LLMs via post-training. Notably, Arena Learning plays a pivotal role in the success of WizardLM-2, and this paper serves both as an exploration of its efficacy and a foundational study for future discussions related to WizardLM-2 and its derivatives.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. BALSAM: A Platform for Benchmarking Arabic Large Language Models

    cs.CL 2025-07 conditional novelty 6.0 of 10

    BALSAM is a new Arabic LLM benchmark with blind test sets, and the paper argues that LLM-based judging should replace n-gram and embedding metrics for scoring it.

  2. CodeEvo: Interaction-Driven Synthesis of Code-centric Data through Hybrid and Iterative Feedback

    cs.SE 2025-07 conditional novelty 6.0 of 10

    CodeEvo uses two interacting LLM agents with keyword-guided instruction evolution and hybrid compiler-plus-LLM feedback to synthesize high-quality instruction-code pairs for fine-tuning code models.

  3. ScIRGen: Synthesize Realistic and Large-Scale RAG Dataset for Scientific Research

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A framework and 61k QA dataset for geoscience RAG, generated from dataset metadata and papers with taxonomy-guided questions and filter-based answer validation.

Pith tools