{"work":{"id":"a097c5d4-6d32-46ee-9826-57d532bbfc9c","openalex_id":"https://openalex.org/W4416036315","doi":"10.18653/v1/2025.emnlp-main.712","arxiv_id":"2409.12122","raw_key":null,"title":"Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement","authors":null,"authors_text":"An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li","year":2024,"venue":"cs.CL","abstract":"In this report, we present a series of math-specific large language models: Qwen2.5-Math and Qwen2.5-Math-Instruct-1.5B/7B/72B. The core innovation of the Qwen2.5 series lies in integrating the philosophy of self-improvement throughout the entire pipeline, from pre-training and post-training to inference: (1) During the pre-training phase, Qwen2-Math-Instruct is utilized to generate large-scale, high-quality mathematical data. (2) In the post-training phase, we develop a reward model (RM) by conducting massive sampling from Qwen2-Math-Instruct. This RM is then applied to the iterative evolution of data in supervised fine-tuning (SFT). With a stronger SFT model, it's possible to iteratively train and update the RM, which in turn guides the next round of SFT data iteration. On the final SFT model, we employ the ultimate RM for reinforcement learning, resulting in the Qwen2.5-Math-Instruct. (3) Furthermore, during the inference stage, the RM is used to guide sampling, optimizing the model's performance.\n  Qwen2.5-Math-Instruct supports both Chinese and English, and possess advanced mathematical reasoning capabilities, including Chain-of-Thought (CoT) and Tool-Integrated Reasoning (TIR). We evaluate our models on 10 mathematics datasets in both English and Chinese, such as GSM8K, MATH, GaoKao, AMC23, and AIME24, covering a range of difficulties from grade school level to math competition problems.","external_url":"https://arxiv.org/abs/2409.12122","cited_by_count":1,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"2409.12122","created_at":"2026-05-09T05:55:30.059603+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement","render_title":"Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement"},"hub":{"state":{"work_id":"a097c5d4-6d32-46ee-9826-57d532bbfc9c","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":186,"external_cited_by_count":1,"distinct_field_count":14,"first_pith_cited_at":"2024-10-10T14:39:33+00:00","last_pith_cited_at":"2026-07-09T07:54:39+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-21T22:19:32.421042+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":18},{"context_role":"baseline","n":5},{"context_role":"method","n":2}],"polarity_counts":[{"context_polarity":"background","n":18},{"context_polarity":"baseline","n":5},{"context_polarity":"use_method","n":2}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement","claims":[{"claim_text":"In this report, we present a series of math-specific large language models: Qwen2.5-Math and Qwen2.5-Math-Instruct-1.5B/7B/72B. The core innovation of the Qwen2.5 series lies in integrating the philosophy of self-improvement throughout the entire pipeline, from pre-training and post-training to inference: (1) During the pre-training phase, Qwen2-Math-Instruct is utilized to generate large-scale, high-quality mathematical data. (2) In the post-training phase, we develop a reward model (RM) by conducting massive sampling from Qwen2-Math-Instruct. This RM is then applied to the iterative evolutio","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Median 0.333 Table 3: Simulator-wise ATE on Omni-MATH. Mean: over ATE>0 problems. Simulator #ATE>0 Mean llm_staged 62 0.300 transform 41 0.398 intensive 34 0.416 bm25_analogy 33 0.220 numeric 18 0.303 transform_st. 14 0.479 sympy 12 0.238 5 Experiments We evaluate CIKA along four research questions (RQ1-RQ4). The base model is Qwen2.5-Math- 7B-Instruct [25], served via vLLM (RTX 4090 cluster, 5 nodes in parallel) with no parameter up- dates. Benchmarks are Omni-MATH-Rule [26] (2,821 problems, pr","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"gent agents capable of comprehensively processing omni-modal inputs, spanning video, audio, and text, to perform complex reasoning, decision-making, and plan- ning [17,51]. Recently, reinforcement learning (RL) post-training [14,24,25] has driven remarkable breakthroughs in large language models (LLMs), empowering them with robust reasoning capabilities to solve intricate mathematical prob- lems [8,10,44] and generate high-quality, functional code [27,61]. Despite these significant advancements ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"UCI Air Quality Nature 1h 7917 10 72 4 3 Solar with Weather Energy 15min 69192 20 96 1 9 ENTSO-e Load Energy 30min 85725 20 48 1 3 5.1.2 Baseline Models and Finetuning Methods.We employ three non-pretrained deep learning-based models which can capture both temporal and cross-channel correlations, making them suitable for all three kinds of forecasting tasks: i) MSD-Mixer [79], which inte- grates a layer-wise decomposition and multi-scale temporal patch- ing structure to capture sub-series variat","claim_type":"baseline","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"mathematical expert model via self-improvement.arXiv preprint arXiv:2409.12122, 2024. [50] Zhe Yang, Yichang Zhang, Tianyu Liu, Jian Yang, Junyang Lin, Chang Zhou, and Zhifang Sui. Can large language models always solve easy problems if they can solve harder ones? In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 1531-1555, 2024. [51] Yixin Ye, Zhen Huang, Yang Xiao, Ethan Chern, Shijie Xia, and Pengfei Liu. LIMO: Less is more for reasoning. InConfe","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"Runji Lin, Shijie Wang, Shuai Bai, Sinan Tan, Tianhang Zhu, Tianhao Li, Tianyu Liu, Wenbin Ge, Xiaodong Deng, Xiaohuan Zhou, Xingzhang Ren, Xinyu Zhang, Xipin Wei, Xuancheng Ren, Xuejing Liu, Yang Fan, Yang Yao, Yichang Zhang, Yu Wan, Yunfei Chu, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, Zhifang Guo, and Zhihao Fan. Qwen2 technical report. CoRR, abs/2407.10671, 2024. [67] An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li, Dayiheng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, Kemi","claim_type":"baseline","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"questions stratified across difficulty levels 2-4), andGSM8K[ 47], a grade-school word problem set (200 questions) serving as an \"easy benchmark\" control where the base model already achieves >94% accuracy. Backbones.We test three LLMs to demonstrate generality:DeepSeek-V3[ 48] (deepseek-chat), a strong open-weight reasoning model;GPT-4o-mini[ 49], OpenAI's compact model; andQwen2.5- 7B[ 50], Alibaba's 7B-parameter model. All models are queried with temperature T= 0.7 ; we collect 48 responses p","claim_type":"method","confidence":0.8,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (15 contexts).","role_counts":[{"n":15,"context_role":"background"},{"n":5,"context_role":"baseline"},{"n":2,"context_role":"method"}]},"error":null,"updated_at":"2026-05-19T21:01:41.683949+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"f6e1a34c-959d-4c3d-b317-1a65d6fe682c","orcid":null,"display_name":"An Yang"},{"id":"54001702-5e02-4890-9876-458e5739ed23","orcid":null,"display_name":"Beichen Zhang"},{"id":"785c6603-fbf9-4a35-bfb2-686be03040d6","orcid":null,"display_name":"Binyuan Hui"},{"id":"f8f90b9f-0e2f-47af-9a78-40f9ab41720a","orcid":null,"display_name":"Bofei Gao"},{"id":"cbf01de2-0d26-46f1-9d5a-d5752c92dae6","orcid":null,"display_name":"Bowen Yu"},{"id":"046bfbcc-b5ba-448d-b266-997ab128d86d","orcid":null,"display_name":"Chengpeng Li"}]},"error":null,"updated_at":"2026-05-19T21:01:41.681322+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T11:09:35.343274+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":38},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":33},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":24},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":23},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":22},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":20},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":20},{"title":"Measuring Mathematical Problem Solving With the MATH Dataset","work_id":"50652ac6-fb7c-4675-a2c2-159c241feb17","shared_citers":18},{"title":"OpenAI o1 System Card","work_id":"68d3c334-0fc9-49e3-b7b0-a69afae933e2","shared_citers":15},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":13},{"title":"Group Sequence Policy Optimization","work_id":"3a98b53b-9f52-4d95-adf7-89353c0a9a65","shared_citers":13},{"title":"Qwen2.5 Technical Report","work_id":"d8432992-4980-4a81-85c7-9fa2c2b87f85","shared_citers":13},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":12},{"title":"Qwen2.5-Coder Technical Report","work_id":"09ba463d-6377-4017-9801-444ffb94b056","shared_citers":12},{"title":"Understanding R1-Zero-Like Training: A Critical Perspective","work_id":"ec354f3b-9484-4a0c-94c8-92d4d0260835","shared_citers":12},{"title":"DeepSeek-V3 Technical Report","work_id":"57d2791d-2219-4c31-a077-afc04b12a75c","shared_citers":10},{"title":"Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?","work_id":"d854765a-e664-41c0-8655-21c4bf2e0cc4","shared_citers":10},{"title":"Tulu 3: Pushing Frontiers in Open Language Model Post-Training","work_id":"28c9dbea-056a-48c2-8000-85f809827e45","shared_citers":10},{"title":"Process Reinforcement through Implicit Rewards","work_id":"c31a2126-86f9-44f3-91f3-208d0fc1463a","shared_citers":9},{"title":"Kimi k1.5: Scaling Reinforcement Learning with LLMs","work_id":"bff96ab1-bd6a-4585-be23-74fdb51969c7","shared_citers":8},{"title":"Let's Verify Step by Step","work_id":"6d05b790-04c5-4fd2-91b2-ba1dfdd5770f","shared_citers":7},{"title":"LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code","work_id":"ea9e51ce-1e75-4182-92d8-4d25f70d2ee4","shared_citers":7},{"title":"Self-Consistency Improves Chain of Thought Reasoning in Language Models","work_id":"8c6d5a6b-b5cc-4105-9c84-9c34bb9375bb","shared_citers":7},{"title":"Chen Wang, Lai Wei, Yanzhi Zhang, Chenyang Shao, Zedong Dan, Weiran Huang, Yuzhi Zhang, and Yue Wang","work_id":"75a5258b-4143-4f2f-99f3-6d950a496305","shared_citers":6}],"time_series":[{"n":5,"year":2025},{"n":55,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T11:09:35.366998+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T11:09:42.275905+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement","claims":[{"claim_text":"In this report, we present a series of math-specific large language models: Qwen2.5-Math and Qwen2.5-Math-Instruct-1.5B/7B/72B. The core innovation of the Qwen2.5 series lies in integrating the philosophy of self-improvement throughout the entire pipeline, from pre-training and post-training to inference: (1) During the pre-training phase, Qwen2-Math-Instruct is utilized to generate large-scale, high-quality mathematical data. (2) In the post-training phase, we develop a reward model (RM) by conducting massive sampling from Qwen2-Math-Instruct. This RM is then applied to the iterative evolutio","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"Median 0.333 Table 3: Simulator-wise ATE on Omni-MATH. Mean: over ATE>0 problems. Simulator #ATE>0 Mean llm_staged 62 0.300 transform 41 0.398 intensive 34 0.416 bm25_analogy 33 0.220 numeric 18 0.303 transform_st. 14 0.479 sympy 12 0.238 5 Experiments We evaluate CIKA along four research questions (RQ1-RQ4). The base model is Qwen2.5-Math- 7B-Instruct [25], served via vLLM (RTX 4090 cluster, 5 nodes in parallel) with no parameter up- dates. Benchmarks are Omni-MATH-Rule [26] (2,821 problems, pr","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"gent agents capable of comprehensively processing omni-modal inputs, spanning video, audio, and text, to perform complex reasoning, decision-making, and plan- ning [17,51]. Recently, reinforcement learning (RL) post-training [14,24,25] has driven remarkable breakthroughs in large language models (LLMs), empowering them with robust reasoning capabilities to solve intricate mathematical prob- lems [8,10,44] and generate high-quality, functional code [27,61]. Despite these significant advancements ","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"UCI Air Quality Nature 1h 7917 10 72 4 3 Solar with Weather Energy 15min 69192 20 96 1 9 ENTSO-e Load Energy 30min 85725 20 48 1 3 5.1.2 Baseline Models and Finetuning Methods.We employ three non-pretrained deep learning-based models which can capture both temporal and cross-channel correlations, making them suitable for all three kinds of forecasting tasks: i) MSD-Mixer [79], which inte- grates a layer-wise decomposition and multi-scale temporal patch- ing structure to capture sub-series variat","claim_type":"baseline","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"mathematical expert model via self-improvement.arXiv preprint arXiv:2409.12122, 2024. [50] Zhe Yang, Yichang Zhang, Tianyu Liu, Jian Yang, Junyang Lin, Chang Zhou, and Zhifang Sui. Can large language models always solve easy problems if they can solve harder ones? In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 1531-1555, 2024. [51] Yixin Ye, Zhen Huang, Yang Xiao, Ethan Chern, Shijie Xia, and Pengfei Liu. LIMO: Less is more for reasoning. InConfe","claim_type":"background","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"Runji Lin, Shijie Wang, Shuai Bai, Sinan Tan, Tianhang Zhu, Tianhao Li, Tianyu Liu, Wenbin Ge, Xiaodong Deng, Xiaohuan Zhou, Xingzhang Ren, Xinyu Zhang, Xipin Wei, Xuancheng Ren, Xuejing Liu, Yang Fan, Yang Yao, Yichang Zhang, Yu Wan, Yunfei Chu, Yuqiong Liu, Zeyu Cui, Zhenru Zhang, Zhifang Guo, and Zhihao Fan. Qwen2 technical report. CoRR, abs/2407.10671, 2024. [67] An Yang, Beichen Zhang, Binyuan Hui, Bofei Gao, Bowen Yu, Chengpeng Li, Dayiheng Liu, Jianhong Tu, Jingren Zhou, Junyang Lin, Kemi","claim_type":"baseline","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"questions stratified across difficulty levels 2-4), andGSM8K[ 47], a grade-school word problem set (200 questions) serving as an \"easy benchmark\" control where the base model already achieves >94% accuracy. Backbones.We test three LLMs to demonstrate generality:DeepSeek-V3[ 48] (deepseek-chat), a strong open-weight reasoning model;GPT-4o-mini[ 49], OpenAI's compact model; andQwen2.5- 7B[ 50], Alibaba's 7B-parameter model. All models are queried with temperature T= 0.7 ; we collect 48 responses p","claim_type":"method","confidence":0.8,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (15 contexts).","role_counts":[{"n":15,"context_role":"background"},{"n":5,"context_role":"baseline"},{"n":2,"context_role":"method"}]},"error":null,"updated_at":"2026-05-19T21:01:41.047666+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement","claims":[{"claim_text":"In this report, we present a series of math-specific large language models: Qwen2.5-Math and Qwen2.5-Math-Instruct-1.5B/7B/72B. The core innovation of the Qwen2.5 series lies in integrating the philosophy of self-improvement throughout the entire pipeline, from pre-training and post-training to inference: (1) During the pre-training phase, Qwen2-Math-Instruct is utilized to generate large-scale, high-quality mathematical data. (2) In the post-training phase, we develop a reward model (RM) by conducting massive sampling from Qwen2-Math-Instruct. This RM is then applied to the iterative evolutio","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T11:09:37.667604+00:00"}},"summary":{"title":"Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement","claims":[{"claim_text":"In this report, we present a series of math-specific large language models: Qwen2.5-Math and Qwen2.5-Math-Instruct-1.5B/7B/72B. The core innovation of the Qwen2.5 series lies in integrating the philosophy of self-improvement throughout the entire pipeline, from pre-training and post-training to inference: (1) During the pre-training phase, Qwen2-Math-Instruct is utilized to generate large-scale, high-quality mathematical data. (2) In the post-training phase, we develop a reward model (RM) by conducting massive sampling from Qwen2-Math-Instruct. This RM is then applied to the iterative evolutio","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":38},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":33},{"title":"Training Verifiers to Solve Math Word Problems","work_id":"acab1aa8-b4d6-40e0-a3ee-25341701dca2","shared_citers":24},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":23},{"title":"Proximal Policy Optimization Algorithms","work_id":"240c67fe-d14d-4520-91c1-38a4e272ca19","shared_citers":22},{"title":"DAPO: An Open-Source LLM Reinforcement Learning System at Scale","work_id":"64019d00-0b11-4bbd-b173-b46c8fad0157","shared_citers":20},{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":20},{"title":"Measuring Mathematical Problem Solving With the MATH Dataset","work_id":"50652ac6-fb7c-4675-a2c2-159c241feb17","shared_citers":18},{"title":"OpenAI o1 System Card","work_id":"68d3c334-0fc9-49e3-b7b0-a69afae933e2","shared_citers":15},{"title":"Evaluating Large Language Models Trained on Code","work_id":"042493e9-b26f-4b4e-bbde-382072ca9b08","shared_citers":13},{"title":"Group Sequence Policy Optimization","work_id":"3a98b53b-9f52-4d95-adf7-89353c0a9a65","shared_citers":13},{"title":"Qwen2.5 Technical Report","work_id":"d8432992-4980-4a81-85c7-9fa2c2b87f85","shared_citers":13},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":12},{"title":"Qwen2.5-Coder Technical Report","work_id":"09ba463d-6377-4017-9801-444ffb94b056","shared_citers":12},{"title":"Understanding R1-Zero-Like Training: A Critical Perspective","work_id":"ec354f3b-9484-4a0c-94c8-92d4d0260835","shared_citers":12},{"title":"DeepSeek-V3 Technical Report","work_id":"57d2791d-2219-4c31-a077-afc04b12a75c","shared_citers":10},{"title":"Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?","work_id":"d854765a-e664-41c0-8655-21c4bf2e0cc4","shared_citers":10},{"title":"Tulu 3: Pushing Frontiers in Open Language Model Post-Training","work_id":"28c9dbea-056a-48c2-8000-85f809827e45","shared_citers":10},{"title":"Process Reinforcement through Implicit Rewards","work_id":"c31a2126-86f9-44f3-91f3-208d0fc1463a","shared_citers":9},{"title":"Kimi k1.5: Scaling Reinforcement Learning with LLMs","work_id":"bff96ab1-bd6a-4585-be23-74fdb51969c7","shared_citers":8},{"title":"Let's Verify Step by Step","work_id":"6d05b790-04c5-4fd2-91b2-ba1dfdd5770f","shared_citers":7},{"title":"LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code","work_id":"ea9e51ce-1e75-4182-92d8-4d25f70d2ee4","shared_citers":7},{"title":"Self-Consistency Improves Chain of Thought Reasoning in Language Models","work_id":"8c6d5a6b-b5cc-4105-9c84-9c34bb9375bb","shared_citers":7},{"title":"Chen Wang, Lai Wei, Yanzhi Zhang, Chenyang Shao, Zedong Dan, Weiran Huang, Yuzhi Zhang, and Yue Wang","work_id":"75a5258b-4143-4f2f-99f3-6d950a496305","shared_citers":6}],"time_series":[{"n":5,"year":2025},{"n":55,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"f6e1a34c-959d-4c3d-b317-1a65d6fe682c","orcid":null,"display_name":"An Yang","source":"manual","import_confidence":0.72},{"id":"54001702-5e02-4890-9876-458e5739ed23","orcid":null,"display_name":"Beichen Zhang","source":"manual","import_confidence":0.72},{"id":"785c6603-fbf9-4a35-bfb2-686be03040d6","orcid":null,"display_name":"Binyuan Hui","source":"manual","import_confidence":0.72},{"id":"f8f90b9f-0e2f-47af-9a78-40f9ab41720a","orcid":null,"display_name":"Bofei Gao","source":"manual","import_confidence":0.72},{"id":"cbf01de2-0d26-46f1-9d5a-d5752c92dae6","orcid":null,"display_name":"Bowen Yu","source":"manual","import_confidence":0.72},{"id":"046bfbcc-b5ba-448d-b266-997ab128d86d","orcid":null,"display_name":"Chengpeng Li","source":"manual","import_confidence":0.72}]}}