Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T01:11:50.343466Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 1 inbound Pith citation observation for arXiv:2605.08472.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T01:11:50.343466Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T16:52:29.295566Z
A source-named dated measurement, never combined with another source.
Source: cited_works
73 of 73 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ff72c57d-92f4-4a47-8303-02b273dab3f7 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Direct Preference Optimization with an Offset
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 675a0eb5-0ba6-4dd0-9d49-1842335af846 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Matharena: Evaluating llms on uncontaminated math competitions, February 2025
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4965bf03-a41e-4d63-a7ce-85511b040d3d · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Unresolved cited work
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bd904d25-fa3e-46cd-8aa9-0523b1bbb46d · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Evaluating Large Language Models Trained on Code
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0428cd94-e17c-42ca-ad2a-d88a0d704612 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Atomic Skills are the Prerequisite: When Reinforcement Learning Synthesizes Compositional Reasoning, and When It Only Amplifies
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation aac59382-4b8a-456a-9ae9-9605a9bb6128 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Training Verifiers to Solve Math Word Problems
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fd2b9ec8-0272-4fd9-b465-31a9521b468c · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4d690ac3-7f72-4110-ad4e-b70f54a63d19 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models The Vendi Score: A Diversity Evaluation Metric for Machine Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 59887040-9301-48f7-b223-8ffec54f4e1f · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 987a00bd-b665-45f1-86e2-6d4247f4054f · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2589dc55-b08f-4d7a-bb31-b9f1338913fa · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ad569163-ad8a-4d7d-87f1-466ebed15892 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models TarGEN: Targeted Data Generation with Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7789266f-52ec-495e-8127-4b7fac36b3a0 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ce05cce0-ee35-484e-822d-823e00d079fc · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b35cf74e-1a9d-4ee5-ada3-20189188e3c4 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models DeepMath-103K: A Large-Scale, Challenging, Decontaminated, and Verifiable Mathematical Dataset for Advancing Reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b9752c4f-a4f7-4230-a2e6-0a434e7047ff · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6364a0eb-cfa0-4754-91d1-44c534e3fdc1 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Unnatural instructions: Tuning language models with (almost) no human labor
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3189ae89-9f0b-4cb3-9d14-6be738866498 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Math-verify: Robust mathematical expression evaluator
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2100ae73-85ef-4301-afc0-06b2ad90b92d · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models GPT-4o System Card
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a422bd89-057c-47b4-bfda-9a446b4402be · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models OpenAI o1 System Card
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9e8e7edf-f295-47ff-9b1f-1ba7ae188728 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models The Art of Scaling Reinforcement Learning Compute for LLMs
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1eeba29b-2ec2-405c-85f1-63dc2c17debc · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Spoc: Search-based pseudocode to code.Advances in Neural Information Processing Systems, 32
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7788295f-2a5b-4967-9d86-56a858c25aae · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Efficient memory management for large language model serving with pagedattention
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e145583c-698b-40c5-8677-aaac3491cc14 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5f346081-10c9-4510-91fd-29809d2213ee · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c03e7620-157c-4327-bd61-18db927190b4 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models The measurement of observer agreement for categorical data.biometrics, pages 159–174
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c99a57c5-8e34-4783-a228-9d9191655941 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Small models struggle to learn from strong reasoners
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6edb31ee-ce2f-412a-acef-be306d12ea27 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7b989002-0fcf-4191-9927-ff9589f81044 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Simpo: Simple preference optimization with a reference-free reward.Advances in Neural Information Processing Systems, 37:124198–124235
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4592ac03-26c8-4464-8822-5e5471a9ba26 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Cross-task gener- alization via natural language crowdsourcing instructions
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 01e7acb6-7832-42c2-b2f0-a548372fbe6e · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Orca 2: Teaching Small Language Models How to Reason
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 29024956-17e6-4c5b-8957-bb8271e22e91 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Mid-training of large language models: A survey
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4b71ea29-b3bc-4005-bc28-ade95fcd600d · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Orca: Progressive Learning from Complex Explanation Traces of GPT-4
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 44ead73a-9baf-47a0-83d3-0f7efcb710e4 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models AMC 12A (2023): Problems and Solutions
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 499f4c98-07c5-4171-a6a8-4fd8ecc010a3 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Olmo 3
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 38add820-83e4-43b8-a773-ecc1659b13b7 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models New embedding models and api updates
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5fa81818-66f8-4cf8-ba06-d13b7452840a · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4f51d288-cc4b-449f-bd5f-b5ac249f82f6 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models How many data samples is an additional instruction worth? InFindings of the Association for Computational Linguistics: EACL 2023, pages 1042–1057
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a6e04f74-1168-4fdf-bafe-ff639c8e4d03 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Princeton science library
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7deff357-4d27-4d81-a351-ee0d43279d2d · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Direct preference optimization: Your language model is secretly a reward model
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a82b28cd-d1b4-458f-ac9b-ae3772e00f66 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models ThinkTuning: Instilling Cognitive Reflections without Distillation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3be10a95-fe9c-4d2a-a596-1b5b5e479f9c · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Triple Preference Optimization: Achieving Better Alignment using a Single Step Optimization
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 62eb4ce1-8673-4f51-bc0c-db154e5e3fec · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Proximal Policy Optimization Algorithms
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6171f08d-2732-40ad-a371-15a5c4404548 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Rl on incorrect synthetic data scales the efficiency of llm math reasoning by eight-fold.Advances in Neural Information Processing Systems, 37:43000–43031
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fd06ad82-9169-46d6-904a-299687785907 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cb78f1e9-767c-4971-b5e6-ed049cea6f10 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models HybridFlow: A Flexible and Efficient RLHF Framework
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0f896885-a606-47d4-a270-dd718fc19e4f · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d8c18b49-bfd4-45a2-b077-4d74ae59e0e7 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c1d08b41-a72c-4882-9aa3-fef12409b799 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models A survey on llm mid-training
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 91c3b67a-ed70-4aec-a4cf-f7afc2a807b7 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Trl: Transformer reinforce- ment learning.https://github.com/huggingface/trl
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 20af2015-3b07-41d5-ae74-883752cb9b33 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4a3322ce-3a85-4435-a2b6-27c1ac169c21 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Self-instruct: Aligning language models with self-generated instruc- tions
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1b9c4119-405b-44cd-9442-0c73ed173a06 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Finetuned Language Models Are Zero-Shot Learners
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b2c38db9-a19f-49f2-9d9d-ce3d189c299c · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Williams
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ca29502b-3249-492b-b520-311988ac9f61 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Wizardlm: Empowering large pre-trained language models to follow complex instructions
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3517e1bb-4afc-45e7-b370-e2d00fa5ba8c · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b486751e-90e9-42e3-b462-095e38d3da75 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 750955e6-33f0-481d-a306-9c33e41762f0 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models arXiv preprint arXiv:2509.25123 , year=
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1080d409-8863-4337-be3d-ec58246afb59 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4b7d1b1c-1615-47e0-94bd-54169a493084 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Star: Bootstrapping reasoning with reasoning.Advances in Neural Information Processing Systems, 35:15476–15488
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c7630bfd-da44-4e11-a038-cf37e83d1801 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models On the interplay of pre-training, mid-training, and rl on reasoning language models, 2025 a
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 02dc47df-a062-4b6c-b96a-c51be6529c44 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models American invitational mathematics examination (aime) 2024
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1983d042-f39c-4416-b394-30a0e1774b7c · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models American invitational mathematics examination (aime) 2025
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 335150b7-a233-4aee-9af1-253c8b86fb5d · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9d044ff9-65c6-4755-9401-dcf9d369aca8 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Unresolved cited work
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f2abb68c-3cf2-4e90-88dd-75cc25ee3e7c · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Unresolved cited work
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5fcea277-bb05-4a79-8c2c-50ac61738c24 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Decision: Yes
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c6405a68-7b39-448b-877c-b50aebbc9754 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models determination hope success
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 609123af-37b3-42ed-901e-e3c921da2c54 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models All the given conditions are satisfied, and the context makes sense
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 38a02fda-7189-4293-a34a-e5d91845e667 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Pappus First construct or create something that helps you explore the problem
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 59137995-f6b9-4baf-988f-37c91dd175d5 · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models First, find what seems to be the right answer or pattern by exploring ex- amples, calculating, or testing possibilities
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f095898f-082b-4c6f-900d-574fdc0f1fcd · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models So, this possibility is inconsistent with the problem statement
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 84da0b5a-0dac-461e-a3ae-39c5f09f4c8d · outbound
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models Now, adding both months together, 48 plus 24 gives 72
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8bbcf149-05dc-45c5-8b39-70200dde9818 · inbound
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.