Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:45:00.496048Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 6 inbound Pith citation observations for arXiv:2412.17256.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T05:45:00.496048Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:48:20.040368Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-09T00:45:49.091631Z
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 57bd1add-80a5-4f4a-bfed-da43f8b57f74 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners without RM
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1a3f4f53-b9cd-43a6-8d16-5130ae0ed74e · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners The Llama 3 Herd of Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1161d321-4de4-4b1d-bf7c-ec018ac1e99d · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Reinforced Self-Training (ReST) for Language Modeling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 189e7789-9db7-4d1e-bfba-94e870164db9 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Measuring Mathematical Problem Solving With the MATH Dataset
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 988c32f2-3d7e-4137-8584-c808bf1b912e · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners V-STaR: Training Verifiers for Self-Taught Reasoners
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20a3b666-2f59-44bb-909f-7832fbc1c03b · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Large Language Models Can Self-Improve
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fa27432-b4e1-4c8f-803d-7575ee8dda83 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Population Based Training of Neural Networks
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d34ce381-1dad-4898-882b-b863d28954f1 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Mistral 7B
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf865fb5-1a51-436a-9246-5f398f760285 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Hyp-RL : Hyperparameter Optimization by Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62626659-c487-4d49-a835-979d4507486f · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Making Large Language Models Better Reasoners with Step-Aware Verifier
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02490e2a-525c-4095-a99c-9a435be3b40e · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Let's Verify Step by Step
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b0b947a-e388-4dd9-818e-146b63b02312 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners SGDR: Stochastic Gradient Descent with Warm Restarts
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb2798a1-0234-499b-9b34-8e1879d888f6 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Iterative Reasoning Preference Optimization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08c31d2a-9b31-4b6a-96c1-cb51afef8579 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fad7db1f-9d20-4785-908b-ce55168e1715 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51da734d-8572-4aed-8ed8-984e51bdb9b9 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b24583ef-ae6f-4a7c-be2e-6dae4539ebcb · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d31095f-0620-4f90-9bc9-d8493252ea8f · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Solving math word problems with process- and outcome-based feedback
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97c6282f-1662-4f36-a741-ab1f761e8f68 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Planning In Natural Language Improves LLM Search For Code Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbc79678-3d02-4d0c-b88a-822269dacd70 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners 12 Published as a conference paper at ICLR 2025 Wikipedia contributors
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 312975e6-0706-4e03-9a73-5d4d1b5c4012 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Progress or Regress? Self-Improvement Reversal in Post-training
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe972380-60d1-4233-b33b-410a1da387aa · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Large Language Models as Optimizers
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7aae4dd8-3a27-4efb-810f-44026970f759 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fddc3b9-30c7-4ffb-beef-85ce43b30eed · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Answer" indicates matching against the ground-truth final answer, and
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b09d34dc-1998-430e-98a0-f12635aa6079 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners For the MATH dataset, we follow previous settings (Lightman et al., 2023; Wang et al., 2024b; Sun et al.,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b7a18344-fcdf-4179-abc9-d898f2afc52a · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners For baselines, we uniformly sample 32 candidate responses per query with a temperature of 0.4
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cbaea97a-41f4-4645-ad90-703476551a4d · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners For the Process Reward Model (PRM), we automatically generate process annotations following the MATH-Shepherd approach (Wang et al., 2024b)
Reference 128
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9f49c1a7-e212-4bc3-a9ba-bc786c6bdd6b · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
Reference 2006
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df973591-8902-41af-a403-faec8186a1c9 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners AutoRL Hyperparameter Landscapes
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 678e6045-3047-4e48-b3d9-247cf1a95873 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners CodeT: Code Generation with Generated Tests
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ca22ffc-ce6f-4c0b-8923-0f88d285d181 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Training Verifiers to Solve Math Word Problems
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8ad65a8-f03b-48cb-98ef-df405ef0036e · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Hyperparameter Tuning for Deep Reinforcement Learning Applications
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7bcffeca-efd5-4ceb-ac2b-23840ab31f17 · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Online Learning Rate Adaptation with Hypergradient Descent
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dbbc932-b27f-4484-87f8-3bc30eb3bbae · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fded7f3-63dd-46f9-83cf-134bc088685f · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Evaluating Large Language Models Trained on Code
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba1be8f6-1e9c-46e5-9890-b523a464221c · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Teaching Large Language Models to Reason with Reinforcement Learning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cb6829f-b254-48fb-a735-126e4665de2f · outbound
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Sample-Efficient Automated Deep Reinforcement Learning
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1b98c00d-2666-435b-a578-81451eae74a0 · inbound
From System 1 to System 2: A Survey of Reasoning Large Language Models B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
Reference 190
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 50fadc9b-6b87-48f1-af34-12ddb1617969 · inbound
A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
Reference 150
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74de4af0-8adf-4a44-9744-56519d94e536 · inbound
Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95a82f25-5c48-47a9-b428-55d1ef7f0dd4 · inbound
A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c2f13c9-6066-49f8-848f-e4594bbe359a · inbound
Do Self-Evolving Agents Forget? Capability Degradation and Preservation in Lifelong LLM Agent Adaptation B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2d9421a3-209a-4604-adcc-b94e054a9bd5 · inbound
MILES: Modular Instruction Memory with Learnable Selection for Self-Improving LLM Reasoning B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.