Pith. sign in

Paper Citation Record · LEDGER

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners

As of 20 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 6 inbound Pith citation observations for arXiv:2412.17256.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.17256 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T05:45:00.496048Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:48:20.040368Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-09T00:45:49.091631Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 57bd1add-80a5-4f4a-bfed-da43f8b57f74 · outbound

This paper cites without RM.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners without RM

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:01.090419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:45:00.496048Z digest=sha256:aeb8abfdc468f7be8f73602d506fe25be3a4aa0439c1929a87e9f76403379f80

Observation 1a3f4f53-b9cd-43a6-8d16-5130ae0ed74e · outbound

This paper cites The Llama 3 Herd of Models.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.344187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.344187Z digest=sha256:fe3e24d61b79ddd60066b35c08795b287e4cc8ef3f786ba146b618026c8aaf84

Observation 1161d321-4de4-4b1d-bf7c-ec018ac1e99d · outbound

This paper cites Reinforced Self-Training (ReST) for Language Modeling.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Reinforced Self-Training (ReST) for Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.354898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.354898Z digest=sha256:3ed895d44e9402dd9781217bf2846b6cec274410410400d5a191bea67a832c28

Observation 189e7789-9db7-4d1e-bfba-94e870164db9 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Measuring Mathematical Problem Solving With the MATH Dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.364355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.364355Z digest=sha256:e72e7e276de087d6d3e15166d21caa764247de945417bcdbbde9220f976eec0f

Observation 988c32f2-3d7e-4137-8584-c808bf1b912e · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.369599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.369599Z digest=sha256:4d786f2d73f13724f313c2d680804e5d9e901749ae7c13e030f988492c888fef

Observation 20a3b666-2f59-44bb-909f-7832fbc1c03b · outbound

This paper cites Large Language Models Can Self-Improve.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Large Language Models Can Self-Improve

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.374724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.374724Z digest=sha256:52f00b9b55a30f81e584ec555950617ba15582064338c2bce6e25a985ff44456

Observation 1fa27432-b4e1-4c8f-803d-7575ee8dda83 · outbound

This paper cites Population Based Training of Neural Networks.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Population Based Training of Neural Networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.379674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.379674Z digest=sha256:b59f6c3b69e4213c1882212d8095ed50cd03135df905c8a9bbe1c0a4cf516283

Observation d34ce381-1dad-4898-882b-b863d28954f1 · outbound

This paper cites Mistral 7B.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Mistral 7B

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.384910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.384910Z digest=sha256:ec58cf9f68acabb55ba8032a24cb7583c0c305bd9dfdf347d835381f5f3ea3c5

Observation cf865fb5-1a51-436a-9246-5f398f760285 · outbound

This paper cites Hyp-RL : Hyperparameter Optimization by Reinforcement Learning.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Hyp-RL : Hyperparameter Optimization by Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.389727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.389727Z digest=sha256:af376c2e02eb7efeb1729308bde0119a3fbba2d3569751fa98b4b0f07ed41b12

Observation 62626659-c487-4d49-a835-979d4507486f · outbound

This paper cites Making Large Language Models Better Reasoners with Step-Aware Verifier.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Making Large Language Models Better Reasoners with Step-Aware Verifier

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.398974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.398974Z digest=sha256:b44a5f69ec28d57884e4a7fcf5850175dfb55bea9c365e79aec1cf1a0b83b0c0

Observation 02490e2a-525c-4095-a99c-9a435be3b40e · outbound

This paper cites Let's Verify Step by Step.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Let's Verify Step by Step

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.403199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.403199Z digest=sha256:e0519e894720120d01808f052dffb890024e89588e8d3015e2ae4ecb7ae3f6f3

Observation 6b0b947a-e388-4dd9-818e-146b63b02312 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.407606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.407606Z digest=sha256:6f19c0cc33896e08e29ea46a137e0631a17636c0bedf8d71d915c47a50529717

Observation bb2798a1-0234-499b-9b34-8e1879d888f6 · outbound

This paper cites Iterative Reasoning Preference Optimization.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Iterative Reasoning Preference Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.416659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.416659Z digest=sha256:34c8d06178124744398ea5c58f1e8152be8d836472cdc61e8856cbc5e2e18efd

Observation 08c31d2a-9b31-4b6a-96c1-cb51afef8579 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.421322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.421322Z digest=sha256:56055184ee57537d3f957e9946f9c6f0e2c84cd20164c587ffbe370f8460aa64

Observation fad7db1f-9d20-4785-908b-ce55168e1715 · outbound

This paper cites A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners A disciplined approach to neural network hyper-parameters: Part 1 -- learning rate, batch size, momentum, and weight decay

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.431542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.431542Z digest=sha256:202449ab89093d6a52512f52ad109471fdcb6f644a26052852b296ca53c7fd1c

Observation 51da734d-8572-4aed-8ed8-984e51bdb9b9 · outbound

This paper cites Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Easy-to-Hard Generalization: Scalable Alignment Beyond Human Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.436365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.436365Z digest=sha256:8997f24ae74612391d2f154ba5a38fcac5a78dd002856ea2f0255a685c0d6e82

Observation b24583ef-ae6f-4a7c-be2e-6dae4539ebcb · outbound

This paper cites DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.441500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.441500Z digest=sha256:795cbcc43a4870fa829d8d79d9122a1a9099b5e5e8d809c18e9288f05c616288

Observation 9d31095f-0620-4f90-9bc9-d8493252ea8f · outbound

This paper cites Solving math word problems with process- and outcome-based feedback.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Solving math word problems with process- and outcome-based feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.446537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.446537Z digest=sha256:b9d5a27654425e0a255a4c7bce30900234f7ffe0dbf3516da86b1b30159c6a98

Observation 97c6282f-1662-4f36-a741-ab1f761e8f68 · outbound

This paper cites Planning In Natural Language Improves LLM Search For Code Generation.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Planning In Natural Language Improves LLM Search For Code Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.451299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.451299Z digest=sha256:df1b0627636cd966666003cb43858a2678f4501a5adbad16ed444d2f74bf1518

Observation cbc79678-3d02-4d0c-b88a-822269dacd70 · outbound

This paper cites 12 Published as a conference paper at ICLR 2025 Wikipedia contributors.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners 12 Published as a conference paper at ICLR 2025 Wikipedia contributors

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:01.184571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:45:00.456151Z digest=sha256:f1ab75063c90e411a4e74b31d1697bda421a48f9073c07b95cebc27c12395f20

Observation 312975e6-0706-4e03-9a73-5d4d1b5c4012 · outbound

This paper cites Progress or Regress? Self-Improvement Reversal in Post-training.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Progress or Regress? Self-Improvement Reversal in Post-training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.460966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.460966Z digest=sha256:e18d51f8d5cb2319572f18fadf263e0892a3b722280d41f9238282d1a451f1b7

Observation fe972380-60d1-4233-b33b-410a1da387aa · outbound

This paper cites Large Language Models as Optimizers.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Large Language Models as Optimizers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.466197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.466197Z digest=sha256:42d29686e0dbce6f4546f106581bdb91e1a8e9bc7c12463551c5044c3fc45ea3

Observation 7aae4dd8-3a27-4efb-810f-44026970f759 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.471178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.471178Z digest=sha256:e8a23fe3d0a2f80bff5fd92eed9ddf1b312a55f4aaf98edb443ef4acc1a3baa5

Observation 3fddc3b9-30c7-4ffb-beef-85ce43b30eed · outbound

This paper cites Answer" indicates matching against the ground-truth final answer, and.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Answer" indicates matching against the ground-truth final answer, and

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:01.163780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:45:00.476231Z digest=sha256:a8f1bb1e8f2ab280252ae9406b51c6e7b90e10924e37cd766f2e8cd002b14d2a

Observation b09d34dc-1998-430e-98a0-f12635aa6079 · outbound

This paper cites For the MATH dataset, we follow previous settings (Lightman et al., 2023; Wang et al., 2024b; Sun et al.,.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners For the MATH dataset, we follow previous settings (Lightman et al., 2023; Wang et al., 2024b; Sun et al.,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:01.145898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:45:00.481101Z digest=sha256:307b7721d42e8db647a3dcaab8527238bcea767b3aa56a55acdbb3090dfbf89a

Observation b7a18344-fcdf-4179-abc9-d898f2afc52a · outbound

This paper cites For baselines, we uniformly sample 32 candidate responses per query with a temperature of 0.4.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners For baselines, we uniformly sample 32 candidate responses per query with a temperature of 0.4

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:01.108400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:45:00.491519Z digest=sha256:c35c64562ed8ec2d9ae8aea2b50a1b459f41e680db8624193afd3dbf73d56909

Observation cbaea97a-41f4-4645-ad90-703476551a4d · outbound

This paper cites For the Process Reward Model (PRM), we automatically generate process annotations following the MATH-Shepherd approach (Wang et al., 2024b).

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners For the Process Reward Model (PRM), we automatically generate process annotations following the MATH-Shepherd approach (Wang et al., 2024b)

Reference 128

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T05:45:01.127958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:45:00.486533Z digest=sha256:8d4b9e56ec23592413603c804b231564e6a301e2793741dc0086d57781e21f92

Observation 9f49c1a7-e212-4bc3-a9ba-bc786c6bdd6b · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 2006

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.426531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.426531Z digest=sha256:f29018ed9a5e161d58bdb0bf09aad6f932e5eee5f0390fdb3f98bc7091388570

Observation df973591-8902-41af-a403-faec8186a1c9 · outbound

This paper cites AutoRL Hyperparameter Landscapes.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners AutoRL Hyperparameter Landscapes

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.411942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.411942Z digest=sha256:1a4b0b6c13a2a41e69d65ec41dfb99c138a529ccee601ddb0b24139118742a24

Observation 678e6045-3047-4e48-b3d9-247cf1a95873 · outbound

This paper cites CodeT: Code Generation with Generated Tests.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners CodeT: Code Generation with Generated Tests

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.323238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.323238Z digest=sha256:67f77b27638f5e00879e768863cd6a238f6b06040f40ec35b131c1fe0900b4b1

Observation 5ca22ffc-ce6f-4c0b-8923-0f88d285d181 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Training Verifiers to Solve Math Word Problems

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.338933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.338933Z digest=sha256:562318164342b00d6535a08056f2b2c8e75db7b55abe4b19dd1af289e5dbc48f

Observation b8ad65a8-f03b-48cb-98ef-df405ef0036e · outbound

This paper cites Hyperparameter Tuning for Deep Reinforcement Learning Applications.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Hyperparameter Tuning for Deep Reinforcement Learning Applications

Reference 2019

Resolution
verified exact
local_arxiv, observed 2026-08-11T05:45:00.793440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:45:00.394591Z digest=sha256:461665d432ed8357163a69ef0e28de829953ac9eab0080a1484df1bcb8eaf30e

Observation 7bcffeca-efd5-4ceb-ac2b-23840ab31f17 · outbound

This paper cites Online Learning Rate Adaptation with Hypergradient Descent.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Online Learning Rate Adaptation with Hypergradient Descent

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.318006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.318006Z digest=sha256:e3ecaf7160c768c7469c3bd5a6fba4ae8fea3e18b97ad5709328e71209d98f38

Observation 7dbbc932-b27f-4484-87f8-3bc30eb3bbae · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.333827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.333827Z digest=sha256:54158dd170b2626a991408b1c32531fb7fcf4a6575b4275fe181cc47d51dd29b

Observation 9fded7f3-63dd-46f9-83cf-134bc088685f · outbound

This paper cites Evaluating Large Language Models Trained on Code.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Evaluating Large Language Models Trained on Code

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.328549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.328549Z digest=sha256:00264fe5647e78fe732067c712a054e03e06f35f97cfaaa926706968a46c3a45

Observation ba1be8f6-1e9c-46e5-9890-b523a464221c · outbound

This paper cites Teaching Large Language Models to Reason with Reinforcement Learning.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Teaching Large Language Models to Reason with Reinforcement Learning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T05:45:00.359505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:45:00.359505Z digest=sha256:0413906eb6416c4aaf08ebb8dc27e322077d45906f82630a934835869a2a719c

Observation 9cb6829f-b254-48fb-a735-126e4665de2f · outbound

This paper cites Sample-Efficient Automated Deep Reinforcement Learning.

B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners Sample-Efficient Automated Deep Reinforcement Learning

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-11T05:45:00.963917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T05:45:00.350019Z digest=sha256:9e0ff1d43435d085490ce5f0cfafcfc9d3627a4a265b1e0007ab520bd1e82777

Pith citing papers

Observation 1b98c00d-2666-435b-a578-81451eae74a0 · inbound

From System 1 to System 2: A Survey of Reasoning Large Language Models cites this paper.

From System 1 to System 2: A Survey of Reasoning Large Language Models B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners

Reference 190

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:36:24.263902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T01:36:23.845366Z digest=sha256:69e2044dc3b402d1bcc19edbd7ad5fe6e81006b5006c73dcbecccc3b90bb55b8

Observation 50fadc9b-6b87-48f1-af34-12ddb1617969 · inbound

A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law cites this paper.

A Survey of Slow Thinking-based Reasoning LLMs using Reinforced Learning and Inference-time Scaling Law B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners

Reference 150

Resolution
unresolved
no resolver link, observed 2026-08-16T00:48:20.040368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:48:20.040368Z digest=sha256:b0b97fa8ccee92f98ceb59cbeefae4807dbf26b7a62c9ecc9b05cacdd223ba38

Observation 74de4af0-8adf-4a44-9744-56519d94e536 · inbound

Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning cites this paper.

Stepwise Reasoning Checkpoint Analysis: A Test Time Scaling Method to Enhance LLMs' Reasoning B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:19.802170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:45:19.802170Z digest=sha256:a3ed94de630786b7db99fd3cb28869ecc5cee2af57a656cb8d082239a06c91b6

Observation 95a82f25-5c48-47a9-b428-55d1ef7f0dd4 · inbound

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula cites this paper.

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T02:43:38.300114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:43:38.300114Z digest=sha256:f184a429b28410467c9f757a6a0ba640e7e33b070ab131909f86aa910fc44a8d

Observation 0c2f13c9-6066-49f8-848f-e4594bbe359a · inbound

Do Self-Evolving Agents Forget? Capability Degradation and Preservation in Lifelong LLM Agent Adaptation cites this paper.

Do Self-Evolving Agents Forget? Capability Degradation and Preservation in Lifelong LLM Agent Adaptation B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:31:29.178804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:10:31.784413Z digest=sha256:c265fb6a0ea5508111db89c0eee89f7627005622abb7b39a2b69cef5d9400d97

Observation 2d9421a3-209a-4604-adcc-b94e054a9bd5 · inbound

MILES: Modular Instruction Memory with Learnable Selection for Self-Improving LLM Reasoning cites this paper.

MILES: Modular Instruction Memory with Learnable Selection for Self-Improving LLM Reasoning B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-09T00:45:49.092945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-09T00:45:00.847714Z digest=sha256:d3a67d85b6bca4521d4a4f339defd60bc50e0406ebdd4b60e67280d6dd346bb4