Pith. sign in

REVIEW 1 cited by

Chain-of-Thought Reasoning is a Policy Improvement Operator

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.08589 v2 pith:M7HU2DM6 submitted 2023-09-15 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords reasoningchain-of-thoughtmodelslanguagemodelproblemssectorsolve
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models have astounded the world with fascinating new capabilities. However, they currently lack the ability to teach themselves new skills, relying instead on large amounts of human-generated training data. We introduce SECToR (Self-Education via Chain-of-Thought Reasoning), a proof-of-concept demonstration that language models can teach themselves new skills using chain-of-thought reasoning. During the self-learning loop, SECToR asks models to solve addition problems using chain-of-thought reasoning before training the next version of the model to solve those same problems directly without using such reasoning. This process often results in an improved model which can, when again augmented with chain-of-thought reasoning, solve even harder problems than the original model, allowing the self-learning loop to continue. Language models trained via SECToR autonomously learn to add up to the longest-length-digit numbers without access to any ground truth examples beyond an initial supervised fine-tuning phase consisting only of numbers with 6 or fewer digits. Our central hypothesis is that chain-of-thought reasoning can act as a policy improvement operator, similarly to how Monte-Carlo Tree Search is used in AlphaZero (Silver et al., 2017). We hope that this research can lead to new directions in which language models can learn to teach themselves without the need for human demonstrations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Ignore the KL Penalty! Boosting Exploration on Critical Tokens to Enhance RL Fine-Tuning

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A certainty-weighted KL penalty that down-weights the penalty on low-confidence tokens improves RL fine-tuning exploration for arithmetic in GPT-2.

Pith tools