Loop-aligned supervision lets a looped Transformer generate CoT chains beyond training length, and those chains improve an auto-regressive CoT model's length generalization.
This task requires computing the minimum number of operations (insert, delete, or replace) needed to transform one sequence into another
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Enhancing Auto-regressive Chain-of-Thought through Loop-Aligned Reasoning
Loop-aligned supervision lets a looped Transformer generate CoT chains beyond training length, and those chains improve an auto-regressive CoT model's length generalization.