A looped padded Transformer with parallel gold-CoT cross-entropy supervision matches explicit CoT accuracy at 3B scale and is 2.5–6.9× faster in the thought phase.
Training large language models to reason in a continuous latent space
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it