Unbiasing Truncated Backpropagation Through Time

arxiv: 1705.08209 · v1 · pith:M5SGWWQ5new · submitted 2017-05-23 · 💻 cs.NE · cs.LG

Unbiasing Truncated Backpropagation Through Time

Corentin Tallec , Yann Ollivier This is my paper

classification 💻 cs.NE cs.LG

keywords truncatedbpttartbpbackpropagationcomputationaltimewhilebenefits

0 comments p. Extension

pith:M5SGWWQ5 Add to your LaTeX paper

What is a Pith Number?

\usepackage{pith}
\pithnumber{M5SGWWQ5}

Prints a linked pith:M5SGWWQ5 badge after your title and writes the identifier into PDF metadata. Compiles on arXiv with no extra files. Learn more

read the original abstract

Truncated Backpropagation Through Time (truncated BPTT) is a widespread method for learning recurrent computational graphs. Truncated BPTT keeps the computational benefits of Backpropagation Through Time (BPTT) while relieving the need for a complete backtrack through the whole data sequence at every step. However, truncation favors short-term dependencies: the gradient estimate of truncated BPTT is biased, so that it does not benefit from the convergence guarantees from stochastic gradient theory. We introduce Anticipated Reweighted Truncated Backpropagation (ARTBP), an algorithm that keeps the computational benefits of truncated BPTT, while providing unbiasedness. ARTBP works by using variable truncation lengths together with carefully chosen compensation factors in the backpropagation equation. We check the viability of ARTBP on two tasks. First, a simple synthetic task where careful balancing of temporal dependencies at different scales is needed: truncated BPTT displays unreliable performance, and in worst case scenarios, divergence, while ARTBP converges reliably. Second, on Penn Treebank character-level language modelling, ARTBP slightly outperforms truncated BPTT.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Generative Recursive Reasoning
cs.AI 2026-05 unverdicted novelty 6.0

GRAM turns recursive latent reasoning into a generative probabilistic model via stochastic trajectories and amortized variational inference, claiming better performance on structured reasoning tasks than deterministic...