Pith. sign in

REVIEW 5 cited by

Why Exposure Bias Matters: An Imitation Learning Perspective of Error Accumulation in Language Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2204.01171 v3 pith:IDTKCDDH submitted 2022-04-03 cs.CL cs.AIcs.LG

Why Exposure Bias Matters: An Imitation Learning Perspective of Error Accumulation in Language Generation

classification cs.CL cs.AIcs.LG
keywords biasexposuregenerationaccumulationhypothesisimitationlanguagelearning
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Current language generation models suffer from issues such as repetition, incoherence, and hallucinations. An often-repeated hypothesis is that this brittleness of generation models is caused by the training and the generation procedure mismatch, also referred to as exposure bias. In this paper, we verify this hypothesis by analyzing exposure bias from an imitation learning perspective. We show that exposure bias leads to an accumulation of errors, analyze why perplexity fails to capture this accumulation, and empirically show that this accumulation results in poor generation quality. Source code to reproduce these experiments is available at https://github.com/kushalarora/quantifying_exposure_bias

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Cross-Tokenizer On-Policy Distillation via Byte-Prefix Marginalization

    cs.LG 2026-07 conditional novelty 7.0

    Byte-Prefix Marginalization maps a teacher's next-token distribution onto the student's vocabulary through shared byte prefixes plus an explicit residual, giving a mass-preserving target for on-policy distillation acr...

  2. Neural operator discovery from heterogeneous trajectories

    cs.LG 2026-07 conditional novelty 6.0

    Trajectory grouping plus a low-dimensional latent bottleneck lets a neural operator discover each system's hidden governing factors and extrapolate to unseen systems.

  3. Protein Autoregressive Modeling via Multiscale Structure Generation

    cs.LG 2026-02 unverdicted novelty 6.0

    PAR is a multi-scale autoregressive transformer framework for protein backbone generation that uses coarse-to-fine prediction, noisy context learning, and flow-based decoding to achieve high-quality unconditional and ...

  4. Flow marching for a generative PDE foundation model

    cs.LG 2025-09 unverdicted novelty 6.0

    Flow Marching jointly samples noise and physical time to learn a velocity field for generative PDE modeling, paired with a latent autoencoder and efficient transformer for large-scale pretraining on 2.5M trajectories.

  5. When Do Autoregressive Sequence Models Forecast Physical Wavefields? A Controlled Study on Synthetic Seismograms

    cs.LG 2026-06 unverdicted novelty 5.0

    Multi-token prediction accounts for nearly all rollout stability gains on synthetic three-component seismograms, with sharp dependence on context covering the full P-S interval and magnitude-based losses unable to pre...