Pith. sign in

REVIEW 2 cited by

Large Language Models Sometimes Generate Purely Negatively-Reinforced Text

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.07567 v2 pith:OMFPDWZC submitted 2023-06-13 cs.LG cs.CL

Large Language Models Sometimes Generate Purely Negatively-Reinforced Text

classification cs.LG cs.CL
keywords examplespasswordslanguagemodelstraininggeneratelargemight
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

When using adversarial training, it is common practice to train against the most egregious failures. However, this might imply using examples with sensitive information (such as leaked passwords or security vulnerabilities) as training data. One might assume that language models trained with gradient descent never generate text snippets which were only present in examples associated with the lowest possible reward. In this paper, we show that this assumption is wrong: in some situations, large language models do learn from such negatively-reinforced examples. We present a specific training setup that enables Pythia-160M to guess passwords 13% more often than it would by guessing randomly, despite only showing it these passwords on examples where the model is incentivized to not output these passwords. Our code is available at www.github.com/FabienRoger/Learning-From-Negative-Examples

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics

    cs.LG 2026-07 conditional novelty 6.0

    Pythia layer Jacobians reorganize during training into an approximately flat infrared TDOS ρ(λ)∼λ^{-0.1} with K(t)∼1/t kernels and a transient memory-self-energy maximum interpreted as critical cognitive-field formation.

  2. Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics

    cs.LG 2026-07 reject novelty 4.0

    Slow relaxation modes in Pythia transformers accumulate toward zero rate during training, yielding a near-flat infrared spectrum and 1/t memory kernels—but the claimed 'critical cognitive field formation' is not direc...