REVIEW 2 cited by
Large Language Models Sometimes Generate Purely Negatively-Reinforced Text
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Large Language Models Sometimes Generate Purely Negatively-Reinforced Text
read the original abstract
When using adversarial training, it is common practice to train against the most egregious failures. However, this might imply using examples with sensitive information (such as leaked passwords or security vulnerabilities) as training data. One might assume that language models trained with gradient descent never generate text snippets which were only present in examples associated with the lowest possible reward. In this paper, we show that this assumption is wrong: in some situations, large language models do learn from such negatively-reinforced examples. We present a specific training setup that enables Pythia-160M to guess passwords 13% more often than it would by guessing randomly, despite only showing it these passwords on examples where the model is incentivized to not output these passwords. Our code is available at www.github.com/FabienRoger/Learning-From-Negative-Examples
Forward citations
Cited by 2 Pith papers
-
Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics
Pythia layer Jacobians reorganize during training into an approximately flat infrared TDOS ρ(λ)∼λ^{-0.1} with K(t)∼1/t kernels and a transient memory-self-energy maximum interpreted as critical cognitive-field formation.
-
Infrared Organization and Critical Cognitive Field Formation in Transformer Dynamics
Slow relaxation modes in Pythia transformers accumulate toward zero rate during training, yielding a near-flat infrared spectrum and 1/t memory kernels—but the claimed 'critical cognitive field formation' is not direc...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.