REVIEW 6 cited by
Forecasting Rare Language Model Behaviors
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Standard language model evaluations can fail to capture risks that emerge only at deployment scale. For example, a model may produce safe responses during a small-scale beta test, yet reveal dangerous information when processing billions of requests at deployment. To remedy this, we introduce a method to forecast potential risks across orders of magnitude more queries than we test during evaluation. We make forecasts by studying each query's elicitation probability -- the probability the query produces a target behavior -- and demonstrate that the largest observed elicitation probabilities predictably scale with the number of queries. We find that our forecasts can predict the emergence of diverse undesirable behaviors -- such as assisting users with dangerous chemical synthesis or taking power-seeking actions -- across up to three orders of magnitude of query volume. Our work enables model developers to proactively anticipate and patch rare failures before they manifest during large-scale deployments.
Forward citations
Cited by 6 Pith papers
-
Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs
With copyable pre-release evidence, any dual-use release rule that keeps legitimate utility q must leave worst-case attacker assistance at least Γ(q)>0, so useful capability, reliable safety, and open access cannot coexist.
-
Where do LLMs Fall Short in CBT-Guided Affective Reasoning?
CBT knowledge alone does not change LLM therapeutic strategy; MCoT guidance yields only ~1.2–1.3% Protocol Leverage Force and models stay biased toward Validation & Reflection.
-
Sound Probabilistic Safety Bounds for Large Language Models
Guided expansion of a few generation-tree branches yields provably valid but extremely small lower bounds on LLM harm probability; in several reported runs the baseline Monte Carlo estimate is orders of magnitude larger.
-
Predicting LLM Safety Before Release by Simulating Deployment
Regenerating model responses on de-identified production conversation prefixes yields pre-deployment misbehavior rate forecasts that track realized production rates within 2-5x and outperform adversarial-prompt baselines.
-
Model Collapse Is Not a Bug but a Feature in Machine Unlearning for LLMs
A new method, Partial Model Collapse, iteratively fine-tunes an LLM on its own self-generated responses to conditionally collapse its output distribution on forget queries, removing private answers without the true la...
-
Rare Event Analysis of Large Language Models
Using annealed transition path sampling plus MBAR reweighting, the authors estimate TinyStories-8M completion probabilities for extreme ARI and log-probability values that are unobservable by direct sampling.
Discussion (0). Continue with ORCID to comment.