Pith. sign in

REVIEW 1 cited by

Understanding Memorisation in LLMs: Dynamics, Influencing Factors, and Implications

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.19262 v1 pith:EQF64M7G submitted 2024-07-27 cs.CL cs.LG

classification cs.CLcs.LG
keywords llmsmemorisationstringsdynamicsframeworkimplicationsrandomdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Understanding whether and to what extent large language models (LLMs) have memorised training data has important implications for the reliability of their output and the privacy of their training data. In order to cleanly measure and disentangle memorisation from other phenomena (e.g. in-context learning), we create an experimental framework that is based on repeatedly exposing LLMs to random strings. Our framework allows us to better understand the dynamics, i.e., the behaviour of the model, when repeatedly exposing it to random strings. Using our framework, we make several striking observations: (a) we find consistent phases of the dynamics across families of models (Pythia, Phi and Llama2), (b) we identify factors that make some strings easier to memorise than others, and (c) we identify the role of local prefixes and global context in memorisation. We also show that sequential exposition to different random strings has a significant effect on memorisation. Our results, often surprising, have significant downstream implications in the study and usage of LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 2 citations worldwide. Full citation record

  1. Generative AI for CAD Automation: Leveraging Large Language Models for 3D Modelling

    cs.HC 2025-07 conditional novelty 3.0 of 10

    An LLM-powered FreeCAD pipeline with error-driven re-prompting succeeds on simple and moderate 3D shapes but fails on highly constrained parameterized models.

Pith tools