LLM self-reports predict behavior selectively: TPB reaches human-level coherence within shared conversations but collapses across sessions for primed behaviors, unlike Big 5, with persona prompting stabilizing reports but not actions.
Language models transmit behavioural traits through hidden signals in data.Nature, 652:615–621
4 Pith papers cite this work. Polarity classification is still indexing.
years
2026 4verdicts
UNVERDICTED 4representative citing papers
Proposes high-temperature synthetic canaries and auxiliary-model auditing to improve empirical privacy measurement for LLM fine-tuning and synthetic data generation.
Interaction-layer antidistillation watermarks use system-prompt-induced behavioral markers like explicit follow-up questions that transfer to distilled student models at 45-89% relative fidelity and can be audited via black-box LLM-as-judge queries.
A learned linear activation bridge achieves high alignment (cosine ~0.97) between Pythia-160M and Pythia-410M states but produces no improvement in downstream multi-hop answering when injected into the receiver.
citing papers explorer
-
Rethinking Psychometric Evaluation of LLMs: When and Why Self-Reports Predict Behavior
LLM self-reports predict behavior selectively: TPB reaches human-level coherence within shared conversations but collapses across sessions for primed behaviors, unlike Big 5, with persona prompting stabilizing reports but not actions.
-
Advancing the State-of-the-Art in Empirical Privacy Auditing
Proposes high-temperature synthetic canaries and auxiliary-model auditing to improve empirical privacy measurement for LLM fine-tuning and synthetic data generation.
-
Asking Back: Interaction-Layer Antidistillation Watermarks
Interaction-layer antidistillation watermarks use system-prompt-induced behavioral markers like explicit follow-up questions that transfer to distilled student models at 45-89% relative fidelity and can be audited via black-box LLM-as-judge queries.
-
A Negative Result on Cross-Model Activation Transfer in a Pythia Multi-Hop Setting
A learned linear activation bridge achieves high alignment (cosine ~0.97) between Pythia-160M and Pythia-410M states but produces no improvement in downstream multi-hop answering when injected into the receiver.