LLM outputs drift measurably when prompts are reworded without changing meaning, instruction-tuned models drift less, and the new PBSS score quantifies this drift using embedding distance.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
When Meaning Stays the Same, but Models Drift: Evaluating Quality of Service under Token-Level Behavioral Instability in LLMs
LLM outputs drift measurably when prompts are reworded without changing meaning, instruction-tuned models drift less, and the new PBSS score quantifies this drift using embedding distance.