LLMs amplify the least-dominant context, so one bad sentence among twenty good ones distorts answers; joint judgment-and-answer fine-tuning (RW-Steering) stabilizes response quality across contamination levels from 0% to 95%.
inappropri- ateness
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Context Engineering for Trustworthiness: Rescorla Wagner Steering Under Mixed and Inappropriate Contexts
LLMs amplify the least-dominant context, so one bad sentence among twenty good ones distorts answers; joint judgment-and-answer fine-tuning (RW-Steering) stabilizes response quality across contamination levels from 0% to 95%.