SafeCtrl-RL frames LLM dialogue as a sequential decision process and uses an RL agent to iteratively refine prompts for safety at inference time.
34 User Prompt Hash 1 Hash fingerprint feature 1
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2026 1verdicts
UNVERDICTED 1representative citing papers
citing papers explorer
-
SafeCtrl-RL: Inference-Time Adaptive Behaviour Control for LLM Dialogue via RL-Driven Prompt Optimisation
SafeCtrl-RL frames LLM dialogue as a sequential decision process and uses an RL agent to iteratively refine prompts for safety at inference time.