The work develops an iterative safe planner that adjusts conformal prediction bounds across policy updates via sensitivity analysis to maintain distribution-free safety guarantees despite interaction-induced distribution shifts.
arXiv preprint arXiv:2505.09427 , year=
4 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
On-policy self-distillation with teacher flip rate yields better safety-reasoning tradeoffs than off-policy or external-teacher baselines across model scales.
ToolChain-CRC is a trajectory-level conformal risk control method for agentic AI that calibrates accept-or-intervene rules on combined step risks, adds drift-aware extensions, and includes an anytime supermartingale alarm.
Step-wise conformal labels plus linear probes recover linearly separable success/failure directions in LLM agents on ScienceWorld and AlfWorld, with preliminary steering gains.
citing papers explorer
-
Safe Planning in Interactive Environments via Iterative Policy Updates and Adversarially Robust Conformal Prediction
The work develops an iterative safe planner that adjusts conformal prediction bounds across policy updates via sensitivity analysis to maintain distribution-free safety guarantees despite interaction-induced distribution shifts.
-
Reducing the Safety Tax in LLM Safety Alignment with On-Policy Self-Distillation
On-policy self-distillation with teacher flip rate yields better safety-reasoning tradeoffs than off-policy or external-teacher baselines across model scales.
-
ToolChain-CRC: Conformal Risk Control for Agentic AI Under Retrieval and Tool-Use Drift
ToolChain-CRC is a trajectory-level conformal risk control method for agentic AI that calibrates accept-or-intervene rules on combined step risks, adds drift-aware extensions, and includes an anytime supermartingale alarm.
-
From Actions to Understanding: Conformal Interpretability of Temporal Concepts in LLM Agents
Step-wise conformal labels plus linear probes recover linearly separable success/failure directions in LLM agents on ScienceWorld and AlfWorld, with preliminary steering gains.