An offline method combines process reward models with tabular dynamic programming to teach LLM agents when to request interventions under a limited budget.
org/CorpusID:247613322
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Self-Regulation and Requesting Interventions
An offline method combines process reward models with tabular dynamic programming to teach LLM agents when to request interventions under a limited budget.