A framework combining GRPO, LLM-as-judge rewards, and OPA-style policy checks improves a 1B medical assistant's domain-scope adherence by 70.9% in a 100-scenario LLM-judge evaluation.
Superintelligence: Paths, Dangers, Strategies
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CY 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
ArGen: Auto-Regulation of Generative AI via GRPO and Policy-as-Code
A framework combining GRPO, LLM-as-judge rewards, and OPA-style policy checks improves a 1B medical assistant's domain-scope adherence by 70.9% in a 100-scenario LLM-judge evaluation.