Pith. sign in

Adversarial reward auditing for active detection and mitigation of reward hacking

7 Pith papers cite this work. Polarity classification is still indexing.

7 Pith papers citing it

citation-role summary

background 3 method 1

citation-polarity summary

years

2026 7

representative citing papers

AI Alignment via Incentives and Correction

cs.LG · 2026-05-02 · unverdicted · novelty 6.0 · 2 refs

AI alignment is reframed as a fixed-point incentive problem in a solver-auditor pipeline, solved via bilevel optimization and bandit search over reward profiles to maintain monitoring and reduce hallucinations in LLM coding tasks.

citing papers explorer

Showing 7 of 7 citing papers.