Pith. sign in

Reward tampering problems and solutions in reinforcement learning: A causal influence diagram perspective

6 Pith papers cite this work, alongside 2 external citations. Polarity classification is still indexing.

6 Pith papers citing it
2 external citations · external index

citation-role summary

background 2

citation-polarity summary

years

2026 6

roles

background 2

polarities

background 2

representative citing papers

AI Alignment via Incentives and Correction

cs.LG · 2026-05-02 · unverdicted · novelty 6.0 · 2 refs

AI alignment is reframed as a fixed-point incentive problem in a solver-auditor pipeline, solved via bilevel optimization and bandit search over reward profiles to maintain monitoring and reduce hallucinations in LLM coding tasks.

Hide to Guide: Learning via Semantic Masking

cs.LG · 2026-05-24 · unverdicted · novelty 5.0

SMEPO applies fine-grained semantic masking to expert guidance in RLVR, turning hard problems into fill-in-the-blank tasks while preserving structure, yielding up to 3.2 point accuracy gains and 4.2x faster training.

citing papers explorer

Showing 6 of 6 citing papers.