Evaluation blindness, silent measurement failure where monitoring looks healthy while systems fail, is formalized, classified into six production classes, and found in 53% of 36 verifiable public incidents.
Constrained episodic reinforcement learning in concave-convex and knapsack settings
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
We propose an algorithm for tabular episodic reinforcement learning with constraints. We provide a modular analysis with strong theoretical guarantees for settings with concave rewards and convex constraints, and for settings with hard constraints (knapsacks). Most of the previous work in constrained reinforcement learning is limited to linear constraints, and the remaining work focuses on either the feasibility question or settings with a single episode. Our experiments demonstrate that the proposed algorithm significantly outperforms these approaches in existing constrained episodic environments.
citation-role summary
citation-polarity summary
fields
cs.LG 1years
2026 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Evaluation Blindness: How Silent Measurement Failures Corrupt AI Systems from Training to Deployment
Evaluation blindness, silent measurement failure where monitoring looks healthy while systems fail, is formalized, classified into six production classes, and found in 53% of 36 verifiable public incidents.