Fitted occupancy-ratio evaluation (FORE) contracts in KL divergence under only occupancy-ratio realizability, enabling offline policy evaluation without Bellman completeness.
Proceedings of The 33rd International Conference on Machine Learning , pages =
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
PEQ-Net uses policy-aware reparameterization of ICE Q-functions and kernel mean embeddings in a shared encoder, followed by LTMLE, to jointly estimate multiple policies while constraining second-order bias for lower variance.
citing papers explorer
-
Fitted Occupancy-Ratio Evaluation without Bellman Completeness
Fitted occupancy-ratio evaluation (FORE) contracts in KL divergence under only occupancy-ratio realizability, enabling offline policy evaluation without Bellman completeness.
-
Smooth Multi-Policy Causal Effect Estimation in Longitudinal Settings
PEQ-Net uses policy-aware reparameterization of ICE Q-functions and kernel mean embeddings in a shared encoder, followed by LTMLE, to jointly estimate multiple policies while constraining second-order bias for lower variance.