A bi-level game-theoretic optimal control plus reinforcement learning framework enables competitor-aware energy management and pit-stop scheduling that exploits aerodynamic drafting in simulated electric endurance races.
Policy invariance under reward transformations: Theory and application to reward shaping
5 Pith papers cite this work. Polarity classification is still indexing.
years
2026 5representative citing papers
STDR infers stage structure from expert videos to supply stage-transition and within-stage progress rewards, improving RL sample efficiency on 14 manipulation tasks.
PBRS-augmented RL trained in simple settings transfers zero-shot to complex UAV environments when wrapped with a CLF-CBF-QP safety filter, yielding shorter missions and formal safety guarantees.
RAID finds multiple diverse game exploits by sequentially training RL agents and masking previously discovered strategies from the reward function.
Fuzzy logic-based adaptive reward shaping improves RL convergence speed, reduces variability, and boosts success rates by up to 5% in drone racing simulations compared to standard rewards.
citing papers explorer
-
Competitor-aware Race Management for Electric Endurance Racing
A bi-level game-theoretic optimal control plus reinforcement learning framework enables competitor-aware energy management and pit-stop scheduling that exploits aerodynamic drafting in simulated electric endurance races.
-
Stage-Transition Dense Reward Modeling for Reinforcement Learning
STDR infers stage structure from expert videos to supply stage-transition and within-stage progress rewards, improving RL sample efficiency on 14 manipulation tasks.
-
Zero-Shot, Safe and Time-Efficient UAV Navigation via Potential-Based Reward Shaping, Control Lyapunov and Barrier Functions
PBRS-augmented RL trained in simple settings transfers zero-shot to complex UAV environments when wrapped with a CLF-CBF-QP safety filter, yielding shorter missions and formal safety guarantees.
-
Reward-Adaptive Iterative Discovery: A Case Study on Automated Game Testing for NHL26
RAID finds multiple diverse game exploits by sequentially training RL agents and masking previously discovered strategies from the reward function.
-
Fuzzy Logic Theory-based Adaptive Reward Shaping for Robust Reinforcement Learning (FARS)
Fuzzy logic-based adaptive reward shaping improves RL convergence speed, reduces variability, and boosts success rates by up to 5% in drone racing simulations compared to standard rewards.