A PPO agent with a hindsight regret reward, block-bootstrapped synthetic data, and a transaction-cost curriculum rebalances a 60/40 portfolio and beats it on return in three out-of-sample periods.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
q-fin.PM 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Regret-Optimized Portfolio Enhancement through Deep Reinforcement Learning and Future Looking Rewards
A PPO agent with a hindsight regret reward, block-bootstrapped synthetic data, and a transaction-cost curriculum rebalances a 60/40 portfolio and beats it on return in three out-of-sample periods.