An iterative R* Decision Transformer that predicts an upper-quantile return-to-go and augments its training set with simulator-filtered high-reward trajectories beats DT, BC, and IQL on the AIGB auto-bidding benchmark.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Optimal Return-to-Go Guided Decision Transformer for Auto-Bidding in Advertisement
An iterative R* Decision Transformer that predicts an upper-quantile return-to-go and augments its training set with simulator-filtered high-reward trajectories beats DT, BC, and IQL on the AIGB auto-bidding benchmark.