← back to paper
arxiv: 2510.24515 · 2 revisions
Learning Ordinal Response Policies in Rank-Based Stochastic Prize-Collecting Games