Offline-trained models, especially reward-weighted behavioral cloning, predict historical oncology trial portfolios more accurately than frontier LLM agents on a new 881-episode benchmark.
When does return-conditioned supervised learning work for offline reinforcement learning? In Advances in Neural Information Processing Systems (NeurIPS), 2022
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents
Offline-trained models, especially reward-weighted behavioral cloning, predict historical oncology trial portfolios more accurately than frontier LLM agents on a new 881-episode benchmark.