← back to paper
arxiv: 2607.27203 · 2 revisions
Do You Really Need to Pretrain Q-Functions for Online RL Fine-Tuning?