← back to paper
arxiv: 2608.02831 · 2 revisions
Reinforcement Learning with Evolving Rubrics as Rewards for Audio Reasoning