← back to paper
arxiv: 2506.15068 · 2 revisions
Semantically-Aware Rewards for Open-Ended R1 Training in Free-Form Generation