← back to paper
arxiv: 2608.03545 · 2 revisions
Hi-TTRL: Regulating Consensus with Hints for Test-Time Reinforcement Learning