← back to paper
arxiv: 2604.20328 · 2 revisions
HyLaR: Hybrid Latent Reasoning with Decoupled Policy Optimization