Pith. sign in

← back to paper

Review history

arxiv: 2607.26991 · 2 revisions

RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

  1. 2026-08-01 CONDITIONAL MODERATE v1.3.0-daily-deepseek novelty 5.0
    50106 ms 23572 in 7210 out 2026-08-01T10:18:25.995725+00:00
  2. 2026-07-30 CONDITIONAL MODERATE v1.2.0-daily-grok45 novelty 6.0
    85532 ms 29308 in 4111 out 2026-07-30T14:42:23.861266+00:00