Pith. sign in

Prioritized level replay

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

fields

cs.LG 2 cs.RO 2

years

2026 3 2025 1

representative citing papers

Robots Need More than VLA and World Models

cs.RO · 2026-06-04 · unverdicted · novelty 5.0

The paper identifies four missing interfaces (data autolabelling, embodiment retargeting, physics-grounded world models, and video-based reward inference) as the central bottleneck beyond VLA scaling for robot intelligence.

Trading Human Curation for Synthetic Augmentation in RLVR

cs.LG · 2026-06-02 · conditional · novelty 5.0

Gated synthetic augmentations of a 10-task human base substitute for ~87 extra human RLVR tasks on aggregate held-out pass@1, with cost-adjusted trade rate ρ_cost in [1.4×, 11.6×].

citing papers explorer

Showing 4 of 4 citing papers.

  • Bridging Performance and Generalization in Reinforcement Learning for Agile Flight cs.RO · 2026-06-25 · unverdicted · none · ref 16

    RL framework for agile drone racing combines task-aware switching and physically informed procedural track generation to achieve 7.4x better zero-shot generalization to unseen tracks while maintaining competitive speeds.

  • Robots Need More than VLA and World Models cs.RO · 2026-06-04 · unverdicted · none · ref 199

    The paper identifies four missing interfaces (data autolabelling, embodiment retargeting, physics-grounded world models, and video-based reward inference) as the central bottleneck beyond VLA scaling for robot intelligence.

  • Trading Human Curation for Synthetic Augmentation in RLVR cs.LG · 2026-06-02 · conditional · none · ref 26

    Gated synthetic augmentations of a 10-task human base substitute for ~87 extra human RLVR tasks on aggregate held-out pass@1, with cost-adjusted trade rate ρ_cost in [1.4×, 11.6×].

  • Learning to Reason at the Frontier of Learnability cs.LG · 2025-02-17 · unverdicted · none · ref 27

    A curriculum sampling questions with high variance in success rate improves reinforcement learning performance for LLM reasoning tasks.