Pith. sign in

Tinyv: Reducing false negatives in verification improves rl for llm reasoning

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it

years

2026 4 2025 1

representative citing papers

Trading Human Curation for Synthetic Augmentation in RLVR

cs.LG · 2026-06-02 · conditional · novelty 5.0

Gated synthetic augmentations of a 10-task human base substitute for ~87 extra human RLVR tasks on aggregate held-out pass@1, with cost-adjusted trade rate ρ_cost in [1.4×, 11.6×].

Trust Region On-Policy Distillation

cs.LG · 2026-05-31 · unverdicted · novelty 5.0

TrOPD stabilizes on-policy distillation for LLMs with trust-region learning, outlier estimation, and off-policy guidance, outperforming prior OPD methods on reasoning and code benchmarks.

citing papers explorer

Showing 5 of 5 citing papers.