pith. sign in

Training language models to follow instructions with human feedback.Ad- vances in neural information processing systems, 35:27730– 27744

2 Pith papers cite this work. Polarity classification is still indexing.

2 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

fields

cs.AI 1 cs.CV 1

years

2026 1 2025 1

roles

background 1

polarities

background 1

representative citing papers

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

cs.AI · 2025-07-29 · unverdicted · novelty 7.0

MixGRPO speeds up GRPO for flow-based image generators by restricting SDE sampling and optimization to a sliding window while using ODE elsewhere, cutting training time by up to 71% with better alignment performance.

citing papers explorer

Showing 2 of 2 citing papers.