Pith. sign in

HelpSteer3-preference: Open human-annotated preference data across diverse tasks and languages

5 Pith papers cite this work. Polarity classification is still indexing.

5 Pith papers citing it

fields

cs.CL 3 cs.LG 2

years

2026 1 2025 4

verdicts

UNVERDICTED 5

representative citing papers

NVIDIA Nemotron 3: Efficient and Open Intelligence

cs.CL · 2025-12-24 · unverdicted · novelty 5.0

NVIDIA releases the Nemotron 3 model family with hybrid Mamba-Transformer architecture, LatentMoE, NVFP4 training, MTP layers, and multi-environment RL post-training for reasoning and agentic tasks.

Reinforcement Learning from Human Feedback

cs.LG · 2025-04-16 · unverdicted · novelty 0.0

An expository book that systematically presents RLHF methods, from reward modeling to direct alignment algorithms, aimed at readers with quantitative backgrounds.

citing papers explorer

Showing 5 of 5 citing papers.