GenAC introduces generative critics with chain-of-thought reasoning and in-context conditioning to improve value approximation and downstream RL performance in LLMs compared to value-based and value-free baselines.
Segmental advantage estimation: Enhancing ppo for long-context llm training
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2verdicts
UNVERDICTED 2representative citing papers
SD-GRPO extends GRPO by computing per-segment advantages via z-normalization of verifiable segment rewards, yielding gains on long-form VL tasks with varying semantic independence across segments.
citing papers explorer
-
Bringing Value Models Back: Generative Critics for Value Modeling in LLM Reinforcement Learning
GenAC introduces generative critics with chain-of-thought reasoning and in-context conditioning to improve value approximation and downstream RL performance in LLMs compared to value-based and value-free baselines.
-
SD-GRPO: Verifiable Segment Decomposition for Long-Form Vision-Language Generation
SD-GRPO extends GRPO by computing per-segment advantages via z-normalization of verifiable segment rewards, yielding gains on long-form VL tasks with varying semantic independence across segments.