OpenRCA 2.0 is the first cross-system RCA benchmark with step-wise causal annotations, revealing that 11 frontier LLMs achieve 20.7% exact root-cause recovery and struggle with causal grounding (61.5% vs 76.0% ungrounded).
Dynamic and generalizable process reward modeling
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 3roles
background 1polarities
background 1representative citing papers
A two-stage probe (hidden-state estimate plus attention-based correction) yields per-step GRPO rewards that survive prefix contamination and beat external-judge and tree-search rewards in the reported benchmarks.
DGPO is a critic-free RL framework that uses bounded Hellinger distance and entropy-gated advantage redistribution to enable fine-grained token-level credit assignment in long CoT generations for LLM alignment, reporting SOTA results on AIME benchmarks.
citing papers explorer
-
OpenRCA 2.0: From Outcome Labels to Causal Process Supervision
OpenRCA 2.0 is the first cross-system RCA benchmark with step-wise causal annotations, revealing that 11 frontier LLMs achieve 20.7% exact root-cause recovery and struggle with causal grounding (61.5% vs 76.0% ungrounded).
-
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization
A two-stage probe (hidden-state estimate plus attention-based correction) yields per-step GRPO rewards that survive prefix contamination and beat external-judge and tree-search rewards in the reported benchmarks.
-
DGPO: Distribution Guided Policy Optimization for Fine Grained Credit Assignment
DGPO is a critic-free RL framework that uses bounded Hellinger distance and entropy-gated advantage redistribution to enable fine-grained token-level credit assignment in long CoT generations for LLM alignment, reporting SOTA results on AIME benchmarks.