Pith. sign in

Dynamic and generalizable process reward modeling

3 Pith papers cite this work. Polarity classification is still indexing.

3 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

fields

cs.AI 2 cs.LG 1

years

2026 3

roles

background 1

polarities

background 1

representative citing papers

OpenRCA 2.0: From Outcome Labels to Causal Process Supervision

cs.AI · 2026-06-25 · unverdicted · novelty 7.0

OpenRCA 2.0 is the first cross-system RCA benchmark with step-wise causal annotations, revealing that 11 frontier LLMs achieve 20.7% exact root-cause recovery and struggle with causal grounding (61.5% vs 76.0% ungrounded).

citing papers explorer

Showing 3 of 3 citing papers.

  • OpenRCA 2.0: From Outcome Labels to Causal Process Supervision cs.AI · 2026-06-25 · unverdicted · none · ref 27

    OpenRCA 2.0 is the first cross-system RCA benchmark with step-wise causal annotations, revealing that 11 frontier LLMs achieve 20.7% exact root-cause recovery and struggle with causal grounding (61.5% vs 76.0% ungrounded).

  • PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization cs.AI · 2026-05-18 · conditional · none · ref 27

    A two-stage probe (hidden-state estimate plus attention-based correction) yields per-step GRPO rewards that survive prefix contamination and beat external-judge and tree-search rewards in the reported benchmarks.

  • DGPO: Distribution Guided Policy Optimization for Fine Grained Credit Assignment cs.LG · 2026-05-05 · unverdicted · none · ref 32 · 2 links

    DGPO is a critic-free RL framework that uses bounded Hellinger distance and entropy-gated advantage redistribution to enable fine-grained token-level credit assignment in long CoT generations for LLM alignment, reporting SOTA results on AIME benchmarks.