Pith. sign in

arXiv preprint arXiv:2505.14674 , year=

8 Pith papers cite this work. Polarity classification is still indexing.

8 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

years

2026 3 2025 5

roles

background 1

polarities

background 1

representative citing papers

Counsel: A Meta-Evaluation Dataset for Agentic Tasks

cs.AI · 2026-06-19 · unverdicted · novelty 7.0

Counsel is a new dataset of LLM-generated process critiques on agent benchmarks paired with human labels on error location and reasoning quality, achieving 0.78 Krippendorff alpha.

Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs

cs.LG · 2025-08-27 · conditional · novelty 6.0

GSR jointly trains LLMs to generate candidate solutions and refine a superior final answer from them, achieving state-of-the-art performance on five mathematical benchmarks while transferring across model scales.

Trust Region On-Policy Distillation

cs.LG · 2026-05-31 · unverdicted · novelty 5.0

TrOPD stabilizes on-policy distillation for LLMs with trust-region learning, outlier estimation, and off-policy guidance, outperforming prior OPD methods on reasoning and code benchmarks.

VRPRM: Process Reward Modeling via Visual Reasoning

cs.LG · 2025-08-05 · conditional · novelty 5.0

VRPRM combines 3.6K CoT-PRM SFT data with 50K non-CoT PRM RL data to train a visual PRM that beats a 400K-data non-thinking PRM and boosts best-of-N accuracy.

citing papers explorer

Showing 8 of 8 citing papers.