Pith. sign in

Self-rewarding correction for mathematical reasoning.arXiv preprint arXiv:2502.19613

12 Pith papers cite this work. Polarity classification is still indexing.

12 Pith papers citing it

citation-role summary

background 1

citation-polarity summary

years

2026 10 2025 2

roles

background 1

polarities

background 1

representative citing papers

Guarded Repair for Harm-Aware Post-hoc Replacement of LLM Mathematical Reasoning

cs.CL · 2026-05-23 · unverdicted · novelty 6.0

GuardedRepair uses guarded best-of-N repair with symbolic checks, semantic diagnostics, and conservative policies to selectively replace LLM reasoning traces, raising GSM8K accuracy from 95.60% to 96.89% and ASDiv from 78.40% to 87.60% without breaking correct cases.

Can LLMs Learn to Reason Robustly under Noisy Supervision?

cs.LG · 2026-04-05 · conditional · novelty 6.0

Online Label Refinement lets LLMs learn robust reasoning from noisy supervision by correcting labels when majority answers show rising rollout success and stable history, delivering 3-4% gains on math and reasoning benchmarks even at high noise levels.

Trust Region On-Policy Distillation

cs.LG · 2026-05-31 · unverdicted · novelty 5.0

TrOPD stabilizes on-policy distillation for LLMs with trust-region learning, outlier estimation, and off-policy guidance, outperforming prior OPD methods on reasoning and code benchmarks.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning

cs.LG · 2026-02-08 · conditional · novelty 4.0

rePIRL learns token-level process rewards for LLM reasoning via a guided-cost-learning-style IRL objective, and shows these rewards improve reasoning policies on math/coding benchmarks.

citing papers explorer

Showing 12 of 12 citing papers.