Pith. sign in

hub

and Lewis, Mike , editor =

21 Pith papers cite this work, alongside 152 external citations. Polarity classification is still indexing.

21 Pith papers citing it
152 external citations · Crossref

hub tools

citation-role summary

background 1 dataset 1

citation-polarity summary

years

2026 20 2025 1

representative citing papers

ICRL: Learning to Internalize Self-Critique with Reinforcement Learning

cs.AI · 2026-05-13 · unverdicted · novelty 6.0

ICRL uses joint RL training of solver and critic with distribution-calibration re-weighting and role-wise advantage estimation to internalize critique into unassisted LLM performance, yielding 6.4-point gains on agentic tasks and 7.0 on math reasoning with Qwen3 models.

An Information-Theoretic Criterion for Efficient Data Synthesis

cs.LG · 2026-05-11 · unverdicted · novelty 6.0

Synthetic data improves models only in information-open generation-training loops with external signals, and coarser signals like binary correctness enable better generalization by converging to the most information-efficient component.

Agentic Reinforced Policy Optimization

cs.LG · 2025-07-26 · unverdicted · novelty 6.0

ARPO adds entropy-based adaptive rollouts and stepwise advantage attribution to RL for LLM agents, outperforming prior trajectory-level methods on 13 benchmarks with half the tool budget.

TASR: Training-Free Adaptive Stopping for Iterative Retrieval

cs.IR · 2026-06-11 · unverdicted · novelty 5.0

TASR provides a training-free predicate that stops iterative retrieval on repeated normalized answers plus calibrated logit margin above 0.25, retaining 94.8% of fixed-k=5 F1 at 62.6% of the calls across 32 configurations.

ReCal: Reward Calibration for RL-based LLM Routing

cs.LG · 2026-06-10 · unverdicted · novelty 5.0

ReCal introduces hierarchical reward decomposition and distribution-aware optimization to address ambiguous credit assignment and optimization bias in RL-based LLM routing.

citing papers explorer

Showing 21 of 21 citing papers.