Pith. sign in

Emergent hierarchical reasoning in llms through reinforcement learning

13 Pith papers cite this work. Polarity classification is still indexing.

13 Pith papers citing it

citation-role summary

background 2

citation-polarity summary

years

2026 12 2025 1

roles

background 2

polarities

background 2

representative citing papers

RL Post-Training Builds Compositional Reasoning Strategies

cs.AI · 2026-07-08 · conditional · novelty 7.0

RL post-training composes primitive rewrite skills into reusable macro and parallel contraction strategies that solve problems inaccessible to the base model under large sampling budgets.

General Preference Reinforcement Learning

cs.LG · 2026-05-18 · unverdicted · novelty 6.0 · 3 refs

GPRL carries a k-dimensional skew-symmetric preference structure into policy updates with per-dimension advantages and a drift monitor, yielding 56.51% length-controlled win rate on AlpacaEval 2.0 from Llama-3-8B-Instruct while outperforming SimPO and SPPO on other benchmarks.

Trust Region On-Policy Distillation

cs.LG · 2026-05-31 · unverdicted · novelty 5.0

TrOPD stabilizes on-policy distillation for LLMs with trust-region learning, outlier estimation, and off-policy guidance, outperforming prior OPD methods on reasoning and code benchmarks.

Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective

cs.LG · 2025-10-11 · unverdicted · novelty 5.0

Derives a token-level entropy change approximation revealing four factors, identifies limitations in prior entropy interventions, and proposes STEER which adaptively reweights tokens to mitigate collapse and improve performance on math and coding benchmarks.

Interpreting FCDNNs via RG on Exponential Family

stat.ML · 2026-05-29 · unverdicted · novelty 4.0

For exponential family data, optimal FC-DNN parameters equal the RG fixed points of the input characteristic parameters, making training equivalent to RG coarse-graining.

citing papers explorer

Showing 13 of 13 citing papers.