Pith. sign in

Towards a unified view of large language model post-training.arXiv preprint arXiv:2509.04419

13 Pith papers cite this work. Polarity classification is still indexing.

13 Pith papers citing it

citation-role summary

background 2 baseline 1

citation-polarity summary

years

2026 11 2025 2

representative citing papers

AIPO: Learning to Reason from Active Interaction

cs.CL · 2026-05-08 · unverdicted · novelty 6.0 · 2 refs

AIPO adds active multi-agent consultation (Verify, Knowledge, Reasoning agents) plus custom importance sampling to RLVR training so LLMs expand their reasoning boundary and then operate without the agents.

GAC: Noise-Aware Adaptive Mixing for Hybrid SFT-RL Post-Training

cs.LG · 2026-05-25 · unverdicted · novelty 5.0

GAC derives adaptive mixing weights for SFT-RL hybrid post-training from online gradient variance and signal disagreement estimates, improving benchmark performance over fixed schedules with under 1% overhead.

citing papers explorer

Showing 13 of 13 citing papers.