APARL combines a pass-rate-based adaptive sampler with KL-regularized DAPO reinforcement learning and reports F1 improvements of 17.19% in-domain and 9.59% out-of-domain for customer service anomaly detection.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.LG 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy
APARL combines a pass-rate-based adaptive sampler with KL-regularized DAPO reinforcement learning and reports F1 improvements of 17.19% in-domain and 9.59% out-of-domain for customer service anomaly detection.