Pith. sign in

Training Reinforcement Learning Agents and Humans With Difficulty-Conditioned Generators

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

We adapt Parameterized Environment Response Model (PERM), a method for training both Reinforcement Learning (RL) Agents and human learners in parameterized environments by directly modeling difficulty and ability. Inspired by Item Response Theory (IRT), PERM aligns environment difficulty with individual ability, creating a Zone of Proximal Development-based curriculum. Remarkably, PERM operates without real-time RL updates and allows for offline training, ensuring its adaptability across diverse students. We present a two-stage training process that capitalizes on PERM's adaptability, and demonstrate its effectiveness in training RL agents and humans in an empirical study.

fields

cs.LG 1

years

2026 1

verdicts

CONDITIONAL 1

representative citing papers

Predicting Task Difficulty Without Rollouts

cs.LG · 2026-08-06 · conditional · novelty 6.0

Pre-rollout task difficulty for agentic benchmarks is predictable from token-level entropy features, with Spearman rho=0.399 in-distribution and 0.225 out-of-distribution.

citing papers explorer

Showing 1 of 1 citing paper.

  • Predicting Task Difficulty Without Rollouts cs.LG · 2026-08-06 · conditional · none · ref 25 · internal anchor

    Pre-rollout task difficulty for agentic benchmarks is predictable from token-level entropy features, with Spearman rho=0.399 in-distribution and 0.225 out-of-distribution.