Pith. sign in

REVIEW 1 cited by

Training Reinforcement Learning Agents and Humans With Difficulty-Conditioned Generators

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.02309 v1 pith:WAGFZAL7 submitted 2023-12-04 cs.AI cs.HCcs.LG

classification cs.AIcs.HCcs.LG
keywords trainingpermagentsabilityadaptabilitydifficultyenvironmenthumans
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We adapt Parameterized Environment Response Model (PERM), a method for training both Reinforcement Learning (RL) Agents and human learners in parameterized environments by directly modeling difficulty and ability. Inspired by Item Response Theory (IRT), PERM aligns environment difficulty with individual ability, creating a Zone of Proximal Development-based curriculum. Remarkably, PERM operates without real-time RL updates and allows for offline training, ensuring its adaptability across diverse students. We present a two-stage training process that capitalizes on PERM's adaptability, and demonstrate its effectiveness in training RL agents and humans in an empirical study.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Predicting Task Difficulty Without Rollouts

    cs.LG 2026-08 conditional novelty 6.0 of 10

    Pre-rollout task difficulty for agentic benchmarks is predictable from token-level entropy features, with Spearman rho=0.399 in-distribution and 0.225 out-of-distribution.

Pith tools