Pith. sign in

R-max-a general polynomial time algorithm for near-optimal reinforcement learning

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2024 1

verdicts

CONDITIONAL 1

representative citing papers

Exploration by Running Away from the Past

cs.LG · 2024-11-21 · conditional · novelty 6.0

RAMP maximizes the divergence between an agent's current and past state distributions, a proxy for Shannon entropy, and outperforms several prior exploration methods on continuous-control benchmarks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Exploration by Running Away from the Past cs.LG · 2024-11-21 · conditional · none · ref 6

    RAMP maximizes the divergence between an agent's current and past state distributions, a proxy for Shannon entropy, and outperforms several prior exploration methods on continuous-control benchmarks.