Pith. sign in

Query The Agent: Improving sample efficiency through epistemic uncertainty estimation

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Curricula for goal-conditioned reinforcement learning agents typically rely on poor estimates of the agent's epistemic uncertainty or fail to consider the agents' epistemic uncertainty altogether, resulting in poor sample efficiency. We propose a novel algorithm, Query The Agent (QTA), which significantly improves sample efficiency by estimating the agent's epistemic uncertainty throughout the state space and setting goals in highly uncertain areas. Encouraging the agent to collect data in highly uncertain states allows the agent to improve its estimation of the value function rapidly. QTA utilizes a novel technique for estimating epistemic uncertainty, Predictive Uncertainty Networks (PUN), to allow QTA to assess the agent's uncertainty in all previously observed states. We demonstrate that QTA offers decisive sample efficiency improvements over preexisting methods.

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

Uncertainty Prioritized Experience Replay

cs.LG · 2025-06-10 · conditional · novelty 6.0

UPER uses ensemble-based epistemic and aleatoric uncertainty to compute an information gain priority for experience replay, outperforming TD-error prioritization on Atari-57.

citing papers explorer

Showing 1 of 1 citing paper.

  • Uncertainty Prioritized Experience Replay cs.LG · 2025-06-10 · conditional · none · ref 1 · internal anchor

    UPER uses ensemble-based epistemic and aleatoric uncertainty to compute an information gain priority for experience replay, outperforming TD-error prioritization on Atari-57.