Pith. sign in

Minimalistic gridworld environment for openai gym

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

citation-role summary

background 1

citation-polarity summary

fields

cs.AI 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

unclear 1

representative citing papers

Learning from Active Human Involvement through Proxy Value Propagation

cs.AI · 2025-02-05 · conditional · novelty 6.0

A reward-free human-in-the-loop RL method that labels human demonstrations with high Q values and intervened agent actions with low Q values, then propagates these values through TD learning to train policies across driving and gridworld tasks.

citing papers explorer

Showing 1 of 1 citing paper.

  • Learning from Active Human Involvement through Proxy Value Propagation cs.AI · 2025-02-05 · conditional · none · ref 4

    A reward-free human-in-the-loop RL method that labels human demonstrations with high Q values and intervened agent actions with low Q values, then propagates these values through TD learning to train policies across driving and gridworld tasks.