Pith. sign in

Exploration in deep reinforcement learning: From single-agent to multiagent domain

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

representative citing papers

HAEPO: History-Aggregated Exploratory Policy Optimization

cs.LG · 2025-08-26 · conditional · novelty 4.0

HAEPO weights each trajectory by its softmax-normalized cumulative log-likelihood, adds entropy and KL penalties, and matches or slightly surpasses PPO, GRPO, and DPO on small RL and summarization tasks.

citing papers explorer

Showing 1 of 1 citing paper.

  • HAEPO: History-Aggregated Exploratory Policy Optimization cs.LG · 2025-08-26 · conditional · none · ref 16

    HAEPO weights each trajectory by its softmax-normalized cumulative log-likelihood, adds entropy and KL penalties, and matches or slightly surpasses PPO, GRPO, and DPO on small RL and summarization tasks.