Pith. sign in

hub

Go-explore: a new approach for hard-exploration problems

17 Pith papers cite this work. Polarity classification is still indexing.

17 Pith papers citing it

hub tools

citation-role summary

background 2 other 1

citation-polarity summary

polarities

background 2 unclear 1

representative citing papers

TriSearch: Learning to Optimize Triangulations via Bistellar Flips

cs.LG · 2026-05-28 · unverdicted · novelty 7.0

TriSearch is an RL framework that optimizes triangulations of polytopes using bistellar flips with a circuit-supported subtriangulation action representation, generalizing zero-shot to larger instances and outperforming prior samplers in 3D and 4D.

Voyager: An Open-Ended Embodied Agent with Large Language Models

cs.AI · 2023-05-25 · unverdicted · novelty 7.0

Voyager achieves superior lifelong learning in Minecraft by combining an automatic exploration curriculum, a library of executable skills, and iterative LLM prompting with environment feedback, yielding 3.3x more unique items and 15.3x faster milestone unlocks than prior methods while generalizing技能

Cognitive Architectures for Language Agents

cs.AI · 2023-09-05 · accept · novelty 6.0

CoALA is a modular cognitive architecture for language agents that organizes memory components, action spaces for internal and external interaction, and a generalized decision-making loop to support more systematic development of capable agents.

Mesh-RL: Coupled subgrid reinforcement learning

cs.LG · 2026-06-24 · unverdicted · novelty 5.0

Mesh-RL applies finite-element-inspired domain decomposition with overlapping subgrids to accelerate temporal-difference learning across distant states in grid-world environments.

PokeRL: Reinforcement Learning for Pokemon Red

cs.LG · 2026-04-12 · unverdicted · novelty 5.0

PokeRL trains PPO agents to finish early Pokemon Red tasks using a loop-aware environment wrapper, multi-layer anti-loop mechanisms, and dense hierarchical rewards.

Optimistic Proximal Policy Optimization

cs.LG · 2019-06-25 · unverdicted · novelty 4.0

OPPO augments PPO with optimistic policy evaluation driven by return uncertainty estimates and shows improved results over prior methods on a tabular sparse-reward task.

citing papers explorer

Showing 17 of 17 citing papers.