Pith. sign in

REVIEW 1 cited by

MAGIC: Learning Macro-Actions for Online POMDP Planning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2011.03813 v4 pith:YIFOLTOH submitted 2020-11-07 cs.RO cs.AIcs.LG

classification cs.ROcs.AIcs.LG
keywords planningmacro-actionsonlinemagicpomdprobotcomputationaldecision
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The partially observable Markov decision process (POMDP) is a principled general framework for robot decision making under uncertainty, but POMDP planning suffers from high computational complexity, when long-term planning is required. While temporally-extended macro-actions help to cut down the effective planning horizon and significantly improve computational efficiency, how do we acquire good macro-actions? This paper proposes Macro-Action Generator-Critic (MAGIC), which performs offline learning of macro-actions optimized for online POMDP planning. Specifically, MAGIC learns a macro-action generator end-to-end, using an online planner's performance as the feedback. During online planning, the generator generates on the fly situation-aware macro-actions conditioned on the robot's belief and the environment context. We evaluated MAGIC on several long-horizon planning tasks both in simulation and on a real robot. The experimental results show that the learned macro-actions offer significant benefits in online planning performance, compared with primitive actions and handcrafted macro-actions.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Meta-learning how to Share Credit among Macro-Actions

    cs.LG 2025-06 conditional novelty 6.0 of 10

    MASP meta-learns a similarity matrix over macro-actions and regularizes Q-values so that similar actions move together, improving exploration and performance in augmented-action-space RL.

Pith tools