Pith. sign in

REVIEW 1 cited by

Entropy-regularized Point-based Value Iteration

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.09388 v1 pith:3YSLD3Y2 submitted 2024-02-14 cs.AI

classification cs.AI
keywords entropy-regularizedinferenceobjectiveduringmodel-basedpoliciesuncertaintyhigher
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Model-based planners for partially observable problems must accommodate both model uncertainty during planning and goal uncertainty during objective inference. However, model-based planners may be brittle under these types of uncertainty because they rely on an exact model and tend to commit to a single optimal behavior. Inspired by results in the model-free setting, we propose an entropy-regularized model-based planner for partially observable problems. Entropy regularization promotes policy robustness for planning and objective inference by encouraging policies to be no more committed to a single action than necessary. We evaluate the robustness and objective inference performance of entropy-regularized policies in three problem domains. Our results show that entropy-regularized policies outperform non-entropy-regularized baselines in terms of higher expected returns under modeling errors and higher accuracy during objective inference.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Convergence of regularized agent-state-based Q-learning in POMDPs

    cs.LG 2025-08 conditional novelty 5.0 of 10

    Regularized agent-state-based Q-learning converges almost surely to the fixed point of a regularized MDP induced by the behavioral policy.

Pith tools