Pith. sign in

REVIEW 2 cited by

Exploratory mean-variance portfolio selection with Choquet regularizers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.03026 v1 pith:GHSCRZWZ submitted 2023-07-06 math.OC math.PR

classification math.OCmath.PR
keywords choquetregularizersexploratorymean-varianceproblemregularizerundercontinuous-time
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we study a continuous-time exploratory mean-variance (EMV) problem under the framework of reinforcement learning (RL), and the Choquet regularizers are used to measure the level of exploration. By applying the classical Bellman principle of optimality, the Hamilton-Jacobi-Bellman equation of the EMV problem is derived and solved explicitly via maximizing statically a mean-variance constrained Choquet regularizer. In particular, the optimal distributions form a location-scale family, whose shape depends on the choices of the Choquet regularizer. We further reformulate the continuous-time Choquet-regularized EMV problem using a variant of the Choquet regularizer. Several examples are given under specific Choquet regularizers that generate broadly used exploratory samplers such as exponential, uniform and Gaussian. Finally, we design a RL algorithm to simulate and compare results under the two different forms of regularizers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A non-zero-sum game with reinforcement learning under mean-variance framework

    math.OC 2025-02 conditional novelty 6.0 of 10

    A two-agent non-zero-sum mean-variance game with Choquet-regularized exploration has a time-consistent Nash equilibrium, explicit in a Gaussian market, and a policy iteration scheme that is claimed to converge uniform...

  2. Exploratory Utility Maximization Problem with Tsallis Entropy

    cs.LG 2025-02 conditional novelty 5.0 of 10

    For CRRA utility with a wealth-scaled Tsallis entropy regularizer, the exploratory optimal strategy is Gaussian for entropy index 1 and Wigner semicircular for index 3, with the same mean as Merton's strategy, and can...

Pith tools