Pith. sign in

REVIEW 2 cited by

Discrete-Time Mean-Variance Strategy Based on Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.15385 v1 pith:NNVN5F32 submitted 2023-12-24 q-fin.MF cs.LGq-fin.PM

classification q-fin.MFcs.LGq-fin.PM
keywords discrete-timemodellearningreinforcementcontinuous-timemean-variancestrategyadditionally
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper studies a discrete-time mean-variance model based on reinforcement learning. Compared with its continuous-time counterpart in \cite{zhou2020mv}, the discrete-time model makes more general assumptions about the asset's return distribution. Using entropy to measure the cost of exploration, we derive the optimal investment strategy, whose density function is also Gaussian type. Additionally, we design the corresponding reinforcement learning algorithm. Both simulation experiments and empirical analysis indicate that our discrete-time model exhibits better applicability when analyzing real-world data than the continuous-time model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-period Asset-liability Management with Reinforcement Learning in a Regime-Switching Market

    math.OC 2025-09 reject novelty 4.0 of 10

    The paper derives and tests an RL-based mean-variance strategy for multi-period asset-liability management with hidden bull/bear regimes, but its filtering step is not valid.

  2. Reinforcement Learning for a Discrete-Time Linear-Quadratic Control Problem with an Application

    stat.ML 2024-12 reject novelty 3.0 of 10

    The paper claims entropy regularization forces the optimal LQ feedback policy to be Gaussian and uses that to solve a mean-variance asset-liability problem, but the proof of the main theorem contains a correlation err...

Pith tools