Pith. sign in

REVIEW 1 cited by

A New DAPO Algorithm for Stock Trading

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.06408 v2 pith:E6BYIPQH submitted 2025-05-09 cs.CE

classification cs.CE
keywords tradingdapoagentalgorithmfinancialhoursoptimizationpolicy
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in reinforcement learning, such as Dynamic Sampling Policy Optimization (DAPO), show strong performance when paired with large language models (LLMs). Motivated by this success, we ask whether similar gains can be realized in financial trading. We design a trading agent that combines an improved Group Relative Policy Optimization (GRPO) algorithm, augmented with ideas from DAPO, with LLM-based risk and sentiment signals extracted from financial news. On the NASDAQ-100 index (FNSPID dataset), our agent attains a cumulative return of 230.49 percent and an information ratio of 0.37, outperforming the CPPO-DeepSeek baseline. It also cuts training time from about 8 hours to 2.5 hours over 100 epochs while markedly reducing RAM usage. The proposed RL-LLM framework offers a scalable path toward data-efficient trading agents. Code: https://github.com/Ruijian-Zha/FinRL-DAPO-SR/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Financial Numerical Prediction and Allocation as Token Generation

    cs.CV 2026-08 conditional novelty 6.0 of 10

    FinATOM uses a causal language model's token vocabulary to output stock forecasts and ETF allocations, and shows policy optimization improves out-of-sample Sharpe.

Pith tools