No evaluated AI agent can fully match professional analysts' newness/importance/direction labels on the new 82-case Frontier Financial Judgement benchmark; GPT-5.5 tops out at 52.4%.
FinMarBa: A Market-Informed Dataset for Financial Sentiment Classification
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
This paper presents a novel hierarchical framework for portfolio optimization, integrating lightweight Large Language Models (LLMs) with Deep Reinforcement Learning (DRL) to combine sentiment signals from financial news with traditional market indicators. Our three-tier architecture employs base RL agents to process hybrid data, meta-agents to aggregate their decisions, and a super-agent to merge decisions based on market data and sentiment analysis. Evaluated on data from 2018 to 2024, after training on 2000-2017, the framework achieves a 26% annualized return and a Sharpe ratio of 1.2, outperforming equal-weighted and S&P 500 benchmarks. Key contributions include scalable cross-modal integration, a hierarchical RL structure for enhanced stability, and open-source reproducibility.
fields
cs.CL 1years
2026 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Frontier Financial Judgement: Can agents tell what might move a stock?
No evaluated AI agent can fully match professional analysts' newness/importance/direction labels on the new 82-case Frontier Financial Judgement benchmark; GPT-5.5 tops out at 52.4%.