A 201,247-decision benchmark shows LLMs produce plausible investment logic (~4/5) but weak event grounding (0.8-2.8/5), gaps that outcome-only metrics hide.
Distilling Analysis from Generative Models for Investment Decisions
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Professionals' decisions are the focus of every field. For example, politicians' decisions will influence the future of the country, and stock analysts' decisions will impact the market. Recognizing the influential role of professionals' perspectives, inclinations, and actions in shaping decision-making processes and future trends across multiple fields, we propose three tasks for modeling these decisions in the financial market. To facilitate this, we introduce a novel dataset, A3, designed to simulate professionals' decision-making processes. While we find current models present challenges in forecasting professionals' behaviors, particularly in making trading decisions, the proposed Chain-of-Decision approach demonstrates promising improvements. It integrates an opinion-generator-in-the-loop to provide subjective analysis based on each news item, further enhancing the proposed tasks' performance.
fields
cs.AI 1years
2026 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents
A 201,247-decision benchmark shows LLMs produce plausible investment logic (~4/5) but weak event grounding (0.8-2.8/5), gaps that outcome-only metrics hide.