A 201,247-decision benchmark shows LLMs produce plausible investment logic (~4/5) but weak event grounding (0.8-2.8/5), gaps that outcome-only metrics hide.
Advances in Neural Information Processing Systems 38 (NeurIPS 2024) , year=
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.AI 1years
2026 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents
A 201,247-decision benchmark shows LLMs produce plausible investment logic (~4/5) but weak event grounding (0.8-2.8/5), gaps that outcome-only metrics hide.