An evaluation of eight LLMs on a hand-built risk-tolerance rubric finds large scoring errors and small demographic differences, but the rubric given to the models contradicts the rubric used as ground truth.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CL 1years
2025 1verdicts
REJECT 1representative citing papers
citing papers explorer
-
Evaluating AI for Finance: Is AI Credible at Assessing Investment Risk?
An evaluation of eight LLMs on a hand-built risk-tolerance rubric finds large scoring errors and small demographic differences, but the rubric given to the models contradicts the rubric used as ground truth.