Pith. sign in

REVIEW 3 cited by

Is ChatGPT a Financial Expert? Evaluating Language Models on Financial Natural Language Processing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.12664 v1 pith:XFZ4G32L submitted 2023-10-19 cs.CL

classification cs.CL
keywords languagefinancialmodelsllmsperformancetaskschatgptdatasets
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The emergence of Large Language Models (LLMs), such as ChatGPT, has revolutionized general natural language preprocessing (NLP) tasks. However, their expertise in the financial domain lacks a comprehensive evaluation. To assess the ability of LLMs to solve financial NLP tasks, we present FinLMEval, a framework for Financial Language Model Evaluation, comprising nine datasets designed to evaluate the performance of language models. This study compares the performance of encoder-only language models and the decoder-only language models. Our findings reveal that while some decoder-only LLMs demonstrate notable performance across most financial tasks via zero-shot prompting, they generally lag behind the fine-tuned expert models, especially when dealing with proprietary datasets. We hope this study provides foundation evaluations for continuing efforts to build more advanced LLMs in the financial domain.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LaQual: An Automated Framework for LLM App Quality Evaluation

    cs.SE 2025-08 reject novelty 5.0 of 10

    LaQual automates LLM app-store quality evaluation through scenario classification, static indicator filtering, and LLM-generated dynamic metrics, with Spearman correlations of about 0.6 against human ratings.

  2. Interpretable LLMs for Credit Risk: A Systematic Review and Taxonomy

    q-fin.RM 2025-06 conditional novelty 4.0 of 10

    A systematic review and taxonomy that organizes LLM-based credit risk research by model architecture, data modality, explainability mechanism, and application domain.

  3. Assessing the Capabilities and Limitations of FinGPT Model in Financial NLP Applications

    cs.CL 2025-07 reject novelty 3.0 of 10

    FinGPT matches GPT-4 on financial sentiment and headline classification, lags on QA and NER, and shows a bullish bias in stock movement prediction.

Pith tools