Pith. sign in

REVIEW 10 cited by

A Survey of Large Language Models in Finance (FinLLMs)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.02315 v1 pith:Y7DIN7FD submitted 2024-02-04 cs.CL q-fin.GN

A Survey of Large Language Models in Finance (FinLLMs)

classification cs.CL q-fin.GN
keywords finllmsfinancialincludinglanguagedatasetsfinancellmsmodels
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Large Language Models (LLMs) have shown remarkable capabilities across a wide variety of Natural Language Processing (NLP) tasks and have attracted attention from multiple domains, including financial services. Despite the extensive research into general-domain LLMs, and their immense potential in finance, Financial LLM (FinLLM) research remains limited. This survey provides a comprehensive overview of FinLLMs, including their history, techniques, performance, and opportunities and challenges. Firstly, we present a chronological overview of general-domain Pre-trained Language Models (PLMs) through to current FinLLMs, including the GPT-series, selected open-source LLMs, and financial LMs. Secondly, we compare five techniques used across financial PLMs and FinLLMs, including training methods, training data, and fine-tuning methods. Thirdly, we summarize the performance evaluations of six benchmark tasks and datasets. In addition, we provide eight advanced financial NLP tasks and datasets for developing more sophisticated FinLLMs. Finally, we discuss the opportunities and the challenges facing FinLLMs, such as hallucination, privacy, and efficiency. To support AI research in finance, we compile a collection of accessible datasets and evaluation benchmarks on GitHub.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. What LLMs explain is not what they believe: Evaluating explanation sufficiency under models' own input beliefs

    cs.LG 2026-06 unverdicted novelty 6.0

    Proposes SCSuff metric for evaluating LLM explanation sufficiency via model-generated alternative inputs, showing explanations are typically insufficient and predictable from hidden states.

  2. MetaPS: Adaptive Programmatic Strategy Selection for Market Agents

    cs.AI 2026-06 unverdicted novelty 6.0

    MetaPS trains models via simulation rollouts to select from programmatic strategy libraries for market agents, yielding better performance than fixed or direct LLM baselines across model sizes.

  3. Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution

    cs.AI 2026-05 unverdicted novelty 6.0

    Causality provides a unifying framework for resolving trade-offs in trustworthy AI by managing invariance conflicts under changes to the data-generating process.

  4. QRAFTI: An Agentic Framework for Empirical Research in Quantitative Finance

    cs.MA 2026-04 unverdicted novelty 6.0

    QRAFTI is a multi-agent framework using tool-calling and reflection-based planning to emulate quant research tasks like factor replication and signal testing on financial data.

  5. Explainable AML Triage with LLMs: Evidence Retrieval and Counterfactual Checks

    cs.AI 2026-03 unverdicted novelty 6.0

    Evidence-grounded LLM triage with structured contracts and counterfactual validation achieves PR-AUC 0.75 and high faithfulness scores on synthetic AML benchmarks.

  6. TriAgent: Divergence-Aware Multi-Agent Committees for Cost-Efficient Financial Sentiment Analysis

    cs.CL 2026-07 conditional novelty 5.0

    A divergence-routed VADER+FinBERT+LLM committee reaches ~0.87 F1 with a 1.5B critic, matching 7B with far less cost, while same-size persona voting regresses to 0.66.

  7. Explainable AML Triage with LLMs: Evidence Retrieval and Counterfactual Checks

    cs.AI 2026-03 unverdicted novelty 5.0

    Evidence-bundled RAG plus citation contracts and counterfactual checks improves synthetic AML triage (PR-AUC 0.75) while raising citation validity and decision faithfulness.

  8. MadEvolve: Evolutionary Optimization of Trading Systems with Large Language Models

    q-fin.TR 2026-05 unverdicted novelty 4.0

    MadEvolve uses LLMs for evolutionary optimization of trading strategies and reports significant backtest improvements on Bitcoin tasks including signal feature evolution and joint strategy optimization.

  9. Trustworthy AI Suffers from Invariance Conflicts and Causality is The Solution

    cs.AI 2026-05 unverdicted novelty 4.0

    Causality resolves trade-offs in trustworthy AI by treating them as invariance conflicts under different data-generating process changes.

  10. Beyond Agent Architecture: Execution Assumptions and Reproducibility in LLM-Based Trading Systems

    cs.AI 2026-06 unverdicted novelty 3.0

    Reproducibility audit of 30 LLM trading papers shows execution assumptions under-reported relative to agent architectures, illustrated by a 10-equity example where frictions compress returns.