Pith. sign in

REVIEW 4 major objections 4 minor 85 references

Interpretable LLMs for Credit Risk: A Systematic Review and Taxonomy

T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper presents the first systematic review and taxonomy of large-language-model approaches to credit risk assessment, built from 60 studies published between 2020 and 2025.

desk verdict Useful taxonomy, but the 'first systematic review of 60 papers' claim doesn't survive contact with Table 2; fix the corpus and this could be a reference map. read the letter →

arxiv 2506.04290 v2 pith:CXS5KZ4V submitted 2025-06-04 q-fin.RM cs.LG

classification q-fin.RMcs.LG
keywords largelanguagemodelscreditriskassessmentsystematicreviewtaxonomyinterpretabilityexplainableAIfinancialNLPPRISMA
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish that LLM-based credit risk assessment is now a recognizable research field with a shape that can be catalogued. It claims to be the first systematic review and taxonomy of that field, built from 60 studies published between 2020 and 2025 and selected through a documented literature-screening flow. The central contribution is a four-part classification: model architectures, data modalities, interpretability mechanisms, and application areas. If the classification holds, researchers and financial institutions gain a common vocabulary for comparing LLM credit-scoring systems and for seeing where work is missing.

What carries the argument

The machinery is the four-axis taxonomy itself. Each collected study is slotted into model architecture, data modality, interpretability mechanism, and application domain, with the PRISMA flowchart, a staged, documented procedure for screening a literature search, used to make the selection repeatable. The taxonomy does the argumentative work: it converts 60 heterogeneous papers into comparable categories, which is what lets the paper claim to be the first reference classification and to list gaps as structural absences rather than individual complaints.

What would settle it

Inspect the 60-entry table and resolve every numbered reference to a distinct, retrievable publication. If any paper appears twice under different numbers, or if a reference resolves to the wrong study, the claimed corpus size and the taxonomy's proportions change; if all 60 entries resolve cleanly to distinct papers, the descriptive claims of the review stand.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that the scattered literature on LLMs in credit risk can be organized into a single taxonomy, and that doing so reveals the field's center of gravity: encoder-only and decoder-only transformers plus domain-specific financial LLMs dominate, with hybrid pipelines and parameter-efficient tuning rising; SHAP and LIME post-hoc explanations still dominate interpretability while chain-of-thought prompting and intrinsically transparent designs are growing; and applications concentrate on retail and SME scoring and news-sentiment signals. The paper further claims that this map exposes structural gaps, including reproducibility, bias, hallucination, efficiency, and missing benchmarks, that should drive the next round of research.

Load-bearing premise

The load-bearing premise is that the 60 studies selected through the described screening flow are distinct, correctly referenced, and representative of the LLM credit risk literature from 2020 to 2025, because every category count and every identified gap in the review is computed from that corpus.

Editorial extensions

If this is right

  • LLMs can score credit from unstructured text such as loan descriptions, news, and analyst reports, not only from financial ratios and payment histories.
  • Post-hoc tools such as SHAP and LIME remain the dominant way to explain LLM credit decisions, but chain-of-thought prompting and intrinsically transparent models are emerging alternatives.
  • A reliable LLM credit-scoring pipeline should be evaluated not just on accuracy but also on fairness, hallucination resistance, reproducibility, and inference cost.
  • Hybrid and retrieval-augmented pipelines, parameter-efficient fine-tuning, and multimodal inputs are the directions the taxonomy identifies as the field's frontier.
  • Standardized benchmarks for LLM credit risk do not yet exist, which makes direct comparison across studies difficult.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same four-axis grid could be carried into neighboring regulated decisions such as insurance underwriting or loan pricing, where the interpretability-bias-reproducibility tensions recur.
  • The dominance of post-hoc explainability suggests a testable expectation: studies using chain-of-thought or intrinsically transparent models should show a smaller gap between the explanation and the true decision logic, but the paper does not measure that gap.
  • Because the corpus mixes peer-reviewed articles, preprints, and working papers, the category proportions should be treated as sensitive to inclusion criteria; re-running the same search protocol at a later date would likely shift counts.
  • The paper notes behavioral and external signals only as a gap, which implies a plausible trajectory: credit scoring may move from static snapshot models toward continuous monitoring as those signals are integrated.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The manuscript presents a systematic review and taxonomy of LLM-based credit risk assessment. It reports a PRISMA-guided selection of 60 papers published between 2020 and 2025 and organizes them along four dimensions: model architectures, data modalities, interpretability mechanisms, and application domains. It also compares 22 existing surveys, answers five research questions, and lists research gaps and future directions. The paper's central claim is that this is the first systematic review and taxonomy focused on LLM-based credit risk with interpretability as a core axis.

Significance. If the 60-paper corpus were complete, deduplicated, and correctly referenced, the paper would be a useful reference: the four-dimensional taxonomy is clear, the survey comparison in Table 1 helps position the work, and the gap analysis points to reasonable research directions. The manuscript also deserves credit for making its corpus enumerable in Table 2 and for stating explicit research questions. However, the current evidence does not support the claimed systematic foundation: duplicate entries and an unresolvable citation in Table 2 invalidate the '60 distinct peer-reviewed studies' premise, and the absence of a coding or extraction table weakens prevalence claims in Section 4 and the Conclusions. The stress-test concern about corpus auditability lands; the central claim is not supported as written, but the problems appear fixable within the scope of a revision.

major comments (4)
  1. [Table 2] Table 2, rows 6 and 32: both entries are the same Pixiu paper (refs [33] and [59] share the same title and arXiv:2306.05443), and rows 7 and 37 are the same Babaei and Giudici paper (refs [34] and [64], same journal, volume, and article number). Rows 11 and 17 share the title 'Optimizing large language models for financial risk assessment in credit unions' with different author attributions, which is at least a strong duplicate signal. Because the central claim is a systematic review of 60 distinct papers, these duplicates mean the corpus count and the PRISMA flow numbers in Figure 2 are not valid as reported. The authors must deduplicate the list, re-count from the 182 initial records, and either restrict the corpus to peer-reviewed publications or drop the 'peer-reviewed' qualifier, since many entries are arXiv or SSRN preprints (e.g., [33], [38], [39], [82]).
  2. [Table 2, row 10] Row 10 lists 'Mehedi Hasan et al. [37]', but reference [37] is Yuyan Chen et al., 'Hallucination detection', so the table entry cannot be mapped to the bibliography. Row 54 lists 'Chafekar et al. [81]' with venue 'Unspecified (likely arXiv or workshop preprint) 2024', while reference [81] is an ICON 2023 conference paper. Every row of the overview table must correspond to exactly one resolvable bibliographic entry; otherwise the extraction underlying the taxonomy cannot be audited.
  3. [Section 3.1 and Figure 2] The PRISMA flow diagram is not auditable. The inclusion and exclusion criteria are not stated; the text only says papers were removed 'with the joint efforts of both authors ... according to the inclusion criteria', and Figure 2 reports no count of records excluded at the identification or screening stage. The transition from 51 papers to 60 by snowballing is also not documented. For a systematic review, the authors should report keyword-search dates, databases searched per query, eligibility criteria, screening decisions, and a list of excluded studies, or else scale back the PRISMA claim.
  4. [Section 4.3 and Conclusions] Claims that post-hoc methods such as SHAP and LIME are 'the most commonly used' and that interpretability techniques are 'prevalent' are not supported by any extraction table or frequency count. The four taxonomy subsections cite illustrative papers rather than a systematic coding of all 60 studies. Adding a per-paper coding matrix (papers by architecture, data type, interpretability mechanism, application, and evaluation metric) with summary counts is necessary to substantiate RQ3 and the prevalence statements in the conclusions.
minor comments (4)
  1. [Conclusions] The text 'SHAP and LIMA (post-hoc) are prevalent interpretability techniques' should read 'SHAP and LIME'; LIME is the method discussed in Section 4.3.
  2. [Figure 2] The arrow from 51 to 60 is attributed to the snowball method in the text, but the diagram does not show the number of records added; label this explicitly and reconcile the counts.
  3. [Section 2] There are multiple grammatical issues, e.g., 'In Joshi et al. study's AI frameworks in credit risk and trading applications are examined' and 'In [7], models such as FinGPT and BloombergGPT are discussed, ... but ignore credit risk applications'; a careful proofreading pass is needed.
  4. [References] Reference [6] lacks publication venue and year, and reference [84] appears to duplicate reference [28] as an SSRN version; clarify whether these are distinct versions or separate records.

Circularity Check

0 steps flagged · score 1.0 of 10

No circular derivation; the sole self-citation is background only, and the 60-paper corpus concerns are correctness/reproducibility issues, not circularity.

full rationale

This paper is a systematic review and taxonomy; it fits no model parameters, derives no quantitative prediction from inputs, and its contribution is an external organization of other studies. The only self-citation is reference [4] (Golec et al., a prior survey on LLM-driven APT detection for 6G), used in Section 1 merely to illustrate that LLMs 'have great potential in financial applications' with their 'high performance in extracting meaning from free text'. That claim is background and is not load-bearing for the taxonomy, the PRISMA procedure, or the four-part classification. The 'first systematic review' claim is supported by comparison with 22 prior surveys in Table 1, not by the authors' own prior results. The skeptic's identified issues—duplicate entries in Table 2 (rows 6/32, 7/37, and the strong title match of rows 11/17), an unresolvable row 10 vs. reference [37], and the venue mismatch for row 54—would weaken the auditability of the claimed 60-paper corpus, but they undermine correctness and reproducibility, not circularity. No quantity is fitted, renamed, or defined in terms of itself. Score 1 reflects the single non-load-bearing self-citation; there is no circular derivation chain.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The review contains no fitted parameters and no invented entities. Its load-bearing premises are methodological: the PRISMA flow produces a complete and deduplicated corpus, the four taxonomy dimensions are a valid way to divide the literature, and the keyword set captures relevant studies. The first premise is contradicted by the duplicate entries in Table 2.

assumptions (3)
  • domain assumption The PRISMA protocol is an appropriate and sufficient method for selecting and reviewing this literature.
    The paper adopts PRISMA in Section 3.1, assuming the resulting corpus is representative of the field.
  • domain assumption The four taxonomy dimensions (model architecture, data modality, explainability, application domain) are a complete and non-overlapping set of categories for LLM credit risk research.
    The taxonomy is constructed with these four headings in Section 4 without an empirical derivation of category completeness.
  • domain assumption The selected keywords in Section 3.1 capture all relevant LLM credit risk studies published 2020-2025.
    The search hinges on five keyword groups; no validation against a gold-standard set of known papers is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Interpretable LLMs for Credit Risk: A Systematic Review and Taxonomy." pith.science (2026). https://pith.science/paper/CXS5KZ4V

@misc{pith2026250604290,
  author       = {Pith},
  title        = {Pith review of: Interpretable LLMs for Credit Risk: A Systematic Review and Taxonomy},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/CXS5KZ4V}},
  note         = {Machine review of arXiv:2506.04290}
}
read the original abstract

Large Language Models (LLM), which have developed in recent years, enable credit risk assessment through the analysis of financial texts such as analyst reports and corporate disclosures. This paper presents the first systematic review and taxonomy focusing on LLMbased approaches in credit risk estimation. We determined the basic model architectures by selecting 60 relevant papers published between 2020-2025 with the PRISMA research strategy. And we examined the data used for scenarios such as credit default prediction and risk analysis. Since the main focus of the paper is interpretability, we classify concepts such as explainability mechanisms, chain of thought prompts and natural language justifications for LLM-based credit models. The taxonomy organizes the literature under four main headings: model architectures, data types, explainability mechanisms and application areas. Based on this analysis, we highlight the main future trends and research gaps for LLM-based credit scoring systems. This paper aims to be a reference paper for artificial intelligence and financial researchers.

Figures

Figures reproduced from arXiv: 2506.04290 by the authors.

Figure 1
Figure 1. shows the organization of the paper. Section 2 examines the current survey studies in the field related finance and LLM , examining their focus and limitations. Then, a comparison analysis is performed by comparing this paper and the literature. Section 3 explains the research questions and article collection strategy. Section 4 provides a taxonomy by classifying LLM￾oriented credit risk assessment systems in four m… view at source ↗
Figure 2
Figure 2. PRISMA flow diagram for the selection of studies on LLM-based credit risk assess [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Taxonomy of Transformer-Based Model Architectures for LLM-Driven Credit Risk [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Taxonomy of Data Modalities Utilized in LLM-Based Credit Risk Systems. [PITH_FULL_IMAGE:figures/full_fig_p010_4.png]
Figure 5
Figure 5. Figure 5: Taxonomy of Interpretability Mechanisms for LLM-Based Credit Models. [PITH_FULL_IMAGE:figures/full_fig_p011_5.png]
Figure 6
Figure 6. Figure 6: Taxonomy of Application Domains for LLMs in Credit Risk Assessment. [PITH_FULL_IMAGE:figures/full_fig_p012_6.png]
Figure 7
Figure 7. Figure 7: Taxonomy of Research Gaps and Future Directions in LLM-Driven Credit Risk Re [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

85 extracted references · 62 canonical work pages

  1. [59]

    Pixiu: A large language model, instruction data and evaluation benchmark for finance.arXiv preprint arXiv:2306.05443, 2023

    Qianqian Xie, Weiguang Han, Xiao Zhang, Yanzhao Lai, Min Peng, Alejandro Lopez- Lira, and Jimin Huang. Pixiu: A large language model, instruction data and evaluation benchmark for finance.arXiv preprint arXiv:2306.05443, 2023

  2. [64]

    Gpt classifications, with application to credit lending

    Golnoosh Babaei and Paolo Giudici. Gpt classifications, with application to credit lending. Machine Learning with Applications, 16:100534, 2024

  3. [38]

    Optimizing large language models for financial risk assessment in credit unions

    Charlie Luca. Optimizing large language models for financial risk assessment in credit unions. 2024

  4. [44]

    Optimizing large language models for financial risk assessment in credit unions

    AYOMIDE JOEL. Optimizing large language models for financial risk assessment in credit unions

  5. [37]

    Hallucination detection: Robustly discerning reliable answers in large language models

    Yuyan Chen, Qiang Fu, Yichen Yuan, Zhihao Wen, Ge Fan, Dayiheng Liu, Dongmei Zhang, Zhixu Li, and Yanghua Xiao. Hallucination detection: Robustly discerning reliable answers in large language models. InProceedings of the 32nd ACM International Conference on Information and Knowledge Management, pages 245–255, 2023

  6. [39]

    Large language model adaptation for financial sentiment analysis.arXiv preprint arXiv:2401.14777, 2024

    Pau Rodriguez Inserte, Mariam Nakhl´ e, Raheel Qader, Gaetan Caillaut, and Jingshu Liu. Large language model adaptation for financial sentiment analysis.arXiv preprint arXiv:2401.14777, 2024

  7. [82]

    Gpt-lgbm: A chatgpt-based integrated framework for credit scoring with textual and structured data.Available at SSRN 4671511, 2023

    Li Yu, Xuefei Bai, and Zhiwei Chen. Gpt-lgbm: A chatgpt-based integrated framework for credit scoring with textual and structured data.Available at SSRN 4671511, 2023

  8. [81]

    Understanding behaviour of large language models for short-term and long-term fairness scenarios

    Chafekar Talha, Hussain Aafiya, and Cheong Chon In. Understanding behaviour of large language models for short-term and long-term fairness scenarios. InProceedings of the 20th International Conference on Natural Language Processing (ICON), pages 52–61, 2023

Show all 85 references
  1. [1]

    Game theory analysis on credit risk assessment in e-commerce.Information Processing & Management, 59(1):102763, 2022

    Zhang Nana, Wei Xiujian, and Zhang Zhongqiu. Game theory analysis on credit risk assessment in e-commerce.Information Processing & Management, 59(1):102763, 2022

  2. [2]

    Data-driven decision-making in credit risk management: The information value of analyst reports.Decision Support Systems, 158:113770, 2022

    Jan Roeder, Matthias Palmer, and Jan Muntermann. Data-driven decision-making in credit risk management: The information value of analyst reports.Decision Support Systems, 158:113770, 2022

  3. [3]

    Comparative investigation of gpt and finbert’s sentiment analysis performance in news across different sectors.Electronics, 14(6):1090, 2025

    Ji-Won Kang and Sun-Yong Choi. Comparative investigation of gpt and finbert’s sentiment analysis performance in news across different sectors.Electronics, 14(6):1090, 2025

  4. [4]

    Llm-driven apt detection for 6g wireless networks: A systematic review and taxonomy

    Muhammed Golec, Yaser Khamayseh, Suhib Bani Melhem, and Abdulmalik Alwarafy. Llm-driven apt detection for 6g wireless networks: A systematic review and taxonomy. arXiv preprint arXiv:2505.18846, 2025

  5. [5]

    Innovative sentiment analysis and prediction of stock price using finbert, gpt-4 and logistic regression: A data-driven approach.Big Data and Cognitive Computing, 8(11):143, 2024

    Olamilekan Shobayo, Sidikat Adeyemi-Longe, Olusogo Popoola, and Bayode Ogunleye. Innovative sentiment analysis and prediction of stock price using finbert, gpt-4 and logistic regression: A data-driven approach.Big Data and Cognitive Computing, 8(11):143, 2024. 14

  6. [6]

    A review on large language models and generative ai in banking

    Daniel Staegemann, Christian Haertel, Christian Daase, Matthias Pohl, Mohammad Ab- dallah, and Klaus Turowski. A review on large language models and generative ai in banking

  7. [7]

    Luke Lee. Enhancing financial inclusion and regulatory challenges: A critical analysis of digital banks and alternative lenders through digital platforms, machine learning, and large language models integration.arXiv preprint arXiv:2404.11898, 2024

  8. [8]

    Satyadhar Joshi. Gen ai for market risk and credit risk learn agentically powered gen ai; gen ai agentic framework for financial risk management.Gen AI Agentic Framework for Financial Risk Management (January 15, 2025), 2025

  9. [9]

    A survey on large language models for critical societal domains: Finance, healthcare, and law.arXiv preprint arXiv:2405.01769, 2024

    Zhiyu Zoey Chen, Jing Ma, Xinlu Zhang, Nan Hao, An Yan, Armineh Nourbakhsh, Xi- anjun Yang, Julian McAuley, Linda Petzold, and William Yang Wang. A survey on large language models for critical societal domains: Finance, healthcare, and law.arXiv preprint arXiv:2405.01769, 2024

  10. [10]

    A systematic review of sentiment analytics in banking headlines.Decision Analytics Jour- nal, page 100584, 2025

    Muhunthan Jayanthakumaran, Nagesh Shukla, Biswajeet Pradhan, and Ghassan Beydoun. A systematic review of sentiment analytics in banking headlines.Decision Analytics Jour- nal, page 100584, 2025

  11. [11]

    A survey of large language models for financial applications: Progress, prospects and challenges.arXiv preprint arXiv:2406.11903, 2024

    Yuqi Nie, Yaxuan Kong, Xiaowen Dong, John M Mulvey, H Vincent Poor, Qingsong Wen, and Stefan Zohren. A survey of large language models for financial applications: Progress, prospects and challenges.arXiv preprint arXiv:2406.11903, 2024

  12. [12]

    Machine learning for identifying risk in financial statements: A survey.ACM Computing Surveys, 57(9):1–37, 2025

    Elias Zavitsanos, Eirini Spyropoulou, George Giannakopoulos, and Georgios Paliouras. Machine learning for identifying risk in financial statements: A survey.ACM Computing Surveys, 57(9):1–37, 2025

  13. [13]

    Layla Abdel-Rahman Aziz and Yuli Andriansyah. The role artificial intelligence in mod- ern banking: an exploration of ai-driven approaches for enhanced fraud prevention, risk management, and regulatory compliance.Reviews of Contemporary Business Analytics, 6(1):110–132, 2023

  14. [14]

    A comprehensive review of gen ai agents: Applications and frameworks in finance, investments and risk domains

    Satyadhar Joshi. A comprehensive review of gen ai agents: Applications and frameworks in finance, investments and risk domains. 2025

  15. [15]

    Large language models for financial and investment management: Applications and benchmarks.Journal of Portfolio Management, 51(2), 2024

    Yaxuan Kong, Yuqi Nie, Xiaowen Dong, John M Mulvey, H Vincent Poor, Qingsong Wen, and Stefan Zohren. Large language models for financial and investment management: Applications and benchmarks.Journal of Portfolio Management, 51(2), 2024

  16. [16]

    Review of gen ai models for financial risk management: Architectural frameworks and implementation strategies.Available at SSRN 5239190, 2025

    Satyadhar Joshi. Review of gen ai models for financial risk management: Architectural frameworks and implementation strategies.Available at SSRN 5239190, 2025

  17. [17]

    Large language models in finance (finllms).Neural Computing and Applications, pages 1–15, 2025

    Jean Lee, Nicholas Stevens, and Soyeon Caren Han. Large language models in finance (finllms).Neural Computing and Applications, pages 1–15, 2025

  18. [18]

    AlSaleh, and Amir Mazhar

    Farhina Sardar Khan, Syed Shahid Mazhar, Kashif Mazhar, Dhoha A. AlSaleh, and Amir Mazhar. Model-agnostic explainable artificial intelligence methods in finance: a system- atic review, recent developments, limitations, challenges and future directions.Artificial Intelligence R...

  19. [19]

    The impact of big data characteristics on credit risk assessment.International Journal of Data Science and Analytics, pages 1–21, 2025

    Amin Karami and Chukwuemeka Igbokwe. The impact of big data characteristics on credit risk assessment.International Journal of Data Science and Analytics, pages 1–21, 2025

  20. [20]

    Large language models in finance: Reasoning.Large Language Models in Finance: Reasoning (December 08, 2024), 2024

    Noguer I Alonso et al. Large language models in finance: Reasoning.Large Language Models in Finance: Reasoning (December 08, 2024), 2024. 15

  21. [21]

    Large language models in finance: A survey

    Yinheng Li, Shaofei Wang, Han Ding, and Hang Chen. Large language models in finance: A survey. InProceedings of the fourth ACM international conference on AI in finance, pages 374–382, 2023

  22. [22]

    Can llms be good financial advisors?: An initial study in personal decision making for optimized outcomes.arXiv preprint arXiv:2307.07422, 2023

    Kausik Lakkaraju, Sai Krishna Revanth Vuruma, Vishal Pallagani, Bharath Muppasani, and Biplav Srivastava. Can llms be good financial advisors?: An initial study in personal decision making for optimized outcomes.arXiv preprint arXiv:2307.07422, 2023

  23. [23]

    Comparative analysis of large language models adapt- ability to detect sentiments in financial domain

    Steve Samson and Jitendra Rout. Comparative analysis of large language models adapt- ability to detect sentiments in financial domain

  24. [24]

    Large language models (llms) for financial security

    Adetoyese Omoseebi, Akorede Jhon, and Winner Olabiyi. Large language models (llms) for financial security. 2024

  25. [25]

    Assessing consistency and reproducibility in the outputs of large language models: Evidence across diverse finance and accounting tasks.arXiv e-prints, pages arXiv–2503, 2025

    Julian Junyan Wang and Victor Xiaoqi Wang. Assessing consistency and reproducibility in the outputs of large language models: Evidence across diverse finance and accounting tasks.arXiv e-prints, pages arXiv–2503, 2025

  26. [26]

    Large language models and generative ai in finance: an analysis of chatgpt, bard, and bing ai.Bard, and Bing AI (July 15, 2023), 2023

    David Krause. Large language models and generative ai in finance: an analysis of chatgpt, bard, and bing ai.Bard, and Bing AI (July 15, 2023), 2023

  27. [27]

    Genai and llm for financial institutions: A corporate strategic survey.Available at SSRN 4988118, 2024

    Jun Xu. Genai and llm for financial institutions: A corporate strategic survey.Available at SSRN 4988118, 2024

  28. [28]

    Credit risk meets large language models: Building a risk indicator from loan descriptions in p2p lending.arXiv preprint arXiv:2401.16458, 2024

    Mario Sanz-Guerrero and Javier Arroyo. Credit risk meets large language models: Building a risk indicator from loan descriptions in p2p lending.arXiv preprint arXiv:2401.16458, 2024

  29. [29]

    Nlp-based application for analyzing private and public banks stocks reaction to news events in the indian stock exchange.Systems, 10(6):233, 2022

    Varun Dogra, Fahd S Alharithi, Roberto Marcelo ´Alvarez, Aman Singh, and Abdulrah- man M Qahtani. Nlp-based application for analyzing private and public banks stocks reaction to news events in the indian stock exchange.Systems, 10(6):233, 2022

  30. [30]

    Extracting structured insights from financial news: An augmented llm driven approach.arXiv preprint arXiv:2407.15788, 2024

    Rian Dolphin, Joe Dursun, Jonathan Chow, Jarrett Blankenship, Katie Adams, and Quin- ton Pike. Extracting structured insights from financial news: An augmented llm driven approach.arXiv preprint arXiv:2407.15788, 2024

  31. [31]

    Overcoming data limitations in credit risk assessment with fingpt-generated synthetic data

    Yuanxi Cai. Overcoming data limitations in credit risk assessment with fingpt-generated synthetic data. In2024 International Conference on Electronics and Devices, Computa- tional Science (ICEDCS), pages 137–141. IEEE, 2024

  32. [32]

    Explainable transformers in financial forecasting

    Vasanthi Govindaraj, Humashankar Vellathur Jaganathan, and P Prakash. Explainable transformers in financial forecasting. 2023

  33. [35]

    Making llms worth every penny: Resource-limited text classification in banking

    Lefteris Loukas, Ilias Stogiannidis, Odysseas Diamantopoulos, Prodromos Malakasiotis, and Stavros Vassos. Making llms worth every penny: Resource-limited text classification in banking. InProceedings of the Fourth ACM International Conference on AI in Finance, pages 392–400, 2023

  34. [36]

    Leveraging xai in prompt-based chatgpt for financial decision support

    Zhengping Liu, Hailing He, Lieping Zhang, and Chen Peng. Leveraging xai in prompt-based chatgpt for financial decision support. InProceedings of the International Conference on Digital Economy, Blockchain and Artificial Intelligence, pages 255–259, 2024. 16

  35. [40]

    Generating plausible counterfactual explanations for deep transformers in financial text classification.arXiv preprint arXiv:2010.12512, 2020

    Linyi Yang, Eoin M Kenny, Tin Lok James Ng, Yi Yang, Barry Smyth, and Ruihai Dong. Generating plausible counterfactual explanations for deep transformers in financial text classification.arXiv preprint arXiv:2010.12512, 2020

  36. [41]

    Deep learning with llm: A new paradigm for financial market prediction and analysis

    Jiarui Rao and Qian Zhang. Deep learning with llm: A new paradigm for financial market prediction and analysis. 2025

  37. [42]

    Design and implementation of an llm system to improve response time for smes technology credit evaluation.International journal of advanced smart convergence, 12(3):51–60, 2023

    Sungwook Yoon. Design and implementation of an llm system to improve response time for smes technology credit evaluation.International journal of advanced smart convergence, 12(3):51–60, 2023

  38. [43]

    Empowering many, biasing a few: Generalist credit scoring through large language models.arXiv preprint arXiv:2310.00566, 2023

    Duanyu Feng, Yongfu Dai, Jimin Huang, Yifang Zhang, Qianqian Xie, Weiguang Han, Zhengyu Chen, Alejandro Lopez-Lira, and Hao Wang. Empowering many, biasing a few: Generalist credit scoring through large language models.arXiv preprint arXiv:2310.00566, 2023

  39. [45]

    Pre- dicting liquidity-aware bond yields using causal gans and deep reinforcement learning with llm evaluation.arXiv preprint arXiv:2502.17011, 2025

    Jaskaran Singh Walia, Aarush Sinha, Srinitish Srinivasan, and Srihari Unnikrishnan. Pre- dicting liquidity-aware bond yields using causal gans and deep reinforcement learning with llm evaluation.arXiv preprint arXiv:2502.17011, 2025

  40. [46]

    Explainable risk classification in financial reports.arXiv preprint arXiv:2405.01881, 2024

    Xue Wen Tan and Stanley Kok. Explainable risk classification in financial reports.arXiv preprint arXiv:2405.01881, 2024

  41. [47]

    Investigating the beneficial impact of segmentation-based modelling for credit scoring.Decision Support Systems, 179:114170, 2024

    Khaoula Idbenjra, Kristof Coussement, and Arno De Caigny. Investigating the beneficial impact of segmentation-based modelling for credit scoring.Decision Support Systems, 179:114170, 2024

  42. [48]

    Instruct-fingpt: Financial sentiment analysis by instruction tuning of general-purpose large language models.arXiv preprint arXiv:2306.12659, 2023

    Boyu Zhang, Hongyang Yang, and Xiao-Yang Liu. Instruct-fingpt: Financial sentiment analysis by instruction tuning of general-purpose large language models.arXiv preprint arXiv:2306.12659, 2023

  43. [49]

    Arafinnlp 2024: The first arabic financial nlp shared task.arXiv preprint arXiv:2407.09818, 2024

    Sanad Malaysha, Mo El-Haj, Saad Ezzini, Mohammed Khalilia, Mustafa Jarrar, Sultan Almujaiwel, Ismail Berrada, and Houda Bouamor. Arafinnlp 2024: The first arabic financial nlp shared task.arXiv preprint arXiv:2407.09818, 2024

  44. [50]

    Unleashing the power of text for credit default prediction: Comparing human-generated and ai-generated texts.SSRN Electronic Journal

    Zongxiao Wu, Yizhe Dong, Yaoyiran Li, and Baofeng Shi. Unleashing the power of text for credit default prediction: Comparing human-generated and ai-generated texts.SSRN Electronic Journal. https://doi. org/10.2139/ssrn, 4601317, 2023

  45. [51]

    Bankruptcy prediction: Data augmentation, llms and the need for au- ditor’s opinion

    Andreas Sideras, Konstantinos Bougiatiotis, Elias Zavitsanos, Georgios Paliouras, and George Vouros. Bankruptcy prediction: Data augmentation, llms and the need for au- ditor’s opinion. InProceedings of the 5th ACM International Conference on AI in Finance, pages 453–460, 2024. 17

  46. [52]

    A comparative analysis of in- struction fine-tuning llms for financial text classification.arXiv preprint arXiv:2411.02476, 2024

    Sorouralsadat Fatemi, Yuheng Hu, and Maryam Mousavi. A comparative analysis of in- struction fine-tuning llms for financial text classification.arXiv preprint arXiv:2411.02476, 2024

  47. [53]

    Scalable fine-tunning strategies for llms in finance domain-specific ap- plication for credit union, 2024

    Kartheek Kalluri. Scalable fine-tunning strategies for llms in finance domain-specific ap- plication for credit union, 2024

  48. [54]

    Open finllm leaderboard: Towards financial ai readiness.arXiv preprint arXiv:2501.10963, 2025

    Shengyuan Colin Lin, Felix Tian, Keyi Wang, Xingjian Zhao, Jimin Huang, Qianqian Xie, Luca Borella, Matt White, Christina Dan Wang, Kairong Xiao, et al. Open finllm leaderboard: Towards financial ai readiness.arXiv preprint arXiv:2501.10963, 2025

  49. [55]

    Zigong 1.0: A large language model for financial credit.arXiv preprint arXiv:2502.16159, 2025

    Yu Lei, Zixuan Wang, Chu Liu, and Tongyao Wang. Zigong 1.0: A large language model for financial credit.arXiv preprint arXiv:2502.16159, 2025

  50. [56]

    Llms for financial advisement: A fair- ness and efficacy study in personal decision making

    Kausik Lakkaraju, Sara E Jones, Sai Krishna Revanth Vuruma, Vishal Pallagani, Bharath C Muppasani, and Biplav Srivastava. Llms for financial advisement: A fair- ness and efficacy study in personal decision making. InProceedings of the Fourth ACM International Conference on AI ...

  51. [57]

    Data-centric fingpt: Democratizing internet-scale data for financial large language models

    Xiao-Yang Liu, Guoxuan Wang, Hongyang Yang, and Daochen Zha. Data-centric fingpt: Democratizing internet-scale data for financial large language models. InNeurIPS Work- shop on Instruction Tuning and Instruction Following, 2023

  52. [58]

    Bridging language models and financial analysis.arXiv preprint arXiv:2503.22693, 2025

    Alejandro Lopez-Lira, Jihoon Kwon, Sangwoon Yoon, Jy-yong Sohn, and Chanyeol Choi. Bridging language models and financial analysis.arXiv preprint arXiv:2503.22693, 2025

  53. [60]

    Advanced default risk prediction in small and medum-sized enterprises using large language models.Applied Sciences, 15(5):2733, 2025

    Haonan Huang, Jing Li, Chundan Zheng, Sikang Chen, Xuanyin Wang, and Xingyan Chen. Advanced default risk prediction in small and medum-sized enterprises using large language models.Applied Sciences, 15(5):2733, 2025

  54. [61]

    Tagging enriched bank transactions using llm-generated topic taxonomies

    Daniel de S Moraes, Polyana B da Costa, Pedro TC Santos, Ivan de JP Pinto, S´ ergio Colcher, Antonio JG Busson, Matheus AS Pinto, Rafael H Rocha, Rennan Gaio, Gabriela Tourinho, et al. Tagging enriched bank transactions using llm-generated topic taxonomies. InBrazilian Symposi...

  55. [62]

    Open-finllms: Open multimodal large language models for financial applications.arXiv preprint arXiv:2408.11878, 2024

    Jimin Huang, Mengxi Xiao, Dong Li, Zihao Jiang, Yuzhe Yang, Yifei Zhang, Lingfei Qian, Yan Wang, Xueqing Peng, Yang Ren, et al. Open-finllms: Open multimodal large language models for financial applications.arXiv preprint arXiv:2408.11878, 2024

  56. [63]

    Dynamic financial sentiment analysis and market forecasting through large language models.International Journal of Human Computations & Intelli- gence, 4(1):397–410, 2025

    Haranadha Reddy Busireddy Seshakagari, Aravindan Umashankar, T Harikala, L Jayas- ree, and Jeffrey Severance. Dynamic financial sentiment analysis and market forecasting through large language models.International Journal of Human Computations & Intelli- gence, 4(1):397–410, 2025

  57. [65]

    Jiaxing Wang, Guoquan Liu, Yang Cheng, Xiaobo Xu, and Zhongyun Li. Leveraging internet-sourced text data for financial analytics in supply chain finance: A large language model-enhanced text mining workflow.IEEE Transactions on Engineering Management, 2025. 18

  58. [66]

    Chatgpt based credit rating and default forecasting.Journal of Data, Information and Management, pages 1–24, 2025

    Jinlin Lin, Sirui Lai, Hao Yu, Rui Liang, and Jerome Yen. Chatgpt based credit rating and default forecasting.Journal of Data, Information and Management, pages 1–24, 2025

  59. [67]

    A novel weighted loss tabtransformer integrating explainable ai for imbalanced credit risk datasets.IEEE Access, 2025

    Kristoko Dwi Hartomo, Christian Arthur, and Yessica Nataliani. A novel weighted loss tabtransformer integrating explainable ai for imbalanced credit risk datasets.IEEE Access, 2025

  60. [68]

    Leveraging large language models for enhancing financial compliance: A focus on anti-money laundering applications

    Yuqi Yan, Tiechuan Hu, and Wenbo Zhu. Leveraging large language models for enhancing financial compliance: A focus on anti-money laundering applications. In2024 4th Inter- national Conference on Robotics, Automation and Artificial Intelligence (RAAI), pages 260–273. IEEE, 2024

  61. [69]

    Systematic evaluation of long-context llms on financial concepts.arXiv preprint arXiv:2412.15386, 2024

    Lavanya Gupta, Saket Sharma, and Yiyun Zhao. Systematic evaluation of long-context llms on financial concepts.arXiv preprint arXiv:2412.15386, 2024

  62. [70]

    Analyse customer behaviour and sentiment using natural language pro- cessing (nlp) techniques to improve customer service and personalize banking experiences

    Dr M Suresh, G Vincent, C Vijai, M Rajendhiran, M Com, AH Vidhyalakshmi, and S Natarajan. Analyse customer behaviour and sentiment using natural language pro- cessing (nlp) techniques to improve customer service and personalize banking experiences. Educational Administration: ...

  63. [71]

    Harnessing earnings reports for stock predictions: A qlora-enhanced llm approach

    Haowei Ni, Shuchen Meng, Xupeng Chen, Ziqing Zhao, Andi Chen, Panfeng Li, Shiyao Zhang, Qifu Yin, Yuanqing Wang, and Yuxi Chan. Harnessing earnings reports for stock predictions: A qlora-enhanced llm approach. In2024 6th International Conference on Data-driven Optimization of ...

  64. [72]

    Large language models and financial market sentiment.Available at SSRN 4584928, 2023

    Shaun A Bond, Hayden Klok, and Min Zhu. Large language models and financial market sentiment.Available at SSRN 4584928, 2023

  65. [73]

    Enhancing auto insurance risk evaluation with transformer and shap.IEEE Access, 2024

    Tiejiang Sun, Jingyun Yang, Jiale Li, Jiaying Chen, Mingyue Liu, Li Fan, and Xukang Wang. Enhancing auto insurance risk evaluation with transformer and shap.IEEE Access, 2024

  66. [74]

    Identifying representation bias in large language models used in financial sentiment analysis

    Alpay Sabuncuoglu and Carsten Maple. Identifying representation bias in large language models used in financial sentiment analysis. In2025 IEEE Symposium on Computational Intelligence for Financial Engineering and Economics (CiFer), pages 1–7. IEEE, 2025

  67. [75]

    Finben: A holistic financial benchmark for large language models.Advances in Neural Information Processing Systems, 37:95716– 95743, 2024

    Qianqian Xie, Weiguang Han, Zhengyu Chen, Ruoyu Xiang, Xiao Zhang, Yueru He, Mengxi Xiao, Dong Li, Yongfu Dai, Duanyu Feng, et al. Finben: A holistic financial benchmark for large language models.Advances in Neural Information Processing Systems, 37:95716– 95743, 2024

  68. [76]

    Research on the application methods of large language model interpretability in fintech scenarios

    Zhiyi Liu, Kai Zhang, Yejie Zheng, and Zheng Sun. Research on the application methods of large language model interpretability in fintech scenarios. In2024 4th International Conference on Computer Communication and Artificial Intelligence (CCAI), pages 526–

  69. [77]

    At-fingpt: Financial risk prediction via an audio-text large language model.Finance Research Letters, 77:106967, 2025

    Yingnan Liu, Ningbo Bu, Zhiqiang Li, Yongmin Zhang, and Zhenyu Zhao. At-fingpt: Financial risk prediction via an audio-text large language model.Finance Research Letters, 77:106967, 2025

  70. [78]

    Secured framework for banking chatbots using ai, ml and nlp

    Rajat Chanda and Sandeep Prabhu. Secured framework for banking chatbots using ai, ml and nlp. In2023 7th International Conference on Intelligent Computing and Control Systems (ICICCS), pages 60–65. IEEE, 2023

  71. [79]

    Ai in investment analysis: Llms for equity stock ratings

    Kassiani Papasotiriou, Srijan Sood, Shayleen Reynolds, and Tucker Balch. Ai in investment analysis: Llms for equity stock ratings. InProceedings of the 5th ACM International Conference on AI in Finance, pages 419–427, 2024. 19

  72. [80]

    Sentiment analysis in finance: From transformers back to explainable lexicons (xlex)

    Maryan Rizinski, Hristijan Peshov, Kostadin Mishev, Milos Jovanovik, and Dimitar Tra- janov. Sentiment analysis in finance: From transformers back to explainable lexicons (xlex). IEEE Access, 12:7170–7198, 2024

  73. [83]

    Seonmi Kim, Seyoung Kim, Yejin Kim, Junpyo Park, Seongjin Kim, Moolkyeol Kim, Chang Hwan Sung, Joohwan Hong, and Yongjae Lee. Llms analyzing the analysts: Do bert and gpt extract more value from financial analyst reports? InProceedings of the Fourth ACM International Conferenc...

  74. [84]

    Credit risk meets large language models: Building a risk indicator from loan descriptions in peer-to-peer lending.Available at SSRN 4979155

    Mario Sanz-Guerrero and Javier Arroyo. Credit risk meets large language models: Building a risk indicator from loan descriptions in peer-to-peer lending.Available at SSRN 4979155

  75. [85]

    Is chatgpt a financial expert? evaluating language models on financial natural language processing.arXiv preprint arXiv:2310.12664, 2023

    Yue Guo, Zian Xu, and Yi Yang. Is chatgpt a financial expert? evaluating language models on financial natural language processing.arXiv preprint arXiv:2310.12664, 2023

  76. [86]

    Beyond black-box ai: A theory of interpretable transformers for asset pricing.Available at SSRN, 2025

    Hasan Fallahgoul. Beyond black-box ai: A theory of interpretable transformers for asset pricing.Available at SSRN, 2025

  77. [87]

    Enhancing financial sentiment analysis via retrieval augmented large language models

    Boyu Zhang, Hongyang Yang, Tianyu Zhou, Muhammad Ali Babar, and Xiao-Yang Liu. Enhancing financial sentiment analysis via retrieval augmented large language models. In Proceedings of the fourth ACM international conference on AI in finance, pages 349–356, 2023. 20

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.