{"id":"5b4b39dd-5561-4cb0-8c87-66f5680f5b21","arxiv_id":"2506.04290","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review and taxonomy that organizes LLM-based credit risk research by model architecture, data modality, explainability mechanism, and application domain.","lead":"This paper reviews 60 studies on LLM-based credit risk and groups them into four categories: model type, data type, interpretability method, and application area. It is a reference survey, not an experimental study, and its value depends on whether its paper selection is trustworthy.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The PRISMA corpus of 60 studies is not deduplicated or auditable: Table 2 contains duplicate entries and at least one unresolvable reference, so the 'systematic review of 60 papers' claim is unsupported as written.","rationale":"I agree with the reader's weakest assumption and verdict. The paper's central contribution is framed as a systematic review of a precisely enumerated corpus: 'the most relevant 60 peer-reviewed studies' collected through PRISMA. If the corpus is not a deduplicated, correctly referenced set of 60 distinct studies, then the central claim—being the first systematic review of this area—lacks its evidentiary foundation. The duplicate references and the unresolvable row 10 are internal, checkable facts, not matters of taste or external consensus, so they are directly load-bearing. I considered whether the taxonomy could stand independently of the corpus count; it might, because the four-way organizational scheme and the per-category summaries (Sections 4.1–4.4) could survive correction of the selection list. That is why I do not recommend rejection. However, the paper should not be accepted as a systematic review until the corpus is corrected, deduplicated, and made auditable with explicit inclusion/exclusion criteria and a reproducible PRISMA flow. The reader's CONDITIONAL verdict is therefore appropriate, and my stress-test does not change it.","tokens_in":15284,"tokens_out":7869,"duration_ms":63721,"concrete_test":"Independently re-run the selection using the five keyword strings in §3.1 against the six named libraries for 2020–2025, then extract DOI/arXiv/SSRN identifiers for every entry in Table 2 and reconcile them with the reference list. Count distinct unique documents after normalizing versions (arXiv vs. SSRN vs. conference proceedings). If the distinct count is below 60—expected because refs [33]/[59], [34]/[64], and likely [38]/[44] collapse—or if row 10's 'Mehedi Hasan et al.' cannot be matched to any reference, the '60 relevant papers' claim and the PRISMA systematic-review claim fail as written and require revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that this is the first systematic review and taxonomy of LLM-based credit risk assessment, built on 'the most relevant 60 peer-reviewed studies' selected via PRISMA (§3.1, Figure 2, Table 2, §6). For that claim to hold, Table 2 must enumerate 60 distinct, correctly identified publications. It does not. Refs [33] and [59] are the same Pixiu paper (same title, same arXiv:2306.05443), appearing as rows 6 and 32. Refs [34] and [64] are the same Babaei and Giudici GPT-classifications paper (same journal, volume, and article number), appearing as rows 7 and 37. Refs [38] and [44] share the same title ('Optimizing large language models for financial risk assessment in credit unions') with different author names, which is at least a strong duplicate signal. Row 10 cites 'Mehedi Hasan et al. [37]', but reference [37] is Chen et al., 'Hallucination detection', so that Table entry cannot be resolved against the reference list. Row 54 also carries a venue/year that contradicts its cited reference ([81] is an ICON 2023 paper, not an unspecified 2024 preprint). Because the inclusion/exclusion criteria and deduplication procedure are described only as 'joint efforts' (§3.1), the 60-paper corpus is not auditable. The taxonomy may still be useful, but the 'first systematic review of 60 papers' claim and the PRISMA label are not supported by the presented evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript presents a systematic review and taxonomy of LLM-based credit risk assessment. It reports a PRISMA-guided selection of 60 papers published between 2020 and 2025 and organizes them along four dimensions: model architectures, data modalities, interpretability mechanisms, and application domains. It also compares 22 existing surveys, answers five research questions, and lists research gaps and future directions. The paper's central claim is that this is the first systematic review and taxonomy focused on LLM-based credit risk with interpretability as a core axis.","tokens_in":15608,"tokens_out":4923,"duration_ms":44477,"significance":"If the 60-paper corpus were complete, deduplicated, and correctly referenced, the paper would be a useful reference: the four-dimensional taxonomy is clear, the survey comparison in Table 1 helps position the work, and the gap analysis points to reasonable research directions. The manuscript also deserves credit for making its corpus enumerable in Table 2 and for stating explicit research questions. However, the current evidence does not support the claimed systematic foundation: duplicate entries and an unresolvable citation in Table 2 invalidate the '60 distinct peer-reviewed studies' premise, and the absence of a coding or extraction table weakens prevalence claims in Section 4 and the Conclusions. The stress-test concern about corpus auditability lands; the central claim is not supported as written, but the problems appear fixable within the scope of a revision.","major_comments":[{"comment":"Table 2, rows 6 and 32: both entries are the same Pixiu paper (refs [33] and [59] share the same title and arXiv:2306.05443), and rows 7 and 37 are the same Babaei and Giudici paper (refs [34] and [64], same journal, volume, and article number). Rows 11 and 17 share the title 'Optimizing large language models for financial risk assessment in credit unions' with different author attributions, which is at least a strong duplicate signal. Because the central claim is a systematic review of 60 distinct papers, these duplicates mean the corpus count and the PRISMA flow numbers in Figure 2 are not valid as reported. The authors must deduplicate the list, re-count from the 182 initial records, and either restrict the corpus to peer-reviewed publications or drop the 'peer-reviewed' qualifier, since many entries are arXiv or SSRN preprints (e.g., [33], [38], [39], [82]).","section":"Table 2"},{"comment":"Row 10 lists 'Mehedi Hasan et al. [37]', but reference [37] is Yuyan Chen et al., 'Hallucination detection', so the table entry cannot be mapped to the bibliography. Row 54 lists 'Chafekar et al. [81]' with venue 'Unspecified (likely arXiv or workshop preprint) 2024', while reference [81] is an ICON 2023 conference paper. Every row of the overview table must correspond to exactly one resolvable bibliographic entry; otherwise the extraction underlying the taxonomy cannot be audited.","section":"Table 2, row 10"},{"comment":"The PRISMA flow diagram is not auditable. The inclusion and exclusion criteria are not stated; the text only says papers were removed 'with the joint efforts of both authors ... according to the inclusion criteria', and Figure 2 reports no count of records excluded at the identification or screening stage. The transition from 51 papers to 60 by snowballing is also not documented. For a systematic review, the authors should report keyword-search dates, databases searched per query, eligibility criteria, screening decisions, and a list of excluded studies, or else scale back the PRISMA claim.","section":"Section 3.1 and Figure 2"},{"comment":"Claims that post-hoc methods such as SHAP and LIME are 'the most commonly used' and that interpretability techniques are 'prevalent' are not supported by any extraction table or frequency count. The four taxonomy subsections cite illustrative papers rather than a systematic coding of all 60 studies. Adding a per-paper coding matrix (papers by architecture, data type, interpretability mechanism, application, and evaluation metric) with summary counts is necessary to substantiate RQ3 and the prevalence statements in the conclusions.","section":"Section 4.3 and Conclusions"}],"minor_comments":[{"comment":"The text 'SHAP and LIMA (post-hoc) are prevalent interpretability techniques' should read 'SHAP and LIME'; LIME is the method discussed in Section 4.3.","section":"Conclusions"},{"comment":"The arrow from 51 to 60 is attributed to the snowball method in the text, but the diagram does not show the number of records added; label this explicitly and reconcile the counts.","section":"Figure 2"},{"comment":"There are multiple grammatical issues, e.g., 'In Joshi et al. study's AI frameworks in credit risk and trading applications are examined' and 'In [7], models such as FinGPT and BloombergGPT are discussed, ... but ignore credit risk applications'; a careful proofreading pass is needed.","section":"Section 2"},{"comment":"Reference [6] lacks publication venue and year, and reference [84] appears to duplicate reference [28] as an SSRN version; clarify whether these are distinct versions or separate records.","section":"References"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the four-axis taxonomy (model architecture, data type, interpretability, application) is a plausible way to organize this literature, and the gap analysis in Section 5 is sensible. What does not hold up is the systematic-review packaging. The PRISMA claim requires 60 distinct, correctly identified papers, and Table 2 doesn't give you that. Refs [33] and [59] are the same Pixiu paper, [34] and [64] are the same Babaei & Giudici paper, [38] and [44] share a title with different authors, and row 10 cites 'Mehedi Hasan et al. [37]' where [37] is actually Chen et al. on hallucination detection. Row 54's venue contradicts its own reference. That's not a nitpick; it means the count of '60 peer-reviewed studies' is not auditable, and the inclusion/exclusion criteria are described only as 'joint efforts.' So the central contribution, as stated, is not supported.\n\nWhat is genuinely useful is the categorization effort itself. The four axes are well chosen for this intersection of NLP, XAI, and credit risk, and the figures and tables give a researcher a quick map of where work exists and where it doesn't. The discussion of gaps (interpretability mostly post-hoc, little causal reasoning, reproducibility issues) matches what I've seen in the area. The paper also does a decent job of comparing prior surveys and showing what was missing.\n\nThe soft spots beyond the corpus: the 'prevalent interpretability techniques' claim in the conclusion isn't backed by a quantitative extraction table; it's asserted. The inclusion of some non-peer-reviewed preprints and internal work undercuts 'peer-reviewed studies.' And the abstract says 'most relevant 60 peer-reviewed studies' while the flow diagram includes preprints.\n\nIs the taxonomy salvageable? Yes. If the authors redo the selection, deduplicate properly, provide a screening table and excluded list, and fix the reference list, this becomes a serviceable reference paper. As written, the systematic-review claim is the load-bearing wall and it has visible cracks.\n\nWho is this for? Researchers entering LLM credit risk who want a quick map of model types and data modalities. It doesn't answer a scientific question and it doesn't change practice. But as an organizing map it has value. I'd give it a serious referee only if the authors are willing to fix the corpus; otherwise it's a desk reject with an invitation to resubmit. My vote: conditionally accept the topic, require the audit trail. I would not cite it in its current form.","headline":"Useful taxonomy, but the 'first systematic review of 60 papers' claim doesn't survive contact with Table 2; fix the corpus and this could be a reference map.","tokens_in":16115,"tokens_out":1732,"would_cite":false,"duration_ms":15709,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper presents the first systematic review and taxonomy of large-language-model approaches to credit risk assessment, built from 60 studies published between 2020 and 2025.","keywords":["large language models","credit risk assessment","systematic review","taxonomy","interpretability","explainable AI","financial NLP","PRISMA"],"falsifier":"Inspect the 60-entry table and resolve every numbered reference to a distinct, retrievable publication. If any paper appears twice under different numbers, or if a reference resolves to the wrong study, the claimed corpus size and the taxonomy's proportions change; if all 60 entries resolve cleanly to distinct papers, the descriptive claims of the review stand.","tokens_in":1475,"feed_emoji":"📊","tokens_out":1961,"duration_ms":73015,"temperature":0.7,"pith_summary":"This paper tries to establish that LLM-based credit risk assessment is now a recognizable research field with a shape that can be catalogued. It claims to be the first systematic review and taxonomy of that field, built from 60 studies published between 2020 and 2025 and selected through a documented literature-screening flow. The central contribution is a four-part classification: model architectures, data modalities, interpretability mechanisms, and application areas. If the classification holds, researchers and financial institutions gain a common vocabulary for comparing LLM credit-scoring systems and for seeing where work is missing.","feed_headline":"First taxonomy maps 60 AI credit-risk studies","feed_subtitle":"A systematic review classifies LLM credit-risk work by architecture, data, interpretability, and use case.","key_machinery":"The machinery is the four-axis taxonomy itself. Each collected study is slotted into model architecture, data modality, interpretability mechanism, and application domain, with the PRISMA flowchart, a staged, documented procedure for screening a literature search, used to make the selection repeatable. The taxonomy does the argumentative work: it converts 60 heterogeneous papers into comparable categories, which is what lets the paper claim to be the first reference classification and to list gaps as structural absences rather than individual complaints.","core_discovery":"On the paper's own terms, the discovery is that the scattered literature on LLMs in credit risk can be organized into a single taxonomy, and that doing so reveals the field's center of gravity: encoder-only and decoder-only transformers plus domain-specific financial LLMs dominate, with hybrid pipelines and parameter-efficient tuning rising; SHAP and LIME post-hoc explanations still dominate interpretability while chain-of-thought prompting and intrinsically transparent designs are growing; and applications concentrate on retail and SME scoring and news-sentiment signals. The paper further claims that this map exposes structural gaps, including reproducibility, bias, hallucination, efficiency, and missing benchmarks, that should drive the next round of research.","pith_inferences":["The same four-axis grid could be carried into neighboring regulated decisions such as insurance underwriting or loan pricing, where the interpretability-bias-reproducibility tensions recur.","The dominance of post-hoc explainability suggests a testable expectation: studies using chain-of-thought or intrinsically transparent models should show a smaller gap between the explanation and the true decision logic, but the paper does not measure that gap.","Because the corpus mixes peer-reviewed articles, preprints, and working papers, the category proportions should be treated as sensitive to inclusion criteria; re-running the same search protocol at a later date would likely shift counts.","The paper notes behavioral and external signals only as a gap, which implies a plausible trajectory: credit scoring may move from static snapshot models toward continuous monitoring as those signals are integrated."],"forward_implications":["LLMs can score credit from unstructured text such as loan descriptions, news, and analyst reports, not only from financial ratios and payment histories.","Post-hoc tools such as SHAP and LIME remain the dominant way to explain LLM credit decisions, but chain-of-thought prompting and intrinsically transparent models are emerging alternatives.","A reliable LLM credit-scoring pipeline should be evaluated not just on accuracy but also on fairness, hallucination resistance, reproducibility, and inference cost.","Hybrid and retrieval-augmented pipelines, parameter-efficient fine-tuning, and multimodal inputs are the directions the taxonomy identifies as the field's frontier.","Standardized benchmarks for LLM credit risk do not yet exist, which makes direct comparison across studies difficult."],"supporting_citations":[{"why":"Supplies the encoder-only and BERT-plus-tree branch: a credit-risk indicator built from P2P loan descriptions.","marker":"[28]"},{"why":"Supports the decoder-only branch by showing GPT classifications work for credit lending with little data.","marker":"[34]"},{"why":"Defines the domain-specific financial LLM benchmark branch (PIXIU/FinMA) for finance instruction data and evaluation.","marker":"[59]"},{"why":"Anchors the FinGPT data-centric branch with LoRA and reinforcement-learning stock price signals.","marker":"[57]"},{"why":"Provides the chain-of-thought and structured-validation interpretability branch for financial news.","marker":"[30]"},{"why":"Anchors intrinsically interpretable designs with a segmentation-based Logit Leaf model for credit scoring.","marker":"[47]"},{"why":"Supplies fairness metrics (ISIP, ISA) for auditing LLM financial advisement, used in the interpretability taxonomy.","marker":"[56]"},{"why":"Supports multimodal and hybrid inputs with ChatGPT-extracted traits combined with structured data in GPT-LGBM.","marker":"[82]"},{"why":"Provides the benchmarking and evaluation branch with the Open FinLLM Leaderboard for standardized comparison.","marker":"[54]"},{"why":"Anchors parameter-efficient tuning with QLoRA on a compact LLaMA model for earnings-based prediction.","marker":"[71]"}],"fun_headline_variants":["60 studies, one taxonomy: interpretable LLMs for credit risk","First systematic map of interpretable LLMs in credit risk","LLM credit-risk taxonomy: what 60 papers reveal","Interpretability gaps in 60 LLM credit-risk papers"],"cache_read_input_tokens":18304,"weakest_assumption_plain":"The load-bearing premise is that the 60 studies selected through the described screening flow are distinct, correctly referenced, and representative of the LLM credit risk literature from 2020 to 2025, because every category count and every identified gap in the review is computed from that corpus.","fun_headline_variants_meta":{"raw":{"variants":["60 studies, one taxonomy: interpretable LLMs for credit risk","First systematic map of interpretable LLMs in credit risk","LLM credit-risk taxonomy: what 60 papers reveal","Interpretability gaps in 60 LLM credit-risk papers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000943,"raw_usage":{"total_tokens":3976,"prompt_tokens":838,"completion_tokens":3138,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":454,"completion_tokens_details":{"reasoning_tokens":3069}},"tokens_in":454,"tokens_out":3138,"duration_ms":20762,"temperature":1.0,"reasoning_tokens":3069,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:54:10.541246+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Inspect the 60-entry table and resolve every numbered reference to a distinct, retrievable publication. If any paper appears twice under different numbers, or if a reference resolves to the wrong study, the claimed corpus size and the taxonomy's proportions change; if all 60 entries resolve cleanly to distinct papers, the descriptive claims of the review stand.","supporting_citations":[{"cited_title":"Data-centric fingpt: Democratizing internet-scale data for financial large language models","cited_arxiv_id":null,"evidence_quote":"Anchors the FinGPT data-centric branch with LoRA and reinforcement-learning stock price signals."},{"cited_title":"Investigating the beneficial impact of segmentation-based modelling for credit scoring.Decision Support Systems, 179:114170, 2024","cited_arxiv_id":null,"evidence_quote":"Anchors intrinsically interpretable designs with a segmentation-based Logit Leaf model for credit scoring."},{"cited_title":"Llms for financial advisement: A fair- ness and efficacy study in personal decision making","cited_arxiv_id":null,"evidence_quote":"Supplies fairness metrics (ISIP, ISA) for auditing LLM financial advisement, used in the interpretability taxonomy."},{"cited_title":"Gpt-lgbm: A chatgpt-based integrated framework for credit scoring with textual and structured data.Available at SSRN 4671511, 2023","cited_arxiv_id":null,"evidence_quote":"Supports multimodal and hybrid inputs with ChatGPT-extracted traits combined with structured data in GPT-LGBM."},{"cited_title":"Harnessing earnings reports for stock predictions: A qlora-enhanced llm approach","cited_arxiv_id":null,"evidence_quote":"Anchors parameter-efficient tuning with QLoRA on a compact LLaMA model for earnings-based prediction."}],"review_version":1}