REVIEW 3 major objections 6 minor 3 references
A Scoping Review of ChatGPT Research in Accounting and Finance
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This review maps 116 ChatGPT-related accounting and finance papers into three themes and finds the field is dominated by potential-application studies.
desk verdict Useful first systematic map of the ChatGPT/LLM literature in accounting and finance, but the headline adoption-maturity percentages are internally inconsistent and need fixing before publication. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a three-part organizing framework. The input component classifies studies by motivation and application area, such as audit, financial reporting, tax, asset pricing, and corporate finance. The process component classifies studies by which LLM capability they leverage, arranged on a ladder from word-embedding generation, information retrieval, and classification up through summarization, prediction, and logical-reasoning decision aids. The output component classifies studies into four adoption-maturity groups: conceptual papers, case studies, potential applications, and value realization. This framework, adapted from earlier technology-adoption reviews and paired with the idea that the stage of adoption shapes the type of research that can be written, is what turns a list of 116 papers into a map of the field and a list of gaps.
What would settle it
Re-run the same scoping exercise with a broader net: include preprint servers, working-paper series, and conference proceedings, and search terms such as 'large language model', 'generative AI', 'FinGPT', and 'BloombergGPT' in several languages, cataloging the adoption-maturity of every hit. If a substantial set of case studies or value-realization studies appears outside the original two databases and term restrictions, the claim that the field is dominated by potential applications is an artifact of search coverage rather than a property of the literature.
Extended reading notes
Core claim
On the paper's own terms, the emerging literature says three things at once. First, almost every accounting and finance domain has produced studies anticipating that LLMs will improve efficiency and effectiveness, with the heaviest concentration in auditing, financial reporting, asset pricing and investment, and corporate finance. Second, when LLMs are actually used as research tools, they frequently outperform dictionary-based and older machine-learning methods on classification, sentiment analysis, and summarization, with sentiment analysis, question-answering, and classification being the most commonly used capabilities. Third, measured by adoption maturity, the literature is dominated by potential applications, with 79.2% of accounting papers and 64.7% of finance papers in that category, and with only one case study and no value-realization accounting studies identified. The review also claims that the heavy volume of potential-application working papers can itself serve as a proxy for accelerated adoption, and that the next leap will come from reimagining processes rather than merely automating existing tasks.
Load-bearing premise
The load-bearing premise is that searching two English-language scholarly databases for only the words 'ChatGPT' or 'GPT' in titles, abstracts, or keyword lists, and dropping papers of five pages or fewer, captures the whole relevant population of LLM research in accounting and finance.
Editorial extensions
If this is right
- If the review's map is right, the next wave of accounting and finance LLM research should shift from demonstrating potential to measuring realized value, using actual adoption events and firm-level performance data.
- The near-total absence of case studies implies that researchers have an opening to document real implementations, including the organizational and regulatory context that shapes success.
- The concentration of LLM use in classification and sentiment analysis suggests that summarization, prediction, and decision-aid capabilities are underused, so those tasks are likely to be the next methodological frontier.
- The gaps the review flags in management accounting, numerical financial reporting, non-English text, and multimodal data are concrete opportunities if the field wants to follow the technology as it matures.
- The finding that LLM-assisted professionals appear more productive than unaided ones points toward substitution of traditional labor by LLM-augmented workflows, a trend worth tracking with archival data.
Reading between the lines
- Editorial inference: because the search covered only two English-language scholarly databases and only the literal terms 'ChatGPT' and 'GPT', the reported gaps, for instance in management accounting or tax, may be artifacts of search coverage rather than true properties of the literature.
- Editorial inference: the repeated result that GPT-4 passes accounting certification exams suggests that certification and assessment bodies may need to redesign exams to measure judgment that machines cannot yet replicate.
- Editorial inference: if LLM sentiment and classification measures reliably beat traditional dictionary methods, then published asset-pricing and disclosure studies built on older textual-analysis measures may need to be re-benchmarked against LLM-based measures.
- Editorial inference: the review's emphasis on text overlooks that current multimodal models accept images and audio; a natural extension is testing whether such models can extract signals from charts, earnings-call recordings, and video that text-only analysis misses.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript is a scoping review of recent research on ChatGPT and related large language models (LLMs) in accounting and finance. The authors describe a four-step review procedure, identify 48 accounting-related and 68 finance-related papers from SSRN and Web of Science up to March 2024, and organize the literature using an input-process-output framework. Their central claims are that the literature falls into three broad themes (applications of LLMs, use of LLMs as research tools, and implications for professionals and organizations) and that most studies are still at the 'potential applications' stage of adoption maturity rather than case studies or value-realization studies. The paper also proposes future research directions and provides a technical appendix on using ChatGPT and related APIs for research.
Significance. If the findings hold, the review provides a timely and useful synthesis of a fast-growing literature and a reasonable framework for future research. The authors are transparent about their search sources and counts, and they explicitly acknowledge important limitations such as chunkization bias in summarization tasks and look-ahead bias in prediction tasks. The technical appendix on model choice, context windows, parameters, prompt engineering, and batch processing is a practical contribution. The three thematic categories are plausible and broadly consistent with the cited studies. However, the headline adoption-maturity percentages contain an internal inconsistency, and the promised quality-assessment step is absent, so the quantitative core of the central claim is not fully auditable as reported.
major comments (3)
- [Section 3.3 (Tables 7 and 8)] The text in Section 3.3 states that potential-application studies 'constitute the majority of studies in accounting (57%) and represent over 62% in finance,' but Tables 7 and 8 report 38/48 = 79.2% for accounting and 44/68 = 64.7% for finance. The table totals match the combined SSRN+WoS sample described in Section 3.2, so the 57% and 62% figures appear to be computed from a different denominator, most likely the SSRN-only subset, without any disclosure. Because this statistic is the paper's most distinctive quantitative claim about the state of the literature, the authors must reconcile the numbers or explicitly report both sample definitions and explain why they differ.
- [Section 3 (methodology, step 3)] The methodology states that the review procedure includes 'executing a quality assessment' as step 3, but the paper never describes any quality-assessment criteria, who performed the assessment, or how its results affected inclusion or interpretation. This is a promised methodological component that is missing. The authors should either report the quality assessment or revise the procedure description to remove it.
- [Section 3.3 and Tables 7–8] The classification of papers into output categories (conceptual, case study, potential application, value realization) is the basis for the headline adoption-maturity finding, yet the paper provides no coding protocol, no inter-rater reliability statistics, and no coding file. Without this information, readers cannot audit the classifications that drive the central claim. I recommend adding a coding appendix with definitions, examples, and reliability checks, or at least a statement explaining why such checks are not feasible for this type of review.
minor comments (6)
- [Section 3.2] The introduction reports 264 SSRN papers meeting the initial criteria, but the retained counts are 37 accounting and 46 finance; please clarify how many papers were excluded at each step and how the economics-network papers were handled in the final sample.
- [Section 2.2] The sentence 'This news series builds upon its predecessor' should read 'This new series builds upon its predecessor.'
- [Section 2.2 and Appendix] Table 2 lists GPT-4 with an 8,192-token context window, while the Appendix states 'the most advanced GPT-4 model has a context window of 128K tokens'; clarify which model variant, such as GPT-4 Turbo, is meant.
- [Section 3.2] 'World of Science' should be 'Web of Science.'
- [Section 5] The paragraph beginning 'The second stream focuses on archival research...' is duplicated almost verbatim; one copy should be removed.
- [Section 3.2] The search strategy relies on 'ChatGPT' or 'GPT' in title, abstract, or keywords and excludes papers of five or fewer pages; this may miss relevant work using terms such as 'large language model' without 'GPT.' This is a legitimate scope choice, but it should be acknowledged as a coverage limitation in Section 3.2 or Section 6.
Circularity Check
No significant circularity: the review's three themes and adoption-maturity counts are classifications of an external literature, and the self-citations serve only as supporting context.
full rationale
The paper's central claims are descriptive classifications of an external body of literature: it identifies three themes (applications, research tools, and implications) and reports the distribution of papers across adoption-maturity categories. These claims are produced by the authors' manual review and categorization of SSRN and Web of Science papers, not by fitting a parameter to a subset of data and then predicting that same subset. The input-process-output framework and the four adoption-maturity groups are organizing definitions in Section 3.1, and the counts in Tables 7 and 8 are the empirical output of applying those definitions; no equation or statistical procedure makes the conclusions equivalent to the inputs by construction. The self-citations (e.g., Stratopoulos and Wang 2022; Stratopoulos, Wang, and Ye 2022) are used for background definitions, adoption-rate comparisons, and an analogy about adoption proxies; even the suggestion that the large number of 'potential applications' papers may be another proxy for predicting the adoption stage is explicitly a hypothesis, not a forced result. The internal inconsistency between the Section 3.3 text (57% and 62%) and Tables 7 and 8 (79.2% and 64.7%) is a correctness and reproducibility concern, not a circularity reduction, because neither figure is derived from the other by construction. No load-bearing step reduces to a self-citation chain or to a fitted input renamed as a prediction, so the appropriate finding is no circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption SSRN and WoS searches using 'ChatGPT' or 'GPT' in title, abstract, or keywords identify the relevant population of LLM research in accounting and finance.
- domain assumption Manual screening of titles and abstracts is a valid way to decide relevance and topic assignment.
- domain assumption The input-process-output framework borrowed from Lee et al. (2023) and the O'Leary adoption-stage model are appropriate lenses for organizing the literature.
- domain assumption When model version is undisclosed, 'ChatGPT' predominantly refers to GPT-3.5 because of timing.
Cite this review
Pith. "Pith review of A Scoping Review of ChatGPT Research in Accounting and Finance." pith.science (2026). https://pith.science/paper/J64V2T7M
@misc{pith2026241205731,
author = {Pith},
title = {Pith review of: A Scoping Review of ChatGPT Research in Accounting and Finance},
year = {2026},
howpublished = {\url{https://pith.science/paper/J64V2T7M}},
note = {Machine review of arXiv:2412.05731}
}
read the original abstract
This paper provides a review of recent publications and working papers on ChatGPT and related Large Language Models (LLMs) in accounting and finance. The aim is to understand the current state of research in these two areas and identify potential research opportunities for future inquiry. We identify three common themes from these earlier studies. The first theme focuses on applications of ChatGPT and LLMs in various fields of accounting and finance. The second theme utilizes ChatGPT and LLMs as a new research tool by leveraging their capabilities such as classification, summarization, and text generation. The third theme investigates implications of LLM adoption for accounting and finance professionals, as well as for various organizations and sectors. While these earlier studies provide valuable insights, they leave many important questions unanswered or partially addressed. We propose venues for further exploration and provide technical guidance for researchers seeking to employ ChatGPT and related LLMs as a tool for their research.
Reference graph
Works this paper leans on
-
[115]
The Benefits of Specific Risk-Factor Disclosures
https://doi.org/10.3390/risks11070115. 49 Hope, Ole-Kristian, Danqi Hu, and Hai Lu. 2016. “The Benefits of Specific Risk-Factor Disclosures.” Review of Accounting Studies 21 (4): 1005–45. https://doi.org/10.1007/s11142-016-9371-1. Hu, Nan, Peng Liang, and Xu Yang. 2023. “Whetting All Your Appetites for Financial Tasks with One Meal from GPT? A Comparison ...
arXiv 2016
-
[2023]
Your Employer Is (Probably) Unprepared for Artificial Intelligence
https://www.economist.com/the-world-ahead/2023/11/13/generative-ai-will-go- mainstream-in-2024. ———. 2023b. “Your Employer Is (Probably) Unprepared for Artificial Intelligence.” The Economist, 2023. https://www.economist.com/finance-and-economics/2023/07/16/your- employer-is-probably-unprepared-for-artificial-intelligence. Eisfeldt, Andrea L., Gregor Schu...
arXiv 2023
-
[2024]
Dividend Announcement and the Value of Sentiment Analysis
“Dividend Announcement and the Value of Sentiment Analysis.” Journal of Management Analytics 0 (0): 1–21. https://doi.org/10.1080/23270012.2024.2306929. Andreou, Panayiotis C., Neophytos Lambertides, and Marina Magidou. 2023. “Stock Price Crash Risk and the Managerial Rhetoric Mechanism: Evidence from R&D Disclosure in 10-K Filings.” SSRN Scholarly Paper....
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.