{"id":"ad03d53a-dc92-41d1-8b8a-eafd0f89a034","arxiv_id":"2412.05731","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic scoping review finds that ChatGPT and LLM research in accounting and finance clusters into three themes: domain applications, LLMs as research tools, and implications for professionals, with most studies still at the potential-application stage.","lead":"This paper reviews recent academic studies on ChatGPT and similar AI language models in accounting and finance. It groups the studies into three themes and lists what remains unknown for future researchers.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline adoption-maturity statistic is internally inconsistent: Tables 7–8 report 79.2%/64.7% potential applications, but §3.3 text says 57%/62%; the quantitative core of the central claim is not auditable as reported.","rationale":"The reader's weakest assumption was search coverage: the SSRN/WoS query with only 'ChatGPT' or 'GPT' may miss relevant LLM research in other repositories or under other terminology. That is a legitimate external-scope limitation, but scoping reviews may reasonably bound their corpus, and the qualitative theme structure could survive a broader search. The more immediate problem is internal: the paper's own results section contradicts its headline tables. Section 3.3 reports 57% and 'over 62%' for potential applications, while Tables 7 and 8 report 79.2% and 64.7% for the same concept. The abstract and the reader's strongest claim use the table numbers, so the most prominent quantitative contribution of the review is not self-consistent. This is directly checkable from the paper's own data, does not depend on an external judgment about what should have been searched, and affects whether the claimed 'most studies are still at potential-application stage' is accurately quantified. The concern does not overturn the broader qualitative finding that the literature is early-stage and application-oriented, so the existing CONDITIONAL verdict remains appropriate; it should be conditional on resolving the denominator and reproducing the counts.","tokens_in":41043,"tokens_out":7183,"duration_ms":68582,"concrete_test":"Recompute the Output Categories from the retained papers: list the 48 accounting and 68 finance papers, flag SSRN vs WoS origin, and recalculate potential-application counts and percentages for the combined sample and each subsample. If 38/48 = 79.2% and 44/68 = 64.7% are not reproducible, or if 57%/62% cannot be attributed to a stated subsample, the paper must correct either the tables or the text and state the denominator. As a secondary check, have a second coder independently classify a random 20% of the sample to measure inter-rater agreement on the four output categories.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The review's most distinctive quantitative finding is that most studies are at the 'potential applications' stage: 79.2% in accounting and 64.7% in finance. But Section 3.3, immediately after presenting Tables 7 and 8, states that potential-application studies 'constitute the majority of studies in accounting (57%) and represent over 62% in finance.' Table 7 shows 38/48 = 79.2%; Table 8 shows 44/68 = 64.7%. The 57% figure is not a rounding of 79.2%, and the text does not state that 57%/62% refer to the SSRN-only subset. If the tables use the combined SSRN+WoS sample and the text uses a different denominator, the definition of the reported sample shifts without disclosure. If both are meant to describe the same sample, one set of numbers is wrong. The reader's strongest claim reproduces the table numbers, so the central claim inherits this ambiguity. The absence of a coding file or inter-rater reliability check makes it impossible to determine which figure reflects the actual classification. This is more load-bearing than the search-coverage concern because it is an internal contradiction in the paper's own headline statistic, not an external scope choice.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript is a scoping review of recent research on ChatGPT and related large language models (LLMs) in accounting and finance. The authors describe a four-step review procedure, identify 48 accounting-related and 68 finance-related papers from SSRN and Web of Science up to March 2024, and organize the literature using an input-process-output framework. Their central claims are that the literature falls into three broad themes (applications of LLMs, use of LLMs as research tools, and implications for professionals and organizations) and that most studies are still at the 'potential applications' stage of adoption maturity rather than case studies or value-realization studies. The paper also proposes future research directions and provides a technical appendix on using ChatGPT and related APIs for research.","tokens_in":41266,"tokens_out":3612,"duration_ms":35693,"significance":"If the findings hold, the review provides a timely and useful synthesis of a fast-growing literature and a reasonable framework for future research. The authors are transparent about their search sources and counts, and they explicitly acknowledge important limitations such as chunkization bias in summarization tasks and look-ahead bias in prediction tasks. The technical appendix on model choice, context windows, parameters, prompt engineering, and batch processing is a practical contribution. The three thematic categories are plausible and broadly consistent with the cited studies. However, the headline adoption-maturity percentages contain an internal inconsistency, and the promised quality-assessment step is absent, so the quantitative core of the central claim is not fully auditable as reported.","major_comments":[{"comment":"The text in Section 3.3 states that potential-application studies 'constitute the majority of studies in accounting (57%) and represent over 62% in finance,' but Tables 7 and 8 report 38/48 = 79.2% for accounting and 44/68 = 64.7% for finance. The table totals match the combined SSRN+WoS sample described in Section 3.2, so the 57% and 62% figures appear to be computed from a different denominator, most likely the SSRN-only subset, without any disclosure. Because this statistic is the paper's most distinctive quantitative claim about the state of the literature, the authors must reconcile the numbers or explicitly report both sample definitions and explain why they differ.","section":"Section 3.3 (Tables 7 and 8)"},{"comment":"The methodology states that the review procedure includes 'executing a quality assessment' as step 3, but the paper never describes any quality-assessment criteria, who performed the assessment, or how its results affected inclusion or interpretation. This is a promised methodological component that is missing. The authors should either report the quality assessment or revise the procedure description to remove it.","section":"Section 3 (methodology, step 3)"},{"comment":"The classification of papers into output categories (conceptual, case study, potential application, value realization) is the basis for the headline adoption-maturity finding, yet the paper provides no coding protocol, no inter-rater reliability statistics, and no coding file. Without this information, readers cannot audit the classifications that drive the central claim. I recommend adding a coding appendix with definitions, examples, and reliability checks, or at least a statement explaining why such checks are not feasible for this type of review.","section":"Section 3.3 and Tables 7–8"}],"minor_comments":[{"comment":"The introduction reports 264 SSRN papers meeting the initial criteria, but the retained counts are 37 accounting and 46 finance; please clarify how many papers were excluded at each step and how the economics-network papers were handled in the final sample.","section":"Section 3.2"},{"comment":"The sentence 'This news series builds upon its predecessor' should read 'This new series builds upon its predecessor.'","section":"Section 2.2"},{"comment":"Table 2 lists GPT-4 with an 8,192-token context window, while the Appendix states 'the most advanced GPT-4 model has a context window of 128K tokens'; clarify which model variant, such as GPT-4 Turbo, is meant.","section":"Section 2.2 and Appendix"},{"comment":"'World of Science' should be 'Web of Science.'","section":"Section 3.2"},{"comment":"The paragraph beginning 'The second stream focuses on archival research...' is duplicated almost verbatim; one copy should be removed.","section":"Section 5"},{"comment":"The search strategy relies on 'ChatGPT' or 'GPT' in title, abstract, or keywords and excludes papers of five or fewer pages; this may miss relevant work using terms such as 'large language model' without 'GPT.' This is a legitimate scope choice, but it should be acknowledged as a coverage limitation in Section 3.2 or Section 6.","section":"Section 3.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is best suited for an accounting or finance field journal rather than a general economics or physics archive. The internal inconsistency in the adoption-maturity percentages and the absence of the promised quality-assessment step should be resolved before publication; the topic is timely and the framework is useful, but the quantitative claims need to be made auditable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a genuinely useful review and likely the first to systematically map the ChatGPT/LLM literature in accounting and finance. The three-theme structure (applications, research tools, adoption implications) is sensible and holds together. The paper also does something practical: it gives researchers a clear landscape of what's been tried, which subfields are crowded, and where the gaps are. The technical appendix on model choice, parameters, look-ahead bias, and batching is a real contribution and probably the part most people will cite.\n\nNow the soft spots, in proportion. The most serious is an internal contradiction in the paper's headline statistic. The text says potential-application studies constitute 57% of accounting and over 62% of finance papers. Tables 7 and 8 report 79.2% and 64.7% for the same categories. Those are not rounding differences. Since this \"most studies are still potential applications\" claim is one of the main takeaways, the paper needs to reconcile the numbers and clarify the denominator. Right now a reader cannot tell which figure describes the sample.\n\nThat problem is compounded by the absence of a coding file or inter-rater reliability check, so the classification itself is not independently auditable. The methodology section promises a quality assessment step, but the paper never reports the results. The search only covers SSRN and Web of Science and only terms \"ChatGPT\" and \"GPT,\" so studies using \"large language model\" without \"GPT\" are likely missed. That is a scope choice, but the conclusions should acknowledge it more explicitly. There is also a duplicated paragraph in Section V (the \"second stream\" description appears twice), which suggests an editing pass is needed.\n\nNone of this is fatal. The thematic mapping is useful on its own, and every issue I listed is fixable in revision. The paper will be a helpful starting point for accounting and finance researchers jumping into this area.\n\nSend it to peer review. A careful referee should ask for the numbers to be reconciled, for the coding procedure to be documented, and for a clean copy. After that, it will be a solid reference.","headline":"Useful first systematic map of the ChatGPT/LLM literature in accounting and finance, but the headline adoption-maturity percentages are internally inconsistent and need fixing before publication.","tokens_in":41778,"tokens_out":2296,"would_cite":true,"duration_ms":23377,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This review maps 116 ChatGPT-related accounting and finance papers into three themes and finds the field is dominated by potential-application studies.","keywords":["ChatGPT","large language models","accounting research","finance research","scoping review","technology adoption","generative AI","textual analysis"],"falsifier":"Re-run the same scoping exercise with a broader net: include preprint servers, working-paper series, and conference proceedings, and search terms such as 'large language model', 'generative AI', 'FinGPT', and 'BloombergGPT' in several languages, cataloging the adoption-maturity of every hit. If a substantial set of case studies or value-realization studies appears outside the original two databases and term restrictions, the claim that the field is dominated by potential applications is an artifact of search coverage rather than a property of the literature.","tokens_in":40843,"feed_emoji":"🤖","tokens_out":7234,"duration_ms":66867,"temperature":0.7,"pith_summary":"This scoping review claims that the early wave of ChatGPT and large language model research in accounting and finance can be mapped onto three themes: applying LLMs to accounting and finance tasks, using LLMs as research tools for classification, summarization, and text generation, and studying the implications for professionals, organizations, education, and labor markets. Organizing 116 papers through an input-process-output lens, the review finds that most work is still at the potential-application stage, with 79.2% of accounting studies and 64.7% of finance studies falling into that category, and only one case study identified. The paper further argues that LLMs usually outperform traditional methods in classification, sentiment analysis, and summarization, and that this concentration of forward-looking application papers is itself a sign of accelerated adoption. A sympathetic reader would care because the review clarifies where the field stands and identifies management accounting, numerical financial reporting, multilingual and multimodal analysis, and actual value-realization evidence as open gaps.","feed_headline":"Most ChatGPT studies still stop at 'potential applications'","feed_subtitle":"A scoping review of 116 papers finds only one case study and little value-realization evidence.","key_machinery":"The machinery is a three-part organizing framework. The input component classifies studies by motivation and application area, such as audit, financial reporting, tax, asset pricing, and corporate finance. The process component classifies studies by which LLM capability they leverage, arranged on a ladder from word-embedding generation, information retrieval, and classification up through summarization, prediction, and logical-reasoning decision aids. The output component classifies studies into four adoption-maturity groups: conceptual papers, case studies, potential applications, and value realization. This framework, adapted from earlier technology-adoption reviews and paired with the idea that the stage of adoption shapes the type of research that can be written, is what turns a list of 116 papers into a map of the field and a list of gaps.","core_discovery":"On the paper's own terms, the emerging literature says three things at once. First, almost every accounting and finance domain has produced studies anticipating that LLMs will improve efficiency and effectiveness, with the heaviest concentration in auditing, financial reporting, asset pricing and investment, and corporate finance. Second, when LLMs are actually used as research tools, they frequently outperform dictionary-based and older machine-learning methods on classification, sentiment analysis, and summarization, with sentiment analysis, question-answering, and classification being the most commonly used capabilities. Third, measured by adoption maturity, the literature is dominated by potential applications, with 79.2% of accounting papers and 64.7% of finance papers in that category, and with only one case study and no value-realization accounting studies identified. The review also claims that the heavy volume of potential-application working papers can itself serve as a proxy for accelerated adoption, and that the next leap will come from reimagining processes rather than merely automating existing tasks.","pith_inferences":["Editorial inference: because the search covered only two English-language scholarly databases and only the literal terms 'ChatGPT' and 'GPT', the reported gaps, for instance in management accounting or tax, may be artifacts of search coverage rather than true properties of the literature.","Editorial inference: the repeated result that GPT-4 passes accounting certification exams suggests that certification and assessment bodies may need to redesign exams to measure judgment that machines cannot yet replicate.","Editorial inference: if LLM sentiment and classification measures reliably beat traditional dictionary methods, then published asset-pricing and disclosure studies built on older textual-analysis measures may need to be re-benchmarked against LLM-based measures.","Editorial inference: the review's emphasis on text overlooks that current multimodal models accept images and audio; a natural extension is testing whether such models can extract signals from charts, earnings-call recordings, and video that text-only analysis misses."],"forward_implications":["If the review's map is right, the next wave of accounting and finance LLM research should shift from demonstrating potential to measuring realized value, using actual adoption events and firm-level performance data.","The near-total absence of case studies implies that researchers have an opening to document real implementations, including the organizational and regulatory context that shapes success.","The concentration of LLM use in classification and sentiment analysis suggests that summarization, prediction, and decision-aid capabilities are underused, so those tasks are likely to be the next methodological frontier.","The gaps the review flags in management accounting, numerical financial reporting, non-English text, and multimodal data are concrete opportunities if the field wants to follow the technology as it matures.","The finding that LLM-assisted professionals appear more productive than unaided ones points toward substitution of traditional labor by LLM-augmented workflows, a trend worth tracking with archival data."],"supporting_citations":[{"why":"Supplies the scoping-review typology and the four-step protocol that the paper follows.","marker":"Paré et al. 2015"},{"why":"Provides the precedent scoping review of an emerging technology whose framework the authors adapt to LLMs.","marker":"Yang Li et al. 2018"},{"why":"Supplies the input-process-output model used to organize motivations, capabilities, and outcomes.","marker":"Lee et al. 2023"},{"why":"Provides the argument that the technology-adoption stage determines the type of research that can be done, shaping the adoption-maturity categories.","marker":"O'Leary 2008"},{"why":"Supports the methodology of creatively collecting literature for emerging topics, justifying the inclusion of working papers.","marker":"Snyder 2019"}],"fun_headline_variants":["79% of accounting ChatGPT papers: potential, not practice","One case study and no value-realization in ChatGPT accounting work","LLMs outperform older methods, yet adoption maturity is low","ChatGPT research needs process reimagination, not automation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that searching two English-language scholarly databases for only the words 'ChatGPT' or 'GPT' in titles, abstracts, or keyword lists, and dropping papers of five pages or fewer, captures the whole relevant population of LLM research in accounting and finance.","fun_headline_variants_meta":{"raw":{"variants":["79% of accounting ChatGPT papers: potential, not practice","One case study and no value-realization in ChatGPT accounting work","LLMs outperform older methods, yet adoption maturity is low","ChatGPT research needs process reimagination, not automation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000948,"raw_usage":{"total_tokens":4014,"prompt_tokens":882,"completion_tokens":3132,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":498,"completion_tokens_details":{"reasoning_tokens":3064}},"tokens_in":498,"tokens_out":3132,"duration_ms":23343,"temperature":1.0,"reasoning_tokens":3064,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T20:24:19.013867+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same scoping exercise with a broader net: include preprint servers, working-paper series, and conference proceedings, and search terms such as 'large language model', 'generative AI', 'FinGPT', and 'BloombergGPT' in several languages, cataloging the adoption-maturity of every hit. If a substantial set of case studies or value-realization studies appears outside the original two databases and term restrictions, the claim that the field is dominated by potential applications is an artifact of search coverage rather than a property of the literature.","supporting_citations":[],"review_version":1}