{"id":"65128b13-40e7-42bd-bf95-feb4074f3d50","arxiv_id":"2412.20438","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of 2018-2023 NLP-in-finance papers reports that asset pricing is the most studied component and that classification, LSTM, and BERT-style models dominate, with persistent data and interpretability limitations.","lead":"This paper reviews about 60 studies to map which natural language processing methods are used in finance. It finds asset pricing is the most studied area and classification models plus LSTM and BERT-style algorithms dominate, while data quality and interpretability remain unsolved.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's frequency claims rest on an undocumented, non-reproducible 60-item corpus; Figure 3's printed percentages sum to 102%, so the headline counts are not independently checkable.","rationale":"The reader identified the same weakest assumption: representativeness of the approximately 60-paper corpus. I agree, and I found additional internal evidence that makes the concern concrete rather than hypothetical: the pie chart in Figure 3, as printed, has percentages summing to 102%. This suggests the coding categories are not mutually exclusive as implied by a pie chart, or there is a counting error; either way the headline statistic (47% hybrid/other) cannot be taken at face value. Since the paper's contribution is descriptive measurement, a non-reproducible sampling frame and an internally inconsistent count are load-bearing. I am not objecting to the qualitative narrative or to the plausibility of the trends; the issue is the epistemic status of the quantified claims. The appropriate remedy is the one the reader already proposed: release the corpus and coding protocol, or reframe the paper as an opinionated narrative. I would therefore keep CONDITIONAL. No formal verification, code, or data accompanies the paper, so there is no independent support to lean on. This is not a finding of fraud; it is a request for the evidence needed to check a measurement.","tokens_in":8166,"tokens_out":3345,"duration_ms":31800,"concrete_test":"Ask the authors to provide the full bibliographic list of the approximately 60 analyzed materials and a per-material coding sheet with three columns: FS component, information-processing technique, and model/algorithm family, together with the exact screening rules used to go from ~250 to 109 to 60. An independent reviewer should then recompute Figures 1–3 and the 47% hybrid share from this coding. If the recomputed percentages total exactly 100% (after handling overlapping labels) and preserve the rankings — asset pricing first, information classification first, LSTM/BERT top — the central claim is supported. If the list cannot be reconstructed or the rankings change when a reproducible Scopus/Web of Science search with explicit inclusion criteria is used instead, the quantitative claim should be downgraded.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 2.1 describes the sample: a Google Scholar query (reproduced only with a placeholder “*FS element mentioned in chapter 1*”), about 250 materials screened down to 109, then “around sixty” classified. The paper never lists the included materials, states inclusion/exclusion criteria, or reports the per-category counts behind Figures 1–3. The central claims in the abstract and Section 2.2 — asset pricing is the most studied component, information classification the most used technique, LSTM/BERT the most used algorithms, and hybrid/other techniques 47% of 60 materials — are therefore measurements whose sampling frame cannot be audited. The problem is not merely cosmetic: if the manual screening favored open-data asset-pricing studies or excluded venues not indexed by the ad-hoc query, every frequency comparison across components and techniques would shift. An internal check compounds this: the percentages printed in Figure 3 sum to 102% (2+10+3+2+2+2+2+5+2+10+12+3+47), implying overlapping categories or an arithmetic/coding error. Either way, the 47% headline share and the technique rankings cannot be verified from the paper as published. The broad qualitative direction (asset pricing frequent, classification common, LSTM/BERT popular) is plausible and consistent with the wider literature, so the concern is about supportability of the quantitative claims, not their truth.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a literature review of natural language processing (NLP) and text-mining applications in the financial system over 2018–2023. It claims that asset pricing, especially stock prediction, is the most studied component; information classification is the most used NLP technique; LSTM and BERT-type models are the most common algorithms; and hybrid/'other' techniques account for 47% of the 60 materials analyzed. It also lists open challenges such as data quality, context adaptation, and model interpretability, and proposes a workflow for analyzing financial text.","tokens_in":8334,"tokens_out":2920,"duration_ms":28991,"significance":"If the quantitative claims were properly supported, the paper would offer a convenient descriptive snapshot of a fast-moving interdisciplinary area. The qualitative challenge list is plausible and consistent with prior reviews, and the proposed workflow could be a useful starting point for practitioners. However, the paper's central contribution is empirical: the frequency distributions in Figures 1–3 are the main results. As it stands, these numbers cannot be checked because the sample is not documented, the search query is incomplete, and at least one figure contains a clear arithmetic inconsistency. The paper therefore currently falls short of the reproducibility standard expected for quantitative survey claims.","major_comments":[{"comment":"The sample construction is not reproducible. The query is given with a placeholder (“*FS element mentioned in chapter 1*”) rather than an actual search string, and no inclusion/exclusion criteria, list of the 109 screened materials, list of the ~60 classified materials, or raw per-category counts are provided. This makes the central descriptive claims (asset pricing most studied, information classification most used, LSTM/BERT most common, 47% hybrid share) unverifiable. The authors should supply the full query, screening protocol, and a supplementary table with the coding of each included material.","section":"Section 2.1"},{"comment":"The percentages printed in Figure 3 sum to 102% (2+10+3+2+2+2+2+5+2+10+12+3+47), indicating either overlapping categories or an arithmetic/coding error. Since the 47% 'other/hybrid' share is a headline result, the figure must be corrected and the underlying counts reported so the sum can be checked.","section":"Figure 3"},{"comment":"Table 1 marks 'restriction to confidential data' as 'solved' in the current research, but the text in Section 3 states that this limitation 'is not fully appointed yet' and later says it 'resonate[s] a need to address them as quickly as possible.' This is an internal contradiction about one of the paper's substantive claims concerning open challenges and should be resolved.","section":"Table 1 and Section 3"},{"comment":"The claim that 'most of the research materials combined probabilistic with vector-space models' is presented without any counts, definitions of 'probabilistic' and 'vector-space,' or coding rules. Without a transparent classification protocol, a reader cannot assess whether this statement is supported by the reviewed materials; it needs to be either operationalized or reformulated as a qualitative observation.","section":"Section 2.2"}],"minor_comments":[{"comment":"The sentence 'There are included citations and patents with the query' is confusing; the method later describes screening 'around 250 scientific materials' of which '35 of them books and the rest papers and articles.' Please clarify whether patents were actually included and how.","section":"Section 2.1"},{"comment":"References [21] and [28] are the same work (Malandri et al., 2018) but are cited as if they were distinct sources. Please merge them and renumber.","section":"References"},{"comment":"The term 'NEUS' appears in the list of techniques but is never defined. GNUS (Generalized News-based Sentiment Analysis) is also mentioned; please make the abbreviations consistent and define each one at first use.","section":"Section 2.2"},{"comment":"The abstract mentions 'bidirectional encoder models' while the body refers to 'BERT types'; please use consistent terminology.","section":"Abstract and Section 2.2"},{"comment":"Figure captions contain stray commas and incomplete phrasing (e.g., 'years 2018-2023,'). Also, 'club inter-domain results' in the Discussion is an unclear phrase; please reword.","section":"Figure captions and text"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a short review in a specialized engineering journal. The main quantitative claims are the paper's raison d'être, and their current lack of auditability is a serious weakness. However, the issue is correctable in principle: the authors could release the full search query, screening protocol, item-level data, and corrected figures. If the journal's standards for empirical survey papers are strict, the threshold for revision should be correspondingly high. I also note the paper reads as a translation or lightly edited draft; a professional language edit would help."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: This is a small, plainly written review of NLP text-mining in finance over 2018-2023. The qualitative message - asset pricing dominates, classification is the most common technique, LSTM/BERT lead - is plausible and matches what most people in the field would say. But the paper's own counts are not verifiable, and one figure doesn't add up.\n\nWhat it does well: it extends Gupta et al. 2020 by covering more recent work, and it packages a comparison of open challenges (Table 1) plus a two-step workflow diagram for someone starting in financial NLP. Those are helpful as orientation materials, not as research results.\n\nThe soft spot is the entire quantitative apparatus. Section 2.1 says a Google Scholar query was used, but the query is printed with a placeholder ('*FS element mentioned in chapter 1*') instead of the actual terms. The reader is told about 250 materials, then 109, then 'around sixty' classified, but no inclusion/exclusion criteria, no list of included papers, and no per-category counts. The histograms in Figures 1-3 are the only evidence for every headline frequency. Figure 3's printed percentages sum to 102%, so either categories overlap or the numbers are miscoded. Either way, the 47% 'other/hybrid' share and the technique rankings cannot be checked. These are not minor nits: the abstract's central claims rest on them.\n\nI also noticed a duplicated reference (21 and 28 are the same Malandri et al. paper) and a suspiciously dated arXiv entry (ref 24 says July 2022 for an arXiv number from 2019). The reference list needs care.\n\nThe qualitative discussion of limitations is sensible, and the distinction between solved, partially solved, and persistent challenges is useful. So the paper is not worthless. But it is not a reliable measurement. If the authors released the corpus and fixed the arithmetic, it could be a passable survey for a teaching context. As is, I would not cite it, and I would desk reject it in its current form with a clear message about reproducibility.","headline":"A plausible but non-reproducible mini-survey of NLP in finance: the challenge table is nice, the counts don't add up.","tokens_in":8902,"tokens_out":4624,"would_cite":false,"duration_ms":40030,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A review of 60 papers finds asset pricing and stock prediction dominate NLP use in finance from 2018-2023, led by classification techniques and LSTM/BERT-type models, with hybrids at 47%.","keywords":["natural language processing","text mining","financial system","asset pricing","stock prediction","information classification","LSTM","BERT"],"falsifier":"Re-run the stated literature search with a published list of included papers and the raw count per financial-system component, technique, and algorithm; if asset pricing no longer dominates or hybrid combinations do not come near 47%, the review's central descriptive claim is falsified.","tokens_in":7907,"feed_emoji":"📈","tokens_out":6079,"duration_ms":55954,"temperature":0.7,"pith_summary":"This paper is a literature review asking which NLP models and text-mining techniques are used across components of the financial system—asset pricing, corporate finance, derivatives, risk management, portfolio theory, public and international finance—between 2018 and 2023. It claims that asset pricing, especially stock prediction, is the most studied component; information classification is the most used information-processing technique; and LSTM- and BERT-type models are the most common algorithms, with about 47% of the 60 surveyed materials combining multiple models or proposing new ones. The authors also say most work mixes probabilistic with vector-space models and textual with numerical data. On the limitations side, they update an earlier challenge list and find that data quality, financial-context adaptation, and model interpretability remain unresolved. A sympathetic reader would care because the review offers a map of where NLP-finance research concentrates and where the bottlenecks are.","feed_headline":"Stock prediction dominates NLP-finance research","feed_subtitle":"A 2018–2023 review of about 60 papers finds classification and LSTM/BERT-type models leading, with hybrids at 47%.","key_machinery":"The carrying object is a three-axis classification scheme applied to the surveyed literature: financial-system component, information-processing NLP technique (retrieval, classification, extraction, or combination), and specific algorithm or model family (probabilistic, vector-space, LSTM, BERT-type, and hybrid variants). This scheme produces the histograms in Figures 1-3 that support the frequency claims, and it organizes the proposed workflow in Figures 4-5, where each language-analysis stage (lexical, syntactic, semantic, pragmatic, discourse) is paired with the persistent limitations the authors identify.","core_discovery":"On its own terms, the paper's central discovery is a frequency distribution over a manually collected corpus of about 60 scientific materials from 2018-2023: asset pricing (particularly stock prediction) is the financial-system component receiving the most NLP attention, corporate finance is second, and information classification is the dominant information-processing technique, with BERT-type and LSTM models the most frequently used algorithms. Roughly 47% of the surveyed materials use hybrid combinations or newly proposed methods such as NumHTML. The paper further finds that most models combine probabilistic and vector-space representations and that text signals are typically fused with numerical data, and it presents an engineering-oriented workflow in which a researcher selects a financial-system component, an NLP technique and model family, and an analysis stage, while checking the limitations attached to that stage.","pith_inferences":["The pattern the paper observes is likely driven by data availability rather than economic importance: open stock and news data make asset-pricing experiments cheap, so the same distribution may not reflect where NLP could add the most value.","Because the paper does not publish its screening criteria or the list of 60 materials, the 47% hybrid figure and component shares are not yet externally auditable; a reproducible protocol with raw counts would settle their stability.","The challenge table implies a concrete research agenda: build sector-specific financial lexicons and annotation sets, adapt models to non-IID and time-varying text streams, and design interpretability tools for finance-domain users.","One testable extension is to run the same classification scheme on the 2024-2025 literature to see whether ChatGPT-era LLMs displace LSTM and BERT as the default financial-text models."],"forward_implications":["If the frequency claims hold, asset pricing and stock prediction are where NLP-finance methods are maturing fastest, leaving public finance, derivatives, and other understudied components as open ground for new applications.","The dominance of information classification suggests that labeling and categorizing financial texts is the proven core, so progress on extraction, retrieval, and combined techniques could shift the field's center of gravity.","The prevalence of LSTM and BERT-type models plus 47% hybrids implies that domain-specific architectures and fine-tuned pretrained models, not generic pipelines, are the current performance frontier.","Persistent limits on data quality, financial lexicons, time-varying distributions, and interpretability mean that gains from larger models will be capped until annotation standards and context-adaptive methods improve.","The proposed engineering workflow gives a practical starting point: pick the component, technique, and analysis stage explicitly, and anticipate the known failure modes at each stage before building the system."],"supporting_citations":[{"why":"It supplies the earlier review whose challenge list Table 1 extends and updates.","marker":"[15]"},{"why":"It provides the language-analysis pipeline from lexical through discourse stages used in the proposed workflow.","marker":"[2]"},{"why":"It documents hybrid deep-learning methods whose frequency the review reports.","marker":"[3]"},{"why":"It is an example of a newly proposed domain-specific model (NumHTML) counted among the 47% hybrid share.","marker":"[5]"},{"why":"It is an example of FinBERT in portfolio optimization, used to illustrate the portfolio-theory component.","marker":"[7]"},{"why":"It is an example of sentiment-plus-ARX/Lasso forecasting used as an asset-pricing and economic-news application.","marker":"[4]"},{"why":"It is cited as evidence that news-based adaptive models improve on overfitting traditional factor models.","marker":"[17]"},{"why":"It is the source for the non-IID, time-varying, low-signal challenges facing asset-pricing models.","marker":"[18]"}],"fun_headline_variants":["Stock prediction tops NLP-finance studies","LSTM and BERT lead NLP in finance review","Hybrid models make up 47% of finance NLP papers","NLP finance review: asset pricing dominates research","Text mining finance: classification main technique"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the roughly 60 papers the authors selected from their keyword search and manual filtering fairly represent NLP research in finance during 2018-2023, even though the inclusion criteria, exclusion rules, and raw counts behind the histograms are not reported.","fun_headline_variants_meta":{"raw":{"variants":["Stock prediction tops NLP-finance studies","LSTM and BERT lead NLP in finance review","Hybrid models make up 47% of finance NLP papers","NLP finance review: asset pricing dominates research","Text mining finance: classification main technique"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00028,"raw_usage":{"total_tokens":1660,"prompt_tokens":943,"completion_tokens":717,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":646}},"tokens_in":559,"tokens_out":717,"duration_ms":6644,"temperature":1.0,"reasoning_tokens":646,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T23:20:36.329361+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the stated literature search with a published list of included papers and the raw count per financial-system component, technique, and algorithm; if asset pricing no longer dominates or hybrid combinations do not come near 47%, the review's central descriptive claim is falsified.","supporting_citations":[{"cited_title":"Comprehensive Review of Text-Mining Applications in Finance","cited_arxiv_id":null,"evidence_quote":"It supplies the earlier review whose challenge list Table 1 extends and updates."},{"cited_title":"Speech and Language Processing an Introduction to Natural Language Processing","cited_arxiv_id":null,"evidence_quote":"It provides the language-analysis pipeline from lexical through discourse stages used in the proposed workflow."},{"cited_title":"Deep Learning in Economics: A Systematic and Critical Review","cited_arxiv_id":null,"evidence_quote":"It documents hybrid deep-learning methods whose frequency the review reports."},{"cited_title":"NumHTML: NumericOriented Hierarchical Transformer Model for Multi-Task Financial Forecasting","cited_arxiv_id":null,"evidence_quote":"It is an example of a newly proposed domain-specific model (NumHTML) counted among the 47% hybrid share."},{"cited_title":"BERT’s Sentiment Score for Portfolio Optimization: A Fine-Tuned View in Black and Litterman Model","cited_arxiv_id":null,"evidence_quote":"It is an example of FinBERT in portfolio optimization, used to illustrate the portfolio-theory component."},{"cited_title":"Forecasting with Economic News","cited_arxiv_id":null,"evidence_quote":"It is an example of sentiment-plus-ARX/Lasso forecasting used as an asset-pricing and economic-news application."},{"cited_title":"A News-based Machine Learning Model for Adaptive Asset Pricing","cited_arxiv_id":"2106.07103","evidence_quote":"It is cited as evidence that news-based adaptive models improve on overfitting traditional factor models."}],"review_version":1}