{"id":"83afb1f0-41ed-48f8-8cb4-6d1e2eaf7ba6","arxiv_id":"2505.05523","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A systematic review of 83 papers identifies five thematic clusters in research on generative AI and entrepreneurship, with calls for macro-level and regulatory research.","lead":"This paper reviews 83 peer-reviewed studies on generative AI and entrepreneurship, using text mining to sort them into five thematic clusters. It offers a map of the field and a research agenda, but adds no new empirical data.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The five-cluster map is not reproducible from the reported method: Section 3.2 omits PCA dimensionality, linkage, distance, and dendrogram cutoff, and inconsistently describes clustering inputs as both TF-IDF vectors and keyword lists, so the central thematic taxonomy could be an artifact.","rationale":"The reader's weakest assumption identifies the clustering pipeline as the load-bearing premise: the five themes are the paper's main empirical contribution, and their validity depends on undocumented, unreported methodological choices. My review confirms this and pinpoints the specific missing parameters and the text's internal inconsistency between TF-IDF on textual content and clustering on keyword lists. The 'first systematic review' claim is also weakly supported because the paper cites prior systematic reviews in the same niche, but the cluster validity issue is the more fundamental threat to the central claim. I agree with the reader's CONDITIONAL verdict: the manuscript should be accepted only after the authors provide the corpus, parameters, and code, and after a robustness check of the clustering; otherwise the five themes, and the future-directions table built on them, lack demonstrable foundation. No ad hominem is intended; the concern is strictly about reproducibility and evidentiary weight.","tokens_in":24562,"tokens_out":2243,"duration_ms":25122,"concrete_test":"Ask the authors to release the 83-paper corpus, the exact preprocessing, and the Python code, then independently re-run the pipeline. As a single decisive check: vary the unspecified parameters (PCA components retained, linkage criterion, distance metric, and cluster cutoff between 2 and 10) and compute the adjusted Rand index (ARI) between the reported five-cluster partition and each alternative partition; also compute silhouette widths for the five-cluster solution. If the ARI drops below about 0.6 under reasonable parameter variations, or if silhouette widths are near zero or negative, the five-cluster taxonomy is not robust and the central mapping claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central contribution is the claim that TF-IDF, PCA, and hierarchical clustering yield five stable thematic clusters organizing the 83-paper corpus (Abstract; Section 3.2; Figure 2). The load-bearing assumption is that this pipeline, as described, actually determines those clusters. Section 3.2 does not report the number of PCA components retained, the linkage criterion (e.g., Ward, complete), the distance metric, or the dendrogram height at which the five-cluster cut is made. Without these, the dendrogram cannot be reconstructed and the cluster labels cannot be checked. The text also shifts between inputs: it first says TF-IDF vectorization transforms textual content, then says hierarchical clustering groups articles 'based on their keyword lists.' If clustering was done on keyword lists rather than TF-IDF abstracts/full texts, the method and results would differ substantially. No stability analysis (e.g., silhouette scores, bootstrap, parameter perturbation) or independent expert validation is reported, so the five themes could reflect the specific undocumented choices rather than meaningful intellectual groupings. This is the weakest link because every downstream finding—distribution by cluster, cluster-specific gaps, and future directions—depends on the validity of those five labels. A secondary but related inconsistency weakens the 'first systematic review' claim: the paper itself cites Dwivedi (2025) as 'A systematic review' and López-Solís et al. (2025) as 'a systematic literature review' in the GenAI-entrepreneurship space, which the conclusion's priority claim does not reconcile.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a systematic literature review of research on generative artificial intelligence (GenAI) and large language models (LLMs) in entrepreneurship. The authors searched Web of Science and Scopus, retained 83 peer-reviewed articles after screening, and applied TF-IDF vectorization, PCA, and hierarchical clustering to identify five thematic clusters. The paper reports descriptive statistics by cluster and year, summarizes the literature within each cluster, discusses ethical concerns, and proposes future research directions organized around the five clusters. The central claims are that these five thematic clusters accurately organize the field and that this is the first systematic literature review of GenAI and LLMs in entrepreneurship research.","tokens_in":24836,"tokens_out":3599,"duration_ms":35805,"significance":"If the clustering analysis is reliable and the novelty claim is accurate, the paper would provide a useful map of a rapidly growing research area and a structured research agenda. The study has several strengths: it follows a documented systematic-review protocol, reports the search and screening flow, gives a worked-out future research directions table, and explicitly addresses ethical dimensions. The unsupervised clustering approach is a reasonable and potentially valuable complement to purely narrative reviews. However, the load-bearing methodological step is not currently reproducible: Section 3.2 omits key parameters, contains an internal inconsistency about whether clustering was performed on TF-IDF vectors or keyword lists, and reports no validation of cluster stability or label validity. Several cited studies are also not clearly about GenAI, which raises questions about corpus composition. These issues must be resolved before the five-cluster taxonomy can be accepted as the paper's main contribution.","major_comments":[{"comment":"The clustering pipeline is not reproducible as described. The authors do not report the number of PCA components retained, the linkage criterion (e.g., Ward, complete, average), the distance metric, or the dendrogram height at which the five-cluster cut is made. Because the number of clusters is selected by inspecting the dendrogram, the five themes in Figure 2 and the cluster distribution in Figure 3 could change under alternative, equally plausible parameter choices. The text also shifts between two different inputs: it first says TF-IDF vectorization transforms textual content, then says hierarchical clustering groups articles 'based on their keyword lists.' If the clustering was actually performed on keyword lists rather than on TF-IDF representations of titles/abstracts/full texts, the method and results would differ substantially. I request that the authors specify the exact input representation, all preprocessing steps, all hyperparameters, the cluster-cut rule, and a stability or validation analysis (e.g., silhouette scores, bootstrap resampling, or independent expert labeling of cluster membership).","section":"Section 3.2"},{"comment":"The claim that this is 'the first systematic literature review of GenAI and LLMs in research on entrepreneurship' is contradicted by the paper's own references. Section 4.3 cites Dwivedi (2025) as 'A systematic review' and Section 4.4 cites López-Solís et al. (2025) as 'a systematic literature review.' Both are described as covering GenAI in entrepreneurship-adjacent domains. The authors should either demonstrate how their scope, corpus, or analytical approach differs from these existing reviews, or qualify the novelty claim accordingly.","section":"Section 6 (and Abstract)"},{"comment":"The search strategy appears over-inclusive relative to the stated focus on GenAI. The keyword list includes models and methods such as BERT, RoBERTa, GANs, and diffusion models, which are not necessarily generative AI in the sense used in the paper, and several included papers do not appear to be about GenAI at all, e.g., Karim et al. (2022) on ICT use, Knieps (2021; 2024) on 5G networks, Sajter (2024) on Croatian economic science, and Basilico and Graf (2023) on regional knowledge spaces. The screening stage is described only as an abstract screening with 27 exclusions; the authors should report inclusion/exclusion criteria in sufficient detail to allow replication and should consider a sensitivity analysis that shows whether the five clusters remain stable when questionable papers are removed.","section":"Section 3.1 (Figure 1) and Sections 4.3–4.6"},{"comment":"The cluster-level results cannot be verified because the paper does not provide a full list of which papers are assigned to which cluster. Figure 2 is a low-resolution dendrogram and Figure 3 gives only cluster sizes. I recommend that the authors include, as a supplementary table or appendix, the cluster membership for all 83 papers, together with the keywords or representative terms used to label each cluster. Without this, the qualitative summaries in Sections 4.2–4.6 cannot be checked against the actual clustering output.","section":"Figure 3 and Section 4.1.1"}],"minor_comments":[{"comment":"The sentence 'not least because of its impact on the preconditions for entrepreneurship' contains a singular/plural mismatch; 'its' should agree with 'GenAI and LLMs.'","section":"Abstract"},{"comment":"The phrase 'international mall and medium-sized enterprises' is a typo; it should read 'small and medium-sized enterprises.'","section":"Section 3.1"},{"comment":"In-text citations and reference entries are inconsistent: 'Abaddi, 2023' in the text corresponds to 'Abaddi, S. (2024)' in the references, and 'Thottoli et al. (2023)' in the text corresponds to a 2025 reference entry. These mismatches should be corrected throughout.","section":"Section 4.2 and References"},{"comment":"The reference list contains duplicate entries, including two entries for Maarouf et al. (2025) and two entries for Duong, C. D. (2024a/2024b) that are not consistently cited. A full bibliography cleanup is needed.","section":"References"},{"comment":"Figure 5 is referenced but not explained in the main text; the authors should describe how the common keywords were computed and what the figure is intended to show.","section":"Section 3.2 and Figure 5"},{"comment":"The paper says 'López-Solís et al. (2025) underscores the continued importance of human oversight' and 'a study by López-Solís et al. (2025)' but the reference list identifies it as a systematic literature review; the text should make clear that this is a review rather than a primary empirical study.","section":"Section 4.4"}],"recommendation":"major_revision","confidential_remarks":"This manuscript sits at the boundary of a bibliometric/descriptive review and a theory-driven systematic review. The main risk is that the five-cluster taxonomy is an artifact of undocumented analytical choices. If the authors can provide full methodological transparency, cluster membership lists, and a stability check, the contribution could be acceptable. If they cannot, the paper's central claims should be substantially softened. I would also encourage the editor to request the analysis code and data as part of the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a readable, well-structured map of an emerging literature, but the clustering pipeline is not reproducible as written and the \"first systematic review\" claim doesn't survive contact with the paper's own reference list.\n\nThe authors assembled 83 Scopus/WoS papers, screened them, and organized them into five themes. The five clusters are plausible, and the later narrative for each cluster is coherent. The future-research table with gaps, directions, and methods is genuinely helpful for someone scouting a niche. The call for more macro-level research on GenAI as an external enabler and for regulatory design is sensible. As an orientation document, this is a decent starting point.\n\nThe method section is the weak spot. Section 3.2 says TF-IDF was applied to textual content, then says hierarchical clustering grouped articles \"based on their keyword lists.\" Those are not the same thing. The paper reports no PCA dimensionality, linkage criterion, distance metric, or dendrogram cutoff, and no code, data, or stability analysis. Anyone trying to reproduce the five clusters would fail. This matters because the downstream description depends on those labels. The risk is moderate rather than fatal: the cluster narratives are plausible and appear to match the cited papers, so the taxonomy is probably not pure noise, but it could shift with different parameter choices.\n\nThe novelty claim also needs work. The conclusion asserts that this is \"the first systematic literature review of GenAI and LLMs in research on entrepreneurship,\" yet the paper itself cites Dwivedi (2025) as a systematic review and López-Solís et al. (2025) as a systematic literature review in this same space. The authors need to explain what their review adds beyond those, or drop the priority claim.\n\nThe search is over-inclusive: the query includes BERT, T5, and \"chatbot,\" which are not necessarily generative, and a few papers in the references (e.g., Al Halbusi et al. on ethical leadership) don't appear to involve GenAI at all. Abstract screening helped, but the final corpus still contains off-target items.\n\nMinor but real: typos, a duplicated reference, and an incomplete citation. These don't change the substance but suggest the manuscript wasn't carefully proofread.\n\nIf the authors release the corpus, code, and parameters, and add a robustness check, this could become a solid reference. As it stands, it is useful for orientation but not a trustworthy methodological template. I would send it to reviewers, with a request for major revision before publication.","headline":"Useful orientation map of GenAI-entrepreneurship research, but the clustering pipeline is not reproducible as reported and the 'first systematic review' claim overreaches.","tokens_in":25393,"tokens_out":2935,"would_cite":false,"duration_ms":31257,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims to be the first systematic literature review of generative AI and large language models in entrepreneurship research, and it organizes 83 peer-reviewed papers into five thematic clusters.","keywords":["GenAI in Entrepreneurship","LLMs in Entrepreneurship","Entrepreneurship","Systematic Literature Review","Clustering","Digital Transformation","Generative Artificial Intelligence","Large Language Models"],"falsifier":"Re-run the same TF-IDF, PCA, and hierarchical clustering recipe on the same 83 papers with different text inputs or similarity settings, have independent entrepreneurship scholars label the resulting groups, and compare; if the five themes do not reappear or the labels cannot be reproduced, the five-cluster map is an artifact of the method.","tokens_in":24372,"feed_emoji":"🤖","tokens_out":6423,"duration_ms":57147,"temperature":0.7,"pith_summary":"This paper sets out to establish that research on generative AI and large language models in entrepreneurship has matured enough to be mapped, and that it organizes into five recognizable themes. A sympathetic reading takes the central claim to be that the 83 peer-reviewed articles retrieved from Web of Science and Scopus form five clusters: digital transformation and behavioural models, GenAI-enhanced education, sustainable innovation and strategic AI impact, business models and market trends, and data-driven technological trends. If the mapping is right, it gives entrepreneurship scholars a baseline of what has been studied and a concrete agenda of what has not, especially macro-level questions about GenAI as an external enabler and about regulation that permits business experimentation.","feed_headline":"Review maps GenAI-entrepreneurship research into five clusters","feed_subtitle":"Analyzing 83 studies, the review finds five themes and flags missing macro-level and regulatory research.","key_machinery":"The mechanism that carries the argument is an unsupervised text-analysis pipeline: TF-IDF vectorization turns the textual content of each paper into a numerical vector; Principal Component Analysis reduces the dimensionality of those vectors; hierarchical clustering then merges similar papers into a tree, and the tree is cut to produce five groups. The paper also uses the notion of external enablers from entrepreneurship theory as the interpretive lens for why GenAI and LLMs matter for entrepreneurship. The pipeline is what generates the five cluster labels; the external-enabler lens is what turns the clusters into a research agenda.","core_discovery":"The paper's central claim is that this is the first systematic literature review devoted specifically to GenAI and LLMs in entrepreneurship research, and that the field's current intellectual landscape consists of five thematic clusters. The largest cluster, Business Models & Market Trends, contains 21 of the 83 papers; Sustainable Innovation & Strategic AI Impact contains 20; Data-Driven Technological Trends contains 18; GenAI-Enhanced Education & Learning Systems contains 13; and Digital Transformation & Behavioural Models contains 11. The authors argue that GenAI and LLMs act as external enablers that change the preconditions for entrepreneurship in two ways: by helping entrepreneurs do current work more effectively and by enabling new ventures, products, and business models. They conclude that existing research has concentrated on micro-level and firm-level effects, and they call for more macro-level research on scope, mechanisms, and roles of GenAI as an external enabler, and on regulatory frameworks that balance ethics and risk with experimentation and innovation.","pith_inferences":["If the clustering is robust, a natural next test would be to compare these keyword-derived themes with a citation-network or co-authorship map of the same 83 papers; topic clusters and community structures do not always coincide.","The paper's own observation that publication counts surged only after 2023 suggests the five clusters may be a snapshot of a pre-paradigmatic field, and later reviews will likely find new clusters or splits.","The external-enabler framing could be made operational by extracting, paper by paper, which specific enabling mechanism the authors credit to GenAI, such as cost reduction, speed, or new opportunity recognition, and testing whether those mechanisms differ across the five clusters.","One testable extension is to apply the same TF-IDF, PCA, and hierarchical clustering recipe to a larger corpus that includes preprints and conference papers; if the five themes persist, the map generalizes beyond peer-reviewed journals."],"forward_implications":["If the five-cluster map is correct, future literature reviews in this area can use it as a starting point rather than starting from scratch.","The cluster sizes imply that business-model and sustainability questions currently dominate, while digital-transformation and behavioural research is comparatively thinner.","The paper's gap analysis says the field needs longitudinal and experimental designs, more diverse samples beyond students in emerging economies, and more attention to macro-level external-enabler and regulatory questions.","The steep growth in publications from about three per year in 2020-2022 to 47 in 2024 suggests the review captures a field in rapid expansion, so the baseline may date quickly."],"supporting_citations":[{"why":"Supplies the systematic review procedure the paper follows for searching, screening, and synthesizing the literature.","marker":"Tranfield et al., 2003"},{"why":"Provides the TF-IDF vectorization method used to turn paper texts into numerical features.","marker":"Leskovec et al., 2019"},{"why":"Supplies the principal component analysis step that reduces the dimensionality of the TF-IDF vectors.","marker":"Lever et al., 2017"},{"why":"Provides the hierarchical clustering technique used to group papers and build the dendrogram.","marker":"Patel et al., 2015"},{"why":"Supplies the external-enabler concept used to frame GenAI and LLMs as environmental factors that enable entrepreneurship.","marker":"Kimjeon and Davidsson, 2022"},{"why":"Establishes the earlier AI-and-entrepreneurship review that the paper positions itself against by focusing specifically on GenAI and LLMs.","marker":"Giuggioli & Pellegrini, 2023"}],"fun_headline_variants":["GenAI entrepreneurship review finds five research clusters","83 papers map five GenAI-entrepreneurship research themes","GenAI entrepreneurship research clusters into five themes","Five clusters emerge in GenAI entrepreneurship, macro view missing","GenAI's entrepreneurship impact: five clusters, missing macro research"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The five themes are genuine divisions in the literature rather than artifacts of the clustering setup, because the number of clusters was chosen from a dendrogram and no stability check or independent expert validation is reported.","fun_headline_variants_meta":{"raw":{"variants":["GenAI entrepreneurship review finds five research clusters","83 papers map five GenAI-entrepreneurship research themes","GenAI entrepreneurship research clusters into five themes","Five clusters emerge in GenAI entrepreneurship, macro view missing","GenAI's entrepreneurship impact: five clusters, missing macro research"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000535,"raw_usage":{"total_tokens":2584,"prompt_tokens":972,"completion_tokens":1612,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":588,"completion_tokens_details":{"reasoning_tokens":1534}},"tokens_in":588,"tokens_out":1612,"duration_ms":13596,"temperature":1.0,"reasoning_tokens":1534,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:15:31.583300+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same TF-IDF, PCA, and hierarchical clustering recipe on the same 83 papers with different text inputs or similarity settings, have independent entrepreneurship scholars label the resulting groups, and compare; if the five themes do not reappear or the labels cannot be reproduced, the five-cluster map is an artifact of the method.","supporting_citations":[],"review_version":1}