{"id":"c62e42b8-1c8a-4cca-8722-2df87ccd474c","arxiv_id":"2505.07393","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Interviews with six Fintech professionals find cautious use of LLMs for routine tasks, a view that existing regulation is inadequate, and a desire for bespoke in-house models.","lead":"This paper reports interviews with six Fintech professionals about how their companies use or plan to use ChatGPT and other large language models. It is worth reading as an early, on-the-ground check on whether a heavily regulated industry will adopt generative AI, and what is blocking it.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Six self-selected informants, only half recorded, cannot bear industry-level claims; the conclusion that 'the Fintech industry remains cautious' needs either a narrower scope or a full audit of the interview evidence.","rationale":"The reader's weakest-assumption analysis identifies exactly the same gap: six self-selected informants cannot support industry-level generalisations, and the absence of recordings for half of the interviews weakens the self-report evidence. My stress-test agrees and sharpens the point by separating two distinct threats. The first is external validity: because the sample is opportunistic, snowball-recruited, all-male, and geographically narrow, the direction and size of any industry-level effect are unknown. The second is internal auditability: with only 3 of 6 interviews recorded, the verbatim quotes in the findings cannot be traced to a preserved record, so the thematic analysis rests partly on unverifiable material. The paper itself signals its limitations ('small scale, exploratory') but the conclusion does not carry those limitations into its wording. Fixing this requires either re-scoping the claims or providing evidence that the analysis is fully traceable. The paper has real strengths: it addresses a genuinely under-researched question, uses a recognised inductive analytic approach, and offers concrete practitioner quotations that give insight into how LLM adoption is being considered in a regulated industry. Those strengths justify a conditional acceptance, but not an unconditional one; the reader's CONDITIONAL verdict is appropriate and should be retained with the explicit condition that the authors either narrow the claims to the informants studied or supply a complete audit trail for quotes and thematic counts.","tokens_in":16172,"tokens_out":5074,"duration_ms":50137,"concrete_test":"Audit the quotes in Sections 4.1-4.4 against the interview record: (i) produce a mapping from each quoted or paraphrased claim to the Table 1 respondent IDs; (ii) for each quote, state whether it came from a recorded interview, from contemporaneous notes, or from memory; (iii) verify that every claim labelled 'all respondents' or '5 out of 6' is actually supported by the coded transcripts. If the authors cannot supply this mapping, the thematic counts and industry-level statements should be relabelled as illustrative observations from six informants, and the conclusion should be revised accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the conclusion that 'the Fintech industry remains cautious and guarded in relation to the adoption of LLMs such as ChatGPT.' The data behind it are six semi-structured interviews recruited through the authors' social network and snowball sampling, all male, all based in Denmark or the UK (with one contradictory row listing 'UK and Sweden' in Table 1), and all selected because they had adopted or intended to adopt LLMs (Section 3.1). That design cannot support an unqualified industry-level claim: the sample is biased toward organisations already engaging with LLMs, so the reported caution may be an artefact of who was reachable through the authors' networks rather than a property of the industry. The concern is compounded by an evidentiary-chain problem: Section 3.1 states that only 3 of the 6 interviews were recorded, while Section 3.2 says all interview data were transcribed and translated. Direct quotes used in Sections 4.1-4.4 from unrecorded interviews cannot be verified as verbatim, and counts such as '5 out of 6' and 'all respondents' are not auditable without a mapping between quotes and respondent IDs. If the authors narrowed every conclusion to 'our six informants' and supplied a respondent-to-quote mapping, the paper would be a modest exploratory contribution; in its current form the central claim overreaches the evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a qualitative, interview-based study of how professionals in the Fintech industry view the adoption and use of large language models (LLMs), particularly ChatGPT. Six informants from Fintech companies in Denmark and the UK were recruited via the authors' social networks and snowball sampling, with each informant already using or intending to use LLMs. The analysis produces four themes: (1) adoption is cautious and limited to routine tasks, (2) existing regulations are seen as not fit for purpose, (3) there is widespread interest in bespoke in-house LLMs, and (4) green/sustainability concerns are deprioritized in LLM development. The paper concludes that 'the Fintech industry remains cautious and guarded in relation to the adoption of LLMs such as ChatGPT.'","tokens_in":16353,"tokens_out":4135,"duration_ms":37578,"significance":"If the findings are considered as a bounded exploratory study, the paper makes a useful contribution by giving voice to practitioners in a regulated industry at an early stage of LLM adoption, complementing hype-driven discourse with empirical interview data. The CSCW framing is appropriate, and the paper is transparent about its small-scale, exploratory nature and its inductive analytic approach. The four themes, if confirmed in broader samples, would offer a valuable starting point for understanding adoption barriers in Fintech. However, the current significance is undercut by the gap between the evidence (six non-random informants, only half of the interviews recorded) and the industry-level conclusions. The paper's value depends on carefully rescoping its claims and providing a more auditable evidentiary chain.","major_comments":[{"comment":"The central conclusion that 'the Fintech industry remains cautious and guarded in relation to the adoption of LLMs such as ChatGPT' overreaches the evidence. The sample consists of six self-selected informants recruited through the authors' social networks and snowball sampling, all male, all based in Denmark or the UK, and all already using or planning to use LLMs (Section 3.1, Table 1). These characteristics make the sample unrepresentative of the industry as a whole, and the paper itself describes the study as 'small scale, exploratory.' Please revise all industry-level statements, including the conclusion and the opening of Section 4, so that they refer explicitly to the six informants or their organizations, or provide a clear methodological argument for why this purposive sample can support broader inference.","section":"Section 3.1, Discussion and conclusion"},{"comment":"The evidentiary chain for the interview quotes is not auditable. Section 3.1 states that only 3 of the 6 interviewees allowed recording, yet Section 3.2 says 'the first author first transcribed and translated the interview data into English.' Direct quotations in Sections 4.1-4.4 cannot be verified as verbatim for the three unrecorded interviews. Please indicate which quotes come from recorded versus unrecorded interviews, provide a mapping between respondent IDs and the quotes used, and clarify how the unrecorded interviews were documented (for example, contemporaneous notes) and how this affects the use of direct quotations.","section":"Sections 3.1 and 3.2"},{"comment":"The analysis section states 'After three such rounds, we established five main categories,' but the Findings section (4.1-4.4) presents only four themes. This internal inconsistency needs to be corrected. If a fifth category was identified and then dropped, explain why; otherwise, align the reported number of categories with the findings actually presented.","section":"Section 3.2 vs Section 4"},{"comment":"Since each organization is represented by a single informant, claims such as 'All 6 fintech companies welcomed the advent of ChatGPT' (Section 4.3) and 'our respondents were uniformly cautious' (Discussion) conflate individual attitudes with organization-level positions. Please specify that the findings reflect the views of individual respondents, not necessarily the policies or positions of their companies, and adjust the language accordingly.","section":"Section 4.3 and Discussion"}],"minor_comments":[{"comment":"Respondent O's 'Organisation base' is listed as 'UK and Sweden,' which contradicts Section 3.1's statement that all participants lived and worked either in Denmark or the United Kingdom. Please clarify the correct geographic scope.","section":"Table 1"},{"comment":"There are several typographical and citation inconsistencies, including 'Kurtzweil' vs. 'Kurzweil,' 'Zanzotti' vs. 'Zanzotto,' 'Annany' vs. 'Ananny,' and 'Mclnerney' vs. 'McInerney.' Please correct these and ensure all reference names are spelled consistently.","section":"Throughout"},{"comment":"The reference list contains formatting issues, including an unnumbered entry for 'Linden, A., & Fenn, J. (2003) Understanding Gartner's hype cycles' and a duplicate reference for Kurzweil (entries [36] and [37]). These should be cleaned up and renumbered properly.","section":"References"},{"comment":"The sentence 'Nothing that looks like real-world evaluation in use has been published, to our knowledge, and very little which solicits expert views...' is too strong given that the paper cites Woodruff et al. (2024), which solicits expert views from knowledge workers. Please soften the claim to avoid an apparent contradiction.","section":"Section 2.2"}],"recommendation":"major_revision","confidential_remarks":"The paper is a plausible exploratory study with timely empirical data, but the current generalizations from six informants to the entire Fintech industry are not supportable. The five-versus-four categories discrepancy and the recording/transcription inconsistency suggest the manuscript needs careful revision rather than minor polishing. I recommend major revision with a clear request to rescope the claims and provide an audit trail for the interview evidence. If the authors are unwilling to narrow the conclusions, rejection would be warranted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nWorth a look if you care about LLM adoption in regulated sectors. This is a small interview study (six Fintech professionals in Denmark/UK) that extends Woodruff et al.'s finding that knowledge workers see generative AI as a tool for menial tasks, not decision-making. The new bits are domain-specific: regulation is the binding constraint, and all six want bespoke in-house models to control data flow. Those are legitimate, plausibly transferable observations, and the paper is honest about its exploratory character.\n\nThe analysis itself is reasonable. The quotes support the stated themes, the coding process is described, and the authors cite Woodruff et al. as the closest precedent. They do not oversell the novelty; if anything, the literature review is heavy on background and light on prior empirical work.\n\nThe soft spots are real but addressable. The biggest is that the conclusions repeatedly refer to \"the Fintech industry\" when the data are six self-selected men recruited through networks, all already using or planning to use LLMs. The paper's own \"small scale, exploratory\" framing does not license the generalizing language in the conclusion. That should be relaxed to \"our informants\" or the sample enlarged. Second, the audio recording chain has a wrinkle: Section 3.1 says only 3 of 6 interviews were recorded, yet Section 3.2 says all interviews were transcribed and translated. Direct quotes from unrecorded interviews are then presented as verbatim. The authors need to clarify whether notes were used and whether quotes are reconstructions. Third, the text promises five main themes in the analysis section and delivers four in the findings. Table 1's anonymized names don't fully match the names used later (Morten, Nicholai, Martin, Ola). These are mechanical problems, not misconduct, but they need cleanup.\n\nI would not block publication on the sample size alone—exploratory work has a place—but the industry-level claims must be scaled back. The citation list also has a few errors (e.g., a reference to Fuchs 2108, and a dangling \"Linden and Fenn\" entry). A serious referee could prompt a solid revision.\n\nWho is this for? CSCW/HCI people interested in how regulated industries talk about AI, and as a small counterpoint to hype. Not a foundation result, but a reasonable empirical data point.\n\nRecommendation: send to peer review with the expectation of major revisions, mainly to align claims with the evidence and fix the internal inconsistencies.","headline":"Small qualitative study with honest limitations; the findings are about six informants, not the Fintech industry, and that gap should be fixed before publication.","tokens_in":16954,"tokens_out":1694,"would_cite":false,"duration_ms":14972,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that Fintech professionals in Denmark and the UK are guarded about adopting LLMs like ChatGPT, limiting use to routine tasks until regulation becomes clearer.","keywords":["large language models","ChatGPT","Fintech","AI adoption","financial regulation","qualitative interviews","professional attitudes","green Fintech"],"falsifier":"A representative survey of Fintech companies in Denmark, Sweden, and the UK showing widespread deployment of LLMs in customer-facing financial advice or discretionary investment decisions would contradict the paper's central claim. So would evidence that existing GDPR or EU AI Act rules are being used as clear, workable adoption guidelines rather than described as unfit.","tokens_in":15915,"feed_emoji":"🤖","tokens_out":6787,"duration_ms":58366,"temperature":0.7,"pith_summary":"The paper tries to establish that Fintech professionals in Denmark and the UK, despite the hype around ChatGPT, are cautious and guarded about adopting large language models. Based on six interviews, it reports four shared views: LLMs are acceptable for routine clerical and analysis tasks but not for consequential decisions; existing regulations are seen as unfit for purpose; firms want bespoke in-house models to control data; and green or sustainability concerns are not a current priority. If true, this matters because it suggests that in a regulated financial industry, regulatory uncertainty rather than model capability is the main barrier to LLM adoption, and that practitioner priorities diverge from academic ethical concerns.","feed_headline":"Fintech stays guarded on ChatGPT, interviews show","feed_subtitle":"Firms limit LLMs to routine tasks and want in-house versions, citing unfit rules.","key_machinery":"The machinery is a small qualitative interview study: six semi-structured interviews with senior professionals representing six Fintech organisations, conducted by the first author and analysed through a general inductive approach in which the two authors grouped responses into themes over three rounds. Fintech is treated as a 'perspicuous setting'—a context where accountability to customers and regulators is unusually explicit—so the observed caution can be interpreted as institutional prudence rather than mere resistance. The four resulting themes carry the argument from interview quotes to the conclusion about industry-wide caution.","core_discovery":"The paper's central claim is that the Fintech industry—at least as represented by professionals in Denmark and the United Kingdom—remains cautious and guarded about adopting large language models such as ChatGPT. The authors ground this claim in four themes from their interviews: adoption is guarded and limited to routine tasks, existing regulations are seen as not fit for purpose, there is a universal desire to build bespoke in-house models to control data, and green concerns are deprioritised. They conclude that LLM potential is recognised, but real-world use is constrained by accountability and legal risk rather than by scepticism about the technology itself. The paper positions this practitioner view as a corrective to hype-driven discussions.","pith_inferences":["The authors do not say this, but the near-universal desire for bespoke in-house models suggests a commercial niche for fine-tuned, domain-specific LLMs where compliance and data control are the selling points.","A testable extension would be a larger survey of Fintech compliance officers to see whether 'regulation is unfit' is a stable industry belief or an artefact of this six-interview sample.","If regulatory uncertainty is the main brake, then countries with clearer AI rules should show measurably different adoption patterns; a cross-jurisdiction comparison could test that."],"forward_implications":["If the paper is right, LLM adoption in Fintech will initially stay confined to back-office and routine tasks, such as drafting documents, summarising data, and customer service analysis, while high-stakes advice and investment decisions remain human-led.","Regulatory clarity, not model performance, will be the decisive factor in speeding up or slowing down adoption; respondents saw current rules as ambiguous and reactive.","Fintech firms are likely to invest in bespoke, in-house models trained or adapted from general LLMs to control data flows and meet compliance requirements.","Sustainability and energy concerns will not, by themselves, deter Fintech adoption of LLMs in the near term, despite the green Fintech agenda.","The gap between academic worries about bias, privacy, and environment and practitioner priorities suggests that industry-focused AI governance should start from accountability and legal risk."],"supporting_citations":[{"why":"Supplies the prior interview evidence across seven industries that the paper extends, especially the finding that knowledge workers expect generative AI to substitute menial tasks rather than make decisions.","marker":"Woodruff et al, 2024"},{"why":"Provides the general inductive analysis method the authors use to condense interview data into the four reported themes.","marker":"Thomas (2003; 2006)"},{"why":"Supplies the idea of accountability that the paper uses to interpret why practitioners are cautious about LLM outputs.","marker":"Garfinkel (1967)"},{"why":"Documents a customised GPT-4 legal assistant, used as evidence that bespoke domain-specific LLMs are already feasible.","marker":"Callister, 2023"},{"why":"Provides the comparison of energy demands between blockchain and AI that frames the paper's discussion of scale and green concerns.","marker":"Sedlmeir et al, 2020"},{"why":"Defines the green Fintech agenda that the respondents deprioritise in their views on LLM development.","marker":"Macchiavello and Siri, 2022"},{"why":"Sets out EU AI policy and GDPR requirements that underlie the respondents' view that existing regulations are not fit for purpose.","marker":"Justo-Hanani, 2022"}],"fun_headline_variants":["Fintech pros wary of ChatGPT, prefer in-house AI","ChatGPT adoption in fintech stays guarded and limited","Rules aren't fit: why fintech holds back on ChatGPT","Fintech pros see LLM potential but not yet adoption-ready","Fintech prefers in-house AI over ChatGPT, citing unfit rules"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's industry-level conclusion rests on six self-selected informants, all male, based in Denmark or the UK and already using or planning to use LLMs, whose interview accounts are taken as describing their firms' actual practices.","fun_headline_variants_meta":{"raw":{"variants":["Fintech pros wary of ChatGPT, prefer in-house AI","ChatGPT adoption in fintech stays guarded and limited","Rules aren't fit: why fintech holds back on ChatGPT","Fintech pros see LLM potential but not yet adoption-ready","Fintech prefers in-house AI over ChatGPT, citing unfit rules"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000744,"raw_usage":{"total_tokens":3277,"prompt_tokens":860,"completion_tokens":2417,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":476,"completion_tokens_details":{"reasoning_tokens":2334}},"tokens_in":476,"tokens_out":2417,"duration_ms":15959,"temperature":1.0,"reasoning_tokens":2334,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:17:30.478256+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A representative survey of Fintech companies in Denmark, Sweden, and the UK showing widespread deployment of LLMs in customer-facing financial advice or discretionary investment decisions would contradict the paper's central claim. So would evidence that existing GDPR or EU AI Act rules are being used as clear, workable adoption guidelines rather than described as unfit.","supporting_citations":[{"cited_title":"G., Rousso -Schindler, S., Smith -Loud, J., & Wilcox, L","cited_arxiv_id":null,"evidence_quote":"Supplies the prior interview evidence across seven industries that the paper extends, especially the finding that knowledge workers expect generative AI to substitute menial tasks rather than make decisions."},{"cited_title":"(1967) Studies in Ethnomethodology, New York: Prentice Hall","cited_arxiv_id":null,"evidence_quote":"Supplies the idea of accountability that the paper uses to interpret why practitioners are cautious about LLM outputs."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents a customised GPT-4 legal assistant, used as evidence that bespoke domain-specific LLMs are already feasible."},{"cited_title":"U., Fridgen, G., & Keller, R","cited_arxiv_id":null,"evidence_quote":"Provides the comparison of energy demands between blockchain and AI that frames the paper's discussion of scale and green concerns."},{"cited_title":"(2022) Sustainable finance and fintech: Can technology contribute to achieving environmental goals? A preliminary assessment of ‘green fintech’and ‘sustainable digital finance’","cited_arxiv_id":null,"evidence_quote":"Defines the green Fintech agenda that the respondents deprioritise in their views on LLM development."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Sets out EU AI policy and GDPR requirements that underlie the respondents' view that existing regulations are not fit for purpose."}],"review_version":1}