Pith. sign in

REVIEW 31 cited by

A Survey of Large Language Models for Financial Applications: Progress, Prospects and Challenges

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.11903 v1 pith:NMV25SAK submitted 2024-06-15 q-fin.GN cs.AIq-fin.CP

classification q-fin.GNcs.AIq-fin.CP
keywords financialapplicationsllmsanalysisapplicationmodelssurveycapabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in large language models (LLMs) have unlocked novel opportunities for machine learning applications in the financial domain. These models have demonstrated remarkable capabilities in understanding context, processing vast amounts of data, and generating human-preferred contents. In this survey, we explore the application of LLMs on various financial tasks, focusing on their potential to transform traditional practices and drive innovation. We provide a discussion of the progress and advantages of LLMs in financial contexts, analyzing their advanced technologies as well as prospective capabilities in contextual understanding, transfer learning flexibility, complex emotion detection, etc. We then highlight this survey for categorizing the existing literature into key application areas, including linguistic tasks, sentiment analysis, financial time series, financial reasoning, agent-based modeling, and other applications. For each application area, we delve into specific methodologies, such as textual analysis, knowledge-based analysis, forecasting, data augmentation, planning, decision support, and simulations. Furthermore, a comprehensive collection of datasets, model assets, and useful codes associated with mainstream applications are presented as resources for the researchers and practitioners. Finally, we outline the challenges and opportunities for future research, particularly emphasizing a number of distinctive aspects in this field. We hope our work can help facilitate the adoption and further development of LLMs in the financial sector.

Discussion (0). Sign in to comment.

Forward citations

Cited by 31 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. AuditFraudBench: Benchmarking Audit Judgment in Detecting Fraudulent Misstatements

    cs.CE 2026-06 unverdicted novelty 7.0 of 10

    AuditFraudBench is a new enforcement-grounded benchmark with three tasks for testing whether LLMs can detect fraudulent misstatements by reasoning over financial figures, disclosure framing, and known manipulation patterns.

  2. From Information to Delegation: Mapping Human-AI Financial Decision Making

    cs.HC 2026-08 conditional novelty 6.0 of 10

    Across 1.5 million ChatGPT and Gemini chats in the US and India, consumers use AI overwhelmingly to inform and shape financial decisions, while delegation of financial execution remains rare.

  3. Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Among 8/8 self-consistent answers on FinQA, residual-stream probes detect wrong answers at 0.68–0.77 AUROC versus 0.55–0.63 for the best cheap output baselines across three 8–9B models.

  4. LLM-Enhanced Dynamic Financial Knowledge Graphs for Cross-Entity Signal Propagation and alpha discovery

    stat.AP 2026-07 conditional novelty 6.0 of 10

    In controlled simulations, community-aware propagation of LLM event signals on dynamic financial knowledge graphs recovers latent communities and prices incrementally beyond direct signals, though live alpha remains untested.

  5. Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language

    cs.CL 2026-07 accept novelty 6.0 of 10

    Current LLMs produce consistent but miscalibrated natural-language descriptors of likelihood and uncertainty from probabilistic predictions and are not yet reliable zero-shot risk communicators.

  6. Consistent but Miscalibrated: Evaluating LLM Limitations for Risk Communication in Natural Language

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Current LLMs are consistent but miscalibrated when selecting verbal descriptors for likelihood and uncertainty of probabilistic predictions, with the bottleneck in verbalization itself.

  7. CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    CLExEval introduces a human-annotated evaluation framework on 40 rare cases that identifies verbosity bias, hidden knowledge paradox, and 68.6% reasoning-to-output mismatch in LLMs while showing LLM-as-a-Judge overest...

  8. MetaPS: Adaptive Programmatic Strategy Selection for Market Agents

    cs.AI 2026-06 unverdicted novelty 6.0 of 10

    MetaPS trains models via simulation rollouts to select from programmatic strategy libraries for market agents, yielding better performance than fixed or direct LLM baselines across model sizes.

  9. Replacing Parameters with Preferences: Federated Alignment of Heterogeneous Vision-Language Models

    cs.AI 2026-05 unverdicted novelty 6.0 of 10

    MoR lets clients train local reward models on private preferences and uses a learned Mixture-of-Rewards with GRPO on the server to align a shared base VLM without exchanging parameters, architectures, or raw data.

  10. Memory in the LLM Era: Modular Architectures and Strategies in a Unified Framework

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    A unified framework for LLM agent memory is benchmarked, with a new hybrid method outperforming state-of-the-art on standard tasks.

  11. Learning to Conceal Risk: Controllable Multi-turn Red Teaming for LLMs in the Financial Domain

    cs.CL 2025-09 unverdicted novelty 6.0 of 10

    CoRT achieves 95% average attack success rate on nine LLMs by using iterative risk-concealing prompts and a controller that scores concealment levels on a new 522-instruction financial risk benchmark.

  12. THEME: Enhancing Thematic Investing with Semantic Stock Representations and Temporal Dynamics

    q-fin.PM 2025-08 conditional novelty 6.0 of 10

    A hierarchical contrastive learning framework that aligns stocks with theme descriptions and refines embeddings with short-term return signals improves thematic retrieval and backtested portfolio metrics.

  13. In-depth Analysis of Graph-based RAG in a Unified Framework

    cs.IR 2025-03 unverdicted novelty 6.0 of 10

    A unified framework and large-scale comparison of graph-based RAG methods on QA tasks yields new high-performing variants obtained by recombining existing components.

  14. ArchRAG: Attributed Community-based Hierarchical Retrieval-Augmented Generation

    cs.IR 2025-02 unverdicted novelty 6.0 of 10

    ArchRAG proposes attributed-community hierarchical indexing and LLM clustering to improve accuracy and lower token usage in graph-based retrieval-augmented generation.

  15. FinAcumen: Financial Multimodal Reasoning via Self-Evolving Experience Memory Harness

    cs.AI 2026-06 unverdicted novelty 5.0 of 10

    FinAcumen introduces selective experience memory that distills prior trajectories into reusable strategies and cautionary rules to improve tool-augmented multimodal financial reasoning.

  16. Closing the Feedback Loop: From Experience Extraction to Insight Governance in Verbal Reinforcement Learning

    cs.AI 2026-06 unverdicted novelty 5.0 of 10

    A feedback-driven three-layer architecture for insight governance in verbal reinforcement learning allows LLM agents to navigate retention-forgetting dilemmas in non-stationary settings, with the curation loop determi...

  17. YouZhi: Towards High-Concurrency Financial LLMs via Adaptive GQA-to-MLA Transition

    cs.CL 2026-06 unverdicted novelty 5.0 of 10

    YouZhi-LLM applies a layer-adaptive GQA-to-MLA transition plus Ascend-specific distillation and fine-tuning to reduce KV-cache size, yielding up to 2.69× higher concurrency and modest gains on financial benchmarks ver...

  18. Beyond Single-Policy: Evaluating Composed Organization-Specific Policy Alignment in LLM Chatbots

    cs.SE 2026-06 unverdicted novelty 5.0 of 10

    COPAL reveals a 33.1% average error rate on composed-policy queries across nine LLM chatbots, showing that existing single-policy benchmarks miss common failures.

  19. Fighting Numerical Hallucinations via Data-centric Compilation for Online Financial QA

    cs.IR 2026-05 unverdicted novelty 5.0 of 10

    DCRC applies data-centric methods with adversarial examples and program synthesis to produce verifiable reasoning programs for financial question answering.

  20. SoK: Security of Autonomous LLM Agents in Agentic Commerce

    cs.CR 2026-04 unverdicted novelty 5.0 of 10

    The paper systematizes security for LLM agents in agentic commerce into five threat dimensions, identifies 12 cross-layer attack vectors, and proposes a layered defense architecture.

  21. MetaGraph: A Large-Scale Meta-Analysis of GenAI in Financial NLP (2022-2025)

    cs.CL 2025-09 unverdicted novelty 5.0 of 10

    MetaGraph uses ontology-guided LLM extraction to build knowledge graphs from 681 papers on GenAI in financial NLP, identifying three distinct phases of development from 2022 to 2025.

  22. MetaGraph: A Large-Scale Meta-Analysis of GenAI in Financial NLP (2022-2025)

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Using LLM extraction on 681 papers, the authors build a public knowledge graph showing financial NLP moved from LLM adoption to limitation-aware, modular system design between 2022 and 2025.

  23. Governing Generative AI Across Financial Institutions: A Framework for Generative AI Risk Control

    q-fin.RM 2026-07 conditional novelty 4.0 of 10

    GAICF maps SR 26-2 model-risk principles into approved-use gates, risk tiers, evidence checks, and output monitoring for generative AI outside the formal model boundary.

  24. MimirRAG: A Multi-Agent RAG Framework for Financial Data Retrieval with Metadata Integration

    cs.LG 2026-05 unverdicted novelty 4.0 of 10

    MimirRAG, a multi-agent RAG framework with metadata integration and table-aware chunking, reaches 89.3% accuracy on FinanceBench and outperforms prior baselines for financial document retrieval.

  25. ComplianceNLP: Knowledge-Graph-Augmented RAG for Multi-Framework Regulatory Gap Detection

    cs.CL 2026-04 unverdicted novelty 4.0 of 10

    ComplianceNLP integrates knowledge-graph-augmented RAG, multi-task legal text extraction, and gap analysis to detect regulatory compliance gaps, reporting 87.7 F1 and real-world efficiency gains over GPT-4o baselines.

  26. Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial

    cs.NI 2025-09 conditional novelty 4.0 of 10

    A survey and tutorial that organizes LLM-enabled wireless network optimization into formulation, solution, and verification stages, with case studies drawn from the authors' own prior papers.

  27. What Factors Affect LLMs and RLLMs in Financial Question Answering?

    cs.CL 2025-07 unverdicted novelty 4.0 of 10

    Prompting and agent methods boost standard LLMs on financial QA by simulating long chain-of-thought reasoning, but reasoning LLMs already have this capability and show limited further gains, while multilingual alignme...

  28. Governing Generative AI Across Financial Institutions: A Framework for Generative AI Risk Control

    q-fin.RM 2026-07 unverdicted novelty 3.0 of 10

    A narrative survey organizes generative AI applications in finance into five capability patterns and maps them to business functions; no new results are reported.

  29. A Review of Large Language Models for Stock Price Forecasting from a Hedge-Fund Perspective

    q-fin.PR 2026-04 unverdicted novelty 3.0 of 10

    This review synthesizes LLM uses in stock forecasting and catalogs key practical pitfalls from a hedge-fund viewpoint.

  30. Data-Driven Function Calling Improvements in Large Language Model for Online Financial QA

    cs.IR 2026-04 unverdicted novelty 3.0 of 10

    A pipeline of dataset construction from prior work, AugFC parameter augmentation, and two-step LLM training improves function calling for financial APIs and is running in production.

  31. Bridging Language Models and Financial Analysis

    q-fin.ST 2025-03 unverdicted novelty 2.0 of 10

    A survey synthesizing recent LLM research and assessing its applicability to financial data analysis.

Pith tools