Pith. sign in

REVIEW 14 cited by

Deficiency of Large Language Models in Finance: An Empirical Examination of Hallucination

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.15548 v1 pith:SDGVMIGH submitted 2023-11-27 cs.CL cs.AIcs.LGq-fin.ST

classification cs.CLcs.AIcs.LGq-fin.ST
keywords hallucinationllmsempiricalfinancialmodelsbehaviorsdeficiencyexamination
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The hallucination issue is recognized as a fundamental deficiency of large language models (LLMs), especially when applied to fields such as finance, education, and law. Despite the growing concerns, there has been a lack of empirical investigation. In this paper, we provide an empirical examination of LLMs' hallucination behaviors in financial tasks. First, we empirically investigate LLM model's ability of explaining financial concepts and terminologies. Second, we assess LLM models' capacity of querying historical stock prices. Third, to alleviate the hallucination issue, we evaluate the efficacy of four practical methods, including few-shot learning, Decoding by Contrasting Layers (DoLa), the Retrieval Augmentation Generation (RAG) method and the prompt-based tool learning method for a function to generate a query command. Finally, our major finding is that off-the-shelf LLMs experience serious hallucination behaviors in financial tasks. Therefore, there is an urgent need to call for research efforts in mitigating LLMs' hallucination.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. OpenHalDet: A Unified Benchmark for Hallucination Detection across Diverse Generation Scenarios

    cs.CL 2026-06 unverdicted novelty 7.0 of 10

    OpenHalDet creates a standardized benchmark and open codebase for comparing hallucination detectors across diverse LLM generation scenarios and access settings.

  2. Structure Guided Retrieval-Augmented Generation for Factual Queries

    cs.IR 2026-04 unverdicted novelty 7.0 of 10

    SG-RAG frames retrieval as subgraph matching to ensure LLMs meet every condition in factual queries and reports large gains over baselines on a new 120k-pair ERQA dataset.

  3. Hallucination Detection in Large Language Models Using Diversion Decoding

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Forcing an LLM away from its greedy answer yields resistance features that train a classifier detecting hallucinations more accurately and cheaply than semantic entropy.

  4. BALTO: Balanced Token-Level Policy Optimization for Hallucination Mitigation

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    BALTO projects claim-level verification into balanced token-level rewards for RL-based hallucination mitigation in LLMs.

  5. You Snooze, You Lose: Automatic Safety Alignment Restoration through Neural Weight Translation

    cs.CR 2026-05 unverdicted novelty 6.0 of 10

    NeWTral is a non-linear weight translation framework using MoE routing that reduces average attack success rate from 70% to 13% on unsafe domain adapters across Llama, Mistral, Qwen, and Gemma models up to 72B while r...

  6. HTDC: Hesitation-Triggered Differential Calibration for Mitigating Hallucination in Large Vision-Language Models

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    HTDC mitigates hallucinations in LVLMs by triggering calibration only at hesitation-prone decoding steps via contrasts with visual-nullification and semantic-nullification probes.

  7. Proof-Carrying Numbers (PCN): A Protocol for Trustworthy Numeric Answers from LLMs via Claim Verification

    cs.CL 2025-09 conditional novelty 6.0 of 10

    Proof-Carrying Numbers binds each displayed LLM number to a structured claim and only marks it verified after a policy-based mechanical check, leaving all other numbers unverified.

  8. ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking

    cs.CV 2025-08 reject novelty 6.0 of 10

    FAITH masks numbers in real 10-K reports to test when financial LLMs hallucinate, and finds even top models err on 10-20% of multi-step calculations.

  9. Towards Lightweight Reliability: Using Soft Prompts for Hallucination Mitigation in Large Language Models

    cs.CL 2026-05 unverdicted novelty 5.0 of 10

    RCSP trains soft prompts with contrastive loss, curriculum learning, and KL regularization to balance hallucination suppression, abstention, and factual recall, yielding higher F-scores than baselines on five QA datas...

  10. Fighting Numerical Hallucinations via Data-centric Compilation for Online Financial QA

    cs.IR 2026-05 unverdicted novelty 5.0 of 10

    DCRC applies data-centric methods with adversarial examples and program synthesis to produce verifiable reasoning programs for financial question answering.

  11. The Alpha Illusion: Reported Alpha from LLM Trading Agents Should Not Be Treated as Deployment Evidence

    cs.CE 2026-05 accept novelty 5.0 of 10

    Reported alpha from end-to-end LLM trading agents does not constitute deployment evidence until it passes structural tests for temporal integrity, frictions, robustness, calibration, execution, and disaggregation.

  12. ALDEN: Boosting Private Data Extraction from Retrieval-Augmented Generation Systems via Active Learning and Distribution Estimation

    cs.IR 2026-04 unverdicted novelty 5.0 of 10

    ALDEN boosts private data extraction rates from RAG systems by combining active learning for query diversification with dynamic estimation of the underlying knowledge-base topic distribution.

  13. LLM Harms: A Taxonomy and Discussion

    cs.CY 2025-12 reject novelty 3.0 of 10

    Proposes a five-bucket taxonomy of LLM harms and calls for dynamic auditing, but the systematic review behind it is not reproducible and contains mismatched citations.

  14. LLM Harms: A Taxonomy and Discussion

    cs.CY 2025-12 unverdicted novelty 3.0 of 10

    This paper proposes a taxonomy of LLM harms in five categories and suggests mitigation strategies plus a dynamic auditing system for responsible development.

Pith tools