Pith. sign in

REVIEW 11 cited by

Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.16974 v4 pith:RUBQPXMF submitted 2025-03-21 q-fin.GN cs.AIcs.CEcs.CLcs.LG

Assessing Consistency and Reproducibility in the Outputs of Large Language Models: Evidence Across Diverse Finance and Accounting Tasks

classification q-fin.GN cs.AIcs.CEcs.CLcs.LG
keywords consistencyoutputsanalysisfinancemodelsreproducibilitytasksaccounting
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This study provides the first comprehensive assessment of consistency and reproducibility in Large Language Model (LLM) outputs in finance and accounting research. We evaluate how consistently LLMs produce outputs given identical inputs through extensive experimentation with 50 independent runs across five common tasks: classification, sentiment analysis, summarization, text generation, and prediction. Using three OpenAI models (GPT-3.5-turbo, GPT-4o-mini, and GPT-4o), we generate over 3.4 million outputs from diverse financial source texts and data, covering MD&As, FOMC statements, finance news articles, earnings call transcripts, and financial statements. Our findings reveal substantial but task-dependent consistency, with binary classification and sentiment analysis achieving near-perfect reproducibility, while complex tasks show greater variability. More advanced models do not consistently demonstrate better consistency and reproducibility, with task-specific patterns emerging. LLMs significantly outperform expert human annotators in consistency and maintain high agreement even where human experts significantly disagree. We further find that simple aggregation strategies across 3-5 runs dramatically improve consistency. We also find that aggregation may come with an additional benefit of improved accuracy for sentiment analysis when using newer models. Simulation analysis reveals that despite measurable inconsistency in LLM outputs, downstream statistical inferences remain remarkably robust. These findings address concerns about what we term "G-hacking," the selective reporting of favorable outcomes from multiple generative AI runs, by demonstrating that such risks are relatively low for finance and accounting tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis

    cs.CV 2026-04 unverdicted novelty 8.0

    DF3DV-1K supplies 1,048 scenes with clean and cluttered image pairs plus a challenging 41-scene subset to benchmark and improve distractor-free radiance field methods.

  2. DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making

    cs.AI 2026-06 conditional novelty 6.0

    Identical replays of financial AI agents often reproduce the same decision while varying the tool-call path, in one prospective study 94-95% decision agreement versus 67-69% exact tool-path agreement.

  3. Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement

    cs.CL 2026-06 conditional novelty 6.0

    A prompt that forces LLMs to separate facts, inferences, and emotions reduces repeated-answer variability (+0.016 to +0.021 SI on a ~0.95 baseline) and, under injected state persistence, cuts decision-flip rate by 82%...

  4. DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis

    cs.CV 2026-04 unverdicted novelty 6.0

    DF3DV-1K supplies 1,048 real scenes with clean/cluttered image pairs and a 41-scene hard subset to benchmark and improve distractor-free radiance-field methods.

  5. Dataset-Level Metrics Attenuate Non-Determinism: A Fine-Grained Non-Determinism Evaluation in Diffusion Language Models

    cs.LG 2026-04 unverdicted novelty 6.0

    Dataset-level metrics in diffusion language models mask substantial sample-level non-determinism that varies with model and system factors, which a new Factor Variance Attribution metric can decompose.

  6. Towards a Science of AI Agent Reliability

    cs.AI 2026-02 conditional novelty 6.0

    Measuring 14 AI agents across two benchmarks, the paper finds 18 months of accuracy gains (≈0.21/yr) bought only small reliability gains (0.03–0.10/yr) under its consistency/robustness/predictability/safety framework.

  7. The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

    cs.LG 2025-05 unverdicted novelty 6.0

    Entropy minimization on self-generated outputs elicits strong reasoning in pretrained LLMs, matching or exceeding supervised RL methods on benchmarks.

  8. From Accuracy to Auditability: A Survey of Determinism in Financial AI Systems

    cs.AI 2026-05 unverdicted novelty 5.0

    Financial AI systems using tabular models, graph networks, and LLM agents exhibit nondeterminism that undermines reproducibility, quantified via experiments on public datasets and addressed by a proposed layered evalu...

  9. Large Language Models as Automatic Annotators and Annotation Adjudicators for Fine-Grained Opinion Analysis

    cs.CL 2026-01 conditional novelty 5.0

    On ASTE/ACOS benchmarks, LLM annotators achieve moderate span-level agreement with humans but low exact-match structure scores, and LLM adjudication gives inconsistent gains—making LLMs assistants rather than replacements.

  10. Large Language Models as Automatic Annotators and Annotation Adjudicators for Fine-Grained Opinion Analysis

    cs.CL 2026-01 conditional novelty 5.0

    LLMs can serve as moderate-fidelity annotators and adjudicators for ASTE/ACOS opinion labels, but exact-match triplet/quadruple accuracy remains far below human relational precision.

  11. Shapley in Context: Explaining Financial Language with Domain Expertise

    q-fin.CP 2026-07 unverdicted novelty 4.0

    Shapley values for LLM explanations in financial text are shown via theory and experiments to produce attributions consistent with financial reasoning.