Pith. sign in

REVIEW 12 cited by

Financial Statement Analysis with Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.17866 v3 pith:OK3EBBUO submitted 2024-07-25 q-fin.ST cs.AIcs.CLq-fin.GNq-fin.PM

Financial Statement Analysis with Large Language Models

classification q-fin.ST cs.AIcs.CLq-fin.GNq-fin.PM
keywords financialanalysisanalystsmodelsearningsfindfuturehuman
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We investigate whether large language models (LLMs) can successfully perform financial statement analysis in a way similar to a professional human analyst. We provide standardized and anonymous financial statements to GPT4 and instruct the model to analyze them to determine the direction of firms' future earnings. Even without narrative or industry-specific information, the LLM outperforms financial analysts in its ability to predict earnings changes directionally. The LLM exhibits a relative advantage over human analysts in situations when the analysts tend to struggle. Furthermore, we find that the prediction accuracy of the LLM is on par with a narrowly trained state-of-the-art ML model. LLM prediction does not stem from its training memory. Instead, we find that the LLM generates useful narrative insights about a company's future performance. Lastly, our trading strategies based on GPT's predictions yield a higher Sharpe ratio and alphas than strategies based on other models. Our results suggest that LLMs may take a central role in analysis and decision-making.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy

    cs.AI 2026-06 conditional novelty 7.0

    Composite LLM scoring saturates on expert investment frameworks while Gate Reconstruction Accuracy still reveals a clear procedural deficit at the frontier.

  2. InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy

    cs.AI 2026-06 unverdicted novelty 7.0

    InvestPhilBench is a new multi-layer benchmark for LLM procedural reasoning in investment philosophy, with BASP metrics showing composite scores saturate while gate reconstruction accuracy reveals procedural deficits.

  3. CrossAlpha: An Annual-Report Benchmark for Cross-Market Factor Researc (with LLM Agents)

    cs.IR 2026-05 unverdicted novelty 7.0

    CrossAlpha is a new open benchmark covering 3600 firms across US, Japan, Taiwan, South Korea and Hong Kong that distills reports into 10-category descriptions, constructs residual cross-market similarity graphs, and t...

  4. CSTrader: A Testbed for Language-Grounded Trading in a Community-Driven Virtual Asset Market

    cs.AI 2026-06 unverdicted novelty 6.0

    CSTrader is a multi-agent LLM trading system for CS2 skins that outperforms a -15.62% market index and single-prompt baselines with up to 7.58% returns by using specialized agents for liquidity, sentiment reversal, an...

  5. ChatGPT as a Time Capsule: The Limits of Price Discovery

    q-fin.GN 2026-04 unverdicted novelty 6.0

    Frozen LLM checkpoints serve as time capsules of public text and generate outlook scores that forecast equity returns and analyst actions beyond contemporaneous valuations.

  6. QRAFTI: An Agentic Framework for Empirical Research in Quantitative Finance

    cs.MA 2026-04 unverdicted novelty 6.0

    QRAFTI is a multi-agent framework using tool-calling and reflection-based planning to emulate quant research tasks like factor replication and signal testing on financial data.

  7. FinTexTS: Financial Text-Paired Time-Series Dataset via Semantic-Based and Multi-Level Pairing

    cs.AI 2026-03 conditional novelty 6.0

    A new dataset and pairing framework links stock prices to semantically relevant news at macro, sector, related-company, and target-company levels, improving stock forecast accuracy over keyword-based pairing.

  8. YouZhi: Towards High-Concurrency Financial LLMs via Adaptive GQA-to-MLA Transition

    cs.CL 2026-06 unverdicted novelty 5.0

    YouZhi-LLM applies a layer-adaptive GQA-to-MLA transition plus Ascend-specific distillation and fine-tuning to reduce KV-cache size, yielding up to 2.69× higher concurrency and modest gains on financial benchmarks ver...

  9. Macro Economists in the Machine: A Multi-Agent LLM Framework for Commodity-Related ETF Portfolio Construction

    q-fin.PM 2026-06 conditional novelty 4.0

    LLM agents (hawkish, dovish, debate) outperform a deterministic z-score rule agent in Sharpe ratio for commodity ETF portfolios by 0.04-0.044, with advantage concentrated in the soft-landing sub-period and preserved u...

  10. Cognitive Agents Powered by Large Language Models for Agile Software Project Management

    cs.SE 2025-08 reject novelty 4.0

    LLM agents acting as Agile roles produced plausible project artifacts in simulation, but the claimed improvements over human teams are unsupported because no comparison or validated metrics are provided.

  11. Predicting Liquidity-Aware Bond Yields using Causal GANs and Deep Reinforcement Learning with LLM Evaluation

    q-fin.CP 2025-02 unverdicted novelty 3.0

    CausalGAN + SAC RL pipeline generates synthetic bond yield data; fine-tuned Qwen2.5-7B LLM produces trading signals, with reported MAE 0.103, 60% profit rate, and LLM score 3.37/5.

  12. Bridging Language Models and Financial Analysis

    q-fin.ST 2025-03 unverdicted novelty 2.0

    A survey synthesizing recent LLM research and assessing its applicability to financial data analysis.