Pith. sign in

REVIEW 16 cited by

INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.18174 v1 pith:FFWK224R submitted 2024-12-24 cs.CE cs.AIq-fin.CP

INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent

classification cs.CE cs.AIq-fin.CP
keywords financialdecision-makingagentagentstaskscomprehensiveinvestorbenchacross
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Recent advancements have underscored the potential of large language model (LLM)-based agents in financial decision-making. Despite this progress, the field currently encounters two main challenges: (1) the lack of a comprehensive LLM agent framework adaptable to a variety of financial tasks, and (2) the absence of standardized benchmarks and consistent datasets for assessing agent performance. To tackle these issues, we introduce \textsc{InvestorBench}, the first benchmark specifically designed for evaluating LLM-based agents in diverse financial decision-making contexts. InvestorBench enhances the versatility of LLM-enabled agents by providing a comprehensive suite of tasks applicable to different financial products, including single equities like stocks, cryptocurrencies and exchange-traded funds (ETFs). Additionally, we assess the reasoning and decision-making capabilities of our agent framework using thirteen different LLMs as backbone models, across various market environments and tasks. Furthermore, we have curated a diverse collection of open-source, multi-modal datasets and developed a comprehensive suite of environments for financial decision-making. This establishes a highly accessible platform for evaluating financial agents' performance across various scenarios.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

    cs.CL 2026-07 conditional novelty 7.0

    Frontier LLMs score about 87-88% on rubric-graded, open-ended business case questions but fully satisfy every rubric criterion on fewer than half of them.

  2. Frontier AI performance across the business disciplines: a case-grounded benchmark of knowledge work and analytical reasoning

    cs.CL 2026-07 conditional novelty 7.0

    On a new 615-question business-case benchmark graded by AI against instructor rubrics, frontier LLMs score 87-88% partial credit but complete only about half the questions.

  3. CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

    cs.AI 2026-06 unverdicted novelty 7.0

    CLQT is a new closed-loop, cost-aware benchmark that diagnoses LLM trading agent capabilities through strategy-consistent metrics and hash-verifiable trails rather than outcome rankings.

  4. InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy

    cs.AI 2026-06 unverdicted novelty 7.0

    InvestPhilBench is a new multi-layer benchmark for LLM procedural reasoning in investment philosophy, with BASP metrics showing composite scores saturate while gate reconstruction accuracy reveals procedural deficits.

  5. InvestPhilBench: A Multi-Layer Benchmark for Evaluating Large Language Model Procedural Reasoning in Expert Investment Philosophy

    cs.AI 2026-06 conditional novelty 7.0

    Composite LLM scoring saturates on expert investment frameworks while Gate Reconstruction Accuracy still reveals a clear procedural deficit at the frontier.

  6. ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?

    cs.AI 2026-08 conditional novelty 6.0

    Across five 100-task agent streams, sequential experience improves normalized reward by 16.9% in 14 of 15 model-domain combinations, but explicit skill maintenance matches pure in-context learning (0.602 vs 0.605) and...

  7. Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable

    cs.CL 2026-07 conditional novelty 6.0

    CM-LRS is a workflow-output-layer reliability score for capital-markets LLM outputs; on five public-document workflows, frontier closed-source models score 4.09–4.31 and Llama 3.3 70B scores 3.15 under four LLM judges.

  8. NextFund: A Unified Performance Tracking Platform for Agentic Portfolio Management

    cs.AI 2026-07 conditional novelty 6.0

    NextFund unifies live multi-market data, multi-agent portfolio reasoning, and end-to-end decision traces with an interactive Trading Arena for fairer LLM agent benchmarking.

  9. CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning

    cs.CL 2026-06 unverdicted novelty 6.0

    CLExEval introduces a human-annotated evaluation framework on 40 rare cases that identifies verbosity bias, hidden knowledge paradox, and 68.6% reasoning-to-output mismatch in LLMs while showing LLM-as-a-Judge overest...

  10. CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

    cs.AI 2026-06 conditional novelty 6.0

    A closed-loop trading benchmark with a five-axis capability scorecard shows that LLM agents' returns rank poorly against diagnostic measures of their reasoning, with the nominal Sharpe winner exposed as a reliability ...

  11. FinDocMRE: A Benchmark for Document-Level Financial Multimodal Reasoning Evaluation

    cs.CE 2026-05 unverdicted novelty 6.0

    FinDocMRE is a new multi-image document-level benchmark spanning 12 financial domains and 5 task types, showing that 11 tested LMMs all score below 65 overall with particular weaknesses in numerical estimation and cro...

  12. QRAFTI: An Agentic Framework for Empirical Research in Quantitative Finance

    cs.MA 2026-04 unverdicted novelty 6.0

    QRAFTI is a multi-agent framework using tool-calling and reflection-based planning to emulate quant research tasks like factor replication and signal testing on financial data.

  13. Emergent Social Intelligence Risks in Generative Multi-Agent Systems

    cs.MA 2026-03 unverdicted novelty 5.0

    Generative multi-agent systems exhibit emergent collusion and conformity behaviors that cannot be prevented by existing agent-level safeguards.

  14. Leakage-Aware Benchmarking of LLM Forecasting: Real-Time Nowcasts as the Decision-Time Input for Macro Factor Ranking

    q-fin.ST 2026-06 unverdicted novelty 3.0

    Leakage-controlled LLM factor ranking yields median Spearman IC of +0.154 that is largely matched by a kNN baseline on the same real-time macro inputs.

  15. Beyond Agent Architecture: Execution Assumptions and Reproducibility in LLM-Based Trading Systems

    cs.AI 2026-06 unverdicted novelty 3.0

    Reproducibility audit of 30 LLM trading papers shows execution assumptions under-reported relative to agent architectures, illustrated by a 10-equity example where frictions compress returns.

  16. A Survey of Scaling in Large Language Model Reasoning

    cs.AI 2025-04 unverdicted novelty 3.0

    A survey categorizing scaling in LLM reasoning across input size, steps, rounds, training, and future directions, noting that scaling can negatively affect performance.