Pith. sign in

Paper Citation Record · LEDGER

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models

As of 12 August 2026, this Paper Citation Record lists 5 of 5 outbound references and 5 inbound Pith citation observations for arXiv:2501.18062.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18062 v1

Coverage vector

measured 5 of 5 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:53:29.689406Z

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:01:16.726294Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:56:25.216379Z

Reference resolution

5 of 5 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e72439a9-86fe-476f-aab3-5a609cab0709 · outbound

This paper cites FinQA: A Dataset of Numerical Reasoning over Financial Data.

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models FinQA: A Dataset of Numerical Reasoning over Financial Data

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T00:53:29.668745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:53:29.668745Z digest=sha256:0328e2c583129a1c2cb88c201c4c45dc36b61ba9028ff436a31ca195d831c5c5

Observation 474df27b-e50d-4e34-bc2a-f91db76556c0 · outbound

This paper cites FinanceBench: A New Benchmark for Financial Question Answering.

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models FinanceBench: A New Benchmark for Financial Question Answering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T00:53:29.678903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:53:29.678903Z digest=sha256:60bed2ecc634b1a109836762f9bb31a971056eade175ce3b7495308343f16021

Observation 2a1c8f9d-32ed-492d-a6d9-ba31d633346c · outbound

This paper cites A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law.

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-10T00:53:29.689406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:53:29.689406Z digest=sha256:7782878c39afb896d8c964e59a6ef30e6fda06821a2746e0ce4dd9e382bc6524

Observation 8a430616-9e27-4c79-a038-68fdfff5f118 · outbound

This paper cites LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models.

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T00:53:29.674023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:53:29.674023Z digest=sha256:65528696342c51ab125b8d0d4a07a9685858e6510c8f8f059f7c3945823c31f9

Observation 7a0ac1f2-210a-47c6-9806-17bf56a4be43 · outbound

This paper cites On Leakage of Code Generation Evaluation Datasets.

FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models On Leakage of Code Generation Evaluation Datasets

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T00:53:29.683526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:53:29.683526Z digest=sha256:4be808ae80826c1cf7a77704ac65f1955a64242f15bc297678c299b3ab840deb

Pith citing papers

Observation b5cbb368-61f5-41b7-a401-2147d79c0faa · inbound

On Path to Multimodal Historical Reasoning: HistBench and HistAgent cites this paper.

On Path to Multimodal Historical Reasoning: HistBench and HistAgent FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:16.726294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:16.726294Z digest=sha256:c0c450eb51064034d41f9ec7c30054e2b21e849f578a4e3aaf5dee2b95e09bbe

Observation 2fbc8455-1593-4cd7-8fff-6dce4cc607e4 · inbound

Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents cites this paper.

Claw-Eval: Towards Trustworthy Evaluation of Autonomous Agents FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:35:51.605474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T18:59:59.383781Z digest=sha256:5b27e11fea04562cda128dc8360703679bfd902798257c8416617ecf8f927ca4

Observation 12905610-f501-471c-8f08-6dbd8eade104 · inbound

BizCompass: Benchmarking the Reasoning Capabilities of LLMs in Business Knowledge and Applications cites this paper.

BizCompass: Benchmarking the Reasoning Capabilities of LLMs in Business Knowledge and Applications FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:56:11.234156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T05:52:40.026883Z digest=sha256:cb4ff7448dc3adacdaf4c5600450520d5d9c2d4753ffa4a4e47dc95fb8833911

Observation 13123e87-942a-4f52-a066-c3c8fae96dff · inbound

LATTICE: Evaluating Decision Support Utility of Crypto Agents cites this paper.

LATTICE: Evaluating Decision Support Utility of Crypto Agents FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:56:25.221197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-07T13:30:46.523784Z digest=sha256:16165e4fa05db69f22a59ae3ef7ddf7cc28683abd8275c755d835642dc6b5ca1

Observation 63f6b8a0-ee55-448e-b67a-77b42fa61a60 · inbound

Fin-Bias: Comprehensive Evaluation for LLM Decision-Making under human bias in Finance Domain cites this paper.

Fin-Bias: Comprehensive Evaluation for LLM Decision-Making under human bias in Finance Domain FinanceQA: A Benchmark for Evaluating Financial Analysis Capabilities of Large Language Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:16:15.789703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-12T02:15:53.024591Z digest=sha256:a67831e6cf456455a5bf977c8b5f4e38e29b0f2d484f5eb00634831ca35fddb6