Pith. sign in

Paper Citation Record · LEDGER

Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 30 inbound Pith citation observations for arXiv:2508.00828.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.00828 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 30 of 30 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T00:22:58.807578Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e3fb759d-a9ee-40c1-b8d5-49ca76f4aaf5 · inbound

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows cites this paper.

Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:48:38.318919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T22:43:48.618334Z digest=sha256:62031c150586167379cb7a11d1e770325ae73eefe0930211e1710e6f0897412f

Observation 9884bc28-e204-49d0-b0b1-909bc64e538c · inbound

FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use cites this paper.

FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T05:54:53.348743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:54:53.348743Z digest=sha256:6eedfa627b10e60702f672cd45b7826084b9c6c284389cc866b59eb5aa85d5ed

Observation 6284342d-4f24-46de-b654-f13aa26fe2ae · inbound

BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows cites this paper.

BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:03.144299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:55:49.453068Z digest=sha256:1047f49a7bf5147fe6b3eafa310f14bd416d66dadbe33275984d188a239b9fae

Observation e171615f-3b16-499a-baad-b554af3ff0d7 · inbound

BizCompass: Benchmarking the Reasoning Capabilities of LLMs in Business Knowledge and Applications cites this paper.

BizCompass: Benchmarking the Reasoning Capabilities of LLMs in Business Knowledge and Applications Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:56:11.236886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:52:40.026883Z digest=sha256:b4373fa288a3b701987211fe33729492f88a1c87281e4e6e17724968002b05b7

Observation 999dfa21-90aa-4b40-b381-cefcf3845150 · inbound

Complete Cyclic Subtask Graphs for Tool-Using LLM Agents: Flexibility, Cost, and Bottlenecks in Multi-Agent Workflows cites this paper.

Complete Cyclic Subtask Graphs for Tool-Using LLM Agents: Flexibility, Cost, and Bottlenecks in Multi-Agent Workflows Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T07:06:52.862114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T07:05:43.392997Z digest=sha256:09ab9c7c41627b7d7dcc4a5b4a26e394aa10efa7a1bd2e4b340b79b9cd4d765d

Observation b7e87a0c-cc3c-436e-987f-328675b5df4f · inbound

The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications cites this paper.

The Price of Agreement: Measuring LLM Sycophancy in Agentic Financial Applications Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:06:25.659426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T03:22:45.216764Z digest=sha256:a350da806c8704a3bed231846f5684005fe4bcd81c9e8eeb9f2e45c8238884a3

Observation 96a8db9d-9314-4862-88a8-186b0d4afc79 · inbound

LATTICE: Evaluating Decision Support Utility of Crypto Agents cites this paper.

LATTICE: Evaluating Decision Support Utility of Crypto Agents Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:56:25.160988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-07T13:30:46.523784Z digest=sha256:e550e9d1cd5fb508e2ee3f398db79c6d2a9f5909841512e227cc64176dad61f9

Observation 66724bfa-d850-4f92-9cef-f5d02839c0d4 · inbound

TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems cites this paper.

TacoMAS: Test-Time Co-Evolution of Topology and Capability in LLM-based Multi-Agent Systems Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T05:51:24.466508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T04:53:54.754878Z digest=sha256:470e58b601d9583acde88c5fa97026c8e8a7f12ae3d21a5f367ae6b3df7c3d20

Observation e65918f3-0ecb-430b-b823-d9ebea59eb22 · inbound

Herculean: An Agentic Benchmark for Financial Intelligence cites this paper.

Herculean: An Agentic Benchmark for Financial Intelligence Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:15:04.687390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T21:06:46.156943Z digest=sha256:4e5358326aaf1a624bfd63c262467f5db1342b43e4de5917209ea230928ad9f0

Observation 6ceb1d54-aa04-4b95-b713-35fb9571c04a · inbound

AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems cites this paper.

AgensFlow: A Coordination-Policy Substrate for Multi-Agent Systems Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:15:49.116976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T16:08:56.103382Z digest=sha256:fb4d5aa1832c6e2c132f50293aba7ef2cc5bd904fc9e5095614fab064dffacb8

Observation 45097e83-8b06-49ae-82ef-ab6f1d3c7248 · inbound

FinCom: A Financial Multi-Agent Demo with Disagree-or-Commit Deliberation cites this paper.

FinCom: A Financial Multi-Agent Demo with Disagree-or-Commit Deliberation Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:36:15.568412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T16:36:33.579549Z digest=sha256:0494996be3cce9b7630dc26e0ea967557149c1b42f59ce3cffa842338575d336

Observation ff2b5e76-4b6d-4462-a8f4-53e185784f01 · inbound

Absorbing Complexity: An Interaction-Native Knowledge Harness for Financial LLM Agents cites this paper.

Absorbing Complexity: An Interaction-Native Knowledge Harness for Financial LLM Agents Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:56:19.783133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:57:00.865047Z digest=sha256:5a7ccf81d16d94a014c6efc70b6fd2a5b6b3ac888eb3a41dccd0f78bfffb4227

Observation 431cc5cb-40b3-4d56-bbb2-8c087d4fe592 · inbound

Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions cites this paper.

Economy of Minds: Emerging Multi-Agent Intelligence with Economic Interactions Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:26:21.796442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:26:50.444070Z digest=sha256:7e173e827663936b6997376fb245272567941d84358af2b85a345525ea551107

Observation d787c9c0-45ef-42d1-887b-6eae9a6afb40 · inbound

BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents cites this paper.

BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:56:34.697572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T09:35:28.694416Z digest=sha256:ebafce9f63d1b63f5b7c824dd6b12e3adf267324e4d9c01f50cfcc05fa1b5ad2

Observation 79b3bbbc-9048-441d-b92a-6d3c79891d38 · inbound

Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments cites this paper.

Continual Learning Bench: Evaluating Frontier AI Systems in Real-World Stateful Environments Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:56:56.868237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T01:45:28.693098Z digest=sha256:207c7207538f78ad506f94a6aa893d01429318dc25037ac9853ea89a858bbc0b

Observation 83ed772e-a5ac-4a60-9eb7-446dda01e3ab · inbound

Flaws in the LLM Automation Narrative cites this paper.

Flaws in the LLM Automation Narrative Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.296220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T10:52:36.252919Z digest=sha256:f2fabab175a1769f07347745775dbf8a9591584abfae61f0596ff26f90be4bee

Observation 27832513-4dfe-43ab-8c5e-5bcde2e0dead · inbound

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application cites this paper.

Agentic Environment Engineering for Large Language Models: A Survey of Environment Modeling, Synthesis, Evaluation, and Application Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 161

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:50:48.360204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T09:46:30.702256Z digest=sha256:91c0bf0ce853e53d6787a91ad459da1fe726566a54c99e1314cf8fff70cc086b

Observation 21d6a224-e484-4044-bf87-3486d8f51bb1 · inbound

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility cites this paper.

AgentBeats: Agentifying Agent Assessment for Openness, Standardization, and Reproducibility Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:33.127177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T06:41:41.799596Z digest=sha256:9647574856e7848e6ed1731fc555b279e6f4ae31bf77a55c568de963316cda25

Observation 44df256c-fa33-45eb-9c09-2057f1d2f58c · inbound

Leakage-Aware Benchmarking of LLM Forecasting: Real-Time Nowcasts as the Decision-Time Input for Macro Factor Ranking cites this paper.

Leakage-Aware Benchmarking of LLM Forecasting: Real-Time Nowcasts as the Decision-Time Input for Macro Factor Ranking Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:45.685029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T09:14:31.166883Z digest=sha256:9b192ef328b757232ccfabaf977200034987fbc61fd4dedbc0ec3738d35b753b

Observation 2d8e58db-f1c9-4ab0-9afc-aff113de5772 · inbound

IPO Finance Agent: Benchmark of LLM Financial Analysts Beyond Finance Agent v2, with Automated Rubric Generation, on the SpaceX (SPCX) IPO cites this paper.

IPO Finance Agent: Benchmark of LLM Financial Analysts Beyond Finance Agent v2, with Automated Rubric Generation, on the SpaceX (SPCX) IPO Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:29:44.669903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:51:42.837201Z digest=sha256:0f5a9f6b9d8c54bae44b1f7c947d14fc72cb84aa139a9b8210f80b24a8af8d2b

Observation 2ed051a4-e875-47b9-b186-5aafe2fc6ebf · inbound

IPO Finance Agent: Benchmark of LLM Financial Analysts Beyond Finance Agent v2, with Automated Rubric Generation, on the SpaceX (SPCX) IPO cites this paper.

IPO Finance Agent: Benchmark of LLM Financial Analysts Beyond Finance Agent v2, with Automated Rubric Generation, on the SpaceX (SPCX) IPO Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:35.925456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-01T06:49:27.958422Z digest=sha256:1c87a5fb2b735c64f3358db88b7c1d008b66ede757aee686c5931610ffa8e921

Observation d041e6c9-afbe-4480-93b2-cd9adaf93e0c · inbound

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? cites this paper.

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:29:56.712681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:34:10.103638Z digest=sha256:a6c1c3ae66c4ccc92084728bdce7b7904303359942bfe1b00c74ce413c01fb42

Observation fdf62176-6edf-42ed-872f-5b928de1a3a2 · inbound

AI Trading: Evaluating Large Language Models for Technical Market Analysis cites this paper.

AI Trading: Evaluating Large Language Models for Technical Market Analysis Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:50.743589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:50.743589Z digest=sha256:caf0086807b5753bdb3f953759e9e6da7bfae666fe67f08a5f5708a68e3af5ad

Observation 61116e77-9e8f-44a3-84c3-56f372496f87 · inbound

FORCE-Bench: A Benchmark, Dataset, and Evaluation Harness for Agentic AI in Enterprise Finance cites this paper.

FORCE-Bench: A Benchmark, Dataset, and Evaluation Harness for Agentic AI in Enterprise Finance Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T07:28:28.105727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:28:28.105727Z digest=sha256:25031f15f7b7f27009f352af3df1ab6b8887ecb54f23ae39314cfb4a18ed2920

Observation 77ef7915-b7a4-4dfe-92d0-0931651c2cb0 · inbound

Frontier Financial Judgement: Can agents tell what might move a stock? cites this paper.

Frontier Financial Judgement: Can agents tell what might move a stock? Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T09:47:15.283030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:47:15.283030Z digest=sha256:36ac4736a2a83753538ce66261325a0c83064571d65ab6a4c100adbfd8b98b1a

Observation 7b9074c3-e3ab-4c76-a185-367336ed36bb · inbound

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation cites this paper.

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T01:16:42.610725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:16:42.610725Z digest=sha256:9842fe8f081f1c38b0cf48840902ef058fa7f5ddc9aed606ba629a2300840d90

Observation f449c170-048b-4f46-b548-ab29ec9b13d4 · inbound

AgentRadio: Passive Awareness for Long-Horizon Multi-Agent Collaboration cites this paper.

AgentRadio: Passive Awareness for Long-Horizon Multi-Agent Collaboration Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-07-31T07:38:10.220810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T07:38:10.220810Z digest=sha256:cef7fb433de77360c87fb28e3980fd0357dad5258172f24e42b4d81a43d09387

Observation 0b3933cf-56e0-441f-b9ee-af1d8ca75f4c · inbound

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements cites this paper.

Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T00:48:02.213824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:48:02.213824Z digest=sha256:d3f5770ee1e5e9425dad500a17d1aaae2ff4d29fd3dd5a46c5ec406e8d4a5024

Observation 9f015896-0a35-463c-a30a-23394db146bc · inbound

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? cites this paper.

AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks? Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-04T01:12:38.779498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:12:38.779498Z digest=sha256:b9261c1f622286b57741f23c3529b4290069080ed9ffc0000dcb4f080a528de2

Observation 19039dca-eb3f-4fa2-ae44-178c439954cd · inbound

FinDeepIndicator: Benchmarking Deep Research Agents in End-to-End Financial Indicator Construction cites this paper.

FinDeepIndicator: Benchmarking Deep Research Agents in End-to-End Financial Indicator Construction Finance Agent Benchmark: Benchmarking LLMs on Real-world Financial Research Tasks

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T00:22:58.807578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T00:22:58.807578Z digest=sha256:35922d9cc4b1084438004c7fd355dd54954c8110cc23694117a332f196cda6fe