Pith. sign in

Paper Citation Record · LEDGER

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks

As of 4 August 2026, this Paper Citation Record lists 22 of 22 outbound references and 1 inbound Pith citation observation for arXiv:2604.10015.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.10015 v2

Coverage vector

measured 22 of 22 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T16:54:02.556652Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T11:50:20.718612Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-02T12:16:14.744948Z

Reference resolution

22 of 22 outbound references displayed

  • verified exact13
  • verified fuzzy5
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 260fbb84-597f-4bb1-83ed-81e998480ddd · outbound

This paper cites Agentpro: Enhancing llm agents with automated process supervision.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks Agentpro: Enhancing llm agents with automated process supervision

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:01:47.713243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:4f4639ec2d4491b24b1a45f2ae615b58f7ab30352e63dd4b52d38be8ae1e83f3

Observation e4c112c7-39a4-4959-bb7a-d90576cb9100 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks LoRA: Low-Rank Adaptation of Large Language Models

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T07:56:00.229236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:939213fb534fdc656278a90b6c8c61ee18a37a2419a24a2134409efdd7c602d8

Observation ebd6c290-d156-4144-a74d-1f03c421c7d5 · outbound

This paper cites FinanceBench: A New Benchmark for Financial Question Answering.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks FinanceBench: A New Benchmark for Financial Question Answering

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:01:13.018607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:2103db860dda59fac84984485d57474a14b1cfd0c1beef78ce137a924a741fc1

Observation 8835f659-29b3-4476-876e-de9178f4025c · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:56:00.163639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:7b4fa2ce18243bd4e2e37f36d82570d23e8c85c6b116c7b41e1ab698103cef12

Observation ee3e1edf-a071-439d-bcd1-75dfa5024cbd · outbound

This paper cites URL https://aclanthology.org/2025.sigdial-1.32/.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks URL https://aclanthology.org/2025.sigdial-1.32/

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:01:47.715889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:9746e74b9d7424deda720e82e08b3364d7ae5abcdc1d0c88b263e4edbb0331b7

Observation bf1952a5-4c2f-4bad-9013-2157c60342ae · outbound

This paper cites LongFuncEval: Measuring the effectiveness of long context models for function calling.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks LongFuncEval: Measuring the effectiveness of long context models for function calling

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:56:00.133797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:febee09fbf2f9716fd94b00e7a7a282a5ddd13ae12f7e2edfb236f7f205873e2

Observation 96d334b5-1a43-49d3-8d27-26bd1daaea1d · outbound

This paper cites Api-bank: A comprehensive benchmark for tool-augmented llms.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks Api-bank: A comprehensive benchmark for tool-augmented llms

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:01:47.710659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:c7575b96e52ac87f4d5c8a85d805af790f2a8a4ad8243d80ec8225d42dbae76c

Observation fc89a28a-8a11-4c57-b0a8-1f4d8f054089 · outbound

This paper cites Findeepforecast: A live multi-agent system for benchmarking deep research agents in financial forecasting.arXiv preprint arXiv:2601.05039.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks Findeepforecast: A live multi-agent system for benchmarking deep research agents in financial forecasting.arXiv preprint arXiv:2601.05039

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:56:00.284676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:88e9b450a2cd14d151c1b1494d24d3a11f6f9898eddd79d3296b6d325ae92a44

Observation 6be4d553-fe58-4197-b3c4-3b7ab3f41f4a · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks AgentBench: Evaluating LLMs as Agents

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:40:05.312349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:152e577d4ba2436054ff41a8f8068b0f507d9cadc6e520f1d03612706819682a

Observation 4a191611-72ef-4637-b4c6-2490531d987d · outbound

This paper cites FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-08-04T02:24:30.509855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:c759e4a538c2e7fdc929b240fee8cd739d0f159a66aa46738b5847c49aa55b2c

Observation 38a83bc3-3063-4c26-8791-1141063cf9e5 · outbound

This paper cites Mo et al.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks Mo et al

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:56:00.273124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:dbf4e781a5514709700975a3f3a7ce196c5e35203d7318bb7669db5c54b8f84d

Observation a6dbd9d4-3b38-4bff-9c59-b5101b7448a3 · outbound

This paper cites ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:56:00.168307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:073b6a9fc518980fdf58ddea5824ee43313ea3acc4d9d9717766f58d39eda7ec

Observation 8461ee18-e8f4-4aa3-8073-351a79d86e96 · outbound

This paper cites ToolGen: Unified Tool Retrieval and Calling via Generation.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks ToolGen: Unified Tool Retrieval and Calling via Generation

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T07:56:00.303273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:545bd0e93214227e35eb9a30688206b2a0a5c045ccc791f8ce3776b1245feec2

Observation c5ef217e-7c33-4b4a-ba8f-63e935a58782 · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:48:26.603034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:b6e84c48d706eaef0cd72660adee2212232f0bf6bb22df30a7399f769d11e7e1

Observation 9d2d87d5-1704-44fb-8600-4bf0b64e7b3c · outbound

This paper cites TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:39:31.128039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:707b71cb53cdd908be3e8dc08cf25ac5cf3e51478db93adba04ef9ffdbc0aea8

Observation 51fc605c-66ab-4213-a560-148c7bd38a92 · outbound

This paper cites The evolution of tool use in LLM agents: From single-tool call to multi-tool orchestration.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks The evolution of tool use in LLM agents: From single-tool call to multi-tool orchestration

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:56:00.239263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:bedfa4d115134872c68833a5cc37b3ffe7b974207ad5a8ed982aa35651e17ba7

Observation 9a1b507f-723e-427c-b929-72e945a0e709 · outbound

This paper cites Survey on Evaluation of LLM-based Agents.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks Survey on Evaluation of LLM-based Agents

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-11T07:56:00.141516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:9888bc96c1a76096c4eea973884cdf8ca39eb9633859a36000615d978ebe6f40

Observation 1f856b85-f1a0-4b18-9b63-feb14b45fb16 · outbound

This paper cites A Appendix A.1 Limitations Deep research agents are advancing at a fast pace.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks A Appendix A.1 Limitations Deep research agents are advancing at a fast pace

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:56:00.290237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:5e93fab2819f7429f25e4bc1145ec6644354852c6faa0653f43c76d6ca761ee7

Observation c24fff9f-11ff-4db7-b7b7-936d60c3fde5 · outbound

This paper cites Finmcp-bench: Benchmarking llm agents for real-world financial tool use under the model context protocol.arXiv preprint arXiv:2603.24943.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks Finmcp-bench: Benchmarking llm agents for real-world financial tool use under the model context protocol.arXiv preprint arXiv:2603.24943

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:56:00.158751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:5284b86fb1bfe0040bc211c42664f393828b9e0007016fbcdd27f3a69f736600

Observation 6fb9238b-ed82-432c-b347-7670ef70957d · outbound

This paper cites Under review.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks Under review

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:01:47.720840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:019ce643c32b10be93f6492061ab4d1f42f29e990c4bc881766bdf53b71984a4

Observation 297d3393-5398-45d5-a8f2-8557dc74e152 · outbound

This paper cites More recently, benchmarks have shifted toward evaluating LLM-based agents’ autonomous behavior in complex, interactive envi- ronments.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks More recently, benchmarks have shifted toward evaluating LLM-based agents’ autonomous behavior in complex, interactive envi- ronments

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T13:01:47.718506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:eefe55fdbca30531160a598e8e912ce9594647e2ab10a929a1093f36db07aac8

Observation 5c207545-2383-41f4-b22a-a49d2767868d · outbound

This paper cites {name}:{description}.

FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks {name}:{description}

Reference 22

Resolution
malformed identifier
raw_fallback, observed 2026-05-17T13:01:47.708208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T16:54:02.556652Z digest=sha256:7be51f5e5356f0fc30931ca03ee0602464ff1609579d0b9349198b588495cbd1

Pith citing papers

Observation 56071d13-9f62-4da9-8248-af9d8e5c6bd6 · inbound

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making cites this paper.

DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making FinTrace: Holistic Trajectory-Level Evaluation of LLM Tool Calling for Long-Horizon Financial Tasks

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-02T11:54:15.255546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-08-02T11:50:20.718612Z digest=sha256:844a0f67bde7c7b307c617333b69f58c9ab4fc957e903b9a87df32a65b477a74