Pith. sign in

Paper Citation Record · LEDGER

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies

As of 15 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 0 inbound Pith citation observations for arXiv:2604.15151.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.15151 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T11:33:01.976659Z

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

16 of 16 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8181e67f-9347-4597-b35a-d129c10f78b5 · outbound

This paper cites Swe-rebench: An automated pipeline for task collection and decontaminated evaluation of software engineering agents.

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies Swe-rebench: An automated pipeline for task collection and decontaminated evaluation of software engineering agents

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:57:20.537514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:33:01.976659Z digest=sha256:ffa66c67a1aff1f20b1cdf7b65782f4058939a005573332e7af29d8d3c94c122

Observation 4ab29d60-07a5-41fa-aa95-909960570f32 · outbound

This paper cites Finance agent benchmark: Benchmarking llms on real-world financial research tasks.

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies Finance agent benchmark: Benchmarking llms on real-world financial research tasks

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:57:20.554355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:33:01.976659Z digest=sha256:664a70952111227818169fa7e1b5a16a7d122b26f4cecde1a33c498194b81131

Observation 25f60a6f-3b31-48c0-a180-26c333628c2f · outbound

This paper cites Finagentbench: A benchmark dataset for agentic retrieval in financial question answering.

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies Finagentbench: A benchmark dataset for agentic retrieval in financial question answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:57:20.552522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:33:01.976659Z digest=sha256:c4a28e1a0f4458b5dff59fe204bff803f63395bd5d65505c28ffd7ea2249ea61

Observation 86be3f85-8a6d-4388-ab8f-2e829b0d2b77 · outbound

This paper cites A survey on LLM-as-a-judge.

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies A survey on LLM-as-a-judge

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:57:20.539304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:33:01.976659Z digest=sha256:2eb10832be0e003b50ff8e0f2cfdbbb967f51876ab6f1d086eedc629474b1348

Observation 14b8ba2f-89ec-4aa6-927f-2421ff4d4c18 · outbound

This paper cites Financebench: A new benchmark for financial question answering.

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies Financebench: A new benchmark for financial question answering

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:57:20.546659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:33:01.976659Z digest=sha256:a5ad44980248ff5326bc60123acec06d54add151e16cfec9c39bb3262a9f6aad

Observation 64d87908-f55b-4baa-827d-b2a16de76415 · outbound

This paper cites Livecodebench: Holistic and contamination-free evaluation of large language models for code.

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies Livecodebench: Holistic and contamination-free evaluation of large language models for code

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:57:20.543176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:33:01.976659Z digest=sha256:783f63c518c2b57dacf78931a7244f44f77e03531cbff6c6f19921d6dec7d254

Observation c30faa27-8454-47e4-bc8c-a4b5b2ef43ce · outbound

This paper cites Jimenez et al.

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies Jimenez et al

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:57:20.544835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:33:01.976659Z digest=sha256:963f4a41dc7d694a3b1abc6fece4e61061c0bee1761e61d9dffb006408f1c5bc

Observation dd4d40bd-a8a7-4b4d-a4e4-bde5a3be2cc3 · outbound

This paper cites Fin-r1: A large language model for financial reasoning through reinforce- ment learning.

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies Fin-r1: A large language model for financial reasoning through reinforce- ment learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:57:20.531578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:33:01.976659Z digest=sha256:d912fb03d88accfe6982c981a06c37b4e477035af12f92194e9dc34620fed262

Observation 0952186a-4f5e-4760-9a0f-70ab78a52170 · outbound

This paper cites Merrill et al.

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies Merrill et al

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:57:20.550300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:33:01.976659Z digest=sha256:3433c436ea29335272a53d894656577219702e9c41e3bed24649032509cf5158

Observation 24fe33f5-1339-41f2-ae36-ab3c58a749a3 · outbound

This paper cites Fino1: On the transferability of reasoning-enhanced llms and reinforce- ment learning to finance.

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies Fino1: On the transferability of reasoning-enhanced llms and reinforce- ment learning to finance

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:57:20.548595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:33:01.976659Z digest=sha256:7c5de8522023585cf3eac58d9cb651300debf2c89913a82309dddcc1be6508d1

Observation 728a2838-8fdc-48fe-a1a0-9c31077527f7 · outbound

This paper cites Quantconnect LEAN documentation.https://www.quantconnect.com/d ocs/.

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies Quantconnect LEAN documentation.https://www.quantconnect.com/d ocs/

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:57:20.529753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:33:01.976659Z digest=sha256:7a9a2d16acbc88b346349fd916f2732e2e1c0965ff9896f217242f15fc3a269e

Observation f339ff9c-b37a-4436-b75e-f6ff630f7060 · outbound

This paper cites Backtrader.https://www.backtrader.com/.

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies Backtrader.https://www.backtrader.com/

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:57:20.535562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:33:01.976659Z digest=sha256:07d2281d0523694b2aede3dbb7c49ce4c9e125506d45c9870af7b7c492980ced

Observation 61928b12-eb9b-4b77-a98a-a805cc371ff4 · outbound

This paper cites Pixiu: A large language model, instruction data and evaluation benchmark for finance.

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies Pixiu: A large language model, instruction data and evaluation benchmark for finance

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:57:20.525814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:33:01.976659Z digest=sha256:cb60aac5e895d871aafa4f2e9bc5f11a49e62e90426c29262b320fc603d001dc

Observation ef8317c5-7366-4cab-aae9-7b2eb45623e4 · outbound

This paper cites Finben: A holistic financial benchmark for large language models.

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies Finben: A holistic financial benchmark for large language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:57:20.527871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:33:01.976659Z digest=sha256:3b1b86649c676faeecb167512ab2175c3dd824661cafed96fe4ae541f84736c9

Observation 27f132f8-556c-4b64-ba77-b90ab4dbbd61 · outbound

This paper cites Judging LLM-as-a-judge with MT-bench and chatbot arena.

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies Judging LLM-as-a-judge with MT-bench and chatbot arena

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:57:20.541322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:33:01.976659Z digest=sha256:b280b5aa176f0817f9113e7ddc62a5867bd18db9c630616202253ee38a1acd72

Observation 668f6e72-613c-4219-9e63-4492b823d1ef · outbound

This paper cites Zipline documentation.https://zipline.ml4trading.io/.

QuantCode-Bench: A Benchmark for Evaluating the Ability of Large Language Models to Generate Executable Algorithmic Trading Strategies Zipline documentation.https://zipline.ml4trading.io/

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T13:57:20.533601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:33:01.976659Z digest=sha256:3266a43a735f1f17e48bdb2db55dd9fbb322e86dd6d925707cf97a1a10b87960

Pith citing papers

No inbound Pith citation observations are available.