Pith. sign in

Paper Citation Record · LEDGER

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis

As of 9 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2607.18260.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.18260 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T13:53:04.260789Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f6492574-59b8-4d31-bc79-d3ef76b92028 · outbound

This paper cites an unresolved cited work.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:02.201912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:02.201912Z digest=sha256:af5aa28de0abe77b3789a2c1c1d8c0de076eaa2f309c0b9e9e6d2aace82f7809

Observation ddc26d80-a986-4b36-a98b-15442a2c33b5 · outbound

This paper cites Notices of the American Mathematical Society , year =.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis Notices of the American Mathematical Society , year =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:02.268546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:02.268546Z digest=sha256:59bcf35b5146459f84651eafff6c227b21f3ee36a79aa888f1f6c22101f43054

Observation ec655c0a-6d43-4d5c-9988-163541020be6 · outbound

This paper cites Mathematics in Computer Science , year =.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis Mathematics in Computer Science , year =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:02.359388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:02.359388Z digest=sha256:7b1050ad9d5f57b79063d4eb5f4194e54297f16c8354cd851359fd4261ba858c

Observation df237279-c32f-4cb1-be20-dab31291a370 · outbound

This paper cites Foundations and Trends in Programming Languages , year =.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis Foundations and Trends in Programming Languages , year =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:02.525440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:02.525440Z digest=sha256:7013120b8d4c2b0a6b2cd28e99c16193cc51d4826df6d486ecdd22c64cc87455

Observation c0254d3a-abb2-444a-825e-32862e5f3a83 · outbound

This paper cites 2021 , eprint =.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis 2021 , eprint =

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:02.620070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:02.620070Z digest=sha256:062362a15855457c2825a7c191aa6e6f027ae8b4af18614748889f5a1c762ae5

Observation a837078f-45ec-4b5a-b04c-841a41469f23 · outbound

This paper cites 2021 , eprint =.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis 2021 , eprint =

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:02.720666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:02.720666Z digest=sha256:eda5ce1b01f71dc466d16efb133361cb464fc712132a71bcaa0ff33b5f230d0d

Observation f06c28be-c323-4151-8228-a53268b3543f · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis Measuring Coding Challenge Competence With APPS

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:02.811964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:02.811964Z digest=sha256:a01d08d50d6652c46d544f3c0d1729da166cd2f5886638369a200e80e7097107

Observation 78be9fde-35fb-43a0-afef-fd4c1bf21bcb · outbound

This paper cites and Yu, Tao , booktitle =.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis and Yu, Tao , booktitle =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:02.900914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:02.900914Z digest=sha256:b585a0cba348ef26606302ff220802a9fa5d62d92b008d2020fb6601d509ff58

Observation d62ff8fa-f42b-46d6-905b-25f4c692b26a · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:03.007019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:03.007019Z digest=sha256:e273fdec9234e7294c51c3bf7ff687115c081b893dd5b2b64305ee0bac9ac628

Observation 8f55fe7d-dd95-472a-aad1-6e788a427f88 · outbound

This paper cites Competition-level code generation with.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis Competition-level code generation with

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:03.122196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:03.122196Z digest=sha256:9738118bd009b3b1b210b84bbaf42acbd978780029fd8a27d519e36efe58c622

Observation aa93c4b2-4f74-4570-8359-570762374dac · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:03.155244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:03.155244Z digest=sha256:be2160809121be09a25b3152ab0b9aeadff640d9c2465a7784bb3105f17f0db3

Observation 9fb3a808-6dc0-4b5d-9a08-8c40b4ec61e9 · outbound

This paper cites CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis CodeXGLUE: A Machine Learning Benchmark Dataset for Code Understanding and Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:03.229863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:03.229863Z digest=sha256:03cf6f7c9fd97bae692d3e176596185c7b7ac9043681115d08f9797c2d5c2a69

Observation d612cc47-3909-41de-b7e0-6c70d99a0fa7 · outbound

This paper cites MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis MultiPL-E: A Scalable and Extensible Approach to Benchmarking Neural Code Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:03.342870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:03.342870Z digest=sha256:2ced8453e5f6c9e3deda753913dd0c42637a4fb2db9217e68d0a81735a2afc27

Observation f9bc12ab-3607-4df5-9788-d0b7c07f46f1 · outbound

This paper cites Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics , year =.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics , year =

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:03.462855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:03.462855Z digest=sha256:ce00783535959ae7881da21e294e772445d163dd2a1c19611ddbdd6d02617579

Observation 70a700fc-ce84-483e-9e2d-b91be1c6ea4a · outbound

This paper cites 2019 , volume =.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis 2019 , volume =

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:03.631101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:03.631101Z digest=sha256:5c54292e41cb523197d8f95b31bb76cdb58b03c61694918b05fe4a62b38a78fc

Observation bcb5b616-d203-4c83-b7bd-5fc62da92ebe · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:03.780302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:03.780302Z digest=sha256:1f507d68fa7ecc2d4fb8410b4f46e81521d36405b008712f5eec3c090db874cc

Observation d0e71e36-f0bd-4e85-b363-8f1f98f49ce5 · outbound

This paper cites MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis MiniF2F: a cross-system benchmark for formal Olympiad-level mathematics

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:03.836999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:03.836999Z digest=sha256:b7352f3e8544a9161f90b12548a753a858fc4c8ff7cd0a8ce5ac8667430b5387

Observation b345c434-eecb-4b91-8189-77adedb24b7a · outbound

This paper cites Qwen2.5-Coder Technical Report.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis Qwen2.5-Coder Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:03.983462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:03.983462Z digest=sha256:60af626877db85a052c2985e9aebea81a7f7e4671dbac0761b5d6502b877ce95

Observation ddd997c8-3ceb-4db1-83b3-d66d8459ef9e · outbound

This paper cites The Llama 3 Herd of Models.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis The Llama 3 Herd of Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T13:53:04.111192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:04.111192Z digest=sha256:8c8b6fc956d107753962b47ab1ba4b81d9ec3610cb54ca5f149ab2284e64ca5f

Observation bdd718c2-ec26-4a05-a36e-f79141b069ac · outbound

This paper cites an unresolved cited work.

FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis Unresolved cited work

Reference 20

Resolution
parse uncertain
no resolver link, observed 2026-08-02T13:53:04.260789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T13:53:04.260789Z digest=sha256:b11709b16b46b5f3ba18353c2ed0668bd9c546ae7546534ede32186297d59dd0

Pith citing papers

No inbound Pith citation observations are available.