Pith. sign in

Paper Citation Record · LEDGER

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents

As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.00172.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00172 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:17:45.470828Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e3c96934-e189-4598-8b84-cb32a3653808 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.198051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.198051Z digest=sha256:8ccdbf740a6a084698ada2ea03c7f51a97a6e987568b0ef5372d5bd54883769d

Observation edf39d07-da64-4c83-a760-260036ed7813 · outbound

This paper cites MARPLE: A Benchmark for Long-Horizon Inference.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents MARPLE: A Benchmark for Long-Horizon Inference

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:17:45.870546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T12:17:44.420152Z digest=sha256:9ca0de453c86bd6e796d0b24c7a367f694d047f55f1c5ab776f47054eb4346a3

Observation 0db5c82e-6129-49d0-b34c-14192f4f7637 · outbound

This paper cites an unresolved cited work.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.778572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.778572Z digest=sha256:09d9a02e207b3ef851b7a98df626ac8adba702192470d04026c1fd263a1e1afd

Observation 9d2ff84f-c45c-443a-a488-d72430be49dd · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents AgentBench: Evaluating LLMs as Agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.888628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.888628Z digest=sha256:13852399eb69bcf3227e34c06a891dd575153548097a2c3ae79938c47aac7188

Observation d664ad6a-ee01-489f-9aac-0f635ddb0cc3 · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents LLM Critics Help Catch LLM Bugs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.997683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.997683Z digest=sha256:131689a08bbdafb053b3e4b5793e33581c61eb70f7f4ab391f4151c15c043e40

Observation 2ff36b65-ec3a-4205-b6ec-6d73fc436d05 · outbound

This paper cites SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:45.190655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:45.190655Z digest=sha256:bff128f1c6e75e63b8f424b13ef5fe53e250f250dce24f155bc7ee8a05226eee

Observation 13c666ef-947e-41ce-be2f-5852f06b73e0 · outbound

This paper cites Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:45.278095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:45.278095Z digest=sha256:79e22dfa2b585d42db438efb884d14c3aa4345a22225bd6abe7f08fe1a361f4f

Observation ab9b37ec-79fd-46c5-afcc-461cd431d942 · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:45.374967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:45.374967Z digest=sha256:489ce4788a26e1883b5dd2b96b8723960ba2927f9be8794e2cd56176a839768d

Observation 34b13d74-94aa-43d4-ae23-797e0a045e74 · outbound

This paper cites SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:45.122364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:45.122364Z digest=sha256:0765ad0c71c76b5639b0052d22a74f8b82c541ec77db345f9131844dd9571af2

Observation 7ff258c3-e82b-42e2-a58b-d3e9d85fff9c · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents Solving Quantitative Reasoning Problems with Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.675554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.675554Z digest=sha256:7c84ac684b07998b1b4854f14d8e2e384fc34827c102c49b8cadf94b47be7486

Observation 0eafdb86-590c-4383-b8d7-23735c733308 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents ReAct: Synergizing Reasoning and Acting in Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:45.470828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:45.470828Z digest=sha256:e1432b7bbc963d1f49027440a4e7036b4d6ab5f59578f22b27b59add42e5540b

Observation 444f1190-f0e7-4f3b-b90d-ce3c4286dd53 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.287496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.287496Z digest=sha256:a1581a5193825be59ae5cb51747c4a4902c78ba23aaf134c1c9687a343bea2f7

Observation d8fe9f11-f709-4ea8-b24b-29bb36acf662 · outbound

This paper cites Measuring AI Ability to Complete Long Software Tasks.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents Measuring AI Ability to Complete Long Software Tasks

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.566544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.566544Z digest=sha256:53b4db050be1be967408903371a789e97ec098e1f56e415793d2c4811e6abc8c

Pith citing papers

No inbound Pith citation observations are available.