Pith. sign in

Paper Citation Record · LEDGER

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents

As of 22 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.00172.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00172 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:17:45.470828Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e3c96934-e189-4598-8b84-cb32a3653808 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.198051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.198051Z digest=sha256:517be7c72a1de56d070bb5c808f9e9c286663bc4a8a37864123ad73266919c0c

Observation edf39d07-da64-4c83-a760-260036ed7813 · outbound

This paper cites MARPLE: A Benchmark for Long-Horizon Inference.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents MARPLE: A Benchmark for Long-Horizon Inference

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:17:45.870546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T12:17:44.420152Z digest=sha256:93ef00d30c4f19a1c89f78cf9b9823a721009a5b3dd6fb0f56c8ca4a99791f8d

Observation 0db5c82e-6129-49d0-b34c-14192f4f7637 · outbound

This paper cites an unresolved cited work.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.778572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.778572Z digest=sha256:814c36656c179e9154c437100027de95b9dac284dc5f52f2aedde4fdf1d2a49e

Observation 9d2ff84f-c45c-443a-a488-d72430be49dd · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents AgentBench: Evaluating LLMs as Agents

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.888628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.888628Z digest=sha256:c61b88fe194e0a8c5962f4e9dff913d27f11ad2db6fbd852430842b2bdb0d345

Observation d664ad6a-ee01-489f-9aac-0f635ddb0cc3 · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents LLM Critics Help Catch LLM Bugs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.997683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.997683Z digest=sha256:0a9dbf14b8d3501a5647030512a9957e6610e38bd1c9b1e9c74d5892ffaf0ce9

Observation 2ff36b65-ec3a-4205-b6ec-6d73fc436d05 · outbound

This paper cites SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:45.190655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:45.190655Z digest=sha256:5510a205323f392190189fc3593601f99fef88a484baec55bf46f1020fd0b9a7

Observation 13c666ef-947e-41ce-be2f-5852f06b73e0 · outbound

This paper cites Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:45.278095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:45.278095Z digest=sha256:17fbed61f4bc3467ad96664636a0ebcbd6d21c8b9a75ec950abfb383f3814e0c

Observation ab9b37ec-79fd-46c5-afcc-461cd431d942 · outbound

This paper cites SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:45.374967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:45.374967Z digest=sha256:52012824b71538597ea32f8cdf9b27d203a7641851ab3cf422a80820502a59da

Observation 34b13d74-94aa-43d4-ae23-797e0a045e74 · outbound

This paper cites SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:45.122364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:45.122364Z digest=sha256:13a9b75db308c14ae96247e1cda91fbf3c719f8e5e93c50465919b9e625c6098

Observation 7ff258c3-e82b-42e2-a58b-d3e9d85fff9c · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents Solving Quantitative Reasoning Problems with Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.675554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.675554Z digest=sha256:cf94b2019843fc3677a00a1a5d1db314b63952865c1a98f09458bffebd176798

Observation 0eafdb86-590c-4383-b8d7-23735c733308 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents ReAct: Synergizing Reasoning and Acting in Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:45.470828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:45.470828Z digest=sha256:bdffb1de28c892b0c68348ae28526093b426decb36423d4b56a72df3a4a191ef

Observation 444f1190-f0e7-4f3b-b90d-ce3c4286dd53 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.287496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.287496Z digest=sha256:96b434e7477293f909a8e72d3ffc1df8a7d25d979c235d197692c7c7a51f5c11

Observation d8fe9f11-f709-4ea8-b24b-29bb36acf662 · outbound

This paper cites Measuring AI Ability to Complete Long Software Tasks.

Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents Measuring AI Ability to Complete Long Software Tasks

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:17:44.566544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:17:44.566544Z digest=sha256:1504d4356a5880d612bab73b3b041cfdcbb83b18e32f6405dca64a474b3ba730

Pith citing papers

No inbound Pith citation observations are available.