Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:17:45.470828Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.00172.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:17:45.470828Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
13 of 13 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e3c96934-e189-4598-8b84-cb32a3653808 · outbound
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edf39d07-da64-4c83-a760-260036ed7813 · outbound
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents MARPLE: A Benchmark for Long-Horizon Inference
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0db5c82e-6129-49d0-b34c-14192f4f7637 · outbound
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d2ff84f-c45c-443a-a488-d72430be49dd · outbound
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents AgentBench: Evaluating LLMs as Agents
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d664ad6a-ee01-489f-9aac-0f635ddb0cc3 · outbound
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents LLM Critics Help Catch LLM Bugs
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ff36b65-ec3a-4205-b6ec-6d73fc436d05 · outbound
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13c666ef-947e-41ce-be2f-5852f06b73e0 · outbound
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab9b37ec-79fd-46c5-afcc-461cd431d942 · outbound
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34b13d74-94aa-43d4-ae23-797e0a045e74 · outbound
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ff258c3-e82b-42e2-a58b-d3e9d85fff9c · outbound
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents Solving Quantitative Reasoning Problems with Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0eafdb86-590c-4383-b8d7-23735c733308 · outbound
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents ReAct: Synergizing Reasoning and Acting in Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 444f1190-f0e7-4f3b-b90d-ce3c4286dd53 · outbound
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8fe9f11-f709-4ea8-b24b-29bb36acf662 · outbound
Breakpoint: Scalable evaluation of system-level reasoning in LLM code agents Measuring AI Ability to Complete Long Software Tasks
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.