Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:09:08.945549Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2608.03764.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T13:09:08.945549Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f18c76b0-c63b-49c9-a1cd-f674f1596b8a · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence.Transactions on Machine Learning Research, 2026
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 323bfd6f-ea54-47e9-ac69-82a1372d9fb0 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks A comprehensive survey of continual learning: Theory, method and application.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46:5362–5383, 2024
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1181aecd-ac95-4f6c-a8f9-eadee7eb2edf · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Reflexion: Language 10 agents with verbal reinforcement learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation fcbed200-7f4a-4a32-b7ea-16b99f87d0d8 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks ExpeL: LLM agents are experiential learners
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dd4e9245-4fe1-43c6-956b-3a4fb1d012ac · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks G¨ odel machines: Fully self-referential optimal universal self-improvers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 14365f4e-6c3d-4859-88e1-193e66514baf · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e58e4473-b390-4895-91ea-d33b71e320b1 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b973ed99-b7d1-436c-ad94-e6aae7b54064 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f7172b8-3e07-4067-818d-4c13b2c0a6df · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bba5e16-9abd-448f-9f42-095f16815c49 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0bff3dce-b501-4881-8b3c-7e55c52d112a · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b890d3ec-6d7e-4343-9ca9-a8ffe333cb3c · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9f14d8b9-53a5-414a-8c41-9eeb6008326e · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00f87561-2eb7-4b66-b1b7-906fecd4cf7b · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Kim, Patrick Chao, Samuel Miserendino, Gildas Chabot, David Li, Michael Sharman, Alexandra Barr, Amelia Glaese, and Jerry Tworek
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dea83764-8c4c-4cdf-91a4-e9634c9da93c · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks SOP-Bench: Complex industrial SOPs for evaluating LLM agents.CoRR, abs/2506.08119, 2025
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d38aca4-dcf1-4b29-a09e-2df355d7f387 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks JobBench: Aligning agent work with human will.CoRR, 2026
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b54a4f99-b720-4dd9-8b09-1b45e80194d1 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks LiveBench: A challenging, contamination-limited LLM benchmark
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d3d0c413-acb6-4dc0-9ffa-a2d08416ec68 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks AntiLeakBench: Preventing data contamination by automatically constructing benchmarks with updated real-world knowledge
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 02b81de1-9fb0-4369-a57c-d9e5093eda5e · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Benchmarking large language models under data contamination: A survey from static to dynamic evaluation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 466b1304-ad71-412e-afaa-f12d86c8e29f · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Voyager: An open-ended embodied agent with large language models.Transactions on Machine Learning Research, 2024
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 432c6128-8d9b-467c-aa7e-f636ec66fb39 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b224103-8426-4d03-94a3-98f4b0c21057 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c020f1da-c0bf-453c-9e38-3b5777b6d402 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks LLMigrate: Transforming "Lazy" Large Language Models into Efficient Source Code Migrators
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddaf9f13-1905-4dc5-8f81-6bc041788c33 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7a5f93ed-3540-4308-b8cd-f2696d2e7fb1 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Reward Is Enough: LLMs Are In-Context Reinforcement Learners
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42494414-8e2c-40b5-8a79-e309ce44dc0a · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 73bf45cd-9e53-4a20-993e-0d8569b9efd5 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8e6ed68-1157-4b7b-923c-3502c1b24722 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d33b646e-a2ec-41a9-8144-9e8dea90379c · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61c8d1c8-f899-4566-9e1c-5e74ed2fafb2 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Pollard, Alistair Johnson, Edward Choi, Yugang Jia, and Jong Ha Lee
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3e56211-5923-4e10-ae1c-c9a8c05b391c · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Harvey LAB: The legal agent benchmark
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9fb19912-667d-41b2-ae54-ad01b4e6a465 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks SpreadsheetBench: Towards challenging real world spreadsheet manipulation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8ad928f2-3a8f-4e0d-82bf-0e40e05a660a · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks InfiAgent-DABench: Evaluating agents on data analysis tasks
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d0420d3e-d12a-4e27-afe8-3df81dc3b39d · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks BIRD-INTERACT: Re-imagining text-to-SQL evaluation for large language models via lens of dynamic interactions.CoRR, abs/2510.05318, 2025
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 927f1840-81d4-4634-84c2-ddc6ccafb9b2 · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks LiveSQLBench: A dynamic and contamination-free benchmark for evaluating LLMs on real-world text-to-SQL tasks.https://github.com/bird-bench/livesqlbench, 2025
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 092207ef-8fef-46f4-ad6b-2ef6e2675dbe · outbound
GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.