Pith. sign in

Paper Citation Record · LEDGER

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks

As of 19 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2608.03764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03764 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:09:08.945549Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact6
  • verified fuzzy16
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f18c76b0-c63b-49c9-a1cd-f674f1596b8a · outbound

This paper cites A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence.Transactions on Machine Learning Research, 2026.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence.Transactions on Machine Learning Research, 2026

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:13.404964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:05.884507Z digest=sha256:3c9a4ec04152c81bafc0055e411f0bf3a40495ecf14b10437d527d7dc91bee81

Observation 323bfd6f-ea54-47e9-ac69-82a1372d9fb0 · outbound

This paper cites A comprehensive survey of continual learning: Theory, method and application.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46:5362–5383, 2024.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks A comprehensive survey of continual learning: Theory, method and application.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46:5362–5383, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:13.187967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:06.043829Z digest=sha256:05e52bf95119221f27cd4bdaddc04530d394f775050cd43a2b0bbd50d59e9446

Observation 1181aecd-ac95-4f6c-a8f9-eadee7eb2edf · outbound

This paper cites Reflexion: Language 10 agents with verbal reinforcement learning.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Reflexion: Language 10 agents with verbal reinforcement learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:12.995487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:06.138087Z digest=sha256:512f3dcaaa7b341977dbc3f14fe786c3c8d788c779a39210ebdcada65b3c796f

Observation fcbed200-7f4a-4a32-b7ea-16b99f87d0d8 · outbound

This paper cites ExpeL: LLM agents are experiential learners.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks ExpeL: LLM agents are experiential learners

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:12.812322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:06.226363Z digest=sha256:aef246fef77e19aaac6f320a351597e3b7aba45a849989aa362808f42ffdb7bf

Observation dd4e9245-4fe1-43c6-956b-3a4fb1d012ac · outbound

This paper cites G¨ odel machines: Fully self-referential optimal universal self-improvers.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks G¨ odel machines: Fully self-referential optimal universal self-improvers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:12.636324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:06.333375Z digest=sha256:7a9d71c2fdd3a2b224f379db3a1c2db058f7664225643613b8f76031e6c703e9

Observation 14365f4e-6c3d-4859-88e1-193e66514baf · outbound

This paper cites Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:06.410037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:06.410037Z digest=sha256:620cd6b67af3db7c77ffad74975bc6f9f13d7c326b66f093dcf79553ac0145f0

Observation e58e4473-b390-4895-91ea-d33b71e320b1 · outbound

This paper cites EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:09:10.439914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:06.494394Z digest=sha256:11e2aa8641f73714bb4c86907426e2a9240a480d7396a38ee8c9daabf6759847

Observation b973ed99-b7d1-436c-ad94-e6aae7b54064 · outbound

This paper cites SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:06.622558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:06.622558Z digest=sha256:33bbc1a80bdb0e6265b44126a533e92320adfd73240db27dca0814bd7bfd32b9

Observation 0f7172b8-3e07-4067-818d-4c13b2c0a6df · outbound

This paper cites SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:06.712982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:06.712982Z digest=sha256:c60b8160e63daf7594809bc93785cf05d835da045ff5f1233e436ef8bddaf829

Observation 7bba5e16-9abd-448f-9f42-095f16815c49 · outbound

This paper cites EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:09:10.241179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:06.823752Z digest=sha256:ce4284567442d81f063b7bf6cabc0dec4212e938d5fe9736afa5f44bfe4c118c

Observation 0bff3dce-b501-4881-8b3c-7e55c52d112a · outbound

This paper cites BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:09:10.035784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:06.907341Z digest=sha256:66413197eddf7f89335f10b7c4112d03987312605e9618f5d46ba96634169d61

Observation b890d3ec-6d7e-4343-9ca9-a8ffe333cb3c · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:12.444324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:06.986986Z digest=sha256:0c4006a49882df3d5a018e5be41a0524a0566236c83707236b05bd8db8ce24b1

Observation 9f14d8b9-53a5-414a-8c41-9eeb6008326e · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:07.068377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:07.068377Z digest=sha256:a310913c2790f0c7a538b446ff62698c355d2230d909d53b9d173d7049fae61d

Observation 00f87561-2eb7-4b66-b1b7-906fecd4cf7b · outbound

This paper cites Kim, Patrick Chao, Samuel Miserendino, Gildas Chabot, David Li, Michael Sharman, Alexandra Barr, Amelia Glaese, and Jerry Tworek.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Kim, Patrick Chao, Samuel Miserendino, Gildas Chabot, David Li, Michael Sharman, Alexandra Barr, Amelia Glaese, and Jerry Tworek

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:12.267688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:07.158436Z digest=sha256:06d6a3f4e77164fdcdbb75d8c540b62bb3d652506ecf5b95533196703bb1bf17

Observation dea83764-8c4c-4cdf-91a4-e9634c9da93c · outbound

This paper cites SOP-Bench: Complex industrial SOPs for evaluating LLM agents.CoRR, abs/2506.08119, 2025.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks SOP-Bench: Complex industrial SOPs for evaluating LLM agents.CoRR, abs/2506.08119, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:07.219151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:07.219151Z digest=sha256:326cd5afa8f32995eaca49a9fda06fcf387c561b95a6b554bcb010737fcb0682

Observation 8d38aca4-dcf1-4b29-a09e-2df355d7f387 · outbound

This paper cites JobBench: Aligning agent work with human will.CoRR, 2026.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks JobBench: Aligning agent work with human will.CoRR, 2026

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:12.087003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:07.334268Z digest=sha256:2f4d72ddd8d580e78fb5173d8087e9cd8847a9da7bd8a9c50c9ef560f873c5ef

Observation b54a4f99-b720-4dd9-8b09-1b45e80194d1 · outbound

This paper cites LiveBench: A challenging, contamination-limited LLM benchmark.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks LiveBench: A challenging, contamination-limited LLM benchmark

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:11.905290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:07.430347Z digest=sha256:d03b7af9ae52dfa27b87395ea478e6b41235fecbb29e04d80368323782ae4533

Observation d3d0c413-acb6-4dc0-9ffa-a2d08416ec68 · outbound

This paper cites AntiLeakBench: Preventing data contamination by automatically constructing benchmarks with updated real-world knowledge.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks AntiLeakBench: Preventing data contamination by automatically constructing benchmarks with updated real-world knowledge

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:11.681345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:07.513842Z digest=sha256:0a9811e164a36aeff3d2245d37d83cc81aa1f7897dd2905a2806c09a8204a9b5

Observation 02b81de1-9fb0-4369-a57c-d9e5093eda5e · outbound

This paper cites Benchmarking large language models under data contamination: A survey from static to dynamic evaluation.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Benchmarking large language models under data contamination: A survey from static to dynamic evaluation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:11.503412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:07.620433Z digest=sha256:4869153429c9b7afe79be81ce524a5daf5a83f4e0301b7749a19bc41a01af7a1

Observation 466b1304-ad71-412e-afaa-f12d86c8e29f · outbound

This paper cites Voyager: An open-ended embodied agent with large language models.Transactions on Machine Learning Research, 2024.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Voyager: An open-ended embodied agent with large language models.Transactions on Machine Learning Research, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:07.681368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:07.681368Z digest=sha256:fc814f8038996b5b84259269da606f3ce50bb14f8f71f2f51cd7162ee15d1bce

Observation 432c6128-8d9b-467c-aa7e-f636ec66fb39 · outbound

This paper cites GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:07.776106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:07.776106Z digest=sha256:79282e6dbcf29445680f2ceb9c4f18b9aee2bd0bd9ee549f6cc238e0d0807a2a

Observation 2b224103-8426-4d03-94a3-98f4b0c21057 · outbound

This paper cites RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:09:09.800507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:07.863829Z digest=sha256:1d6c0e0fb3caa662ecac494421231e83b8b263074f3dd4029f96e47bd1742af3

Observation c020f1da-c0bf-453c-9e38-3b5777b6d402 · outbound

This paper cites LLMigrate: Transforming "Lazy" Large Language Models into Efficient Source Code Migrators.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks LLMigrate: Transforming "Lazy" Large Language Models into Efficient Source Code Migrators

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:07.958750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:07.958750Z digest=sha256:b33a69c52b8099663d8613dfc990cb328c678192510c54720d7dbbb985a702a5

Observation ddaf9f13-1905-4dc5-8f81-6bc041788c33 · outbound

This paper cites Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:11.316441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:08.041173Z digest=sha256:0a32f9d7e0ca47911abd5b924a72e65aaffec2d299168af195619bb7dbbaa659

Observation 7a5f93ed-3540-4308-b8cd-f2696d2e7fb1 · outbound

This paper cites Reward Is Enough: LLMs Are In-Context Reinforcement Learners.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Reward Is Enough: LLMs Are In-Context Reinforcement Learners

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:08.132427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:08.132427Z digest=sha256:ec2d32cc5b9badb42fc2a65cf2af1903a5e404cfc52c49ec7d3f62e03999ae47

Observation 42494414-8e2c-40b5-8a79-e309ce44dc0a · outbound

This paper cites SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:09:09.574586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:08.210473Z digest=sha256:3c4d1c778065b0f7dd1aad243bf687cab86a2fa5e66a4e56c85c2c5af3168477

Observation 73bf45cd-9e53-4a20-993e-0d8569b9efd5 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:08.293502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:08.293502Z digest=sha256:0f4af7708eeb3552390403db3862d21a5fdfefcabfb2d3edf4b804add26f5cf2

Observation c8e6ed68-1157-4b7b-923c-3502c1b24722 · outbound

This paper cites CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:09:09.387826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:08.411553Z digest=sha256:5b527d4d6b0568f73fefc2839ef2788342c77cd782f3f7b780a1ea394e0a8069

Observation d33b646e-a2ec-41a9-8144-9e8dea90379c · outbound

This paper cites MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:08.491553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:08.491553Z digest=sha256:ff2157554a4e2798d76a5b9c56658dee972860ba3da9de29efe4b3185d6779ce

Observation 61c8d1c8-f899-4566-9e1c-5e74ed2fafb2 · outbound

This paper cites Pollard, Alistair Johnson, Edward Choi, Yugang Jia, and Jong Ha Lee.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Pollard, Alistair Johnson, Edward Choi, Yugang Jia, and Jong Ha Lee

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:08.590394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:08.590394Z digest=sha256:a2e23e96267cd7605a4fb094542f166974261658915e3802f609eff7ed2cdabf

Observation d3e56211-5923-4e10-ae1c-c9a8c05b391c · outbound

This paper cites Harvey LAB: The legal agent benchmark.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Harvey LAB: The legal agent benchmark

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:11.114346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:08.651464Z digest=sha256:222099a025ad11236caf5344098b0373562201f91d88c1f09d3fe816993eabdc

Observation 9fb19912-667d-41b2-ae54-ad01b4e6a465 · outbound

This paper cites SpreadsheetBench: Towards challenging real world spreadsheet manipulation.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks SpreadsheetBench: Towards challenging real world spreadsheet manipulation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:10.949683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:08.730253Z digest=sha256:2e0a97d18d12c6ddf404d40ab7de222cd114820248eb61612b58f4464204b1b6

Observation 8ad928f2-3a8f-4e0d-82bf-0e40e05a660a · outbound

This paper cites InfiAgent-DABench: Evaluating agents on data analysis tasks.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks InfiAgent-DABench: Evaluating agents on data analysis tasks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:10.769801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:08.783111Z digest=sha256:5904c71b0ff1f3001e5c82b8d927ccc64134408a191504f0a96bf7fa7b5f5142

Observation d0420d3e-d12a-4e27-afe8-3df81dc3b39d · outbound

This paper cites BIRD-INTERACT: Re-imagining text-to-SQL evaluation for large language models via lens of dynamic interactions.CoRR, abs/2510.05318, 2025.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks BIRD-INTERACT: Re-imagining text-to-SQL evaluation for large language models via lens of dynamic interactions.CoRR, abs/2510.05318, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:08.848101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:08.848101Z digest=sha256:af7952cfd9c352f0049dd62bf18ac0847442b91f10bc6f8cd5350a3d9b8a8904

Observation 927f1840-81d4-4634-84c2-ddc6ccafb9b2 · outbound

This paper cites LiveSQLBench: A dynamic and contamination-free benchmark for evaluating LLMs on real-world text-to-SQL tasks.https://github.com/bird-bench/livesqlbench, 2025.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks LiveSQLBench: A dynamic and contamination-free benchmark for evaluating LLMs on real-world text-to-SQL tasks.https://github.com/bird-bench/livesqlbench, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:10.620259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-05T13:09:08.898514Z digest=sha256:3889e2bceb97e3dc7ca29ef53541da2b9c916a4531c91e5713c7a5daeea999d5

Observation 092207ef-8fef-46f4-ad6b-2ef6e2675dbe · outbound

This paper cites WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:08.945549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:08.945549Z digest=sha256:f6cfb8d1e49a4b36612f9a25bc2468c690adeacb10cf3d81f921d30d2a88b79b

Pith citing papers

No inbound Pith citation observations are available.