Pith. sign in

Paper Citation Record · LEDGER

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks

As of 19 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2608.03764.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03764 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T13:09:08.945549Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact6
  • verified fuzzy16
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f18c76b0-c63b-49c9-a1cd-f674f1596b8a · outbound

This paper cites A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence.Transactions on Machine Learning Research, 2026.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks A survey of self-evolving agents: What, when, how, and where to evolve on the path to artificial super intelligence.Transactions on Machine Learning Research, 2026

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:13.404964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:05.884507Z digest=sha256:cc6ba6023f44c4043e6c2f2602c27b7d0ebb8ef850cb72524545172bd9fa0843

Observation 323bfd6f-ea54-47e9-ac69-82a1372d9fb0 · outbound

This paper cites A comprehensive survey of continual learning: Theory, method and application.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46:5362–5383, 2024.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks A comprehensive survey of continual learning: Theory, method and application.IEEE Transactions on Pattern Analysis and Machine Intelligence, 46:5362–5383, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:13.187967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:06.043829Z digest=sha256:544a4e7bdcb0753239a2a7f620264643475012f6eeb342c9c91c588902f447a9

Observation 1181aecd-ac95-4f6c-a8f9-eadee7eb2edf · outbound

This paper cites Reflexion: Language 10 agents with verbal reinforcement learning.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Reflexion: Language 10 agents with verbal reinforcement learning

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:12.995487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:06.138087Z digest=sha256:707309040bf86235a98b368dbc70bc775a05dcf2f8081828fda20cecf159de92

Observation fcbed200-7f4a-4a32-b7ea-16b99f87d0d8 · outbound

This paper cites ExpeL: LLM agents are experiential learners.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks ExpeL: LLM agents are experiential learners

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:12.812322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:06.226363Z digest=sha256:0323c79845a701419bd1b67d0505ae44b40f6eed0fe4ef1778770310db7b2ac4

Observation dd4e9245-4fe1-43c6-956b-3a4fb1d012ac · outbound

This paper cites G¨ odel machines: Fully self-referential optimal universal self-improvers.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks G¨ odel machines: Fully self-referential optimal universal self-improvers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:12.636324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:06.333375Z digest=sha256:e36ec23b2510f43548b2426558cfeb212ea909ba2c7bcf3aa5023d9fba13cdee

Observation 14365f4e-6c3d-4859-88e1-193e66514baf · outbound

This paper cites Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:06.410037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:06.410037Z digest=sha256:620cd6b67af3db7c77ffad74975bc6f9f13d7c326b66f093dcf79553ac0145f0

Observation e58e4473-b390-4895-91ea-d33b71e320b1 · outbound

This paper cites EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks EdgeBench: Unveiling Scaling Laws of Learning from Real-World Environments

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:09:10.439914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:06.494394Z digest=sha256:d3faa334a48a402ae6378e81ba293af31a7adc53fba6618cec3edff8af379fc4

Observation b973ed99-b7d1-436c-ad94-e6aae7b54064 · outbound

This paper cites SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks SkillFlow:Benchmarking Lifelong Skill Discovery and Evolution for Autonomous Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:06.622558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:06.622558Z digest=sha256:33bbc1a80bdb0e6265b44126a533e92320adfd73240db27dca0814bd7bfd32b9

Observation 0f7172b8-3e07-4067-818d-4c13b2c0a6df · outbound

This paper cites SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:06.712982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:06.712982Z digest=sha256:c60b8160e63daf7594809bc93785cf05d835da045ff5f1233e436ef8bddaf829

Observation 7bba5e16-9abd-448f-9f42-095f16815c49 · outbound

This paper cites EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:09:10.241179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:06.823752Z digest=sha256:e8bd4f738bd55a2dfb5af9c2f4a49206afff869ffa35c5357eca9c6a1aeadc46

Observation 0bff3dce-b501-4881-8b3c-7e55c52d112a · outbound

This paper cites BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks BenchTrace: A Benchmark for Testing Reflection Ability and Controlled Evolution in LLM Agents

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:09:10.035784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:06.907341Z digest=sha256:2b0519b73e3265fba6289a263ca1afaf82cae81abe8341bf04628bde609aabc6

Observation b890d3ec-6d7e-4343-9ca9-a8ffe333cb3c · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik R

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:12.444324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:06.986986Z digest=sha256:9218d6da1b8ba025bf29a0d74417be030a50579347cf9e820e7724a7645d3e05

Observation 9f14d8b9-53a5-414a-8c41-9eeb6008326e · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:07.068377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:07.068377Z digest=sha256:a310913c2790f0c7a538b446ff62698c355d2230d909d53b9d173d7049fae61d

Observation 00f87561-2eb7-4b66-b1b7-906fecd4cf7b · outbound

This paper cites Kim, Patrick Chao, Samuel Miserendino, Gildas Chabot, David Li, Michael Sharman, Alexandra Barr, Amelia Glaese, and Jerry Tworek.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Kim, Patrick Chao, Samuel Miserendino, Gildas Chabot, David Li, Michael Sharman, Alexandra Barr, Amelia Glaese, and Jerry Tworek

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:12.267688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:07.158436Z digest=sha256:a8aebe10c5c935902e0f5b8e195187330b2db222e42ea09ad19c83131fb7d934

Observation dea83764-8c4c-4cdf-91a4-e9634c9da93c · outbound

This paper cites SOP-Bench: Complex industrial SOPs for evaluating LLM agents.CoRR, abs/2506.08119, 2025.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks SOP-Bench: Complex industrial SOPs for evaluating LLM agents.CoRR, abs/2506.08119, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:07.219151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:07.219151Z digest=sha256:326cd5afa8f32995eaca49a9fda06fcf387c561b95a6b554bcb010737fcb0682

Observation 8d38aca4-dcf1-4b29-a09e-2df355d7f387 · outbound

This paper cites JobBench: Aligning agent work with human will.CoRR, 2026.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks JobBench: Aligning agent work with human will.CoRR, 2026

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:12.087003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:07.334268Z digest=sha256:70ca1f1e6db771b71c29936fd310695419e790397f38d990023749da027ab5c0

Observation b54a4f99-b720-4dd9-8b09-1b45e80194d1 · outbound

This paper cites LiveBench: A challenging, contamination-limited LLM benchmark.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks LiveBench: A challenging, contamination-limited LLM benchmark

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:11.905290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:07.430347Z digest=sha256:29f4ab9e8db93a867d673a71aab36c4860e8765214ff409918698fb31edbf41d

Observation d3d0c413-acb6-4dc0-9ffa-a2d08416ec68 · outbound

This paper cites AntiLeakBench: Preventing data contamination by automatically constructing benchmarks with updated real-world knowledge.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks AntiLeakBench: Preventing data contamination by automatically constructing benchmarks with updated real-world knowledge

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:11.681345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:07.513842Z digest=sha256:fcc80fe979d492b0910da06bffdc9f2e4456b8a1a57ae6c73d137e08ce0bc87e

Observation 02b81de1-9fb0-4369-a57c-d9e5093eda5e · outbound

This paper cites Benchmarking large language models under data contamination: A survey from static to dynamic evaluation.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Benchmarking large language models under data contamination: A survey from static to dynamic evaluation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:11.503412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:07.620433Z digest=sha256:bfbc4e7fcb0b2cc52b74446cdfc2784ff8fea7431e5c554dbcc02a24f127f5fd

Observation 466b1304-ad71-412e-afaa-f12d86c8e29f · outbound

This paper cites Voyager: An open-ended embodied agent with large language models.Transactions on Machine Learning Research, 2024.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Voyager: An open-ended embodied agent with large language models.Transactions on Machine Learning Research, 2024

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:07.681368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:07.681368Z digest=sha256:fc814f8038996b5b84259269da606f3ce50bb14f8f71f2f51cd7162ee15d1bce

Observation 432c6128-8d9b-467c-aa7e-f636ec66fb39 · outbound

This paper cites GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:07.776106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:07.776106Z digest=sha256:79282e6dbcf29445680f2ceb9c4f18b9aee2bd0bd9ee549f6cc238e0d0807a2a

Observation 2b224103-8426-4d03-94a3-98f4b0c21057 · outbound

This paper cites RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks RSIBench-Data: Benchmarking Data-Centric Research for Recursive Self-Improvement

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:09:09.800507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:07.863829Z digest=sha256:b3faa50cb7d5606fec668359158fafd8bfb97a3e3cf72e1481f906450f8d16ef

Observation c020f1da-c0bf-453c-9e38-3b5777b6d402 · outbound

This paper cites LLMigrate: Transforming "Lazy" Large Language Models into Efficient Source Code Migrators.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks LLMigrate: Transforming "Lazy" Large Language Models into Efficient Source Code Migrators

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:07.958750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:07.958750Z digest=sha256:b33a69c52b8099663d8613dfc990cb328c678192510c54720d7dbbb985a702a5

Observation ddaf9f13-1905-4dc5-8f81-6bc041788c33 · outbound

This paper cites Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:11.316441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:08.041173Z digest=sha256:a610bca53c8ce79cabfb092dce2a6992c198db1f7c8c9d991da796c7bb2934ac

Observation 7a5f93ed-3540-4308-b8cd-f2696d2e7fb1 · outbound

This paper cites Reward Is Enough: LLMs Are In-Context Reinforcement Learners.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Reward Is Enough: LLMs Are In-Context Reinforcement Learners

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:08.132427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:08.132427Z digest=sha256:ec2d32cc5b9badb42fc2a65cf2af1903a5e404cfc52c49ec7d3f62e03999ae47

Observation 42494414-8e2c-40b5-8a79-e309ce44dc0a · outbound

This paper cites SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:09:09.574586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:08.210473Z digest=sha256:1c5048a17a0728e83f89bc490ef1db0dd979fa72598eda8a15d42ce1c0149518

Observation 73bf45cd-9e53-4a20-993e-0d8569b9efd5 · outbound

This paper cites $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:08.293502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:08.293502Z digest=sha256:0f4af7708eeb3552390403db3862d21a5fdfefcabfb2d3edf4b804add26f5cf2

Observation c8e6ed68-1157-4b7b-923c-3502c1b24722 · outbound

This paper cites CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:09:09.387826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:08.411553Z digest=sha256:a5345c444a57c01cb2f1411d5fa49d85459288ef4e7aafbd8b9de702f6cc1792

Observation d33b646e-a2ec-41a9-8144-9e8dea90379c · outbound

This paper cites MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:08.491553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:08.491553Z digest=sha256:ff2157554a4e2798d76a5b9c56658dee972860ba3da9de29efe4b3185d6779ce

Observation 61c8d1c8-f899-4566-9e1c-5e74ed2fafb2 · outbound

This paper cites Pollard, Alistair Johnson, Edward Choi, Yugang Jia, and Jong Ha Lee.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Pollard, Alistair Johnson, Edward Choi, Yugang Jia, and Jong Ha Lee

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:08.590394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:08.590394Z digest=sha256:a2e23e96267cd7605a4fb094542f166974261658915e3802f609eff7ed2cdabf

Observation d3e56211-5923-4e10-ae1c-c9a8c05b391c · outbound

This paper cites Harvey LAB: The legal agent benchmark.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks Harvey LAB: The legal agent benchmark

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:11.114346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:08.651464Z digest=sha256:d745ad34ea28c12ac54f6f0433754b61ac7869730ba13eabe133beecf81ae114

Observation 9fb19912-667d-41b2-ae54-ad01b4e6a465 · outbound

This paper cites SpreadsheetBench: Towards challenging real world spreadsheet manipulation.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks SpreadsheetBench: Towards challenging real world spreadsheet manipulation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:10.949683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:08.730253Z digest=sha256:16905cc7399f202a1a6bc54e08236c166a7567fb18193add3d1c20a4406d7faf

Observation 8ad928f2-3a8f-4e0d-82bf-0e40e05a660a · outbound

This paper cites InfiAgent-DABench: Evaluating agents on data analysis tasks.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks InfiAgent-DABench: Evaluating agents on data analysis tasks

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:10.769801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:08.783111Z digest=sha256:0b179b98169c9678b7e544c462448e7e7a6d5f7abb59decbb4bc04642c663421

Observation d0420d3e-d12a-4e27-afe8-3df81dc3b39d · outbound

This paper cites BIRD-INTERACT: Re-imagining text-to-SQL evaluation for large language models via lens of dynamic interactions.CoRR, abs/2510.05318, 2025.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks BIRD-INTERACT: Re-imagining text-to-SQL evaluation for large language models via lens of dynamic interactions.CoRR, abs/2510.05318, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:08.848101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:08.848101Z digest=sha256:af7952cfd9c352f0049dd62bf18ac0847442b91f10bc6f8cd5350a3d9b8a8904

Observation 927f1840-81d4-4634-84c2-ddc6ccafb9b2 · outbound

This paper cites LiveSQLBench: A dynamic and contamination-free benchmark for evaluating LLMs on real-world text-to-SQL tasks.https://github.com/bird-bench/livesqlbench, 2025.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks LiveSQLBench: A dynamic and contamination-free benchmark for evaluating LLMs on real-world text-to-SQL tasks.https://github.com/bird-bench/livesqlbench, 2025

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T13:09:10.620259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T13:09:08.898514Z digest=sha256:e62682930b7e073dc588c39755c2eaeb6a573621c5883524083d980946d8d02b

Observation 092207ef-8fef-46f4-ad6b-2ef6e2675dbe · outbound

This paper cites WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting.

GDPevo: Evaluating Agent Self-Evolution on Real Business Tasks WorkBench: a Benchmark Dataset for Agents in a Realistic Workplace Setting

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T13:09:08.945549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:09:08.945549Z digest=sha256:f6cfb8d1e49a4b36612f9a25bc2468c690adeacb10cf3d81f921d30d2a88b79b

Pith citing papers

No inbound Pith citation observations are available.