Pith. sign in

Paper Citation Record · LEDGER

Evo-Bench: Can Language Models Improve Agent Harness?

As of 16 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2608.09096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.09096 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:22:26.870215Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 21a47d1a-1229-447d-8e70-15b6cd4707a0 · outbound

This paper cites HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry.

Evo-Bench: Can Language Models Improve Agent Harness? HarnessX: A Composable, Adaptive, and Evolvable Agent Harness Foundry

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.628134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.628134Z digest=sha256:bc58e416ee2b55379f34323da979cc00591ce48cded06c692724c4c5740ce0aa

Observation b94458d9-0d25-41da-bf49-88c57a490cd2 · outbound

This paper cites 11 Evo-Bench: Can Language Models Improve Agent Harness? DeepSeek-AI.

Evo-Bench: Can Language Models Improve Agent Harness? 11 Evo-Bench: Can Language Models Improve Agent Harness? DeepSeek-AI

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.636498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.636498Z digest=sha256:b1683af7c40e0cc9a823b49d18408b914c1c45584ce1f418ff2e0bcf0c6e44ce

Observation 192edf02-eef4-49db-98f1-2d6ce7af8406 · outbound

This paper cites SIA: Self Improving AI with Harness & Weight Updates.

Evo-Bench: Can Language Models Improve Agent Harness? SIA: Self Improving AI with Harness & Weight Updates

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.644350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.644350Z digest=sha256:2e19c506fb7ce8c11f0cb7899bcca8f90ecc035b7110c3956f5bb5fe30032531

Observation 542b4358-d5d5-4cfc-baaf-7ee47c9cd90d · outbound

This paper cites Automated Design of Agentic Systems.

Evo-Bench: Can Language Models Improve Agent Harness? Automated Design of Agentic Systems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.648963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.648963Z digest=sha256:43dd2552e0570af72d2598ecf146cd5fd2bf197b72bc0fbad205d3bb91488fe9

Observation 5f998b68-8bfa-4d4e-b156-2a4bcfd2b75a · outbound

This paper cites MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation.

Evo-Bench: Can Language Models Improve Agent Harness? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.653121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.653121Z digest=sha256:852b41af2f2e1092cfc40b2de61e2f79f66489a16d2a8d2b8d5059591bf7eec9

Observation 7d91f3af-8392-4574-a5e8-d0ebd3653329 · outbound

This paper cites MiniMax Sparse Attention.

Evo-Bench: Can Language Models Improve Agent Harness? MiniMax Sparse Attention

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.661868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.661868Z digest=sha256:f83d17d789134f11b7a44bc3a0e799cc7f6cbb55e3fb024f6aa2a171400d4192

Observation 9a3a6329-b9da-45e3-85f6-f93c83c20bf1 · outbound

This paper cites Meta-Harness: End-to-End Optimization of Model Harnesses.

Evo-Bench: Can Language Models Improve Agent Harness? Meta-Harness: End-to-End Optimization of Model Harnesses

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.665740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.665740Z digest=sha256:e1fae18261c104e8827fbb741335b291781843c0a080201ca045fa95c45b7bb9

Observation 60e17c74-e933-49df-965a-8b850b4db3d5 · outbound

This paper cites ClawEnvKit: Automatic Environment Generation for Claw-Like Agents.

Evo-Bench: Can Language Models Improve Agent Harness? ClawEnvKit: Automatic Environment Generation for Claw-Like Agents

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.670161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.670161Z digest=sha256:8cfe49c54e37d6769d00211058d3a990e88d9e31e5fa9126b651b94fca726d01

Observation f1e14da7-cf84-4e04-974c-c61c67ca037d · outbound

This paper cites Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents.

Evo-Bench: Can Language Models Improve Agent Harness? Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.673955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.673955Z digest=sha256:66237c504915af0caf27680c0ceaf2c500f5be23913af26ccfd580fd7fb404c2

Observation 0dc64604-d4fd-4563-9004-aaf627cdceee · outbound

This paper cites The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?.

Evo-Bench: Can Language Models Improve Agent Harness? The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.679367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.679367Z digest=sha256:e5ba5e42a0933d85433446fe0495a2abb96ce5d18b8aaf70627597d29b5f427b

Observation 41b0c0d4-d6b0-45ed-9a06-eb006d0cb72c · outbound

This paper cites AlphaEvolve: A coding agent for scientific and algorithmic discovery.

Evo-Bench: Can Language Models Improve Agent Harness? AlphaEvolve: A coding agent for scientific and algorithmic discovery

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.683106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.683106Z digest=sha256:5f6378d54131165d04f1ca5f199123fc010e27cefc1202d668512c84860fbfa9

Observation ae91a44e-5ae1-41f6-b4ef-1a20507e0424 · outbound

This paper cites 16, 2025; cloud-agent research preview announced May 16,.

Evo-Bench: Can Language Models Improve Agent Harness? 16, 2025; cloud-agent research preview announced May 16,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.686828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.686828Z digest=sha256:f4699de5f1b9a05d081488098fb4d3bbd499265ad15d040e27b8baf4684281cd

Observation 98c760b8-a68b-4f7c-acaf-fef3972e3371 · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

Evo-Bench: Can Language Models Improve Agent Harness? GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.691164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.691164Z digest=sha256:623222d69bd698101c470897aa1e71a4493dff1403cfad8ac298ca9e0da316d1

Observation c95b22f6-3d20-4ced-b076-616f405caaf0 · outbound

This paper cites Qwen3.6-27B: Flagship-level coding in a 27B dense model, April 2026a.

Evo-Bench: Can Language Models Improve Agent Harness? Qwen3.6-27B: Flagship-level coding in a 27B dense model, April 2026a

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.696626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.696626Z digest=sha256:ac7a3f24bcfe8157fd4b4068308bd69aa53c33f9d7bdf808139a2c5dde3f1c5d

Observation cfc401ee-8430-4d30-90e9-ab59b4f4e620 · outbound

This paper cites PaperBench: Evaluating AI's Ability to Replicate AI Research.

Evo-Bench: Can Language Models Improve Agent Harness? PaperBench: Evaluating AI's Ability to Replicate AI Research

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.700377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.700377Z digest=sha256:6d1f69afffbb8729503e451b112db403b0dd28e65a108a39b4ec2272d6296aab

Observation f739b9d8-5d58-4ebe-a8d2-e9c415888516 · outbound

This paper cites Gemma 4 Technical Report.

Evo-Bench: Can Language Models Improve Agent Harness? Gemma 4 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.705407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.705407Z digest=sha256:8df4cb784b87255f502df830436da723a1b06bddcc3a1dc54b329d63213c3ed9

Observation 0174f3a2-cb08-46a2-8fcf-1240ab09803f · outbound

This paper cites MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling.

Evo-Bench: Can Language Models Improve Agent Harness? MiroThinker: Pushing the Performance Boundaries of Open-Source Research Agents via Model, Context, and Interactive Scaling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.710180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.710180Z digest=sha256:dde0db009e4063339ae97356f379296f9ddbdcc8c049966d7b4ab9a4ff73b236

Observation 108dbe65-6a23-495c-9c94-9e21f0963f76 · outbound

This paper cites VeRO: A Harness for Agents to Optimize Agents.

Evo-Bench: Can Language Models Improve Agent Harness? VeRO: A Harness for Agents to Optimize Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.714960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.714960Z digest=sha256:eb3c10b57007ce975e93044fd974238b875cb8619c70cb9a7d5e56ea4f955955

Observation def9e4d5-b117-45cf-9016-6c6da82b2285 · outbound

This paper cites Apex-agents.arXiv preprint arXiv:2601.14242,.

Evo-Bench: Can Language Models Improve Agent Harness? Apex-agents.arXiv preprint arXiv:2601.14242,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.730007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.730007Z digest=sha256:5c2fdae12343ca288308afcad345540770ad9335039e07de06a767ce5a753e3d

Observation 4308d9a8-9172-47a9-b55f-c1eacca0b19e · outbound

This paper cites Rethinking the Evaluation of Harness Evolution for Agents.

Evo-Bench: Can Language Models Improve Agent Harness? Rethinking the Evaluation of Harness Evolution for Agents

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.761061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.761061Z digest=sha256:2fa195352a06805f9a1d54b81af8c429e60ea3b5c6ea6db0cc495a0353311118

Observation 6c557bd6-57e4-4451-8b03-2bd022e647fa · outbound

This paper cites BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents.

Evo-Bench: Can Language Models Improve Agent Harness? BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.804133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.804133Z digest=sha256:dc18302d758a8bde64ff0c3c8970a26f4655d064f71aa2bbd495ded14d64ac5e

Observation 40910431-7f02-4f95-ab99-34b730e701de · outbound

This paper cites RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts.

Evo-Bench: Can Language Models Improve Agent Harness? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.843491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.843491Z digest=sha256:86ff313781fc5ee4caf236c27f4a9eab0fa58640360e561a55297231ebc27fbd

Observation f1477e90-55bf-4874-b626-25d1c27f2edf · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Evo-Bench: Can Language Models Improve Agent Harness? $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.848374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.848374Z digest=sha256:d8a9b952b346baa7a07e18ec23b707a77c4de82d0218405e7fcc05f24538d383

Observation f0df5ac0-0714-4f99-9677-2070dfceb393 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Evo-Bench: Can Language Models Improve Agent Harness? $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.851933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.851933Z digest=sha256:15f8a42fc8b3ba54688310d9cc9806f2631e8ef4cff4c355dd12a935276f77f8

Observation 3206dd85-11ef-448a-a448-3c52425eeef9 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

Evo-Bench: Can Language Models Improve Agent Harness? GLM-5: from Vibe Coding to Agentic Engineering

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.856839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.856839Z digest=sha256:c356735588d2a14fbe05fa5cc3c067e8bb6b0bd620d29ea36ec04bd5f5433261

Observation ef8392bb-4f72-42ca-baca-929458df42b2 · outbound

This paper cites Self-Harness: Harnesses That Improve Themselves.

Evo-Bench: Can Language Models Improve Agent Harness? Self-Harness: Harnesses That Improve Themselves

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.860842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.860842Z digest=sha256:5b0d8c367a57186d6ec3c8787ae1773fedd214feedd5fcd1d94155c3d6b0f40a

Observation 77191e59-d1ca-4303-b48b-1e2363aa591e · outbound

This paper cites 16 A.2 Details of the Policy Harness.

Evo-Bench: Can Language Models Improve Agent Harness? 16 A.2 Details of the Policy Harness

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.865501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.865501Z digest=sha256:2e30aaa056de47b321fd15d7dc0829471612c5b41d402d8537a972a062e969f5

Observation 3f3e5601-2092-4b19-952a-7e640d960733 · outbound

This paper cites Domain tools, planning, memory, and verification are left for evolution.

Evo-Bench: Can Language Models Improve Agent Harness? Domain tools, planning, memory, and verification are left for evolution

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.870215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.870215Z digest=sha256:0e6021d1876d9d377db35e94cd9a4b47e373a8440d88e0d4a266813628c1806e

Observation f98ced87-c1ac-48b5-ab92-2d30adcc738a · outbound

This paper cites SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment.

Evo-Bench: Can Language Models Improve Agent Harness? SEA-Eval: A Benchmark for Evaluating Self-Evolving Agents Beyond Episodic Assessment

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.657509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.657509Z digest=sha256:f7aab35e4539222d0908af75188bc6e2fb32d568226a128037f3e5b5671745a9

Observation a1d1bcf8-d99b-489d-a183-23f4ca6af170 · outbound

This paper cites EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer.

Evo-Bench: Can Language Models Improve Agent Harness? EvoAgentBench: Benchmarking Agent Self-Evolution via Ability Transfer

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.640162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.640162Z digest=sha256:37883e68704daeb6317e311d46d898d039153f65e74face22145d607a1173958

Observation 65bbff88-96b1-4ed7-974c-958d79cbae98 · outbound

This paper cites 24, 2025; general availability May 22,.

Evo-Bench: Can Language Models Improve Agent Harness? 24, 2025; general availability May 22,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.620002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.620002Z digest=sha256:e5fca8556e51ee0611e2dc55ec4c5a0ef6384217038b08efebb4943076e7af93

Observation fffa250f-dc3b-462d-8df6-8ba7d009ff76 · outbound

This paper cites Mle-bench: Evaluating machine learn- ing agents on machine learning engineering.

Evo-Bench: Can Language Models Improve Agent Harness? Mle-bench: Evaluating machine learn- ing agents on machine learning engineering

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.624677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.624677Z digest=sha256:34f71d5b1e5eb43a283d984ba5c1b2d81bc12783a0d074f8b8fb9b038d693c43

Pith citing papers

No inbound Pith citation observations are available.