Pith. sign in

Paper Citation Record · LEDGER

MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 76 inbound Pith citation observations for arXiv:2310.03302.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.03302 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 76 of 76 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 76 of 76 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:59:33.392950Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-05T17:51:14.743354Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e22682f9-6d22-4950-8ab2-4df6f2cbf0ab · inbound

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code cites this paper.

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 164

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T17:34:42.767663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T17:34:42.565806Z digest=sha256:f8fb0e5351ef9c01e0fbd012347e841c7b92c76dbaa41c3793e3b10c05edd2a3

Observation dd32d7fe-6786-494d-9a52-773dd83e9a5c · inbound

Probing the Capacity of Language Model Agents to Operationalize Disparate Experiential Context Despite Distraction cites this paper.

Probing the Capacity of Language Model Agents to Operationalize Disparate Experiential Context Despite Distraction MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T17:12:27.467241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:12:27.467241Z digest=sha256:85b7bf3a8e1f75792b209423c05e4ede9303730b9a66d83080c33998f5bb384c

Observation 482e5e49-c928-48db-80dd-957436abcce4 · inbound

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts cites this paper.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.146051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.146051Z digest=sha256:52b67f054e12aeaf4811aa960da3ece681b90c1446f5d827f605ebc4843a5c02

Observation 2e654b03-1055-4c8e-81a8-94424a24131e · inbound

How Well Can Modern LLMs Act as Agent Cores in Radiology Environments? cites this paper.

How Well Can Modern LLMs Act as Agent Cores in Radiology Environments? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T17:03:41.279896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:03:41.279896Z digest=sha256:8ad5d54254632df8253ecfb2af4104fc72015731eb21f2e6aa35af079e9c8e66

Observation de8015ab-8bd0-4a98-80ea-3f67f0c5bd17 · inbound

A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future cites this paper.

A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 243

Resolution
unresolved
no resolver link, observed 2026-08-11T12:33:41.068131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:33:41.068131Z digest=sha256:528e27a1f01704dc078275655ab57d252b5854b3ee3599f6ff9d47aad3fb0c2e

Observation 2362b601-a1e3-4c51-bafb-9189be1bc3b5 · inbound

A Survey on LLM-based Multi-Agent System: Recent Advances and New Frontiers in Application cites this paper.

A Survey on LLM-based Multi-Agent System: Recent Advances and New Frontiers in Application MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T05:29:24.434408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:29:24.434408Z digest=sha256:0db41417ff04d711f10c07a662decdf7b1959d0445e94c6c8f64dcf6dbcd91f6

Observation 0b1e7f17-7f48-4c2c-952d-dd87aa842518 · inbound

HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems cites this paper.

HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T19:46:23.635504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:46:23.635504Z digest=sha256:5ada145ea84a9c4476b50c9014752667d70996cef7eb74b2332796b5736daf19

Observation 06221fec-513e-4339-bd7c-6c3668e36722 · inbound

IRIS: Interactive Research Ideation System for Accelerating Scientific Discovery cites this paper.

IRIS: Interactive Research Ideation System for Accelerating Scientific Discovery MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T10:59:33.392950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T10:59:33.392950Z digest=sha256:3c83df011301bb166498249a647f1c29278942abb415dc55f14a9f9c8a8fc1af

Observation 33a3b63b-61a8-4645-be24-bb57c4251e68 · inbound

MLZero: A Multi-Agent System for End-to-end Machine Learning Automation cites this paper.

MLZero: A Multi-Agent System for End-to-end Machine Learning Automation MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:12:30.978625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:12:30.978625Z digest=sha256:da7dd5dab9203341030d3cbe83fde595b723e0f25ae6dde7f8b5da6c29ff9906

Observation a9e4d109-9fac-4863-ae07-f3aa906b5f2d · inbound

MM-Agent: LLM as Agents for Real-world Mathematical Modeling Problem cites this paper.

MM-Agent: LLM as Agents for Real-world Mathematical Modeling Problem MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:43.182959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:43.182959Z digest=sha256:5c82e51cc38562dc438e49a6cf9952c03ca2c3fdd81af37f188deeccac1ec4f9

Observation 497203fd-20c9-4fac-8251-62b635e883c4 · inbound

CIKT: A Collaborative and Iterative Knowledge Tracing Framework with Large Language Models cites this paper.

CIKT: A Collaborative and Iterative Knowledge Tracing Framework with Large Language Models MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:46:02.358310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:46:02.358310Z digest=sha256:028c8d98823672d5109e7ffd67430efb144b09dc844a88e308995a93ed38bf9c

Observation 36be6c58-142b-41d6-af9d-7998f68f4ee6 · inbound

RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving cites this paper.

RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:16.028495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:16.028495Z digest=sha256:ea57eff61f9a2b4b4cb09adb4123b53ff08b3b3d1dfe850225111c55f27d0581

Observation 5cc0eecf-07f5-42c9-8155-d58f6c755ec1 · inbound

Co-Saving: Resource Aware Multi-Agent Collaboration for Software Development cites this paper.

Co-Saving: Resource Aware Multi-Agent Collaboration for Software Development MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:27:34.074882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:27:34.074882Z digest=sha256:4b91e7d3ee88ffd0d59972627b8c906c1ac38482352916c4fae17a60a8ba52c3

Observation cdc97199-e562-4ef9-abd7-04dff8b60230 · inbound

Predicting Empirical AI Research Outcomes with Language Models cites this paper.

Predicting Empirical AI Research Outcomes with Language Models MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:55.906410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:55.906410Z digest=sha256:1e427204cadc76f3fb3e5c080707e6fede70d816259e7986679c50397842ed8e

Observation c0913741-ea1a-45d3-8443-1b3869f10af8 · inbound

Agentomics-ML: Autonomous Machine Learning Experimentation Agent for Genomic and Transcriptomic Data cites this paper.

Agentomics-ML: Autonomous Machine Learning Experimentation Agent for Genomic and Transcriptomic Data MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:59.699319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:59.699319Z digest=sha256:5a60937056109b52f9531fa1b8e9052b68f9e7a7c4e12587ce2d22dac9478dfe

Observation 9240ddd9-22e9-456f-af9c-378e3fe489c8 · inbound

Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research cites this paper.

Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T00:55:16.903112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:55:16.903112Z digest=sha256:80a4eba6344980e4b48e27c0736df1d1b51ae56397e817f092a9a9441132b4d6

Observation 824bc8cf-a92d-44e2-b1bf-97ae87bf6406 · inbound

ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning cites this paper.

ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:28:26.876986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:28:26.876986Z digest=sha256:b05dae91f9ddb36bafe5b23de92091391cb06378cf60e0b61401fd55497866e8

Observation 34a0a56a-b63d-4394-94e3-d06fd10fc306 · inbound

Deep Research Agents: A Systematic Examination And Roadmap cites this paper.

Deep Research Agents: A Systematic Examination And Roadmap MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:56.670441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:56.670441Z digest=sha256:ca5058f4329b899c5f3064f1d2d5518489c55b7ac348a0dfb27c2e53e18a473d

Observation 4af3d1df-4469-45d9-87ff-caa730185580 · inbound

THE-Tree: Can Tracing Historical Evolution Enhance Scientific Verification and Reasoning? cites this paper.

THE-Tree: Can Tracing Historical Evolution Enhance Scientific Verification and Reasoning? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:24:47.622941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:24:47.622941Z digest=sha256:bc4de2370a989d58742e601db98044ed4486401ec5cb838e3138db298a949b74

Observation 43152c97-83f1-4231-afce-e6da05582938 · inbound

Agent Identity Evals: Measuring Agentic Identity cites this paper.

Agent Identity Evals: Measuring Agentic Identity MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T14:57:06.640490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:57:06.640490Z digest=sha256:130220db2f3075ffff53fd9f9b31e17000b72e05a6ece17d7d9339d4df32c615

Observation a62dac5d-51ee-4299-b37e-9688afd25f0a · inbound

KompeteAI: Accelerated Autonomous Multi-Agent System for End-to-End Pipeline Generation for Machine Learning Problems cites this paper.

KompeteAI: Accelerated Autonomous Multi-Agent System for End-to-End Pipeline Generation for Machine Learning Problems MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T22:22:51.733291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T22:22:19.478156Z digest=sha256:980ecdbab844738f8e41390070012d119a10379deacafccc2e9c84aa2454f923

Observation 47afa8af-5f77-41d0-8d91-daac13edd68a · inbound

Reinforcement Learning for Machine Learning Engineering Agents cites this paper.

Reinforcement Learning for Machine Learning Engineering Agents MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:24:02.231780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:24:02.231780Z digest=sha256:279fa45afff261f58b487e95fb6e61524d2694c7ac898dc4e65f864dfa9b120c

Observation 227d29d8-f707-437a-8499-71de11a6b9e8 · inbound

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization cites this paper.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.629076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.629076Z digest=sha256:45bf7f708fc9ec5446f5de2b9c62fc4cc415a4d6c43c72c71e5a06d4ecf81eff

Observation 19e65b4c-d119-4f78-a9e3-fd558d2d7f9d · inbound

Can We Predict Before Executing Machine Learning Agents? cites this paper.

Can We Predict Before Executing Machine Learning Agents? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T15:43:03.592807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T15:42:25.852911Z digest=sha256:aec553cbb6e8c71822a582a85edf336d0b66d838aff2d96c5208a8deb89c1fc9

Observation b4090343-b136-41b7-b2cd-14bd4c35b2d0 · inbound

AceGRPO: Adaptive Curriculum Enhanced Group Relative Policy Optimization for Autonomous Machine Learning Engineering cites this paper.

AceGRPO: Adaptive Curriculum Enhanced Group Relative Policy Optimization for Autonomous Machine Learning Engineering MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:32:27.339639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T06:32:22.038300Z digest=sha256:14e8b64fc833dbdf8d8818ab1c85e637b50c2892afb02b09b8bd25f70a1c1fd9

Observation d2b289d4-6ea3-4b36-b315-08cd9ed0c454 · inbound

Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search cites this paper.

Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:50:12.229176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T17:49:46.383559Z digest=sha256:b4751834c92c65df95ee8c3fc7ea71726a8c61b370de9ca117f5ced9bf1dd499

Observation dd51afc6-27d5-4d56-9342-87524914d2e2 · inbound

Pioneer Agent: Continual Improvement of Small Language Models in Production cites this paper.

Pioneer Agent: Continual Improvement of Small Language Models in Production MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:05:57.368134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-10T17:48:40.520740Z digest=sha256:c84ada8f114c4763bbf0f573c38f7d72af8d1ef779c3962c9c3d49cbb03834e1

Observation cd3f2020-b5ed-4450-bbe0-7742b0f2ee6c · inbound

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? cites this paper.

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:03.453162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T15:55:34.768853Z digest=sha256:8a0b975fb3f6a8de43ed8196fb4ec7e9cc50b9dd6de626adee46ad40fe1f3081

Observation 223bdcbf-c044-43d3-a82e-c276b1d41722 · inbound

TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration cites this paper.

TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:46:35.049357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T12:34:29.808503Z digest=sha256:4888509920a8c2a3af92c07cf3354e9b396fbfcf4bcd3f477a0e761786dc1699

Observation 8e9faee6-5976-4c31-b9fa-35f042e7b416 · inbound

LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review cites this paper.

LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:56:33.865350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T19:52:49.324500Z digest=sha256:18cf0bfd07f5364677f97bc49f36f1819db6883022e6ce3fc556c1c39564f061

Observation 147f34cf-774e-4056-a5f2-7dbc53eb5c5c · inbound

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale cites this paper.

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:01:13.385053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T05:59:01.010437Z digest=sha256:cc08be947c94fb271a8928da0c697072191d19c8f87e380f7cbf9d56feba322e

Observation 086c5d67-6e5b-4777-889d-4c4053514e26 · inbound

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale cites this paper.

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-05T17:51:14.745073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-05T17:45:55.631459Z digest=sha256:5238891d4cea959ad387cbd8b92809a9997068f6eb6bb0fee463ecc6345a48d6

Observation 9faebeff-3442-4c9e-920a-b387b6f37505 · inbound

Read, Grep, and Synthesize: Diagnosing Cross-Domain Seed Exposure for LLM Research Ideation cites this paper.

Read, Grep, and Synthesize: Diagnosing Cross-Domain Seed Exposure for LLM Research Ideation MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:07:00.071290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-13T01:06:02.275470Z digest=sha256:0edb852843cea21e303abd7c6a04e663cc19f6cb9fe46bfe957eabb28e3cbf6b

Observation 83d892e4-5187-48a9-84a4-4ec6ee350ed5 · inbound

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse cites this paper.

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:42:59.121140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-14T20:23:34.514071Z digest=sha256:8c778b36f0fb85fec482eaf3250b85fbacc3fe6a0b97792037f8e52b5aa8ba71

Observation 72e9c313-cd13-4bd5-a23a-2b5b36a1f247 · inbound

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse cites this paper.

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T04:55:04.193134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T04:51:17.519200Z digest=sha256:fe2f6f75ce1e6eaa55733f64a54680bb8e61d25548d31ec70485e3d1b0f69003

Observation 6ab426c4-9ec5-48d3-b3f7-f1bbf567ca0d · inbound

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction cites this paper.

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:05:06.747450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-15T06:04:03.605898Z digest=sha256:ac9806cad38ce0877280a5e29a3ef85fa6f4e0e9bb7bd39d6af8d366a33836fe

Observation 7d92aaab-5869-465d-b174-a9f8e419e1e8 · inbound

BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks cites this paper.

BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:32:43.734709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-19T19:31:32.334837Z digest=sha256:ba8a1813ef9cceb3f680a0ac662f732d216e4642c85ef03ea071db7e74680528

Observation faa7c95a-d138-42c4-95fb-b654ee60a4e1 · inbound

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility cites this paper.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:03:43.962049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:1695ee27e62601bbe7c6e7f64a8c80073cc2db18dcb73cf8b1d8555daa8f94a9

Observation d0dab6a4-af03-438e-b8d0-050a0e77b66f · inbound

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics cites this paper.

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:28:21.467445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T14:25:15.565386Z digest=sha256:a61c01b585522caa759631d9fc6895e53fe0e7c2856ccec06bbe6d20ce6c3caf

Observation 65df9179-0ebe-4a8f-acc9-3fa4c16117b7 · inbound

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics cites this paper.

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.912674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T19:00:30.961402Z digest=sha256:eeec9de13beacb04aae2abff0c54cbc93436a37846de6a3f4a9f5a96ee911428

Observation 9d3d2142-eb7d-4ff4-8b8b-ced9c26b255a · inbound

How Far Are We From True Auto-Research? cites this paper.

How Far Are We From True Auto-Research? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T09:58:11.238700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-20T09:56:16.160551Z digest=sha256:75616c648741c5f1c01f030f6f432c40e118dc8c0d9498a9667a8f075b0831aa

Observation 792e23f6-c3ff-4f18-b586-b2a53000ffce · inbound

AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery cites this paper.

AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 119

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:21.832940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-25T04:46:43.679185Z digest=sha256:164ecac0cadfced49fde79792fb3d952254d72759fb69dcc4d66a714f583cea7

Observation f42692fa-8b91-4298-acb2-9a463988e008 · inbound

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence cites this paper.

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T21:23:59.113747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-29T21:19:03.281629Z digest=sha256:9e3037f5c92625b59c23b60569ad1e1f3c8b742973aa3f59a11e9d5285c9fc05

Observation 10091f2b-0aad-42d7-bf17-f83624152935 · inbound

AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation cites this paper.

AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:43:25.675420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T12:39:00.460842Z digest=sha256:c4ab2d7257326b50bdbef6f7db6dc659b315c15fbbf223947a9d08665303da11

Observation adf244bb-13de-4a2d-be08-e6597990a58a · inbound

LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis cites this paper.

LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T09:03:16.040070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T08:56:24.657380Z digest=sha256:381fba8121178987b240d15058b88e8f3e0e7fbd2a83e8eb161fc194b0bba1f6

Observation a251f68d-354f-4671-9add-bf046b693955 · inbound

Can Generalist Agents Automate Data Curation? cites this paper.

Can Generalist Agents Automate Data Curation? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:56:35.140323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T09:32:39.361415Z digest=sha256:af23679b4a38848e87bb1389409051bab6e85f1aaf3c6d8684358aa0c2152b28

Observation 137e84eb-005c-41d3-b6ad-e2b7e31e2fdf · inbound

Towards Persistent Case-Based Memory for Autonomous Data Science: A CBR-Augmented R&D-Agent with a Locally Deployable Small Language Model cites this paper.

Towards Persistent Case-Based Memory for Autonomous Data Science: A CBR-Augmented R&D-Agent with a Locally Deployable Small Language Model MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T10:06:51.697112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T05:19:14.975753Z digest=sha256:c64afb54adcc33555ff26924d0804054148a6a1c6ecb6e10fb1deff9f3695e4e

Observation 2e5178ee-23cf-4342-8f27-b7e1bedac843 · inbound

MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery cites this paper.

MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:46:59.786067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T00:57:02.959467Z digest=sha256:0f756e9919599dbd5f50fe73a199ad5dca85f9458285c976c1222ed71dde1bdb

Observation 15c51748-9d40-4cd7-bb6b-27f5de7401c7 · inbound

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research cites this paper.

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:33:15.805420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-29T08:24:42.412763Z digest=sha256:5b788b68a40ae24a3165def3e921b0791be1cde9d5317945520e60d02961ae27

Observation 0f774748-35f1-4c5d-8cef-ef19ab480933 · inbound

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research cites this paper.

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:39:16.575746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-04T00:30:00.449270Z digest=sha256:9085bb210e2778fb5a6974aab895482320446efafefe6401efdb662c999354b5

Observation 701e14d2-be9f-4733-a559-bf63a93985bb · inbound

Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories cites this paper.

Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:37:37.772034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T13:43:02.919248Z digest=sha256:8493e9aaf53727d9e1955861fa461d2ea6f56efa3653c4f1f7a35226dd453433

Observation 84420d36-0c62-49d9-9907-497a96646daf · inbound

Trustworthy Self-Composable Big-Data-as-a-Service: An LLM-Orchestrated Multi-Agent Framework for Automated Data Engineering, AutoML, MLOps Deployment, and Drift-Aware Lifecycle Optimization cites this paper.

Trustworthy Self-Composable Big-Data-as-a-Service: An LLM-Orchestrated Multi-Agent Framework for Automated Data Engineering, AutoML, MLOps Deployment, and Drift-Aware Lifecycle Optimization MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:29:02.943758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T22:07:45.909694Z digest=sha256:a882cb424954caf3db38e1528376cec14a860aaf0d86866f2fe4db83c3b27b54

Observation 9cd2f965-2ec6-4874-b5dd-4df2b11b08cd · inbound

Agentic AutoResearch forSpace Autonomy: An Auditable, LLM-Driven Research Agent for Aerospace Control Problems cites this paper.

Agentic AutoResearch forSpace Autonomy: An Auditable, LLM-Driven Research Agent for Aerospace Control Problems MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:59:32.715993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T17:29:59.914576Z digest=sha256:8dc3d71cf78f4f9346018ab7173fbf0bee06fa984fed4502a644139f4b21f4f1

Observation a9e68e22-d9e6-4cc4-86ef-bfa3f0a25ec7 · inbound

PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement cites this paper.

PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:59:43.504947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T10:36:14.840581Z digest=sha256:fd77308c5b7d411de38e02d8b0a405fc21bd76ecf941b0d0f489d9939934250e

Observation f90db1d0-05fc-4791-aa6b-ce31a1fafd1c · inbound

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? cites this paper.

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:57.799757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-26T00:13:14.940915Z digest=sha256:94b2c099291fdda23828583d90f9b6f69eaab03fa89284a127a8b3c5503fa65e

Observation 34e973d5-9efb-483a-94d1-008b7dde1658 · inbound

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? cites this paper.

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T12:32:39.706055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T12:32:39.706055Z digest=sha256:5540e0df6f35fef34cc92005811c47d0351d612828ec3436f76d6015b1f055dd

Observation 9c1c4257-7653-40c1-84d5-31e249cf05cf · inbound

Glite ARF: Verifier-Driven Research with Parallel LLM Coding Agents cites this paper.

Glite ARF: Verifier-Driven Research with Parallel LLM Coding Agents MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:06:02.745752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-29T01:15:00.436786Z digest=sha256:d694da5056b5af4884fa7f988cdfda964b2084cc180ffcf1f9c9eecdd529e9ef

Observation ebac2fff-1b08-4520-86b8-79572d0a36ec · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:14:21.196292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T07:10:38.909339Z digest=sha256:44ab3fa18734748beb2455751d7dc115c18ab40c1f1910958fd298c315b4febe

Observation 5df1d11f-51c6-4bf4-85fa-92d80987df8b · inbound

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks cites this paper.

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-15T10:24:53.345620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:24:53.345620Z digest=sha256:86a5621be9b71e806f2433c57e92e91c986e0292b8f9a27c6c7bdde3657440a1

Observation 7ee6a2fc-46ff-4d74-b42a-5d17c42349a2 · inbound

FARS: A Fully Automated Research System Deployed at Scale cites this paper.

FARS: A Fully Automated Research System Deployed at Scale MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:35:42.748977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-01T05:18:53.840963Z digest=sha256:46ec356bfed3b82075fc5c32af4d5352a0190eaeac27a7f5700cb453b7165ca2

Observation 490b25af-c307-4926-b17b-7874354c2ff5 · inbound

FARS: A Fully Automated Research System Deployed at Scale cites this paper.

FARS: A Fully Automated Research System Deployed at Scale MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T16:55:37.417202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:55:37.417202Z digest=sha256:44e0e13f603c3d10100b5a1db2fb7648c2b42ff9a40511079400e4036be350da

Observation d6bea006-f3cf-4edf-8e8c-13f3d8637773 · inbound

Compression, structure, and executor capability: a controlled real-cost decomposition of language-model agent skill optimisation cites this paper.

Compression, structure, and executor capability: a controlled real-cost decomposition of language-model agent skill optimisation MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T05:14:41.315225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:14:41.315225Z digest=sha256:05ebcd506b514e94a7f0fb0032878f7b1407213fb7ece98e7667a222040b01f7

Observation 5ec7eccd-f664-4f90-aedf-6adcbf74213f · inbound

ArchEval: Measuring AI Agents as Computer Architects cites this paper.

ArchEval: Measuring AI Agents as Computer Architects MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T01:16:50.927415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:16:50.927415Z digest=sha256:108e76ed5e4a42e144a1cc0bb29ae65d491b28379b87df45dd73a1b5f0f71783

Observation a4cb8144-1aa6-4c15-8402-fc415491b0be · inbound

When Does Restricting a Coding Agent to execute_code Help? A Regime $\times$ Agent-Design Ablation cites this paper.

When Does Restricting a Coding Agent to execute_code Help? A Regime $\times$ Agent-Design Ablation MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T10:46:01.272433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:46:01.272433Z digest=sha256:d3dcc546e481e01a1aafd80fd9b0ad5ccb7086f540fa64e7dee1dbae43eaaf68

Observation 8915d0f7-10cf-4049-800f-45908588bdda · inbound

StarCodex: Dynamic Coding Harness for Starlink Measurement Analysis and Experiment Automation cites this paper.

StarCodex: Dynamic Coding Harness for Starlink Measurement Analysis and Experiment Automation MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T23:04:49.393664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:04:49.393664Z digest=sha256:6b857693c54733458f206fb7b3a0e5014147f8ef6d60d5b0ce60ae584ff5f9d4

Observation 75e97664-e545-4253-a3ab-ac63e16441ba · inbound

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows cites this paper.

SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T03:35:43.019419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:35:43.019419Z digest=sha256:87d7bf529048b9755d2c5dcec43fc8774a273a3ae2c9e00d45b60eae1aa47883

Observation 668008f8-d1d1-4c27-9fb9-3c117585f4d5 · inbound

Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering cites this paper.

Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T01:39:47.207858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:39:47.207858Z digest=sha256:76978ac2f05172d62c63afaba4dac92f27dd515610f3262ae0b9b3b740d91f93

Observation 6e976994-74ce-40cc-be05-cb29731d9f6d · inbound

RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems cites this paper.

RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T11:00:13.304321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T11:00:13.304321Z digest=sha256:ce874461aa72c3787826d6ae3fc42b8683bdb676c330bbd6242726671676cbe0

Observation bfabdb40-308d-4ade-9774-d158da25aaf2 · inbound

Agentic Bayesian Optimization through Surrogate-Augmented Autoresearch cites this paper.

Agentic Bayesian Optimization through Surrogate-Augmented Autoresearch MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T00:51:32.809698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:51:32.809698Z digest=sha256:9704c37cf093ca4742ac54778b6d5e76c85095935fa9b3842bc21c93ae9d406e

Observation b1873b8c-4f30-4464-b184-fb473de7d847 · inbound

Towards a new paradigm of scientific discovery with socialized artificial intelligence cites this paper.

Towards a new paradigm of scientific discovery with socialized artificial intelligence MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-15T15:02:48.155720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:02:48.155720Z digest=sha256:0a2cfaea43d1c07e805e4f93a9a58ff19aa4d887a8f1d4c8346ed916b62fb20a

Observation f6a76272-45cc-4970-8ec7-f3dfe9430d26 · inbound

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details cites this paper.

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 186

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:40.147268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:40.147268Z digest=sha256:b0dfb482e8de94946c3563155702484f38f4c50abdf9d2682cca4c76208db95a

Observation b1f909b1-6832-467f-b923-e52a84a786e0 · inbound

Evo-Bench: Can Language Models Improve Agent Harness? cites this paper.

Evo-Bench: Can Language Models Improve Agent Harness? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T23:44:24.114790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:44:24.114790Z digest=sha256:cd33cf96ebea8aad7b58d17fed8275673725c9521beffe3e027e3ea52e520107

Observation 5f998b68-8bfa-4d4e-b156-2a4bcfd2b75a · inbound

Evo-Bench: Can Language Models Improve Agent Harness? cites this paper.

Evo-Bench: Can Language Models Improve Agent Harness? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.653121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.653121Z digest=sha256:8bfc9d56eda121e0e6c24fc6c4562dcbbf942ec8f2f30997875a08a2b31c28e1

Observation 98d398fb-3d0b-4ab0-8dfb-c691452f1b71 · inbound

verdi: retrieval is not transfer for continual world model optimization cites this paper.

verdi: retrieval is not transfer for continual world model optimization MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-11T15:34:09.138177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:34:09.138177Z digest=sha256:b9627d0a263ac76c7baa2df6b0a2de85ad9f270f50e52131935791a06600e745

Observation 2a32d19f-9c42-4377-8d46-6418a4e133d2 · inbound

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? cites this paper.

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T14:25:46.440154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:25:46.440154Z digest=sha256:8328122f32e09af3f6e16edda3ba4c89a176a9153d322e359518463a9e7d6c1f

Observation 4f0fc178-9d5c-4096-9224-1b83e2470f95 · inbound

VALG: An Agentic System for ML Theory Research cites this paper.

VALG: An Agentic System for ML Theory Research MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:44:14.910309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:44:14.910309Z digest=sha256:f846830b99d374bd16f20e29f1ae7511632c97255a2a497cfbc31730852f7a1c