Pith. sign in

Paper Citation Record · LEDGER

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

As of 7 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 100 inbound Pith citation observations for arXiv:2410.07095.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.07095 v6

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T19:11:20.600633Z

measured 136 of 136 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 100 of 127 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:57:36.403917Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact15
  • verified fuzzy16
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

9
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 10058669-c66c-4d10-931e-d0de89239b7b · outbound

This paper cites Anthropic's Responsible Scaling Policy , Version 1.0, September 2023.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Anthropic's Responsible Scaling Policy , Version 1.0, September 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:03:26.251461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:7e095d424ffb891ba44955cf1ea27152f11f60b453408ac756f449b98b88323c

Observation b41e221d-f8cb-4ff6-9116-74d91283c89c · outbound

This paper cites Program Synthesis with Large Language Models.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Program Synthesis with Large Language Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:13:21.578703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:d70dfe622a32191d2d4f14dc94430e5400821898e31962aadd1bab6429d0133c

Observation 69eb685f-89f0-40b6-97f7-cdb6cb72e6bd · outbound

This paper cites Quantifying Memorization Across Neural Language Models.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Quantifying Memorization Across Neural Language Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:13:21.574142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:054507496ad565c906d4bce9be9238284cdb2d384a0e08c47cc9ebfc9a31fb1e

Observation 31abeee7-f11b-4692-9173-46311c65a0b1 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Evaluating Large Language Models Trained on Code

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:13:21.570148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:d97ab56585ec136fe08c7e4693c8a28f2f3159d24a03d5e1bf9391b021479b4f

Observation a2285312-d1d7-4a8c-b93e-ec90e7c1b685 · outbound

This paper cites Cognition Introducing Devin , the first AI software engineer, March 2024.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Cognition Introducing Devin , the first AI software engineer, March 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:03:26.247456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:600497336054f961d8a3b5aaba58f5e60ac0d0c8f2becc9851f93d4fab2a3d31

Observation 063ad94b-622d-44f6-9392-3df1945fcdd7 · outbound

This paper cites Openvaccine: Covid-19 mrna vaccine degradation prediction.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Openvaccine: Covid-19 mrna vaccine degradation prediction

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:03:26.239077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:6b67950ea987550a71d233093ff80fc55b1ad7ac9113cf7038b5fffcdde6f937

Observation bfc1c359-73ad-4401-95e5-ac006cf958c6 · outbound

This paper cites ConStat: Performance-Based Contamination Detection in Large Language Models.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering ConStat: Performance-Based Contamination Detection in Large Language Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:13:21.549550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:5c15b264a813d1c194afbdae1e34bd3efcae28325044324d46f8e0fdf1694dc6

Observation 0a4c73d4-f7b1-4862-912d-044d64b9f07f · outbound

This paper cites GitHub Copilot Workspace : Welcome to the Copilot -native developer environment, April 2024.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering GitHub Copilot Workspace : Welcome to the Copilot -native developer environment, April 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:03:26.234961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:3676bf8ae5beabc91b1f95e72e08a0d82f911157c20fe4efd3b9ed59b4b44589

Observation 91bcb36d-b566-4dc5-85bc-cba24f532b4f · outbound

This paper cites Code Droid Technical Report , June 2024.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Code Droid Technical Report , June 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:03:26.231007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:26dd8b9222baf3118a9a1b9c76e861516e05a946bc3ad4f9d46852d162092e36

Observation 1ad5e1cc-4025-4189-834a-920f3b668c5c · outbound

This paper cites AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:13:21.517261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:574dbbb57d97d6ea6ab33fa63a36279d926c8ff66d79fdbe131c1d84b7533fd2

Observation b17d8e9a-d934-49d7-88d9-6219f8ef0a17 · outbound

This paper cites Frontier Safety Framework , May 2024.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Frontier Safety Framework , May 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:03:26.226806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:02efe8c262513b5afb8f20b5608410f8ff7eeb288f3439e194d8a4094c562506

Observation 43b31cb2-ca58-4eef-8763-1347390b8110 · outbound

This paper cites Measuring Coding Challenge Competence With APPS.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Measuring Coding Challenge Competence With APPS

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:13:21.521250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:3f4e71f8c90e1cca58b753aeca2c15013a9648780a1dbf61d50b7230c115704f

Observation 4a595fff-28c7-4351-a9f6-de5ac44b11e9 · outbound

This paper cites AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T19:13:21.526637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:7bccf44968aab06a3e576099248f394ba657c3cc616f297abe816e72cfa4518d

Observation b0ced4d7-6faf-4949-b66b-16f20112a0c3 · outbound

This paper cites MLAgentBench : Evaluating Language Agents on Machine Learning Experimentation.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering MLAgentBench : Evaluating Language Agents on Machine Learning Experimentation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:03:26.219044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:a1c33f278f59e573519a14506cbd2067cfb889b499647e00b7081d47f3cdc73b

Observation 8db45d11-7b65-4b61-804c-b0b81bf12438 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:13:21.544518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:9cd86ab2b17b1ccf95e2996303c91609f15ddec365b2f86f105e1059b3aabd69

Observation 8af70357-a0e3-45bc-929d-a3a2381a43b2 · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T19:13:21.566443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:f657a491361219975574fd657a696375a211593cd4165ad0c46a7cd59b0bafd5

Observation 556b8617-b887-420c-9d9f-a9e5dc970094 · outbound

This paper cites DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:13:21.510007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:c00f79ff399d62f177c0a3f2ec39ff78d0cf6ea28332255edf0ff98aa765b88c

Observation 4fc3763e-a96e-4bdf-9174-369cbc7c776e · outbound

This paper cites Kaggle Progression System Kaggle.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Kaggle Progression System Kaggle

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:03:26.215181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:566fa74e027c9c6e997def04586e2e0c53d7ca56111da2915c402097320204a7

Observation e80d8849-0494-4cf5-97a1-9a33bc0c445a · outbound

This paper cites Research: quantifying GitHub Copilot ’s impact on developer productivity and happiness, September 2022.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Research: quantifying GitHub Copilot ’s impact on developer productivity and happiness, September 2022

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:03:26.211519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:6ad29778bebce101851a199b068b4850919b608cde9106c8374e1f63a5ae2af2

Observation 0d7dd57a-ebc5-419b-a74b-fff200b39d49 · outbound

This paper cites AI Agents That Matter.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering AI Agents That Matter

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:13:21.505537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:f1d37e186147658834ae87f1381f703ce560e89b2b3fce00b6c197783040a793

Observation a88b45dc-14c6-4881-8967-babcd2da0a5d · outbound

This paper cites Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando De Freitas, Koray Kavukcuoglu, and Oriol Vinyals.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando De Freitas, Koray Kavukcuoglu, and Oriol Vinyals

Reference 21

Resolution
verified exact
doi, observed 2026-05-23T19:13:20.478391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:8f3b07715a60cf277bf504a9b72f10f7dabbb310cd2a2a93689efb160fd92d55

Observation 129dc745-ebc3-4563-ac56-0c49e86d4d97 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering AgentBench: Evaluating LLMs as Agents

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:13:21.557714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:b51eba4834d43a4ef53b8339a318d4863ae242caaadec2360e344c571ba3de61

Observation d67ec729-61b6-428c-999c-2ea8b6b34c40 · outbound

This paper cites Vesuvius challenge - ink detection.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Vesuvius challenge - ink detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:03:26.207090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:6c0373668621bace89155f501d1945a4a46ebad3d1847ad450bc7d7fe02c0c3e

Observation 740a3b7a-a86a-4816-8326-2bc938fef9e8 · outbound

This paper cites Discovering and exploring cases of educational source code plagiarism with Dolos.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Discovering and exploring cases of educational source code plagiarism with Dolos

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:03:26.203336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:33edfd91f95146766e0473846f5bfd19d84cc400a1023292d58ab5635cb9dae0

Observation 08f8be6f-122d-49e2-832e-4b75c9d471ac · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering GAIA: a benchmark for General AI Assistants

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-23T19:13:21.553404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:122bae39f9ed063cd469df8d267a36076d77821bb76c6c26181658e3a6cee484

Observation eabc68ac-2f55-445e-9fca-cbee9fb43411 · outbound

This paper cites Preparedness Framework , December 2023.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Preparedness Framework , December 2023

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:03:26.198351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:be24de26618d0d5094ccc2ae0d72bc63315fee11a215e043817561d944713ba3

Observation fb6f2bfa-73e9-4738-933f-8410d816ac27 · outbound

This paper cites Introducing Weco AIDE , April 2024.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Introducing Weco AIDE , April 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:03:26.194782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:1650b864bf2c8a77440b7b8b132d7e54471b04bca90f071b5d68249cd06e055d

Observation 230b2e79-42ba-4507-bf2f-bb3b8e302bd2 · outbound

This paper cites ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:13:21.535616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:2f24cf6126d76da15166d5df17a18f8054139466ea960dd3357184c050f57f29

Observation f2d1aebf-c91e-43e6-bca1-a6111eda7e80 · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T19:13:21.562094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:500eeb353f7fb6d26c7cb5dbba1e85cd156280d9ac9554c228e0e61b25bf3950

Observation ef46fc68-42e6-4160-adc9-00f48e3a7f2e · outbound

This paper cites The shift from models to compound ai systems.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering The shift from models to compound ai systems

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:03:26.191266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:e4bf366c54642dc8ca037c65047ec6b0dd95baedde78e78f02f7fc94e804d0ed

Observation 52ade589-ecb0-4e0a-8f4b-9fb84951d706 · outbound

This paper cites AutoCodeRover: Autonomous Program Improvement.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering AutoCodeRover: Autonomous Program Improvement

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:13:21.530857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:a4090204b5613e5f3ada6afce5429a5db43a1dd82c016fe0061672300cf40771

Observation a36e9827-74bb-41e2-9818-050d571d840b · outbound

This paper cites Can GPT-4 Perform Neural Architecture Search?.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Can GPT-4 Perform Neural Architecture Search?

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:13:21.540246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:088b6b4e41df6a6cf6522663e73eb81df6414cb2577243a0e67a680c9c897d40

Observation 6f895805-6347-4eac-800d-13ca628d0fbd · outbound

This paper cites write newline.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering write newline

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:03:26.187338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:478fe3709b38e8f8aecfec607f276d8ab5889b5c3628538e0e5a8ddd49c4ef1f

Observation f847085e-e9ab-47cc-9d4b-77a13c06e902 · outbound

This paper cites @esa (Ref.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering @esa (Ref

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T20:03:26.243426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:f153186cbfb2b948d28fa74f256532fd740fdb5925976e45d21d22ca7d4fc8ba

Observation 3db40eef-97a1-40e9-bc8f-41c2a5fa4593 · outbound

This paper cites an unresolved cited work.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-23T20:03:26.183215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:72e5c5dba3a2f9513aa5bd50ecfca4a5d47e2da66c3cc117c6c524f1e2ada2fa

Observation ccc572ff-0c54-41e8-8e34-f6e8b3856200 · outbound

This paper cites an unresolved cited work.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-23T20:03:26.178485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:e3f9bbd120e986210c55ae02b3670c825e50f4c9c819f9e3e8bc6bfef30abec9

Pith citing papers

Observation 0949488f-6b45-47ab-99f4-ccf2509dce71 · inbound

Frontier Models are Capable of In-context Scheming cites this paper.

Frontier Models are Capable of In-context Scheming MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-16T14:22:01.708833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T14:22:01.616448Z digest=sha256:ac53114118627edd4729d302fc93f1d2120fa0b3d3dee3df57c1aa66f52de3eb

Observation 2c8a8636-7558-4f7c-b51b-62ff7e988de6 · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.382147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:2e43187a42efdd02afff07d8caedd56f7d48aacda783b1ef3708be13a4790121

Observation baf3ef74-fbc2-4445-8a92-07d2c91f88e8 · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.266922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:81bff72dee696e34fc2b7b2e3b2ca337c5add8a9710e9948c77bd56e6e6b4baf

Observation f883961b-a295-4de7-95c4-00bcd68091d6 · inbound

Large Language Model Agent: A Survey on Methodology, Applications and Challenges cites this paper.

Large Language Model Agent: A Survey on Methodology, Applications and Challenges MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 144

Resolution
verified exact
local_arxiv, observed 2026-05-22T21:52:10.592130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T21:51:34.309870Z digest=sha256:df4949de2223ba5818cca91b99050b898978a42f60e732634ee1632e8056a5de

Observation 46a8763d-d797-452b-adba-009b1d51cbe2 · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 127

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:57:38.488809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:8d7f2938af5fcdac5bebd4c7b87304cbf5d9b45f3b7a90aeaa43ec4e720861fd

Observation 800e8d9e-d336-4dbf-b08d-e2ba52b7c0aa · inbound

From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization cites this paper.

From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:57:36.403917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:57:36.403917Z digest=sha256:7b384bb1c5dee231fa9268828a5005e639b54af315e1c0f6e514e12e9fad513f

Observation 94911a56-4025-4428-a8db-5db3c2374852 · inbound

Large Language Models for Planning: A Comprehensive and Systematic Survey cites this paper.

Large Language Models for Planning: A Comprehensive and Systematic Survey MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:11:51.566184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:11:51.566184Z digest=sha256:83b82aff69790807e45555bf5e540c4709e2f54ebe96837145910dd8c533a3be

Observation db9ab53c-6279-496e-9327-d61a9c33d19d · inbound

RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving cites this paper.

RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:16.130724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:16.130724Z digest=sha256:8b6ad26d652de0602f8355d0ea120dd4f3632df5b14639f88cba664c79533753

Observation 9cd5b506-3428-41fe-843e-60bc3e7b709a · inbound

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks cites this paper.

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:50.576064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:50.576064Z digest=sha256:e7f91ee4fbe9ed13ec12c8d1b547a83f7e3176ea2cfafc66e6f623be4350193a

Observation 089f937e-d284-4d45-9032-4f160e40483b · inbound

AI Scientists Fail Without Strong Implementation Capability cites this paper.

AI Scientists Fail Without Strong Implementation Capability MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:01.526190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:01.526190Z digest=sha256:7ef28adfc196697c436805d00b44656eeb29fd610bf1e2b9ea5378c36c0e73ce

Observation 2ed01d53-5a9b-436b-82be-455d36c104a1 · inbound

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments cites this paper.

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:57:16.254409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T11:57:08.314088Z digest=sha256:ea3129613ed8935450a0d4d55ace84c90dd3a646a87182a0abc65886d25d7f2d

Observation 7765159c-2546-4784-b802-f379cc58522b · inbound

TextAtari: 100K Frames Game Playing with Language Agents cites this paper.

TextAtari: 100K Frames Game Playing with Language Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:57.932392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:57.932392Z digest=sha256:45ac6566a9afdbcaca5c1d6b63eb229a65eee11f45e44d402f3fa6344914f14f

Observation 00a96ab0-57aa-4f08-a269-a73cf4789e00 · inbound

Agentomics-ML: Autonomous Machine Learning Experimentation Agent for Genomic and Transcriptomic Data cites this paper.

Agentomics-ML: Autonomous Machine Learning Experimentation Agent for Genomic and Transcriptomic Data MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:22:59.340621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:22:59.340621Z digest=sha256:f663e35d930a3ad098aecd76308e62294535dcea558491ac5f148de54f999ea7

Observation b37b2185-8b21-4127-a86d-11fa9a8c23d5 · inbound

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems cites this paper.

OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:23:50.616857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:23:50.616857Z digest=sha256:11ff0e124d74d2eede4644dbc7bdd9033f794a810544c588babac8fd95d68c43

Observation b9aa265e-c0c0-4265-b9a9-2354de16d98e · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 172

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.820876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.820876Z digest=sha256:d681dc2937170a7c0ee103960c973ec64493164b4c82ea91c45e1550d53b77d1

Observation 81445514-0b9c-4510-bfc1-5ad5d93c31d6 · inbound

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents cites this paper.

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:07:39.426511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:07:39.384613Z digest=sha256:3ff4b2855505f80361cef49c6cce27d3ba19b761657b3c8a1508389db5f5b796

Observation 1333024c-de16-4666-82d9-44d41a199367 · inbound

Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research cites this paper.

Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T00:55:16.898887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:55:16.898887Z digest=sha256:4b8d6cc881199dd897f14125a6f2fd031a40acd59fd42408e64a6e6178541ec0

Observation ccc444aa-7de8-44e9-b689-0b6714a020eb · inbound

RExBench: Can coding agents autonomously implement AI research extensions? cites this paper.

RExBench: Can coding agents autonomously implement AI research extensions? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:37:08.913299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T07:33:39.675929Z digest=sha256:fd272f3302f30e05588983f22636929e6cc54e72baff8ae2cc99c47ed2f20d77

Observation dbeba95d-0dd7-46bf-9482-91eef5049960 · inbound

DABstep: Data Agent Benchmark for Multi-step Reasoning cites this paper.

DABstep: Data Agent Benchmark for Multi-step Reasoning MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:37:50.650169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:37:50.650169Z digest=sha256:34d0371bbf3277707212eca34db36cffa1802993f4c75d2de8b073284227bfd4

Observation 6fe5516d-1166-42c1-9628-e810171168ca · inbound

AI4Research: A Survey of Artificial Intelligence for Scientific Research cites this paper.

AI4Research: A Survey of Artificial Intelligence for Scientific Research MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:12.210151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:12.210151Z digest=sha256:da0a17388fa547262d971ac0ea3829242c1b2bba21ae7b23a57959ef4f415061

Observation c3080b94-bee8-43c4-a758-4a0b10a01473 · inbound

Establishing Best Practices for Building Rigorous Agentic Benchmarks cites this paper.

Establishing Best Practices for Building Rigorous Agentic Benchmarks MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:21.069258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:21.069258Z digest=sha256:6d213024a80ae6e7f0b9cf27c3b156eb899f60883392454de848ca849d811c15

Observation 150e9c98-f8c7-42fb-922b-8887deee6850 · inbound

AblationBench: Evaluating Automated Planning of Ablations in Empirical AI Research cites this paper.

AblationBench: Evaluating Automated Planning of Ablations in Empirical AI Research MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T19:00:08.554622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:00:08.554622Z digest=sha256:0829c9cdb44ee1656bef5acc50f6205c6ec6919e7a437a0223dab2ca49d3516e

Observation cc15c45f-a2ad-4758-adad-4896e7be1a57 · inbound

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity cites this paper.

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:29.955772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:29.955772Z digest=sha256:fd2483e91390fb43b654ae415be3a50e8f45a2539fe8470ee1b65d79e69681cc

Observation 2e26fafb-125c-42f3-9fa6-98a8e1106f7f · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 73

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T22:23:15.658139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:a5b415c0500226f7752a95698441a70104b565dbe430b393ca5dfb3fe1be0ecd

Observation 96bba6a7-7066-4c86-8529-f0aae9c18d72 · inbound

How Far Are AI Scientists from Changing the World? cites this paper.

How Far Are AI Scientists from Changing the World? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T10:55:14.616722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:55:14.616722Z digest=sha256:58abff54ec97b3f306391e5a54432fd92c60388aff2fa9b306be0e806d715894

Observation 2c8caff6-8dcb-45c4-9c85-80b1ad948731 · inbound

TextQuests: How Good are LLMs at Text-Based Video Games? cites this paper.

TextQuests: How Good are LLMs at Text-Based Video Games? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T10:32:33.646495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:32:33.646495Z digest=sha256:d0246123471d97cb550dacc49e838f3393eaf72b900301e26b6b0d9c983c5619

Observation 5b912a33-b4a6-4a11-bbfa-cfa0ac414b7e · inbound

Preliminary suggestions for rigorous GPAI model evaluations cites this paper.

Preliminary suggestions for rigorous GPAI model evaluations MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:20:12.437386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:20:12.437386Z digest=sha256:588b1b8a227adb9774942fa98506c9b251a02f9ab0a578cfd9fb0dd677c7993d

Observation 46c82f76-f455-4970-8219-272ba1e8053b · inbound

KompeteAI: Accelerated Autonomous Multi-Agent System for End-to-End Pipeline Generation for Machine Learning Problems cites this paper.

KompeteAI: Accelerated Autonomous Multi-Agent System for End-to-End Pipeline Generation for Machine Learning Problems MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-18T22:22:51.728367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T22:22:19.478156Z digest=sha256:5661f28403c84ceb7157242e3a5875b25110a2367c9356d55b223159532ca0fa

Observation e0f69ee3-1813-43c4-9ef2-efff9eb485ec · inbound

Reliable Weak-to-Strong Monitoring of LLM Agents cites this paper.

Reliable Weak-to-Strong Monitoring of LLM Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T15:53:47.697485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:53:47.697485Z digest=sha256:4c7e29d04af00590bf74be55e202cd0885f420ac389a5e8afaf0e55df63397dc

Observation 9966b5d1-224c-4fea-9a94-78f61bc54f1f · inbound

Reinforcement Learning for Machine Learning Engineering Agents cites this paper.

Reinforcement Learning for Machine Learning Engineering Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:24:02.235163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:24:02.235163Z digest=sha256:9e813a59fb5eae5fa68888e2460bffae7d8701e7ba467d8a6e97a9adea0d548d

Observation bc36038c-3468-43e8-a1e8-9d5666dab71a · inbound

MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining cites this paper.

MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T18:11:42.733862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T18:09:18.157131Z digest=sha256:96e2e356d5b7dedb36d542f544596b3eeda629c94479986ce480bcc17fbb6d75

Observation 66fb61dd-5ec4-455b-9983-47a6a7b6d04c · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-18T00:02:25.003117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:52486fdd869729d86259094aa95a2016dfad7b727d51efde6cf45a0595202281

Observation 8efd4654-b1fb-4d81-bad9-179da1230ff3 · inbound

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization cites this paper.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.611481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.611481Z digest=sha256:41a11ea7c6309afac119a1eb83d916b5bbeaf9eaaa2f427280af4f5353abc69d

Observation 5c91b463-b855-4a13-a201-70374c10804e · inbound

What Makes AI Research Replicable? Executable Knowledge Graphs as Scientific Knowledge Representations cites this paper.

What Makes AI Research Replicable? Executable Knowledge Graphs as Scientific Knowledge Representations MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T06:00:57.896540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T05:56:45.755148Z digest=sha256:8d9c30a2a7c666feab7e73ae82406a5963fcae36560ade66d76e9d74c367435b

Observation cdcbd144-bca0-4e94-9a72-46a5a6bc8518 · inbound

BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers? cites this paper.

BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T08:59:45.080256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:59:45.080256Z digest=sha256:6f064aee603b962bb9eee52ec73f039f9c7422f6729b65814622529d548c1912

Observation 2290bce9-fdff-431d-9d56-67285c1d0f87 · inbound

End-to-end PDDL Planning with Hardcoded and Dynamic Agents cites this paper.

End-to-end PDDL Planning with Hardcoded and Dynamic Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T23:38:41.675007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-16T23:34:20.411940Z digest=sha256:2c686975c8149580b2b24d0b57111582ef0561877bfd183022d2f21b4dabefe2

Observation 45a15d5e-76cb-48b3-8d58-1c69d11df068 · inbound

Automated Modernization of Machine Learning Engineering Notebooks for Reproducibility cites this paper.

Automated Modernization of Machine Learning Engineering Notebooks for Reproducibility MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:43:05.501625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:43:05.501625Z digest=sha256:9217367e503b17642203b26efbb2f60015479e09200141f7df14a6edf4a9eab8

Observation 89386d3a-3e0e-46dc-aef9-00b813d0419a · inbound

AceGRPO: Adaptive Curriculum Enhanced Group Relative Policy Optimization for Autonomous Machine Learning Engineering cites this paper.

AceGRPO: Adaptive Curriculum Enhanced Group Relative Policy Optimization for Autonomous Machine Learning Engineering MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T06:32:27.285326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T06:32:22.038300Z digest=sha256:1a8668618ca0b620602ea4190be7e9ee88dc71bf19ad002fed6695f071a09b55

Observation 73cfc36c-eea1-4951-b0be-8f9df34bd852 · inbound

iML: Executable, Problem-Grounded, and Broadly Exploratory Code-Driven AutoML cites this paper.

iML: Executable, Problem-Grounded, and Broadly Exploratory Code-Driven AutoML MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T23:25:33.220360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:25:33.220360Z digest=sha256:634003df787995e4f39ef7738535d755fa6f3458c44cc47844d10687a1115c3f

Observation fc560ce5-11dd-48aa-a399-67b7b6147fd9 · inbound

Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search cites this paper.

Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:50:12.253270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T17:49:46.383559Z digest=sha256:e1c31f3c9bed52736ce9daafc0f912b5d4b3ea7799689da6bc614aeb8d322ac5

Observation 4118b16a-0b55-44a3-a83a-f8d953c886b1 · inbound

AI Can Learn Scientific Taste cites this paper.

AI Can Learn Scientific Taste MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T18:14:52.244321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:14:52.244321Z digest=sha256:2431ba911be9ddc17fb432d4c55dc4b4609b3918aba73ac4aa6501376966dc53

Observation b09ef985-2053-4c07-b40f-c57aa800b28c · inbound

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments cites this paper.

From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:23:27.135290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T01:20:03.181903Z digest=sha256:a5f42668a5f822b81c6ed7258c9052ef244d4ca92fdf1e607d9a23fd4ce62462

Observation 2b64becb-1446-4d02-b3a9-1c144242b0f0 · inbound

TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration cites this paper.

TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T16:50:50.552104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T16:50:50.552104Z digest=sha256:03087355203d11f5c67199d676fbbf52c44c9219a5df1c19e1e024838e587829

Observation 548d6f2c-9474-40b7-b595-4bdfdf214bb8 · inbound

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation cites this paper.

Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:15:50.744720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T20:02:40.980538Z digest=sha256:0f166f23921292a5057c5deabe193535acf29bc51ca085030c6a923f4afb8622

Observation 1485e507-b83d-48cb-9c8d-5c7be3e37f64 · inbound

FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks cites this paper.

FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:30:49.944041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T19:49:32.983778Z digest=sha256:1ae268b68b2ab1904ba21386cf2a149e1de8135426e0c7c27c4237fdc32608d7

Observation 5401f046-ac6e-40d6-9445-7369c63ebad1 · inbound

In-Place Test-Time Training cites this paper.

In-Place Test-Time Training MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:49.058898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T19:07:47.174513Z digest=sha256:72a34e5655562b7bcbc48a9fceeee4ae7a7cae99fd3039460f2045f2ffe46699

Observation 8c76dcaa-1c19-4cae-90fd-690547baf310 · inbound

Pioneer Agent: Continual Improvement of Small Language Models in Production cites this paper.

Pioneer Agent: Continual Improvement of Small Language Models in Production MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:05:57.568178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T17:48:40.520740Z digest=sha256:bff8165e6dac1fa7beeba86a16f75eb0eed55e4e937b97f4e7fbcdf85308910f

Observation ea68e041-dfcd-410b-a9b8-b2602f0a0fb3 · inbound

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? cites this paper.

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:03.503133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:55:34.768853Z digest=sha256:3a861b575327799131fb66a9df8512d53eef4d117a7b94ca1a2d903802d98a82

Observation 78afe47a-b5d9-48f7-b914-0cc5619ce32b · inbound

Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks cites this paper.

Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:31:01.230008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:28:03.449453Z digest=sha256:2b85c9eac19ba2219feb8decd4e141d928b29f7df559f0920fc7802ba4592ed1

Observation 64982a15-6358-4aa5-9359-19eef36708b0 · inbound

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization cites this paper.

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:01:01.084341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:17:32.290531Z digest=sha256:17c0714c7a0bb5deaf615ca628e7c00caa2d0585a2b7b2da6c5bd979b8351598

Observation 680be77e-6b91-4632-bc70-5246ade11647 · inbound

TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration cites this paper.

TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:46:34.784533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T12:34:29.808503Z digest=sha256:d0282ed888e2e17afe6d0bdf6c580544912137e623b32c0ff259ec6a78844864

Observation a624fea8-8603-4239-beba-02a5a84879da · inbound

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale cites this paper.

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:01:13.394419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:59:01.010437Z digest=sha256:3f25aa8bb49138efe584a95d0bfdb56ee02ab5d7e0b4c142584bb3475095e095

Observation b08bc7e2-ff39-416d-9d0a-85ab7a1f7825 · inbound

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale cites this paper.

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-05T17:51:14.765804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-05T17:45:55.631459Z digest=sha256:f78c36bcfab7333a40bde8201ad92b8a19485c39e6a753b1136b809e4199cd32

Observation e2136a57-a6db-45e2-8324-eec869a62ade · inbound

Evaluation-driven Scaling for Scientific Discovery cites this paper.

Evaluation-driven Scaling for Scientific Discovery MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:26:07.565713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T03:39:52.204043Z digest=sha256:664fb1f0c499c672881286bfe8edd45f82138baac3175ac1e3f604b35bb2d706

Observation 80be0d3b-badd-4671-9d41-49cbf429c6c6 · inbound

On Benchmark Hacking in ML Contests: Modeling, Insights and Design cites this paper.

On Benchmark Hacking in ML Contests: Modeling, Insights and Design MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:21:12.738010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T09:18:34.040257Z digest=sha256:7f32dc507e0787ccdaeab0ed151540022ddc1f3c8ecbdc928a67ee2cc8d0afad

Observation 7cc74e0f-2b06-412a-a1ad-f53a219a33d8 · inbound

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning cites this paper.

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:16:09.185864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T20:22:58.061772Z digest=sha256:7f54523c285b1857e01392fb4b0b76c77f5c80d68bf869204c615cb2f4eaabdf

Observation fc6af751-d40f-4062-b432-34e55189aefd · inbound

AcademiClaw: When Students Set Challenges for AI Agents cites this paper.

AcademiClaw: When Students Set Challenges for AI Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:29.695068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T19:24:29.696454Z digest=sha256:2caba41bab64c0dd4c8e7f726fcfdba2b776fca0a26c245e368f6a85cc9ab638

Observation 313fca62-7306-40cf-9175-65e1b277c37b · inbound

PACEvolve++: Improving Test-time Learning for Evolutionary Search Agents cites this paper.

PACEvolve++: Improving Test-time Learning for Evolutionary Search Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:56.387556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:54:39.349292Z digest=sha256:16035db9d6e4e50b53614483d468c9b4268a05d3748d8da014087910bf0c4e37

Observation 8f610c96-c7e3-462e-9038-b3956a66b57a · inbound

TeamBench: Evaluating Agent Coordination under Enforced Role Separation cites this paper.

TeamBench: Evaluating Agent Coordination under Enforced Role Separation MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:55.525903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T00:55:51.358828Z digest=sha256:97d18bb9247fb2196410252ec835f08a43ce8c9753f34888460a04de226aeeb8

Observation eb03d3b3-e7ad-4c7b-baa3-d79464aaf430 · inbound

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces cites this paper.

OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 142

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T03:01:18.592393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T02:57:15.521594Z digest=sha256:70b498f46e783b6fb65065c855d81fab578888e3011905f26144456e625924b9

Observation 6e96672d-fb47-41b1-87d5-efd84d4ccf42 · inbound

DataMaster: Data-Centric Autonomous AI Research cites this paper.

DataMaster: Data-Centric Autonomous AI Research MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:36:24.801807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T04:09:40.731228Z digest=sha256:82a3b63ad60cc8fdabe5f6ccb6031b9ce2692de30d6d23119aa56a398da248e2

Observation f1c856d2-85f2-4075-a859-4985e291ed36 · inbound

DataMaster: Data-Centric Autonomous AI Research cites this paper.

DataMaster: Data-Centric Autonomous AI Research MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:12:59.133651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T21:11:22.202161Z digest=sha256:1cf3686087db609f21794d28bdfa8da237eb7a053515f3a8a6389c3c7202039c

Observation cfd3bf6c-cbdc-4ae8-b32b-7bda4deff2cd · inbound

SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces cites this paper.

SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:12:17.914966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T05:10:04.763149Z digest=sha256:486cc04a6025edacc5e062cab77e33310b257a6539571148e58b2b292f68e58d

Observation f1f887bd-a17f-4a92-b2ef-065878966ef9 · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:32:56.909283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:ddecf829d3d962a7343322d0f84899562f0d414206a1eb1c2678ca132b66e8da

Observation 4ba5177d-4501-49f5-b6e8-2acd5155a599 · inbound

Europe and the Geopolitics of AGI: The Need for a Preparedness Plan cites this paper.

Europe and the Geopolitics of AGI: The Need for a Preparedness Plan MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-14T17:49:23.579524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-14T17:43:38.233339Z digest=sha256:90669f6bd07d76fe23c40a76fbcbb4e4ddbe294e10d6b761270fd0c2883088e4

Observation df6e7b88-e280-4dcd-8151-e0557fa17296 · inbound

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction cites this paper.

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T06:05:06.706952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-15T06:04:03.605898Z digest=sha256:e1f268e9f955aa409bec0decc134a3f18a58b675c39076b2f2b2bb22cb3f74b5

Observation a239fdf7-7476-420a-abdd-f2561ed0e672 · inbound

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale cites this paper.

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:03:28.882759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T02:02:25.597640Z digest=sha256:18440aca3eff3384d5b827e91d137a6b8b16db162e5b18a17c90f599085e718c

Observation 1516a089-4cdb-4397-a03d-9b2f3b7c0b19 · inbound

Graphs of Research: Citation Evolution Graphs as Supervision for Research Idea Generation cites this paper.

Graphs of Research: Citation Evolution Graphs as Supervision for Research Idea Generation MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T20:55:04.045232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:52:32.567107Z digest=sha256:707c5ac2cb498ccde5aa490766dca6c035e23520574d2a9d87a034db02cc1e68

Observation 97b19e26-c772-4343-9579-bc8081f1aac1 · inbound

SMCEvolve: Principled Scientific Discovery via Sequential Monte Carlo Evolution cites this paper.

SMCEvolve: Principled Scientific Discovery via Sequential Monte Carlo Evolution MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:22:39.499555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T16:22:12.940259Z digest=sha256:f288c08bf117632f15bf9544585ab020f306103828335c089b37aff931d16698

Observation a711806f-cefe-41e7-aff4-214ea632e467 · inbound

BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks cites this paper.

BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-19T19:32:43.762882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T19:31:32.334837Z digest=sha256:87cca825f1c76120b4507ddea5f7e9f205751ce14979a45cea7e9b61ed9c8429

Observation 92bd12e4-4819-4baa-8d3b-88a8a9a2219f · inbound

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility cites this paper.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T20:03:44.137430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:5541c559c31f2699fb20bb8ec59734a7b1513a68987f80c4e1c5418658c1456a

Observation f29b7ab8-9d8a-4738-8d69-a85e6ac2b3f2 · inbound

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics cites this paper.

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-20T14:28:21.432887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T14:25:15.565386Z digest=sha256:f9fbd6f58457e981f325a0e4e49d0e7e61cd4a287f30de4dc4e493a73d0f3887

Observation 26dca1fb-2a0c-44d7-84b5-6d5f62d7a0c3 · inbound

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics cites this paper.

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:05:00.903660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T19:00:30.961402Z digest=sha256:ec0049aaf738feceaf84ef4b3100f05e2c9e3ff82bacdefe04f94c60b983ef33

Observation 2d271753-d42e-4655-b891-26878c699197 · inbound

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents cites this paper.

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T23:12:51.589025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-19T23:09:56.130802Z digest=sha256:9d5cb1899413345356dbceefe5ec78096d62da1968d24b7698c9f8ed64eb2d97

Observation ae64da6d-0fa5-465b-bcb9-14f40f62d585 · inbound

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents cites this paper.

DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T13:08:17.865759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T13:05:38.485058Z digest=sha256:4e61bbdfd097e79cb650d2ad377e3638aff935dc0c6dc93c8d66a6e04079e676

Observation 47cadefa-3de7-42a4-8968-5415fc8a6fe2 · inbound

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games cites this paper.

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:28:17.264615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T12:24:06.062957Z digest=sha256:9c67689ad8adde8461c5d86f44dc79444f0518e3661130aada25e5b9007c5cb1

Observation 00bc18d5-ec31-494c-8b6c-b2308cea62e8 · inbound

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games cites this paper.

WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:45:23.240288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-25T05:45:04.573722Z digest=sha256:42208392be9d941b0cbfb76987f3d0bb14d0612838b7e9e1cc342efa43d7c33e

Observation f4293ff5-e60d-4736-9d3e-1c6fdbe8ad82 · inbound

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows cites this paper.

DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:18:11.752386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T10:16:38.920528Z digest=sha256:2b62122f4d5981e542154ee802c9abc715ce86c23d8075de606838e54b76b85d

Observation 5f4e81b0-4561-4a35-b96a-98a815f18d14 · inbound

How Far Are We From True Auto-Research? cites this paper.

How Far Are We From True Auto-Research? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T09:58:11.265272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T09:56:16.160551Z digest=sha256:e64c037a62292d62469899e4acf32eefcfa047f65857359e81ca484ec094a957

Observation 41b8d2f9-b593-42da-a0a8-1c5783be82f9 · inbound

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents cites this paper.

What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:43:05.771032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:41:23.712146Z digest=sha256:a27a815ac60a583628486707052f7b12bece34947a36a20cd25ec3d088e7e296

Observation 02af144c-51c0-4476-bf9c-11c9f6a69597 · inbound

AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration cites this paper.

AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-20T05:23:03.562193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:22:56.341333Z digest=sha256:c19a7ca0f89c63009d5e1cbf6aa2fd3e3bf8cefeb06b93c6a79ce86bf198a5e2

Observation 1d4d4ea8-3b43-47c9-9d54-b10f82bdb35e · inbound

AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration cites this paper.

AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-01T15:05:47.884299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:07:57.905347Z digest=sha256:8b7147f2e147fa7f1bb62339b88f7f5c338e7a33c281cf3c610acee7832c718e

Observation e2ee0fa7-77a2-43b0-be19-d29bd853ec47 · inbound

What Do Evolutionary Coding Agents Evolve? cites this paper.

What Do Evolutionary Coding Agents Evolve? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-05-20T03:48:03.327687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T03:44:18.658541Z digest=sha256:136539847ef4aa5b5433e5541508f6bc676d67de94cb1e467994bc88d232d988

Observation bbab3210-536e-4bbe-ae63-fdb65a033849 · inbound

WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents cites this paper.

WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 6

Resolution
malformed identifier
local_arxiv, observed 2026-05-21T07:54:49.600155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T07:54:42.190296Z digest=sha256:e79096e1e1338fe31a663d41ac847f9d111584a83e9628ef9418863f7b1b1bec

Observation 104b269d-9077-47f1-80db-b816e75ff523 · inbound

WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents cites this paper.

WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 6

Resolution
malformed identifier
local_arxiv, observed 2026-06-30T18:14:59.143256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:07:22.464758Z digest=sha256:e00c27a69c1ebc79191180a9381cbd73606be11eee48ebf8150ecd25576f6b63

Observation 14b3caf7-8f5f-4193-9edf-63d6d36a4d44 · inbound

Declarative Data Services: Structured Agentic Discovery for Composing Data Systems cites this paper.

Declarative Data Services: Structured Agentic Discovery for Composing Data Systems MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-21T05:19:39.287742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T05:16:46.549921Z digest=sha256:0074b7ee33d87a55e0d3272c7b140780475a595f7605bf16692d215f9dc9c147

Observation 7c85de97-2215-434f-8825-ee600e61f827 · inbound

Declarative Data Services: Structured Agentic Discovery for Composing Data Systems cites this paper.

Declarative Data Services: Structured Agentic Discovery for Composing Data Systems MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-01T15:05:48.125602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T17:46:34.881102Z digest=sha256:630ab62c1f3ee8d863b6d5764dd4dda59c283c2dcc4a22334f406d55443358ef

Observation 85350ddd-64b2-4a6c-b328-e22f820dc274 · inbound

IdleSpec: Exploiting Idle Time via Speculative Planning for LLM Agents cites this paper.

IdleSpec: Exploiting Idle Time via Speculative Planning for LLM Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:24:40.483673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T06:21:56.971204Z digest=sha256:613f4e5fd4e89ecdc86f9e0f0742234ed7f8df9cd4fc2697e92042c7d742727f

Observation b4578c69-a663-4e2f-ae8e-ab40b82ddf6a · inbound

AION: Next-Generation Tasks and Practical Harness for Time Series cites this paper.

AION: Next-Generation Tasks and Practical Harness for Time Series MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T11:54:38.596441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T11:24:42.704735Z digest=sha256:2eec5b87f8f9fccbf8f369fa3685d5db242113a6b18b264f41bf88c65989963a

Observation 1fc5fd31-db11-4b37-89bb-27d93630b493 · inbound

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence cites this paper.

ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-29T21:23:59.052296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-29T21:19:03.281629Z digest=sha256:725e4d88677a12e6d8d9e4953eeebbb7fc12b482f8ba578d9dfe6892ee2573ab

Observation 11e75caa-4fdc-43c9-8be3-57473bdf22ff · inbound

SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones? cites this paper.

SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:23:15.805279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T08:13:42.770740Z digest=sha256:a82a094e57686f5c540c0682233b7e4d9c4a6ed330f20f47451487429dd1b975

Observation d4cb0844-9fd9-46a9-a2ab-7445ea2a6490 · inbound

Business Utility of Large Language Models as Exploratory Data Analysis Agents cites this paper.

Business Utility of Large Language Models as Exploratory Data Analysis Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T23:35:06.995731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T23:25:17.416071Z digest=sha256:1d739741f5c810fe324a989a93208472b0f8f47a4b1d6904c811a2aebd2b39a0

Observation 7c39c3e1-b733-4c05-a127-d018a7747b8f · inbound

VESTA: Visual Exploration with Statistical Tool Agents cites this paper.

VESTA: Visual Exploration with Statistical Tool Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:56:10.520983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T21:58:11.339217Z digest=sha256:16e969d21b1e3ccae63b469642e181877459d2225c40ae24e5dcaa2420a430ec

Observation ac48592d-2c15-40fc-9674-b48ffeda9bf4 · inbound

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models cites this paper.

AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T23:06:21.108710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T14:38:40.017263Z digest=sha256:0cd05be3497ee83194f98db5635ebe1f0afcffc6f717a7797f23a3e40ef21097

Observation ba8327f8-3fcf-45f8-97e5-32533a365df6 · inbound

Can Generalist Agents Automate Data Curation? cites this paper.

Can Generalist Agents Automate Data Curation? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:56:35.137529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T09:32:39.361415Z digest=sha256:e4d1eadacc4f2c6ed07166f99262faa0f3e7f9730cfdb88320eb44f7a99f37e5

Observation bb561dac-0119-4a07-b03e-764941fb9271 · inbound

The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development? cites this paper.

The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-02T07:56:47.678461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T06:29:42.665765Z digest=sha256:93a82ed79dd1cac5b2d2068c5f8beb094d4475f170dfd6fca315f303fdc269d2

Observation c2edd2e1-602e-4a3b-aa4e-74a69caf9de2 · inbound

Towards Persistent Case-Based Memory for Autonomous Data Science: A CBR-Augmented R&D-Agent with a Locally Deployable Small Language Model cites this paper.

Towards Persistent Case-Based Memory for Autonomous Data Science: A CBR-Augmented R&D-Agent with a Locally Deployable Small Language Model MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-02T10:06:51.727212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T05:19:14.975753Z digest=sha256:888fbe41c3ce01cbd98889d78add1ed511539c15f2e3a8bcc6c7cfc85308eedf

Observation 53dc74d2-7240-4d29-8666-bc7c511b3a01 · inbound

Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories cites this paper.

Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:37:37.769214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T13:43:02.919248Z digest=sha256:d89c104b621646029091a89f5015bacfb5ee5f7acf57314535b87350920d739c

Observation 63a7b378-6de1-4e7d-8e33-4a76c8fd6d57 · inbound

Search Discipline for Long-Horizon Research Agents cites this paper.

Search Discipline for Long-Horizon Research Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-27T13:30:57.161480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T12:51:07.758929Z digest=sha256:e65b3febe8d0653e5a9ee119f9e1b35fa74b8e39efd62e85723c13b8cbadc88f

Observation 63ca5120-9179-411c-96a6-ee7e58da9f14 · inbound

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement cites this paper.

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 115

Resolution
verified exact
local_arxiv, observed 2026-06-27T09:40:47.023998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T09:34:41.800309Z digest=sha256:6bd99a2653a391b1ada494424e68fa81728e18c3ba02b64072d555c97a5154c7