Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T19:11:20.600633Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 100 inbound Pith citation observations for arXiv:2410.07095.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T19:11:20.600633Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:57:36.403917Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
36 of 36 outbound references displayed
External citation measurements
9
pith, observed 2026-08-05T02:28:24.338817Z
Observation 10058669-c66c-4d10-931e-d0de89239b7b · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Anthropic's Responsible Scaling Policy , Version 1.0, September 2023
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b41e221d-f8cb-4ff6-9116-74d91283c89c · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Program Synthesis with Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 69eb685f-89f0-40b6-97f7-cdb6cb72e6bd · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Quantifying Memorization Across Neural Language Models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 31abeee7-f11b-4692-9173-46311c65a0b1 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Evaluating Large Language Models Trained on Code
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2285312-d1d7-4a8c-b93e-ec90e7c1b685 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Cognition Introducing Devin , the first AI software engineer, March 2024
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 063ad94b-622d-44f6-9392-3df1945fcdd7 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Openvaccine: Covid-19 mrna vaccine degradation prediction
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bfc1c359-73ad-4401-95e5-ac006cf958c6 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering ConStat: Performance-Based Contamination Detection in Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a4c73d4-f7b1-4862-912d-044d64b9f07f · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering GitHub Copilot Workspace : Welcome to the Copilot -native developer environment, April 2024
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91bcb36d-b566-4dc5-85bc-cba24f532b4f · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Code Droid Technical Report , June 2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ad5e1cc-4025-4189-834a-920f3b668c5c · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering AgentQuest: A Modular Benchmark Framework to Measure Progress and Improve LLM Agents
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b17d8e9a-d934-49d7-88d9-6219f8ef0a17 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Frontier Safety Framework , May 2024
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43b31cb2-ca58-4eef-8763-1347390b8110 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Measuring Coding Challenge Competence With APPS
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a595fff-28c7-4351-a9f6-de5ac44b11e9 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering AgentCoder: Multi-Agent-based Code Generation with Iterative Testing and Optimisation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0ced4d7-6faf-4949-b66b-16f20112a0c3 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering MLAgentBench : Evaluating Language Agents on Machine Learning Experimentation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8db45d11-7b65-4b61-804c-b0b81bf12438 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8af70357-a0e3-45bc-929d-a3a2381a43b2 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 556b8617-b887-420c-9d9f-a9e5dc970094 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering DSBench: How Far Are Data Science Agents from Becoming Data Science Experts?
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4fc3763e-a96e-4bdf-9174-369cbc7c776e · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Kaggle Progression System Kaggle
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e80d8849-0494-4cf5-97a1-9a33bc0c445a · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Research: quantifying GitHub Copilot ’s impact on developer productivity and happiness, September 2022
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0d7dd57a-ebc5-419b-a74b-fff200b39d49 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering AI Agents That Matter
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a88b45dc-14c6-4881-8967-babcd2da0a5d · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Mankowitz, Esme Sutherland Robson, Pushmeet Kohli, Nando De Freitas, Koray Kavukcuoglu, and Oriol Vinyals
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 129dc745-ebc3-4563-ac56-0c49e86d4d97 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering AgentBench: Evaluating LLMs as Agents
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d67ec729-61b6-428c-999c-2ea8b6b34c40 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Vesuvius challenge - ink detection
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 740a3b7a-a86a-4816-8326-2bc938fef9e8 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Discovering and exploring cases of educational source code plagiarism with Dolos
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08f8be6f-122d-49e2-832e-4b75c9d471ac · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering GAIA: a benchmark for General AI Assistants
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eabc68ac-2f55-445e-9fca-cbee9fb43411 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Preparedness Framework , December 2023
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fb6f2bfa-73e9-4738-933f-8410d816ac27 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Introducing Weco AIDE , April 2024
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 230b2e79-42ba-4507-bf2f-bb3b8e302bd2 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2d1aebf-c91e-43e6-bca1-a6111eda7e80 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering OpenHands: An Open Platform for AI Software Developers as Generalist Agents
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ef46fc68-42e6-4160-adc9-00f48e3a7f2e · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering The shift from models to compound ai systems
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52ade589-ecb0-4e0a-8f4b-9fb84951d706 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering AutoCodeRover: Autonomous Program Improvement
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a36e9827-74bb-41e2-9818-050d571d840b · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Can GPT-4 Perform Neural Architecture Search?
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6f895805-6347-4eac-800d-13ca628d0fbd · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering write newline
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f847085e-e9ab-47cc-9d4b-77a13c06e902 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering @esa (Ref
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3db40eef-97a1-40e9-bc8f-41c2a5fa4593 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ccc572ff-0c54-41e8-8e34-f6e8b3856200 · outbound
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0949488f-6b45-47ab-99f4-ccf2509dce71 · inbound
Frontier Models are Capable of In-context Scheming MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c8a8636-7558-4f7c-b51b-62ff7e988de6 · inbound
Humanity's Last Exam MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation baf3ef74-fbc2-4445-8a92-07d2c91f88e8 · inbound
Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f883961b-a295-4de7-95c4-00bcd68091d6 · inbound
Large Language Model Agent: A Survey on Methodology, Applications and Challenges MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 144
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 46a8763d-d797-452b-adba-009b1d51cbe2 · inbound
From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 127
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 800e8d9e-d336-4dbf-b08d-e2ba52b7c0aa · inbound
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94911a56-4025-4428-a8db-5db3c2374852 · inbound
Large Language Models for Planning: A Comprehensive and Systematic Survey MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db9ab53c-6279-496e-9327-d61a9c33d19d · inbound
RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9cd5b506-3428-41fe-843e-60bc3e7b709a · inbound
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 089f937e-d284-4d45-9032-4f160e40483b · inbound
AI Scientists Fail Without Strong Implementation Capability MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ed01d53-5a9b-436b-82be-455d36c104a1 · inbound
VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7765159c-2546-4784-b802-f379cc58522b · inbound
TextAtari: 100K Frames Game Playing with Language Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00a96ab0-57aa-4f08-a269-a73cf4789e00 · inbound
Agentomics-ML: Autonomous Machine Learning Experimentation Agent for Genomic and Transcriptomic Data MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b37b2185-8b21-4127-a86d-11fa9a8c23d5 · inbound
OPT-BENCH: Evaluating LLM Agent on Large-Scale Search Spaces Optimization Problems MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9aa265e-c0c0-4265-b9a9-2354de16d98e · inbound
Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 172
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81445514-0b9c-4510-bfc1-5ad5d93c31d6 · inbound
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1333024c-de16-4666-82d9-44d41a199367 · inbound
Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccc444aa-7de8-44e9-b689-0b6714a020eb · inbound
RExBench: Can coding agents autonomously implement AI research extensions? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dbeba95d-0dd7-46bf-9482-91eef5049960 · inbound
DABstep: Data Agent Benchmark for Multi-step Reasoning MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fe5516d-1166-42c1-9628-e810171168ca · inbound
AI4Research: A Survey of Artificial Intelligence for Scientific Research MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3080b94-bee8-43c4-a758-4a0b10a01473 · inbound
Establishing Best Practices for Building Rigorous Agentic Benchmarks MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 150e9c98-f8c7-42fb-922b-8887deee6850 · inbound
AblationBench: Evaluating Automated Planning of Ablations in Empirical AI Research MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc15c45f-a2ad-4758-adad-4896e7be1a57 · inbound
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e26fafb-125c-42f3-9fa6-98a8e1106f7f · inbound
A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96bba6a7-7066-4c86-8529-f0aae9c18d72 · inbound
How Far Are AI Scientists from Changing the World? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c8caff6-8dcb-45c4-9c85-80b1ad948731 · inbound
TextQuests: How Good are LLMs at Text-Based Video Games? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b912a33-b4a6-4a11-bbfa-cfa0ac414b7e · inbound
Preliminary suggestions for rigorous GPAI model evaluations MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46c82f76-f455-4970-8219-272ba1e8053b · inbound
KompeteAI: Accelerated Autonomous Multi-Agent System for End-to-End Pipeline Generation for Machine Learning Problems MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e0f69ee3-1813-43c4-9ef2-efff9eb485ec · inbound
Reliable Weak-to-Strong Monitoring of LLM Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9966b5d1-224c-4fea-9a94-78f61bc54f1f · inbound
Reinforcement Learning for Machine Learning Engineering Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc36038c-3468-43e8-a1e8-9d5666dab71a · inbound
MachineLearningLM: Scaling Many-shot In-context Learning via Continued Pretraining MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 66fb61dd-5ec4-455b-9983-47a6a7b6d04c · inbound
A Survey of Reinforcement Learning for Large Reasoning Models MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8efd4654-b1fb-4d81-bad9-179da1230ff3 · inbound
Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c91b463-b855-4a13-a201-70374c10804e · inbound
What Makes AI Research Replicable? Executable Knowledge Graphs as Scientific Knowledge Representations MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cdcbd144-bca0-4e94-9a72-46a5a6bc8518 · inbound
BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2290bce9-fdff-431d-9d56-67285c1d0f87 · inbound
End-to-end PDDL Planning with Hardcoded and Dynamic Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 45a15d5e-76cb-48b3-8d58-1c69d11df068 · inbound
Automated Modernization of Machine Learning Engineering Notebooks for Reproducibility MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89386d3a-3e0e-46dc-aef9-00b813d0419a · inbound
AceGRPO: Adaptive Curriculum Enhanced Group Relative Policy Optimization for Autonomous Machine Learning Engineering MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73cfc36c-eea1-4951-b0be-8f9df34bd852 · inbound
iML: Executable, Problem-Grounded, and Broadly Exploratory Code-Driven AutoML MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc560ce5-11dd-48aa-a399-67b7b6147fd9 · inbound
Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4118b16a-0b55-44a3-a83a-f8d953c886b1 · inbound
AI Can Learn Scientific Taste MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b09ef985-2053-4c07-b40f-c57aa800b28c · inbound
From Pixels to Digital Agents: An Empirical Study on the Taxonomy and Technological Trends of Reinforcement Learning Environments MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2b64becb-1446-4d02-b3a9-1c144242b0f0 · inbound
TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 548d6f2c-9474-40b7-b595-4bdfdf214bb8 · inbound
Multilingual Prompt Localization for Agent-as-a-Judge: Language and Backbone Sensitivity in Requirement-Level Evaluation MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1485e507-b83d-48cb-9c8d-5c7be3e37f64 · inbound
FrontierFinance: A Long-Horizon Computer-Use Benchmark of Real-World Financial Tasks MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5401f046-ac6e-40d6-9445-7369c63ebad1 · inbound
In-Place Test-Time Training MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c76dcaa-1c19-4cae-90fd-690547baf310 · inbound
Pioneer Agent: Continual Improvement of Small Language Models in Production MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ea68e041-dfcd-410b-a9b8-b2602f0a0fb3 · inbound
SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 78afe47a-b5d9-48f7-b914-0cc5619ce32b · inbound
Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 64982a15-6358-4aa5-9359-19eef36708b0 · inbound
Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 680be77e-6b91-4632-bc70-5246ade11647 · inbound
TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a624fea8-8603-4239-beba-02a5a84879da · inbound
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b08bc7e2-ff39-416d-9d0a-85ab7a1f7825 · inbound
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2136a57-a6db-45e2-8324-eec869a62ade · inbound
Evaluation-driven Scaling for Scientific Discovery MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 80be0d3b-badd-4671-9d41-49cbf429c6c6 · inbound
On Benchmark Hacking in ML Contests: Modeling, Insights and Design MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7cc74e0f-2b06-412a-a1ad-f53a219a33d8 · inbound
Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc6af751-d40f-4062-b432-34e55189aefd · inbound
AcademiClaw: When Students Set Challenges for AI Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 313fca62-7306-40cf-9175-65e1b277c37b · inbound
PACEvolve++: Improving Test-time Learning for Evolutionary Search Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8f610c96-c7e3-462e-9038-b3956a66b57a · inbound
TeamBench: Evaluating Agent Coordination under Enforced Role Separation MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb03d3b3-e7ad-4c7b-baa3-d79464aaf430 · inbound
OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 142
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e96672d-fb47-41b1-87d5-efd84d4ccf42 · inbound
DataMaster: Data-Centric Autonomous AI Research MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f1c856d2-85f2-4075-a859-4985e291ed36 · inbound
DataMaster: Data-Centric Autonomous AI Research MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cfd3bf6c-cbdc-4ae8-b32b-7bda4deff2cd · inbound
SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f1f887bd-a17f-4a92-b2ef-065878966ef9 · inbound
Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4ba5177d-4501-49f5-b6e8-2acd5155a599 · inbound
Europe and the Geopolitics of AGI: The Need for a Preparedness Plan MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df6e7b88-e280-4dcd-8151-e0557fa17296 · inbound
Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a239fdf7-7476-420a-abdd-f2561ed0e672 · inbound
FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1516a089-4cdb-4397-a03d-9b2f3b7c0b19 · inbound
Graphs of Research: Citation Evolution Graphs as Supervision for Research Idea Generation MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97b19e26-c772-4343-9579-bc8081f1aac1 · inbound
SMCEvolve: Principled Scientific Discovery via Sequential Monte Carlo Evolution MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a711806f-cefe-41e7-aff4-214ea632e467 · inbound
BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 92bd12e4-4819-4baa-8d3b-88a8a9a2219f · inbound
MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f29b7ab8-9d8a-4738-8d69-a85e6ac2b3f2 · inbound
FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26dca1fb-2a0c-44d7-84b5-6d5f62d7a0c3 · inbound
FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2d271753-d42e-4655-b891-26878c699197 · inbound
DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ae64da6d-0fa5-465b-bcb9-14f40f62d585 · inbound
DiagEval: Trajectory-Conditioned Diagnosis for Reliable Software Evaluation with GUI Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47cadefa-3de7-42a4-8968-5415fc8a6fe2 · inbound
WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 00bc18d5-ec31-494c-8b6c-b2308cea62e8 · inbound
WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f4293ff5-e60d-4736-9d3e-1c6fdbe8ad82 · inbound
DecisionBench: A Benchmark for Emergent Delegation in Long-Horizon Agentic Workflows MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f4e81b0-4561-4a35-b96a-98a815f18d14 · inbound
How Far Are We From True Auto-Research? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 41b8d2f9-b593-42da-a0a8-1c5783be82f9 · inbound
What and When to Distill: Selective Hindsight Distillation for Multi-Turn Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 02af144c-51c0-4476-bf9c-11c9f6a69597 · inbound
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d4d4ea8-3b43-47c9-9d54-b10f82bdb35e · inbound
AutoResearchClaw: Self-Reinforcing Autonomous Research with Human-AI Collaboration MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2ee0fa7-77a2-43b0-be19-d29bd853ec47 · inbound
What Do Evolutionary Coding Agents Evolve? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bbab3210-536e-4bbe-ae63-fdb65a033849 · inbound
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 104b269d-9077-47f1-80db-b816e75ff523 · inbound
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 14b3caf7-8f5f-4193-9edf-63d6d36a4d44 · inbound
Declarative Data Services: Structured Agentic Discovery for Composing Data Systems MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c85de97-2215-434f-8825-ee600e61f827 · inbound
Declarative Data Services: Structured Agentic Discovery for Composing Data Systems MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 85350ddd-64b2-4a6c-b328-e22f820dc274 · inbound
IdleSpec: Exploiting Idle Time via Speculative Planning for LLM Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4578c69-a663-4e2f-ae8e-ab40b82ddf6a · inbound
AION: Next-Generation Tasks and Practical Harness for Time Series MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1fc5fd31-db11-4b37-89bb-27d93630b493 · inbound
ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11e75caa-4fdc-43c9-8be3-57473bdf22ff · inbound
SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4cb0844-9fd9-46a9-a2ab-7445ea2a6490 · inbound
Business Utility of Large Language Models as Exploratory Data Analysis Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c39c3e1-b733-4c05-a127-d018a7747b8f · inbound
VESTA: Visual Exploration with Statistical Tool Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ac48592d-2c15-40fc-9674-b48ffeda9bf4 · inbound
AutoMedBench: Towards Medical AutoResearch with Agentic AI Models MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba8327f8-3fcf-45f8-97e5-32533a365df6 · inbound
Can Generalist Agents Automate Data Curation? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bb561dac-0119-4a07-b03e-764941fb9271 · inbound
The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development? MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c2edd2e1-602e-4a3b-aa4e-74a69caf9de2 · inbound
Towards Persistent Case-Based Memory for Autonomous Data Science: A CBR-Augmented R&D-Agent with a Locally Deployable Small Language Model MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53dc74d2-7240-4d29-8666-bc7c511b3a01 · inbound
Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63a7b378-6de1-4e7d-8e33-4a76c8fd6d57 · inbound
Search Discipline for Long-Horizon Research Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 63ca5120-9179-411c-96a6-ee7e58da9f14 · inbound
Toward Generalist Autonomous Research via Hypothesis-Tree Refinement MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.