Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 76 inbound Pith citation observations for arXiv:2310.03302.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T10:59:33.392950Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-05T17:51:14.743354Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation e22682f9-6d22-4950-8ab2-4df6f2cbf0ab · inbound
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 164
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dd32d7fe-6786-494d-9a52-773dd83e9a5c · inbound
Probing the Capacity of Language Model Agents to Operationalize Disparate Experiential Context Despite Distraction MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 482e5e49-c928-48db-80dd-957436abcce4 · inbound
RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e654b03-1055-4c8e-81a8-94424a24131e · inbound
How Well Can Modern LLMs Act as Agent Cores in Radiology Environments? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de8015ab-8bd0-4a98-80ea-3f67f0c5bd17 · inbound
A Review of Multimodal Explainable Artificial Intelligence: Past, Present and Future MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 243
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2362b601-a1e3-4c51-bafb-9189be1bc3b5 · inbound
A Survey on LLM-based Multi-Agent System: Recent Advances and New Frontiers in Application MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b1e7f17-7f48-4c2c-952d-dd87aa842518 · inbound
HackerRank-ASTRA: Evaluating Correctness & Consistency of Large Language Models on cross-domain multi-file project problems MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06221fec-513e-4339-bd7c-6c3668e36722 · inbound
IRIS: Interactive Research Ideation System for Accelerating Scientific Discovery MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33a3b63b-61a8-4645-be24-bb57c4251e68 · inbound
MLZero: A Multi-Agent System for End-to-end Machine Learning Automation MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9e4d109-9fac-4863-ae07-f3aa906b5f2d · inbound
MM-Agent: LLM as Agents for Real-world Mathematical Modeling Problem MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 497203fd-20c9-4fac-8251-62b635e883c4 · inbound
CIKT: A Collaborative and Iterative Knowledge Tracing Framework with Large Language Models MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36be6c58-142b-41d6-af9d-7998f68f4ee6 · inbound
RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cc0eecf-07f5-42c9-8155-d58f6c755ec1 · inbound
Co-Saving: Resource Aware Multi-Agent Collaboration for Software Development MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdc97199-e562-4ef9-abd7-04dff8b60230 · inbound
Predicting Empirical AI Research Outcomes with Language Models MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0913741-ea1a-45d3-8443-1b3869f10af8 · inbound
Agentomics-ML: Autonomous Machine Learning Experimentation Agent for Genomic and Transcriptomic Data MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9240ddd9-22e9-456f-af9c-378e3fe489c8 · inbound
Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 824bc8cf-a92d-44e2-b1bf-97ae87bf6406 · inbound
ML-Master: Towards AI-for-AI via Integration of Exploration and Reasoning MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34a0a56a-b63d-4394-94e3-d06fd10fc306 · inbound
Deep Research Agents: A Systematic Examination And Roadmap MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4af3d1df-4469-45d9-87ff-caa730185580 · inbound
THE-Tree: Can Tracing Historical Evolution Enhance Scientific Verification and Reasoning? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43152c97-83f1-4231-afce-e6da05582938 · inbound
Agent Identity Evals: Measuring Agentic Identity MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a62dac5d-51ee-4299-b37e-9688afd25f0a · inbound
KompeteAI: Accelerated Autonomous Multi-Agent System for End-to-End Pipeline Generation for Machine Learning Problems MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 47afa8af-5f77-41d0-8d91-daac13edd68a · inbound
Reinforcement Learning for Machine Learning Engineering Agents MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 227d29d8-f707-437a-8499-71de11a6b9e8 · inbound
Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19e65b4c-d119-4f78-a9e3-fd558d2d7f9d · inbound
Can We Predict Before Executing Machine Learning Agents? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b4090343-b136-41b7-b2cd-14bd4c35b2d0 · inbound
AceGRPO: Adaptive Curriculum Enhanced Group Relative Policy Optimization for Autonomous Machine Learning Engineering MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d2b289d4-6ea3-4b36-b315-08cd9ed0c454 · inbound
Reasoning as Gradient: Scaling MLE Agents Beyond Tree Search MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dd51afc6-27d5-4d56-9342-87524914d2e2 · inbound
Pioneer Agent: Continual Improvement of Small Language Models in Production MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cd3f2020-b5ed-4450-bbe0-7742b0f2ee6c · inbound
SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 223bdcbf-c044-43d3-a82e-c276b1d41722 · inbound
TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8e9faee6-5976-4c31-b9fa-35f042e7b416 · inbound
LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 147f34cf-774e-4056-a5f2-7dbc53eb5c5c · inbound
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 086c5d67-6e5b-4777-889d-4c4053514e26 · inbound
EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9faebeff-3442-4c9e-920a-b387b6f37505 · inbound
Read, Grep, and Synthesize: Diagnosing Cross-Domain Seed Exposure for LLM Research Ideation MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 83d892e4-5187-48a9-84a4-4ec6ee350ed5 · inbound
Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 72e9c313-cd13-4bd5-a23a-2b5b36a1f247 · inbound
Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 6ab426c4-9ec5-48d3-b3f7-f1bbf567ca0d · inbound
Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7d92aaab-5869-465d-b174-a9f8e419e1e8 · inbound
BioXArena: Benchmarking LLM Agents on Multi-Modal Biomedical Machine Learning Tasks MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation faa7c95a-d138-42c4-95fb-b654ee60a4e1 · inbound
MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d0dab6a4-af03-438e-b8d0-050a0e77b66f · inbound
FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 65df9179-0ebe-4a8f-acc9-3fa4c16117b7 · inbound
FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9d3d2142-eb7d-4ff4-8b8b-ced9c26b255a · inbound
How Far Are We From True Auto-Research? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 792e23f6-c3ff-4f18-b586-b2a53000ffce · inbound
AutoResearch AI: Towards AI-Powered Research Automation for Scientific Discovery MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 119
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f42692fa-8b91-4298-acb2-9a463988e008 · inbound
ScientistOne: Towards Human-Level Autonomous Research via Chain-of-Evidence MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 10091f2b-0aad-42d7-bf17-f83624152935 · inbound
AutoScientists: Self-Organizing Agent Teams for Long-Running Scientific Experimentation MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation adf244bb-13de-4a2d-be08-e6597990a58a · inbound
LongDS-Bench: On the Failure of Long-Horizon Agentic Data Analysis MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a251f68d-354f-4671-9add-bf046b693955 · inbound
Can Generalist Agents Automate Data Curation? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 137e84eb-005c-41d3-b6ad-e2b7e31e2fdf · inbound
Towards Persistent Case-Based Memory for Autonomous Data Science: A CBR-Augmented R&D-Agent with a Locally Deployable Small Language Model MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 2e5178ee-23cf-4342-8f27-b7e1bedac843 · inbound
MLEvolve: A Self-Evolving Framework for Automated Machine Learning Algorithm Discovery MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 15c51748-9d40-4cd7-bb6b-27f5de7401c7 · inbound
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0f774748-35f1-4c5d-8cef-ef19ab480933 · inbound
ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 701e14d2-be9f-4733-a559-bf63a93985bb · inbound
Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 84420d36-0c62-49d9-9907-497a96646daf · inbound
Trustworthy Self-Composable Big-Data-as-a-Service: An LLM-Orchestrated Multi-Agent Framework for Automated Data Engineering, AutoML, MLOps Deployment, and Drift-Aware Lifecycle Optimization MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9cd2f965-2ec6-4874-b5dd-4df2b11b08cd · inbound
Agentic AutoResearch forSpace Autonomy: An Auditable, LLM-Driven Research Agent for Aerospace Control Problems MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a9e68e22-d9e6-4cc4-86ef-bfa3f0a25ec7 · inbound
PaperClaw: Harnessing Agents for Autonomous Research and Human-in-the-Loop Refinement MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f90db1d0-05fc-4791-aa6b-ce31a1fafd1c · inbound
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 34e973d5-9efb-483a-94d1-008b7dde1658 · inbound
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c1c4257-7653-40c1-84d5-31e249cf05cf · inbound
Glite ARF: Verifier-Driven Research with Parallel LLM Coding Agents MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ebac2fff-1b08-4520-86b8-79572d0a36ec · inbound
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5df1d11f-51c6-4bf4-85fa-92d80987df8b · inbound
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ee6a2fc-46ff-4d74-b42a-5d17c42349a2 · inbound
FARS: A Fully Automated Research System Deployed at Scale MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 490b25af-c307-4926-b17b-7874354c2ff5 · inbound
FARS: A Fully Automated Research System Deployed at Scale MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6bea006-f3cf-4edf-8e8c-13f3d8637773 · inbound
Compression, structure, and executor capability: a controlled real-cost decomposition of language-model agent skill optimisation MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ec7eccd-f664-4f90-aedf-6adcbf74213f · inbound
ArchEval: Measuring AI Agents as Computer Architects MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4cb8144-1aa6-4c15-8402-fc415491b0be · inbound
When Does Restricting a Coding Agent to execute_code Help? A Regime $\times$ Agent-Design Ablation MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8915d0f7-10cf-4049-800f-45908588bdda · inbound
StarCodex: Dynamic Coding Harness for Starlink Measurement Analysis and Experiment Automation MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75e97664-e545-4253-a3ab-ac63e16441ba · inbound
SQBench: A Benchmark for Evaluating Task Delivery by Language-Model Agents in Production-Oriented Workflows MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 668008f8-d1d1-4c27-9fb9-3c117585f4d5 · inbound
Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e976994-74ce-40cc-be05-cb29731d9f6d · inbound
RecHarness: A Bandit-Routed Agentic Harness for Self-Evolving Recommender Systems MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfabdb40-308d-4ade-9774-d158da25aaf2 · inbound
Agentic Bayesian Optimization through Surrogate-Augmented Autoresearch MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1873b8c-4f30-4464-b184-fb473de7d847 · inbound
Towards a new paradigm of scientific discovery with socialized artificial intelligence MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6a76272-45cc-4970-8ec7-f3dfe9430d26 · inbound
Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 186
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1f909b1-6832-467f-b923-e52a84a786e0 · inbound
Evo-Bench: Can Language Models Improve Agent Harness? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f998b68-8bfa-4d4e-b156-2a4bcfd2b75a · inbound
Evo-Bench: Can Language Models Improve Agent Harness? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98d398fb-3d0b-4ab0-8dfb-c691452f1b71 · inbound
verdi: retrieval is not transfer for continual world model optimization MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a32d19f-9c42-4377-8d46-6418a4e133d2 · inbound
DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments? MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f0fc178-9d5c-4096-9224-1b83e2470f95 · inbound
VALG: An Agentic System for ML Theory Research MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.