Pith. sign in

Paper Citation Record · LEDGER

MLGym: A New Framework and Benchmark for Advancing AI Research Agents

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 36 inbound Pith citation observations for arXiv:2502.14499.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.14499 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:42:43.090307Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

4
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 24af59e5-79f5-48b6-a9b5-fdcb8693094c · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 122

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:57:38.473587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:811e10e2858081fb83a27201849dde60c247ae40ee39b3b2f17f65b862df5431

Observation f76081b5-bdb8-4bf5-9fc4-b728fddc0f48 · inbound

MM-Agent: LLM as Agents for Real-world Mathematical Modeling Problem cites this paper.

MM-Agent: LLM as Agents for Real-world Mathematical Modeling Problem MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:43.090307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:43.090307Z digest=sha256:1597f91ed01e2cc4c613ad99ba5e48838d1c7a3c34ac22d727361a3c5af92563

Observation 52d62860-0f53-413c-8dc2-a3af42c12106 · inbound

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks cites this paper.

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:53.333828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:53.333828Z digest=sha256:cf91aa1d3eda40f5553d05115afd4cf3e9b3c436b49dc1edefb3a776ce2a8c61

Observation 13645ed6-a5b8-4d94-942a-6771b241bbe6 · inbound

TextAtari: 100K Frames Game Playing with Language Agents cites this paper.

TextAtari: 100K Frames Game Playing with Language Agents MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:57.997567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:57.997567Z digest=sha256:b291b2c8885e14203271df591279a72afedbcba055511e971a19f5c624628add

Observation 463cd5b2-c0d1-4b07-9eca-843fb111b21c · inbound

LLM-First Search: Self-Guided Exploration of the Solution Space cites this paper.

LLM-First Search: Self-Guided Exploration of the Solution Space MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:47.225949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:47.225949Z digest=sha256:76c82a06f9b617427b36babb8558ef56a36422bbf3ef98a28386261acd5d9bd5

Observation 79bccf71-c36d-4f0c-ba6c-fac4d942ced4 · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.472486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.472486Z digest=sha256:e28f09f863be8a56044790cc07ba81c82419c69a6fc4582947d874d75a2d0d18

Observation 8a9e3bef-3d00-48e7-bb34-5fb578784dcc · inbound

Kaleidoscopic Teaming in Multi Agent Simulations cites this paper.

Kaleidoscopic Teaming in Multi Agent Simulations MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:49.121636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:35:49.121636Z digest=sha256:21fb52207d69f40ac29a394b60b0e29e2dacce9ec3091329fb5a42e1592bd21e

Observation 0860c8ab-af57-4a9f-bc7e-58eb1e415f44 · inbound

Deep Research Agents: A Systematic Examination And Roadmap cites this paper.

Deep Research Agents: A Systematic Examination And Roadmap MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:56.783174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:56.783174Z digest=sha256:ef38404aa7ffbce890c98e35e2b1a224b557fad4facd51989914d10a5db8ab91

Observation 9824b926-79df-498e-ab20-0cee30bbb560 · inbound

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems cites this paper.

A Comprehensive Survey of Self-Evolving AI Agents: A New Paradigm Bridging Foundation Models and Lifelong Agentic Systems MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:21:42.565484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T23:21:42.029285Z digest=sha256:8e875075e74b5013f2a477b7e2e10ea2ecf644092c9af1b8aa79fab0003ed29b

Observation 9593ae03-a769-400c-bd17-8ebad3f56a97 · inbound

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence cites this paper.

DeepTRACE: Auditing Deep Research AI Systems for Tracking Reliability Across Citations and Evidence MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T12:11:34.231652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:11:34.231652Z digest=sha256:9dc10655fdb9295bb1bddfa69e4bc2e43bb52192decf85285fec0b45db01eb41

Observation e9e7f922-74b4-4365-a969-2385d37e313f · inbound

AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments cites this paper.

AgentCE-Bench: Agent Configurable Evaluation with Scalable Horizons and Controllable Difficulty under Lightweight Environments MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:30:49.503880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:07:46.077831Z digest=sha256:bd8914772ac1ef0422cb9dbc376162a79826bda5e32a0e1d4d654ba1ac326d6d

Observation f16f896e-a368-43f4-8557-60b61edd8e3f · inbound

AI-Driven Research for Databases cites this paper.

AI-Driven Research for Databases MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:55:59.963254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:52:17.486258Z digest=sha256:9e491b3dd5582c7561004b8ba30939809c4049aea4d96fdff9151e88fb902013

Observation 1c404655-6b82-45f4-bfab-9a9a3c85b18f · inbound

TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration cites this paper.

TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:46:35.120334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T12:34:29.808503Z digest=sha256:a1ebd04c1a58cc94064658f53661bf123ac56c0a72a25f31045ad2dcf3d46ead

Observation e22e6dff-ce38-43e4-bfff-2fbdc3bd2b30 · inbound

EO-Gym: A Multimodal, Interactive Environment for Earth Observation Agents cites this paper.

EO-Gym: A Multimodal, Interactive Environment for Earth Observation Agents MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-09T22:13:58.361015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T15:03:09.166127Z digest=sha256:6f1e849e61dbbf35b2dcedd5dcd6d4b9f4153c9ee5e965c9fdb9ce0cd42dba26

Observation 74cfeb6c-6026-46a0-ad7b-a6d6a568ebf1 · inbound

Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight cites this paper.

Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:56.553489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T00:54:25.549158Z digest=sha256:df733e085bfbf8a535f7e5336099d821955b577ddab658f3621d53624d5bf568

Observation 16597c5b-5b3f-4237-8ce9-65cc80a6bd3c · inbound

Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight cites this paper.

Behavior Cue Reasoning: Monitorable Reasoning Improves Efficiency and Safety through Oversight MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:29:52.723765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T08:29:09.122055Z digest=sha256:82a0cf1bce15577145e303a5d4587c8d9e49429bb65fff219eec72652d506581

Observation 3343f866-4dd0-4b75-b324-dea8922abc84 · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:21:25.265962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T01:13:35.990078Z digest=sha256:cc3ec823eff1e3f505e8a8405a78def85a748e857ccb99b254ec2d5b4ca20728

Observation e6a903ea-d8a9-4c5f-a5a4-2ad2b5968289 · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T13:25:46.030035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T23:12:57.154537Z digest=sha256:f8a7a7b64b0a4e58add641e5eaaf0fe88584044e3fbd1fc3e39ccf03d10d8427

Observation 6ef9c2b0-10b2-4462-942c-e288a2fa4801 · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-12T17:14:49.310598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:14:49.310598Z digest=sha256:34166d70a5371216645383295ad9bce61bf7cc7456401893833a855f00255052

Observation bb6bf2b1-c166-4a23-b57c-4b04936e025b · inbound

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility cites this paper.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:03:44.085167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:ce44c1295b5695b90e30e339eaa620fa9e31f89b572ea373f48b66bd2e52eb40

Observation 7f48d165-c0ec-4440-8ddd-e272d23f420d · inbound

AI for Auto-Research: Roadmap & User Guide cites this paper.

AI for Auto-Research: Roadmap & User Guide MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:33:12.639602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T10:30:50.256635Z digest=sha256:57210e70ed11d442a1212239e045cf603b920359d992de8fed8304739270d5ef

Observation 6dc2c65c-838e-40c2-8d4f-b500548d840c · inbound

AI for Auto-Research: Roadmap & User Guide cites this paper.

AI for Auto-Research: Roadmap & User Guide MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 135

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:40.325453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:40.325453Z digest=sha256:c2244dac4555a681a870a8237c19b7458dc207b3c3b60b1004881aabe56aadd9

Observation e24dbd26-fbef-49fe-82f7-367b0c4149cb · inbound

SAGE: A Quantitative Evaluation of Socialized Evolution in Agent Ecosystems cites this paper.

SAGE: A Quantitative Evaluation of Socialized Evolution in Agent Ecosystems MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:26:29.555018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:02:25.870386Z digest=sha256:f27d63b8493d00719fc8295dbac8b75e73d4fa9eeb5c2377a28da79d9553a0be

Observation 9cf0516c-52f2-4d55-afa1-a277f036a1c3 · inbound

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research cites this paper.

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:33:15.813204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-29T08:24:42.412763Z digest=sha256:3ef8b9a5cee18468f5d2b654098d20179d12e37b1214c529c18148203aceb46f

Observation 2ddd90b3-696c-4d5e-941f-f6a3714ce425 · inbound

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research cites this paper.

ResearchClawBench: A Benchmark for End-to-End Autonomous Scientific Research MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:39:16.507175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-04T00:30:00.449270Z digest=sha256:e1ed5979ea88dbb886f60c59d848416a52e65965e36bff0f2bbcf8383e0e25e7

Observation be25b1b2-71ba-4e48-bac1-dffb1720bbf0 · inbound

InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees cites this paper.

InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:07:37.266203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T14:06:59.471772Z digest=sha256:7f6ee4d976611d117cd35e4a5a960177407161b906f2b9e54c17bdd2aa6d0639

Observation 9362f5de-c93a-4a7d-ae64-b221944a01d1 · inbound

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales cites this paper.

Benchmarking AI Agents for Addressing Scientific Challenges Across Scales MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:28:04.333266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T09:34:09.347912Z digest=sha256:ba9306b01a23b0f4ffb72d6949514a0f9929e09289b5c172a4b3f9eac5119fc1

Observation fe5d0459-1a22-417e-a3b0-1d2c71d1afa2 · inbound

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data? cites this paper.

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data? MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:58:22.222379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T07:17:43.188693Z digest=sha256:71d9fa7ed38ce75f74892f701bf015483d424b29d6c465513a12a09a73c5b888

Observation 21be2fa4-d534-4af3-9fd3-422ffd310c30 · inbound

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data? cites this paper.

TerraBench: Can Agents Reason Over Heterogeneous Earth-System Data? MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:37:25.449237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-02T22:29:58.492918Z digest=sha256:f8c454e64709e912ea017886020d1254a0dff5e769d4f1f62df93c40fd9b8fd7

Observation 0cd66105-af08-4e2f-87c2-a77ade7a068b · inbound

Environment-Grounded Automated Prompt Optimization for LLM Game Agents cites this paper.

Environment-Grounded Automated Prompt Optimization for LLM Game Agents MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T00:40:18.662768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-27T00:30:42.133192Z digest=sha256:22f9341b4c41f614c69b9e3f62c58d2f79de5d259b1800eced1b9ba36da63799

Observation 954b2b2f-e649-4138-a210-40e2995b1809 · inbound

Discovering Crystal Structure Prediction Algorithms with an AI Co-Scientist cites this paper.

Discovering Crystal Structure Prediction Algorithms with an AI Co-Scientist MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:19:46.994246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T09:03:00.675456Z digest=sha256:c33525b68a57dc4633a5d6089874447cdf72790e03cb7edd021a05f32ae3dc24

Observation 82ea1e40-c1c7-43a5-aaf0-47e24585bd20 · inbound

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? cites this paper.

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:57.820516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-26T00:13:14.940915Z digest=sha256:5f907d5055fd6ad21ae9b95ca41d6b6a368700265551764f0765ca2593e9e4de

Observation 6fab8d45-91dc-469c-8eeb-8e5e0e8c7039 · inbound

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? cites this paper.

NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-12T12:32:39.706055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T12:32:39.706055Z digest=sha256:512564ef22ebd7d7f81cbe514f82a2b289b83ce348b59c4b00a2a3a0331122cf

Observation db995256-b2bf-4c09-bff4-0a297d14f7e0 · inbound

Autoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation Data cites this paper.

Autoresearch with Coding Agents: Generalizers and Metric-Maximizers on Quran Recitation Data MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T16:16:31.350666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:16:31.350666Z digest=sha256:752ce37e67e200cddf9748fb686f0ed885c92288393b6b87287a91567bc1ff10

Observation 90d01cd4-8504-4389-8829-49cd3e2e4d33 · inbound

One Run Is Not an Idea: The Implementation Lottery in Automated Research cites this paper.

One Run Is Not an Idea: The Implementation Lottery in Automated Research MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T12:50:04.327582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T12:50:04.327582Z digest=sha256:681fce804a6bf85341d29b9b957aee0740aa8d99e7bffff4a02c1c03d6e98523

Observation 49b22e7f-9d05-42f0-907f-9ffcb04fa167 · inbound

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details cites this paper.

Is Inter-Seed Cross-Play Enough? Evaluating the Robustness of Zero-Shot Coordination Algorithms to Implementation Details MLGym: A New Framework and Benchmark for Advancing AI Research Agents

Reference 169

Resolution
unresolved
no resolver link, observed 2026-08-05T15:25:40.067169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:25:40.067169Z digest=sha256:7a46f4879019cdc5668b55eba9a99c32722987bc89630a039d6d8f3a138cae00