Pith. sign in

Paper Citation Record · LEDGER

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization

As of 20 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 3 inbound Pith citation observations for arXiv:2509.09321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09321 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:21:20.656965Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:19:04.202588Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:08:43.128824Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 47be8e18-2e3d-4eff-a005-713060cace01 · outbound

This paper cites DeepSeek-V3 Technical Report.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization DeepSeek-V3 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.617362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.617362Z digest=sha256:aecda8239ad13ae2857f1cd47c70cb38f121854ba5851de104ec9a17443b5f96

Observation 0c1fd238-f871-4ab2-8b44-2a48fd10c74a · outbound

This paper cites Based on thebest solution.pyand the special instructionfield, we apply an LLM-as-judge method.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization Based on thebest solution.pyand the special instructionfield, we apply an LLM-as-judge method

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.652518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.652518Z digest=sha256:5b60c46952ebfc1789bdef87c575fc08e5008feaf4164c8bac3baaace3d49c87

Observation e1b0b085-c7ee-4717-aafa-2f5efa5e9772 · outbound

This paper cites Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.625233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.625233Z digest=sha256:89d51bfdfa4727502a4c2e30647aabe183019329a168b6af628d9d8391c23ef9

Observation 227d29d8-f707-437a-8499-71de11a6b9e8 · outbound

This paper cites MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.629076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.629076Z digest=sha256:45bf7f708fc9ec5446f5de2b9c62fc4cc415a4d6c43c72c71e5a06d4ecf81eff

Observation 3108a0c6-682c-41e6-81c6-1097e8878488 · outbound

This paper cites AIDE: AI-Driven Exploration in the Space of Code.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization AIDE: AI-Driven Exploration in the Space of Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.632584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.632584Z digest=sha256:60a39356033cd6a976ab0acc63e32970b69be8bcd8aadcc2c28285e7f9b19352

Observation 4cbe0489-1d73-4c7c-95f9-bff113c376f8 · outbound

This paper cites https://openai.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization https://openai

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.636967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.636967Z digest=sha256:d08f875a63418ff73425c3d0422cc04e71e885e7b08fb6f5fd6477f7ef42ec37

Observation c8745975-a680-4800-ab74-c32d932b8cb5 · outbound

This paper cites arXiv:2506.10974.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization arXiv:2506.10974

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.640845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.640845Z digest=sha256:7f15186908cae82fe8cca8f70a4ff6c8b4a53c646dc375b2de439cab66f5c62b

Observation 0e597af1-c846-4930-8e80-0df76d12aebf · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.644500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.644500Z digest=sha256:6034c7ae4ef33f88dbe23a1adf263f9bf5281713fa74e95173e15e6eceae552d

Observation d546f07d-c57c-410d-8f82-3cd9487f3cc8 · outbound

This paper cites A dash (“-”) indicates that nosubmission.csvwas generated or the generated file failed to meet the required format for evaluation.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization A dash (“-”) indicates that nosubmission.csvwas generated or the generated file failed to meet the required format for evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.656965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.656965Z digest=sha256:3ba0a47c43f4d9caabe1e4bc179baabfacbe2e59c283a9c90176ab3d170b8a35

Observation a64d702a-bb40-4664-8400-a711ddc705c5 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization ReAct: Synergizing Reasoning and Acting in Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.648156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.648156Z digest=sha256:ccf7bf777d166ad53679ed44137ec420423caed5a55389bafa3312a419b9a38d

Observation 59f460dc-f957-45a0-9e56-290c1927528e · outbound

This paper cites Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.621399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.621399Z digest=sha256:705c0494d23d784658dfa9124a5e0f624d3e10a2b99d2e32b6efbdc17ac6b7af

Observation 8efd4654-b1fb-4d81-bad9-179da1230ff3 · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.611481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.611481Z digest=sha256:a2687cec5b557ac68f93292a23c3fd1b00ec79a9011fba935e4f7700524fba99

Pith citing papers

Observation 83bb601f-f2dc-4200-ba0d-534fbcd87854 · inbound

Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery cites this paper.

Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T02:57:46.718690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:57:46.718690Z digest=sha256:55e3734e59982ba93cefc499c753c704dcecd8e44f7807b7e2fbc9e8887ca8ab

Observation ac74188f-114d-418d-97c0-6dd30932082b · inbound

From Question Answering to Task Completion: A Survey on Agent System and Harness Design cites this paper.

From Question Answering to Task Completion: A Survey on Agent System and Harness Design Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:08:43.130648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T04:40:30.985824Z digest=sha256:2b878b92216cfe5e76e953c729c59925280560fd6a74295ecfa94054e4ee227f

Observation c77656e4-d561-4551-af1c-1a4ce5632115 · inbound

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision cites this paper.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:04.202588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:04.202588Z digest=sha256:97b6481f9efc7ecf43254c0de6241f5caf1ba0d7c27c042b09da392665d7429f