Pith. sign in

Paper Citation Record · LEDGER

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization

As of 9 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 3 inbound Pith citation observations for arXiv:2509.09321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.09321 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T19:21:20.656965Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T07:19:04.202588Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:08:43.128824Z

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 47be8e18-2e3d-4eff-a005-713060cace01 · outbound

This paper cites DeepSeek-V3 Technical Report.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization DeepSeek-V3 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.617362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.617362Z digest=sha256:c429505431f50d374198f02553704238bafd5e552f897555e1df3872caf64b05

Observation 0c1fd238-f871-4ab2-8b44-2a48fd10c74a · outbound

This paper cites Based on thebest solution.pyand the special instructionfield, we apply an LLM-as-judge method.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization Based on thebest solution.pyand the special instructionfield, we apply an LLM-as-judge method

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.652518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.652518Z digest=sha256:366ded0388e08d532c11164828a0aa36bfd4094b807060e8e60439494b757ead

Observation e1b0b085-c7ee-4717-aafa-2f5efa5e9772 · outbound

This paper cites Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.625233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.625233Z digest=sha256:1913d4fb7768ac3f6fdc490455ae62be1cfc866ba519e11c9dae5b193749bca6

Observation 227d29d8-f707-437a-8499-71de11a6b9e8 · outbound

This paper cites MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.629076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.629076Z digest=sha256:d2e078efefb6f385f94681cf7131c20959ac4bbd1213b343e9711d1477d494b6

Observation 3108a0c6-682c-41e6-81c6-1097e8878488 · outbound

This paper cites AIDE: AI-Driven Exploration in the Space of Code.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization AIDE: AI-Driven Exploration in the Space of Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.632584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.632584Z digest=sha256:ba9e87456209b67704856544ab5a19b76f66edfd3df07e07bbd8cb51e70851cb

Observation 4cbe0489-1d73-4c7c-95f9-bff113c376f8 · outbound

This paper cites https://openai.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization https://openai

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.636967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.636967Z digest=sha256:c50b827dcaec7955e29b859aa67a4e5e88b4b1d27bd86655ff30856e20110f16

Observation c8745975-a680-4800-ab74-c32d932b8cb5 · outbound

This paper cites arXiv:2506.10974.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization arXiv:2506.10974

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.640845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.640845Z digest=sha256:91ddad06e13f621e8e3474cc7671c481eec8a2bebfd12f218ce67ca25451aab5

Observation 0e597af1-c846-4930-8e80-0df76d12aebf · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.644500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.644500Z digest=sha256:2d8a7d29b7989ad603103a3c53f5976144c905361f1d8986771d6f4452db255b

Observation d546f07d-c57c-410d-8f82-3cd9487f3cc8 · outbound

This paper cites A dash (“-”) indicates that nosubmission.csvwas generated or the generated file failed to meet the required format for evaluation.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization A dash (“-”) indicates that nosubmission.csvwas generated or the generated file failed to meet the required format for evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.656965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.656965Z digest=sha256:bb5a0e6a733695494478636d5ab9017bec2992678fd57724a2ca3d38089f34c7

Observation a64d702a-bb40-4664-8400-a711ddc705c5 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization ReAct: Synergizing Reasoning and Acting in Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.648156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.648156Z digest=sha256:3b0f19deb89adf56193b315627a42ac1822d741f69d8c9a4b87945dc318ff9e7

Observation 59f460dc-f957-45a0-9e56-290c1927528e · outbound

This paper cites Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization Empowering Many, Biasing a Few: Generalist Credit Scoring through Large Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.621399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.621399Z digest=sha256:d5eac5c10c9c87d027a3a624b40a5facb989a822716721a11beab7d7e8c1faf8

Observation 8efd4654-b1fb-4d81-bad9-179da1230ff3 · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T19:21:20.611481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T19:21:20.611481Z digest=sha256:c12cc5089cebb4c1f58197e6317fd0a3da78586b528e354ece5b18092f9eae3a

Pith citing papers

Observation 83bb601f-f2dc-4200-ba0d-534fbcd87854 · inbound

Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery cites this paper.

Can Large Language Models Derive New Knowledge? A Dynamic Benchmark for Biological Knowledge Discovery Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T02:57:46.718690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:57:46.718690Z digest=sha256:6a6bd653c2d2b52ccb40adc778e3258aa8c3ead7b6e2d366156306b7e5324047

Observation ac74188f-114d-418d-97c0-6dd30932082b · inbound

From Question Answering to Task Completion: A Survey on Agent System and Harness Design cites this paper.

From Question Answering to Task Completion: A Survey on Agent System and Harness Design Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:08:43.130648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T04:40:30.985824Z digest=sha256:d591ed15fd938135922ab25176a09ee4b648c354fd9a199d9bbdde8f32945b61

Observation c77656e4-d561-4551-af1c-1a4ce5632115 · inbound

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision cites this paper.

ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision Towards Adaptive ML Benchmarks: Web-Agent-Driven Construction, Domain Expansion, and Metric Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T07:19:04.202588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:19:04.202588Z digest=sha256:1389e5238c0c1628b2f94cc4c0b24873adf0b66c4144a9c30f6e27a1aed40aae