Pith. sign in

Paper Citation Record · LEDGER

ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2311.09835.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.09835 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:30:31.039397Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T14:05:46.736275Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ebf24730-2b79-4bb7-ac86-c752e9e58ab9 · inbound

MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework cites this paper.

MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:43:18.933691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T03:43:18.632292Z digest=sha256:3bec1667b71724488f8e74e93732b9511275c49eab0a7048fcecab6b546635b3

Observation 230b2e79-42ba-4507-bf2f-bb3b8e302bd2 · inbound

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering cites this paper.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:13:21.535616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:6afbd633e5b9c5d2afbaf7a36940ed43c6ce633c36a5e43d5294d1b024bea5a8

Observation 0a3c29a7-66ad-42f9-8b46-abb6bc0485fa · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 173

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:52:16.430624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:0e521a7616bf3311e6464219d3830d16eada56a6231784a4ea7758113954eca5

Observation c6f21ed6-1294-43ac-9006-709465b70578 · inbound

AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research cites this paper.

AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:30:31.039397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:30:31.039397Z digest=sha256:86ebfca653f485cd6e1032be686fbb6e7aa635da3228907f5ffbd0060ca40dc8

Observation 744c909d-1d45-4ef6-aa4d-29e09ab1a074 · inbound

Reinforcement Learning for Machine Learning Engineering Agents cites this paper.

Reinforcement Learning for Machine Learning Engineering Agents ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T12:24:02.277658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:24:02.277658Z digest=sha256:0c861e07600575c0e850854d1881bff8d7096aaf52ce46f6dfbc06846e561b02

Observation 174c2508-feb8-4a56-86a8-86a342144568 · inbound

Compiling Large Multi-Modal Requirement Documents into Runnable Software Systems: From an Agentic Test-Driven Perspective cites this paper.

Compiling Large Multi-Modal Requirement Documents into Runnable Software Systems: From an Agentic Test-Driven Perspective ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T23:31:16.475615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:31:16.475615Z digest=sha256:204a61a985a3a7610e16cad502352d04fc5dc2c4934278b45b76c342509a7466

Observation 18816a07-66c5-4ba8-a943-1a7a84a17885 · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:25.084754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:13:35.990078Z digest=sha256:c14f37bdfb1c6e372ed07d4be519e4247b28da95847dcc78e5dc6c33af166f5f

Observation cecd6380-a041-4259-aa06-cedd57d3e95b · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:25:46.005528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T23:12:57.154537Z digest=sha256:03312f65066bc71782df3b5835181e97b26761a2d3d78398b21a55f27239767f

Observation ef75ebca-0ea7-44ac-8c98-c58ed522f8cb · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 94

Resolution
unresolved
no resolver link, observed 2026-07-12T17:14:49.310598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:14:49.310598Z digest=sha256:af125c45cfe84617bd16b1ba2da93b30e9a843ab33156f969b079581724d2dd1

Observation 637a3010-188f-4e81-8dbc-ff3578f8ba1d · inbound

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows cites this paper.

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:22.501265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T05:56:36.312877Z digest=sha256:7006528107a5ebd2c4115cad389da216aea3b96ad842107c55acea6dcbeb02f4

Observation 26cc257b-1017-4935-887b-1bbe3b2806a1 · inbound

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows cites this paper.

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:46.737909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T22:18:45.189576Z digest=sha256:56f3e3b8769f908bf597ff67337444e89aea378f1735ddaf94f537e59cfefc5d

Observation 3e1dffac-109e-4453-9e1c-48b967604fc4 · inbound

Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering cites this paper.

Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-31T01:39:48.754483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:39:48.754483Z digest=sha256:f9ba6f04d3bb810695207d3c9903d580f4a158d872323887df778e074fbf0eef

Observation b8946c5c-24df-47d6-8697-164497746df3 · inbound

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers cites this paper.

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:14.160911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:14.160911Z digest=sha256:e3abe04f4c3ad4fe67cd0f95522ecb931f6a9d485410c4fd715ddad3f6bb6984