Pith. sign in

Paper Citation Record · LEDGER

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI

As of 22 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2501.15627.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.15627 v2

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T14:11:16.960772Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:18:43.057747Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T12:24:02.722405Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e71574fb-b897-40ff-8d9f-a3458143e4f2 · outbound

This paper cites an unresolved cited work.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:11:17.272660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.871707Z digest=sha256:c6700a1f0cb0d931ae0cd79e8d858121d2d2bddbd6089c0cde76eb2ebad9f3e0

Observation e2d90d53-12a0-45b0-b1c6-1a4a5ee49ab4 · outbound

This paper cites B., Mann, B., Ryder, N., et al.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI B., Mann, B., Ryder, N., et al

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:11:17.262917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.875811Z digest=sha256:c5dbb181814fd1d4ef2f75767d7e2aea9df4e5d9917b9c60d4b3610b680b4b00

Observation d99a925f-c492-4f43-8b98-82f9547d4415 · outbound

This paper cites an unresolved cited work.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:11:17.251934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.879416Z digest=sha256:69b575fdc59dd482778a17d67a30527f0e39fc42512a05690c83f542d31ac64a

Observation 71d960eb-2ca7-49a1-874d-c956f056d2a5 · outbound

This paper cites an unresolved cited work.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:11:17.238508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.882694Z digest=sha256:9048a4d000c67ce38edcac03b7cbd8165bf546a58c6299f4ec81408408bcafae

Observation 5072a14a-2ad9-4b90-b39d-977e5b9573a6 · outbound

This paper cites an unresolved cited work.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:11:17.226327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.886012Z digest=sha256:f76b76d9a314d5fa82c799ffd85271103aa5c5fd5578c9b019390fd14567f9f2

Observation 39c04393-42f3-4f9b-9f54-84de34b4a8c6 · outbound

This paper cites an unresolved cited work.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:11:17.215528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.889355Z digest=sha256:d26a88b78d5fff7d74a4230d2fc86f636416bc081f297f5876194e25bbb46867

Observation 2823e9a8-2701-4776-8b1b-d0b5fffb7f50 · outbound

This paper cites an unresolved cited work.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:11:17.205479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.892807Z digest=sha256:506266afe55fcebc1fd167bd93a9a875e786067e6b021c2c4631ef0d95b7a595

Observation 2fc2b982-75bf-42bd-90ae-647afd7d4b2d · outbound

This paper cites ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI ClinicalBERT: Modeling Clinical Notes and Predicting Hospital Readmission

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T14:11:16.895743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:11:16.895743Z digest=sha256:b07ed87398ec54139eea223fcd40678c8121bc25e134e885666bcddf1fee0593

Observation 1c9655e3-f3d6-4d10-9b1d-dd4b0b1244e7 · outbound

This paper cites an unresolved cited work.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:11:17.195411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.899097Z digest=sha256:98e7bfb0f7b596be3c0c65b09ea3841c80eeb84ea4d17ca7aa56c946911d9b31

Observation a1f0bde9-7a6e-460f-a158-51a2fe9c4775 · outbound

This paper cites I., & Mitchell, T.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI I., & Mitchell, T

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:11:17.184951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.902442Z digest=sha256:37f75eaf6b9e26cf30ffbcc24c2dc14d2cce1407e02631439652d9addd2ee67a

Observation 97099659-ae3e-483b-8fc3-141d589a4dda · outbound

This paper cites an unresolved cited work.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T14:11:16.905833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:11:16.905833Z digest=sha256:e7a83c529477a329e36dde7225ba9aa54bd805f51430c13e2ec1fa123ab8293b

Observation 19b0b44a-e66a-4cf5-ab76-617a81c09dcd · outbound

This paper cites an unresolved cited work.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:11:17.167158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.909005Z digest=sha256:e91ff6f02bb5289ed7e1bfbaedcb0080f0f1c9e43f89270cea8f010547da1e0c

Observation dd4ac28e-f27f-44cd-a2de-df6488633e1b · outbound

This paper cites an unresolved cited work.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:11:17.156506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.912473Z digest=sha256:419f57ea458c8fea44b80c07d27b5cce3dee8cb23b47919f614a4de41bf3f699

Observation 35456de7-2428-4833-8abe-1b84011777e8 · outbound

This paper cites Direct speech-to-speech translation with a sequence-to-sequence model.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Direct speech-to-speech translation with a sequence-to-sequence model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T14:11:16.915616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:11:16.915616Z digest=sha256:29f82aed41f69f92f060be454a3f1911be6bede047e71064b9644b6dbc28e600

Observation de65e791-2b36-486d-ba3d-ba130315b2eb · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:11:16.919552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:11:16.919552Z digest=sha256:47c03c559c7fd9d0627b707264eab45a1cfbfa601d439f588c2032bfa9002826

Observation 4879da6c-4d51-425d-bcca-ed2ea660295d · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T14:11:16.923011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:11:16.923011Z digest=sha256:2f5b1870d2a1002857200c5d2509422d65616c0e4d1fa83b5474651438b2a655

Observation f0177ed7-f761-4ff3-b7e7-1eddb433ec6c · outbound

This paper cites an unresolved cited work.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:11:17.146946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.927035Z digest=sha256:af16b903d70f00d09a8674d3a77ea47f6a98e07b747c805b33fd0f1fd303c45f

Observation 69eddccf-a3ef-4517-a999-dfc9532d3cac · outbound

This paper cites an unresolved cited work.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:11:17.136544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.930618Z digest=sha256:6b006c164a5724e814c4634fba54e630a628eee7a44162ce4dc7da79e5ac27db

Observation 02311cf2-784f-431e-a237-615d4c8fb339 · outbound

This paper cites an unresolved cited work.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:11:17.124286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.934348Z digest=sha256:d4a399a31f921c2bd2e6567269b1416ce0f766b6bfc8e3ade1f6d408899e4388

Observation 36318977-ea7c-4ee8-a667-b60999375aad · outbound

This paper cites an unresolved cited work.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:11:17.112994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.938801Z digest=sha256:9da2c50edc66c5dd9c00fa4552b5df13216af4ae9c20a532b1b4286a6fc45b04

Observation 5686f064-125e-41de-88c2-a0ffc791202b · outbound

This paper cites an unresolved cited work.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:11:17.100887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.942887Z digest=sha256:b6df517b0cc6ec5c340da8dd453d1ac5735d0b2644240fa1954f63cc767f30a2

Observation 6924ce18-cc38-4f58-bb96-05578bb0eda7 · outbound

This paper cites an unresolved cited work.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-10T14:11:17.090510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.946556Z digest=sha256:46c01f37473fd41d354388e8141cbc6999730fb344fb888499dbfb0fc2c04339

Observation 540962e0-6f97-456a-b433-f6464c81b3bc · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T14:11:16.950061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:11:16.950061Z digest=sha256:2637bc0fb7ddd9853ddeb605cbf6f31a0b5609fdb2a5ea84c27e2d2e06db301f

Observation b84227a2-2fc1-4ad0-859c-3f2bff1c6329 · outbound

This paper cites Learning to Reason with LLMs.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Learning to Reason with LLMs

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T14:11:17.080099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.953986Z digest=sha256:205d5321cbc998dbcc147f5a4de4b8c917ff37b03ca786b9d8a5d4765353f903

Observation 75f288d7-23ac-4a58-80da-1a4ca3ef2fae · outbound

This paper cites From Human Days to Machine Seconds: Automatically Answering and Generating Machine Learning Final Exams.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI From Human Days to Machine Seconds: Automatically Answering and Generating Machine Learning Final Exams

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T14:11:16.999605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-10T14:11:16.957322Z digest=sha256:1a917713377e7d0f7726106deb06f76d3411ce4bef42840babbb8c151a916c2c

Observation eb125f94-3cf4-43b8-82cf-be3025d6a083 · outbound

This paper cites Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI.

HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI Understanding and Benchmarking Artificial Intelligence: OpenAI's o3 Is Not AGI

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T14:11:16.960772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:11:16.960772Z digest=sha256:ac3cb34743a95fd5f2f6d7ddc182abbcd1d94769a41a54948a92de652c7c303f

Pith citing papers

Observation ae28873d-c04b-4a9a-975d-5ca945b11b12 · inbound

Make Still Further Progress: Chain of Thoughts for Tabular Data Leaderboard cites this paper.

Make Still Further Progress: Chain of Thoughts for Tabular Data Leaderboard HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:43.057747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:18:43.057747Z digest=sha256:c9f56d133bf128ce1193e063ea53ce1730455d75977b114bf5d0e4c84c07ed74

Observation 4ad69974-7d04-4d41-9bf4-127080bd50a7 · inbound

Reinforcement Learning for Machine Learning Engineering Agents cites this paper.

Reinforcement Learning for Machine Learning Engineering Agents HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:24:02.725488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T12:24:02.292197Z digest=sha256:5be88338d81e2a359c0ae21b38257beb8c58f473a503f198cbdcd8e620676c98