Pith. sign in

Paper Citation Record · LEDGER

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities

As of 17 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 1 inbound Pith citation observation for arXiv:2412.06745.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.06745 v2

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:24:17.901468Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-22T21:09:30.462709Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T21:12:08.583301Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9971c8b-1223-4725-a55e-fad83c9db937 · outbound

This paper cites llava-7b 3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities llava-7b 3

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.264742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.733490Z digest=sha256:ad5d6db551cc8226493f4298e1ea3c7c963bc063e2c8dc7d8a5e91ef45f1b6da

Observation 24ae2197-60bb-4d1a-96e5-c76da6fe78af · outbound

This paper cites A comprehensive evaluation of these alternatives could offer new insight for aggregating model performance.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities A comprehensive evaluation of these alternatives could offer new insight for aggregating model performance

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.111732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.894669Z digest=sha256:f8696b5d652c5a16798e068aa4e53bd256d5fe5d33a548758115faa07e1bb70a

Observation fdcd4fad-ad4c-4133-9df6-9b3e23eefa6b · outbound

This paper cites an unresolved cited work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:24:18.100731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.898103Z digest=sha256:ed8a5ee378f429ef5bf563104310bc3d736927e7c5cde9645a21df8669252b8f

Observation 03900b93-7987-4443-9004-1852de35bb7c · outbound

This paper cites an unresolved cited work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:24:18.090126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.901468Z digest=sha256:a52d49dac2cc15b5192b36190ceda99994ccbd813d4daf2a7ecf84279e9ecfc2

Observation a73e2cf4-f347-4d47-9b62-aabc3282662f · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.693587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.693587Z digest=sha256:93c3a13f66c59904b007252b7c36a9921737178e947a62b7ffae8878066d88a4

Observation 61dd9087-020f-4142-9b45-bfd8f4805819 · outbound

This paper cites tinyBenchmarks: evaluating LLMs with fewer examples.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities tinyBenchmarks: evaluating LLMs with fewer examples

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.697517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.697517Z digest=sha256:178fd4ec2c949e06bec3d1036ce1104a0a5d8c9aa73583334e8e598338b8c7d2

Observation e589acc7-12b6-4795-a779-6ffc28164486 · outbound

This paper cites Efficient Lifelong Model Evaluation in an Era of Rapid Progress.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Efficient Lifelong Model Evaluation in an Era of Rapid Progress

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T19:24:18.008433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.701768Z digest=sha256:f85861a8395661be857ac2ec9fb1cfeffaec7e478cb9897ae9da008de4aaa0ab

Observation c1183266-9857-418a-a685-8f5b28542b99 · outbound

This paper cites When is it Better to Compare than to Score?.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities When is it Better to Compare than to Score?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.709449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.709449Z digest=sha256:a4f815820abeb78733ca0f7b86578f93c9664743da6e3b7adbb5cfca9e56491d

Observation 0a0c7732-bf17-4a9b-a8d5-76da4ebe76bf · outbound

This paper cites How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities How Many Unicorns Are in This Image? A Safety Evaluation Benchmark for Vision LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.717055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.717055Z digest=sha256:a885202cffeb8cbbe318a9193715c0e659c8713c0d7379645ec78d9ad4d1b8a2

Observation 1e01ed4f-49e3-4e02-820c-b462d300196d · outbound

This paper cites Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Benchmark Self-Evolving: A Multi-Agent Framework for Dynamic LLM Evaluation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.721018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.721018Z digest=sha256:ac09436e1b45be111d5d37ebb510ae4533be54d422731d1c3438b5c3a19a1962

Observation ee0eebcf-5173-4998-9590-78a7c320bf4e · outbound

This paper cites Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Top Leaderboard Ranking = Top Coding Proficiency, Always? EvoEval: Evolving Coding Benchmarks via LLM

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.725599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.725599Z digest=sha256:9db6cdfef67e98a969b0b3bf041fd8c44eae6651f21bc9954bd75e8bc144798c

Observation f212106c-a053-4858-99fc-9fc7dff01d6e · outbound

This paper cites Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Skill-Mix: a Flexible and Expandable Family of Evaluations for AI models

Reference 14

Resolution
malformed identifier
no resolver link, observed 2026-08-11T19:24:17.729268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.729268Z digest=sha256:e59a31bb4c12d97dfa82b44e653c133d8424116c868db181bf7233423d915c0b

Observation 6051a52e-b1f5-4209-ba17-827c9e4d50ec · outbound

This paper cites internlm-xcomposer 3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities internlm-xcomposer 3

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.255171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.737268Z digest=sha256:57640acfd9beb4c565495abc635aba8e9ead489089cb16bae17b28217783f619

Observation a0ea9fba-5106-4252-beda-39ee94da3414 · outbound

This paper cites idefics_80b_instruct3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities idefics_80b_instruct3

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.245573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.740911Z digest=sha256:7bb243a17b2f4a5dd1b0d4352a9388cfcccc988a64436526f09b160fb538b95f

Observation cd287553-a19d-42ad-bac6-52eef6bc42d5 · outbound

This paper cites idefics-80b3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities idefics-80b3

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.235013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.744497Z digest=sha256:b607bdec3525785aec7e78cc3c8943ffa42591f09f742415ec4ff495f8f579a3

Observation dde5673d-dbd5-4dcc-b881-a5277231c45c · outbound

This paper cites claude3_sonnet3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities claude3_sonnet3

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.223053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.748186Z digest=sha256:a23dd13802f88cb8f2bba938651708357556430286413d9e1e3db430f83e5fa3

Observation 3a770b11-7318-4ed6-a2a8-462d5ce5f172 · outbound

This paper cites an unresolved cited work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:24:18.212567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.752064Z digest=sha256:3cad78ca3f224e0c240dad7775a1cb5cfb713741bfaecca9ec5c3a2f4dccca0d

Observation 51a05230-5484-440f-94a9-8464647a9396 · outbound

This paper cites an unresolved cited work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:24:18.190924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.760166Z digest=sha256:f1f81d945d9aa3eeea9a6f236864cb52bf08241ea4a89c7f548839d3f3d3896a

Observation 48f12b08-3461-42bd-8a28-ac6ab2969a12 · outbound

This paper cites Who is the Byronic hero?.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Who is the Byronic hero?

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.180056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.763565Z digest=sha256:b3b3da9978c2e2d6cfccead2e0515d206695df5a6da90aec98d468f29ab4bb6d

Observation d3ce9320-7062-4fa4-a463-8a6d5331f4dc · outbound

This paper cites palmyra-vision-3 3.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities palmyra-vision-3 3

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.169222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.766901Z digest=sha256:c51c3d031666de02bdfc07934c2a1937002c6c15431fa3ea72b7afffc1d0c4cd

Observation 6e757455-8564-43db-b8f6-733c7f5dec94 · outbound

This paper cites _ bought a….

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities _ bought a…

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.158438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.770148Z digest=sha256:2b2be96cbd532d16102e4b075e87d3a0c61a138ae8d4c70aa8468909e72d6680

Observation 2209fb4d-30d5-4d2c-86d2-de1240b84379 · outbound

This paper cites an unresolved cited work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T19:24:18.201532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.773522Z digest=sha256:1eb639b255a4ed17e226c54a95cfdc9ff91c88d79874d12f4270882e5c4ac0de

Observation 39b74299-4a1e-4453-b2f9-95d8a1c956b2 · outbound

This paper cites decreasing the heat energy of a gas? During an isothermal expansion, a confined ideal gas does 150 J of work.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities decreasing the heat energy of a gas? During an isothermal expansion, a confined ideal gas does 150 J of work

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.147195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.776707Z digest=sha256:afbb7f00862546606e2c997175d0202dbde62985c7d55c8b0d147fa1d3a9e203

Observation aa7140e3-5b19-462e-b605-662464a3b8c0 · outbound

This paper cites high-quality.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities high-quality

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.135343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.883115Z digest=sha256:8c1411b623780dfb95d0fc11b17851cbd16d75194457ce833a1e28a9148f3c29

Observation 7d46f22b-a73f-4e5a-891b-c4695d6d0f8f · outbound

This paper cites These pools can be greatly expanded and diversified by expanding to incorporatingall existingLLM and LMM benchmarks.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities These pools can be greatly expanded and diversified by expanding to incorporatingall existingLLM and LMM benchmarks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:24:18.123424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.891223Z digest=sha256:7699491ac4988834f7e18318f58ba7ec3ddafb412c86dda98830632f0e978595

Observation 7b64b2cd-5d17-4a7c-9983-a194ef74fa47 · outbound

This paper cites Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Functional Benchmarks for Robust Evaluation of Reasoning Performance, and the Reasoning Gap

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.713429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.713429Z digest=sha256:e4f7215056fa00ef1181a68ba101bad99f4e9d32e2aecb78d0758963483f1266

Observation 50460282-d1c5-4bbb-818c-f5cb97887a06 · outbound

This paper cites DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities DSPy: Compiling Declarative Language Model Calls into Self-Improving Pipelines

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.689075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.689075Z digest=sha256:735cc6e54f47617ffff76e2ef1192d49487514bdc3c1d37af7033c92841e67d8

Observation ab076295-4cb9-4034-a2f6-215f70230b11 · outbound

This paper cites Memorization vs. Generalization: Quantifying Data Leakage in NLP Performance Evaluation.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Memorization vs. Generalization: Quantifying Data Leakage in NLP Performance Evaluation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.676182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.676182Z digest=sha256:a688503656077db72cb1e7f00b6473eaaa3bed981b9c00d199e252d4962fd7af

Observation 4ae36895-c721-44f3-bfd2-e0190f64c863 · outbound

This paper cites Data Contamination Report from the 2024 CONDA Shared Task.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Data Contamination Report from the 2024 CONDA Shared Task

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.705665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.705665Z digest=sha256:8ccc50accae0752ecc2ec4af92473bb0d1072af74514e12820cb8e90a875cc7b

Observation a964547f-d5ea-47ed-bef3-e0f263c30130 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T19:24:17.680849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:24:17.680849Z digest=sha256:32750fbbdd1ad6cab263dec4babcd6eab6499553e4f2309cc20ed0d89363fe04

Observation c8bc9018-c127-4fa4-aca2-b78c1748846f · outbound

This paper cites Towards Reliable Assessments of Demographic Disparities in Multi-Label Image Classifiers.

ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities Towards Reliable Assessments of Demographic Disparities in Multi-Label Image Classifiers

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T19:24:18.057411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T19:24:17.684622Z digest=sha256:708a4b7ad5abecdc49079d645696942cb6c5cfb8fe1aa9943a73bf40887c4a49

Pith citing papers

Observation 82bba862-11c9-4694-8553-0fa2e4bc86c0 · inbound

Efficient Portfolio Selection through Preference Aggregation with Quicksort and the Bradley--Terry Model cites this paper.

Efficient Portfolio Selection through Preference Aggregation with Quicksort and the Bradley--Terry Model ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-22T21:12:08.585633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-22T21:09:30.462709Z digest=sha256:a4042a4876457e4e2bff6440958fced7e098c480a88c58d9f2da46b2fd70c9ec