Pith. sign in

Paper Citation Record · LEDGER

Human-Centric Evaluation for Foundation Models

As of 7 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 3 inbound Pith citation observations for arXiv:2506.01793.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01793 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:36:10.926833Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:59:53.097164Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T10:25:41.538596Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0925fa6d-86b5-4107-91d1-39b216dd577b · outbound

This paper cites ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate.

Human-Centric Evaluation for Foundation Models ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.031077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.031077Z digest=sha256:2367590f0a4faca5e728e2fc81843af45d2e993ccdc9edc9d5b4d2476ff2a6c5

Observation 6a0d0b08-ad7f-49d6-b64d-7f052af9b507 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.102859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.102859Z digest=sha256:85b8168ac1c8fa26fde5864f4abae99d746692ce0eb30679d7d18ed96c8d2e19

Observation 10f8423d-b43e-4101-9f64-c10296dc7f76 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.224321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.224321Z digest=sha256:55c397ee37ee0e1b41600f7a3bdd28944c92786ddaf92857ab6a7d2420bc8a46

Observation 7cb0088f-27c9-4efb-89af-fd601b966365 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.474041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.474041Z digest=sha256:f25ddb520834abda9f97909e79b71c6bd490db0a6832420d4f50b316a67dac41

Observation cb08326d-a2b4-4860-bddf-5e5d0d273b09 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Human-Centric Evaluation for Foundation Models Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.593518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.593518Z digest=sha256:af8f2c88133c82b55c37ffd91666556bb111bbd0484a3a954acd47b5b78db347

Observation c9ef97ce-3767-40b3-b057-acc3948f46e0 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.706899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.706899Z digest=sha256:bfcc271804d3aaeb1301ab367e0e6863e3f52e70f233b09273b4f8c473e8094b

Observation 42158aae-ee92-4301-92d9-8ff88a539e89 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Human-Centric Evaluation for Foundation Models Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.792029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.792029Z digest=sha256:152e3ce94b4f4e708d74edadd45199b6a6183a5c1c9bb8faad33e0c09201e33d

Observation 8ee8cdb2-dfa9-42af-9b33-144ce5b9461d · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:36:12.166425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:36:08.874604Z digest=sha256:7eb7f1554d8e3b594e15c6ebc91bba52d4f826591e28727e9b819f8d1a4a0b2b

Observation a17b21af-4cf2-4fc7-a023-0c07623c5ad2 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Human-Centric Evaluation for Foundation Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.961655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.961655Z digest=sha256:4ae695fc6474560fcbf6f633bb675b38c2fa960e5e593f9390385ea5205fe35e

Observation 8b0f499c-efa2-4a89-b27c-39659bdb72cf · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

Human-Centric Evaluation for Foundation Models Evaluating Large Language Models: A Comprehensive Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:09.124209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:09.124209Z digest=sha256:95fabcbfa5868c32724f754f57c2ec800eb866993cfd8110f63832a9d601404a

Observation d97275f3-aeb7-4259-beb4-4f5b339d4c49 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Human-Centric Evaluation for Foundation Models Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:09.278576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:09.278576Z digest=sha256:9cc28d4bb6dd00166f72792c7bf6855372f39cc24d82b75e0ff347ff7ed2130a

Observation 5b1ff2b2-0971-427a-8e64-d414fca976d8 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:36:12.015503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:36:09.435590Z digest=sha256:7a922ba6b31d830314affd77e816e1ee3ac03cddf071183e6c1770b061765052

Observation 25ade072-4507-4075-b16b-4c97191acf3a · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

Human-Centric Evaluation for Foundation Models SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:09.544344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:09.544344Z digest=sha256:fa2c73775ea58b6a466e87b9697449eb1d5dc2ca7571f52786d832adc22f3fab

Observation 151c695c-ddd2-49fe-9a92-79b3978cd901 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:09.708758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:09.708758Z digest=sha256:60108e60a5a75ab0c80346219f88c38c94af5247b5e6aff68a522ab8fd7c11ce

Observation d031f4f1-005b-4baf-8be1-021ff3e3a4fa · outbound

This paper cites Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena.

Human-Centric Evaluation for Foundation Models Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:10.030772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:10.030772Z digest=sha256:57f89ba53132b79b7acf4de4564ac17e447eb1d75d4f395d058d517f01123760

Observation 919e515e-177e-4bd9-aeb6-75efb7ded9e5 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:36:11.856502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:36:10.154302Z digest=sha256:e02633a226522967c47b1be96ed62d5c4464b03081d985ba3cd0faa082cc922c

Observation 260f860c-a932-449e-8e9a-da62a1004ff0 · outbound

This paper cites Humanity's Last Exam.

Human-Centric Evaluation for Foundation Models Humanity's Last Exam

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:10.296684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:10.296684Z digest=sha256:24ec3ca11ba628c1672d7d56e18bd29820be29de079e37a5fb290932298d2c24

Observation b1bfe8b1-1829-4482-91fa-ee62fedbcd31 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:10.369987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:10.369987Z digest=sha256:5a6667652a7393de5d0842dd34b9723cc643b45500f62ff3cbc4662725b06ec9

Observation 0c8d298b-56b9-4d13-bb50-b23deab3d747 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:10.453562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:10.453562Z digest=sha256:6fc6759c01481d78f966935e51df9cc386a8241b6078ce9f4f8a5c9741da24da

Observation 12fd8df5-1cc1-44d1-8fb9-895e0f64c910 · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Human-Centric Evaluation for Foundation Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:10.624298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:10.624298Z digest=sha256:e49d24782b9b5fca4d9cb2bbc1c136378012d00b22c600375cde3a5e31bfdc8e

Observation 34896e7b-0bd2-4806-bb20-04bc90db23c2 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:36:11.406036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:36:10.706120Z digest=sha256:ba9d13e51208059f14672ae598256dacd9b084bac13c2dfdd9569095329c8868

Observation e1633281-b1ea-4bbf-b403-de3c5dbe2028 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Human-Centric Evaluation for Foundation Models LLaMA: Open and Efficient Foundation Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:10.812719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:10.812719Z digest=sha256:276678d9a85d0c4e6a0d053a328f69320ad9eef3bf32e65a1078577de6f0b53d

Observation 6efeae68-bae3-4bbb-98a0-2fe772516df6 · outbound

This paper cites an unresolved cited work.

Human-Centric Evaluation for Foundation Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:36:11.186918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:36:10.926833Z digest=sha256:c9dc89cdd6b8a6af2d0397286e90b7a2305f1f781c0e5ac42545a0904cf9abce

Observation dd24c201-aef2-4daa-9cec-1c0213776a6f · outbound

This paper cites FinQA: A Dataset of Numerical Reasoning over Financial Data.

Human-Centric Evaluation for Foundation Models FinQA: A Dataset of Numerical Reasoning over Financial Data

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:08.350199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:08.350199Z digest=sha256:48bef2b25818055f3ba73a3e8e35745a3ad7d9d37f59a7eef7f9c505a1dc1fa7

Observation fcb0f193-1d3f-4d51-bf99-f699c382f548 · outbound

This paper cites Holistic Evaluation of Language Models.

Human-Centric Evaluation for Foundation Models Holistic Evaluation of Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:09.881359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:09.881359Z digest=sha256:e0e918df00368aadbdbf5fca33dc7625ef8964624370ea1a694e08bb0c4993ed

Observation b7d07220-01ed-4edf-855e-d4ed56d2bb6a · outbound

This paper cites Nature Machine Intelligence 5, 1 (2023), 46–57.

Human-Centric Evaluation for Foundation Models Nature Machine Intelligence 5, 1 (2023), 46–57

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:36:11.617747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:36:10.530656Z digest=sha256:8b88b9bc693c7b48faadee6f3b32c41a0e2bfa4693ee90750af6d88081d5410c

Pith citing papers

Observation 9e614170-5bc0-409a-a2c2-1da423bb0e56 · inbound

Affordance Benchmark for MLLMs cites this paper.

Affordance Benchmark for MLLMs Human-Centric Evaluation for Foundation Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:53.097164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:59:53.097164Z digest=sha256:27949d5b1cc0f8c4c98cb4f4f35111ada5b1e476d333d6d555355af626c70356

Observation 6a1415e8-4082-45cf-a4ea-ad53bda88d4d · inbound

First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows cites this paper.

First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows Human-Centric Evaluation for Foundation Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:21:06.094847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T03:47:54.377400Z digest=sha256:1828ab098a04277fbe15fd10765944e19d755c404c6de1133178b79a13fca05a

Observation 548330fb-bdcc-40f8-b0eb-a81748eef3eb · inbound

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning cites this paper.

CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning Human-Centric Evaluation for Foundation Models

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T10:25:41.540358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-07-01T05:33:52.027771Z digest=sha256:ae7c4820e0ecda1ccbda50af1845645c6dbe722814df1da79b73a64f5b199da1