Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:36:10.926833Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 3 inbound Pith citation observations for arXiv:2506.01793.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:36:10.926833Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:59:53.097164Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T10:25:41.538596Z
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0925fa6d-86b5-4107-91d1-39b216dd577b · outbound
Human-Centric Evaluation for Foundation Models ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a0d0b08-ad7f-49d6-b64d-7f052af9b507 · outbound
Human-Centric Evaluation for Foundation Models Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10f8423d-b43e-4101-9f64-c10296dc7f76 · outbound
Human-Centric Evaluation for Foundation Models Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb0088f-27c9-4efb-89af-fd601b966365 · outbound
Human-Centric Evaluation for Foundation Models Unresolved cited work
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb08326d-a2b4-4860-bddf-5e5d0d273b09 · outbound
Human-Centric Evaluation for Foundation Models Training Verifiers to Solve Math Word Problems
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9ef97ce-3767-40b3-b057-acc3948f46e0 · outbound
Human-Centric Evaluation for Foundation Models Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42158aae-ee92-4301-92d9-8ff88a539e89 · outbound
Human-Centric Evaluation for Foundation Models Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ee8cdb2-dfa9-42af-9b33-144ce5b9461d · outbound
Human-Centric Evaluation for Foundation Models Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a17b21af-4cf2-4fc7-a023-0c07623c5ad2 · outbound
Human-Centric Evaluation for Foundation Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b0f499c-efa2-4a89-b27c-39659bdb72cf · outbound
Human-Centric Evaluation for Foundation Models Evaluating Large Language Models: A Comprehensive Survey
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d97275f3-aeb7-4259-beb4-4f5b339d4c49 · outbound
Human-Centric Evaluation for Foundation Models Measuring Massive Multitask Language Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b1ff2b2-0971-427a-8e64-d414fca976d8 · outbound
Human-Centric Evaluation for Foundation Models Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25ade072-4507-4075-b16b-4c97191acf3a · outbound
Human-Centric Evaluation for Foundation Models SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 151c695c-ddd2-49fe-9a92-79b3978cd901 · outbound
Human-Centric Evaluation for Foundation Models Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d031f4f1-005b-4baf-8be1-021ff3e3a4fa · outbound
Human-Centric Evaluation for Foundation Models Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 919e515e-177e-4bd9-aeb6-75efb7ded9e5 · outbound
Human-Centric Evaluation for Foundation Models Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 260f860c-a932-449e-8e9a-da62a1004ff0 · outbound
Human-Centric Evaluation for Foundation Models Humanity's Last Exam
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1bfe8b1-1829-4482-91fa-ee62fedbcd31 · outbound
Human-Centric Evaluation for Foundation Models Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c8d298b-56b9-4d13-bb50-b23deab3d747 · outbound
Human-Centric Evaluation for Foundation Models Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12fd8df5-1cc1-44d1-8fb9-895e0f64c910 · outbound
Human-Centric Evaluation for Foundation Models Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34896e7b-0bd2-4806-bb20-04bc90db23c2 · outbound
Human-Centric Evaluation for Foundation Models Unresolved cited work
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e1633281-b1ea-4bbf-b403-de3c5dbe2028 · outbound
Human-Centric Evaluation for Foundation Models LLaMA: Open and Efficient Foundation Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6efeae68-bae3-4bbb-98a0-2fe772516df6 · outbound
Human-Centric Evaluation for Foundation Models Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd24c201-aef2-4daa-9cec-1c0213776a6f · outbound
Human-Centric Evaluation for Foundation Models FinQA: A Dataset of Numerical Reasoning over Financial Data
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcb0f193-1d3f-4d51-bf99-f699c382f548 · outbound
Human-Centric Evaluation for Foundation Models Holistic Evaluation of Language Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7d07220-01ed-4edf-855e-d4ed56d2bb6a · outbound
Human-Centric Evaluation for Foundation Models Nature Machine Intelligence 5, 1 (2023), 46–57
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9e614170-5bc0-409a-a2c2-1da423bb0e56 · inbound
Affordance Benchmark for MLLMs Human-Centric Evaluation for Foundation Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a1415e8-4082-45cf-a4ea-ad53bda88d4d · inbound
First, Do No Harm (With LLMs): Mitigating Racial Bias via Agentic Workflows Human-Centric Evaluation for Foundation Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 548330fb-bdcc-40f8-b0eb-a81748eef3eb · inbound
CLExEval: A Human-in-the-Loop Framework for Qualitative Evaluation of LLM Clinical Reasoning Human-Centric Evaluation for Foundation Models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.