Pith. sign in

Paper Citation Record · LEDGER

Position: AI Evaluation Should Learn from How We Test Humans

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2306.10512.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.10512 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:36:09.647330Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T22:00:41.508241Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 68778ea7-cdff-46c2-b429-b893d4500153 · inbound

Psychometric-Based Evaluation for Theorem Proving with Large Language Models cites this paper.

Psychometric-Based Evaluation for Theorem Proving with Large Language Models Position: AI Evaluation Should Learn from How We Test Humans

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T17:36:09.647330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:36:09.647330Z digest=sha256:124d5bc05126f60c78c01b5db3820ccef630ff22461dec61340ba199df4cecda

Observation 05ed1da9-f499-4840-9920-0ff28eff1f3c · inbound

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs cites this paper.

AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Position: AI Evaluation Should Learn from How We Test Humans

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T13:40:25.866299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:40:25.866299Z digest=sha256:212472dc7468a2b54f7dbbcec65177ab24e5186a17f61cd4aba9037628347e0e

Observation 3fc827b1-b721-4472-8bcc-617ffe7d9033 · inbound

Fluid Language Model Benchmarking cites this paper.

Fluid Language Model Benchmarking Position: AI Evaluation Should Learn from How We Test Humans

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T17:10:43.454204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:10:43.454204Z digest=sha256:2eecaaa153fe447f6752f27a5c26ef88b8b6a5bd896e1b16ed1e69b8dbe94c1c

Observation 9ce79528-af9e-4bc9-a45c-9b7b9f3488d5 · inbound

Position: AI Evaluations Should be Grounded on a Theory of Capability cites this paper.

Position: AI Evaluations Should be Grounded on a Theory of Capability Position: AI Evaluation Should Learn from How We Test Humans

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:00:41.510890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-21T21:57:55.834632Z digest=sha256:6a1b90d017405a027f73f767cf8437d074c3d2edd6a1cb7038f8060e7dc36063

Observation 3cbdb515-ecc6-42aa-84ee-abe526a966b9 · inbound

Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks cites this paper.

Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks Position: AI Evaluation Should Learn from How We Test Humans

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-04T08:10:10.570262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:10:10.570262Z digest=sha256:32a469146465c66e93dd53f9ed3a449c04a3ceee76eb31f15b3be80cfb16f542

Observation b116e399-813e-4881-a96a-e464b8e0b6a1 · inbound

FairTree: Subgroup Fairness Auditing of Machine Learning Models with Bias-Variance Decomposition cites this paper.

FairTree: Subgroup Fairness Auditing of Machine Learning Models with Bias-Variance Decomposition Position: AI Evaluation Should Learn from How We Test Humans

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:31:03.899427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T03:32:02.079410Z digest=sha256:d312a9fcc59d02b25bcb5b514f5c23dc1c465c5129968b63d69e8ac2db213570

Observation 5504ef60-7b19-44d4-9713-039aaca1b0b4 · inbound

Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation cites this paper.

Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation Position: AI Evaluation Should Learn from How We Test Humans

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:06:27.320015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T08:02:49.168872Z digest=sha256:b2d24298a5179580edb7e5c03472a8be4622274c81c6178d76b51f4faa074fbd