Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2306.10512.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T17:36:09.647330Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T22:00:41.508241Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 68778ea7-cdff-46c2-b429-b893d4500153 · inbound
Psychometric-Based Evaluation for Theorem Proving with Large Language Models Position: AI Evaluation Should Learn from How We Test Humans
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05ed1da9-f499-4840-9920-0ff28eff1f3c · inbound
AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs Position: AI Evaluation Should Learn from How We Test Humans
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fc827b1-b721-4472-8bcc-617ffe7d9033 · inbound
Fluid Language Model Benchmarking Position: AI Evaluation Should Learn from How We Test Humans
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ce79528-af9e-4bc9-a45c-9b7b9f3488d5 · inbound
Position: AI Evaluations Should be Grounded on a Theory of Capability Position: AI Evaluation Should Learn from How We Test Humans
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3cbdb515-ecc6-42aa-84ee-abe526a966b9 · inbound
Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks Position: AI Evaluation Should Learn from How We Test Humans
Reference 2015
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b116e399-813e-4881-a96a-e464b8e0b6a1 · inbound
FairTree: Subgroup Fairness Auditing of Machine Learning Models with Bias-Variance Decomposition Position: AI Evaluation Should Learn from How We Test Humans
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5504ef60-7b19-44d4-9713-039aaca1b0b4 · inbound
Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation Position: AI Evaluation Should Learn from How We Test Humans
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.